Spatial IntelligenceVersion 1.5

InSpatio-World 1.5

Predict the next novel view from a single image, image set, panorama, or video, and explore wider, more consistent scenes across space and time.

Overview

Beyond the frame. Into the world.

InSpatio-World 1.5 accepts a single image, an image set, a panorama, or a video, opening more ways to explore scenes across space and time.

Through next-view prediction, the model infers unseen space from existing observations. Wider viewpoint changes support scene roaming and bullet-time creation, while consistency helps preserve subjects, structure, and details as the camera moves.

Version 1.0 built explorable dynamic 4D worlds from reference video, using State-Anchored World Modeling to maintain spatiotemporal consistency. Version 1.5 expands on that foundation with more input types and a wider exploration range.

Specifications

Model TypeSpatial intelligence world model
OutputExplorable novel views and spatiotemporal scenes
ExplorationReal-time scene roaming
InputsSingle image, image set, panorama, video
Key AdvancesWider viewpoint changes and stronger consistency
Version1.5

Key Capabilities

Wide-Range Scene Roaming

Enter a scene from a single image and move beyond the original frame while keeping its appearance and spatial structure coherent.

Exploration Across Space and Time

Use video input to choose where and when to observe a dynamic scene as the action unfolds.

Flexible Inputs and Bullet Time

Start from an image, image set, panorama, or video; use synchronized images to create a coherent camera path around a subject.

Consistency Across Views

Preserve subject identity, spatial structure, and scene details through large camera moves and changes over time.

Method

From Next-View Prediction to Spatial Intelligence

An image is only one observation of a world. InSpatio-World 1.5 infers scene structure from existing views, predicts how the same world appears from a new position and angle, and maintains consistency as viewpoints change. Version 1.0 established video-conditioned dynamic exploration through State-Anchored World Modeling; 1.5 extends that foundation to more inputs and a wider range of spatial and temporal exploration.

Version 1.0 Benchmark

The Real-Time Foundation of 1.0

The 1.3B-parameter model in version 1.0 ranked first among real-time methods on WorldScore-Dynamic and ran at 24 FPS on a single GPU. These are results for 1.0, not performance claims for 1.5.

InSpatio-World 1.0 WorldScore-Dynamic benchmark chart
* Version 1.0 leaderboard rankings and scores are current as of July 9, 2026.

From Content Creation to Spatial Intelligence

Version 1.5 Applications

Controlled Cinematography
Immersive Experiences
Spatial Intelligence
Embodied AI
Interactive Media

Get Started

Explore InSpatio-World 1.5

$ git clone https://huggingface.co/inspatio/world-1.5

Watch the 1.5 demos on the project page and visit the model page. The previous version’s open-source code remains available: InSpatio-World 1.0 on GitHub

Frequently Asked Questions

What inputs does InSpatio-World 1.5 support?

It supports single images, image sets, panoramas, and videos for scene roaming, bullet-time creation, and exploration across space and time.

What changed from 1.0 to 1.5?

Version 1.0 built explorable dynamic worlds from reference video. Version 1.5 adds image and panorama inputs, wider viewpoint changes, and stronger consistency across views.

What is next-view prediction?

The model infers scene structure from existing observations and predicts how the same world appears from a new position and angle while preserving subjects and spatial relationships.

What can 1.5 be used for?

Applications include controlled cinematography, immersive experiences, spatial intelligence research, and dynamic environment modeling for embodied AI.