---
title: inspatio-world-v1.5
canonical_url: "https://www.modelscope.ai/models/InSpatio/inspatio-world-v1.5"
md_url: "https://www.modelscope.ai/models/InSpatio/inspatio-world-v1.5.md"
repository: InSpatio/inspatio-world-v1.5
last_updated: 2026-09-30
pipeline_tag: any-to-any
tasks:
  - any-to-any
parameters: 1.4B
tensor_type:
  - F32
library_name:
  - safetensors
downloads: 8
stars: 0
---

# inspatio-world-v1.5

> inspatio-world-v1.5 - An open-source model by InSpatio on ModelScope. 🤖 ModelScope &nbsp;&nbsp;| &nbsp;&nbsp;🤗 HuggingFace &nbsp;&nbsp;| &nbsp;&nbsp;💻 GitHub &nbsp;&nbsp;| &nbsp;&nbsp;🌐 Project Page &nbsp;&nbsp;| &nbsp;&nbsp;📑 Paper &nbsp;&nbsp;|…

- **Repository**: InSpatio/inspatio-world-v1.5
- **Tasks**: any-to-any
- **Parameters**: 1.4B
- **Downloads**: 8
- **Stars**: 0
- **Last updated**: 2026-09-30

Source: https://www.modelscope.ai/models/InSpatio/inspatio-world-v1.5

---

<h1 align="center">InSpatio-World-1.5</h1>

<p align="center">
    🤖 <a href="https://modelscope.cn/models/InSpatio/inspatio-world-v1.5">ModelScope</a>&nbsp;&nbsp;|
    &nbsp;&nbsp;🤗 <a href="https://huggingface.co/inspatio/world-1.5">HuggingFace</a>&nbsp;&nbsp;|
    &nbsp;&nbsp;💻 <a href="https://github.com/inspatio/inspatio-world-v1.5">GitHub</a>&nbsp;&nbsp;|
    &nbsp;&nbsp;🌐 <a href="https://inspatio.github.io/inspatio-world-1.5/">Project Page</a>&nbsp;&nbsp;|
    &nbsp;&nbsp;📑 <a href="https://arxiv.org/abs/2604.07209">Paper</a>&nbsp;&nbsp;|
    &nbsp;&nbsp;🖥️ <a href="https://world.inspatio.com/">Live Demo</a>
</p>

## Introduction

**InSpatio-World** is a real-time 4D world simulator via spatiotemporal autoregressive modeling. This repository provides the **InSpatio-World-1.5 1.3B checkpoint**.

- **Multiple input modes** — Run inference from a single image, four images, or a video.
- **Target camera trajectories** — Supply a target camera trajectory to generate video from the requested viewpoints.
- **Depth and source camera estimation** — For video inputs, Depth Anything 3 estimates missing depth and per-frame source camera intrinsics and extrinsics.
- **Real-time examples** — See the single-image and multi-image inference results in the showcase below.

For more details, see the [GitHub repo](https://github.com/inspatio/inspatio-world-v1.5), [project page](https://inspatio.github.io/inspatio-world-v1.5/), and [paper](https://arxiv.org/abs/2604.07209).

## Quick Start

### Installation

Requirements: **Python 3.10** and **CUDA 12.6**.

```bash
git clone https://github.com/inspatio/inspatio-world-v1.5.git
cd inspatio-world-v1.5

conda env create -f environment.yml
conda activate inspatio_world_test

python -m pip install --no-deps depth-anything-3==0.1.1
python -m pip install modelscope
```

The official `environment.yml` includes the packages DA3 uses on this inference path. Installing DA3 separately with `--no-deps` avoids installing its unrelated app and benchmark dependencies. On Hopper GPUs, the official repository recommends FlashAttention-3 (FA3) for faster attention.

### Run the Demo

```bash
bash run_example.sh
```

This runs the six scenes listed in `examples/manifest.json`. Results are saved to:

```text
output/<output_id>/
├── source.mp4
├── render.mp4
├── mask.mp4
└── pred.mp4
```

Image and multi-image scenes reuse uint16 depth PNGs. For video inputs, DA3 estimates missing depth and per-frame source camera intrinsics and extrinsics. The target trajectory is supplied by the user. The default depth mode is `auto`; use `--depth_mode existing` to require existing depth, or `--depth_mode estimate` to regenerate it.

### Single-Image and Multi-Image Inputs

Prepare a scene using the official [image and multi-image examples](https://github.com/inspatio/inspatio-world-v1.5/blob/main/examples/README.md), then run:

```bash
bash run_inference.sh --scene_dir /path/to/my_scene
```

Inputs must be **832×480**. A scene includes RGB input, depth maps, a prompt, source camera intrinsics and poses, a target camera trajectory, and scene metadata. See the linked examples for the file layout and camera matrix conventions.

### Video Inputs

Provide the RGB video, prompt text, and one target OpenCV world-to-camera **4×4 matrix per frame** (16 row-major numbers per line):

```bash
bash run_inference.sh --video /path/to/input.mp4 \
  --prompt "Your prompt" --target_traj /path/to/target_tcw.txt
```

The runner reads the video's frame rate and frame count, resizes it to **832×480** if needed, and estimates depth and per-frame source cameras with DA3. The target trajectory must use the same first-frame-normalized coordinate system and displacement scale as the estimated source cameras.

## Showcase

<p align="center">
    <img src="https://modelscope.cn/api/v1/models/InSpatio/inspatio-world-v1.5/repo?Revision=master&amp;FilePath=assets%2Freadme%2Fimage_prediction.gif" width="48%" alt="Real-time inference results with a single input image"/>
    <img src="https://modelscope.cn/api/v1/models/InSpatio/inspatio-world-v1.5/repo?Revision=master&amp;FilePath=assets%2Freadme%2Fmulti_image_prediction.gif" width="48%" alt="Real-time inference results with four input images"/>
</p>
<p align="center"><em>Real-time inference results: single-image input (left) and four-image input (right).</em></p>

Try the [live demo](https://world.inspatio.com/) for an interactive experience.

## License

The official repository's code is licensed under the [Apache-2.0 License](https://github.com/inspatio/inspatio-world-v1.5/blob/main/LICENSE). The upstream README states that this license applies only to code in its library; dependencies such as [Depth Anything 3](https://github.com/ByteDance-Seed/depth-anything-3) are separately licensed.

## Citation

If you use InSpatio-World in your research, please use the BibTeX entry provided by the official repository:

```bibtex
@misc{inspatio-world,
    title={INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling},
    author={InSpatio Team},
    journal={arXiv preprint arXiv: 2604.07209},
    year={2026}
}
```

## Acknowledgement

InSpatio-World uses a backbone based on [Wan2.1](https://github.com/Wan-Video/Wan2.1), with its training code referencing [Self-Forcing](https://github.com/guandeh17/Self-Forcing). We thank the Self-Forcing and Wan teams for their work and open-source contributions, and acknowledge [Depth Anything 3](https://github.com/ByteDance-Seed/depth-anything-3) and [ReCamMaster](https://github.com/KlingAIResearch/ReCamMaster) for their contributions.
