TL;DR: SpatialCrafter builds generative 3D proxies from a single image to enable consistent, controllable video generation.
- [2026-08] Paper released on arXiv.
- Inference code && Pre-trained model weights
- Dataset (115K scenes)
- Training code
# TODO
# TODOThis project is released under the Apache 2.0 License.
The dataset is released for non-commercial research purposes only.
@article{spatialcrafter2026,
title = {SpatialCrafter: Single Image World Modeling with Generative 3D Proxies},
author = {Author names omitted for double-blind review},
journal = {arXiv preprint},
year = {2026}
}We thank the teams behind TRELLIS and WanVideo for open-sourcing their models, and the authors of SpatialGen for the website template.