SANA-WM_streaming#

SANA-WM_streaming is the chunk-causal, camera-controlled NVlabs/Sana world model release. It produces video progressively across autoregressive chunks with a chunk-causal Stage-1 DiT, streaming LTX-2 refiner, and streaming VAE decode path. FlashDreams exposes it through the cam2v-sana-wm-streaming application.

The sibling full-sequence release has a separate model card: SANA-WM_bidirectional.

SANA-WM streaming FlashDreams sample clip.

Requirements#

  • PyTorch: >= 2.9.

  • Precision: BF16 by default. FP8 Stage-1/refiner inference is available on Hopper or newer GPUs (sm_90+), and FP4 is available on Blackwell (sm_100+). These upstream precision flags belong to SANA-WM_streaming.

Installation#

# from the repo root
uv sync --package flashdreams-sana-wm --extra dev

Interactive Cam2V application#

Launch the V2 application to drive SANA-WM with live keyboard controls. The model adapter passes controls through the SANA-WM action remapper and appends each generated block to the camera conditioning history.

uv run --no-sync flashdreams-run-v2 cam2v-sana-wm-streaming \
    --mode webrtc --host 0.0.0.0 --port 8089 -- \
    --example-data

The application uses the checkpoint fixed resolution of 1280x704 and ten 24-frame blocks by default. Use --total-blocks after -- to change the rollout length.

Use --example-data to download the official demo_0.png and paired prompt to the FlashDreams example-data cache. Explicit image and prompt arguments override those example inputs.

Profiling benchmark#

The charts below compare steady-state generation latency per produced chunk for FlashDreams SANA-WM_streaming and the official SANA-WM_streaming implementation under matched settings. Warmup runs and the first decoded chunk are excluded from the headline metric. These GB300 latency runs show the official implementation faster than FlashDreams for BF16, FP8, and FP4.

In these charts, Official Impl means the pinned NVlabs/Sana upstream implementation measured by the FlashDreams benchmark harness under matched settings. It is not the SANA-WM 80-scene benchmark result published by the model authors.

BF16 steady-state milliseconds per produced chunk on one NVIDIA GB300 GPU: official 1,170.29 ms, FlashDreams 1,957.93 ms.

FP8 steady-state milliseconds per produced chunk on one NVIDIA GB300 GPU: official 1,482.92 ms, FlashDreams 2,392.67 ms.

FP4 steady-state milliseconds per produced chunk on one NVIDIA GB300 GPU: official 1,594.33 ms, FlashDreams 4,118.36 ms.

All charts use the same demo image/prompt, w-80,dw-40,w-80,aw-40 action path, 241 requested frames, one discarded warmup run, and three measured runs. The benchmark runs recorded FlashDreams commit bd0816e and upstream commit 6298508.

Citation#

If you use SANA-WM, please cite the original SANA work:

@misc{xie2024sana,
      title={SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers},
      author={Enze Xie and Junsong Chen and Junyu Chen and Han Cai and Haotian Tang and Yujun Lin and Zhekai Zhang and Muyang Li and Ligeng Zhu and Yao Lu and Song Han},
      year={2024},
      eprint={2410.10629},
      archivePrefix={arXiv},
      primaryClass={cs.CV}
}