SANA-WM_bidirectional#

SANA-WM_bidirectional is the full-sequence, camera-controlled NVlabs/Sana world model release. Given a first frame, a text prompt, and a camera trajectory, it renders a video clip in a single bidirectional pass. FlashDreams exposes it through the PIPELINE_SANA_WM_BIDIRECTIONAL pipeline configuration, with a native Stage-1 DiT and an LTX-2 refiner.

The sibling streaming release has a separate model card: SANA-WM_streaming.

SANA-WM bidirectional FlashDreams sample clip.

Requirements#

  • PyTorch: >= 2.9.

  • Precision: BF16 by default. The pipeline configuration also exposes opt-in FP8 and FP4 execution paths, but the upstream-vs-FlashDreams benchmark for SANA-WM_bidirectional is BF16-only because upstream SANA-WM_bidirectional does not support those precision flags.

Installation#

# from the repo root
uv sync --package flashdreams-sana-wm --extra dev

Programmatic pipeline access#

The bidirectional model is available as a pipeline configuration:

from sana_wm.config import PIPELINE_SANA_WM_BIDIRECTIONAL

pipeline = PIPELINE_SANA_WM_BIDIRECTIONAL.setup().to("cuda").eval()

Profiling benchmark#

The BF16 chart below compares steady-state in-process generation latency per generated clip for FlashDreams SANA-WM_bidirectional and the official SANA-WM_bidirectional implementation under matched settings on one NVIDIA GB300 GPU. FlashDreams measured 34,182.39 ms per clip versus 56,932.83 ms for the official implementation.

In this chart, Official Impl means the pinned NVlabs/Sana upstream implementation measured by the FlashDreams benchmark harness under matched settings. It is not the SANA-WM 80-scene benchmark result published by the model authors.

This chart shows steady-state in-process generation latency per generated clip in milliseconds for a 121-frame full-pipeline BF16 run (Stage-1 DiT + LTX-2 refiner + SANA VAE decode). The measured row used one NVIDIA GB300 GPU, one live warmup generation, and three measured generations. Model construction, checkpoint loading, video writing, and frame dumps are outside the timing boundary. The benchmark runs recorded FlashDreams commit bd0816e and upstream commit 6298508.

Citation#

If you use SANA-WM, please cite the original SANA work:

@misc{xie2024sana,
      title={SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers},
      author={Enze Xie and Junsong Chen and Junyu Chen and Han Cai and Haotian Tang and Yujun Lin and Zhekai Zhang and Muyang Li and Ligeng Zhu and Yao Lu and Song Han},
      year={2024},
      eprint={2410.10629},
      archivePrefix={arXiv},
      primaryClass={cs.CV}
}