UnStep

Training-Free Acceleration of Causal Video Diffusion
with Fewer Steps Than Distillation

  • Youssef Mansour
  • Enis Simsar
  • Fadime Sener
  • Markos Georgopoulos
  • Albert Pumarola
  • Ali Thabet
  • Edgar Schoenfeld

Meta Superintelligence Labs

Single H100

50 FPSTwo steps
60 FPSOne step

Single GB200

64 FPSTwo steps
77 FPSOne step
Training-free Inference-only acceleration

The same model. A faster way to run it.

Fewer steps.
Not a new checkpoint.

Few-step video models are fast, but not fast enough. UnStep is an inference-only wrapper that accelerates existing causal video models with fewer denoising steps, a bounded attention cache, and a faster transformer and decoder runtime.

Two training-free mechanisms recover lost quality: refining the existing clean-cache pass and applying truncated SVD to attention value and output projections. The result is faster generation from the same pretrained checkpoint.

A new speed regime for causal video diffusion.

Five-second video generation.
832 × 480. Batch size one.

Figure 1: UnStep's speed-quality operating points compared with causal video generation methods on H100 and GB200.
Figure 1. End-to-end throughput and VBench quality on H100 and GB200, as reported in the paper. Full-size figure

Wrap the checkpoint.
Keep the quality.

One wrapper, applied to Self Forcing, Causal Forcing, and LongLive 1.0 without retraining.

H100 results from the paper. Throughput is generation FPS, not video playback rate. Headline speeds are rounded.

Default two-step UnStep results on the full 946-prompt VBench T2V suite
Checkpoint FPS VBench total
Self Forcing 17.0 84.31
+ UnStep 49.8 84.56
Causal Forcing 17.0 84.82
+ UnStep 49.8 84.93
LongLive 1.0 17.0 83.31
+ UnStep 49.8 83.42

Accelerate. Reduce. Recover.

All at inference.
No finetuning. No redistillation.

UnStep method: optimize the DiT and VAE runtime, reduce denoising steps and the KV window, then recover quality with clean-cache refinement and V/O truncated SVD.
01

Optimize the runtime

Efficient attention calls and KV indexing, fused Triton RoPE, and VAE precision, memory-layout, and convolution optimizations.

02

Do less computation

Keep four denoising steps for the first chunk. Use two for later chunks, with a bounded temporal KV window.

03

Recover the quality

Refine and reuse the existing clean-cache pass, then apply truncated SVD to the attention value and output projections.

Same input. Different runtime.

Selected text-to-video examples.

Self Forcing17 FPS
Self Forcing + UnStep50 FPS

Image-to-video

a group of hot air balloons flying over a valley

Reference image
Hot air balloons over a valley
Self Forcing 17 FPS
Self Forcing + UnStep 50 FPS

Citation

Cite UnStep

@misc{mansour2026unsteptrainingfreeaccelerationcausal,
  title = {UnStep: Training-Free Acceleration of Causal Video Diffusion with Fewer Steps Than Distillation},
  author = {Youssef Mansour and Enis Simsar and Fadime Sener and Markos Georgopoulos and Albert Pumarola and Ali Thabet and Edgar Schoenfeld},
  year = {2026},
  eprint = {2609.32518},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.32518}
}