CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching
Abstract
Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-free caching can reduce this cost, yet existing policies make reuse decisions primarily from model-internal denoising dynamics and do not explicitly account for control transitions. Actually, interactive generation explicitly exposes a signal they do not use: the controls for a chunk arrive before it is denoised, so a schedule derived from them costs no forward pass. To this end, we analyze adjacent chunks under different control regimes and find that structural similarity drops around action changes, while low-frequency structure remains more persistent than high-frequency detail. Motivated by these observations, we propose CtrlCache, a training-free control-aware caching framework that adapts computation to the current control sequence. Specifically, the action-aware scheduling and refresh policy detects action changes across and within chunks, and labels each chunk as initial, transition, turning, or steady state. At one selected interior denoising step, initial and transition chunks retain full computation, while turning and steady chunks reuse the transformer residual from the most recent fully computed step in the same chunk. To exploit the persistence of low-frequency structure during steady interaction, we further introduce a frequency-mixed history prior guidance that incorporates complementary information from the preceding clean latent without an additional DiT forward pass. Evaluated on Matrix-Game 2.0 and LingBot-World v1/v2, CtrlCache achieves 1.21x to 1.41x DiT-backbone speedups without model retraining while improving WBench Overall scores over original inference across all three models.
Community
Hi everyone! We present CtrlCache, a training-free caching method for interactive video world models.
Key idea: in interactive generation, the user's controls for a chunk are known before the chunk is denoised. CtrlCache uses them to decide when cached transformer computation can be reused (steady motion / turning) and when it must be refreshed (action changes), and adds a frequency-mixed history prior during steady interaction.
On Matrix-Game 2.0 and LingBot-World v1/v2, it gives 1.21×–1.41× DiT-backbone speedups while improving WBench Overall over the original models.
Project page with side-by-side videos: https://wrecklong.github.io/CtrlCache/
Code: https://github.com/wrecklong/CtrlCache (coming soon)
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- TRAC: Trajectory-aware Reuse and Adaptive Correction for Efficient Autoregressive Video Generation (2026)
- DriveCache: Action-Aware Caching for Driving World Model Inference (2026)
- EpaCache: Error-Propagation-Aware Caching for Accelerating Diffusion-Based Visual Generation (2026)
- DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency (2026)
- LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation (2026)
- UnStep: Training-Free Acceleration of Causal Video Diffusion with Fewer Steps Than Distillation (2026)
- ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.08777 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper