FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow
Abstract
FlashRender accelerates generative video rendering via representation alignment, a mean-flow objective, and on-policy distillation to achieve high-quality few-step camera-controlled synthesis.
We present FlashRender, a few-step generative rendering framework that retakes a source video along a target camera trajectory in seconds. We identify sampling-step-dependent camera control as a prominent manifestation of discretization error in existing multi-step generative rendering models and show that resolving this inconsistency substantially lowers denoising trajectory curvature, facilitating subsequent step distillation. To this end, we introduce Representation Transformation and Alignment (RETA), which aligns hidden source-video representations with target-video features from a frozen visual geometry model. This directly encodes the geometric transformation within the source-video stream, enabling sampling-step-consistent camera control. We then fine-tune the model with the MeanFlow objective on the lower-curvature denoising trajectory induced by RETA, allowing the model to more effectively address discretization error. Finally, we apply on-policy flow map distillation to correct self-rollout errors under fixed few-step sampling. Extensive experiments show that RETA, MeanFlow, and on-policy flow map distillation play complementary roles in few-step generative rendering. Together, they enable our approach to match multi-step baselines in video quality and geometric consistency at 25x lower sampling cost while achieving superior camera controllability, even under out-of-distribution target camera trajectories.
Community
FlashRender retakes an input video along a target camera trajectory in seconds, using only 4 NFE.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- 4DStreamCtrl: Interactive Video Generation with Online 4D Control (2026)
- Wonder: Video World Model Done Better (2026)
- PE-Field 4D: Video Generation Models as Canvas (2026)
- TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting (2026)
- Manifold4D: Denoising on Point Cloud Rendered Manifolds for Video Re-shooting (2026)
- SpatialCrafter: Single Image World Modeling with Generative 3D Proxies (2026)
- EditaLive! Unified Character Video Editing for Live Streaming (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.03563 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper