SQuad: Sub-Quadratic Attention Distillation for Efficient Video Generation
Abstract
SQuad distills pretrained video diffusion transformers into sub-quadratic attention via two-stage flow-matching and distribution-matching distillation, cutting attention cost and latency while preserving generation quality.
Video Diffusion Transformers (DiTs) spend most of their compute inside the Self-Attention operation, whose cost grows quadratically, O(n^2), with the number of latent tokens n. For the task of video generation, the token count is large, so this term dominates runtime and memory, and thereby caps the resolution and duration we can generate. Linear O(n) and low-rank O(nk) surrogates of Self-Attention trade the full softmax QK^T for cheaper kernels, but rarely recover the original's expressivity, leaving a stubborn quality gap. Motivated by this, we propose SQuad, a Sub-Quadratic Attention Distillation framework that achieves a complexity of O(nn) in the resulting distilled Attention, naturally balancing the efficiency v/s expressivity trade-off. Instead of training our own Video DiT from scratch, which is prohibitively expensive, we fit a pretrained full softmax Self-Attention DiT into our proposed SQuad-Attention one by distilling the former in two stages: Flow-Matching Supervised Fine-Tuning (SFT), followed by improved Distribution Matching Distillation (DMD2) which additionally makes the sampling more efficient. On the Wan~2.2 5B text-to-video model, SQuAD matches the quadratic teacher on VBench (83.20 v/s 83.08) while cutting the per-step per-block attention FLOPs by sim67times and attention latency by sim11times, and end-to-end DiT latency by 2times, all while also generating a video in only 6 Neural Functional Evaluations (NFEs) instead of the default 100.
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper