Abstract
Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits requires training samples that capture both visits, often spanning minutes. Yet dense attention over the full span incurs quadratic costs, making long-span supervision expensive. Memorizon breaks this coupling: long spans are needed for supervision, but not for attention, since the two visits can share a forward pass without including every intervening frame. A training sample covers a span of any length but is scored only on its last k chunks. Instead of tokenizing the history before them, each scored chunk retrieves its own top-K latents by camera co-visibility, and the union of these requests forms a shared bank. The bank is bounded by kK, so the sequence stays bounded however long the span; at the shortest span the recipe is exactly conventional training. Adding the bank raises the cost of a step once; beyond that, a longer span costs little, and going from 100 to 400 s adds 12% to the step time. Against a sliding-window baseline, retrieval raises revisit consistency on every split, and a span long enough to reach the first visit of each return adds a further 24% to 30%, at some cost in image quality; beyond that span, more length no longer helps. Filling the bank from another episode lowers revisit correlation by 83%, so the model uses what it retrieves. Project page: https://tingtingliao.github.io/memorizon
Community
Memorizon trains a camera-controlled video world model on spans far longer than its context window. Each training sample packs up to 100-400s of history into a fixed-size sequence: every chunk retrieves the past frames whose camera frustums overlap its own. Training therefore costs the same as a 10 s window, yet the model learns to remember scenes it saw minutes earlier. Built on Wan2.2-TI2V-5B and distilled to 4 steps with Self-Forcing, it keeps returning views consistent over minute-long interactive rollouts.
๐ Project page: https://tingtingliao.github.io/memorizon/
๐ป Code: https://github.com/TingtingLiao/memorizon
๐ค Weights: https://huggingface.co/Luffuly/memorizon
Get this paper in your agent:
hf papers read 2610.00544 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper