Memorizon: Training World Models Beyond Their Context Window
Paper โข 2610.00544 โข Published โข 4
Tingting Liao ยท
Xuezhi Liang ยท
Hao Li ยท
Guangyi Liu
Institute of Foundation Models (IFM), MBZUAI
| Model | Memorizon, 4-step distilled (Self-Forcing + DMD, CFG distilled in) |
| Base | Wan2.2-TI2V-5B |
| Input | one image + keyboard actions or a camera trajectory |
| Output | 864ร480, 16 fps, generated one second at a time |
| Sampling | 4 steps, no CFG (read from the config) |
git clone https://github.com/TingtingLiao/memorizon.git && cd memorizon && pip install -e .
python scripts/generate.py --checkpoint Luffuly/memorizon --image photo.jpg \
--actions "w*16 l*24 w*16 l*24" --output out.mp4
from memorizon import MemorizonPipeline, actions_to_c2w, save_video
pipe = MemorizonPipeline.from_pretrained("Luffuly/memorizon")
video = pipe("photo.jpg", actions_to_c2w("w*16 l*24 w*16 l*24"))
save_video(video, "out.mp4")
| Key | Action | Key | Action |
|---|---|---|---|
w s |
forward / backward ยท 0.25 m | j l |
turn left / right ยท 7.5ยฐ |
a d |
left / right ยท 0.25 m | i k |
look up / down ยท 7.5ยฐ |
. |
stay | *N |
repeat N times |
Without a prompt, the image is captioned by Qwen3-VL-8B in the training format
(memorizon/caption.py). Custom prompts should follow it โ a perspective prefix,
then one paragraph:
First-person perspective โ character not visible. A winding paved path curves beneath a canopy of vibrant pink cherry blossoms, โฆ
Instead of actions, pass --trajectory a [T, 4, 4] array of camera-to-world poses,
T = 1 + 4 ร seconds, in the first frame's coordinates (OpenCV axes), translations
in metres / 4.
@article{memorizon2026,
title = {Memorizon: Training World Models Beyond Their Context Window},
author = {Tingting Liao, Xuezhi Liang, Hao Li, Guangyi Liu},
journal = {arXiv preprint arXiv:2610.00544},
year = {2026}
}
Base model
Wan-AI/Wan2.2-TI2V-5B-Diffusers