U-Space: Uncovering When and Why Uncertainty Arises in Language Models Paper • 2610.09087 • Published 4 days ago • 30
MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers Paper • 2610.06801 • Published 5 days ago • 29
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement Paper • 2610.11959 • Published 2 days ago • 58
Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position Paper • 2610.10114 • Published 3 days ago • 31
nanoMuse: An Open-Source Personal Agent for Every Device You Own Paper • 2610.08699 • Published 4 days ago • 98
In-Distribution Forcing for Long Video Generation at Test Time Paper • 2610.03120 • Published 8 days ago • 49
How to Loop MoE: Flatten the Experts, Untie the Attention Paper • 2609.35751 • Published 12 days ago • 13