Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Paper • 2608.02711 • Published 24 days ago • 92
Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations Paper • 2608.01628 • Published 25 days ago • 23
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 28 days ago • 40
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models Paper • 2607.24957 • Published Jul 27 • 19
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 149
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 88
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published Jul 8 • 64
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space Paper • 2607.05373 • Published Jul 6 • 66
MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation Paper • 2606.26087 • Published Jun 24 • 35
World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible Paper • 2606.13652 • Published Jun 11 • 16