Transferring the Intelligence of VLMs to Robotic Control Paper • 2609.22966 • Published 7 days ago • 117
GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots Paper • 2606.29705 • Published Jun 29 • 16
ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Paper • 2606.27313 • Published Jun 25 • 38
Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack Paper • 2606.14409 • Published Jun 12 • 16
GEM: Generative Supervision Helps Embodied Intelligence Paper • 2605.28548 • Published May 27 • 32
Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models Paper • 2603.18118 • Published Mar 18 • 12
Transferring the Intelligence of VLMs to Robotic Control Paper • 2609.22966 • Published 7 days ago • 117
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published Jul 16 • 120
ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Paper • 2606.27313 • Published Jun 25 • 38
GEM: Generative Supervision Helps Embodied Intelligence Paper • 2605.28548 • Published May 27 • 32
HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents Paper • 2604.07430 • Published Apr 8 • 130
HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents Paper • 2604.07430 • Published Apr 8 • 130
Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training Paper • 2603.12255 • Published Mar 12 • 91
Unleashing Text-to-Image Diffusion Models for Visual Perception Paper • 2303.02153 • Published Mar 3, 2023