AutoGUIWorld: Image Generators as Visual World Models for GUI Agent Paper • 2610.01215 • Published 8 days ago • 61
Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning Paper • 2609.33781 • Published 12 days ago • 46
Raven: The Harness of Harnesses for Composable Agentic Intelligence Paper • 2609.33439 • Published 12 days ago • 569
APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants Paper • 2609.37559 • Published 10 days ago • 46
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning Paper • 2609.35505 • Published 11 days ago • 26
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders Paper • 2609.31620 • Published 14 days ago • 160
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation Paper • 2609.30221 • Published 15 days ago • 47
Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes Paper • 2609.25247 • Published 18 days ago • 10
Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model Paper • 2609.18323 • Published 23 days ago • 133
JEV-as-a-Judge: Accept When Confident, Escalate When Unsure Paper • 2609.26550 • Published 17 days ago • 42
Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts Paper • 2609.06011 • Published Sep 5 • 18
ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning Paper • 2609.22323 • Published 23 days ago • 10
Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies Paper • 2609.24682 • Published 18 days ago • 12
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents Paper • 2609.22000 • Published 21 days ago • 80
WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing Paper • 2609.20423 • Published 22 days ago • 48
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 23 days ago • 84
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 28 days ago • 266
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data Paper • 2609.05405 • Published Sep 4 • 45