PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation Paper • 2609.38597 • Published 7 days ago • 32
Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation Paper • 2610.01092 • Published 5 days ago • 27
DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence Paper • 2609.39222 • Published 6 days ago • 42
It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them Paper • 2609.37863 • Published 7 days ago • 38
AutoDataBench: A Data-centric Testbed for Accelerating Auto Research Paper • 2609.40097 • Published 6 days ago • 40
Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models Paper • 2609.33355 • Published 9 days ago • 50
4Director: Controlling Video World Models with Rigid 3D Geometry Paper • 2610.02160 • Published 5 days ago • 35
SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation Paper • 2610.02201 • Published 5 days ago • 36
FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation Paper • 2609.38839 • Published 6 days ago • 92
Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training Paper • 2609.40111 • Published 6 days ago • 50
LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models Paper • 2609.32264 • Published 10 days ago • 48
Beyond the Current Scene: Event-Referential Grasping with Active View Selection Paper • 2609.39375 • Published 6 days ago • 50
Scaling and Distilling Text Embeddings for Better Diffusibility Paper • 2610.01016 • Published 5 days ago • 51
Persona Dosing: Calibrated Activation Steering for Graded Trait Control Paper • 2609.36388 • Published 8 days ago • 51
Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding Paper • 2609.32019 • Published 11 days ago • 53
More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models Paper • 2609.38827 • Published 6 days ago • 60
AutoGUIWorld: Image Generators as Visual World Models for GUI Agent Paper • 2610.01215 • Published 5 days ago • 56
Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation Paper • 2609.38024 • Published 7 days ago • 58
InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation Paper • 2610.02196 • Published 5 days ago • 57
OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software Paper • 2609.39903 • Published 6 days ago • 61