RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 9 days ago • 277
Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models Paper • 2609.39820 • Published 10 days ago • 29
Language Models that Play Chess and Explain Their Moves Paper • 2610.03695 • Published 8 days ago • 37
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 14 days ago • 326
Post-Training Leaves Behavioral Shadows on Unrelated Decisions Paper • 2609.29233 • Published 16 days ago • 274 • 2
Selecting The Most Informative Tokens in Natural Language Autoencoders Paper • 2609.37040 • Published 11 days ago • 18
Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression Paper • 2609.36322 • Published 12 days ago • 115
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies Paper • 2609.38155 • Published 11 days ago • 115
VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models Paper • 2609.32607 • Published 14 days ago • 155