RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 10 days ago • 279
SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation Paper • 2610.02304 • Published 10 days ago • 55
HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents Paper • 2610.03574 • Published 9 days ago • 60
Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers Paper • 2610.00531 • Published 11 days ago • 61
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 15 days ago • 326
EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation Paper • 2609.38157 • Published 12 days ago • 21
HiRAE: Hierarchical Representation Autoencoding with Residual Budgets Paper • 2609.37775 • Published 12 days ago • 26
EVO-WAM: Evolving World Action Models through Video-Action Verification Paper • 2609.38057 • Published 12 days ago • 49
MoreThought/Fable-5.1-Max-Reasoning-Filtered-10000x Viewer • Updated about 3 hours ago • 15k • 5.17k • 295