RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 6 days ago • 268
Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features Paper • 2609.33463 • Published 10 days ago • 11
Equal Ranking Quality, Different Decisions: Measuring and Reducing Order Dependence in LLM Scorers Paper • 2608.26762 • Published 11 days ago • 17
Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models Paper • 2609.39820 • Published 7 days ago • 15
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 11 days ago • 324