RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 7 days ago • 275
What Makes Recurrence Effective in Looped Language Models? Paper • 2609.36636 • Published 9 days ago • 10
PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond? Paper • 2609.34314 • Published 10 days ago • 4
WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation Paper • 2609.37687 • Published 9 days ago • 6
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 12 days ago • 325
orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF Text Generation • 27B • Updated 6 days ago • 20.6k • 467