The Embedder's Dilemma: LLMs Are Better, but at What Cost? Paper • 2608.12875 • Published 14 days ago • 14
DarwinX: Evolving Agent Harnesses Through Natural Selection Paper • 2608.07545 • Published 27 days ago • 113
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published 21 days ago • 62
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published Jul 26 • 106
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents Paper • 2607.20709 • Published Jul 22 • 35
Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness Paper • 2607.19322 • Published Jul 21 • 11
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing Paper • 2607.07953 • Published Jul 8 • 16
Quantifying and Expanding the Theoretical Capacity of Late-Interaction Retrieval Models Paper • 2607.05803 • Published Jul 7 • 10
LLM-as-a-Verifier: A General-Purpose Verification Framework Paper • 2607.05391 • Published Jul 6 • 18
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Paper • 2606.18216 • Published Jun 16 • 65
Prompt-Level Distillation: A Non-Parametric Alternative to Model Fine-Tuning for Efficient Reasoning Paper • 2602.21103 • Published Jun 2 • 9
CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning Paper • 2605.28742 • Published May 27 • 4
Reinforcement Learning from Rich Feedback with Distributional DAgger Paper • 2606.05152 • Published Jun 3 • 3
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters Paper • 2606.02437 • Published Jun 1 • 243
Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention Paper • 2605.29548 • Published May 28 • 13