LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures Paper • 2610.04292 • Published 8 days ago • 38
Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI Paper • 2609.38143 • Published 12 days ago • 85
Learning from Teacher Continuations at Student States Paper • 2609.36246 • Published 13 days ago • 41
Selecting Diverse SFT Traces Improves Post-RL Generalization Paper • 2609.33780 • Published 14 days ago • 38
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published Sep 3 • 188
Evaluating the Hidden Costs of Personalization in Large Language Models Paper • 2608.28833 • Published Aug 28 • 30
PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments Paper • 2608.14441 • Published Aug 14 • 29
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published Aug 7 • 114
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published Aug 12 • 118
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published Jul 30 • 189
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Paper • 2606.22388 • Published Jun 21 • 96
Brick-Composer: Using MLLMs for Assembly with Diverse Bricks Paper • 2606.05445 • Published Jun 3 • 8
AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints Paper • 2606.05622 • Published Jun 4 • 45
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe Paper • 2604.13016 • Published Apr 14 • 116
CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing Paper • 2605.02910 • Published May 6 • 24
UserRL: Training Interactive User-Centric Agent via Reinforcement Learning Paper • 2509.19736 • Published Sep 24, 2025 • 13