Agentic RAG Evaluation: Budget Allocation Across Questions, Trajectories, and Reads Paper • 2610.05034 • Published 6 days ago • 19
Agentic RAG Evaluation: Budget Allocation Across Questions, Trajectories, and Reads Paper • 2610.05034 • Published 6 days ago • 19
Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts Paper • 2610.00314 • Published 11 days ago • 102
Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts Paper • 2610.00314 • Published 11 days ago • 102
Rules to Tools: Executable Checks for LLM Agents in Scientific Computing Paper • 2610.00313 • Published 11 days ago • 11
Rules to Tools: Executable Checks for LLM Agents in Scientific Computing Paper • 2610.00313 • Published 11 days ago • 11
ControlScope: Workflow Revision and Reliability in LLM Agents Paper • 2609.34313 • Published 12 days ago • 2
ControlScope: Workflow Revision and Reliability in LLM Agents Paper • 2609.34313 • Published 12 days ago • 2
WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents Paper • 2609.27490 • Published 17 days ago • 14
WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents Paper • 2609.27490 • Published 17 days ago • 14
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents Paper • 2609.09219 • Published Sep 7 • 16
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents Paper • 2609.09219 • Published Sep 7 • 16
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents Paper • 2609.09219 • Published Sep 7 • 16
Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA Paper • 2608.22856 • Published Aug 24 • 4
Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA Paper • 2608.22856 • Published Aug 24 • 4
Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA Paper • 2608.22856 • Published Aug 24 • 4
Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes Paper • 2605.05724 • Published May 7 • 16
Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes Paper • 2605.05724 • Published May 7 • 16
Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes Paper • 2605.05724 • Published May 7 • 16
SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks Paper • 2604.20087 • Published Apr 22 • 15