Your Agent Aced the Task. Will It Do It Again?
• 105
Enterprise AI and ML, Foundation Models, Responsible AI
DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents
Evaluating Autonomous AI Agents for Industry 4.0 Tasks
2 dozen real-life agents optimized for open models & stack
Benchmark asset operation performance in the browser
Configurable Generalist Agent, leader in AppWorld Benchmark
2 dozen real-life agents optimized for open models & stack
Benchmark asset operation performance in the browser
Configurable Generalist Agent, leader in AppWorld Benchmark
Evaluating Autonomous AI Agents for Industry 4.0 Tasks