Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence Paper • 2608.31075 • Published 3 days ago • 26
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces Paper • 2604.05172 • Published Apr 6 • 25
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks Paper • 2602.12670 • Published Feb 13 • 65