Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions Paper • 2608.29109 • Published 14 days ago • 15
Using Grounded Theory for Agent Behavior Analysis at Scale Paper • 2608.30391 • Published 12 days ago • 19
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published Aug 7 • 111
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 263
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval Paper • 2608.01481 • Published Aug 2 • 72
CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Paper • 2608.02589 • Published Aug 3 • 25
TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning Paper • 2608.04007 • Published Aug 4 • 19
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications Paper • 2607.28617 • Published Jul 30 • 37