CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty Paper • 2601.22027 • Published Jan 29 • 86
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 Paper • 2608.27370 • Published 5 days ago • 32
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction Paper • 2608.26005 • Published 6 days ago • 173
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 8 days ago • 205
OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs Paper • 2608.21360 • Published 11 days ago • 30
Towards Quantifying Benchmark Optimization in ASR Models Paper • 2608.19936 • Published 12 days ago • 12
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? Paper • 2608.19799 • Published 12 days ago • 65
EnvHarness: Awakening Static Worlds for Agent Learning Paper • 2608.19880 • Published 12 days ago • 273
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence Paper • 2608.11341 • Published 21 days ago • 65
Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation Paper • 2607.27372 • Published Jul 29 • 19
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training Paper • 2606.30406 • Published Jun 29 • 25
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms Paper • 2607.26497 • Published Jul 30 • 52
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published Jul 30 • 185