view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 15 days ago • 122
view article Article BenchMIRT: What are LLM benchmarks actually measuring? allenai • 16 days ago • 22
view article Article Measuring benchmark optimization in speech recognition +5 tlebryk02, bezzam, aliceebaird, dayllon, jpc, jens-hume-ai, tzirakis • 28 days ago • 66
view article Article State of Open Models: Summer 2026 Observations +1 AdinaY, multimodalart, irenesolaiman • Aug 14 • 203
view article Article Introducing North Mini Code: Cohere’s First Model For Developers CohereLabs • Jun 9 • 86
view article Article Task-Seeded Synthetic Q&A Generation for Nemotron Pretraining nvidia • Jun 4 • 17
view article Article Harness, Scaffold, and the AI Agent Terms Worth Getting Right sergiopaniego, ariG23498 • May 25 • 147
The ATOM Report: Measuring the Open Language Model Ecosystem Paper • 2604.07190 • Published Apr 8 • 5
view article Article Keep the Tokens Flowing: Lessons from 16 Open-Source RL Libraries +7 aminediroHF, qgallouedec, kashif, lewtun, edbeeching, albertvillanova, nouamanetazi, lvwerra, sergiopaniego • Mar 10 • 189
view article Article Ulysses Sequence Parallelism: Training with Million-Token Contexts kashif, stas • Mar 9 • 33