Running on CPU Upgrade Featured 3.3k The Smol Training Playbook 📚 3.3k The secrets to building world-class LLMs
Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation Paper • 2608.29846 • Published 16 days ago • 16
It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning Paper • 2609.00638 • Published 14 days ago • 74