Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It Paper • 2609.32444 • Published 5 days ago • 24
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 13 days ago • 138
deepseek-ai/DeepSeek-V4.1-Flash Image-Text-to-Text • 763B • Updated about 5 hours ago • 748k • • 3.95k
NeoHorse-1 Collection NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness • 3 items • Updated 22 days ago • 8
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 23 days ago • 327
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published Aug 31 • 63
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp Image-Text-to-Text • 305B • Updated about 1 month ago • 966k • • 930