DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation Paper • 2610.04596 • Published 7 days ago • 19
Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents Paper • 2610.01892 • Published 9 days ago • 27
andersonbcdefg/red_teaming_reward_modeling_pairwise_no_as_an_ai Viewer • Updated Jun 1, 2023 • 35.3k • 264 • 13
LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches Paper • 2610.06647 • Published 5 days ago • 42
The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models Paper • 2610.02191 • Published 9 days ago • 14
AdityaaXD/Multi-Agent_Reinforcement_Learning_Trading_System_Data Viewer • Updated Feb 1 • 5.28k • 252 • 12
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies Paper • 2609.38155 • Published 11 days ago • 115
AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation Paper • 2609.35530 • Published 12 days ago • 22
n1ghtf4l1/Agentic-Diagnostic-Reasoning-with-Multimodal-SLMs-via-Reinforcement-Learning Updated Nov 21, 2025 • 170 • 15
leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset Viewer • Updated Aug 12, 2025 • 14k • 163 • 11
Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR Paper • 2609.37868 • Published 11 days ago • 65
What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling Paper • 2609.34981 • Published 11 days ago • 138
tomyimkc/repro-optimal-regret-for-policy-optimization-in-contextual-bandits-traces Traces • Updated Jul 26 • 1 • 100 • 3
open-source-metrics/reinforcement-learning-checkpoint-downloads Viewer • Updated Oct 6, 2022 • 367 • 174 • 9