arxiv:2609.35505
Shangzhe Li
DVA13304
AI & ML interests
Reinforcement Learning, Imitation Learning, Learning Theory
Recent Activity
authored a paper about 23 hours ago
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It authored a paper 1 day ago
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM ReasoningOrganizations
None yet