Papers
arxiv:2609.35505

An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning

Published on Sep 28
· Submitted by
Shangzhe Li
on Sep 29
Authors:
,
,
,
,
,

Abstract

We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that brings optimistic exploration and off-policy data reuse from value-based RL into policy distillation. LSPD preserves policy diversity through exploration while improving rollout efficiency by repeatedly learning from previously collected trajectories. Our theoretical analysis connects LSPD to optimistic value-based learning and shows that its idealized formulation achieves a sharp mathcal O(log K) regret bound under online exploration. Empirically, LSPD consistently outperforms existing distillation baselines across six mathematical reasoning benchmarks and diverse teacher-student settings, with average gains of +1.59 points in Avg@16. Remarkably, through Pass@k evaluations up to k=64, we found that LSPD better preserves policy diversity by achieving stronger performance as k grows. Its fully off-policy variant achieves comparable performance to vanilla OPD using only the first 25% of rollout batches. Together, these results provide an RL perspective on OPD that offers both a principled interpretation and a practical route toward more effective and rollout-efficient language model distillation.

Community

Paper author Paper submitter
•
edited about 10 hours ago

LSPD: An RL View of On-Policy Distillation

On-policy distillation (OPD) is commonly implemented with policy-gradient updates. But its reverse-KL objective also has an exact interpretation as KL-regularized reinforcement learning, with a reward defined by the teacher relative to a reference policy. This raises a natural question: can value-based RL principles make policy distillation more sample-efficient?

We introduce Least Square Policy Distillation (LSPD), motivated by optimistic reward estimation and KL-regularized policy improvement. Rather than using teacher feedback only to drive policy-gradient updates, LSPD directly matches student and teacher log probabilities through a robust least-squares objective. This formulation enables learning from accumulated experience, bringing a central advantage of value-based RL to LLM distillation. Our theoretical analysis establishes logarithmic reverse-KL regret for an idealized optimistic formulation.

Across six mathematical reasoning benchmarks and three teacher–student settings, LSPD improves reasoning performance and multi-sample solution coverage. Its replay-based variant, LSPD-RB, achieves performance comparable to vanilla OPD with approximately one-quarter of the rollout batches. The broader message: viewing distillation as an RL problem opens up algorithmic choices beyond on-policy policy gradients—and opportunities to learn more from each interaction with the teacher.

📄 Paper · 💻 Code

Paper author Paper submitter
•
This comment has been hidden

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.35505
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.35505 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.35505 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.35505 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.