OPDHub

A Survey of On-Policy Distillation for Large Language Models

23214311912151979392025-0102030405060708091011122026-010203040506079

On-Policy Distillation papers, monthly · from 2025-01 · 190 papers

Loss
FKL34 RKL43 Symmetric18 f-Divergence1 KL+RL42 Preference2 Other57
Teacher
External-Teacher109 Self48 Privileged-Info34 Verifier4 Other-Signal2
Topic
Math16 Code5 Reasoning43 Agent22 Multimodal21 Alignment12 Inference4 Preference2 General72
Cadence
per-step157 per-outer-iter37 once-before-training3
Size
<3B68 3-10B93 >30B3 Unknown33
Year
20234 20243 202517 2026173
Recent
2026-0639 2026-0579 2026-0419
 

§4 Objective Functions and Optimization

§4.1 — Fixed Divergence Objectives

  • Reinforcement Learning from Rich Feedback with Distributional DAggerarXiv 2026-06§4.1FKLQwen3-8B → Self; Distributional DAgger via forward cross-entropy: monotonic-improvement objective with future-aware credit assignment, an OPD analogue of RL distributional bootstrapping.
  • OPD+: Rethinking the Advantage Design for On-Policy DistillationarXiv 2026-06§4.1f-DivergenceQwen3-8B → Qwen3-8B-Base; Corrects advantage estimation in on-policy distillation via f-divergence gradient analysis
  • Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual GroundingarXiv 2026-06§4.1RKLQwen3-VL-8B-Instruct → Qwen3-VL-2B-Instruct; Decomposes VLM on-policy distillation into language prior and visual grounding, steering gradients toward visual subspace
  • Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future GuidancearXiv 2026-06§4.1RKLQwen3-30B-A3B-Instruct-2507 → Qwen3-4B-Instruct-2507; Trajectory-aware OPD using OT-based near-future guidance to fix token-level reasoning correction failures
  • Anti-Self-Distillation for Reasoning RL via Pointwise Mutual InformationarXiv 2026-05§4.1KL+RLQwen3-4B/8B/14B/30B → Self; reverses divergence direction to boost deliberation tokens via pMI sign flip; entropy-triggered gate
  • KL for a KL: On-Policy Distillation with Control Variate BaselinearXiv 2026-05§4.1RKLQwen3-1.7B/4B-Base → Qwen3-1.7B/4B-Inst (self-distill), OLMo-3-7B
  • Surgical Post-Training: Proximal On-Policy Distillation for Reasoning with Knowledge RetentionarXiv 2026-03§4.1OtherQwen3-8B → Self (oracle-rectified); Oracle-rectified proximal on-policy data + reward-based BCE; 4k math pairs, 16-min training on 8xH800
  • Distillation of Large Language Models via Concrete Score MatchingarXiv 2025-09§4.1Other CodeGPT-2 0.1B–0.3B → GPT-2 1.5B / OpenLLaMA-7B
  • DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMsarXiv 2025-03§4.1Symmetric CodeQwen2-1.5B / Gemma-2-2B → Qwen2-7B / Gemma-2-9B
  • DistiLLM: Towards Streamlined Distillation for Large Language ModelsarXiv 2024-02§4.1Symmetric CodeGPT-2 (student) → GPT-2 XL (teacher)
  • On-Policy Distillation of Language Models: Learning from Self-Generated MistakesarXiv 2023-06§4.1FKLT5-Small/Base/Large → T5-XL 3B

§4.2 — Adaptive Divergence Objectives

  • Stabilizing On-Policy Distillation for MLLM Reasoning with Global NormalizationarXiv 2026-06§4.2OtherNormalization of KL signal to batch-relative advantages is an adaptive modification of the distillation divergence (§4.2)…
  • Trust Region On-Policy DistillationarXiv 2026-06§4.2SymmetricSkywork-OR1-Math-7B → DeepSeek-R1-Distill-Qwen-1.5B; Trust-region OPD with outlier estimation and off-policy guidance for stable reasoning distillation
  • RAFT: Data Refinement and Adaptive Distillation for Domain Fine-Tuning with Alleviated ForgettingarXiv 2026-06§4.2FKLSmolLM3-3B → Self; Two-stage framework coupling data refinement with on-policy distillation to mitigate forgetting in domain SFT
  • Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy DistillationarXiv 2026-05§4.2KL+RLSkyWork-OR1-Math-7B → DeepSeek-R1-Distill-Qwen-1.5B; Identifies Supervision Fidelity Decay in OPD and proposes Lookahead Group Reward to combat it
  • Not All Disagreement Is Learnable: Token Teachability in On-Policy DistillationarXiv 2026-05§4.2FKLQwen3-8B → 4B; Binary teachability mask selects 5-10% tokens for budgeted RKL, filtering unreliable teacher signals
  • Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM ReasoningarXiv 2026-05§4.2SymmetricQwen3-4B → Self (privileged); Entropy-routed direction-adaptive self-distillation reversing teacher pressure at high-entropy tokens.
  • When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for ReasoningarXiv 2026-05§4.2FKL CodeQwen3-4B → Self; Position-weighted clipped FKL: later reasoning tokens get higher weight due to accumulated teacher error
  • Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token LevelarXiv 2026-05§4.2KL+RLQwen3-8B-Base / Qwen3-4B-Base → Qwen3-32B / Qwen3-8B; replaces negative RL with localized divergence minimization (AOPD)
  • Scaling Reasoning Efficiently via Relaxed On-Policy DistillationarXiv 2026-03§4.2RKLDeepSeek-R1-Distill-Qwen-1.5B → SkyWork-OR1-7B/32B
  • Entropy-Aware On-Policy Distillation of Language ModelsarXiv 2026-03§4.2SymmetricQwen3-0.6B/1.7B/4B → Qwen3-8B
  • Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM TrainingarXiv 2026-02§4.2Other CodeQwen2.5-7B / DeepSeek-R1-Distill-7B → Self (on-policy SFT)
  • Distribution-Aligned Sequence Distillation for Superior Long-CoT ReasoningarXiv 2026-01§4.2Other CodeQwen3-4B → gpt-oss-120b / Qwen3-Next-80B-A3B-Thinking (DASD)
  • Stable On-Policy Distillation through Adaptive Target ReformulationarXiv 2026-01§4.2SymmetricQwen2-0.5B-Instruct → Qwen2-7B-Instruct

§4.3 — RL-Augmented Objectives

  • RLCSD: Reinforcement Learning with Contrastive On-Policy Self-DistillationarXiv 2026-06§4.3OtherSelf (correct-hint) + Self (wrong-hint) → Student; Contrastive RKL inside GRPO: two-path loss pulls student toward correct-hint teacher and away from wrong-hint teacher simultaneously
  • Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal DataarXiv 2026-06§4.3OtherGPT-5.4 → Qwen3-4B-Base; Calibration-coupled GRPO + masked judge distillation: external judge scores distilled into self-evaluation tokens only, leaving the answer untouched.
  • StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement LearningarXiv 2026-05§4.3OtherQwen2.5-3B / Qwen3-1.7B → Self; Advantage-integrated OPD: teacher-student log-ratio fused into GRPO advantage for agentic tasks
  • OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM ReasoningarXiv 2026-05§4.3OtherQwen3-32B → Qwen3-4B; Bayesian token-level credit via oracle-conditioned likelihood ratios in PPO-style update.
  • AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit AssignmentarXiv 2026-05§4.3OtherQwen2.5-7B / Qwen3-8B → Self; CIG (pointwise KL) modulates PPO advantage; meta-reflective teacher conditions on privileged info
  • Teacher-Guided Policy Optimization for On-Policy Reasoning Distillation under Large Policy DivergencearXiv 2026-05§4.3KL+RLQwen2.5-Math-1.5B / Qwen2.5-Math-7B → Qwen3-30B-A3B / R1-Distill-Qwen-32B; dense directional teacher guidance on student rollouts; fixes uninformative RKL negatives (NLP2CT/NEU)
  • Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-TrainingarXiv 2026-05§4.3RKLSparse RL on teacher (GRPO) → dense OPD bridge to student; Qwen3/Llama reward-density allocation rule
  • Combining On-Policy Optimization and Distillation for Long-Context Reasoning in Large Language ModelsarXiv 2026-05§4.3KL+RLQwen3-1.7B → Qwen3-32B; augments GRPO with dense OPD teacher guidance for long-context; introduces LongBlocks benchmark
  • CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy OptimizationarXiv 2026-05§4.3KL+RLQwen2.5-Math-1.5B / Qwen2.5-Math-7B → Qwen2.5-Math-7B / Qwen2.5-Math-1.5B; bidirectional co-distillation (Google)
  • MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent DebatearXiv 2026-05§4.3Symmetric CodeQwen3-1.7B/4B/8B/14B → Multi-teacher debate; confidence-weighted token supervision (OPAD)
  • Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge DistillationarXiv 2026-03§4.3Other CodeQwen2.5-1.5B / Gemma-2-2B / Qwen3-1.7B → Qwen2.5-14B / Gemma-2-9B / Qwen3-8B
  • Reinforcement-aware Knowledge Distillation for LLM ReasoningarXiv 2026-02§4.3KL+RLQwen3-0.6B–8B → Qwen3-8B / Qwen3-32B (RLAD)
  • X-KD: General Experiential Knowledge Distillation for Large Language ModelsarXiv 2026-02§4.3KL+RLT5-Small/Base → T5-Large 780M
  • Learning beyond Teacher: Generalized On-Policy Distillation with Reward ExtrapolationarXiv 2026-02§4.3KL+RL CodeQwen3-4B-Non-Thinking → Self-RL teachers / Qwen3-30B-A3B (G-OPD)
  • KEPO: Knowledge-Enhanced Preference Optimization for Multimodal Reasoning with Applications to Medical VQAarXiv 2026-02§4.3KL+RLQwen3-VL-2B / Qwen3-VL-8B → Qwen3-VL-32B (KEPO)
  • Rethinking Large Language Model Distillation: A Constrained Markov Decision Process PerspectivearXiv 2025-09§4.3KL+RLQwen2.5-1.5B-Math / Llama-3.2-3B → Qwen2.5-7B-Math / Llama-3.2-11B
  • From Correction to Mastery: Reinforced Distillation of Large Language Model AgentsarXiv 2025-09§4.3Other CodeQwen2.5-7B / Qwen3-8B → Self (SCoRe, agent)
  • KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement LearningarXiv 2025-06§4.3KL+RLR1-Distill-Qwen-1.5B → Skywork-OR1-Math-7B (KDRL)
  • RLKD: Distilling LLMs' Reasoning via Reinforcement LearningarXiv 2025-05§4.3Other CodeQwen2.5-Math-7B / R1-Distill-Qwen-7B → DeepSeek-R1 traces (RLKD)
  • KETCHUP: K-Step Return Estimation for Sequential Knowledge DistillationarXiv 2025-04§4.3OtherT5-Base 250M → FLAN-T5-XL 3B
  • Learning to Reason under Off-Policy GuidancearXiv 2025-04§4.3Other CodeQwen2.5-Math-7B → Self; RM: off-policy DeepSeek-R1
  • AlignDistil: Token-Level Language Model Alignment as Adaptive Policy DistillationarXiv 2025-03§4.3RKL CodeQwen2-1.5B / Qwen2.5-1.5B-Instruct → Self (AlignDistil)

§5 Signal Source and Teacher Architecture

§5.1 — White-Box Logit Supervision

  • Breaking the Tokenizer Barrier: On-Policy Distillation across Model FamiliesarXiv 2026-06§5.1OtherCross-family Teacher → Cross-family Student; Token-mapping enables on-policy distillation across model families with different tokenizers
  • OPRD: On-Policy Representation DistillationarXiv 2026-06§5.1OtherExternal Teacher → Student; Extends OPD from logit space to hidden-state representation alignment, reducing Monte Carlo KL variance over large vocabularies
  • DuDi: Dual-Signal Distillation with Cross-Lingual VerbalizerarXiv 2026-06§5.1KL+RL CodeQwen2.5-3B-Instruct → Qwen2.5-0.5B; Dual-signal distillation: online sequence-level SPIN objective combined with off-policy + on-policy token-level KD via a cross-lingual verbalizer.
  • Pair-In, Pair-Out: Latent Multi-Token Prediction for Efficient LLMsarXiv 2026-05§5.1FKLQwen3.5-9B → compressed Qwen3.5 (latent MTP); On-policy distillation stage with reverse-KL on student rollouts + auxiliary confidence-head BCE loss, used to recover accuracy of latent multi-token-prediction compressor trained on DAPO-Math + Codeforces
  • Multi-Rollout On-Policy Distillation via Peer Successes and FailuresarXiv 2026-05§5.1SymmetricQwen3-8B → Qwen3-32B; peer-conditioned teacher signals from success/failure rollout groups; more faithful supervision (CMU)
  • On-Policy Distillation with Best-of-N Teacher Rollout SelectionarXiv 2026-05§5.1Symmetric CodeDeepSeek-R1-Distill-Qwen-1.5B → JustRL-DeepSeek-1.5B / DeepSeek-R1-Distill-Qwen-7B; samples teacher trajectory pool, selects via correctness-first / alignment-second priority
  • Reasoning Compression with Mixed-Policy DistillationarXiv 2026-05§5.1RKLQwen3-1.7B → Qwen3-8B; teacher rewrites student's verbose trajectories concisely; distills compressed reasoning
  • SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy DistillationarXiv 2026-05§5.1RKLQwen2.5-7B-Inst / Phi-4-mini → Phi-4-mini / Gemma-2-2B-IT; multi-token continuation units for cross-tokenizer OPD
  • A Dual-Space Framework for General Knowledge Distillation of Large Language ModelsarXiv 2025-04§5.1FKL CodeGPT-2 120M / TinyLLaMA-1.1B → GPT-2 1.5B / Qwen2-1.5B
  • PromptKD: Distilling Student-Friendly Knowledge for Generative Language Models via Prompt TuningarXiv 2024-02§5.1RKL CodeGPT-2 120M–760M / OPT/Llama-7B → GPT-2 XL / OPT-13B / Llama-13B
  • MiniLLM: On-Policy Distillation of Large Language ModelsarXiv 2023-06§5.1RKL CodeGPT-2 120M–760M → GPT-2 1.5B / GPT-J 6B / OPT-13B

§5.2 — Black-Box and API-Constrained

  • OmniOPD: Logit-Free On-Policy Distillation via Speculative VerificationarXiv 2026-06§5.2OtherQwen3-32B → Qwen3-1.7B; Logit-free on-policy distillation using chunk-level Monte Carlo semantic verification from black-box teachers
  • Rubric-based On-policy DistillationarXiv 2026-05§5.2Other CodeGPT-5.2 / Qwen3-30B-A3B → Qwen3-4B / Gemma3-4B; structured semantic rubrics replace teacher logits; 10× sample efficiency
  • Pre-alignment via Black-box On-policy Distillation for Multimodal RLarXiv 2026-04§5.2Other CodeQwen3-VL-8B → Self; PRISM adversarial MoE discriminator, logit-free OPD as pre-alignment before RLVR
  • OVD: On-policy Verbal DistillationarXiv 2026-01§5.2OtherQwen2.5-3B / LLaMA-3.2-3B → QwQ-32B (verbal feedback)
  • Black-Box On-Policy Distillation of Large Language ModelsarXiv 2025-11§5.2OtherLlama-3.1-8B / Qwen2.5-3B–14B → GPT-5-Chat (black-box)
  • ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM DistillationarXiv 2025-09§5.2PreferenceTinyLlama-1.1B-Instruct / InternLM2.5-1.8B-Chat → InternLM2.5-7B-Chat; ORPO-Distill: black-box cross-architecture distillation via ORPO (SFT + log-odds margin); teacher CoT y_P sampled K=8 times offline, on-policy student y_N per outer-iter (mixed policy fraction phi)

§5.3 — Self-Distillation

  • Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer FeedbackarXiv 2026-06§5.3OtherQwen3-8B ↔ Qwen3-8B (peer); OPCoD: two coupled on-policy self-distillation loops, each self-teacher conditioned on own correct rollout + peer NL feedback; cognizance gating + feedback anchoring; cross-domain mutual Pareto improvement
  • Rubric-Guided Self-Distillation: Post-Training Without Rubric VerifiersarXiv 2026-06§5.3OtherSelf (w/ rubric) → Self; Rubric as privileged context for same-model teacher; JSD distillation eliminates external LLM verifier from open-ended post-training
  • When Context Returns: Toward Robust Internalization in On-Policy DistillationarXiv 2026-06§5.3OtherSelf (w/ privileged context) → Self; FKL no-context anchoring regularizer prevents context-induced degradation when privileged context is re-introduced at inference
  • HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-DistillationarXiv 2026-06§5.3OtherSelf (w/ hindsight env obs) → Self (no hindsight); Turn-level KL self-distillation using future environment observations as privileged context for multi-turn agents
  • Beyond Absolute Imitation: Anchored Residual Guidance for Privileged On-Policy DistillationarXiv 2026-06§5.3OtherOracle (privileged) → Student; AR-OPD: anchor + oracle residual prevents hindsight leakage; +2.3 vs full OPD, +7.9 vs SFT, -21.7% leakage
  • Self-Distillation Policy Optimization via Visual Feedback: Bridging Code and Visual ArtifactsarXiv 2026-06§5.3OtherSelf (w/ rendered artifact) → Code-LLM; Visual-SDPO: rendered visual artifacts as privileged feedback + statement-weighted KL + GRPO for code-to-visualization
  • PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit AssignmentarXiv 2026-06§5.3OtherSelf (w/ GT) → Self (w/o GT); PBSD: Bayesian self-distillation converts sparse trajectory rewards to turn-level credits for long-horizon RL
  • SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher SamplingarXiv 2026-06§5.3OtherTeacher → Student; SG-OPD: sign-consistency gating + phased teacher sampling for verifier-guided OPD; +1.98/+7.50 on math
  • Thinking Without Images: Internalizing Visual Manipulation with On-Policy Self-DistillationarXiv 2026-06§5.3OtherSelf (w/ cropped tiles) → Self (w/o crops); Privileged cropped-image self-teacher internalizes zoom-in reasoning so no image crops are needed at inference
  • Teaching the Way, Not the Answer: Privileged Tutoring Distillation for Multimodal Policy OptimizationarXiv 2026-06§5.3OtherPrivileged teacher → 2B-8B VLM; PTD-PO: Top-K JSD privileged tutoring with spatial+reasoning hints for multimodal policy optimization
  • Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy DistillationarXiv 2026-06§5.3OtherAR-LM (frozen) → Diffusion-LM; OPDLM: self-distillation converts AR LM to diffusion LM on-policy; 15x-7000x fewer training tokens
  • Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-DistillationarXiv 2026-06§5.3OtherSymbolic-state-privileged Teacher → Visual VLM Student; Two-stage MGSD: symbolic game state as privileged context bridges perception-reasoning modality gap for spatial planning
  • Self-Distilled Policy GradientarXiv 2026-06§5.3KL+RLQwen3-4B → Self; Full-vocabulary reverse-KL self-distillation gated by positive advantage; teacher = same model conditioned on ground-truth privileged info.
  • World Models Meet Language Models: On the Complementarity of Concrete and Abstract ReasoningarXiv 2026-06§5.3FKLQwen3.5-9B → Qwen3.6-27B (privileged-info teacher) / Gemini-3.1-Pro; Privileged-info self-distillation: future videos + ground-truth answers as privileged context teach MLLM when to invoke / verify / rely on world-model rollouts; D_KL teacher term
  • Constitutional On-Policy Safe DistillationarXiv 2026-06§5.3KL+RLQwen3-VL-4B → Self; On-policy self-distillation with safety-constitution privileged context as teacher; cross-SFT cold-start aligns base/instruct teachers.
  • COMAP: Co-Evolving World Models and Agent Policies for LLM AgentsarXiv 2026-06§5.3OtherQwen3-4B → Self; Co-evolving textual world models and agent policies via on-policy self-distillation and future-aware reflection
  • Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable OversightarXiv 2026-06§5.3RKLQwen3-4B-base → Self; On-policy critique distillation using weak model critiques to improve strong models
  • Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language ModelsarXiv 2026-05§5.3RKLQwen3-8B → Self; On-policy self-distillation aligning RAW-SHARDED multi-turn answers with FULL-context teacher behavior
  • Skill-Conditioned Gated Self-Distillation for LLM ReasoningarXiv 2026-05§5.3OtherQwen3-1.7B/4B/8B → Self (skill-conditioned); Skill-conditioned multi-teacher pool with outcome-validated teacher polarity; bounded gated distillation objective.
  • ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across DomainsarXiv 2026-05§5.3Symmetric CodeQwen3-4B/8B (self-teacher) → Qwen3-4B/8B; Error-focused reflection + quote-localized self-distillation; reflector extracts corrective idea and error span, distillation loss applied only from error onward
  • MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-DistillationarXiv 2026-05§5.3FKLQwen2.5-3B / Qwen2.5-7B / Llama-3.1-8B → Self; EMA self-teacher + GJD/RKL for multi-turn dialogue; history-cleaned prompts prevent conversation drift
  • Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented ReasoningarXiv 2026-05§5.3FKLQwen2.5-7B → Self (privileged context); Alternating GRPO with offline self-distillation between rounds for search-augmented reasoning self-evolution.
  • Unlocking Proactivity in Task-Oriented DialoguearXiv 2026-05§5.3FKLQwen3-4B → Self (privileged view); Asymmetric self-distillation from privileged user-concern view plus state-transition policy gradient for proactive TOD.
  • On-Policy Consistency Training Improves LLM Safety with Minimal Capability DegradationarXiv 2026-05§5.3FKLLlama-3.1-8B / Qwen2.5-7B / Qwen3-8B → Self; Per-token reverse KL on contrastive prompt pairs for safety alignment (anti-sycophancy, jailbreak defense)
  • AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged SignalsarXiv 2026-05§5.3FKL CodeSelf → Self (multi-view PI); Multi-view on-policy self-distillation decomposing privileged teacher signals into geometric consensus + gated residuals
  • It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMsarXiv 2026-05§5.3RKLQwen2.5-7B → Self; Complementary self-distillation: two feedback-conditioned self-teachers (utility / privacy) provide joint reverse-KL token-level supervision over on-policy rollouts for contextual integrity alignment
  • Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-DistillationarXiv 2026-05§5.3FKL CodeQwen3.5-4B/9B → Self (crop→full-image); regional-to-global self-distillation with on-policy rollouts + token-level JSD; VLM self-distillation: crop-conditioned teacher distills fine-grained visual details to full-image student via on-policy JSD
  • SD-Search: On-Policy Hindsight Self-Distillation for Search-Augmented ReasoningarXiv 2026-05§5.3SymmetricQwen2.5-3B (hindsight-conditioned) → Qwen2.5-3B; On-policy hindsight self-distillation for step-level search query supervision in RL agents
  • HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon AgentsarXiv 2026-05§5.3FKLQwen3-4B-Instruct-2507 (EMA + feedback-conditioned) → Qwen3-4B-Instruct-2507; Targeted self-distillation applying feedback-conditioned teacher only at failure-relevant turns in long-horizon agent tr
  • Self-Supervised On-Policy Distillation for Reasoning Language ModelsarXiv 2026-05§5.3FKLQwen3-8B (stop-gradient self) → Qwen3-8B; Self-supervised on-policy distillation using intra-group correct-wrong contrast as dense process supervision
  • Self-Distilled Agentic Reinforcement LearningarXiv 2026-05§5.3KL+RLQwen2.5/Qwen3 → Self; sigmoid-gated OPSD auxiliary with RL; asymmetric positive/negative teacher signal; +9.4%/+10.2%/+7.0% over GRPO on ALFWorld/WebShop/SearchQA
  • Learning from Language Feedback via Variational Policy DistillationarXiv 2026-05§5.3RKLLLM → Self (co-evolved); Variational EM co-optimizes teacher+student; adaptive trust-region teacher update from language feedback; outperforms RLVR+SDPO on code/science reasoning (Salesforce)
  • RESD: Learning with Rare Success but Rich Feedback via Reflection-Enhanced Self-DistillationarXiv 2026-05§5.3RKLLLM agents → Self; retrospective reflection on failures generates corrective self-supervision + persistent playbook; outperforms GRPO 8× (Amazon/UCSD)
  • OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM ReasoningarXiv 2026-05§5.3RKLQwen3-8B → Self; outcome rewards contrast correct vs. failed on-policy trajectories to calibrate teacher logits
  • From Generic Correlation to Input-Specific Credit in On-Policy Self DistillationarXiv 2026-05§5.3KL+RLQwen3-8B → Self; pMI decomposition of self-distillation reward; batch-contrastive baseline isolates input-specific credit
  • Adaptive Teacher Exposure for Self-Distillation in LLM ReasoningarXiv 2026-05§5.3FKLQwen3-1.7B/4B/8B → Self; learnable Beta-policy controller for teacher exposure ratio (ByteDance)
  • Efficient LLM Reasoning via Variational Posterior Guidance with Efficiency AwarenessarXiv 2026-05§5.3KL+RLDeepSeek-R1-Distill-Qwen-1.5B/7B / DeepSeek-R1-Distill-Llama-8B → Self (dual-stream); VPG-EA: posterior (answer-conditioned) and prior streams share params; advantage-gated forward KL distillation; cross-view validation filters pseudo-efficient paths
  • Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVRarXiv 2026-05§5.3KL+RLQwen3-8B → Self; RLRT: inverted self-distillation signal; reinforces student's reasoning tokens via GRPO
  • Crosslingual On-Policy Self-Distillation for Multilingual ReasoningarXiv 2026-05§5.3FKL CodeQwen3-8B → Self; English translation + reference solution as privileged context for 17 low-resource languages
  • UniSD: Towards a Unified Self-Distillation Framework for Large Language ModelsarXiv 2026-05§5.3Symmetric CodeLlama-3.1/Qwen2.5/Phi-3 families → Self (EMA teacher + multi-teacher agreement + divergence clipping)
  • Preference-Based Self-Distillation: Beyond KL Matching via Reward RegularizationarXiv 2026-05§5.3PreferenceQwen3-1.7B/4B/8B → Self (context-augmented); PBSD: black-box OPD via DPO with privileged-context self-teacher y+ vs on-policy student y-; per-step rollouts; reward-regularized preference gap (no logits available, hence DPO not KL)
  • Multilingual Safety Alignment via Self-DistillationarXiv 2026-05§5.3RKLQwen2.5-7B / Llama-3-8B → Self; MSD: English CoT as privileged context + Dual-Perspective Safety Weighting
  • Healthcare AI GYM for Medical AgentsarXiv 2026-05§5.3KL+RLQwen3-8B → Self; TT-OPD: EMA teacher + outcome-privileged hints + turn-level KL for multi-turn agentic distillation
  • Learn where to Click from Yourself: On-Policy Self-Distillation for GUI GroundingarXiv 2026-05§5.3RKLQwen3-VL-8B → Self; GUI-SD: visual privileged context (bounding box + Gaussian soft mask) + entropy-guided token weighting
  • Partial-Solution Adaptive Interpolated Training for Self-Distilled ReasonersarXiv 2026-04§5.3FKLQwen3-4B/8B → Self; PAINT: rollout-reference overlap + energy interpolation on OPSD
  • OPSDL: On-Policy Self-Distillation for Long-Context Language ModelsarXiv 2026-04§5.3RKLQwen2.5-Instruct-7B–32B → Self (short-context as privileged teacher for long-context, per-token reverse-KL)
  • π-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External DataarXiv 2026-04§5.3KL+RL CodeQwen3-4B / Qwen3-4B-Instruct / Qwen3-8B → Self (QCP as privileged context for dense supervision)
  • Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense SupervisionarXiv 2026-04§5.3RKLQwen3-4B-Instruct / Olmo-3-7B-Instruct → Self
  • Self-Distilled RLVRarXiv 2026-04§5.3KL+RLQwen3-VL-4B / Qwen3-VL-8B → Self (RLSD)
  • Unifying Group-Relative and Self-Distillation Policy Optimization via Sample RoutingarXiv 2026-04§5.3KL+RLQwen3-4B / Qwen3-8B → Self (SRPO)
  • HDPO: Hybrid Distillation Policy Optimization via Privileged Self-DistillationarXiv 2026-03§5.3KL+RLQwen2.5-Math-1.5B-Instruct → Self (HDPO)
  • Online Experiential Learning for Language ModelsarXiv 2026-03§5.3RKLQwen3-1.7B / Qwen3-4B / Qwen3-8B → Self (OEL)
  • CRISP: Compressed Reasoning via Iterative Self-Policy DistillationarXiv 2026-03§5.3RKL CodeQwen3-VL-8B → Self; GUI-SD: visual privileged context (bounding box + Gaussian soft mask) + entropy-guided token weighting
  • GATES: Self-Distillation under Privileged Context with Consensus GatingarXiv 2026-02§5.3OtherQwen3-4B → Self; oracle: Qwen2.5-32B (privileged gating)
  • On-Policy Context Distillation for Language ModelsarXiv 2026-02§5.3RKL CodeQwen3-1.7B/4B/8B → Qwen3-8B (thinking, OPCD)
  • Multi-Token Prediction via Self-DistillationarXiv 2026-02§5.3FKL CodeLlama-3.1-8B → Self (online distillation for 3× faster decoding)
  • Privileged Information Distillation for Language ModelsarXiv 2026-02§5.3KL+RLQwen3-4B / Qwen3-8B → Self (privileged info, π-Distill)
  • Expanding the Capabilities of Reinforcement Learning via Text FeedbackarXiv 2026-02§5.3OtherLlama-3.1-8B-Instruct → Qwen3-235B (text feedback RL)
  • Reinforcement Learning via Self-DistillationarXiv 2026-01§5.3RKL CodeQwen3-8B → Self (SDPO, iterative)
  • Self-Distillation Enables Continual LearningarXiv 2026-01§5.3RKLQwen2.5-7B-Instruct → Self (demonstration-conditioned teacher, SDFT)
  • Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language ModelsarXiv 2026-01§5.3Symmetric CodeQwen3-4B / Qwen3-8B → Self (reasoning distillation)

§5.3.1

  • VISD: Enhancing Video Reasoning via Structured Self-DistillationarXiv 2026-05§5.3.1OtherVISD: video-aware judge decomposes quality (correctness/grounding/consistency) as structured privileged info; direction–magnitude decoupling for stable RL+SD integration; VideoLLM → Self

§5.3.2

  • Training with Harnesses: On-Policy Harness Self-Distillation for Complex ReasoningarXiv 2026-05§5.3.2FKL CodeQwen3-8B → Self; harness-augmented model (draft-verify / plan-solve) as teacher; +10.83% over OPSD on HMMT25 (PKU)

§6 Training Efficiency and Stabilization

§6 — Efficiency, Stability, and Compute

  • Escaping the KL Agreement Trap in On-Policy DistillationarXiv 2026-06§6OtherTeacher → Student; KAT: online rollout truncation at KL agreement trap regions (degraded prefixes teacher locally accepts) restores useful supervision and improves training efficiency
  • Trajectory-Refined DistillationarXiv 2026-06§6OtherTeacher → Student; TRD: trajectory-level teacher correction of prefix-failure fragmented gradients in OPD
  • Rethinking Continual Experience Internalization for Self-Evolving LLM AgentsarXiv 2026-06§6RKLQwen3-4B-Instruct → Self; Compares on-policy vs off-policy continual experience internalization; principle-level granularity + step-wise injection stabilize multi-iteration self-evolution.
  • Physics-Guided Policy Optimization with Self-DistillationarXiv 2026-06§6RKLQwen3-8B → Self; Physics-guided self-distillation: information-modulated step-size multiplier reweights gradients by mutual information between student predictions and feedback-conditioned teacher.
  • When Should the Teacher Move? Temporal Coupling and Stability in Self On-Policy DistillationarXiv 2026-06§6KL+RLQwen3-8B → Self; Studies when the self-teacher should refresh in self-OPD; introduces isolation gate (minimum freeze) + reward-ratchet gate to prevent unstable bootstrapping.
  • Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy DistillationarXiv 2026-06§6OtherQwen3-4B-Non-Thinking → Qwen3-30B-A3B-Instruct; FiRe-OPD: trajectory filtering by teacher log-prob + soft token reweighting; PPO-clipped weighted loss for OPD
  • SafeSteer: Localized On-Policy Distillation for Efficient Safety AlignmentarXiv 2026-06§6RKLQwen3-4B-Instruct → Self; Localized on-policy distillation confined to safety tokens via activation steering teacher
  • Are Full Rollouts Necessary for On-Policy Distillation?arXiv 2026-05§6RKLJustRL-R1-1.5B → R1-Distill-1.5B; Horizon-control strategies (POPD, TOPD) improve OPD efficiency by truncating rollouts
  • Trust-Region Behavior Blending for On-Policy DistillationarXiv 2026-05§6RKLQwen3-1.7B-Base / Qwen3-0.6B-Base → Qwen3-8B / Qwen3-4B; Trust-region warmup curriculum: behavior policy under student-centered KL constraint stabilizes early-stage OPD; standard reverse-KL distill loss unchanged
  • Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain PreservationarXiv 2026-05§6FKLQwen3-8B → Qwen3-4B; Dual teacher conflict-aware distillation with 3+1 alternating schedule for domain preservation
  • Less is More: Early Stopping Rollout for On-Policy DistillationarXiv 2026-05§6RKLQwen3-32B → 8B; Early-stopped rollouts at 40-60% length for 2x efficiency with maintained RKL distillation quality
  • Visual-Advantage On-Policy Distillation for Vision-Language ModelsarXiv 2026-05§6FKLQwen3-VL-8B → Qwen3-VL-2B; Visual-advantage reweighting for token-level on-policy VLM distillation with reverse KL.
  • Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning DistillationarXiv 2026-05§6FKLQwen3-32B → Qwen3-4B; MOTAB: monitors student on-policy trajectories via adaptive entropy boundary; backtracks to safe state for teacher correction to mitigate dual exposure biases in reasoning distillation
  • f-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware ControlarXiv 2026-05§6OtherQwen2.5-Math-72B → Qwen2.5-Math-7B / Qwen3-Coder-30B-A3B → Qwen3-8B; freshness-aware async OPD; Freshness-aware control for async OPD: sample-level staleness scoring + adaptive buffer refresh + rollout-anchored KL
  • DeltaPrompts: Escaping the Zero-Delta Trap in Multimodal DistillationarXiv 2026-05§6RKLQwen3-VL-8B-Thinking → Qwen3-VL-235B-Thinking; answer-divergence-guided prompt synthesis for OPD; 15% relative gain; 200k high-divergence prompts (NVIDIA)
  • Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM ReasoningarXiv 2026-05§6KL+RLQwen3-4B/8B → Self; entropy-guided confidence gate + causal-lookahead variant; advances accuracy-length frontier
  • GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-DistillationarXiv 2026-05§6OtherQwen3-4B/8B → Self; on-policy student vs GT-conditioned teacher divergence for adaptive segment boundaries; up to +20% over GRPO
  • Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy DistillationarXiv 2026-05§6RKLQwen3-8B → Self; module-allocation + update-direction perspectives: OPD identifies critical reasoning modules early
  • TRACE: Distilling Where It Matters via Token-Routed Self On-Policy AlignmentarXiv 2026-05§6KL+RLQwen3-8B → Self; token-routed self-OPD: FKL on key spans + optional RKL on error spans + GRPO elsewhere; +2.76pp (NJU/Alibaba)
  • Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon ReasoningarXiv 2026-05§6KL+RLR1-Distill-Qwen-1.5B / Qwen3-1.7B → R1-Distill-Qwen-7B / Qwen3-4B; top-k overlap drift detection for adaptive rollout truncation; 37-68% speedup
  • Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective PackingarXiv 2026-05§6FKLNPD: async generation + Δ-IFD filtering for 8.1× speedup; openPangu-Embedded-1B → 68.73% SOTA
  • Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective RecipearXiv 2026-05§6KL+RL CodeQwen3-1.7B/4B → Multi-teacher; Uni-OPD: student exploration (difficulty+correctness-aware) + teacher reliability
  • Co-Evolving Policy DistillationarXiv 2026-04§6KL+RLQwen3-VL-4B (image / text / video branches) → Qwen3-VL-4B (mutual peer); CoPD: bidirectional parallel RLVR branches + interleaved mutual OPD co-evolution
  • TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous AgentsarXiv 2026-04§6FKL CodeQwen2.5-0.5B/1.5B/3B/7B / Qwen3-0.6B/1.7B/4B → Qwen2.5-7B-GRPO / Qwen3-30B-A3B-Instruct; TCOD: temporal curriculum for autonomous agent OPD
  • TIP: Token Importance in On-Policy DistillationarXiv 2026-04§6RKL CodeQwen3-4B / Llama-3.1-8B / Qwen2.5-1.5B → Qwen3-8B / Llama-70B / Qwen2.5-14B
  • Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy DistillationarXiv 2026-04§6RKLQwen3-4B-Base / Qwen3-8B-Base → Qwen3-8B / Qwen3-32B / QwQ-32B; Lightning-OPD: offline precomputed teacher logprobs eliminate live teacher server; theoretical equivalence to online OPD under teacher consistency
  • SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive WeightingarXiv 2026-04§6FKL CodeDeepSeek-R1-Distill-Qwen-1.5B / Qwen3-1.7B → SkyWork-OR1-Math-7B / Qwen3-8B
  • Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language ModelsarXiv 2026-04§6RKLQwen2.5-Math-1.5B/7B → DeepSeek-R1-Distill-7B / OpenThinker3-7B
  • PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student CompetencearXiv 2026-03§6SymmetricQwen3-8B → Qwen3-14B; Qwen2.5-Math-7B-Instruct → Self
  • Fast and Effective On-policy Distillation from Reasoning PrefixesarXiv 2026-02§6RKLQwen3-1.7B / Qwen3-8B → Qwen3-8B (teacher)
  • SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMsarXiv 2025-10§6FKLQwen2-1.5B / Gemma-2-2B / Danube2-1.8B → Qwen2-7B / Gemma-2-9B / Mistral-7B
  • AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive SwitchingarXiv 2025-10§6FKLQwen2.5-0.5B / Llama-3.1-1B / Gemma-2B → Qwen2.5-3B / Llama-3.1-3B / Gemma-7B
  • Lion: Adversarial Distillation of Proprietary Large Language ModelsarXiv 2023-05§6OtherLion-7B / Lion-13B (LLaMA) → ChatGPT (gpt-3.5-turbo, black-box API); Adversarial black-box distillation: imitation-discrimination-generation loop iteratively identifies hard instructions via student-teacher gap; early-era black-box OPD canonical reference (HoF-tier)

§7 Understanding OPD

§7.1 — Success Conditions and Empirical Analyses

  • Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy DistillationarXiv 2026-06§7.1OtherAnalysis: OPD parameter updates are coordinate-sparse, FFN-heavy, and oriented off-principal directions; dense supervision produces sparse, structured weight changes
  • On the Geometry of On-Policy DistillationarXiv 2026-06§7.1OtherGeometry of OPD: subspace locking in parameter-space trajectories; OPD avoids principal directions vs SFT/RL
  • Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and WhyarXiv 2026-05§7.1OtherQwen3-1.7B/8B → Self; training-free per-token diagnostic; ideal per-node gradient + gradient alignment score (Apple)
  • OPSD Compresses What RLVR Teaches: A Post-RL Compaction Stage for Reasoning ModelsarXiv 2026-05§7.1OtherQwen3-8B / R1-Distill-7B / AceReason-7B → Self (OPSD as compression-not-correction in thinking-enabled reasoning)
  • Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and RecipearXiv 2026-04§7.1Other CodeQwen3-1.7B → DeepSeek-R1-Distill-7B / Qwen3-4B etc.

§7.2 — Failure Modes and Diagnostics

  • Prefix Teach, Suffix Fade: Local Teachability Collapse in Strong-to-Weak On-Policy DistillationarXiv 2026-05§7.2OtherQwen3-1.7B/4B/8B → Qwen3-14B; BIC change-point release rule; dense OPD supervision degrades in suffix when teacher margin vanishes
  • The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and FixesarXiv 2026-05§7.2OtherQwen3-8B → Self; distribution mismatch, biased TopK RKL gradients, PI aggregation collapse (UIUC)
  • Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy DistillationarXiv 2026-05§7.2OtherQwen3-8B → Self; persistent high-loss tokens (~18%) that resist teacher correction; structural residuals
  • The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured OutputsarXiv 2026-05§7.2RKLQwen3-1.7B/8B → Qwen3-8B; closed-form clip-safety threshold for reward extrapolation in structured JSON (NTU)
  • The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy DistillationarXiv 2026-04§7.2KL+RL CodeQwen3-0.6B–32B → Self (CaOPD: miscalibration scaling law + calibration-aware OPD)
  • Revisiting On-Policy Distillation: Empirical Failure Modes and Simple FixesarXiv 2026-03§7.2OtherQwen2.5-7B-Instruct → OpenThinker3-7B / GiGPO-Qwen2.5-7B
  • Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?arXiv 2026-03§7.2Other CodeQwen3-8B / DeepSeek-Distill-7B / Olmo3-7B → Self

§7.3 — Unified Theoretical Perspectives

  • A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and AlgorithmsarXiv 2025-12§7.3OtherTheoretical (no specific models)

§8 Applications, Systems, and Emerging Domains

§8.1 — Industrial Deployment

  • Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-DistillationarXiv 2026-05§8.1FKL CodeR1-Distill-1.5B / Qwen3-0.6B–8B → Self (OPSA); frozen teacher = same model + safety privileged context; per-token KL on student rollouts; teacher flip rate for context search
  • KAT-Coder-V2 Technical ReportarXiv 2026-03§8.1KL+RLKAT-Coder-V2 → 5 domain specialists; proprietary multi-expert agentic pipeline
  • Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy DistillationarXiv 2026-03§8.1KL+RLNemotron-Cascade-2-30B-A3B → Self (multi-ckpt MOPD)
  • ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget ReasoningarXiv 2026-01§8.1RKLDeepSeek-Distill-Qwen-1.5B / Qwen3-4B-Thinking / Nemotron-7B → Self (multi-teacher OPD fusion)
  • MiMo-V2-Flash Technical ReportarXiv 2026-01§8.1KL+RL CodeMiMo-V2-Flash 309B MoE → Self (multi-teacher MOPD)
  • Qwen3 Technical ReportarXiv 2025-05§8.1RKL CodeQwen3 series → Qwen3 (larger, on-policy logit KD)

§8.2 — Emerging Domains

  • Reward-Weighted On-Policy Distillation with an Open Property-Equivalence Verifier for NL-to-SVA GenerationarXiv 2026-05§8.2FKLQwen2.5-Coder-7B → CodeV-SVA-14B; verifier-reward-weighted FKL on student rollouts; new SOTA on NL2SVA
  • Revisiting DAgger in the Era of LLM-AgentsarXiv 2026-05§8.2OtherQwen3-4B-Instruct-2507 / Qwen3-8B → Qwen3-Coder-30B-A3B-Instruct (DAgger); turn-level student-teacher interpolation for SWE agents; +3.9pp on SWE-bench Verified
  • ProteinOPD: Towards Effective and Efficient Preference Alignment for Protein DesignarXiv 2026-05§8.2Symmetric CodeProtein PLM → Multi-teacher; geometric consensus of weighted preference-specific teachers; 8× faster than RL (THU/IDEA)
  • SOD: Step-wise On-policy Distillation for Small Language Model AgentsarXiv 2026-05§8.2KL+RL CodeQwen3-0.6B/1.7B → Qwen3-4B; step-wise OPD for agentic tasks; progressive trajectory distillation
  • LiteGUI: Distilling Compact GUI Agents with Reinforcement LearningarXiv 2026-05§8.2KL+RLQwen3-VL-32B → 2B–3B GUI agents; guided OPD + dual-level GRPO; ScreenSpot-Pro, OS-World, Lite-Bench
  • HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search AgentsarXiv 2026-05§8.2KL+RL CodeExternal teacher → HyperEyes-30B (Qwen3-VL-30B); micro-level OPD provides dense token-level supervision on failed rollouts
  • Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM AgentsarXiv 2026-04§8.2KL+RLQwen3-4B-Instruct → Self (Skill-SDL)
  • On-Policy Distillation of Language Models for Autonomous Vehicle Motion PlanningarXiv 2026-04§8.2SymmetricQwen3-1.7B → Qwen3-8B
  • HY-Embodied-0.5: Embodied Foundation Models for Real-World AgentsarXiv 2026-04§8.2FKL CodeHY-Embodied-0.5 MoE-A32B (large) → HY-Embodied-0.5 MoT-2B (small)
  • DP-OPD: Differentially Private On-Policy Distillation for Language ModelsarXiv 2026-04§8.2SymmetricDistilGPT-2 82M → GPT-2 Large 774M (+ DP-SGD)
  • VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy DistillationarXiv 2026-03§8.2RKLOpenVLA-OFT → SimpleVLA-RL (frozen expert teacher); dense token-level RKL on student-generated trajectories for robot manipulation (LIBERO / RoboTwin2.0)
  • X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMsarXiv 2026-03§8.2RKLQwen3-Omni-A3B → Qwen3-A3B-Instruct (text teacher)
  • OpenClaw-RL: Train Any Agent Simply by TalkingarXiv 2026-03§8.2KL+RL CodeHindsight-guided OPD for agentic scenarios; task completion reward → policy distillation
  • Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy DistillationarXiv 2026-02§8.2RKLQwen3-VL-8B → Qwen3-VL-32B (Video-OPD)
  • CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal DistillationarXiv 2026-01§8.2KL+RLQwen2-Audio-7B / Step-Audio2-mini → Self (cross-modal)
  • VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy DistillationarXiv 2025-10§8.2KL+RLQwen2.5-VL-3B → Qwen3-8B (text reasoning teacher)

§8.3 — System-Level Integration

  • Test-Time SpeculationarXiv 2026-05§8.3FKLQwen3-8B / Llama-3.1-8B → Qwen3-32B / Llama-3.1-70B; online OPD at inference-time; up to 72% acceptance-length gain
  • Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved SamplingarXiv 2024-10§8.3FKLGemma-2B-IT / Qwen2-0.5B-IT → Gemma-7B-IT / Qwen2-7B-IT
  • DistillSpec: Improving Speculative Decoding via Knowledge DistillationarXiv 2023-10§8.3SymmetricT5-Small → T5-XL (on-policy KD for speculative decoding)

Citation

@article{song2026opdsurvey,
  title  = {A Survey of On-Policy Distillation for Large Language Models},
  author = {Mingyang Song and Mao Zheng},
  journal= {arXiv preprint arXiv:2604.00626},
  year   = {2026}
}