The emergence of reasoning models such as OpenAI o1/o3 and DeepSeek-R1 sparked widespread efforts to distill dense reasoning traces into smaller 7B and 14B student models, unlocking striking performance leaps on math and coding benchmarks. However, a fundamental mechanistic question has lingered: does On-Policy Distillation (OPD) inject new parametric factual knowledge into the student, or does it merely impart compositional reasoning skills? Researchers from Hong Kong University of Science and Technology (HKUST) address this mystery in a landmark study (arXiv:2610.09639). Utilizing a controlled synthetic framework across four models from three distinct architectures, the authors independently isolate factual memory from multi-step compositional ability. The experiments establish that reverse-KL OPD reliably transfers compositional problem-solving skills across unseen task structures, but transfers minimal factual knowledge. Replacing reverse-KL with forward-KL restores factual memorization, while student rollouts specifically refine multi-step logical execution. These findings demonstrate that on-policy distillation does not expand a model's knowledge boundary, but rather teaches it to mobilize and organize facts it already possesses.

Key Takeaways

  • ✓HKUST demonstrates through controlled synthetic frameworks that on-policy distillation (OPD) transfers compositional skills but minimal factual knowledge
  • ✓Proves reverse-KL optimizes multi-step logical execution while forward-KL is required for parametric factual memorization
  • ✓Establishes a principled two-stage recipe for reasoning agents: forward-KL for factual ingestion, followed by reverse-KL for reasoning activation
🧭

Turn your technical choice into a development budget

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Background and the Problem

The rise of frontier reasoning architectures (such as DeepSeek-R1 and OpenAI o-series) catalyzed an industry-wide push toward reasoning distillation. By collecting long Chain-of-Thought (Long-CoT) rollouts and applying On-Policy Distillation (OPD), practitioners transfer advanced reasoning capabilities into compact 7B and 14B models, producing dramatic performance leaps on GSM8K and coding leaderboards. However, a foundational ambiguity has persisted: does distillation actually infuse new parametric world facts into student models, or does it exclusively teach compositional execution heuristics? In domain-specific applications lacking foundational knowledge, does reasoning distillation simply produce fluent hallucinations?

Architecture and How It Works

To disentangle knowledge acquisition from reasoning skill, researchers from Hong Kong University of Science and Technology establish a synthetic capability control platform (arXiv:2610.09639):

  1. Orthogonal Decoupling of Facts vs. Compositional Skills: Formulates synthetic knowledge graphs and multi-hop reasoning chains where student baselines are systematically benchmarked against teachers possessing either isolated new facts, isolated multi-hop skills, or both.
  2. Divergence Formulation Ablation: Evaluates four models across three architectural families under reverse-KL OPD (standard in RLVR and on-policy distillation) versus forward-KL objectives (standard in SFT).
  3. Dual Validation on Empirical Benchmarks: Evaluates performance across recent factual QA datasets and competitive mathematics benchmarks (AIME/AMC) to track factual retention and compositional accuracy.

Benchmarks and Measured Results

Empirical investigations reveal a sharp mechanistic asymmetry:

  1. Robust Skill Transfer with Zero Factual Acquisition: Under reverse-KL OPD, students faithfully absorb the teacher's compositional reasoning patterns across novel structural topologies. However, performance on factual recall tasks evaluating teacher-exclusive knowledge remains near zero.
  2. Reverse vs. Forward KL Divergence Mechanics: Substituting reverse-KL with forward-KL restores factual memorization. Conversely, student rollouts specifically improve the planning and sequential execution of multi-step chains.
  3. Consistent Behavior on Competition Math: On mathematical benchmarks, reverse-KL OPD delivers double-digit percentage gains in multi-step problem solving without inducing measurable expansion in the model's factual memory store.

Getting Started for Developers

The HKUST findings provide clear architectural principles for post-training reasoning models. When customizing domain-specific coding agents or enterprise reasoning models, engineers must not rely on reasoning distillation to teach unfamiliar APIs or internal system facts. Instead, teams should implement a strict two-stage post-training pipeline:

  1. Knowledge Implantation Stage: Apply forward-KL (supervised fine-tuning or continued pre-training) to seed domain facts, syntax, and schema invariants into model weights.
  2. Reasoning Activation Stage: Once factual representations are established, execute reverse-KL on-policy distillation or RLVR to cultivate multi-step planning, reflection, and compositional problem-solving. This disciplined separation eliminates structurally fluent hallucinations in enterprise agents.
Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.