Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
arXiv:2604.13016v2 Announce Type: replace-cross Abstract: On-policy distillation (OPD) has become a core technique in the post-training of large language models,…
