Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
arXiv:2605.06387v2 Announce Type: replace-cross Abstract: On-policy distillation (OPD) trains a student on its own trajectories with token-level teacher feedback and…
