On the Position Bias of On-Policy Distillation
arXiv:2606.22600v2 Announce Type: replace-cross Abstract: On-Policy Distillation (OPD) improves the learning efficiency of standard reinforcement learning through dense, token-level supervision…
arXiv:2606.22600v2 Announce Type: replace-cross Abstract: On-Policy Distillation (OPD) improves the learning efficiency of standard reinforcement learning through dense, token-level supervision…
arXiv:2606.23927v1 Announce Type: new Abstract: Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making…
Nature Machine Intelligence, Published online: 23 June 2026; doi:10.1038/s42256-026-01269-x Recent breakthroughs in mathematical research show that AI is transforming the…
Nature Machine Intelligence, Published online: 23 June 2026; doi:10.1038/s42256-026-01263-3 Nassour, Berberich and colleagues present a soft robotic hand exoskeleton that…
Nature Machine Intelligence, Published online: 22 June 2026; doi:10.1038/s42256-026-01252-6 An, Luo, Zhang and colleagues present Turbo, a transformer-based reinforcement learning…