Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning
arXiv:2605.08061v1 Announce Type: new Abstract: We argue that decomposing reward into weighted, verifiable criteria and using an LLM judge to…
arXiv:2605.08061v1 Announce Type: new Abstract: We argue that decomposing reward into weighted, verifiable criteria and using an LLM judge to…
arXiv:2605.06755v1 Announce Type: cross Abstract: Reinforcement learning is widely used to improve the reasoning ability of large language models, especially…
arXiv:2605.06729v1 Announce Type: cross Abstract: We present the E$Delta$-MHC-Geo Transformer, a novel architecture that unifies Manifold-Constrained Hyper-Connections (mHC), Deep Delta…
arXiv:2605.06671v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated strong potential for many mathematical problems. However, their performance…
arXiv:2605.07671v1 Announce Type: cross Abstract: Eliciting truthful reports from autonomous agents is a core problem in scalable AI oversight: a…
arXiv:2605.07655v1 Announce Type: cross Abstract: Searching a multi-biometric database of a billion records for a country-level identity system requires pushing…
arXiv:2605.06387v2 Announce Type: replace-cross Abstract: On-policy distillation (OPD) trains a student on its own trajectories with token-level teacher feedback and…
arXiv:2605.05329v1 Announce Type: new Abstract: Safety policies define what constitutes safe and unsafe AI outputs, guiding data annotation and model…
arXiv:2605.05097v2 Announce Type: replace-cross Abstract: LLMs are trained once, then deployed into a world that never stops changing. External memory…