Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning
arXiv:2606.10968v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become standard for improving LLM reasoning. However, existing…
arXiv:2606.10968v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become standard for improving LLM reasoning. However, existing…
arXiv:2606.12169v1 Announce Type: cross Abstract: High-stakes clinical use of large vision-language models (LVLMs) requires reasoning that is grounded in visual…
arXiv:2606.11207v1 Announce Type: new Abstract: We present SemantiClean, a modular framework for extracting structured semantic signals from e-commerce session data…
arXiv:2606.11272v1 Announce Type: cross Abstract: Federated Learning (FL) enables collaborative and privacy-preserving model training across distributed clients, but most existing…
Nature Machine Intelligence, Published online: 11 June 2026; doi:10.1038/s42256-026-01250-8 Xiong et al. introduce ConfSeq, a molecular conformation description language that…
arXiv:2606.09377v2 Announce Type: replace-cross Abstract: Formal neural network verification — proving that a network satisfies safety properties for *all* inputs…
arXiv:2604.15414v2 Announce Type: replace-cross Abstract: Continual reinforcement learning must balance retention with adaptation, yet many methods still rely on emph{single-model…
arXiv:2606.11140v1 Announce Type: cross Abstract: Data assimilation (DA) in subsurface flow entails calibrating model parameters to match observed data, typically…
arXiv:2511.01927v2 Announce Type: replace-cross Abstract: Solving large-scale Generalized Eigenvalue Problems (GEPs) is a fundamental yet computationally prohibitive task in science…