When Does Muon Help Agentic Reinforcement Learning?
arXiv:2607.16169v2 Announce Type: replace-cross Abstract: Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training…
arXiv:2607.16169v2 Announce Type: replace-cross Abstract: Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training…
arXiv:2607.16195v1 Announce Type: new Abstract: We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels…
arXiv:2607.15280v1 Announce Type: new Abstract: Sequential diagnosis requires balancing diagnostic accuracy against resource costs through iterative information gathering. Existing Large…
arXiv:2510.12229v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have been shown to internalize human-like biases during finetuning, yet the…
arXiv:2604.06742v2 Announce Type: replace-cross Abstract: The evolution of Large Language Models (LLMs) has catalyzed a paradigm shift towards intent-driven software…
arXiv:2512.03054v3 Announce Type: replace-cross Abstract: Federated Learning (FL) holds the potential to advance equality in health by enabling diverse institutions…
arXiv:2205.04599v2 Announce Type: replace-cross Abstract: Explainable Artificial Intelligence (XAI) is essential for trustworthy AI in healthcare, yet many existing methods…
arXiv:2607.08423v2 Announce Type: replace Abstract: The rapid integration of Large Vision-Language Models (VLMs) into critical infrastructure promises to revolutionize personalized…
arXiv:2607.15095v2 Announce Type: replace-cross Abstract: The formation of political coalitions is a complex negotiation driven by both concrete policy objectives…
arXiv:2607.16130v1 Announce Type: cross Abstract: AI governance increasingly requires judgments about whether an AI system remains adequately trustworthy over time,…