FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
arXiv:2601.07853v1 Announce Type: cross Abstract: Financial agents powered by large language models (LLMs) are increasingly deployed for investment analysis, risk…
CLewR: Curriculum Learning with Restarts for Machine Translation Preference Learning
arXiv:2601.05858v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated competitive performance in zero-shot multilingual machine translation (MT). Some…
Collective Communication for 100k+ GPUs
arXiv:2510.20171v4 Announce Type: replace-cross Abstract: The increasing scale of large language models (LLMs) necessitates highly efficient collective communication frameworks, particularly…
Simulation-Free PSRO: Removing Game Simulation from Policy Space Response Oracles
arXiv:2601.05279v1 Announce Type: cross Abstract: Policy Space Response Oracles (PSRO) combines game-theoretic equilibrium computation with learning and is effective in…
IIB-LPO: Latent Policy Optimization via Iterative Information Bottleneck
arXiv:2601.05870v1 Announce Type: cross Abstract: Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) for Large Language Model (LLM) reasoning…
Memorization in Large Language Models in Medicine: Prevalence, Characteristics, and Implications
arXiv:2509.08604v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated significant potential in medicine, with many studies adapting them…
