DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory
arXiv:2507.07855v3 Announce Type: replace-cross Abstract: Normative theories allow one to elicit key parts of a ML algorithm from first principles,…
DIVER-1 : Deep Integration of Vast Electrophysiological Recordings at Scale
arXiv:2512.19097v2 Announce Type: replace-cross Abstract: Unifying the vast heterogeneity of brain signals into a single foundation model is a longstanding…
Attributing and situating knowledge cannot be left to language models
Nature Machine Intelligence, Published online: 06 February 2026; doi:10.1038/s42256-026-01193-0 Attributing and situating knowledge cannot be left to language models
Authorization of prognostic AI medical devices
Nature Machine Intelligence, Published online: 06 February 2026; doi:10.1038/s42256-025-01171-y Less than 2% of artificial intelligence devices authorized by the US…
Knowledge Model Prompting Increases LLM Performance on Planning Tasks
arXiv:2602.03900v1 Announce Type: new Abstract: Large Language Models (LLM) can struggle with reasoning ability and planning tasks. Many prompting techniques…
Not All Negative Samples Are Equal: LLMs Learn Better from Plausible Reasoning
arXiv:2602.03516v2 Announce Type: replace-cross Abstract: Learning from negative samples holds great promise for improving Large Language Model (LLM) reasoning capability,…
SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization
arXiv:2602.04811v1 Announce Type: cross Abstract: True self-evolution requires agents to act as lifelong learners that internalize novel experiences to solve…
Beyond Rewards in Reinforcement Learning for Cyber Defence
arXiv:2602.04809v1 Announce Type: cross Abstract: Recent years have seen an explosion of interest in autonomous cyber defence agents trained to…
CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
arXiv:2602.03048v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a key approach for enhancing LLM…
CreditAudit: 2D Auditing for LLM Evaluation and Selection
arXiv:2602.02515v1 Announce Type: new Abstract: Leaderboard scores on public benchmarks have been steadily rising and converging, with many frontier language…
