Watermarking Diffusion Language Models
arXiv:2509.24368v1 Announce Type: cross Abstract: We introduce the first watermark tailored for diffusion language models (DLMs), an emergent LLM paradigm…
MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training
arXiv:2509.22199v2 Announce Type: replace-cross Abstract: Vision Language Action (VLA) models derive their generalization capability from diverse training data, yet collecting…
HiddenBench: Assessing Collective Reasoning in Multi-Agent LLMs via Hidden Profile Tasks
arXiv:2505.11556v2 Announce Type: replace-cross Abstract: Multi-agent systems built on large language models (LLMs) promise enhanced problem-solving through distributed information integration,…
Capacity-Aware Planning and Scheduling in Budget-Constrained Multi-Agent MDPs: A Meta-RL Approach
arXiv:2410.21249v2 Announce Type: replace-cross Abstract: We study capacity- and budget-constrained multi-agent MDPs (CB-MA-MDPs), a class that captures many maintenance and…
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
arXiv:2507.22168v2 Announce Type: replace-cross Abstract: Current benchmarks for evaluating Large Language Models (LLMs) often do not exhibit enough writing style…
Mamba Integrated with Physics Principles Masters Long-term Chaotic System Forecasting
arXiv:2505.23863v2 Announce Type: replace-cross Abstract: Long-term forecasting of chaotic systems remains a fundamental challenge due to the intrinsic sensitivity to…
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
arXiv:2507.13681v2 Announce Type: replace-cross Abstract: Multi-turn dialogues are essential in many real-world applications of large language models, such as chatbots…
An Approach to Checking Correctness for Agentic Systems
arXiv:2509.20364v1 Announce Type: new Abstract: This paper presents a temporal expression language for monitoring AI agent behavior, enabling systematic error-detection…
SIM-CoT: Supervised Implicit Chain-of-Thought
arXiv:2509.20317v2 Announce Type: replace-cross Abstract: Implicit Chain-of-Thought (CoT) methods offer a token-efficient alternative to explicit CoT reasoning in Large Language…
Adoption, usability and perceived clinical value of a UK AI clinical reference platform (iatroX): a mixed-methods formative evaluation of real-world usage and a 1,223-respondent user survey
arXiv:2509.21188v1 Announce Type: cross Abstract: Clinicians face growing information overload from biomedical literature and guidelines, hindering evidence-based care. Retrieval-augmented generation…
