Multi-Sourced, Multi-Agent Evidence Retrieval for Fact-Checking
arXiv:2603.00267v1 Announce Type: new Abstract: Misinformation spreading over the Internet poses a significant threat to both societies and individuals, necessitating…
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
arXiv:2602.24009v2 Announce Type: replace-cross Abstract: Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making robustness estimates stale…
From Variance to Invariance: Qualitative Content Analysis for Narrative Graph Annotation
arXiv:2603.01930v1 Announce Type: cross Abstract: Narratives in news discourse play a critical role in shaping public understanding of economic events,…
Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
arXiv:2603.01919v1 Announce Type: cross Abstract: Access to frontier large language models (LLMs), such as GPT-5 and Gemini-2.5, is often hindered…
Pseudo Contrastive Learning for Diagram Comprehension in Multimodal Models
arXiv:2602.23589v2 Announce Type: replace-cross Abstract: Recent multimodal models such as Contrastive Language-Image Pre-training (CLIP) have shown remarkable ability to align…
HumanMCP: A Human-Like Query Dataset for Evaluating MCP Tool Retrieval Performance
arXiv:2602.23367v1 Announce Type: new Abstract: Model Context Protocol (MCP) servers contain a collection of thousands of open-source standardized tools, linking…
Conformalized Neural Networks for Federated Uncertainty Quantification under Dual Heterogeneity
arXiv:2602.23296v2 Announce Type: replace-cross Abstract: Federated learning (FL) faces challenges in uncertainty quantification (UQ). Without reliable UQ, FL systems risk…
Offline-to-Online Multi-Agent Reinforcement Learning with Offline Value Function Memory and Sequential Exploration
arXiv:2410.19450v2 Announce Type: replace Abstract: Offline-to-Online Reinforcement Learning has emerged as a powerful paradigm, leveraging offline data for initialization and…
