Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
arXiv:2604.06742v2 Announce Type: replace-cross Abstract: The evolution of Large Language Models (LLMs) has catalyzed a paradigm shift towards intent-driven software…
arXiv:2604.06742v2 Announce Type: replace-cross Abstract: The evolution of Large Language Models (LLMs) has catalyzed a paradigm shift towards intent-driven software…
arXiv:2512.03054v3 Announce Type: replace-cross Abstract: Federated Learning (FL) holds the potential to advance equality in health by enabling diverse institutions…
arXiv:2205.04599v2 Announce Type: replace-cross Abstract: Explainable Artificial Intelligence (XAI) is essential for trustworthy AI in healthcare, yet many existing methods…
arXiv:2607.08423v2 Announce Type: replace Abstract: The rapid integration of Large Vision-Language Models (VLMs) into critical infrastructure promises to revolutionize personalized…
arXiv:2607.15095v2 Announce Type: replace-cross Abstract: The formation of political coalitions is a complex negotiation driven by both concrete policy objectives…
arXiv:2607.16130v1 Announce Type: cross Abstract: AI governance increasingly requires judgments about whether an AI system remains adequately trustworthy over time,…
arXiv:2607.16131v1 Announce Type: cross Abstract: Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence…
arXiv:2607.15263v2 Announce Type: replace-cross Abstract: Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit…
arXiv:2607.13162v2 Announce Type: replace-cross Abstract: What a language model will and will not do is largely set during post-training, but…
arXiv:2607.14543v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used as the reasoning backbone of embodied agents, enabling robots…