DOA Estimation with Lightweight Network on LLM-Aided Simulated Acoustic Scenes
arXiv:2511.08012v1 Announce Type: cross Abstract: Direction-of-Arrival (DOA) estimation is critical in spatial audio and acoustic signal processing, with wide-ranging applications…
On the generalization of language models from in-context learning and finetuning: a controlled study
arXiv:2505.00661v3 Announce Type: replace-cross Abstract: Large language models exhibit exciting capabilities, yet can show surprisingly narrow generalization from finetuning. E.g.…
Analysing Environmental Efficiency in AI for X-Ray Diagnosis
arXiv:2511.07436v1 Announce Type: new Abstract: The integration of AI tools into medical applications has aimed to improve the efficiency of…
Surgical Agent Orchestration Platform for Voice-directed Patient Data Interaction
arXiv:2511.07392v2 Announce Type: replace-cross Abstract: In da Vinci robotic surgery, surgeons’ hands and eyes are fully engaged in the procedure,…
Beyond the Pixels: VLM-based Evaluation of Identity Preservation in Reference-Guided Synthesis
arXiv:2511.08087v1 Announce Type: cross Abstract: Evaluating identity preservation in generative models remains a critical yet unresolved challenge. Existing metrics rely…
Dynamic Sparsity: Challenging Common Sparsity Assumptions for Learning World Models in Robotic Reinforcement Learning Benchmarks
arXiv:2511.08086v1 Announce Type: cross Abstract: The use of learned dynamics models, also known as world models, can improve the sample…
Hard vs. Noise: Resolving Hard-Noisy Sample Confusion in Recommender Systems via Large Language Models
arXiv:2511.07295v2 Announce Type: replace-cross Abstract: Implicit feedback, employed in training recommender systems, unavoidably confronts noise due to factors such as…
Evidence-Bound Autonomous Research (EviBound): A Governance Framework for Eliminating False Claims
arXiv:2511.05524v1 Announce Type: new Abstract: LLM-based autonomous research agents report false claims: tasks marked “complete” despite missing artifacts, contradictory metrics,…
SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models
arXiv:2511.05459v2 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) for software engineering has been limited by narrow task coverage,…
