LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents
arXiv:2607.27690v2 Announce Type: replace-cross Abstract: We introduce LabEvolver, a training-free framework that equips safe and grounded wet-lab agents with episodic…
Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models
arXiv:2607.26119v1 Announce Type: new Abstract: Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their…
Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition
arXiv:2607.25961v2 Announce Type: replace-cross Abstract: Ambivalence and hesitancy (A/H) are conflicting affective states that precede the delay or abandonment of…
BayesAME: Bayesian Active Model Evaluation
arXiv:2607.27023v1 Announce Type: cross Abstract: Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need…
SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception
arXiv:2607.26985v1 Announce Type: cross Abstract: Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow wall-clock…
Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction
arXiv:2607.25718v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks.…
