PreAct-Bench: Benchmarking Predictive Monitoring in LLMs
arXiv:2606.09890v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents capable of executing multi-step action…
arXiv:2606.09890v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents capable of executing multi-step action…
arXiv:2606.04409v2 Announce Type: replace-cross Abstract: Modern deep neural networks usually have large parameter scales and nonlinear hierarchical structures, and they…
arXiv:2606.08948v1 Announce Type: cross Abstract: Comprehensive estimation of dietary micronutrients from food images could improve clinical nutrition care, but training…
arXiv:2602.10172v2 Announce Type: replace-cross Abstract: Reconstructing the early universe from the evolved present-day universe is a challenging and computationally demanding…
arXiv:2605.23595v3 Announce Type: replace-cross Abstract: The rapid advancement of machine learning has led to an unprecedented expansion of model ecosystems,…
arXiv:2606.06554v2 Announce Type: replace-cross Abstract: Reliable polymer identification is essential for ensuring the quality and safety of recycled plastics, yet…
arXiv:2606.08710v1 Announce Type: cross Abstract: Modernization of legacy scientific codes is often necessary to keep up with the ever-evolving changes…
arXiv:2606.08712v1 Announce Type: cross Abstract: Purpose: Spatial transcriptomics (ST) enables gene expression measurements within the tissue context. However, these measurements…
arXiv:2606.07379v2 Announce Type: replace-cross Abstract: A growing failure mode in agent evaluation and training is that models can achieve high…
arXiv:2606.07549v1 Announce Type: new Abstract: Recent advances in Multimodal Large Language Models (MLLMs) and agent workflows have shown strong promise…