Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset
arXiv:2608.06165v2 Announce Type: replace-cross Abstract: Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music…
HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses
arXiv:2608.06984v1 Announce Type: cross Abstract: Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills,…
PHASE-Tree: Modeling Character-State Evolution in Long-Horizon Role-Playing Dialogue
arXiv:2608.06975v1 Announce Type: cross Abstract: Long-horizon role-playing demands that characters remain recognizable as they evolve with the narrative. Yet existing…
HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection
arXiv:2608.05771v2 Announce Type: replace-cross Abstract: Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation, yet detector performance…
Agentic Nesting: A New Methodology for Existing Enterprise Application Integration and Services
arXiv:2608.05159v1 Announce Type: new Abstract: Enterprise operations extensively rely on multiple heterogeneous business systems and information applications, which also result…
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
arXiv:2608.05131v2 Announce Type: replace-cross Abstract: On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal…
Spectral Aliasing Pretext: A novel task for Self-Supervised fault diagnosis in rotating machinery
arXiv:2608.05705v1 Announce Type: cross Abstract: Deep learning is a new way for machinery fault diagnosis but requires extensive labeled data,…
Answer First, Reason Later: Commitment Order in Diffusion LLMs
arXiv:2608.05687v1 Announce Type: cross Abstract: Masked diffusion language models (dLLMs) can commit tokens in any order — a freedom marketed…
