Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection
arXiv:2608.20169v2 Announce Type: replace-cross Abstract: We present a novel approach to efficient LLM harness optimization through adaptive validation task selection.…
What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs
arXiv:2608.00013v2 Announce Type: replace-cross Abstract: Choosing the right large language model (LLM) backbone is the most consequential decision when building…
SA-RSQ: A Versatile Sparse Representation Framework for Multi-modal Recommender Systems
arXiv:2608.22979v1 Announce Type: new Abstract: Deploying high-dimensional multimodal features in industrial recommender systems incurs substantial storage and latency overhead. Hard…
STAR-OPD: Structured Aspect-Cascade-Aware On-Policy Reward Distillation for ABSA Quadruple Extraction
arXiv:2608.20831v1 Announce Type: cross Abstract: Aspect-based sentiment analysis (ABSA) quadruple extraction requires jointly predicting target, aspect, opinion, and sentiment over…
Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts
arXiv:2607.18970v3 Announce Type: replace-cross Abstract: Agent Skills have become persistent behavioral artifacts across independent AI agent systems. They combine natural-language…
CoST: Semantic-Aware Urban Understanding via Spatial-Temporal Alignment
arXiv:2608.21041v1 Announce Type: cross Abstract: Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and…
Beyond End-to-End Success: Diagnosing Failures in Long-Horizon Security LLM Agents
arXiv:2608.20563v1 Announce Type: cross Abstract: Long-horizon security LLM agents must carry information and decisions across many dependent interactions, where later…
Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability
arXiv:2608.11238v2 Announce Type: replace Abstract: Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence,…
SDAD: Spec-Driven Agentic Development for the AI-Native SDLC
arXiv:2608.20341v1 Announce Type: new Abstract: Frontier coding agents backed by large language models with context windows from hundreds of thousands…
