BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation
arXiv:2601.06401v2 Announce Type: replace Abstract: Large language models are becoming increasingly significant in financial applications. Nevertheless, prevailing benchmarks are largely…
Technical Report on the CVPR 2026@AdvML Workshop Challenge
arXiv:2607.11560v1 Announce Type: cross Abstract: Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning.…
CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
arXiv:2512.02551v3 Announce Type: replace-cross Abstract: In this paper, we propose CUDA-L2, a system that combines large language models (LLMs) and…
LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
arXiv:2607.09322v2 Announce Type: replace Abstract: In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making. Prior…
Multi-Scale Convolution with Optimal Transport Attention Effect on Multivariate Time Series
arXiv:2607.10740v1 Announce Type: cross Abstract: The analysis of Multivariate Time Series (MTS) plays an important role in a lot of…
Lightning Fast Matching Dependency Discovery with Desbordante
arXiv:2607.10771v1 Announce Type: cross Abstract: Matching dependency is a generalization of the functional dependency concept, which allows users to apply…
A Sovereign, Open-Source Foundation Model for German and English
arXiv:2607.09424v2 Announce Type: replace-cross Abstract: We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model…
Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA
arXiv:2605.17932v2 Announce Type: replace-cross Abstract: Prompt compression reduces inference cost and context length in large language models, but prior evaluations…
The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning
arXiv:2607.11436v1 Announce Type: new Abstract: Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable…
Evaluating Retrieval-Augmented Generation vs. Long-Context Input for Clinical Reasoning over EHRs
arXiv:2508.14817v2 Announce Type: replace-cross Abstract: Objective: To evaluate whether retrieval-augmented generation (RAG) can serve as an efficient alternative to long-context…
