Inclusive Easy-to-Read Generation for Individuals with Cognitive Impairments
arXiv:2510.00691v1 Announce Type: cross Abstract: Ensuring accessibility for individuals with cognitive impairments is essential for autonomy, self-determination, and full citizenship.…
Vector-Valued Reproducing Kernel Banach Spaces for Neural Networks and Operators
arXiv:2509.26371v2 Announce Type: replace-cross Abstract: Recently, there has been growing interest in characterizing the function spaces underlying neural networks. While…
Blueprint-Bench: Comparing spatial intelligence of LLMs, agents and image models
arXiv:2509.25229v1 Announce Type: new Abstract: We introduce Blueprint-Bench, a benchmark designed to evaluate spatial reasoning capabilities in AI models through…
BRIDGE — Building Reinforcement-Learning Depth-to-Image Data Generation Engine for Monocular Depth Estimation
arXiv:2509.25077v2 Announce Type: replace-cross Abstract: Monocular Depth Estimation (MDE) is a foundational task for computer vision. Traditional methods are limited…
From MNIST to ImageNet: Understanding the Scalability Boundaries of Differentiable Logic Gate Networks
arXiv:2509.25933v1 Announce Type: cross Abstract: Differentiable Logic Gate Networks (DLGNs) are a very fast and energy-efficient alternative to conventional feed-forward…
The Impact of Scaling Training Data on Adversarial Robustness
arXiv:2509.25927v1 Announce Type: cross Abstract: Deep neural networks remain vulnerable to adversarial examples despite advances in architectures and training paradigms.…
UI-UG: A Unified MLLM for UI Understanding and Generation
arXiv:2509.24361v2 Announce Type: replace-cross Abstract: Although Multimodal Large Language Models (MLLMs) have been widely applied across domains, they are still…
Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
arXiv:2509.22746v1 Announce Type: new Abstract: Current visual reasoning methods mainly focus on exploring specific reasoning modes. Although improvements can be…
Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers
arXiv:2509.22445v2 Announce Type: replace-cross Abstract: The Minimum Description Length (MDL) principle offers a formal framework for applying Occam’s razor in…
From Satellite to Street: A Hybrid Framework Integrating Stable Diffusion and PanoGAN for Consistent Cross-View Synthesis
arXiv:2509.24369v1 Announce Type: cross Abstract: Street view imagery has become an essential source for geospatial data collection and urban analytics,…
