UI-UG: A Unified MLLM for UI Understanding and Generation
arXiv:2509.24361v2 Announce Type: replace-cross Abstract: Although Multimodal Large Language Models (MLLMs) have been widely applied across domains, they are still…
arXiv:2509.24361v2 Announce Type: replace-cross Abstract: Although Multimodal Large Language Models (MLLMs) have been widely applied across domains, they are still…
arXiv:2509.25927v1 Announce Type: cross Abstract: Deep neural networks remain vulnerable to adversarial examples despite advances in architectures and training paradigms.…
arXiv:2509.25933v1 Announce Type: cross Abstract: Differentiable Logic Gate Networks (DLGNs) are a very fast and energy-efficient alternative to conventional feed-forward…
arXiv:2509.25077v2 Announce Type: replace-cross Abstract: Monocular Depth Estimation (MDE) is a foundational task for computer vision. Traditional methods are limited…
arXiv:2509.25229v1 Announce Type: new Abstract: We introduce Blueprint-Bench, a benchmark designed to evaluate spatial reasoning capabilities in AI models through…
arXiv:2509.22199v2 Announce Type: replace-cross Abstract: Vision Language Action (VLA) models derive their generalization capability from diverse training data, yet collecting…
arXiv:2509.24368v1 Announce Type: cross Abstract: We introduce the first watermark tailored for diffusion language models (DLMs), an emergent LLM paradigm…
arXiv:2509.24369v1 Announce Type: cross Abstract: Street view imagery has become an essential source for geospatial data collection and urban analytics,…
arXiv:2509.22445v2 Announce Type: replace-cross Abstract: The Minimum Description Length (MDL) principle offers a formal framework for applying Occam’s razor in…
arXiv:2509.22746v1 Announce Type: new Abstract: Current visual reasoning methods mainly focus on exploring specific reasoning modes. Although improvements can be…