Location-Aware Pretraining for Medical Difference Visual Question Answering

ByAdmin

Apr 23, 2026

THE AI TODAY

arXiv:2603.04950v2 Announce Type: replace-cross
Abstract: Differential medical VQA models compare multiple images to identify clinically meaningful changes and rely on vision encoders to capture fine-grained visual differences that reflect radiologists’ comparative diagnostic workflows. However, vision encoders trained using standard contrastive or classification objectives often fail to capture the subtle variations needed to distinguish true disease progression from acquisition-related variability. To address this limitation, we introduce a location-aware pretraining framework that incorporates automatic referring expressions (AREF), grounded captioning (GCAP), and conditional automatic referring expressions (CAREF). These tasks promote the learning of fine-grained, spatially grounded visual representations. When integrated with a language model, our approach achieves state-of-the-art performance on medical difference VQA by accurately identifying and reasoning about clinically relevant changes in chest X-ray images.

By Admin

AI RESEARCH

Location-Aware Pretraining for Medical Difference Visual Question Answering

ByAdmin

By Admin

Related Post

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

Matrix-Free Photoacoustic Image Reconstruction via Sensor-Token Self-Attention

Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation

You missed

DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model

Inferring Missing Trajectory Data with Temporal Convolutional Networks

Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks