Cross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models
arXiv:2609.20139v1 Announce Type: cross Abstract: Vision-language models (VLMs) are fragile under image corruption. We find that the wording of the…
