arxivcs.CV2026-07-14
MQAdapter: Multi-Modal Quantum Adapter for Coarse-to-Fine VLM Fine-tuning
Yumiao Zhao, Bo Jiang, Min Lu, Xiao Wang, Jin Tang
Large-scale Vision-Language Models have demonstrated impressive transfer learning capabilities across a wide range of tasks. For few-shot classification, we observe that VLMs exhibit a notable ability to filter candidate categories and thus achieve high Top-K accuracy. However, t…