Deniz Kizaroğlu, Unified Local-Global Prompt Learning for Few-Shot Vision-Language Adaptation via Optimal Transport
This thesis proposes a unified framework for few-shot vision-language adaptation. Addressing the limitations of holistic image matching in models like CLIP, we introduce a dual-branch architecture combining global prompts with a locality-aware pathway. This local branch utilizes Value-Value (V-V) attention and Optimal Transport (OT) to enforce balanced, discriminative alignments between fine-grained image patches and class-specific prompts. Extensive evaluation on 11 benchmarks demonstrates state-of-the-art average accuracy. Furthermore, the framework exhibits superior Out-of-Distribution (OOD) robustness, offering a configurable trade-off between task specialization and generalized robustness.
Date: 05.01.2026 / 13:00 Place: A-212









