Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages

IRIS

An old-school recipe for training a classifier is to (i) learn a good feature extractor and (ii) optimize a linear layer atop. When only a handful of samples are available per category, as in Few-Shot Adaptation (FSA), data are insufficient to fit a large number of parameters, rendering the above impractical. This is especially true with large pre-trained Vision-Language Models (VLMs), which motivated successful research at the intersection of Parameter-Efficient Fine-tuning (PEFT) and FSA. In this work, we start by analyzing the learning dynamics of PEFT techniques when trained on few-shot data from only a subset of categories, referred to as the "base"classes. We show that such dynamics naturally splits into two distinct phases: (i) task-level feature extraction and (ii) specialization to the available concepts. To accommodate this dynamic, we then depart from prompt- or adapter-based methods and tackle FSA differently. Specifically, given a fixed computational budget, we split it to (i) learn a task-specific feature extractor via PEFT and (ii) train a linear classifier on top. We call this scheme Two-Stage Few-Shot Adaptation (2SFS). Differently from established methods, our scheme enables a novel form of selective inference at a category level, i.e., at test time, only novel categories are embedded by the adapted text encoder, while embeddings of base categories are available within the classifier. Results with fixed hyperparameters across two settings, three backbones, and eleven datasets, show that 2SFS matches or surpasses the state-of-the-art, while established methods degrade significantly across settings..

Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages / Farina, M., Mancini, M., Iacca, G., Ricci, E.. - (2025), pp. 29989-29998. (2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 Nashville 10th June-17th June 2025) [10.1109/cvpr52734.2025.02791].

Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages

Farina, Matteo;Mancini, Massimiliano;Iacca, Giovanni;Ricci, Elisa

2025-01-01

Abstract

An old-school recipe for training a classifier is to (i) learn a good feature extractor and (ii) optimize a linear layer atop. When only a handful of samples are available per category, as in Few-Shot Adaptation (FSA), data are insufficient to fit a large number of parameters, rendering the above impractical. This is especially true with large pre-trained Vision-Language Models (VLMs), which motivated successful research at the intersection of Parameter-Efficient Fine-tuning (PEFT) and FSA. In this work, we start by analyzing the learning dynamics of PEFT techniques when trained on few-shot data from only a subset of categories, referred to as the "base"classes. We show that such dynamics naturally splits into two distinct phases: (i) task-level feature extraction and (ii) specialization to the available concepts. To accommodate this dynamic, we then depart from prompt- or adapter-based methods and tackle FSA differently. Specifically, given a fixed computational budget, we split it to (i) learn a task-specific feature extractor via PEFT and (ii) train a linear classifier on top. We call this scheme Two-Stage Few-Shot Adaptation (2SFS). Differently from established methods, our scheme enables a novel form of selective inference at a category level, i.e., at test time, only novel categories are embedded by the adapted text encoder, while embeddings of base categories are available within the classifier. Results with fixed hyperparameters across two settings, three backbones, and eleven datasets, show that 2SFS matches or surpasses the state-of-the-art, while established methods degrade significantly across settings..

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno di pubblicazione (Date of publication)
	
				2025
			
	Titolo del volume (Proceedings title)
	
				Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
			
	Luogo di edizione (Place of publication)
	
				New York, NY, USA
			
	Casa editrice (Publisher)
	
				IEEE Computer Society
			
	Codice Scopus (Scopus Identifier)
	
				2-s2.0-105017049694
			
	Codice WOS (WOS identifier)
	
				WOS:001601181100551
			
	Tutti gli autori
	
						Farina, Matteo; Mancini, Massimiliano; Iacca, Giovanni; Ricci, Elisa
					
	Citazione
	
				Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages / Farina, M., Mancini, M., Iacca, G., Ricci, E.. - (2025), pp. 29989-29998. (2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 Nashville 10th June-17th June 2025) [10.1109/cvpr52734.2025.02791].

File in questo prodotto:

Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/464534

Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni

ND

10

2

5

social impact