Vision-Language Models (VLMs) have shown significant promise in Visual Question Answering (VQA) tasks by leveraging web-scale multimodal datasets. However, these models often struggle with continual learning due to catastrophic forgetting when adapting to new tasks. As an effective remedy to mitigate catastrophic forgetting, rehearsal strategy uses the data of past tasks upon learning new task. However, such strategy incurs the need of storing past data, which might not be feasible due to hardware constraints or privacy concerns. In this work, we propose the first data-free method that leverages the language generation capability of a VLM, instead of relying on external models, to produce pseudo-rehearsal data for addressing continual VQA. Our proposal, named as GaB, generates pseudo-rehearsal data by posing previous task questions on new task data. Yet, despite being effective, the distribution of generated questions skews towards the most frequently posed questions due to the limited and task-specific training data. To mitigate this issue, we introduce a pseudo-rehearsal balancing module that aligns the generated data towards the ground-truth data distribution using either the question meta-statistics or an unsupervised clustering method. We evaluate our proposed method on two recent benchmarks, i.e. VQACL- VQAv2 and CLOVE-function benchmarks. GaB outperforms all the data-free baselines with substantial improvement in maintaining VQA performance across evolving tasks, while being on-par with methods with access to the past data. Code and models are available at https://github.com/Deepayan137/GaB.

One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering / Das, D., Talon, D., Mancini, M., Wang, Y., Ricci, E.. - (2025), pp. 5635-5645. (WACV Tucson, AZ, USA 26 February 2025 - 06 March 2025) [10.1109/WACV61041.2025.00550].

One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering

Deepayan Das
;
Davide Talon;Massimiliano Mancini;Yiming Wang;Elisa Ricci
2025-01-01

Abstract

Vision-Language Models (VLMs) have shown significant promise in Visual Question Answering (VQA) tasks by leveraging web-scale multimodal datasets. However, these models often struggle with continual learning due to catastrophic forgetting when adapting to new tasks. As an effective remedy to mitigate catastrophic forgetting, rehearsal strategy uses the data of past tasks upon learning new task. However, such strategy incurs the need of storing past data, which might not be feasible due to hardware constraints or privacy concerns. In this work, we propose the first data-free method that leverages the language generation capability of a VLM, instead of relying on external models, to produce pseudo-rehearsal data for addressing continual VQA. Our proposal, named as GaB, generates pseudo-rehearsal data by posing previous task questions on new task data. Yet, despite being effective, the distribution of generated questions skews towards the most frequently posed questions due to the limited and task-specific training data. To mitigate this issue, we introduce a pseudo-rehearsal balancing module that aligns the generated data towards the ground-truth data distribution using either the question meta-statistics or an unsupervised clustering method. We evaluate our proposed method on two recent benchmarks, i.e. VQACL- VQAv2 and CLOVE-function benchmarks. GaB outperforms all the data-free baselines with substantial improvement in maintaining VQA performance across evolving tasks, while being on-par with methods with access to the past data. Code and models are available at https://github.com/Deepayan137/GaB.
2025
2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
Los Alamitos, CA, USA
IEEE Computer Society
979-8-3315-1083-1
Das, Deepayan; Talon, Davide; Mancini, Massimiliano; Wang, Yiming; Ricci, Elisa
One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering / Das, D., Talon, D., Mancini, M., Wang, Y., Ricci, E.. - (2025), pp. 5635-5645. (WACV Tucson, AZ, USA 26 February 2025 - 06 March 2025) [10.1109/WACV61041.2025.00550].
File in questo prodotto:
File Dimensione Formato  
Das_One_VLM_to_Keep_it_Learning_Generation_and_Balancing_for_WACV_2025_paper.pdf

accesso aperto

Tipologia: Post-print referato (Refereed author’s manuscript)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 2.83 MB
Formato Adobe PDF
2.83 MB Adobe PDF Visualizza/Apri
One_VLM_to_Keep_it_Learning_Generation_and_Balancing_for_Data-free_Continual_Visual_Question_Answering.pdf

Solo gestori archivio

Tipologia: Versione editoriale (Publisher’s layout)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 2.73 MB
Formato Adobe PDF
2.73 MB Adobe PDF   Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/472132
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 6
  • ???jsp.display-item.citation.isi??? 3
  • OpenAlex 6
social impact