Instruction-based image editing models offer increased personalization opportunities in generative tasks. However, properly evaluating their results is challenging, and most of the existing metrics lag in terms of alignment with human judgment and explainability. To tackle these issues, we introduce DICE (DIfference Coherence Estimator), a model designed to detect localized differences between the original and the edited image and to assess their relevance to the given modification request. DICE consists of two key components: a difference detector and a coherence estimator, both built on an autoregressive Multimodal Large Language Model (MLLM) and trained using a strategy that leverages self-supervision, distillation from inpainting networks, and full supervision. Through extensive experiments, we evaluate each stage of our pipeline, comparing different MLLMs within the proposed framework. We demonstrate that DICE effectively identifies coherent edits, effectively evaluating images generated by different editing models with a strong correlation with human judgment. We publicly release our source code, models, and data at project page.

What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models / Baraldi, L., Bucciarelli, D., Betti, F., Cornia, M., Baraldi, L., Sebe, N., Cucchiara, R.. - (2025), pp. 16217-16226. (International Conference on Computer Vision Honolulu October 2025) [10.1109/iccv51701.2025.01505].

What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models

Betti, Federico;Sebe, Nicu;
2025-01-01

Abstract

Instruction-based image editing models offer increased personalization opportunities in generative tasks. However, properly evaluating their results is challenging, and most of the existing metrics lag in terms of alignment with human judgment and explainability. To tackle these issues, we introduce DICE (DIfference Coherence Estimator), a model designed to detect localized differences between the original and the edited image and to assess their relevance to the given modification request. DICE consists of two key components: a difference detector and a coherence estimator, both built on an autoregressive Multimodal Large Language Model (MLLM) and trained using a strategy that leverages self-supervision, distillation from inpainting networks, and full supervision. Through extensive experiments, we evaluate each stage of our pipeline, comparing different MLLMs within the proposed framework. We demonstrate that DICE effectively identifies coherent edits, effectively evaluating images generated by different editing models with a strong correlation with human judgment. We publicly release our source code, models, and data at project page.
2025
2025 IEEE/CVF International Conference on Computer Vision (ICCV)
New York
IEEE
979-8-3315-8775-8
979-8-3315-8776-5
Baraldi, Lorenzo; Bucciarelli, Davide; Betti, Federico; Cornia, Marcella; Baraldi, Lorenzo; Sebe, Nicu; Cucchiara, Rita
What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models / Baraldi, L., Bucciarelli, D., Betti, F., Cornia, M., Baraldi, L., Sebe, N., Cucchiara, R.. - (2025), pp. 16217-16226. (International Conference on Computer Vision Honolulu October 2025) [10.1109/iccv51701.2025.01505].
File in questo prodotto:
File Dimensione Formato  
Baraldi_What_Changed_Detecting_and_Evaluating_Instruction-Guided_Image_Edits_with_Multimodal_ICCV_2025_paper.pdf

accesso aperto

Tipologia: Versione editoriale (Publisher’s layout)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 1.7 MB
Formato Adobe PDF
1.7 MB Adobe PDF Visualizza/Apri
What_Changed_Detecting_and_Evaluating_Instruction-Guided_Image_Edits_with_Multimodal_Large_Language_Models.pdf

Solo gestori archivio

Tipologia: Versione editoriale (Publisher’s layout)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 1.51 MB
Formato Adobe PDF
1.51 MB Adobe PDF   Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/486957
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex 0
social impact