Open-vocabulary object detection (OvOD) is set to revolutionize security screening by enabling systems to recognize any item in X-ray scans. However, developing effective OvOD models for X-ray imaging presents unique challenges due to data scarcity and the modality gap that prevents direct adoption of RGB-based solutions. To overcome these limitations, we propose RAXO, a training-free framework that repurposes off-the-shelf RGB OvOD detectors for robust X-ray detection. RAXO builds high-quality X-ray class descriptors using a dual-source retrieval strategy. It gathers relevant RGB images from the web and enriches them via a novel X-ray material transfer mechanism, eliminating the need for labeled databases. These visual descriptors replace text-based classification in OvOD, leveraging intra-modal feature distances for robust detection. Extensive experiments demonstrate that RAXO consistently improves OvOD performance, providing an average mAP increase of up to 17.0 points over base detectors. To further support research in this emerging field, we also introduce DET-COMPASS, a new benchmark featuring bounding box annotations for over 300 object categories, enabling large-scale evaluation of OvOD in X-ray. Code and dataset available at: https://pagf188.github.io/RAXO/.

Superpowering Open-Vocabulary Object Detectors for X-ray Vision / Garcia-Fernandez, P., Vaquero, L., Liu, M., Xue, F., Cores, D., Sebe, N., Mucientes, M., Ricci, E.. - (2025), pp. 20770-20779. (International Conference on Computer Vision Honolulu October 2025) [10.1109/iccv51701.2025.01931].

Superpowering Open-Vocabulary Object Detectors for X-ray Vision

Liu, Mingxuan;Xue, Feng;Sebe, Nicu;Ricci, Elisa
2025-01-01

Abstract

Open-vocabulary object detection (OvOD) is set to revolutionize security screening by enabling systems to recognize any item in X-ray scans. However, developing effective OvOD models for X-ray imaging presents unique challenges due to data scarcity and the modality gap that prevents direct adoption of RGB-based solutions. To overcome these limitations, we propose RAXO, a training-free framework that repurposes off-the-shelf RGB OvOD detectors for robust X-ray detection. RAXO builds high-quality X-ray class descriptors using a dual-source retrieval strategy. It gathers relevant RGB images from the web and enriches them via a novel X-ray material transfer mechanism, eliminating the need for labeled databases. These visual descriptors replace text-based classification in OvOD, leveraging intra-modal feature distances for robust detection. Extensive experiments demonstrate that RAXO consistently improves OvOD performance, providing an average mAP increase of up to 17.0 points over base detectors. To further support research in this emerging field, we also introduce DET-COMPASS, a new benchmark featuring bounding box annotations for over 300 object categories, enabling large-scale evaluation of OvOD in X-ray. Code and dataset available at: https://pagf188.github.io/RAXO/.
2025
2025 IEEE/CVF International Conference on Computer Vision (ICCV)
New York
IEEE
979-8-3315-8775-8
Garcia-Fernandez, Pablo; Vaquero, Lorenzo; Liu, Mingxuan; Xue, Feng; Cores, Daniel; Sebe, Nicu; Mucientes, Manuel; Ricci, Elisa
Superpowering Open-Vocabulary Object Detectors for X-ray Vision / Garcia-Fernandez, P., Vaquero, L., Liu, M., Xue, F., Cores, D., Sebe, N., Mucientes, M., Ricci, E.. - (2025), pp. 20770-20779. (International Conference on Computer Vision Honolulu October 2025) [10.1109/iccv51701.2025.01931].
File in questo prodotto:
File Dimensione Formato  
Garcia-Fernandez_Superpowering_Open-Vocabulary_Object_Detectors_for_X-ray_Vision_ICCV_2025_paper.pdf

accesso aperto

Tipologia: Post-print referato (Refereed author’s manuscript)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 1.4 MB
Formato Adobe PDF
1.4 MB Adobe PDF Visualizza/Apri
Superpowering_Open-Vocabulary_Object_Detectors_for_X-ray_Vision.pdf

Solo gestori archivio

Tipologia: Versione editoriale (Publisher’s layout)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 1.11 MB
Formato Adobe PDF
1.11 MB Adobe PDF   Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/486955
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex 1
social impact