Off-the-shelf visual representations have been widely applied in various tasks. However, as image retrieval involves a compact representation, the above practice does not obtain convincing performance, especially in realistic scenarios of novel domains and categories. In this paper, we make the first attempt to address it by organizing a generalized image retrieval task and proposing an off-the-shelf quantizer. Challenges of realizing them include two perspectives: visual inconsistency across domains and hidden semantics of unknown categories, which corrupt compact features. To tackle the former issue, we propose a cross-aligned contrastive learning objective for model training. It simultaneously reduces quantization error and domain gap, encouraging the model to generate domain-invariant quantization codes. To tackle the latter one, we design a 2nd order codebook which holds representative information of seen categories. Novel images are extracted by performing a compositional projection via all 2nd order codewords, which improves generalization ability on unknowns. Combining them, we obtain a significant performance gain compared to the current state-of-the-art, where an up to 8.0% increase in mAP is observed. Code is available at: https://github.com/ZHlo-404/GIR-OTSQ.
Generalized Image Retrieval with Off-The-Shelf Quantizer / Zeng, P., Duan, Y., Zhu, X., Song, J., Gao, L., Sebe, N., Shen, H.. - In: INTERNATIONAL JOURNAL OF COMPUTER VISION. - ISSN 0920-5691. - 134:7(2026). [10.1007/s11263-026-02857-5]
Generalized Image Retrieval with Off-The-Shelf Quantizer
Song, Jingkuan;Sebe, Nicu;
2026-01-01
Abstract
Off-the-shelf visual representations have been widely applied in various tasks. However, as image retrieval involves a compact representation, the above practice does not obtain convincing performance, especially in realistic scenarios of novel domains and categories. In this paper, we make the first attempt to address it by organizing a generalized image retrieval task and proposing an off-the-shelf quantizer. Challenges of realizing them include two perspectives: visual inconsistency across domains and hidden semantics of unknown categories, which corrupt compact features. To tackle the former issue, we propose a cross-aligned contrastive learning objective for model training. It simultaneously reduces quantization error and domain gap, encouraging the model to generate domain-invariant quantization codes. To tackle the latter one, we design a 2nd order codebook which holds representative information of seen categories. Novel images are extracted by performing a compositional projection via all 2nd order codewords, which improves generalization ability on unknowns. Combining them, we obtain a significant performance gain compared to the current state-of-the-art, where an up to 8.0% increase in mAP is observed. Code is available at: https://github.com/ZHlo-404/GIR-OTSQ.| File | Dimensione | Formato | |
|---|---|---|---|
|
s11263-026-02857-5.pdf
Solo gestori archivio
Tipologia:
Versione editoriale (Publisher’s layout)
Licenza:
Tutti i diritti riservati (All rights reserved)
Dimensione
4.34 MB
Formato
Adobe PDF
|
4.34 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione



