Recognizing arbitrary or previously unseen categories is essential for comprehensive real-world 3D scene understanding. Currently, all existing methods rely on 2D or textual modalities during training or together at inference. This highlights the clear absence of a model capable of processing 3D data alone for learning semantics end-to-end, along with the necessary data to train such a model. Meanwhile, 3D Gaussian Splatting (3DGS) has emerged as the de facto standard for 3D scene representation across various vision tasks. However, effectively integrating semantic reasoning into 3DGS in a generalizable manner remains an open challenge. To address these limitations, we introduce SceneSplat in Fig. 1, to our knowledge the first large-scale 3D indoor scene understanding approach that operates natively on 3DGS. Furthermore, we propose a self-supervised learning scheme that unlocks rich 3D feature learning from unlabeled scenes. To power the proposed methods, we introduce SceneSplat-7K, the first large-scale 3DGS dataset for indoor scenes, comprising 7916 scenes derived from seven established datasets, such as ScanNet and Matterport3D. Generating SceneSplat-7K required computational resources equivalent to 150 GPU days on an L4 GPU, enabling standardized benchmarking for 3DGS-based reasoning for indoor scenes. Our exhaustive experiments on SceneSplat-7K demonstrate the significant benefit of the proposed method over the established baselines. Our code, model, and datasets will be released at SceneSplat.

SceneSplat: Gaussian Splatting-Based Scene Understanding with Vision-Language Pretraining / Li, Y., Ma, Q.i., Yang, R., Li, H., Ma, M., Ren, B., Popovic, N., Sebe, N., Konukoglu, E., Gevers, T., Van Gool, L., Oswald, M.R., Paudel, D.P.. - (2025), pp. 4961-4972. (International Conference on Computer Vision Honolulu October 2025) [10.1109/iccv51701.2025.00472].

SceneSplat: Gaussian Splatting-Based Scene Understanding with Vision-Language Pretraining

Ren, Bin;Sebe, Nicu;
2025-01-01

Abstract

Recognizing arbitrary or previously unseen categories is essential for comprehensive real-world 3D scene understanding. Currently, all existing methods rely on 2D or textual modalities during training or together at inference. This highlights the clear absence of a model capable of processing 3D data alone for learning semantics end-to-end, along with the necessary data to train such a model. Meanwhile, 3D Gaussian Splatting (3DGS) has emerged as the de facto standard for 3D scene representation across various vision tasks. However, effectively integrating semantic reasoning into 3DGS in a generalizable manner remains an open challenge. To address these limitations, we introduce SceneSplat in Fig. 1, to our knowledge the first large-scale 3D indoor scene understanding approach that operates natively on 3DGS. Furthermore, we propose a self-supervised learning scheme that unlocks rich 3D feature learning from unlabeled scenes. To power the proposed methods, we introduce SceneSplat-7K, the first large-scale 3DGS dataset for indoor scenes, comprising 7916 scenes derived from seven established datasets, such as ScanNet and Matterport3D. Generating SceneSplat-7K required computational resources equivalent to 150 GPU days on an L4 GPU, enabling standardized benchmarking for 3DGS-based reasoning for indoor scenes. Our exhaustive experiments on SceneSplat-7K demonstrate the significant benefit of the proposed method over the established baselines. Our code, model, and datasets will be released at SceneSplat.
2025
2025 IEEE/CVF International Conference on Computer Vision (ICCV)
New York
IEEE
979-8-3315-8775-8
Li, Yue; Ma, Qi; Yang, Runyi; Li, Huapeng; Ma, Mengjiao; Ren, Bin; Popovic, Nikola; Sebe, Nicu; Konukoglu, Ender; Gevers, Theo; Van Gool, Luc; Oswald,...espandi
SceneSplat: Gaussian Splatting-Based Scene Understanding with Vision-Language Pretraining / Li, Y., Ma, Q.i., Yang, R., Li, H., Ma, M., Ren, B., Popovic, N., Sebe, N., Konukoglu, E., Gevers, T., Van Gool, L., Oswald, M.R., Paudel, D.P.. - (2025), pp. 4961-4972. (International Conference on Computer Vision Honolulu October 2025) [10.1109/iccv51701.2025.00472].
File in questo prodotto:
File Dimensione Formato  
Li_SceneSplat_Gaussian_Splatting-based_Scene_Understanding_with_Vision-Language_Pretraining_ICCV_2025_paper.pdf

accesso aperto

Tipologia: Post-print referato (Refereed author’s manuscript)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 6.66 MB
Formato Adobe PDF
6.66 MB Adobe PDF Visualizza/Apri
SceneSplat_Gaussian_Splatting-Based_Scene_Understanding_with_Vision-Language_Pretraining.pdf

Solo gestori archivio

Tipologia: Versione editoriale (Publisher’s layout)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 6.54 MB
Formato Adobe PDF
6.54 MB Adobe PDF   Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/486953
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex 1
social impact