We propose a novel spatial-temporal graph Mamba (STG-Mamba) for the music-guided dance video synthesis task, i.e., to translate the input music to a dance video. STG-Mamba consists of two translation mappings: music-to-skeleton translation and skeleton-to-video translation. In the music-to-skeleton translation, we introduce a novel spatial-temporal graph Mamba (STGM) block to effectively construct skeleton sequences from the input music, capturing dependencies between joints in both the spatial and temporal dimensions. For the skeleton-to-video translation, we propose a novel self-supervised regularization network to translate the generated skeletons, along with a conditional image, into a dance video. Lastly, we collect a new skeleton-to-video translation dataset from the Internet, containing 54,944 video clips. Extensive experiments demonstrate that STG-Mamba achieves significantly better results than existing methods.

Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis / Tang, H., Shao, L., Zhang, Z., Van Gool, L., Sebe, N.. - In: IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE. - ISSN 0162-8828. - 47:11(2025), pp. 9626-9636. [10.1109/TPAMI.2025.3588237]

Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis

Hao Tang;Nicu Sebe
2025-01-01

Abstract

We propose a novel spatial-temporal graph Mamba (STG-Mamba) for the music-guided dance video synthesis task, i.e., to translate the input music to a dance video. STG-Mamba consists of two translation mappings: music-to-skeleton translation and skeleton-to-video translation. In the music-to-skeleton translation, we introduce a novel spatial-temporal graph Mamba (STGM) block to effectively construct skeleton sequences from the input music, capturing dependencies between joints in both the spatial and temporal dimensions. For the skeleton-to-video translation, we propose a novel self-supervised regularization network to translate the generated skeletons, along with a conditional image, into a dance video. Lastly, we collect a new skeleton-to-video translation dataset from the Internet, containing 54,944 video clips. Extensive experiments demonstrate that STG-Mamba achieves significantly better results than existing methods.
2025
11
Tang, Hao; Shao, Ling; Zhang, Zhenyu; Van Gool, Luc; Sebe, Nicu
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis / Tang, H., Shao, L., Zhang, Z., Van Gool, L., Sebe, N.. - In: IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE. - ISSN 0162-8828. - 47:11(2025), pp. 9626-9636. [10.1109/TPAMI.2025.3588237]
File in questo prodotto:
File Dimensione Formato  
2507.06689v1 (1)-compressed.pdf

Solo gestori archivio

Tipologia: Altro materiale allegato (Other attachments)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 397.21 kB
Formato Adobe PDF
397.21 kB Adobe PDF   Visualizza/Apri
Spatial-Temporal_Graph_Mamba_for_Music-Guided_Dance_Video_Synthesis.pdf

Solo gestori archivio

Tipologia: Versione editoriale (Publisher’s layout)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 8.02 MB
Formato Adobe PDF
8.02 MB Adobe PDF   Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/464941
Citazioni
  • ???jsp.display-item.citation.pmc??? 1
  • Scopus 2
  • ???jsp.display-item.citation.isi??? 1
  • OpenAlex 5
social impact