Existing text-to-motion (T2M) generation methods primarily rely on regression-based objectives, such as minimizing positional errors. However, they lack effective semantic supervision and correction mechanisms, often leading to substantial misalignment between text and motion. To address this, we propose Aligned Text-to-Motion (ATM), a semantics-aware generation framework that automatically identifies and corrects text-motion misalignment. ATM incorporates two key components: (1) Inter-motion alignment, which detects semantic contradictions across motions and applies adaptive corrections based on the degree of semantic discrepancy, flexibly handing diverse mis-alignments and ensuring global text-motion consistency; (2) Intra-motion alignment, which refines locally missing or inaccurate motion semantics in an unsupervised manner by inferring semantic proxies, effectively addressing the absence of localized textual annotations. ATM is model-agnostic and can be seamlessly integrated into various T2M methods as a plug-and-play module. Extensive experiments on HumanML3D and KIT demonstrate that ATM consistently improves both generation quality and text-motion alignment. Code is available at https://github.com/ke-han-aca/ATM.git.

ATM: Enhanced Alignment for Text-to-Motion Generation / Han, K.e., Lyu, Y., Yu, W., Sebe, N.. - (2026), pp. 6862-6872. (IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Tucson, Arizona March 2026) [10.1109/wacv61042.2026.00663].

ATM: Enhanced Alignment for Text-to-Motion Generation

Han, Ke;Sebe, Nicu
2026-01-01

Abstract

Existing text-to-motion (T2M) generation methods primarily rely on regression-based objectives, such as minimizing positional errors. However, they lack effective semantic supervision and correction mechanisms, often leading to substantial misalignment between text and motion. To address this, we propose Aligned Text-to-Motion (ATM), a semantics-aware generation framework that automatically identifies and corrects text-motion misalignment. ATM incorporates two key components: (1) Inter-motion alignment, which detects semantic contradictions across motions and applies adaptive corrections based on the degree of semantic discrepancy, flexibly handing diverse mis-alignments and ensuring global text-motion consistency; (2) Intra-motion alignment, which refines locally missing or inaccurate motion semantics in an unsupervised manner by inferring semantic proxies, effectively addressing the absence of localized textual annotations. ATM is model-agnostic and can be seamlessly integrated into various T2M methods as a plug-and-play module. Extensive experiments on HumanML3D and KIT demonstrate that ATM consistently improves both generation quality and text-motion alignment. Code is available at https://github.com/ke-han-aca/ATM.git.
2026
IEEE/CVF Winter Conference on Applications of Computer Vision
New York
IEEE
979-8-3315-5511-5
979-8-3315-5512-2
Han, Ke; Lyu, Yueming; Yu, Weichen; Sebe, Nicu
ATM: Enhanced Alignment for Text-to-Motion Generation / Han, K.e., Lyu, Y., Yu, W., Sebe, N.. - (2026), pp. 6862-6872. (IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Tucson, Arizona March 2026) [10.1109/wacv61042.2026.00663].
File in questo prodotto:
File Dimensione Formato  
Han_ATM_Enhanced_Alignment_for_Text-to-Motion_Generation_WACV_2026_paper.pdf

accesso aperto

Tipologia: Versione editoriale (Publisher’s layout)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 1.74 MB
Formato Adobe PDF
1.74 MB Adobe PDF Visualizza/Apri
ATM_Enhanced_Alignment_for_Text-to-Motion_Generation.pdf

Solo gestori archivio

Tipologia: Versione editoriale (Publisher’s layout)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 4.09 MB
Formato Adobe PDF
4.09 MB Adobe PDF   Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/486930
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex 0
social impact