Skeleton-based gait recognizers excel at modeling spatial configurations but often underuse explicit motion dynamics that are crucial under appearance changes. We introduce a plug-and-play Wavelet Feature Stream that augments any skeleton backbone with time–frequency dynamics of joint velocities. Concretely, per-joint velocity sequences are transformed by the continuous wavelet transform (CWT) into multi-scale scalograms, from which a lightweight multi-scale CNN learns discriminative dynamic cues. The resulting descriptor is fused with the backbone representation for classification, requiring no changes to the backbone architecture or additional supervision. Across CASIA-B, the proposed stream delivers consistent gains on strong skeleton backbones (e.g., GaitMixer, GaitFormer, GaitGraph) and establishes a new skeleton-based state of the art when attached to Gait-Mixer. The improvements are especially pronounced under covariate shifts such as carrying bags (BG) and wearing coats (CL), highlighting the complementarity of explicit time– frequency modeling and standard spatio–temporal encoders.

Explicit Time-Frequency Dynamics for Skeleton-Based Gait Recognition / Ko, S., Song, Y., Chung, E., Quagliato, L., Lee, T., Noh, J.. - (2026), pp. 906-910. (ICASSP 2026 Barcelona, Spain May 4–8, 2026) [10.48550/arXiv.2604.03002].

Explicit Time-Frequency Dynamics for Skeleton-Based Gait Recognition

Quagliato, L.;
2026-01-01

Abstract

Skeleton-based gait recognizers excel at modeling spatial configurations but often underuse explicit motion dynamics that are crucial under appearance changes. We introduce a plug-and-play Wavelet Feature Stream that augments any skeleton backbone with time–frequency dynamics of joint velocities. Concretely, per-joint velocity sequences are transformed by the continuous wavelet transform (CWT) into multi-scale scalograms, from which a lightweight multi-scale CNN learns discriminative dynamic cues. The resulting descriptor is fused with the backbone representation for classification, requiring no changes to the backbone architecture or additional supervision. Across CASIA-B, the proposed stream delivers consistent gains on strong skeleton backbones (e.g., GaitMixer, GaitFormer, GaitGraph) and establishes a new skeleton-based state of the art when attached to Gait-Mixer. The improvements are especially pronounced under covariate shifts such as carrying bags (BG) and wearing coats (CL), highlighting the complementarity of explicit time– frequency modeling and standard spatio–temporal encoders.
2026
ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing
New York, USA
IEEE Institute of Electrical and Electronics Engineers Inc.
979-8-3315-6701-9
Ko, S.; Song, Y.; Chung, E.; Quagliato, L.; Lee, T.; Noh, J.
Explicit Time-Frequency Dynamics for Skeleton-Based Gait Recognition / Ko, S., Song, Y., Chung, E., Quagliato, L., Lee, T., Noh, J.. - (2026), pp. 906-910. (ICASSP 2026 Barcelona, Spain May 4–8, 2026) [10.48550/arXiv.2604.03002].
File in questo prodotto:
File Dimensione Formato  
2026 [EWHA] ICASSP 2026 Gait analysis.pdf

Solo gestori archivio

Descrizione: IEEE ICASSP 2026 - conference paper
Tipologia: Versione editoriale (Publisher’s layout)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 1.93 MB
Formato Adobe PDF
1.93 MB Adobe PDF   Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/501271
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact