Video-to-Piano (V2P) requires precise modeling of symbolic structures such as note onsets, dynamics, and sustain, which standard end-to-end video-to-audio methods often fail to capture due to limited temporal and structural control. We propose a hierarchical V2P framework that introduces MIDI as an intermediate representation, with progressive MIDI prediction (pitch, velocity, sustain) guiding waveform synthesis. To better capture performance dynamics, we design a multi-view MIDI predictor that accepts different levels of video inputs: the top view generates coarse MIDI, the addition of front and side views enables precise velocity estimation, and incorporating the pedal view captures fine-grained sustain control. Furthermore, we propose a control branch that injects the predicted MIDI into V2A/V2M models, significantly improving temporal alignment and expressive quality. Experiments on public benchmarks demonstrate that our approach achieves high-fidelity piano audio closely synchronized with video content.

Towards Multi-View Hierarchical Video-to-Piano Generation with MIDI Guidance / Liu, C., Chen, Z., Chen, G., Ding, C., Sebe, N.. - (2026), pp. 12002-12006. (ICASSP Barcelona May 2026) [10.1109/icassp55912.2026.11465075].

Towards Multi-View Hierarchical Video-to-Piano Generation with MIDI Guidance

Liu, Chang;Sebe, Nicu
2026-01-01

Abstract

Video-to-Piano (V2P) requires precise modeling of symbolic structures such as note onsets, dynamics, and sustain, which standard end-to-end video-to-audio methods often fail to capture due to limited temporal and structural control. We propose a hierarchical V2P framework that introduces MIDI as an intermediate representation, with progressive MIDI prediction (pitch, velocity, sustain) guiding waveform synthesis. To better capture performance dynamics, we design a multi-view MIDI predictor that accepts different levels of video inputs: the top view generates coarse MIDI, the addition of front and side views enables precise velocity estimation, and incorporating the pedal view captures fine-grained sustain control. Furthermore, we propose a control branch that injects the predicted MIDI into V2A/V2M models, significantly improving temporal alignment and expressive quality. Experiments on public benchmarks demonstrate that our approach achieves high-fidelity piano audio closely synchronized with video content.
2026
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
New York
IEEE
979-8-3315-6701-9
979-8-3315-6702-6
Liu, Chang; Chen, Zihao; Chen, Gongyu; Ding, Chaofan; Sebe, Nicu
Towards Multi-View Hierarchical Video-to-Piano Generation with MIDI Guidance / Liu, C., Chen, Z., Chen, G., Ding, C., Sebe, N.. - (2026), pp. 12002-12006. (ICASSP Barcelona May 2026) [10.1109/icassp55912.2026.11465075].
File in questo prodotto:
File Dimensione Formato  
Towards_Multi-View_Hierarchical_Video-to-Piano_Generation_with_MIDI_Guidance.pdf

Solo gestori archivio

Tipologia: Versione editoriale (Publisher’s layout)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 3.8 MB
Formato Adobe PDF
3.8 MB Adobe PDF   Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/491830
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex 0
social impact