Video-to-Piano (V2P) requires precise modeling of symbolic structures such as note onsets, dynamics, and sustain, which standard end-to-end video-to-audio methods often fail to capture due to limited temporal and structural control. We propose a hierarchical V2P framework that introduces MIDI as an intermediate representation, with progressive MIDI prediction (pitch, velocity, sustain) guiding waveform synthesis. To better capture performance dynamics, we design a multi-view MIDI predictor that accepts different levels of video inputs: the top view generates coarse MIDI, the addition of front and side views enables precise velocity estimation, and incorporating the pedal view captures fine-grained sustain control. Furthermore, we propose a control branch that injects the predicted MIDI into V2A/V2M models, significantly improving temporal alignment and expressive quality. Experiments on public benchmarks demonstrate that our approach achieves high-fidelity piano audio closely synchronized with video content.
Towards Multi-View Hierarchical Video-to-Piano Generation with MIDI Guidance / Liu, C., Chen, Z., Chen, G., Ding, C., Sebe, N.. - (2026), pp. 12002-12006. (ICASSP Barcelona May 2026) [10.1109/icassp55912.2026.11465075].
Towards Multi-View Hierarchical Video-to-Piano Generation with MIDI Guidance
Liu, Chang;Sebe, Nicu
2026-01-01
Abstract
Video-to-Piano (V2P) requires precise modeling of symbolic structures such as note onsets, dynamics, and sustain, which standard end-to-end video-to-audio methods often fail to capture due to limited temporal and structural control. We propose a hierarchical V2P framework that introduces MIDI as an intermediate representation, with progressive MIDI prediction (pitch, velocity, sustain) guiding waveform synthesis. To better capture performance dynamics, we design a multi-view MIDI predictor that accepts different levels of video inputs: the top view generates coarse MIDI, the addition of front and side views enables precise velocity estimation, and incorporating the pedal view captures fine-grained sustain control. Furthermore, we propose a control branch that injects the predicted MIDI into V2A/V2M models, significantly improving temporal alignment and expressive quality. Experiments on public benchmarks demonstrate that our approach achieves high-fidelity piano audio closely synchronized with video content.| File | Dimensione | Formato | |
|---|---|---|---|
|
Towards_Multi-View_Hierarchical_Video-to-Piano_Generation_with_MIDI_Guidance.pdf
Solo gestori archivio
Tipologia:
Versione editoriale (Publisher’s layout)
Licenza:
Tutti i diritti riservati (All rights reserved)
Dimensione
3.8 MB
Formato
Adobe PDF
|
3.8 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione



