Beyond Voice Activity Detection: Hybrid Audio Segmentation for Direct Speech Translation

IRIS

The audio segmentation mismatch between training data and those seen at run-time is a major problem in direct speech translation. Indeed, while systems are usually trained on manually segmented corpora, in real use cases they are often presented with continuous audio requiring automatic (and sub-optimal) segmentation. After comparing existing techniques (VAD-based, fixed-length and hybrid segmentation methods), in this paper we propose enhanced hybrid solutions to produce better results without sacrificing latency. Through experiments on different domains and language pairs, we show that our methods outperform all the other techniques, reducing by at least 30% the gap between the traditional VAD-based approach and optimal manual segmentation.

Beyond Voice Activity Detection: Hybrid Audio Segmentation for Direct Speech Translation / Gaido, M., Negri, M., Cettolo, M., Turchi, M.. - (2021), pp. 55-62. (4th International Conference on Natural Language and Speech Processing, ICNLSP 2021 Trento, Italy 12-13 November 2021).

Beyond Voice Activity Detection: Hybrid Audio Segmentation for Direct Speech Translation

Gaido M.;Negri M.;Cettolo M.;Turchi M.

2021-01-01

Abstract

The audio segmentation mismatch between training data and those seen at run-time is a major problem in direct speech translation. Indeed, while systems are usually trained on manually segmented corpora, in real use cases they are often presented with continuous audio requiring automatic (and sub-optimal) segmentation. After comparing existing techniques (VAD-based, fixed-length and hybrid segmentation methods), in this paper we propose enhanced hybrid solutions to produce better results without sacrificing latency. Through experiments on different domains and language pairs, we show that our methods outperform all the other techniques, reducing by at least 30% the gap between the traditional VAD-based approach and optimal manual segmentation.

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno di pubblicazione (Date of publication)
	
				2021
			
	Titolo del volume (Proceedings title)
	
				ICNLSP 2021 - Proceedings of the 4th International Conference on Natural Language and Speech Processing
			
	Luogo di edizione (Place of publication)
	
				Trento, Italy
			
	Casa editrice (Publisher)
	
				Association for Computational Linguistics (ACL)
			
	Codice Scopus (Scopus Identifier)
	
				2-s2.0-85136321161
			
	Tutti gli autori
	
						Gaido, M.; Negri, M.; Cettolo, M.; Turchi, M.
					
	Citazione
	
				Beyond Voice Activity Detection: Hybrid Audio Segmentation for Direct Speech Translation / Gaido, M., Negri, M., Cettolo, M., Turchi, M.. - (2021), pp. 55-62. (4th International Conference on Natural Language and Speech Processing, ICNLSP 2021 Trento, Italy 12-13 November 2021).

File in questo prodotto:

Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/369992

Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni

ND

16

ND

ND

social impact