Beyond task success: A closer look at jointly learning to see, ask, and GuessWhat

IRIS

We propose a grounded dialogue state encoder which addresses a foundational issue on how to integrate visual grounding with dialogue system components. As a test-bed, we focus on the GuessWhat?! game, a two-player game where the goal is to identify an object in a complex visual scene by asking a sequence of yes/no questions. Our visually-grounded encoder leverages synergies between guessing and asking questions, as it is trained jointly using multi-task learning. We further enrich our model via a cooperative learning regime. We show that the introduction of both the joint architecture and cooperative learning lead to accuracy improvements over the baseline system. We compare our approach to an alternative system which extends the baseline with reinforcement learning. Our in-depth analysis shows that the linguistic skills of the two models differ dramatically, despite approaching comparable performance levels. This points at the importance of analyzing the linguistic output of competing systems beyond numeric comparison solely based on task success

Beyond task success: A closer look at jointly learning to see, ask, and GuessWhat / Shekhar, Ravi; Venkatesh, Aashish; Baumgärtner, Tim; Bruni, Elia; Plank, Barbara; Bernardi, Raffaella; Fernández, Raquel. - ELETTRONICO. - (2019), pp. 2578-2587. (Intervento presentato al convegno NAACL HLT 2019 tenutosi a Minneapolis, MN nel 2nd-5th June 2019) [10.18653/v1/N19-1265].

Beyond task success: A closer look at jointly learning to see, ask, and GuessWhat

Shekhar Ravi;Venkatesh Aashish;Baumgärtner Tim;Bruni Elia;Plank Barbara;Bernardi Raffaella;Fernández Raquel

2019-01-01

Abstract

We propose a grounded dialogue state encoder which addresses a foundational issue on how to integrate visual grounding with dialogue system components. As a test-bed, we focus on the GuessWhat?! game, a two-player game where the goal is to identify an object in a complex visual scene by asking a sequence of yes/no questions. Our visually-grounded encoder leverages synergies between guessing and asking questions, as it is trained jointly using multi-task learning. We further enrich our model via a cooperative learning regime. We show that the introduction of both the joint architecture and cooperative learning lead to accuracy improvements over the baseline system. We compare our approach to an alternative system which extends the baseline with reinforcement learning. Our in-depth analysis shows that the linguistic skills of the two models differ dramatically, despite approaching comparable performance levels. This points at the importance of analyzing the linguistic output of competing systems beyond numeric comparison solely based on task success

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno di pubblicazione (Date of publication)
	
				2019
			
	Titolo del volume (Proceedings title)
	
				NAACL HLT 2019: The 2019 Conferenceof the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Proceedings of the Conference Vol. 1: Long and Short Papers
			
	Luogo di edizione (Place of publication)
	
				Stroudsburg, PA
			
	Casa editrice (Publisher)
	
				ACL
			
	ISBN
	
				978-1-950737-13-0
			
	Codice Scopus (Scopus Identifier)
	
				2-s2.0-85074818683
			
	Codice WOS (WOS identifier)
	
				WOS:000900116902076
			
	Tutti gli autori
	
						Shekhar, Ravi; Venkatesh, Aashish; Baumgärtner, Tim; Bruni, Elia; Plank, Barbara; Bernardi, Raffaella; Fernández, Raquel
					
	Citazione
	
				Beyond task success: A closer look at jointly learning to see, ask, and GuessWhat / Shekhar, Ravi; Venkatesh, Aashish; Baumgärtner, Tim; Bruni, Elia; Plank, Barbara; Bernardi, Raffaella; Fernández, Raquel. - ELETTRONICO. - (2019), pp. 2578-2587. (Intervento presentato al  convegno NAACL HLT 2019 tenutosi a Minneapolis, MN nel 2nd-5th June 2019) [10.18653/v1/N19-1265].
			
	Appare nelle tipologie:
	
				04.1 Saggio in atti di convegno (Paper in Proceedings)

File in questo prodotto:

File	Dimensione	Formato
naacl19.pdf accesso aperto Tipologia: Versione editoriale (Publisher’s layout) Licenza: Creative commons Dimensione 898.27 kB Formato Adobe PDF Visualizza/Apri	898.27 kB	Adobe PDF	Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/250569

Citazioni

ND

39

22

ND

social impact