Anglicized words and misspelled cognates in native language identification

IRIS

In this paper, we present experiments that estimate the impact of specific lexical choices of people writing in a second language (L2). In particular, we look at misspelled words that indicate lexical uncertainty on the part of the author, and separate them into three categories: misspelled cognates, “L2-ed” (in our case, anglicized) words, and all other spelling errors. We test the assumption that such errors contain clues about the native language of an essay’s author through the task of native language identification. The results of the experiments show that the information brought by each of these categories is complementary. We also note that while the distribution of such features changes with the proficiency level of the writer, their contribution towards native language identification remains significant at all levels.

Anglicized words and misspelled cognates in native language identification / Markov, I., Nastase, V., Strapparava, C.. - ELETTRONICO. - (2019), pp. 275-284. (Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications @ ACL 2019 Florence, Italy August 2019) [10.18653/v1/W19-4429].

Anglicized words and misspelled cognates in native language identification

Ilia Markov;Vivi Nastase;Carlo Strapparava

2019-01-01

Abstract

In this paper, we present experiments that estimate the impact of specific lexical choices of people writing in a second language (L2). In particular, we look at misspelled words that indicate lexical uncertainty on the part of the author, and separate them into three categories: misspelled cognates, “L2-ed” (in our case, anglicized) words, and all other spelling errors. We test the assumption that such errors contain clues about the native language of an essay’s author through the task of native language identification. The results of the experiments show that the information brought by each of these categories is complementary. We also note that while the distribution of such features changes with the proficiency level of the writer, their contribution towards native language identification remains significant at all levels.

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno di pubblicazione (Date of publication)
	
				2019
			
	Titolo del volume (Proceedings title)
	
				Proceedings of the Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications
			
	Luogo di edizione (Place of publication)
	
				USA
			
	Casa editrice (Publisher)
	
				Association for Computational Linguistics
			
	ISBN
	
				978-1-950737-34-5
			
	Codice Scopus (Scopus Identifier)
	
				2-s2.0-85120977900
			
	Codice WOS (WOS identifier)
	
				WOS:000521943400029
			
	Tutti gli autori
	
						Markov, Ilia; Nastase, Vivi; Strapparava, Carlo
					
	Citazione
	
				Anglicized words and misspelled cognates in native language identification / Markov, I., Nastase, V., Strapparava, C.. - ELETTRONICO. - (2019), pp. 275-284. (Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications @ ACL 2019 Florence, Italy August 2019) [10.18653/v1/W19-4429].
			
	Appare nelle tipologie:
	
				04.1 Saggio in atti di convegno (Paper in Proceedings)

File in questo prodotto:

Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/343610

Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni

ND

3

2

2

social impact