In this paper, we present experiments that estimate the impact of specific lexical choices of people writing in a second language (L2). In particular, we look at misspelled words that indicate lexical uncertainty on the part of the author, and separate them into three categories: misspelled cognates, “L2-ed” (in our case, anglicized) words, and all other spelling errors. We test the assumption that such errors contain clues about the native language of an essay’s author through the task of native language identification. The results of the experiments show that the information brought by each of these categories is complementary. We also note that while the distribution of such features changes with the proficiency level of the writer, their contribution towards native language identification remains significant at all levels.
Anglicized words and misspelled cognates in native language identification / Markov, Ilia; Nastase, Vivi; Strapparava, Carlo. - ELETTRONICO. - (2019), pp. 275-284. ((Intervento presentato al convegno Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications @ ACL 2019 tenutosi a Florence, Italy nel August 2019 [10.18653/v1/W19-4429].
Scheda prodotto non validato
I dati visualizzati non sono stati ancora sottoposti a validazione formale da parte dello Staff di IRIS, ma sono stati ugualmente trasmessi al Sito Docente Cineca (Loginmiur).
Titolo: | Anglicized words and misspelled cognates in native language identification | |
Autori: | Markov, Ilia; Nastase, Vivi; Strapparava, Carlo | |
Autori Unitn: | ||
Titolo del volume contenente il saggio: | Proceedings of the Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications | |
Luogo di edizione: | USA | |
Casa editrice: | Association for Computational Linguistics | |
Anno di pubblicazione: | 2019 | |
Codice identificativo Scopus: | 2-s2.0-85120977900 | |
Codice identificativo WOS: | WOS:000521943400029 | |
ISBN: | 978-1-950737-34-5 | |
Handle: | http://hdl.handle.net/11572/343610 | |
Citazione: | Anglicized words and misspelled cognates in native language identification / Markov, Ilia; Nastase, Vivi; Strapparava, Carlo. - ELETTRONICO. - (2019), pp. 275-284. ((Intervento presentato al convegno Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications @ ACL 2019 tenutosi a Florence, Italy nel August 2019 [10.18653/v1/W19-4429]. | |
Appare nelle tipologie: | 04.1 Saggio in atti di convegno (Paper in proceedings) |