Lexical semantic changes spanning centuries can reveal the complicated developing process of language and social culture. In recent years, natural language processing (NLP) methods have been applied in this field to provide insight into the diachronic frequency change for word senses from large-scale historical corpus, for instance, analyzing which senses appear, increase, or decrease at which times. However, there is still a lack of Chinese diachronic corpus and dataset in this field to support supervised learning and text mining, and at the method level, few existing works analyze the Chinese semantic changes at the level of morpheme. This paper constructs a diachronic Chinese dataset for semantic tracking applications spanning 3000 years and extends the existing framework to the level of Chinese characters and morphemes, which contains four main steps of contextual sense representation, sense identification, morpheme sense mining, and diachronic semantic change representation. The experiment shows the effectiveness of our method in each step. Finally, in an interesting statistic, we discover the strong positive correlation of frequency and changing trend between monosyllabic word sense and the corresponding morpheme.

Diachronic Semantic Tracking for Chinese Words and Morphemes over Centuries / Chi, Y., Giunchiglia, F., Xu, H.. - In: ELECTRONICS. - ISSN 2079-9292. - 13:9(2024). (2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT 2024 bra 2024) [10.3390/electronics13091728].

Diachronic Semantic Tracking for Chinese Words and Morphemes over Centuries

Giunchiglia F.;
2024-01-01

Abstract

Lexical semantic changes spanning centuries can reveal the complicated developing process of language and social culture. In recent years, natural language processing (NLP) methods have been applied in this field to provide insight into the diachronic frequency change for word senses from large-scale historical corpus, for instance, analyzing which senses appear, increase, or decrease at which times. However, there is still a lack of Chinese diachronic corpus and dataset in this field to support supervised learning and text mining, and at the method level, few existing works analyze the Chinese semantic changes at the level of morpheme. This paper constructs a diachronic Chinese dataset for semantic tracking applications spanning 3000 years and extends the existing framework to the level of Chinese characters and morphemes, which contains four main steps of contextual sense representation, sense identification, morpheme sense mining, and diachronic semantic change representation. The experiment shows the effectiveness of our method in each step. Finally, in an interesting statistic, we discover the strong positive correlation of frequency and changing trend between monosyllabic word sense and the corresponding morpheme.
2024
2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT 2024
ST ALBAN-ANLAGE 66, CH-4052 BASEL, SWITZERLAND
Multidisciplinary Digital Publishing Institute (MDPI)
Chi, Y.; Giunchiglia, F.; Xu, H.
Diachronic Semantic Tracking for Chinese Words and Morphemes over Centuries / Chi, Y., Giunchiglia, F., Xu, H.. - In: ELECTRONICS. - ISSN 2079-9292. - 13:9(2024). (2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT 2024 bra 2024) [10.3390/electronics13091728].
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/464122
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 5
  • ???jsp.display-item.citation.isi??? 6
  • OpenAlex 4
social impact