Batsuren, Khuyagbaatar (2018) Understanding and Exploiting Language Diversity. PhD thesis, University of Trento.
PDF (Disclaimer of Khuyagbaatar Batsuren) - Disclaimer Restricted to Repository staff only until 9999. 1092Kb | ||
| PDF (Doctoral Thesis of Khuyagbaatar Batsuren) - Doctoral Thesis 6Mb |
Abstract
Languages are well known to be diverse on all structural levels, from the smallest (phonemic) to the broadest (pragmatic). We propose a set of formal, quantitative measures for the language diversity of linguistic phenomena, the resource incompleteness, and resource incorrectness. We apply all these measures to lexical semantics where we show how evidence of a high degree of universality within a given language set can be used to extend lexico-semantic resources in a precise, diversity-aware manner. We demonstrate our approach on several case studies: First is on polysemes and homographs among cases of lexical ambiguity. Contrarily to past research that focused solely on exploiting systematic polysemy, the notion of universality provides us with an automated method also capable of predicting irregular polysemes. Second is to automatically identify cognates from the existing lexical resource across different orthographies of genetically unrelated languages. Contrarily to past research that focused on detecting cognates from 225 concepts of Swadesh list, we captured 3.1 million cognate pairs across 40 different orthographies and 335 languages by exploiting the existing wordnet-like lexical resources.
Item Type: | Doctoral Thesis (PhD) |
---|---|
Doctoral School: | Information and Communication Technology |
PhD Cycle: | 31 |
Subjects: | Area 09 - Ingegneria industriale e dell'informazione > ING-INF/05 SISTEMI DI ELABORAZIONE DELLE INFORMAZIONI |
Uncontrolled Keywords: | Language diversity, computational lexical semantics, computational linguistics |
Repository Staff approval on: | 07 Dec 2018 10:36 |
Repository Staff Only: item control page