Dissertation/ Thesis

Unsupervised Morphological Segmentation and Part-of-Speech Tagging for Low-Resource Scenarios

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Unsupervised Morphological Segmentation and Part-of-Speech Tagging for Low-Resource Scenarios
Συγγραφείς: Eskander, Ramy
Έτος έκδοσης: 2021
Συλλογή: Columbia University: Academic Commons
Θεματικοί όροι: Computer science, Speech processing systems--Computer programs, Automatic speech recognition--Computer programs, Machine translating, Question-answering systems
Περιγραφή: With the high cost of manually labeling data and the increasing interest in low-resource languages, for which human annotators might not be even available, unsupervised approaches have become essential for processing a typologically diverse set of languages, whether high-resource or low-resource. In this work, we propose new fully unsupervised approaches for two tasks in morphology: unsupervised morphological segmentation and unsupervised cross-lingual part-of-speech (POS) tagging, which have been two essential subtasks for several downstream NLP applications, such as machine translation, speech recognition, information extraction and question answering. We propose a new unsupervised morphological-segmentation approach that utilizes Adaptor Grammars (AGs), nonparametric Bayesian models that generalize probabilistic context-free grammars (PCFGs), where a PCFG models word structure in the task of morphological segmentation. We implement the approach as a publicly available morphological-segmentation framework, MorphAGram, that enables unsupervised morphological segmentation through the use of several proposed language-independent grammars. In addition, the framework allows for the use of scholar knowledge, when available, in the form of affixes that can be seeded into the grammars. The framework handles the cases when the scholar-seeded knowledge is either generated from language resources, possibly by someone who does not know the language, as weak linguistic priors, or generated by an expert in the underlying language as strong linguistic priors. Another form of linguistic priors is the design of a grammar that models language-dependent specifications. We also propose a fully unsupervised learning setting that approximates the effect of scholar-seeded knowledge through self-training. Moreover, since there is no single grammar that works best across all languages, we propose an approach that picks a nearly optimal configuration (a learning setting and a grammar) for an unseen language, a language that is not part ...
Τύπος εγγράφου: thesis
Γλώσσα: English
DOI: 10.7916/d8-jd2d-9p51
Διαθεσιμότητα: https://doi.org/10.7916/d8-jd2d-9p51
Αριθμός Καταχώρησης: edsbas.AD2E7ED4
Βάση Δεδομένων: BASE