Dissertation/ Thesis
Unsupervised Morphological Segmentation and Part-of-Speech Tagging for Low-Resource Scenarios
| Τίτλος: | Unsupervised Morphological Segmentation and Part-of-Speech Tagging for Low-Resource Scenarios |
|---|---|
| Συγγραφείς: | Eskander, Ramy |
| Έτος έκδοσης: | 2021 |
| Συλλογή: | Columbia University: Academic Commons |
| Θεματικοί όροι: | Computer science, Speech processing systems--Computer programs, Automatic speech recognition--Computer programs, Machine translating, Question-answering systems |
| Περιγραφή: | With the high cost of manually labeling data and the increasing interest in low-resource languages, for which human annotators might not be even available, unsupervised approaches have become essential for processing a typologically diverse set of languages, whether high-resource or low-resource. In this work, we propose new fully unsupervised approaches for two tasks in morphology: unsupervised morphological segmentation and unsupervised cross-lingual part-of-speech (POS) tagging, which have been two essential subtasks for several downstream NLP applications, such as machine translation, speech recognition, information extraction and question answering. We propose a new unsupervised morphological-segmentation approach that utilizes Adaptor Grammars (AGs), nonparametric Bayesian models that generalize probabilistic context-free grammars (PCFGs), where a PCFG models word structure in the task of morphological segmentation. We implement the approach as a publicly available morphological-segmentation framework, MorphAGram, that enables unsupervised morphological segmentation through the use of several proposed language-independent grammars. In addition, the framework allows for the use of scholar knowledge, when available, in the form of affixes that can be seeded into the grammars. The framework handles the cases when the scholar-seeded knowledge is either generated from language resources, possibly by someone who does not know the language, as weak linguistic priors, or generated by an expert in the underlying language as strong linguistic priors. Another form of linguistic priors is the design of a grammar that models language-dependent specifications. We also propose a fully unsupervised learning setting that approximates the effect of scholar-seeded knowledge through self-training. Moreover, since there is no single grammar that works best across all languages, we propose an approach that picks a nearly optimal configuration (a learning setting and a grammar) for an unseen language, a language that is not part ... |
| Τύπος εγγράφου: | thesis |
| Γλώσσα: | English |
| DOI: | 10.7916/d8-jd2d-9p51 |
| Διαθεσιμότητα: | https://doi.org/10.7916/d8-jd2d-9p51 |
| Αριθμός Καταχώρησης: | edsbas.AD2E7ED4 |
| Βάση Δεδομένων: | BASE |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://doi.org/10.7916/d8-jd2d-9p51# Name: EDS - BASE (ns324271) Category: fullText Text: View record from BASE |
|---|---|
| Header | DbId: edsbas DbLabel: BASE An: edsbas.AD2E7ED4 RelevancyScore: 828 AccessLevel: 3 PubType: Dissertation/ Thesis PubTypeId: dissertation PreciseRelevancyScore: 828.347900390625 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Unsupervised Morphological Segmentation and Part-of-Speech Tagging for Low-Resource Scenarios – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Eskander%2C+Ramy%22">Eskander, Ramy</searchLink> – Name: DatePubCY Label: Publication Year Group: Date Data: 2021 – Name: Subset Label: Collection Group: HoldingsInfo Data: Columbia University: Academic Commons – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Computer+science%22">Computer science</searchLink><br /><searchLink fieldCode="DE" term="%22Speech+processing+systems--Computer+programs%22">Speech processing systems--Computer programs</searchLink><br /><searchLink fieldCode="DE" term="%22Automatic+speech+recognition--Computer+programs%22">Automatic speech recognition--Computer programs</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+translating%22">Machine translating</searchLink><br /><searchLink fieldCode="DE" term="%22Question-answering+systems%22">Question-answering systems</searchLink> – Name: Abstract Label: Description Group: Ab Data: With the high cost of manually labeling data and the increasing interest in low-resource languages, for which human annotators might not be even available, unsupervised approaches have become essential for processing a typologically diverse set of languages, whether high-resource or low-resource. In this work, we propose new fully unsupervised approaches for two tasks in morphology: unsupervised morphological segmentation and unsupervised cross-lingual part-of-speech (POS) tagging, which have been two essential subtasks for several downstream NLP applications, such as machine translation, speech recognition, information extraction and question answering. We propose a new unsupervised morphological-segmentation approach that utilizes Adaptor Grammars (AGs), nonparametric Bayesian models that generalize probabilistic context-free grammars (PCFGs), where a PCFG models word structure in the task of morphological segmentation. We implement the approach as a publicly available morphological-segmentation framework, MorphAGram, that enables unsupervised morphological segmentation through the use of several proposed language-independent grammars. In addition, the framework allows for the use of scholar knowledge, when available, in the form of affixes that can be seeded into the grammars. The framework handles the cases when the scholar-seeded knowledge is either generated from language resources, possibly by someone who does not know the language, as weak linguistic priors, or generated by an expert in the underlying language as strong linguistic priors. Another form of linguistic priors is the design of a grammar that models language-dependent specifications. We also propose a fully unsupervised learning setting that approximates the effect of scholar-seeded knowledge through self-training. Moreover, since there is no single grammar that works best across all languages, we propose an approach that picks a nearly optimal configuration (a learning setting and a grammar) for an unseen language, a language that is not part ... – Name: TypeDocument Label: Document Type Group: TypDoc Data: thesis – Name: Language Label: Language Group: Lang Data: English – Name: DOI Label: DOI Group: ID Data: 10.7916/d8-jd2d-9p51 – Name: URL Label: Availability Group: URL Data: https://doi.org/10.7916/d8-jd2d-9p51 – Name: AN Label: Accession Number Group: ID Data: edsbas.AD2E7ED4 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.AD2E7ED4 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.7916/d8-jd2d-9p51 Languages: – Text: English Subjects: – SubjectFull: Computer science Type: general – SubjectFull: Speech processing systems--Computer programs Type: general – SubjectFull: Automatic speech recognition--Computer programs Type: general – SubjectFull: Machine translating Type: general – SubjectFull: Question-answering systems Type: general Titles: – TitleFull: Unsupervised Morphological Segmentation and Part-of-Speech Tagging for Low-Resource Scenarios Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Eskander, Ramy IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2021 Identifiers: – Type: issn-locals Value: edsbas – Type: issn-locals Value: edsbas.oa |
| ResultId | 1 |