Dissertation/ Thesis

Unsupervised Morphological Segmentation and Part-of-Speech Tagging for Low-Resource Scenarios

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Unsupervised Morphological Segmentation and Part-of-Speech Tagging for Low-Resource Scenarios
Συγγραφείς: Eskander, Ramy
Έτος έκδοσης: 2021
Συλλογή: Columbia University: Academic Commons
Θεματικοί όροι: Computer science, Speech processing systems--Computer programs, Automatic speech recognition--Computer programs, Machine translating, Question-answering systems
Περιγραφή: With the high cost of manually labeling data and the increasing interest in low-resource languages, for which human annotators might not be even available, unsupervised approaches have become essential for processing a typologically diverse set of languages, whether high-resource or low-resource. In this work, we propose new fully unsupervised approaches for two tasks in morphology: unsupervised morphological segmentation and unsupervised cross-lingual part-of-speech (POS) tagging, which have been two essential subtasks for several downstream NLP applications, such as machine translation, speech recognition, information extraction and question answering. We propose a new unsupervised morphological-segmentation approach that utilizes Adaptor Grammars (AGs), nonparametric Bayesian models that generalize probabilistic context-free grammars (PCFGs), where a PCFG models word structure in the task of morphological segmentation. We implement the approach as a publicly available morphological-segmentation framework, MorphAGram, that enables unsupervised morphological segmentation through the use of several proposed language-independent grammars. In addition, the framework allows for the use of scholar knowledge, when available, in the form of affixes that can be seeded into the grammars. The framework handles the cases when the scholar-seeded knowledge is either generated from language resources, possibly by someone who does not know the language, as weak linguistic priors, or generated by an expert in the underlying language as strong linguistic priors. Another form of linguistic priors is the design of a grammar that models language-dependent specifications. We also propose a fully unsupervised learning setting that approximates the effect of scholar-seeded knowledge through self-training. Moreover, since there is no single grammar that works best across all languages, we propose an approach that picks a nearly optimal configuration (a learning setting and a grammar) for an unseen language, a language that is not part ...
Τύπος εγγράφου: thesis
Γλώσσα: English
DOI: 10.7916/d8-jd2d-9p51
Διαθεσιμότητα: https://doi.org/10.7916/d8-jd2d-9p51
Αριθμός Καταχώρησης: edsbas.AD2E7ED4
Βάση Δεδομένων: BASE
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://doi.org/10.7916/d8-jd2d-9p51#
    Name: EDS - BASE (ns324271)
    Category: fullText
    Text: View record from BASE
Header DbId: edsbas
DbLabel: BASE
An: edsbas.AD2E7ED4
RelevancyScore: 828
AccessLevel: 3
PubType: Dissertation/ Thesis
PubTypeId: dissertation
PreciseRelevancyScore: 828.347900390625
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Unsupervised Morphological Segmentation and Part-of-Speech Tagging for Low-Resource Scenarios
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Eskander%2C+Ramy%22">Eskander, Ramy</searchLink>
– Name: DatePubCY
  Label: Publication Year
  Group: Date
  Data: 2021
– Name: Subset
  Label: Collection
  Group: HoldingsInfo
  Data: Columbia University: Academic Commons
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Computer+science%22">Computer science</searchLink><br /><searchLink fieldCode="DE" term="%22Speech+processing+systems--Computer+programs%22">Speech processing systems--Computer programs</searchLink><br /><searchLink fieldCode="DE" term="%22Automatic+speech+recognition--Computer+programs%22">Automatic speech recognition--Computer programs</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+translating%22">Machine translating</searchLink><br /><searchLink fieldCode="DE" term="%22Question-answering+systems%22">Question-answering systems</searchLink>
– Name: Abstract
  Label: Description
  Group: Ab
  Data: With the high cost of manually labeling data and the increasing interest in low-resource languages, for which human annotators might not be even available, unsupervised approaches have become essential for processing a typologically diverse set of languages, whether high-resource or low-resource. In this work, we propose new fully unsupervised approaches for two tasks in morphology: unsupervised morphological segmentation and unsupervised cross-lingual part-of-speech (POS) tagging, which have been two essential subtasks for several downstream NLP applications, such as machine translation, speech recognition, information extraction and question answering. We propose a new unsupervised morphological-segmentation approach that utilizes Adaptor Grammars (AGs), nonparametric Bayesian models that generalize probabilistic context-free grammars (PCFGs), where a PCFG models word structure in the task of morphological segmentation. We implement the approach as a publicly available morphological-segmentation framework, MorphAGram, that enables unsupervised morphological segmentation through the use of several proposed language-independent grammars. In addition, the framework allows for the use of scholar knowledge, when available, in the form of affixes that can be seeded into the grammars. The framework handles the cases when the scholar-seeded knowledge is either generated from language resources, possibly by someone who does not know the language, as weak linguistic priors, or generated by an expert in the underlying language as strong linguistic priors. Another form of linguistic priors is the design of a grammar that models language-dependent specifications. We also propose a fully unsupervised learning setting that approximates the effect of scholar-seeded knowledge through self-training. Moreover, since there is no single grammar that works best across all languages, we propose an approach that picks a nearly optimal configuration (a learning setting and a grammar) for an unseen language, a language that is not part ...
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: thesis
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.7916/d8-jd2d-9p51
– Name: URL
  Label: Availability
  Group: URL
  Data: https://doi.org/10.7916/d8-jd2d-9p51
– Name: AN
  Label: Accession Number
  Group: ID
  Data: edsbas.AD2E7ED4
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.AD2E7ED4
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.7916/d8-jd2d-9p51
    Languages:
      – Text: English
    Subjects:
      – SubjectFull: Computer science
        Type: general
      – SubjectFull: Speech processing systems--Computer programs
        Type: general
      – SubjectFull: Automatic speech recognition--Computer programs
        Type: general
      – SubjectFull: Machine translating
        Type: general
      – SubjectFull: Question-answering systems
        Type: general
    Titles:
      – TitleFull: Unsupervised Morphological Segmentation and Part-of-Speech Tagging for Low-Resource Scenarios
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Eskander, Ramy
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2021
          Identifiers:
            – Type: issn-locals
              Value: edsbas
            – Type: issn-locals
              Value: edsbas.oa
ResultId 1