Dissertation/ Thesis

Enabling Structured Navigation of Longform Spoken Dialog with Automatic Summarization

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Enabling Structured Navigation of Longform Spoken Dialog with Automatic Summarization
Συγγραφείς: Li, Daniel
Έτος έκδοσης: 2022
Συλλογή: Columbia University: Academic Commons
Θεματικοί όροι: Computer science, Speech processing systems--Computer programs, Dialogue, Podcasts, Information technology, Electronic information resource searching
Περιγραφή: Longform spoken dialog is a rich source of information that is present in all facets of everyday life, taking the form of podcasts, debates, and interviews; these mediums contain important topics ranging from healthcare and diversity to current events, economics and politics. Individuals need to digest informative content to know how to vote, decide how to stay safe from COVID-19, and how to increase diversity in the workplace. Unfortunately compared to text, spoken dialog can be challenging to consume as it is slower than reading and difficult to skim or navigate. Although an individual may be interested in a given topic, they may be unwilling to commit the required time necessary to consume long form auditory media given the uncertainty as to whether such content will live up to their expectations. Clearly, there exists a need to provide access to the information spoken dialog provides in a manner through which individuals can quickly and intuitively access areas of interest without investing large amounts of time. From Human Computer Interaction, we apply the idea of information foraging, which theorizes how people browse and navigate to satisfy an information need, to the longform spoken dialog domain. Information foraging states that people do not browse linearly. Rather people “forage” for information similar to how animals sniff around for food, scanning from area to area, constantly deciding whether to keep investigating their current area or to move on to greener pastures. This is an instance of the classic breadth vs. depth dilemma. People rely on perceived structure and information cues to make these decisions. Unfortunately speech, either spoken or transcribed, is unstructured and lacks information cues, making it difficult for users to browse and navigate. We create a longform spoken dialog browsing system that utilizes automatic summarization and speech modeling to structure longform dialog to present information in a manner that is both intuitive and flexible towards different user browsing needs. ...
Τύπος εγγράφου: thesis
Γλώσσα: English
DOI: 10.7916/cg58-qj86
Διαθεσιμότητα: https://doi.org/10.7916/cg58-qj86
Αριθμός Καταχώρησης: edsbas.BAE543DA
Βάση Δεδομένων: BASE
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://doi.org/10.7916/cg58-qj86#
    Name: EDS - BASE (ns324271)
    Category: fullText
    Text: View record from BASE
Header DbId: edsbas
DbLabel: BASE
An: edsbas.BAE543DA
RelevancyScore: 842
AccessLevel: 3
PubType: Dissertation/ Thesis
PubTypeId: dissertation
PreciseRelevancyScore: 842.276306152344
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Enabling Structured Navigation of Longform Spoken Dialog with Automatic Summarization
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Li%2C+Daniel%22">Li, Daniel</searchLink>
– Name: DatePubCY
  Label: Publication Year
  Group: Date
  Data: 2022
– Name: Subset
  Label: Collection
  Group: HoldingsInfo
  Data: Columbia University: Academic Commons
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Computer+science%22">Computer science</searchLink><br /><searchLink fieldCode="DE" term="%22Speech+processing+systems--Computer+programs%22">Speech processing systems--Computer programs</searchLink><br /><searchLink fieldCode="DE" term="%22Dialogue%22">Dialogue</searchLink><br /><searchLink fieldCode="DE" term="%22Podcasts%22">Podcasts</searchLink><br /><searchLink fieldCode="DE" term="%22Information+technology%22">Information technology</searchLink><br /><searchLink fieldCode="DE" term="%22Electronic+information+resource+searching%22">Electronic information resource searching</searchLink>
– Name: Abstract
  Label: Description
  Group: Ab
  Data: Longform spoken dialog is a rich source of information that is present in all facets of everyday life, taking the form of podcasts, debates, and interviews; these mediums contain important topics ranging from healthcare and diversity to current events, economics and politics. Individuals need to digest informative content to know how to vote, decide how to stay safe from COVID-19, and how to increase diversity in the workplace. Unfortunately compared to text, spoken dialog can be challenging to consume as it is slower than reading and difficult to skim or navigate. Although an individual may be interested in a given topic, they may be unwilling to commit the required time necessary to consume long form auditory media given the uncertainty as to whether such content will live up to their expectations. Clearly, there exists a need to provide access to the information spoken dialog provides in a manner through which individuals can quickly and intuitively access areas of interest without investing large amounts of time. From Human Computer Interaction, we apply the idea of information foraging, which theorizes how people browse and navigate to satisfy an information need, to the longform spoken dialog domain. Information foraging states that people do not browse linearly. Rather people “forage” for information similar to how animals sniff around for food, scanning from area to area, constantly deciding whether to keep investigating their current area or to move on to greener pastures. This is an instance of the classic breadth vs. depth dilemma. People rely on perceived structure and information cues to make these decisions. Unfortunately speech, either spoken or transcribed, is unstructured and lacks information cues, making it difficult for users to browse and navigate. We create a longform spoken dialog browsing system that utilizes automatic summarization and speech modeling to structure longform dialog to present information in a manner that is both intuitive and flexible towards different user browsing needs. ...
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: thesis
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.7916/cg58-qj86
– Name: URL
  Label: Availability
  Group: URL
  Data: https://doi.org/10.7916/cg58-qj86
– Name: AN
  Label: Accession Number
  Group: ID
  Data: edsbas.BAE543DA
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.BAE543DA
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.7916/cg58-qj86
    Languages:
      – Text: English
    Subjects:
      – SubjectFull: Computer science
        Type: general
      – SubjectFull: Speech processing systems--Computer programs
        Type: general
      – SubjectFull: Dialogue
        Type: general
      – SubjectFull: Podcasts
        Type: general
      – SubjectFull: Information technology
        Type: general
      – SubjectFull: Electronic information resource searching
        Type: general
    Titles:
      – TitleFull: Enabling Structured Navigation of Longform Spoken Dialog with Automatic Summarization
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Li, Daniel
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2022
          Identifiers:
            – Type: issn-locals
              Value: edsbas
            – Type: issn-locals
              Value: edsbas.oa
ResultId 1