Dissertation/ Thesis

Harmonizing Audio and Human Interaction: Enhancement, Analysis, and Application of Audio Signals via Machine Learning Approaches

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Harmonizing Audio and Human Interaction: Enhancement, Analysis, and Application of Audio Signals via Machine Learning Approaches
Συγγραφείς: Xu, Ruilin
Έτος έκδοσης: 2024
Συλλογή: Columbia University: Academic Commons
Θεματικοί όροι: Computer science, Machine learning, Three-dimensional imaging, Automatic speech recognition--Computer programs, Ambient sounds, Sound--Reverberation, Computer simulation, Dance--Computer programs, Computer animation--Computer programs
Περιγραφή: In this thesis, we tackle key challenges in processing audio signals, specifically focusing on speech and music. These signals are crucial for human interaction with both the environment and machines. Our research addresses three core topics: speech denoising, speech dereverberation, and music-dance generation, each of which plays a vital role in enhancing the harmony between audio and human interaction. Leveraging machine learning and human-centric approaches inspired by classical algorithms, we develop methods to mitigate common audio degradations, such as additive noise and multiplicative reverberation, delivering high-quality audio suitable for human use and applications. Furthermore, we introduce a real-time, music-responsive system for generating 3D dance animations, advancing the integration of audio signals with human engagement. The first focus of our thesis is the elimination of additive noise from audio signals by focusing on short pauses, or silent intervals, in human speech. These brief pauses provide key insights into the noise profile, enabling our model to dynamically reduce ambient noise from speech. Tested across diverse datasets, our method outperforms traditional and audiovisual denoising techniques, showcasing its effectiveness and adaptability across different languages and even musical contexts. In the second work of our research, we address reverberation removal from audio signals, a task traditionally reliant on knowing the environment's exact impulse response—a requirement often impractical in real-world settings. Our novel solution combines the strengths of classical and learning-based approaches, tailored for online communication contexts. This human-centric method includes a one-time personalization step, adapting to specific environments and human speakers. The two-stage model, integrating feature-based Wiener deconvolution and network refinement, has shown through extensive experiments to outperform current methods, both in effectiveness and user preference. Transitioning from ...
Τύπος εγγράφου: thesis
Γλώσσα: English
DOI: 10.7916/7bfv-6g22
Διαθεσιμότητα: https://doi.org/10.7916/7bfv-6g22
Αριθμός Καταχώρησης: edsbas.94D1AB6
Βάση Δεδομένων: BASE
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://doi.org/10.7916/7bfv-6g22#
    Name: EDS - BASE (ns324271)
    Category: fullText
    Text: View record from BASE
Header DbId: edsbas
DbLabel: BASE
An: edsbas.94D1AB6
RelevancyScore: 872
AccessLevel: 3
PubType: Dissertation/ Thesis
PubTypeId: dissertation
PreciseRelevancyScore: 871.7080078125
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Harmonizing Audio and Human Interaction: Enhancement, Analysis, and Application of Audio Signals via Machine Learning Approaches
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Xu%2C+Ruilin%22">Xu, Ruilin</searchLink>
– Name: DatePubCY
  Label: Publication Year
  Group: Date
  Data: 2024
– Name: Subset
  Label: Collection
  Group: HoldingsInfo
  Data: Columbia University: Academic Commons
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Computer+science%22">Computer science</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Three-dimensional+imaging%22">Three-dimensional imaging</searchLink><br /><searchLink fieldCode="DE" term="%22Automatic+speech+recognition--Computer+programs%22">Automatic speech recognition--Computer programs</searchLink><br /><searchLink fieldCode="DE" term="%22Ambient+sounds%22">Ambient sounds</searchLink><br /><searchLink fieldCode="DE" term="%22Sound--Reverberation%22">Sound--Reverberation</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+simulation%22">Computer simulation</searchLink><br /><searchLink fieldCode="DE" term="%22Dance--Computer+programs%22">Dance--Computer programs</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+animation--Computer+programs%22">Computer animation--Computer programs</searchLink>
– Name: Abstract
  Label: Description
  Group: Ab
  Data: In this thesis, we tackle key challenges in processing audio signals, specifically focusing on speech and music. These signals are crucial for human interaction with both the environment and machines. Our research addresses three core topics: speech denoising, speech dereverberation, and music-dance generation, each of which plays a vital role in enhancing the harmony between audio and human interaction. Leveraging machine learning and human-centric approaches inspired by classical algorithms, we develop methods to mitigate common audio degradations, such as additive noise and multiplicative reverberation, delivering high-quality audio suitable for human use and applications. Furthermore, we introduce a real-time, music-responsive system for generating 3D dance animations, advancing the integration of audio signals with human engagement. The first focus of our thesis is the elimination of additive noise from audio signals by focusing on short pauses, or silent intervals, in human speech. These brief pauses provide key insights into the noise profile, enabling our model to dynamically reduce ambient noise from speech. Tested across diverse datasets, our method outperforms traditional and audiovisual denoising techniques, showcasing its effectiveness and adaptability across different languages and even musical contexts. In the second work of our research, we address reverberation removal from audio signals, a task traditionally reliant on knowing the environment's exact impulse response—a requirement often impractical in real-world settings. Our novel solution combines the strengths of classical and learning-based approaches, tailored for online communication contexts. This human-centric method includes a one-time personalization step, adapting to specific environments and human speakers. The two-stage model, integrating feature-based Wiener deconvolution and network refinement, has shown through extensive experiments to outperform current methods, both in effectiveness and user preference. Transitioning from ...
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: thesis
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.7916/7bfv-6g22
– Name: URL
  Label: Availability
  Group: URL
  Data: https://doi.org/10.7916/7bfv-6g22
– Name: AN
  Label: Accession Number
  Group: ID
  Data: edsbas.94D1AB6
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.94D1AB6
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.7916/7bfv-6g22
    Languages:
      – Text: English
    Subjects:
      – SubjectFull: Computer science
        Type: general
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: Three-dimensional imaging
        Type: general
      – SubjectFull: Automatic speech recognition--Computer programs
        Type: general
      – SubjectFull: Ambient sounds
        Type: general
      – SubjectFull: Sound--Reverberation
        Type: general
      – SubjectFull: Computer simulation
        Type: general
      – SubjectFull: Dance--Computer programs
        Type: general
      – SubjectFull: Computer animation--Computer programs
        Type: general
    Titles:
      – TitleFull: Harmonizing Audio and Human Interaction: Enhancement, Analysis, and Application of Audio Signals via Machine Learning Approaches
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Xu, Ruilin
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2024
          Identifiers:
            – Type: issn-locals
              Value: edsbas
            – Type: issn-locals
              Value: edsbas.oa
ResultId 1