Dissertation/ Thesis
Harmonizing Audio and Human Interaction: Enhancement, Analysis, and Application of Audio Signals via Machine Learning Approaches
| Τίτλος: | Harmonizing Audio and Human Interaction: Enhancement, Analysis, and Application of Audio Signals via Machine Learning Approaches |
|---|---|
| Συγγραφείς: | Xu, Ruilin |
| Έτος έκδοσης: | 2024 |
| Συλλογή: | Columbia University: Academic Commons |
| Θεματικοί όροι: | Computer science, Machine learning, Three-dimensional imaging, Automatic speech recognition--Computer programs, Ambient sounds, Sound--Reverberation, Computer simulation, Dance--Computer programs, Computer animation--Computer programs |
| Περιγραφή: | In this thesis, we tackle key challenges in processing audio signals, specifically focusing on speech and music. These signals are crucial for human interaction with both the environment and machines. Our research addresses three core topics: speech denoising, speech dereverberation, and music-dance generation, each of which plays a vital role in enhancing the harmony between audio and human interaction. Leveraging machine learning and human-centric approaches inspired by classical algorithms, we develop methods to mitigate common audio degradations, such as additive noise and multiplicative reverberation, delivering high-quality audio suitable for human use and applications. Furthermore, we introduce a real-time, music-responsive system for generating 3D dance animations, advancing the integration of audio signals with human engagement. The first focus of our thesis is the elimination of additive noise from audio signals by focusing on short pauses, or silent intervals, in human speech. These brief pauses provide key insights into the noise profile, enabling our model to dynamically reduce ambient noise from speech. Tested across diverse datasets, our method outperforms traditional and audiovisual denoising techniques, showcasing its effectiveness and adaptability across different languages and even musical contexts. In the second work of our research, we address reverberation removal from audio signals, a task traditionally reliant on knowing the environment's exact impulse response—a requirement often impractical in real-world settings. Our novel solution combines the strengths of classical and learning-based approaches, tailored for online communication contexts. This human-centric method includes a one-time personalization step, adapting to specific environments and human speakers. The two-stage model, integrating feature-based Wiener deconvolution and network refinement, has shown through extensive experiments to outperform current methods, both in effectiveness and user preference. Transitioning from ... |
| Τύπος εγγράφου: | thesis |
| Γλώσσα: | English |
| DOI: | 10.7916/7bfv-6g22 |
| Διαθεσιμότητα: | https://doi.org/10.7916/7bfv-6g22 |
| Αριθμός Καταχώρησης: | edsbas.94D1AB6 |
| Βάση Δεδομένων: | BASE |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://doi.org/10.7916/7bfv-6g22# Name: EDS - BASE (ns324271) Category: fullText Text: View record from BASE |
|---|---|
| Header | DbId: edsbas DbLabel: BASE An: edsbas.94D1AB6 RelevancyScore: 872 AccessLevel: 3 PubType: Dissertation/ Thesis PubTypeId: dissertation PreciseRelevancyScore: 871.7080078125 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Harmonizing Audio and Human Interaction: Enhancement, Analysis, and Application of Audio Signals via Machine Learning Approaches – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Xu%2C+Ruilin%22">Xu, Ruilin</searchLink> – Name: DatePubCY Label: Publication Year Group: Date Data: 2024 – Name: Subset Label: Collection Group: HoldingsInfo Data: Columbia University: Academic Commons – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Computer+science%22">Computer science</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Three-dimensional+imaging%22">Three-dimensional imaging</searchLink><br /><searchLink fieldCode="DE" term="%22Automatic+speech+recognition--Computer+programs%22">Automatic speech recognition--Computer programs</searchLink><br /><searchLink fieldCode="DE" term="%22Ambient+sounds%22">Ambient sounds</searchLink><br /><searchLink fieldCode="DE" term="%22Sound--Reverberation%22">Sound--Reverberation</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+simulation%22">Computer simulation</searchLink><br /><searchLink fieldCode="DE" term="%22Dance--Computer+programs%22">Dance--Computer programs</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+animation--Computer+programs%22">Computer animation--Computer programs</searchLink> – Name: Abstract Label: Description Group: Ab Data: In this thesis, we tackle key challenges in processing audio signals, specifically focusing on speech and music. These signals are crucial for human interaction with both the environment and machines. Our research addresses three core topics: speech denoising, speech dereverberation, and music-dance generation, each of which plays a vital role in enhancing the harmony between audio and human interaction. Leveraging machine learning and human-centric approaches inspired by classical algorithms, we develop methods to mitigate common audio degradations, such as additive noise and multiplicative reverberation, delivering high-quality audio suitable for human use and applications. Furthermore, we introduce a real-time, music-responsive system for generating 3D dance animations, advancing the integration of audio signals with human engagement. The first focus of our thesis is the elimination of additive noise from audio signals by focusing on short pauses, or silent intervals, in human speech. These brief pauses provide key insights into the noise profile, enabling our model to dynamically reduce ambient noise from speech. Tested across diverse datasets, our method outperforms traditional and audiovisual denoising techniques, showcasing its effectiveness and adaptability across different languages and even musical contexts. In the second work of our research, we address reverberation removal from audio signals, a task traditionally reliant on knowing the environment's exact impulse response—a requirement often impractical in real-world settings. Our novel solution combines the strengths of classical and learning-based approaches, tailored for online communication contexts. This human-centric method includes a one-time personalization step, adapting to specific environments and human speakers. The two-stage model, integrating feature-based Wiener deconvolution and network refinement, has shown through extensive experiments to outperform current methods, both in effectiveness and user preference. Transitioning from ... – Name: TypeDocument Label: Document Type Group: TypDoc Data: thesis – Name: Language Label: Language Group: Lang Data: English – Name: DOI Label: DOI Group: ID Data: 10.7916/7bfv-6g22 – Name: URL Label: Availability Group: URL Data: https://doi.org/10.7916/7bfv-6g22 – Name: AN Label: Accession Number Group: ID Data: edsbas.94D1AB6 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.94D1AB6 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.7916/7bfv-6g22 Languages: – Text: English Subjects: – SubjectFull: Computer science Type: general – SubjectFull: Machine learning Type: general – SubjectFull: Three-dimensional imaging Type: general – SubjectFull: Automatic speech recognition--Computer programs Type: general – SubjectFull: Ambient sounds Type: general – SubjectFull: Sound--Reverberation Type: general – SubjectFull: Computer simulation Type: general – SubjectFull: Dance--Computer programs Type: general – SubjectFull: Computer animation--Computer programs Type: general Titles: – TitleFull: Harmonizing Audio and Human Interaction: Enhancement, Analysis, and Application of Audio Signals via Machine Learning Approaches Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Xu, Ruilin IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2024 Identifiers: – Type: issn-locals Value: edsbas – Type: issn-locals Value: edsbas.oa |
| ResultId | 1 |