Academic Journal

Lightweight Real-Time Recurrent Models for Speech Enhancement and Automatic Speech Recognition.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Lightweight Real-Time Recurrent Models for Speech Enhancement and Automatic Speech Recognition.
Συγγραφείς: Dhahbi, Sami, Saleem, Nasir, Gunawan, Teddy Surya, Bourouis, Sami, Ali, Imad, Trigui, Aymen, Algarni, Abeer D.
Πηγή: International Journal of Interactive Multimedia & Artificial Intelligence; Jun2024, Vol. 8 Issue 6, p74-85, 12p
Θεματικοί όροι: Automatic speech recognition, Intelligibility of speech, Speech enhancement, Artificial neural networks, Recurrent neural networks, Speech
Περίληψη: Traditional recurrent neural networks (RNNs) encounter difficulty in capturing long-term temporal dependencies. However, lightweight recurrent models for speech enhancement are important to improve noisy speech, while being computationally efficient and able to capture long-term temporal dependencies efficiently. This study proposes a lightweight hourglass-shaped model for speech enhancement (SE) and automatic speech recognition (ASR). Simple recurrent units (SRU) with skip connections are implemented where attention gates are added to the skip connections, highlighting the important features and spectral regions. The model operates without relying on future information that is well-suited for real-time processing. Combined acoustic features and two training objectives are estimated. Experimental evaluations using the short time speech intelligibility (STOI), perceptual evaluation of speech quality (PESQ), and word error rates (WERs) indicate better intelligibility, perceptual quality, and word recognition rates. The composite measures further confirm the performance of residual noise and speech distortion. With the TIMIT database, the proposed model improves the STOI and PESQ by 16.21% and 0.69 (31.1%) whereas with the LibriSpeech database, the model improves STOI by 16.41% and PESQ by 0.71 (32.9%) over the noisy speech. Further, our model outperforms other deep neural networks (DNNs) in seen and unseen conditions. The ASR performance is measured using the Kaldi toolkit and achieves 15.13% WERs in noisy backgrounds. [ABSTRACT FROM AUTHOR]
Copyright of International Journal of Interactive Multimedia & Artificial Intelligence is the property of Universidad Internacional de La Rioja (UNIR) and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Βάση Δεδομένων: Complementary Index
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://resolver.ebsco.com/c/fiv2js/result?sid=EBSCO:edb&genre=article&issn=19891660&ISBN=&volume=8&issue=6&date=20240601&spage=74&pages=74-85&title=International Journal of Interactive Multimedia & Artificial Intelligence&atitle=Lightweight%20Real-Time%20Recurrent%20Models%20for%20Speech%20Enhancement%20and%20Automatic%20Speech%20Recognition.&aulast=Dhahbi%2C%20Sami&id=DOI:10.9781/ijimai.2024.04.003
    Name: Full Text Finder (for New FTF UI) (ns324271)
    Category: fullText
    Text: Full Text Finder
    MouseOverText: Full Text Finder
Header DbId: edb
DbLabel: Complementary Index
An: 177785890
RelevancyScore: 966
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 965.707580566406
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Lightweight Real-Time Recurrent Models for Speech Enhancement and Automatic Speech Recognition.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Dhahbi%2C+Sami%22">Dhahbi, Sami</searchLink><br /><searchLink fieldCode="AR" term="%22Saleem%2C+Nasir%22">Saleem, Nasir</searchLink><br /><searchLink fieldCode="AR" term="%22Gunawan%2C+Teddy+Surya%22">Gunawan, Teddy Surya</searchLink><br /><searchLink fieldCode="AR" term="%22Bourouis%2C+Sami%22">Bourouis, Sami</searchLink><br /><searchLink fieldCode="AR" term="%22Ali%2C+Imad%22">Ali, Imad</searchLink><br /><searchLink fieldCode="AR" term="%22Trigui%2C+Aymen%22">Trigui, Aymen</searchLink><br /><searchLink fieldCode="AR" term="%22Algarni%2C+Abeer+D%2E%22">Algarni, Abeer D.</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: International Journal of Interactive Multimedia & Artificial Intelligence; Jun2024, Vol. 8 Issue 6, p74-85, 12p
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Automatic+speech+recognition%22">Automatic speech recognition</searchLink><br /><searchLink fieldCode="DE" term="%22Intelligibility+of+speech%22">Intelligibility of speech</searchLink><br /><searchLink fieldCode="DE" term="%22Speech+enhancement%22">Speech enhancement</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+neural+networks%22">Artificial neural networks</searchLink><br /><searchLink fieldCode="DE" term="%22Recurrent+neural+networks%22">Recurrent neural networks</searchLink><br /><searchLink fieldCode="DE" term="%22Speech%22">Speech</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Traditional recurrent neural networks (RNNs) encounter difficulty in capturing long-term temporal dependencies. However, lightweight recurrent models for speech enhancement are important to improve noisy speech, while being computationally efficient and able to capture long-term temporal dependencies efficiently. This study proposes a lightweight hourglass-shaped model for speech enhancement (SE) and automatic speech recognition (ASR). Simple recurrent units (SRU) with skip connections are implemented where attention gates are added to the skip connections, highlighting the important features and spectral regions. The model operates without relying on future information that is well-suited for real-time processing. Combined acoustic features and two training objectives are estimated. Experimental evaluations using the short time speech intelligibility (STOI), perceptual evaluation of speech quality (PESQ), and word error rates (WERs) indicate better intelligibility, perceptual quality, and word recognition rates. The composite measures further confirm the performance of residual noise and speech distortion. With the TIMIT database, the proposed model improves the STOI and PESQ by 16.21% and 0.69 (31.1%) whereas with the LibriSpeech database, the model improves STOI by 16.41% and PESQ by 0.71 (32.9%) over the noisy speech. Further, our model outperforms other deep neural networks (DNNs) in seen and unseen conditions. The ASR performance is measured using the Kaldi toolkit and achieves 15.13% WERs in noisy backgrounds. [ABSTRACT FROM AUTHOR]
– Name: Abstract
  Label:
  Group: Ab
  Data: <i>Copyright of International Journal of Interactive Multimedia & Artificial Intelligence is the property of Universidad Internacional de La Rioja (UNIR) and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edb&AN=177785890
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.9781/ijimai.2024.04.003
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 12
        StartPage: 74
    Subjects:
      – SubjectFull: Automatic speech recognition
        Type: general
      – SubjectFull: Intelligibility of speech
        Type: general
      – SubjectFull: Speech enhancement
        Type: general
      – SubjectFull: Artificial neural networks
        Type: general
      – SubjectFull: Recurrent neural networks
        Type: general
      – SubjectFull: Speech
        Type: general
    Titles:
      – TitleFull: Lightweight Real-Time Recurrent Models for Speech Enhancement and Automatic Speech Recognition.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Dhahbi, Sami
      – PersonEntity:
          Name:
            NameFull: Saleem, Nasir
      – PersonEntity:
          Name:
            NameFull: Gunawan, Teddy Surya
      – PersonEntity:
          Name:
            NameFull: Bourouis, Sami
      – PersonEntity:
          Name:
            NameFull: Ali, Imad
      – PersonEntity:
          Name:
            NameFull: Trigui, Aymen
      – PersonEntity:
          Name:
            NameFull: Algarni, Abeer D.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 06
              Text: Jun2024
              Type: published
              Y: 2024
          Identifiers:
            – Type: issn-print
              Value: 19891660
          Numbering:
            – Type: volume
              Value: 8
            – Type: issue
              Value: 6
          Titles:
            – TitleFull: International Journal of Interactive Multimedia & Artificial Intelligence
              Type: main
ResultId 1