Academic Journal

Speech Emotion Recognition Using Deep Learning Transfer Models and Explainable Techniques.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Speech Emotion Recognition Using Deep Learning Transfer Models and Explainable Techniques.
Συγγραφείς: Kim, Tae-Wan, Kwak, Keun-Chang
Πηγή: Applied Sciences (2076-3417); Feb2024, Vol. 14 Issue 4, p1553, 23p
Θεματικοί όροι: Emotion recognition, Deep learning, Transfer of training, Speech, Gaussian distribution, Signal processing
Περίληψη: This study aims to establish a greater reliability compared to conventional speech emotion recognition (SER) studies. This is achieved through preprocessing techniques that reduce uncertainty elements, models that combine the structural features of each model, and the application of various explanatory techniques. The ability to interpret can be made more accurate by reducing uncertain learning data, applying data in different environments, and applying techniques that explain the reasoning behind the results. We designed a generalized model using three different datasets, and each speech was converted into a spectrogram image through STFT preprocessing. The spectrogram was divided into the time domain with overlapping to match the input size of the model. Each divided section is expressed as a Gaussian distribution, and the quality of the data is investigated by the correlation coefficient between distributions. As a result, the scale of the data is reduced, and uncertainty is minimized. VGGish and YAMNet are the most representative pretrained deep learning networks frequently used in conjunction with speech processing. In dealing with speech signal processing, it is frequently advantageous to use these pretrained models synergistically rather than exclusively, resulting in the construction of ensemble deep networks. And finally, various explainable models (Grad CAM, LIME, occlusion sensitivity) are used in analyzing classified results. The model exhibits adaptability to voices in various environments, yielding a classification accuracy of 87%, surpassing that of individual models. Additionally, output results are confirmed by an explainable model to extract essential emotional areas, converted into audio files for auditory analysis using Grad CAM in the time domain. Through this study, we enhance the uncertainty of activation areas that are generated by Grad CAM. We achieve this by applying the interpretable ability from previous studies, along with effective preprocessing and fusion models. We can analyze it from a more diverse perspective through other explainable techniques. [ABSTRACT FROM AUTHOR]
Copyright of Applied Sciences (2076-3417) is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Βάση Δεδομένων: Complementary Index
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://resolver.ebsco.com/c/fiv2js/result?sid=EBSCO:edb&genre=article&issn=20763417&ISBN=&volume=14&issue=4&date=20240215&spage=1553&pages=1553-1575&title=Applied Sciences (2076-3417)&atitle=Speech%20Emotion%20Recognition%20Using%20Deep%20Learning%20Transfer%20Models%20and%20Explainable%20Techniques.&aulast=Kim%2C%20Tae-Wan&id=DOI:10.3390/app14041553
    Name: Full Text Finder (for New FTF UI) (ns324271)
    Category: fullText
    Text: Full Text Finder
    MouseOverText: Full Text Finder
Header DbId: edb
DbLabel: Complementary Index
An: 175652541
RelevancyScore: 950
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 949.948608398438
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Speech Emotion Recognition Using Deep Learning Transfer Models and Explainable Techniques.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Kim%2C+Tae-Wan%22">Kim, Tae-Wan</searchLink><br /><searchLink fieldCode="AR" term="%22Kwak%2C+Keun-Chang%22">Kwak, Keun-Chang</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: Applied Sciences (2076-3417); Feb2024, Vol. 14 Issue 4, p1553, 23p
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Emotion+recognition%22">Emotion recognition</searchLink><br /><searchLink fieldCode="DE" term="%22Deep+learning%22">Deep learning</searchLink><br /><searchLink fieldCode="DE" term="%22Transfer+of+training%22">Transfer of training</searchLink><br /><searchLink fieldCode="DE" term="%22Speech%22">Speech</searchLink><br /><searchLink fieldCode="DE" term="%22Gaussian+distribution%22">Gaussian distribution</searchLink><br /><searchLink fieldCode="DE" term="%22Signal+processing%22">Signal processing</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: This study aims to establish a greater reliability compared to conventional speech emotion recognition (SER) studies. This is achieved through preprocessing techniques that reduce uncertainty elements, models that combine the structural features of each model, and the application of various explanatory techniques. The ability to interpret can be made more accurate by reducing uncertain learning data, applying data in different environments, and applying techniques that explain the reasoning behind the results. We designed a generalized model using three different datasets, and each speech was converted into a spectrogram image through STFT preprocessing. The spectrogram was divided into the time domain with overlapping to match the input size of the model. Each divided section is expressed as a Gaussian distribution, and the quality of the data is investigated by the correlation coefficient between distributions. As a result, the scale of the data is reduced, and uncertainty is minimized. VGGish and YAMNet are the most representative pretrained deep learning networks frequently used in conjunction with speech processing. In dealing with speech signal processing, it is frequently advantageous to use these pretrained models synergistically rather than exclusively, resulting in the construction of ensemble deep networks. And finally, various explainable models (Grad CAM, LIME, occlusion sensitivity) are used in analyzing classified results. The model exhibits adaptability to voices in various environments, yielding a classification accuracy of 87%, surpassing that of individual models. Additionally, output results are confirmed by an explainable model to extract essential emotional areas, converted into audio files for auditory analysis using Grad CAM in the time domain. Through this study, we enhance the uncertainty of activation areas that are generated by Grad CAM. We achieve this by applying the interpretable ability from previous studies, along with effective preprocessing and fusion models. We can analyze it from a more diverse perspective through other explainable techniques. [ABSTRACT FROM AUTHOR]
– Name: Abstract
  Label:
  Group: Ab
  Data: <i>Copyright of Applied Sciences (2076-3417) is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edb&AN=175652541
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.3390/app14041553
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 23
        StartPage: 1553
    Subjects:
      – SubjectFull: Emotion recognition
        Type: general
      – SubjectFull: Deep learning
        Type: general
      – SubjectFull: Transfer of training
        Type: general
      – SubjectFull: Speech
        Type: general
      – SubjectFull: Gaussian distribution
        Type: general
      – SubjectFull: Signal processing
        Type: general
    Titles:
      – TitleFull: Speech Emotion Recognition Using Deep Learning Transfer Models and Explainable Techniques.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Kim, Tae-Wan
      – PersonEntity:
          Name:
            NameFull: Kwak, Keun-Chang
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 15
              M: 02
              Text: Feb2024
              Type: published
              Y: 2024
          Identifiers:
            – Type: issn-print
              Value: 20763417
          Numbering:
            – Type: volume
              Value: 14
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: Applied Sciences (2076-3417)
              Type: main
ResultId 1