Speech emotion recognition using CNN pretrained model.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Speech emotion recognition using CNN pretrained model.
Συγγραφείς: Alsaffar, Ali A., Ramo, Fawziya M.
Πηγή: AIP Conference Proceedings; 2025, Vol. 3211 Issue 1, p1-12, 12p
Θεματικοί όροι: Machine learning, Convolutional neural networks, Object recognition (Computer vision), Emotion recognition, Feature extraction, Deep learning, Automatic speech recognition
Περίληψη: Transfer learning is the process of using previously acquired knowledge to accelerate development in a new environment. Pre-trained convolutional neural network (CNN) models are commonly used in image and object recognition programs and can also be applied to verbal emotion recognition. This study involves acquiring the features of audio files in the form of MFCC images with a sample rate of 22050, which are then sent to both a CNN model and CNN pre_trained models. Our contribution is to combine machine learning technology with deep learning technology to distinguish emotions in speech. In our proposed system, we used CNN model called (SERDL) and pre-trained models based on machine learning called (SERTL).in SERTL model, The features extraction process was performed in the first part of the CNN layers represented by the hidden layers. The second part of classification represented by fully connected layers has been removed and replaced with machine learning classification techniques such as SVM, DT, and RF. Our proposed system achieved 99% accuracy in speech emotion recognition for both the CNN model and the Transfer learning model on English dataset (TESS Dataset). For the Multi Dataset (English, French, and Italian), we obtained 94% accuracy using the SERTL_ML model based on ResNet50 network and 89% accuracy when using the SERDL_SL model based on CNN network model, SERTL primary model achieved high accuracy in classification using SVM and RF classifiers. [ABSTRACT FROM AUTHOR]
Copyright of AIP Conference Proceedings is the property of American Institute of Physics and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Βάση Δεδομένων: Complementary Index
FullText Text:
  Availability: 0
Header DbId: edb
DbLabel: Complementary Index
An: 184977069
RelevancyScore: 1009
AccessLevel: 6
PubType: Conference
PubTypeId: conference
PreciseRelevancyScore: 1008.54284667969
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Speech emotion recognition using CNN pretrained model.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Alsaffar%2C+Ali+A%2E%22">Alsaffar, Ali A.</searchLink><br /><searchLink fieldCode="AR" term="%22Ramo%2C+Fawziya+M%2E%22">Ramo, Fawziya M.</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: AIP Conference Proceedings; 2025, Vol. 3211 Issue 1, p1-12, 12p
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Convolutional+neural+networks%22">Convolutional neural networks</searchLink><br /><searchLink fieldCode="DE" term="%22Object+recognition+%28Computer+vision%29%22">Object recognition (Computer vision)</searchLink><br /><searchLink fieldCode="DE" term="%22Emotion+recognition%22">Emotion recognition</searchLink><br /><searchLink fieldCode="DE" term="%22Feature+extraction%22">Feature extraction</searchLink><br /><searchLink fieldCode="DE" term="%22Deep+learning%22">Deep learning</searchLink><br /><searchLink fieldCode="DE" term="%22Automatic+speech+recognition%22">Automatic speech recognition</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Transfer learning is the process of using previously acquired knowledge to accelerate development in a new environment. Pre-trained convolutional neural network (CNN) models are commonly used in image and object recognition programs and can also be applied to verbal emotion recognition. This study involves acquiring the features of audio files in the form of MFCC images with a sample rate of 22050, which are then sent to both a CNN model and CNN pre_trained models. Our contribution is to combine machine learning technology with deep learning technology to distinguish emotions in speech. In our proposed system, we used CNN model called (SERDL) and pre-trained models based on machine learning called (SERTL).in SERTL model, The features extraction process was performed in the first part of the CNN layers represented by the hidden layers. The second part of classification represented by fully connected layers has been removed and replaced with machine learning classification techniques such as SVM, DT, and RF. Our proposed system achieved 99% accuracy in speech emotion recognition for both the CNN model and the Transfer learning model on English dataset (TESS Dataset). For the Multi Dataset (English, French, and Italian), we obtained 94% accuracy using the SERTL_ML model based on ResNet50 network and 89% accuracy when using the SERDL_SL model based on CNN network model, SERTL primary model achieved high accuracy in classification using SVM and RF classifiers. [ABSTRACT FROM AUTHOR]
– Name: Abstract
  Label:
  Group: Ab
  Data: <i>Copyright of AIP Conference Proceedings is the property of American Institute of Physics and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edb&AN=184977069
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1063/5.0257175
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 12
        StartPage: 1
    Subjects:
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: Convolutional neural networks
        Type: general
      – SubjectFull: Object recognition (Computer vision)
        Type: general
      – SubjectFull: Emotion recognition
        Type: general
      – SubjectFull: Feature extraction
        Type: general
      – SubjectFull: Deep learning
        Type: general
      – SubjectFull: Automatic speech recognition
        Type: general
    Titles:
      – TitleFull: Speech emotion recognition using CNN pretrained model.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Alsaffar, Ali A.
      – PersonEntity:
          Name:
            NameFull: Ramo, Fawziya M.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 03
              Text: 2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 0094243X
          Numbering:
            – Type: volume
              Value: 3211
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: AIP Conference Proceedings
              Type: main
ResultId 1