Conference
Speech emotion recognition using CNN pretrained model.
| Τίτλος: | Speech emotion recognition using CNN pretrained model. |
|---|---|
| Συγγραφείς: | Alsaffar, Ali A., Ramo, Fawziya M. |
| Πηγή: | AIP Conference Proceedings; 2025, Vol. 3211 Issue 1, p1-12, 12p |
| Θεματικοί όροι: | Machine learning, Convolutional neural networks, Object recognition (Computer vision), Emotion recognition, Feature extraction, Deep learning, Automatic speech recognition |
| Περίληψη: | Transfer learning is the process of using previously acquired knowledge to accelerate development in a new environment. Pre-trained convolutional neural network (CNN) models are commonly used in image and object recognition programs and can also be applied to verbal emotion recognition. This study involves acquiring the features of audio files in the form of MFCC images with a sample rate of 22050, which are then sent to both a CNN model and CNN pre_trained models. Our contribution is to combine machine learning technology with deep learning technology to distinguish emotions in speech. In our proposed system, we used CNN model called (SERDL) and pre-trained models based on machine learning called (SERTL).in SERTL model, The features extraction process was performed in the first part of the CNN layers represented by the hidden layers. The second part of classification represented by fully connected layers has been removed and replaced with machine learning classification techniques such as SVM, DT, and RF. Our proposed system achieved 99% accuracy in speech emotion recognition for both the CNN model and the Transfer learning model on English dataset (TESS Dataset). For the Multi Dataset (English, French, and Italian), we obtained 94% accuracy using the SERTL_ML model based on ResNet50 network and 89% accuracy when using the SERDL_SL model based on CNN network model, SERTL primary model achieved high accuracy in classification using SVM and RF classifiers. [ABSTRACT FROM AUTHOR] |
| Copyright of AIP Conference Proceedings is the property of American Institute of Physics and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Βάση Δεδομένων: | Complementary Index |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: edb DbLabel: Complementary Index An: 184977069 RelevancyScore: 1009 AccessLevel: 6 PubType: Conference PubTypeId: conference PreciseRelevancyScore: 1008.54284667969 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Speech emotion recognition using CNN pretrained model. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Alsaffar%2C+Ali+A%2E%22">Alsaffar, Ali A.</searchLink><br /><searchLink fieldCode="AR" term="%22Ramo%2C+Fawziya+M%2E%22">Ramo, Fawziya M.</searchLink> – Name: TitleSource Label: Source Group: Src Data: AIP Conference Proceedings; 2025, Vol. 3211 Issue 1, p1-12, 12p – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Convolutional+neural+networks%22">Convolutional neural networks</searchLink><br /><searchLink fieldCode="DE" term="%22Object+recognition+%28Computer+vision%29%22">Object recognition (Computer vision)</searchLink><br /><searchLink fieldCode="DE" term="%22Emotion+recognition%22">Emotion recognition</searchLink><br /><searchLink fieldCode="DE" term="%22Feature+extraction%22">Feature extraction</searchLink><br /><searchLink fieldCode="DE" term="%22Deep+learning%22">Deep learning</searchLink><br /><searchLink fieldCode="DE" term="%22Automatic+speech+recognition%22">Automatic speech recognition</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Transfer learning is the process of using previously acquired knowledge to accelerate development in a new environment. Pre-trained convolutional neural network (CNN) models are commonly used in image and object recognition programs and can also be applied to verbal emotion recognition. This study involves acquiring the features of audio files in the form of MFCC images with a sample rate of 22050, which are then sent to both a CNN model and CNN pre_trained models. Our contribution is to combine machine learning technology with deep learning technology to distinguish emotions in speech. In our proposed system, we used CNN model called (SERDL) and pre-trained models based on machine learning called (SERTL).in SERTL model, The features extraction process was performed in the first part of the CNN layers represented by the hidden layers. The second part of classification represented by fully connected layers has been removed and replaced with machine learning classification techniques such as SVM, DT, and RF. Our proposed system achieved 99% accuracy in speech emotion recognition for both the CNN model and the Transfer learning model on English dataset (TESS Dataset). For the Multi Dataset (English, French, and Italian), we obtained 94% accuracy using the SERTL_ML model based on ResNet50 network and 89% accuracy when using the SERDL_SL model based on CNN network model, SERTL primary model achieved high accuracy in classification using SVM and RF classifiers. [ABSTRACT FROM AUTHOR] – Name: Abstract Label: Group: Ab Data: <i>Copyright of AIP Conference Proceedings is the property of American Institute of Physics and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edb&AN=184977069 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1063/5.0257175 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 12 StartPage: 1 Subjects: – SubjectFull: Machine learning Type: general – SubjectFull: Convolutional neural networks Type: general – SubjectFull: Object recognition (Computer vision) Type: general – SubjectFull: Emotion recognition Type: general – SubjectFull: Feature extraction Type: general – SubjectFull: Deep learning Type: general – SubjectFull: Automatic speech recognition Type: general Titles: – TitleFull: Speech emotion recognition using CNN pretrained model. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Alsaffar, Ali A. – PersonEntity: Name: NameFull: Ramo, Fawziya M. IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 03 Text: 2025 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 0094243X Numbering: – Type: volume Value: 3211 – Type: issue Value: 1 Titles: – TitleFull: AIP Conference Proceedings Type: main |
| ResultId | 1 |