Academic Journal
Masked-speech Recognition Using Human and Synthetic Cloned Speech.
| Title: | Masked-speech Recognition Using Human and Synthetic Cloned Speech. |
|---|---|
| Authors: | Calandruccio, Lauren, Hariri, Mohsen, Buss, Emily, Chaudhary, Vipin |
| Source: | Trends in Hearing; 12/8/2025, Vol. 29, p1-14, 14p |
| Subject Terms: | Masking (Psychology), Automatic speech recognition, Scale analysis (Psychology), Task performance, Data analysis, Research funding, Artificial intelligence, Statistical sampling, Intelligibility of speech, Descriptive statistics, Experimental design, Physiological aspects of speech, Speech evaluation, Statistics, Human voice, Speech perception, Semantics, Confidence intervals, Data analysis software, Regression analysis |
| Abstract: | Voice cloning is used to generate synthetic speech that mimics vocal characteristics of human talkers. This experiment used voice cloning to compare human and synthetic speech for intelligibility, human-likeness, and perceptual similarity, all tested in young adults with normal hearing. Masked-sentence recognition was evaluated using speech produced by five human talkers and their synthetically generated voice clones presented in speech-shaped noise at −6 dB signal-to-noise ratio. There were two types of sentences: semantically meaningful and nonsense. Human and automatic speech recognition scoring was used to evaluate performance. Participants were asked to rate human-likeness and determine whether pairs of sentences were produced by the same versus different people. As expected, sentence-recognition scores were worse for nonsense sentences compared to meaningful sentences, but they were similar for speech produced by human talkers and voice clones. Human-likeness scores were also similar for speech produced by human talkers and their voice clones. Participants were very good at identifying differences between voices but were less accurate at distinguishing between human/clone pairs, often leaning towards thinking they were produced by the same person. Reliability scoring by automatic speech recognition agreed with human reliability scoring for 98% of keywords and was minimally dependent on the context of the target sentences. Results provide preliminary support for the use of voice clones when evaluating the recognition of human and synthetic speech. More generally, voice synthesis and automatic speech recognition are promising tools for evaluating speech recognition in human listeners. [ABSTRACT FROM AUTHOR] |
| Copyright of Trends in Hearing is the property of Sage Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Complementary Index |
| FullText | Links: – Type: other Text: Availability: 0 |
|---|---|
| Header | DbId: edb DbLabel: Complementary Index An: 190255144 RelevancyScore: 1041 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 1040.81286621094 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Masked-speech Recognition Using Human and Synthetic Cloned Speech. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Calandruccio%2C+Lauren%22">Calandruccio, Lauren</searchLink><br /><searchLink fieldCode="AR" term="%22Hariri%2C+Mohsen%22">Hariri, Mohsen</searchLink><br /><searchLink fieldCode="AR" term="%22Buss%2C+Emily%22">Buss, Emily</searchLink><br /><searchLink fieldCode="AR" term="%22Chaudhary%2C+Vipin%22">Chaudhary, Vipin</searchLink> – Name: TitleSource Label: Source Group: Src Data: Trends in Hearing; 12/8/2025, Vol. 29, p1-14, 14p – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Masking+%28Psychology%29%22">Masking (Psychology)</searchLink><br /><searchLink fieldCode="DE" term="%22Automatic+speech+recognition%22">Automatic speech recognition</searchLink><br /><searchLink fieldCode="DE" term="%22Scale+analysis+%28Psychology%29%22">Scale analysis (Psychology)</searchLink><br /><searchLink fieldCode="DE" term="%22Task+performance%22">Task performance</searchLink><br /><searchLink fieldCode="DE" term="%22Data+analysis%22">Data analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Research+funding%22">Research funding</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+intelligence%22">Artificial intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+sampling%22">Statistical sampling</searchLink><br /><searchLink fieldCode="DE" term="%22Intelligibility+of+speech%22">Intelligibility of speech</searchLink><br /><searchLink fieldCode="DE" term="%22Descriptive+statistics%22">Descriptive statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Experimental+design%22">Experimental design</searchLink><br /><searchLink fieldCode="DE" term="%22Physiological+aspects+of+speech%22">Physiological aspects of speech</searchLink><br /><searchLink fieldCode="DE" term="%22Speech+evaluation%22">Speech evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Statistics%22">Statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Human+voice%22">Human voice</searchLink><br /><searchLink fieldCode="DE" term="%22Speech+perception%22">Speech perception</searchLink><br /><searchLink fieldCode="DE" term="%22Semantics%22">Semantics</searchLink><br /><searchLink fieldCode="DE" term="%22Confidence+intervals%22">Confidence intervals</searchLink><br /><searchLink fieldCode="DE" term="%22Data+analysis+software%22">Data analysis software</searchLink><br /><searchLink fieldCode="DE" term="%22Regression+analysis%22">Regression analysis</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Voice cloning is used to generate synthetic speech that mimics vocal characteristics of human talkers. This experiment used voice cloning to compare human and synthetic speech for intelligibility, human-likeness, and perceptual similarity, all tested in young adults with normal hearing. Masked-sentence recognition was evaluated using speech produced by five human talkers and their synthetically generated voice clones presented in speech-shaped noise at −6 dB signal-to-noise ratio. There were two types of sentences: semantically meaningful and nonsense. Human and automatic speech recognition scoring was used to evaluate performance. Participants were asked to rate human-likeness and determine whether pairs of sentences were produced by the same versus different people. As expected, sentence-recognition scores were worse for nonsense sentences compared to meaningful sentences, but they were similar for speech produced by human talkers and voice clones. Human-likeness scores were also similar for speech produced by human talkers and their voice clones. Participants were very good at identifying differences between voices but were less accurate at distinguishing between human/clone pairs, often leaning towards thinking they were produced by the same person. Reliability scoring by automatic speech recognition agreed with human reliability scoring for 98% of keywords and was minimally dependent on the context of the target sentences. Results provide preliminary support for the use of voice clones when evaluating the recognition of human and synthetic speech. More generally, voice synthesis and automatic speech recognition are promising tools for evaluating speech recognition in human listeners. [ABSTRACT FROM AUTHOR] – Name: Abstract Label: Group: Ab Data: <i>Copyright of Trends in Hearing is the property of Sage Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edb&AN=190255144 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1177/23312165251403080 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 14 StartPage: 1 Subjects: – SubjectFull: Masking (Psychology) Type: general – SubjectFull: Automatic speech recognition Type: general – SubjectFull: Scale analysis (Psychology) Type: general – SubjectFull: Task performance Type: general – SubjectFull: Data analysis Type: general – SubjectFull: Research funding Type: general – SubjectFull: Artificial intelligence Type: general – SubjectFull: Statistical sampling Type: general – SubjectFull: Intelligibility of speech Type: general – SubjectFull: Descriptive statistics Type: general – SubjectFull: Experimental design Type: general – SubjectFull: Physiological aspects of speech Type: general – SubjectFull: Speech evaluation Type: general – SubjectFull: Statistics Type: general – SubjectFull: Human voice Type: general – SubjectFull: Speech perception Type: general – SubjectFull: Semantics Type: general – SubjectFull: Confidence intervals Type: general – SubjectFull: Data analysis software Type: general – SubjectFull: Regression analysis Type: general Titles: – TitleFull: Masked-speech Recognition Using Human and Synthetic Cloned Speech. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Calandruccio, Lauren – PersonEntity: Name: NameFull: Hariri, Mohsen – PersonEntity: Name: NameFull: Buss, Emily – PersonEntity: Name: NameFull: Chaudhary, Vipin IsPartOfRelationships: – BibEntity: Dates: – D: 08 M: 12 Text: 12/8/2025 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 23312165 Numbering: – Type: volume Value: 29 Titles: – TitleFull: Trends in Hearing Type: main |
| ResultId | 1 |