Λεπτομέρειες βιβλιογραφικής εγγραφής
| Τίτλος: |
Masked-speech Recognition Using Human and Synthetic Cloned Speech. |
| Συγγραφείς: |
Calandruccio, Lauren, Hariri, Mohsen, Buss, Emily, Chaudhary, Vipin |
| Πηγή: |
Trends in Hearing; 12/8/2025, Vol. 29, p1-14, 14p |
| Θεματικοί όροι: |
Masking (Psychology), Automatic speech recognition, Scale analysis (Psychology), Task performance, Data analysis, Research funding, Artificial intelligence, Statistical sampling, Intelligibility of speech, Descriptive statistics, Experimental design, Physiological aspects of speech, Speech evaluation, Statistics, Human voice, Speech perception, Semantics, Confidence intervals, Data analysis software, Regression analysis |
| Περίληψη: |
Voice cloning is used to generate synthetic speech that mimics vocal characteristics of human talkers. This experiment used voice cloning to compare human and synthetic speech for intelligibility, human-likeness, and perceptual similarity, all tested in young adults with normal hearing. Masked-sentence recognition was evaluated using speech produced by five human talkers and their synthetically generated voice clones presented in speech-shaped noise at −6 dB signal-to-noise ratio. There were two types of sentences: semantically meaningful and nonsense. Human and automatic speech recognition scoring was used to evaluate performance. Participants were asked to rate human-likeness and determine whether pairs of sentences were produced by the same versus different people. As expected, sentence-recognition scores were worse for nonsense sentences compared to meaningful sentences, but they were similar for speech produced by human talkers and voice clones. Human-likeness scores were also similar for speech produced by human talkers and their voice clones. Participants were very good at identifying differences between voices but were less accurate at distinguishing between human/clone pairs, often leaning towards thinking they were produced by the same person. Reliability scoring by automatic speech recognition agreed with human reliability scoring for 98% of keywords and was minimally dependent on the context of the target sentences. Results provide preliminary support for the use of voice clones when evaluating the recognition of human and synthetic speech. More generally, voice synthesis and automatic speech recognition are promising tools for evaluating speech recognition in human listeners. [ABSTRACT FROM AUTHOR] |
|
Copyright of Trends in Hearing is the property of Sage Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) |
| Βάση Δεδομένων: |
Complementary Index |