Academic Journal

Masked-speech Recognition Using Human and Synthetic Cloned Speech.

Bibliographic Details
Title: Masked-speech Recognition Using Human and Synthetic Cloned Speech.
Authors: Calandruccio, Lauren, Hariri, Mohsen, Buss, Emily, Chaudhary, Vipin
Source: Trends in Hearing; 12/8/2025, Vol. 29, p1-14, 14p
Subject Terms: Masking (Psychology), Automatic speech recognition, Scale analysis (Psychology), Task performance, Data analysis, Research funding, Artificial intelligence, Statistical sampling, Intelligibility of speech, Descriptive statistics, Experimental design, Physiological aspects of speech, Speech evaluation, Statistics, Human voice, Speech perception, Semantics, Confidence intervals, Data analysis software, Regression analysis
Abstract: Voice cloning is used to generate synthetic speech that mimics vocal characteristics of human talkers. This experiment used voice cloning to compare human and synthetic speech for intelligibility, human-likeness, and perceptual similarity, all tested in young adults with normal hearing. Masked-sentence recognition was evaluated using speech produced by five human talkers and their synthetically generated voice clones presented in speech-shaped noise at −6 dB signal-to-noise ratio. There were two types of sentences: semantically meaningful and nonsense. Human and automatic speech recognition scoring was used to evaluate performance. Participants were asked to rate human-likeness and determine whether pairs of sentences were produced by the same versus different people. As expected, sentence-recognition scores were worse for nonsense sentences compared to meaningful sentences, but they were similar for speech produced by human talkers and voice clones. Human-likeness scores were also similar for speech produced by human talkers and their voice clones. Participants were very good at identifying differences between voices but were less accurate at distinguishing between human/clone pairs, often leaning towards thinking they were produced by the same person. Reliability scoring by automatic speech recognition agreed with human reliability scoring for 98% of keywords and was minimally dependent on the context of the target sentences. Results provide preliminary support for the use of voice clones when evaluating the recognition of human and synthetic speech. More generally, voice synthesis and automatic speech recognition are promising tools for evaluating speech recognition in human listeners. [ABSTRACT FROM AUTHOR]
Copyright of Trends in Hearing is the property of Sage Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Complementary Index
FullText Links:
  – Type: other
Text:
  Availability: 0
Header DbId: edb
DbLabel: Complementary Index
An: 190255144
RelevancyScore: 1041
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 1040.81286621094
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Masked-speech Recognition Using Human and Synthetic Cloned Speech.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Calandruccio%2C+Lauren%22">Calandruccio, Lauren</searchLink><br /><searchLink fieldCode="AR" term="%22Hariri%2C+Mohsen%22">Hariri, Mohsen</searchLink><br /><searchLink fieldCode="AR" term="%22Buss%2C+Emily%22">Buss, Emily</searchLink><br /><searchLink fieldCode="AR" term="%22Chaudhary%2C+Vipin%22">Chaudhary, Vipin</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: Trends in Hearing; 12/8/2025, Vol. 29, p1-14, 14p
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Masking+%28Psychology%29%22">Masking (Psychology)</searchLink><br /><searchLink fieldCode="DE" term="%22Automatic+speech+recognition%22">Automatic speech recognition</searchLink><br /><searchLink fieldCode="DE" term="%22Scale+analysis+%28Psychology%29%22">Scale analysis (Psychology)</searchLink><br /><searchLink fieldCode="DE" term="%22Task+performance%22">Task performance</searchLink><br /><searchLink fieldCode="DE" term="%22Data+analysis%22">Data analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Research+funding%22">Research funding</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+intelligence%22">Artificial intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+sampling%22">Statistical sampling</searchLink><br /><searchLink fieldCode="DE" term="%22Intelligibility+of+speech%22">Intelligibility of speech</searchLink><br /><searchLink fieldCode="DE" term="%22Descriptive+statistics%22">Descriptive statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Experimental+design%22">Experimental design</searchLink><br /><searchLink fieldCode="DE" term="%22Physiological+aspects+of+speech%22">Physiological aspects of speech</searchLink><br /><searchLink fieldCode="DE" term="%22Speech+evaluation%22">Speech evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Statistics%22">Statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Human+voice%22">Human voice</searchLink><br /><searchLink fieldCode="DE" term="%22Speech+perception%22">Speech perception</searchLink><br /><searchLink fieldCode="DE" term="%22Semantics%22">Semantics</searchLink><br /><searchLink fieldCode="DE" term="%22Confidence+intervals%22">Confidence intervals</searchLink><br /><searchLink fieldCode="DE" term="%22Data+analysis+software%22">Data analysis software</searchLink><br /><searchLink fieldCode="DE" term="%22Regression+analysis%22">Regression analysis</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Voice cloning is used to generate synthetic speech that mimics vocal characteristics of human talkers. This experiment used voice cloning to compare human and synthetic speech for intelligibility, human-likeness, and perceptual similarity, all tested in young adults with normal hearing. Masked-sentence recognition was evaluated using speech produced by five human talkers and their synthetically generated voice clones presented in speech-shaped noise at −6 dB signal-to-noise ratio. There were two types of sentences: semantically meaningful and nonsense. Human and automatic speech recognition scoring was used to evaluate performance. Participants were asked to rate human-likeness and determine whether pairs of sentences were produced by the same versus different people. As expected, sentence-recognition scores were worse for nonsense sentences compared to meaningful sentences, but they were similar for speech produced by human talkers and voice clones. Human-likeness scores were also similar for speech produced by human talkers and their voice clones. Participants were very good at identifying differences between voices but were less accurate at distinguishing between human/clone pairs, often leaning towards thinking they were produced by the same person. Reliability scoring by automatic speech recognition agreed with human reliability scoring for 98% of keywords and was minimally dependent on the context of the target sentences. Results provide preliminary support for the use of voice clones when evaluating the recognition of human and synthetic speech. More generally, voice synthesis and automatic speech recognition are promising tools for evaluating speech recognition in human listeners. [ABSTRACT FROM AUTHOR]
– Name: Abstract
  Label:
  Group: Ab
  Data: <i>Copyright of Trends in Hearing is the property of Sage Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edb&AN=190255144
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1177/23312165251403080
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 14
        StartPage: 1
    Subjects:
      – SubjectFull: Masking (Psychology)
        Type: general
      – SubjectFull: Automatic speech recognition
        Type: general
      – SubjectFull: Scale analysis (Psychology)
        Type: general
      – SubjectFull: Task performance
        Type: general
      – SubjectFull: Data analysis
        Type: general
      – SubjectFull: Research funding
        Type: general
      – SubjectFull: Artificial intelligence
        Type: general
      – SubjectFull: Statistical sampling
        Type: general
      – SubjectFull: Intelligibility of speech
        Type: general
      – SubjectFull: Descriptive statistics
        Type: general
      – SubjectFull: Experimental design
        Type: general
      – SubjectFull: Physiological aspects of speech
        Type: general
      – SubjectFull: Speech evaluation
        Type: general
      – SubjectFull: Statistics
        Type: general
      – SubjectFull: Human voice
        Type: general
      – SubjectFull: Speech perception
        Type: general
      – SubjectFull: Semantics
        Type: general
      – SubjectFull: Confidence intervals
        Type: general
      – SubjectFull: Data analysis software
        Type: general
      – SubjectFull: Regression analysis
        Type: general
    Titles:
      – TitleFull: Masked-speech Recognition Using Human and Synthetic Cloned Speech.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Calandruccio, Lauren
      – PersonEntity:
          Name:
            NameFull: Hariri, Mohsen
      – PersonEntity:
          Name:
            NameFull: Buss, Emily
      – PersonEntity:
          Name:
            NameFull: Chaudhary, Vipin
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 08
              M: 12
              Text: 12/8/2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 23312165
          Numbering:
            – Type: volume
              Value: 29
          Titles:
            – TitleFull: Trends in Hearing
              Type: main
ResultId 1