Automatic Speech Recognition in Primary Progressive Apraxia of Speech.

Tetzloff, Katerina A; Wiepert, Daniela; Botha, Hugo; Duffy, Joseph R; Clark, Heather M; Whitwell, Jennifer L; Josephs, Keith A; Utianski, Rene L

Tetzloff, Katerina A; Wiepert, Daniela; Botha, Hugo; Duffy, Joseph R; Clark, Heather M; Whitwell, Jennifer L; Josephs, Keith A; Utianski, Rene L.

Afiliação

Tetzloff KA; Department of Neurology, Mayo Clinic, Rochester, MN.
Wiepert D; Department of Neurology, Mayo Clinic, Rochester, MN.
Botha H; Department of Neurology, Mayo Clinic, Rochester, MN.
Duffy JR; Department of Neurology, Mayo Clinic, Rochester, MN.
Clark HM; Department of Neurology, Mayo Clinic, Rochester, MN.
Whitwell JL; Department of Radiology, Mayo Clinic, Rochester, MN.
Josephs KA; Department of Neurology, Mayo Clinic, Rochester, MN.
Utianski RL; Department of Neurology, Mayo Clinic, Rochester, MN.

J Speech Lang Hear Res ; 67(9): 2964-2976, 2024 Sep 12.

Article em En | MEDLINE | ID: mdl-39265154

ABSTRACT

ABSTRACT

INTRODUCTION:

Transcribing disordered speech can be useful when diagnosing motor speech disorders such as primary progressive apraxia of speech (PPAOS), who have sound additions, deletions, and substitutions, or distortions and/or slow, segmented speech. Since transcribing speech can be a laborious process and requires an experienced listener, using automatic speech recognition (ASR) systems for diagnosis and treatment monitoring is appealing. This study evaluated the efficacy of a readily available ASR system (wav2vec 2.0) in transcribing speech of PPAOS patients to determine if the word error rate (WER) output by the ASR can differentiate between healthy speech and PPAOS and/or among its subtypes, whether WER correlates with AOS severity, and how the ASR's errors compare to those noted in manual transcriptions.

METHOD:

Forty-five patients with PPAOS and 22 healthy controls were recorded repeating 13 words, 3 times each, which were transcribed manually and using wav2vec 2.0. The WER and phonetic and prosodic speech errors were compared between groups, and ASR results were compared against manual transcriptions.

RESULTS:

Mean overall WER was 0.88 for patients and 0.33 for controls. WER significantly correlated with AOS severity and accurately distinguished between patients and controls but not between AOS subtypes. The phonetic and prosodic errors from the ASR transcriptions were also unable to distinguish between subtypes, whereas errors calculated from human transcriptions were. There was poor agreement in the number of phonetic and prosodic errors between the ASR and human transcriptions.

CONCLUSIONS:

This study demonstrates that ASR can be useful in differentiating healthy from disordered speech and evaluating PPAOS severity but does not distinguish PPAOS subtypes. ASR transcriptions showed weak agreement with human transcriptions; thus, ASR may be a useful tool for the transcription of speech in PPAOS, but the research questions posed must be carefully considered within the context of its limitations. SUPPLEMENTAL

MATERIAL:

https//doi.org/10.23641/asha.26359417.

Assuntos

Interface para o Reconhecimento da Fala; Humanos; Masculino; Feminino; Idoso; Pessoa de Meia-Idade; Fala/fisiologia; Apraxias/diagnóstico; Medida da Produção da Fala/métodos; Fonética; Afasia Primária Progressiva/diagnóstico; Estudos de Casos e Controles

Texto completo

Adicionar na Minha BVS

Imprimir

XML

PubMed Links

Buscar no Google

Texto completo: 1 Coleções: 01-internacional Base de dados: MEDLINE Assunto principal: Interface para o Reconhecimento da Fala Limite: Aged / Female / Humans / Male / Middle aged Idioma: En Revista: J Speech Lang Hear Res Assunto da revista: AUDIOLOGIA / PATOLOGIA DA FALA E LINGUAGEM Ano de publicação: 2024 Tipo de documento: Article País de afiliação: Mongólia País de publicação: Estados Unidos

Texto completo

Adicionar na Minha BVS

Imprimir

XML

PubMed Links

Buscar no Google