Skip to main navigation Skip to search Skip to main content

Bridging the Gap in Children’s Speech Recognition: Zero-Speech Approaches with Speech Modifications and ASR architectures

  • National Institute of Technology, Sikkim
  • AMD Silo AI
  • University of Southern California

Research output: Chapter in Book/Report/Conference proceedingConference article in proceedingsScientificpeer-review

7 Downloads (Pure)

Abstract

Pretrained end-to-end (E2E) automatic speech recognition (ASR) models, such as Wav2Vec2, HuBERT and WavLM, have achieved near-human performance on adult speech in zero-resource settings. However, their performance in children’s speech remains poor in zero-resource scenarios. To substantially improve performance in children ASR fine-tuning with little in-domain data is required, which might be untenable given the lack of labeled data. In this context, we wonder how without using children’s speech can we bridge the performance gap? In this work, we address this challenge by (1) reviewing modifications applicable in zero-resource scenarios, (2) leveraging in-domain text resources for adaptation, and (3) comparing both E2E ASR architectures and hybrid HMM/DNN Kaldi-based systems. Our observations serve as important takeaways for building children ASR with minimal resources.

Original languageEnglish
Title of host publication2025 33rd European Signal Processing Conference, EUSIPCO 2025 - Proceedings
PublisherEuropean Association For Signal Processing
Pages351-355
Number of pages5
ISBN (Electronic)978-94-645936-2-4
DOIs
Publication statusPublished - 2025
MoE publication typeA4 Conference publication
EventEuropean Signal Processing Conference - Palermo, Italy
Duration: 8 Sept 202512 Sept 2025
Conference number: 33

Conference

ConferenceEuropean Signal Processing Conference
Abbreviated titleEUSIPCO
Country/TerritoryItaly
CityPalermo
Period08/09/202512/09/2025

Keywords

  • ASR
  • children speech
  • end-to-end
  • HMM/DNN

Fingerprint

Dive into the research topics of 'Bridging the Gap in Children’s Speech Recognition: Zero-Speech Approaches with Speech Modifications and ASR architectures'. Together they form a unique fingerprint.

Cite this