Abstract
Pretrained end-to-end (E2E) automatic speech recognition (ASR) models, such as Wav2Vec2, HuBERT and WavLM, have achieved near-human performance on adult speech in zero-resource settings. However, their performance in children’s speech remains poor in zero-resource scenarios. To substantially improve performance in children ASR fine-tuning with little in-domain data is required, which might be untenable given the lack of labeled data. In this context, we wonder how without using children’s speech can we bridge the performance gap? In this work, we address this challenge by (1) reviewing modifications applicable in zero-resource scenarios, (2) leveraging in-domain text resources for adaptation, and (3) comparing both E2E ASR architectures and hybrid HMM/DNN Kaldi-based systems. Our observations serve as important takeaways for building children ASR with minimal resources.
| Original language | English |
|---|---|
| Title of host publication | 2025 33rd European Signal Processing Conference, EUSIPCO 2025 - Proceedings |
| Publisher | European Association For Signal Processing |
| Pages | 351-355 |
| Number of pages | 5 |
| ISBN (Electronic) | 978-94-645936-2-4 |
| DOIs | |
| Publication status | Published - 2025 |
| MoE publication type | A4 Conference publication |
| Event | European Signal Processing Conference - Palermo, Italy Duration: 8 Sept 2025 → 12 Sept 2025 Conference number: 33 |
Conference
| Conference | European Signal Processing Conference |
|---|---|
| Abbreviated title | EUSIPCO |
| Country/Territory | Italy |
| City | Palermo |
| Period | 08/09/2025 → 12/09/2025 |
Keywords
- ASR
- children speech
- end-to-end
- HMM/DNN
Fingerprint
Dive into the research topics of 'Bridging the Gap in Children’s Speech Recognition: Zero-Speech Approaches with Speech Modifications and ASR architectures'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver