Abstrakti
Pretrained end-to-end (E2E) automatic speech recognition (ASR) models, such as Wav2Vec2, HuBERT and WavLM, have achieved near-human performance on adult speech in zero-resource settings. However, their performance in children’s speech remains poor in zero-resource scenarios. To substantially improve performance in children ASR fine-tuning with little in-domain data is required, which might be untenable given the lack of labeled data. In this context, we wonder how without using children’s speech can we bridge the performance gap? In this work, we address this challenge by (1) reviewing modifications applicable in zero-resource scenarios, (2) leveraging in-domain text resources for adaptation, and (3) comparing both E2E ASR architectures and hybrid HMM/DNN Kaldi-based systems. Our observations serve as important takeaways for building children ASR with minimal resources.
| Alkuperäiskieli | Englanti |
|---|---|
| Otsikko | 2025 33rd European Signal Processing Conference, EUSIPCO 2025 - Proceedings |
| Kustantaja | European Association For Signal Processing |
| Sivut | 351-355 |
| Sivumäärä | 5 |
| ISBN (elektroninen) | 978-94-645936-2-4 |
| DOI - pysyväislinkit | |
| Tila | Julkaistu - 2025 |
| OKM-julkaisutyyppi | A4 Artikkeli konferenssijulkaisussa |
| Tapahtuma | European Signal Processing Conference - Palermo, Italia Kesto: 8 syysk. 2025 → 12 syysk. 2025 Konferenssinumero: 33 |
Conference
| Conference | European Signal Processing Conference |
|---|---|
| Lyhennettä | EUSIPCO |
| Maa/Alue | Italia |
| Kaupunki | Palermo |
| Ajanjakso | 08/09/2025 → 12/09/2025 |
Sormenjälki
Sukella tutkimusaiheisiin 'Bridging the Gap in Children’s Speech Recognition: Zero-Speech Approaches with Speech Modifications and ASR architectures'. Ne muodostavat yhdessä ainutlaatuisen sormenjäljen.Siteeraa tätä
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver