!!Projects per year
Abstrakti
In this paper, we propose a software toolkit for easier end-to-end training of deep learning based spoken language identification models across several speech datasets. We apply our toolkit to implement three baseline models, one speaker recognition model, and three x-vector architecture variations, which are trained on three datasets previously used in spoken language identification experiments. All models are trained separately on each dataset (closed task) and on a combination of all datasets (open task), after which we compare if the open task training yields better language embeddings. We begin by training all models end-to-end as discriminative classifiers of spectral features, labeled by language. Then, we extract language embedding vectors from the trained end-to-end models, train separate Gaussian Naive Bayes classifiers on the vectors, and compare which model provides best language embeddings for the back-end classifier. Our experiments show that the open task condition leads to improved language identification performance on only one of the datasets. In addition, we discovered that increasing x-vector model robustness with random frequency channel dropout significantly reduces its end-to-end classification performance on the test set, while not affecting back-end classification performance of its embeddings. Finally, we note that two baseline models consistently outperformed all other models.
| Alkuperäiskieli | Englanti |
|---|---|
| Otsikko | Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH |
| Kustantaja | International Speech Communication Association (ISCA) |
| Sivut | 467-471 |
| Sivumäärä | 5 |
| Vuosikerta | 2020-October |
| DOI - pysyväislinkit | |
| Tila | Julkaistu - 2020 |
| OKM-julkaisutyyppi | A4 Artikkeli konferenssijulkaisussa |
| Tapahtuma | Interspeech - Shanghai, Kiina Kesto: 25 lokak. 2020 → 29 lokak. 2020 Konferenssinumero: 21 http://www.interspeech2020.org/ |
Julkaisusarja
| Nimi | Interspeech |
|---|---|
| Kustantaja | International Speech Communication Association |
| ISSN (painettu) | 2308-457X |
Conference
| Conference | Interspeech |
|---|---|
| Lyhennettä | INTERSPEECH |
| Maa/Alue | Kiina |
| Kaupunki | Shanghai |
| Ajanjakso | 25/10/2020 → 29/10/2020 |
| www-osoite |
Rahoitus
This work was supported by EU’s Horizon 2020 research and innovation programme via the project MeMAD (GA 780069). Computational resources were provided by Aalto Science-IT.
Sormenjälki
Sukella tutkimusaiheisiin 'Releasing a toolkit and comparing the performance of language embeddings across various spoken language identification datasets'. Ne muodostavat yhdessä ainutlaatuisen sormenjäljen.Projektit
- 1 Päättynyt
-
MeMAD: Methods for Managing Audiovisual Data: Combining Automatic Efficiency with Human Accuracy
Kurimo, M. (Vastuullinen johtaja), Grönroos, S.-A. (Projektin jäsen), Grósz, T. (Projektin jäsen), Brander, T. (Projektin jäsen), Porjazovski, D. (Projektin jäsen), Virkkunen, A. (Projektin jäsen), Choudhary, S. (Projektin jäsen), Xu, Z. (Projektin jäsen), Rouhe, A. (Projektin jäsen), Lindgren, M. (Projektin jäsen) & Raitio, R. (Projektin jäsen)
27/12/2017 → 31/03/2021
Projekti: EU: Framework programmes funding
Laitteet
Siteeraa tätä
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver