Skip to main navigation Skip to search Skip to main content

Teflon Pronunciation Assessment Challenge

  • Giampiero Salvi (Creator)
  • Tamás Grósz (Creator)

Dataset

Description

Speech data for the Teflon Pronunciation Assessment Challenge. This dataset contains recordings of isolated words in Norwegian spoken by children in the age 4-12. Some of the children are motherthongue Norwegian, others are immigrants with different linguisti backgrounds. The fine names were randomized to hide the identity of the speakers. The pronunciation of each utterance has been assessed by at least one assessor on a scale 1-5 with the following meaning: Not at all identifiable as the target word Difficult to identify as the target word Slight phonemic error(s) Subphonemic error(s) or "unexpected variants" Prototypical, adult-like The data is divided into train and test. The corresponding CSV files contain, for each utterance, information about the spoken word and the score. Scores are only given for the training data.
Date made available31 Oct 2024
PublisherZenodo

Dataset Licences

  • CC-BY-4.0
  • Collecting Linguistic Resources for Assessing Children's Pronunciation of Nordic Languages

    Olstad, A. M. H., Smolander, A., Strömbergsson, S., Ylinen, S., Lehtonen, M., Kurimo, M., Getman, Y., Grosz, T., Cao, X., Svendsen, T. & Salvi, G., 2024, 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings. Calzolari, N., Kan, M.-Y., Hoste, V., Lenci, A., Sakti, S. & Xue, N. (eds.). European language resources distribution agency, p. 3529-3537 9 p. (International conference on computational linguistics)(LREC proceedings).

    Research output: Chapter in Book/Report/Conference proceedingConference article in proceedingsScientificpeer-review

    Open Access

Cite this