Projekteja vuodessa
Abstrakti
In this study, formant tracking is investigated by refining the formants tracked by an existing data-driven tracker, DeepFormants, using the formants estimated in a model-driven manner by linear prediction (LP) -based methods. As LP-based formant estimation methods, conventional covariance analysis (LP-COV) and the recently proposed quasi-closed phase forward-backward (QCP-FB) analysis are used. In the proposed refinement approach, the contours of the three lowest for-
mants are first predicted by the data-driven DeepFormants tracker, and the predicted formants are replaced frame-wise with local spectral peaks shown by the model-driven LP-based methods. The refinement procedure can be plugged into the DeepFormants tracker with no need for any new data learning. Two refined DeepFormants trackers were compared with the original DeepFormants and with five known traditional trackers using the popular vocal tract resonance (VTR) corpus. The results indicated that the data-driven DeepFormants trackers outperformed the conventional trackers and that the best performance was obtained by refining the formants predicted by DeepFormants using QCP-FB analysis. In addition, by tracking formants using VTR speech that was corrupted
by additive noise, the study showed that the refined DeepFormants trackers were more resilient to noise than the reference trackers. In general, these results suggest that LP-based model-driven approaches, which have traditionally been used in formant estimation, can be combined with a modern data-driven tracker easily with no further training to improve the tracker’s performance.
mants are first predicted by the data-driven DeepFormants tracker, and the predicted formants are replaced frame-wise with local spectral peaks shown by the model-driven LP-based methods. The refinement procedure can be plugged into the DeepFormants tracker with no need for any new data learning. Two refined DeepFormants trackers were compared with the original DeepFormants and with five known traditional trackers using the popular vocal tract resonance (VTR) corpus. The results indicated that the data-driven DeepFormants trackers outperformed the conventional trackers and that the best performance was obtained by refining the formants predicted by DeepFormants using QCP-FB analysis. In addition, by tracking formants using VTR speech that was corrupted
by additive noise, the study showed that the refined DeepFormants trackers were more resilient to noise than the reference trackers. In general, these results suggest that LP-based model-driven approaches, which have traditionally been used in formant estimation, can be combined with a modern data-driven tracker easily with no further training to improve the tracker’s performance.
Alkuperäiskieli | Englanti |
---|---|
Artikkeli | 101515 |
Sivumäärä | 11 |
Julkaisu | Computer Speech and Language |
Vuosikerta | 81 |
Varhainen verkossa julkaisun päivämäärä | 24 maalisk. 2023 |
DOI - pysyväislinkit | |
Tila | Julkaistu - kesäk. 2023 |
OKM-julkaisutyyppi | A1 Julkaistu artikkeli, soviteltu |
Sormenjälki
Sukella tutkimusaiheisiin 'Refining a Deep Learning-based Formant Tracker using Linear Prediction Methods'. Ne muodostavat yhdessä ainutlaatuisen sormenjäljen.Projektit
- 1 Aktiivinen
-
HEART: Speech-based biomarking of heart failure
Alku, P., Javanmardi, F., Mittapalle, K., Tirronen, S., Kadiri, S., Pohjalainen, H. & Kodali, M.
01/09/2020 → 31/08/2024
Projekti: Academy of Finland: Other research funding
Lehtileikkeet
-
Researchers from Aalto University Detail New Studies and Findings in the Area of Information Technology (Refining a Deep Learning-based Formant Tracker Using Linear Prediction Methods)
22/09/2023
1 kohde/ Medianäkyvyys
Lehdistö/media: Esiintyminen mediassa