Description
A dataset of parallel fullband speech simultaneously recorded by various microphones in the Wilska Multichannel Anechoic Chamber at the Acoustics Lab of Aalto University, Espoo, Finland.
The source data comes from the VCTK dataset. It was played back using a high quality professional fullrange loudspeaker and recorded by two headset microphones and six studio microphones. The total amount of data per microphone is 23 hours, 11 minutes and 13 seconds. All data was recorded at 48kHz and 24-bit, and later on compressed in .flac format (loseless compression).
Date structure
The data is structured as follows:
headset-mics: Folder containing headset microphone recordings. Each folder contains data recorded by a single microphone.
studio-mics: Folder containing studio microphone recordings. Each folder contains data recorded by a single microphone.
Each folder contains the same number of files. Files with the same name across folders are parallel recordings, meaning that they have identical content and the same number of samples, but captured by each corresponding microphone.
Each file uses the following naming convention:
pXXX_YYY_ZZ.flac
where:
pXXX is the speaker ID, matching exactly the IDs used in the VCTK dataset.
YYY is the utterance ID, also consistent with the VCTK dataset.
ZZ is the segment ID. Utterances were trimmed to remove silent sections, and some were split into multiple segments as a result (e.g., p225_009_00.flac and p225_009_01.flac).
Each folder contains a meta.csv file with metadata for its corresponding subset, including a SHA256 hash to verify data integrity. These were generated using sndls with the following command:
sndls audio --csv meta.csv -e .flac -r --sha256
The source data comes from the VCTK dataset. It was played back using a high quality professional fullrange loudspeaker and recorded by two headset microphones and six studio microphones. The total amount of data per microphone is 23 hours, 11 minutes and 13 seconds. All data was recorded at 48kHz and 24-bit, and later on compressed in .flac format (loseless compression).
Date structure
The data is structured as follows:
headset-mics: Folder containing headset microphone recordings. Each folder contains data recorded by a single microphone.
studio-mics: Folder containing studio microphone recordings. Each folder contains data recorded by a single microphone.
Each folder contains the same number of files. Files with the same name across folders are parallel recordings, meaning that they have identical content and the same number of samples, but captured by each corresponding microphone.
Each file uses the following naming convention:
pXXX_YYY_ZZ.flac
where:
pXXX is the speaker ID, matching exactly the IDs used in the VCTK dataset.
YYY is the utterance ID, also consistent with the VCTK dataset.
ZZ is the segment ID. Utterances were trimmed to remove silent sections, and some were split into multiple segments as a result (e.g., p225_009_00.flac and p225_009_01.flac).
Each folder contains a meta.csv file with metadata for its corresponding subset, including a SHA256 hash to verify data integrity. These were generated using sndls with the following command:
sndls audio --csv meta.csv -e .flac -r --sha256
| Koska saatavilla | 2 toukok. 2025 |
|---|---|
| Julkaisija | Zenodo |
Dataset Licences
- CC-BY-SA-4.0
Siteeraa tätä
- DataSetCite