BanglaVowelDataset: True AC and Synthetic BC Bangla Vowel Datasets in Speech Information System
Ohidujjaman, Bejoy Munshi, Md. Mainul Hasan, M M Huda, Suman Ahmmed, Hasan Sarwar
The BanglaVowelDataset [1] resolves the unavailability of Bangla AC and synthetic BC vowel data in the speech information system. We recorded raw Bangla air-conducted (AC) vowels, with five male and five female speakers participating in the recording system, set up in a soundproof room. Male and female speakers were trained by experts to learn the proper pronunciation of Bangla vowels. The positions of the equipment are adjusted at specific distances. The speaker was sitting in front of the AC microphone, away from 30 cm. Afterward, we converted the recorded true AC vowels into synthetic BC vowels using an all-pole source filter model to obtain a synthetic Bangla BC dataset. The length of each recorded raw Bangla vowel ranges from 600 to 780 ms. The sampling frequency, audio resolution correspond to 44100 Hz and 32-bit, respectively. Raw speech data are recorded in .wav format. We set the stereo in Audacity audio recording software while recording vowels to have a more natural and immersive experience. We obtained a total of 120 true AC vowels from five male and five female speakers, as the Bangla alphabet has twelve vowels. The recoded each vowel is cropped to avoid the unvoiced portions. Thus, the cropped vowel length ranges from 500 to 650 ms. In addition, we created a synthetic dataset of the BC Bangla vowels containing 120 vowels sampled at 8000 Hz, having discarded the unvoiced portions. In this paper, we focus on two datasets: raw data consisting of 120 recorded Bangla AC vowels and synthetic data containing 120 synthetic Bangla BC vowels. There are twelve vowels in the Bangla alphabet due to diphthongs and variations. In the literature, the Bangla vowel speech dataset is not available yet. Consequently, this dataset provides a significant opportunity for further research. Bangla vowel dataset is used for noisy and clean environments comparatively, for speech recognition and speaker identification. Commonly, AC speech is severely affected by ambient noise, whereas BC speech suffers less. Thus, the appropriate method is suggested for noise robustness. This dataset has potential for pitch (F0) detection, spectrum estimation, and condition number determination in speech signal processing and analysis. Performance of distinct approaches, including machine learning, deep learning, and statistical methods, is evaluated using raw AC and synthetic BC vowels. The complete dataset is publicly accessible on the Mendeley Data repository, organized hierarchically with separate directories for all Bangla AC and BC vowels.