CORTEXA
← Browse
openalexMendeley Data2026-07-23Cited by 0

UFTBESD: University of Frontier Technology, Bangladesh-Bangla Emotional Speech Dataset

Md Rayhan Ali, Sadman Saeef, Mohammad Miftahul Islam Irfan Mohammad, Maliha Khan, Suchi Hasan, Sajib Das, Tanjim Taharat Aurpa, Sharad Hasan

UFTBESD (University of Frontier Technology, Bangladesh - Bangla Emotional Speech Dataset) is a Bangla-language speech emotion recognition dataset developed to capture realistic emotional speech under everyday acoustic conditions. The dataset consists of 1,400 audio recordings collected from 100 native Bangla speakers aged 15 to 38 years (mean age 22.6 years), with a balanced gender distribution (50 male and 50 female), recruited from all eight administrative divisions of Bangladesh. Each participant read two predefined Bangla sentences, with each sentence recorded once under each of seven target emotional states: angry, disgust, fear, happy, neutral, sad, and surprise. This yields 100 speakers x 2 sentences x 7 emotions = 1,400 audio clips. Every speaker contributes exactly 14 recordings, and every emotion category contains exactly 200 recordings (100 female, 100 male). All recordings were made using participants' own consumer smartphones in real indoor and outdoor environments (dormitory rooms, kitchens, courtyards, roadsides, etc.), without acoustic isolation or signal conditioning at any stage. A substantial portion of the corpus therefore contains natural background noise, device-specific microphone characteristics, and room reverberation, making it well suited to real-world speech emotion recognition research. Audio files are stored uncompressed as WAV (44.1 kHz, 16-bit PCM, mono), with durations ranging from 0.74 to 6.04 seconds (mean 2.68 s). This version corrects folder and file naming inconsistencies present in the previous release and adds a full metadata CSV file (one row per recording), covering speaker ID, pseudonym, age, gender, division, district, native dialect, education level, sentence text (Bangla and English), emotion label, recording environment, device brand/model, and audio properties. UFTBESD is intended to support the development and evaluation of Bangla speech emotion recognition systems and can be used with common machine learning and deep learning architectures such as CNN, LSTM, BiLSTM, and transformer-based models.

View free PDFSource page

Related papers

openalexMendeley Data2026-07-25

BanglaVowelDataset: True AC and Synthetic BC Bangla Vowel Datasets in Speech Information System

Ohidujjaman, Bejoy Munshi, Md. Mainul Hasan, M M Huda, Suman Ahmmed, Hasan Sarwar

The BanglaVowelDataset [1] resolves the unavailability of Bangla AC and synthetic BC vowel data in the speech information system. We recorded raw Bangla air-conducted (AC) vowels, with five male and five female speakers participating in the recording system, set up in a soundproo…

View free PDFSource page
openalexMendeley Data2026-07-23

MachDep-PhysX: a survey dataset on automation dependency and physical fitness decline among university students

Md Fahim Ferdous, Fardia Akter Omi, Md. Mehedi Hasan

This dataset contains responses from a structured questionnaire-based survey designed to investigate the relationship between automation/technology dependency and physical fitness decline among university undergraduate students in Bangladesh. The survey was administered as a bili…

View free PDFSource page
openalexMendeley Data2026-07-23

Title: Bridge or Substitute? Generative AI, Self-Diagnosis and Health Equity among Nigerian University Students

Suraj ibrahim

This dataset contains anonymised responses from an online qualitative survey examining how Nigerian university students use generative artificial intelligence (GenAI) tools such as ChatGPT and Meta AI for self-diagnosis and health management. The survey was completed by 196 respo…

View free PDFSource page
openalexMendeley Data2026-07-23

NightLPD-BD: Annotated Bangladeshi Nighttime Vehicle License Plate Dataset

Md. Jobayer Ahmed, Md Ali Emam Al Mamun, Md Naimul Islam Nuhash

NightLPD-BD is a real-world nighttime vehicle license plate dataset comprising 401 images collected from urban roadside environments in Dhaka, Bangladesh. The dataset includes manually annotated license plate regions using both bounding box and polygon segmentation annotations, p…

View free PDFSource page
openalexMendeley Data2026-07-23

Bengali Idiom Detection: A BIO-Annotated Dataset for Figurative Language Processing

Rahul Chandra Shil, Utsab Kumar Saha, Farzana Alam, Kazi Tanvir, Kamruddin Nur

This dataset includes a corpus of the Bengali language for idiom identification and sequence tagging, in which idioms are identified using token-based BIO tagging. The dataset contains 30,100 sentences covering both idiomatic and non-idiomatic usages across diverse contexts. Ther…

View free PDFSource page