CORTEXA
← Browse
openalexMendeley Data2026-07-23Cited by 0

Bengali Idiom Detection: A BIO-Annotated Dataset for Figurative Language Processing

Rahul Chandra Shil, Utsab Kumar Saha, Farzana Alam, Kazi Tanvir, Kamruddin Nur

This dataset includes a corpus of the Bengali language for idiom identification and sequence tagging, in which idioms are identified using token-based BIO tagging. The dataset contains 30,100 sentences covering both idiomatic and non-idiomatic usages across diverse contexts. There are 20–25 sentences per idiom. For each idiom: • 60% of the sentences contain the idiomatic usage of the expression. • 40% of the sentences represent literal or non-idiomatic usage and are designed as hard negative samples. This dataset helps machine learning models distinguish between literal and figurative usages of idiomatic expressions. The corpus was developed through a systematic data collection and annotation process. Canonical Bengali idioms were collected from publicly available and authoritative resources, including the National Curriculum and Textbook Board (NCTB) textbook বাংলা ভাষার ব্যাকরণ ও নির্মিতি (Secondary Level, Grades 9–10) and the Bengali Wiktionary (উইকিঅভিধান) idiom category. These resources were selected to ensure broad coverage of commonly used Bengali idiomatic expressions while relying on openly accessible educational and lexical sources. LLM-based tools were used to facilitate the generation of repetitive sentence templates, improving contextual diversity and linguistic coverage. All generated sentences were manually reviewed, refined, and validated by the annotation team before inclusion in the dataset. Data Structure: The following JSON structure is used for every row in the dataset: • A unique ID for each instance • A Bengali sentence containing idiomatic or literal expressions • Tokenized representation of the sentence • BIO-formatted labels marking idiomatic spans (B-IDIOM, I-IDIOM, O) • A binary label indicating the presence of an idiom in the sentence Example format: { "id": 3310, "sentence": "মহান লেখকের মৃত্যু আমাদের জন্য এক ইন্দ্রপতন", "tokens": ["মহান", "লেখকের", "মৃত্যু", "আমাদের", "জন্য", "এক", "ইন্দ্রপতন"], "labels": ["O", "O", "O", "O", "O", "O", "B-IDIOM"], "is_idiom": 1 } To promote linguistic diversity and robustness, each idiom appears in multiple sentence variants with different syntactic structures and contextual usages. This corpus is suitable for sequence labeling, token classification, sentence classification, Bengali NLP using transformer-based language models, and figurative language research in low-resource settings. Bengali Idiom Lexicon : This database comprises a carefully curated collection of nearly 1,300 Bengali idioms along with their meanings. Each record consists of a unique identifier, the idiom, and its meaning in the Bangla language. Example: { "id": 1290, "idiom": "হাতে আকাশ পাওয়া", "meaning": "অভাবিতভাবে কিছু পাওয়া" } All data is provided in UTF-8 JSON Lines (.jsonl) format for easy integration into NLP pipelines and machine learning workflows.

View free PDFSource page

Related papers

openalexMendeley Data2026-07-23

NightLPD-BD: Annotated Bangladeshi Nighttime Vehicle License Plate Dataset

Md. Jobayer Ahmed, Md Ali Emam Al Mamun, Md Naimul Islam Nuhash

NightLPD-BD is a real-world nighttime vehicle license plate dataset comprising 401 images collected from urban roadside environments in Dhaka, Bangladesh. The dataset includes manually annotated license plate regions using both bounding box and polygon segmentation annotations, p…

View free PDFSource page
openalexMendeley Data2026-07-23

Longitudinal Indoor Air Quality Dataset Collected Using a Low-Cost Multi-Sensor IoT Monitoring Platform

Md Abubakar Siddique

This repository contains a longitudinal indoor air quality (IAQ) dataset collected with a Raspberry Pi 5–based multi-sensor IoT monitoring platform deployed in an indoor laboratory. The monitoring campaign spans approximately 15.6 days of continuous post-initialization operation…

View free PDFSource page
openalexMendeley Data2026-07-23

UFTBESD: University of Frontier Technology, Bangladesh-Bangla Emotional Speech Dataset

Md Rayhan Ali, Sadman Saeef, Mohammad Miftahul Islam Irfan Mohammad, Maliha Khan, Suchi Hasan, Sajib Das, et al.

UFTBESD (University of Frontier Technology, Bangladesh - Bangla Emotional Speech Dataset) is a Bangla-language speech emotion recognition dataset developed to capture realistic emotional speech under everyday acoustic conditions. The dataset consists of 1,400 audio recordings col…

View free PDFSource page
openalexMendeley Data2026-07-25

BanglaVowelDataset: True AC and Synthetic BC Bangla Vowel Datasets in Speech Information System

Ohidujjaman, Bejoy Munshi, Md. Mainul Hasan, M M Huda, Suman Ahmmed, Hasan Sarwar

The BanglaVowelDataset [1] resolves the unavailability of Bangla AC and synthetic BC vowel data in the speech information system. We recorded raw Bangla air-conducted (AC) vowels, with five male and five female speakers participating in the recording system, set up in a soundproo…

View free PDFSource page
openalexMendeley Data2026-07-23

Research Data for “Two-Fraction Kinetic Modelling of Ethyl Levulinate Production from Oil Palm Empty Fruit Bunches via Dimethyl Carbonate-Assisted Ethanolysis

heriyanti heriyanti

This dataset contains the experimental and processed data supporting the study entitled “Two-Fraction Kinetic Modelling of Ethyl Levulinate Production from Oil Palm Empty Fruit Bunches via Dimethyl Carbonate-Assisted Ethanolysis.” The dataset includes temperature- and time-depend…

View free PDFSource page