openalexMendeley Data2026-07-23
Bengali Idiom Detection: A BIO-Annotated Dataset for Figurative Language Processing
Rahul Chandra Shil, Utsab Kumar Saha, Farzana Alam, Kazi Tanvir, Kamruddin Nur
This dataset includes a corpus of the Bengali language for idiom identification and sequence tagging, in which idioms are identified using token-based BIO tagging. The dataset contains 30,100 sentences covering both idiomatic and non-idiomatic usages across diverse contexts. Ther…