CORTEXA
← Browse
openalexFrontiers in Artificial Intelligence2026-07-24Cited by 0

Do language families matter? Evaluating LLMs for sentiment analysis through a hierarchical cross-lingual lens

Muhamet Kastrati, Abdul Manaf, Ali Shariq Imran, Zenun Kastrati, Sher Muhammad Daudpota, Marenglen Biba

Social media sentiment analysis has become one of the most significant instruments for understanding the opinion of the population in the spheres of healthcare, politics, and education. Yet, large language models (LLMs) remain unevenly distributed in their linguistic coverage, failing to adequately serve a large portion of the world's languages. This study evaluates five state-of-the-art LLMs: GPT-4o, Gemini 2.0 Flash, DeepSeek-V3, Mistral Large, and Claude 3.7 Sonnet on three-class sentiment classification across 36 datasets spanning 36 languages, with emphasis on low- and medium-resource settings, using zero-shot and few-shot prompting without task-specific fine-tuning. In addition to the traditional measures of performance per language, the study presents a hierarchical analysis of languages based on a genealogical tree of Indo-European, Afro-Asiatic, Niger-Congo, Turkic, Austronesian, and English Creole language families, so that it is possible to identify the systematic patterns of performance superiority and inferiority among the language families. The findings show that few-shot prompting improves the results of a vast majority of languages, with several models approaching or surpassing the performance of the state-of-the-art benchmark of task-specific models. The GPT-4o and Claude achieved the highest performance in the high-resource and medium-resource settings, and Gemini is a competent trade-off that allows balancing the performance and the computational cost. Although it has lower zero-shot performance, Mistral benefits the most from few-shot prompting and becomes highly competitive in the few-shot setting. Despite these developments, the level of performance on low-resource languages, such as Oromo, Xitsonga, Azerbaijani, and Twi, remains significantly lower, underscoring that progress in multilingual LLMs requires moving beyond English-centric evaluation toward genuinely representative and globally inclusive benchmarks.

View free PDFSource page

Related papers

openalexFrontiers in Artificial Intelligence2026-07-23

HIDANet: a lightweight deep learning framework for Vannamei post-larval stage classification and morphometric estimation with background bias validation

Sugunapriya A, Markkandan S

Introduction Quality control of hatchery production relies on accurate developmental staging of the Pacific white shrimp Litopenaeus vannamei post-larvae (PL), but current methods rely on subjective manual visual evaluation that leads to observer bias and inconsistency. Methods I…

View free PDFSource page
openalexFrontiers in Artificial Intelligence2026-07-23

The VIBE-HI framework: a conceptual model for evaluating vibe coding appropriateness, quality, and safety in health informatics

Ahmed Alqheedan, Saleh Alzughaibi

Background Vibe coding—generating software through natural-language prompts to large language models without reviewing the underlying code—has moved rapidly from consumer technology into peer-reviewed clinical applications. By early 2026, clinicians had published vibe-coded teach…

View free PDFSource page
openalexFrontiers in Artificial Intelligence2026-07-23

From mechanistic models to artificial intelligence: exploring the potential of digital twins in geriatric oncology

Panagiotis Karampelesis, Spyros Denazis, Odysseas Koufopavlou, Evangelia I. Zacharaki

This survey explores how machine learning and artificial intelligence (AI) can be integrated with mechanistic models to create more accurate, dynamic, predictive, and personalized representations of biological systems, commonly referred to as digital twins (DTs). Mechanistic mode…

View free PDFSource page
openalexFrontiers in Artificial Intelligence2026-07-24

Mapping seasonal dynamics of forage and cereal crops in a hyper-arid environment using Sentinel-1 and Sentinel-2 time series

Areej Alwahas, Kasper Johansen, Jorge Rodriguez, Matthew F. McCabe

Introduction In arid and hyper-arid regions, agriculture depends heavily on irrigation, making crop type monitoring important for water allocation, monitoring crop management policies, and providing the information required to forecast food supply. However, field labels are often…

View free PDFSource page
openalexFrontiers in Artificial Intelligence2026-07-24

AI-based secure event-driven serverless architecture for scalable digital civic participation platform

Aizhan Kassymova, Abdul Razaque, Raissa Uskenbayeva, Z. B. Kalpeyeva, Aizhan Anartayeva

Introduction With the growing digitalization of urban governance and the increasing demand for transparency, sustainability and secure decision-making, the need for scalable and intelligent digital civic platforms has been raised. However, current e-participation systems are ofte…

View free PDFSource page
openalexFrontiers in Artificial Intelligence2026-07-23

Hybrid fuzzy C-means and deep learning framework for intelligent fault classification in solar PV systems

Vignesh V, R. Senthil Kumar, G. Suganeshwari

Photovoltaic (PV) systems have proven themselves to be a viable alternative energy source; however, there are multiple faults related to PV systems which cause energy losses and low efficiencies. Manual or rule-based algorithms are traditionally used for fault diagnosis, which ar…

View free PDFSource page