Do language families matter? Evaluating LLMs for sentiment analysis through a hierarchical cross-lingual lens
Muhamet Kastrati, Abdul Manaf, Ali Shariq Imran, Zenun Kastrati, Sher Muhammad Daudpota, Marenglen Biba
Social media sentiment analysis has become one of the most significant instruments for understanding the opinion of the population in the spheres of healthcare, politics, and education. Yet, large language models (LLMs) remain unevenly distributed in their linguistic coverage, failing to adequately serve a large portion of the world's languages. This study evaluates five state-of-the-art LLMs: GPT-4o, Gemini 2.0 Flash, DeepSeek-V3, Mistral Large, and Claude 3.7 Sonnet on three-class sentiment classification across 36 datasets spanning 36 languages, with emphasis on low- and medium-resource settings, using zero-shot and few-shot prompting without task-specific fine-tuning. In addition to the traditional measures of performance per language, the study presents a hierarchical analysis of languages based on a genealogical tree of Indo-European, Afro-Asiatic, Niger-Congo, Turkic, Austronesian, and English Creole language families, so that it is possible to identify the systematic patterns of performance superiority and inferiority among the language families. The findings show that few-shot prompting improves the results of a vast majority of languages, with several models approaching or surpassing the performance of the state-of-the-art benchmark of task-specific models. The GPT-4o and Claude achieved the highest performance in the high-resource and medium-resource settings, and Gemini is a competent trade-off that allows balancing the performance and the computational cost. Although it has lower zero-shot performance, Mistral benefits the most from few-shot prompting and becomes highly competitive in the few-shot setting. Despite these developments, the level of performance on low-resource languages, such as Oromo, Xitsonga, Azerbaijani, and Twi, remains significantly lower, underscoring that progress in multilingual LLMs requires moving beyond English-centric evaluation toward genuinely representative and globally inclusive benchmarks.