CORTEXA
← Browse
arxivcs.CLcs.AI2026-07-20

Multilingual Sentence Embeddings for Linguistic-Integrated Reliability Audit

Ummugul Bezirhan, Ji Yoon Jung, Matthias von Davier

Multilingual assessment systems commonly rely on translation for scoring and quality-control processes. We evaluate whether multilingual sentence embeddings can replace translated English input for Linguistic-Integrated Reliability Auditing (LiRA) across 11 PIRLS constructed-response items and three embedding models. Native-language embeddings reproduced translation-based reliability estimates closely while recovering responses excluded after translation failure, with no meaningful change in reliability.

View free PDFSource page

Related papers

arxivcs.CLcs.AI2026-06-25

SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

Subham Kumar, Prakrithi Shivaprakash, Abhishek Manoharan, Astut Kurariya, Diptadhi Mukherjee, Prabhat Chand, et al.

Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare context remains largely unknown. In this study, we first conduct the systematic audit of ASR performance on r…

View free PDFSource page
arxivcs.CLcs.AIcs.LGcs.SD2026-07-07

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts

Andrei-George Durdun, Victor Constantinescu, Radu Tudor Ionescu

Automatically recognizing the sentiment, positive or negative, from speech is a challenging task, requiring both the analysis of vocal inflections and the interpretation of uttered words. Recent solutions rely on audio foundation models to solve the task, but it remains unclear i…

View free PDFSource page
arxivcs.AIcs.CL2026-07-07

Integrating knowledge graphs and multilingual scholarly corpora for domain-adaptive LLMs in SSH

Adam Faci, Alessio Miaschi, Anne Combe, Pascal Cuxac, Francesca Frontini, Nicolas Larrousse, et al.

The integration of Large Language Models (LLMs) into scientific research workflows, particularly for bibliographic discovery and literature synthesis, raises significant methodological, epistemic and regulatory challenges for the Social Sciences and Humanities (SSH), especially w…

View free PDFSource page
arxivcs.CLcs.AI2026-07-17

Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection

Indraveni Chebolu, Rohan Singh, Arnab Mallick, Harmesh Rana

Moderation systems increasingly rely on external toxicity tools, but those tools are unreliable under code-mixing, transliteration, slang, and language mismatch. We study the \emph{conditional reliability} of toxicity priors in Indian multilingual and code-mixed short text: Engli…

View free PDFSource page
arxivcs.CLcs.AI2026-07-09

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

Zongyou Yang, Yinghan Hou, Xiaokun Yang

An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed. We treat this evaluator-replacement ambiguity as a measurement-validity problem. Across four judgment datasets, we compare two upgrade paths available in practic…

View free PDFSource page
arxivcs.CLcs.AIcs.LG2026-07-05

Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders

Lucas Hideki Takeuchi Okamura, Alexandre Alcoforado, Anna Helena Reali Costa

Portuguese remains underrepresented in text embedding evaluation, despite being one of the most widely spoken languages in the world. As a result, embedding models are often selected based on English or multilingual metrics, while their effectiveness in Portuguese remains unclear…

View free PDFSource page