CORTEXA
← Browse
arxivcs.AIcs.CYstat.APstat.ME2026-06-27

Aristotelian Virtue Profiling of LLMs through Ethical Dilemmas

Ioannis Tzachristas, John Pavlopoulos

Large Language Models (LLMs) often face ethical tradeoffs in which several responses may be defensible but express different priorities, such as fairness, honesty, courage, or restraint. We introduce VirtueMap, a framework for describing these patterns through an Aristotelian virtue-ethics lens. Instead of asking for a single correct answer, VirtueMap asks humans or LLMs to rank all five responses to each of seven general, non-lethal, non-political, and non-religious ethical dilemmas. To define the reference orderings used for scoring, we first proposed, for each dilemma and virtue, an ordering of the five responses from most to least expressive of that virtue. We then collected more than 100 respondent evaluations per ordering and retained it as operational ground truth only when at least 95% confirmed it. Rankings are scored against these retained orderings using normalized Borda alignment, yielding profiles over Practical Wisdom, Justice, Truthfulness, Courage, and Temperance. We apply VirtueMap to nine LLM families in a repeated-run evaluation and find high mean rank consistency (90.3%), with the largest differences appearing on Courage, Temperance, and Justice. We also release an interactive website that computes profiles locally in the browser and compares respondents with measured LLM profiles.

View free PDFSource page

Related papers

arxivstat.MEcs.AIcs.MSq-fin.STstat.AP2026-07-07

tsbootstrap: Distribution-Free Uncertainty Quantification and Conformal Prediction for Time Series

Sankalp Gilda

Finance, sensing, and demand streams violate the exchangeability that IID conformal prediction and the IID bootstrap assume, and existing libraries implement either a general resampling engine or conformal calibration without the other. tsbootstrap provides block, residual, sieve…

View free PDFSource page
arxivcs.AIcs.CLcs.CYcs.MAstat.ME2026-07-03

Silicon Sampling via Cross-Survey Transfer

Chan-Tung Ku, Chan Hsu, Pei-Cing Huang, Frank Cheng-shan Liu, I-Ling Cheng, Yihuang Kang

Silicon sampling-using large language models (LLMs) to simulate human survey respondents-has emerged as a promising approach for augmenting traditional survey research. However, most evaluations rely on distributional comparisons rather than individual-level prediction, which ris…

View free PDFSource page
arxivcs.HCcs.AIcs.CYcs.ETcs.RO2026-07-14

Practical Judgment, Virtue, and Intuition in the Use of Opaque AI-Enabled Systems

Nathan G. Wood, Andrew P. Rebera

AI-enabled systems are seeing increasing deployment across numerous domains, with many being "black boxes" with respect to core functions and capabilities. I.e., many systems take inputs and give outputs, but without users having any ability to see how the former lead to the latt…

View free PDFSource page
arxivcs.CLcs.AIcs.CY2026-07-06

The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment

Haonan Huang

Large language models (LLMs) increasingly issue judgments read as binary verdicts, and a growing literature reports such judgments shifting under logically irrelevant changes of wording - among them an amplified yes-no bias on moral dilemmas, absent in humans. A single framing ca…

View free PDFSource page