CORTEXA
← Browse
arxivcs.HC2026-07-10

Reading the Eyes in VR: Multimodal Modeling of Social Intelligence

Mohammad Fahim Abrar, Shayla Sharmin, Roghayeh Leila Barmaki

Social intelligence, the ability to interpret others' emotions, beliefs, and intentions, is often assessed with the Reading the Mind in the Eyes Test (RMET), in which participants infer mental states from images of the eye region. Yet RMET is typically presented on paper or desktop displays, where viewing geometry can vary across participants, and it rarely includes immediate feedback. We investigated whether presentation medium and brief trial-level feedback influence RMET behavior. We implemented RMET in Unity for both desktop and Virtual Reality (VR), using VR to hold stimulus distance and field of view constant without changing the items. We conducted a 2x2 mixed study with 20 participants, with device (VR vs. desktop) manipulated between subjects and feedback (immediate correctness cue vs. none) manipulated within subjects. Eye-tracking and EEG data were recorded and synchronized with behavioral logs. We analyzed fixation-based gaze measures, EEG signals, response time, accuracy, and subjective measures. Immediate feedback was associated with longer fixation durations and higher EEG-based engagement, while no significant differences were observed in task completion time or total correct answers. Presentation medium did not produce reliable differences in the objective measures, but VR received higher usability ratings and was also rated as more effortful. These results provide initial evidence that RMET can be studied as a process-aware assessment task in controlled VR and desktop settings.

View free PDFSource page

Related papers

arxivcs.AIcs.CLcs.HC2026-07-16

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

Patrick Phuoc Do, Chau M. Ta, Chaoli Wang

Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualizat…

View free PDFSource page
arxivcs.CVcs.HC2026-07-22

MV-Bench: Benchmarking Multimodal Large Language Models for Coordinated Multi-View Interface Construction

Yue Zhao, Hongxu Liu, Feiyu Wang, Xiaoyu Yang, Tong Ge, Zhen Yang, et al.

Multimodal large language models (MLLMs) are increasingly expected to automate visualization development by generating code directly from visual designs. However, existing evaluations mainly focus on single-chart generation and overlook coordinated multi-view interface constructi…

View free PDFSource page
arxivcs.ROcs.HC2026-07-11

Navigating the Crowd: Non-linear MPC with Social Forces Dynamics for Human-Aware Robot Navigation

Stefano Trepella, Andrea Ostuni, Mauro Martini, Pablo Pueyo, Noé Pérez-Higueras, Marcello Chiaberge, et al.

Safe and socially compliant navigation remains a fundamental challenge for autonomous robots operating in human-populated environments. Beyond collision avoidance, robots must anticipate human motion and respect personal space to ensure human comfort. Model Predictive Control (MP…

View free PDFSource page
arxivcs.HCcs.ET2026-07-16

Assessing Learning Processes with Multimodal Data in Virtual Reality Learning Environments

Eileen McGivney, Oluwatomilade Olarinde, Erica Kleinman, Kaylah Facey, Rana Jahani, Shripad Agashe, et al.

Assessing learning in virtual reality (VR) environments typically relied on traditional pre-post content retention tests, revealing little about the process of learnng within such immersive environments. Multimodal data from player activity in VR is promising to better measure le…

View free PDFSource page
arxivcs.CLcs.AIcs.HCcs.LG2026-07-09

LEXIC: Lightweight Eye-tracking eXtension via Injected Complexity

Sumin Lee, Kyeonghun Kim, Subeen Lee, Jiwon Yang, Tien Nguyen, Ken Ying-Kai Liao, et al.

On the recent EyeBench benchmark, predicting reading comprehension from eye movements exposes a stark gap: text-aware models using pretrained language models reach 56--63% AUROC, while gaze-only models operate at chance. We ask how far a gaze-only model can be pushed by lightweig…

View free PDFSource page
arxivcs.HCcs.AI2026-07-21

Public perceptions of AI-driven decision-making in healthcare: A structural equation modeling approach

Leonie Westerbeek, Ernesto de Leon, Julia C. M. van Weert

Artificial intelligence (AI) is increasingly integrated into healthcare to support diagnostics, decision-making, and administrative processes. However, the successful implementation of AI depends not only on technical performance but also on public perceptions of its helpfulness,…

View free PDFSource page