CORTEXA
← Browse
arxivcs.HC2026-06-30Cited by 0

Evaluating Interactivity: Toward Automated Assessment of AI-Generated Explorable Explanations

Xiaozao Wang, Zhewei Wang, Hongyi Wen

While large language models now enable rapid generation of interactive learning materials, evaluating the interaction quality of these explorable explanations remains an open challenge. Existing benchmarks largely focus on code executability or visual fidelity, providing limited insight into dynamic interaction behaviors such as learner-controlled state transitions and context-sensitive system responses, which are factors that critically shape learners' conceptual understanding. We present EE-Eval, an automated evaluation framework that formalizes interactivity as a finite space of learner-controllable states and transitions, represented as a Finite State Machine (FSM). By extracting FSMs from AI-generated explorable explanations, EE-Eval externalizes implicit interaction logic into an explicit, machine-interpretable graph. Evaluation is performed by comparing each generated FSM to an ideal FSM that encodes pedagogical intent, using a combination of graph-based metrics and embedding-based comparison of states, actions, and feedback to measure their structural and semantic similarity. Across thousands of generated explorable explanations spanning 127 concepts and produced by 6 AI models, EE-Eval consistently differentiates interaction quality beyond surface-level criteria such as functional correctness or visual quality, and exhibits substantially stronger alignment with human judgments of interactivity and pedagogical effectiveness than existing baselines. By framing interactivity as testable behavioral models rather than an emergent byproduct of LLM generation, EE-Eval transforms evaluation into a reflective diagnostic tool, enabling pedagogically grounded and actionable human-AI collaboration in creating interactive educational content.

View free PDFSource page

Related papers

arxivcs.IRcs.HC2026-07-03

AI Overviews in Academic Search: Evaluating AI-generated Summaries of Search Results in a Domain-specific Search Engine

Kevin Schott, Kanishka Silva, Ingo Frommholz, Philipp Mayr, Dagmar Kern, Daniel Hienert

Evaluating search engine results pages (SERPs) to assess result relevance is a demanding step in academic search. In a formative mixed-methods design study, we examine AI-generated SERP-level summaries as a support feature in an academic search engine for social science informati…

View free PDFSource page
arxivcs.HC2026-06-29

Concept Catalyst: Exploring Scrutable Interfaces to Structure K-12 Teacher Interactions with Generative AI

Gennie Mansi, Sunni Newton, Roxanne Moore, Meltem Alemdar, Mark Riedl

Purpose: This paper explores how to align AI-based tools with teachers' classroom needs by using scrutable interfaces -- interfaces that link an easily manipulable knowledge representation to an underlying AI model, so users can change the system's outputs without understanding i…

View free PDFSource page
arxivcs.HC2026-07-07

Exploring the Interaction of Explanation Styles, Context, and Trust of AI Privacy Redaction in AI-mediated Interactions

Roshni Kaushik, Maarten Sap, Koichi Onoue

AI-mediated communication is increasingly being utilized to help facilitate interactions; however, in privacy sensitive domains, an AI mediator has the additional challenge of considering how to preserve privacy. In these contexts, a mediator may redact or withhold information, r…

View free PDFSource page
arxivcs.HCcs.AIcs.CY2026-06-26

Generative AI Literacy Training Improves Intelligence Analysts' Discrimination of Real and AI-Generated Images

Negar Kamali, Candice Rockell Gerstner, Jessica Hullman, Matthew Groh

Across social and online platforms, people are increasingly exposed to AI-generated images. As a consequence, the task of distinguishing AI-generated from authentic images is becoming a central challenge for information ecosystems. While humans perform better than chance, accurac…

View free PDFSource page