A Conversational Ethics Bot (CALEB) Versus GPT-4o for Medical Ethics Education: Transcript Analysis and Extended Reality Feasibility Study.

TitleA Conversational Ethics Bot (CALEB) Versus GPT-4o for Medical Ethics Education: Transcript Analysis and Extended Reality Feasibility Study.
Publication TypeJournal Article
Year of Publication2026
AuthorsJotwani R, Barot VD, Goldstein PA, Sigaras A, Sriram S, White RS, Mukherjee D, Gabbay E, Chan JM, Schnall JSteigerwal, Rubin JE
JournalJ Med Ext Real
Volume3
Pagination29941520261469342
Date Published2026 Jan-Dec
ISSN2994-1520
Abstract

Ethics education in medical training remains difficult to standardize and sustain. Many curricula still rely heavily on didactic teaching rather than immersive ethical reasoning. Although large language models (LLMs) can generate structured analyses of ethical dilemmas, they are not designed to facilitate embodied, conversational engagement that mirrors real-world ethics discussions. We developed CALEB (Conversational Agent Learning Ethics Bot), a domain-specific, extended reality (XR)-enabled conversational agent designed to simulate pragmatic, case-based moral deliberation. CALEB integrates a curated medical ethics knowledge base, structured persona design, and bounded generative architecture to promote focused, dialogical interaction. We conducted a two-phase feasibility evaluation. In phase 1, CALEB and GPT-4o accessed through the ChatGPT interface were compared using standardized transcript outputs generated from matched medical ethics prompts. In phase 2, medical ethicists interacted with CALEB in a live XR setting and provided formative post-session feedback. GPT-4o generated substantially longer and comprehensive responses. In contrast, CALEB produced significantly shorter but more principle-dense responses and was the only system to consistently demonstrate first-person and emotionally interpretive language. Computational analysis revealed higher empathy scores for CALEB in interpretive dimensions (p < 0.001). In the live XR phase, post-session feedback suggested that CALEB was perceived more favorably in an interactive, embodied setting than in transcript-only review. This study provides evidence that domain-specific agents like CALEB are feasible and pedagogically distinct from general-purpose LLMs. Foundational and specialized systems may serve complementary roles in advancing scalable, interactive medical ethics education.

DOI10.1177/29941520261469342
Alternate JournalJ Med Ext Real
PubMed ID42558094
PubMed Central IDPMC13396201