Constarium
← Search

Table · dataset · 2026

Supplementary file 2_A cross-sectional evaluation of large language model chatbot interfaces for patient-facing herpes zoster information: safety, information quality, and readability.docx

Listed in figshare

Background<p>Large language model (LLM) chatbots are increasingly used to obtain health information.

Description

However, fluent and clinically plausible responses may still contain safety-relevant omissions, inadequate source attribution and disclosure, or difficult-to-read text.</p>Objective<p>To evaluate the safety, accuracy, empathy, information quality, response-level source attribution and disclosure, overall quality, and readability of five LLM chatbot interfaces answering lay-oriented questions about herpes zoster.</p>Methods<p>This cross-sectional evaluation used 46 standardised English-language questions across seven herpes zoster domains.

Each question was submitted once to ChatGPT-5.5 Instant, DeepSeek-V4-Pro, Doubao-Seed-2.0-Pro, Gemini 3.5 Flash, and Qwen3.7-Plus under consumer-access conditions on June 2–3, 2026. Five dermatologists independently assessed safety, accuracy, empathy, DISCERN, Ensuring Quality Information for Patients (EQIP), Journal of the American Medical Association (JAMA) benchmark criteria, and the Global Quality Score (GQS). Six readability indices were calculated.

Read the rest (4 more)

Paired comparisons and inter-rater agreement were evaluated using prespecified statistical methods.</p>Results<p>The five interfaces generated 230 complete responses. Under the study conditions, 66 responses (28.7%) were classified as potentially unsafe using the prespecified ≥3/5-rater majority threshold, with no significant difference between interfaces (Cochran’s Q = 1.661, df = 4, p = 0.798). Sensitivity analyses using alternative thresholds showed consistent results.

Accuracy and empathy differed across interfaces (both p < 0.001; Kendall’s W = 0.127 and 0.211, respectively), although effect sizes were small. Information-quality and overall-quality measures also showed outcome-specific differences. All six readability indices differed across interfaces (all p < 0.001); DeepSeek-V4-Pro generally produced text estimated to be easier to read, whereas Gemini 3.5 Flash produced text estimated to be more difficult to read.

Safety agreement was substantial (Fleiss’ κ = 0.682), and ICC (2,1) values for other manually rated outcomes ranged from 0.832 to 0.889.</p>Conclusion<p>In this standardised benchmark, LLM chatbot interfaces provided generally favourable accuracy and information-quality scores but showed clinically relevant safety limitations, sparse source attribution and disclosure, and readability challenges. No interface consistently performed best across all outcomes.

These findings represent single first responses obtained under the tested conditions and do not establish response stability across repeated queries. LLM-generated herpes zoster information should therefore be interpreted cautiously and should not replace individualised professional assessment or professionally reviewed patient information.</p>

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Immunology 72% · Text 75%
Provenance · 1 source records, 18 field assertions
SourceKeyLast seenRaw
figshareoai:figshare.com:article/338290969 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · figshare comconnector:figshare_com@1.0.0
concepts[field].anzsrc:group:3204enrichment · figshare comtaxonomy-embedding@1.0.0title+keywords+description (72%)
concepts[field].local:field:astronomymapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:chemistrymapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:computer-science-aimapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:earth-environmentalmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:economics-financemapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:engineeringmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:humanitiesmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:life-sciencesmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:medicine-healthmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:ocean-atmosphericmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:social-sciencemapping · figshare comconnector:figshare_com@1.0.0
concepts[modality].local:modality:textenrichment · figshare comkeyword-concept-rules@1.0.0title+description (75%)
descriptionsource · figshare comconnector:figshare_com@1.0.0/metadata/dc/description
licensesource · figshare comconnector:figshare_com@1.0.0/metadata/dc/rights
publication_datesource · figshare comconnector:figshare_com@1.0.0
titlesource · figshare comconnector:figshare_com@1.0.0/metadata/dc/title