Constarium
← Search

Table · dataset · 2026

Supplementary file 1_Public-facing health information on achalasia generated by large language models: a multidimensional comparative study.docx

Listed in figshare

Background<p>Large language models (LLMs) are increasingly used by the public to obtain health information, but their ability to provide reliable information for achalasia remains unclear.

Description

We therefore compared six LLMs in answering public questions about achalasia, an uncommon esophageal motility disorder, focusing on safety, accuracy, empathy, information quality and reliability, and readability.</p>Methods<p>In this cross-sectional comparative study, 40 patient-oriented questions about achalasia were submitted once to each of six LLMs (ChatGPT 5.5, Claude Opus 4.7, DeepSeek V4 Pro, Gemini 3.1 Pro Thinking, Grok-4.3, and Qwen3-Max), between May 7 and May 12, 2026.

Three gastroenterologists independently assessed 240 responses for safety, accuracy, empathy, information quality, and readability using DISCERN, EQIP, the JAMA benchmark criteria assessing authorship, attribution, disclosure, and currency, and the Global Quality Score (GQS). Readability was evaluated using six established readability indices.</p>Results<p>Overall, 27 of 240 responses (11.3%) were classified as potentially unsafe, with proportions ranging from 7.5 to 15.0% across models; the overall between-model difference in safety was not statistically significant.

Read the rest (2 more)

Accuracy differed significantly among models, although the effect size was small. Larger differences were observed in empathy, information reliability and quality, and readability. ChatGPT 5.5 and Gemini 3.1 Pro Thinking achieved higher empathy scores, whereas Claude Opus 4.7 showed greater reading difficulty.

JAMA benchmark scores were low overall, indicating limited source transparency.</p>Conclusion<p>Current LLMs provided generally accurate and often useful answers to public questions about achalasia. However, some differences were found in safety, empathy, transparency, information quality, and readability. Public-facing LLM responses to public questions about achalasia should be consistent with guidelines, clearly cite sources, communicate with empathy, and use plain language.</p>

Links

Where it is published

Catalogue records · 1

Topics

Provenance · 1 source records, 17 field assertions
SourceKeyLast seenRaw
figshareoai:figshare.com:article/337715988 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · figshare comconnector:figshare_com@1.0.0
concepts[field].anzsrc:group:4201enrichment · figshare comtaxonomy-embedding@1.0.0title+keywords+description (71%)
concepts[field].local:field:astronomymapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:chemistrymapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:computer-science-aimapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:earth-environmentalmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:economics-financemapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:engineeringmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:humanitiesmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:life-sciencesmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:medicine-healthmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:ocean-atmosphericmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:social-sciencemapping · figshare comconnector:figshare_com@1.0.0
descriptionsource · figshare comconnector:figshare_com@1.0.0/metadata/dc/description
licensesource · figshare comconnector:figshare_com@1.0.0/metadata/dc/rights
publication_datesource · figshare comconnector:figshare_com@1.0.0
titlesource · figshare comconnector:figshare_com@1.0.0/metadata/dc/title