Constarium
← Search

Table · dataset · 2026

Table 1_Benchmarking large language models on a Chinese radiation oncology technology examination-preparation question set: accuracy, consensus, and efficiency for AI assisted education.docx

Listed in figshare and Loughborough Research Repository — shown once because both records carry DOI 10.3389/fonc.2026.1946351.s001

Purpose<p>This study evaluated GPT-4o, GPT-5.4, and two DeepSeek platform configurations using 1, 053 examination-preparation questions from a publicly and commercially available 2025 exercise collection for the Chinese National Radiation Oncology Technology Qualification Examination (Intermediate Level).

Description

We assessed accuracy rates, identical-response coverage, accuracy consensus, and model-reported latency estimates.</p>Methods<p>Four configurations—GPT-4o, GPT-5.4, DS-Fast, and DS-Expert—were tested using an identical zero-shot Chinese prompt.

Accuracy rates were summarized with Wilson 95% confidence intervals. Pairwise differences in the primary Chinese-language analysis were assessed using exact McNemar’s tests with Holm adjustment. A descriptive language-controlled analysis evaluated English translations of the same questions using an equivalent English prompt.</p>Results<p>Overall accuracy rates were 56.7% (95% CI, 53.7–59.7) for GPT-4o, 66.1% (63.2–68.9) for GPT-5.4, 98.4% (97.4–99.0) for DS-Fast, and 99.8% (99.3–99.9) for DS-Expert.

Read the rest (3 more)

All six overall pairwise comparisons remained significant after Holm adjustment. DS-Fast and DS-Expert produced identical responses for 1, 034 of 1, 053 questions (98.2%), all of which were correct. GPT-5.4 and DS-Expert showed 100% conditional accuracy among shared responses, but their identical-response coverage was only 65.9% (694/1, 053).

Model-reported latency estimates were lowest for DS-Fast (0.68 ± 1.00 seconds) and highest for DS-Expert (3.69 ± 2.24 seconds). Under the English-translated condition, overall accuracy rates were 62.6%, 61.5%, 66.5%, and 76.9%, respectively. DS-Expert remained the most accurate configuration, although its advantage over the GPT models was reduced.</p>Conclusion<p>DeepSeek configurations outperformed the GPT models on this Chinese-language examination-preparation question set, with DS-Expert achieving the highest accuracy.

However, performance varied substantially by question language. These findings support further evaluation for answer verification and practice-question review, but they do not establish educational effectiveness.</p>

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Oncology and carcinogenesis 72% · Tabular 65%
Provenance · 2 source records, 30 field assertions
SourceKeyLast seenRaw
figshareoai:figshare.com:article/339709189 d agoJSON v1
Loughborough Research Repositoryoai:figshare.com:article/339709188 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · figshare comconnector:figshare_com@1.0.0
concepts[field].anzsrc:group:3211enrichment · figshare comtaxonomy-embedding@1.0.0title+keywords+description (72%)
concepts[field].local:field:astronomymapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:chemistrymapping · repository lboro ac ukconnector:repository_lboro_ac_uk@1.0.0
concepts[field].local:field:chemistrymapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:computer-science-aimapping · repository lboro ac ukconnector:repository_lboro_ac_uk@1.0.0
concepts[field].local:field:computer-science-aimapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:earth-environmentalmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:earth-environmentalmapping · repository lboro ac ukconnector:repository_lboro_ac_uk@1.0.0
concepts[field].local:field:economics-financemapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:economics-financemapping · repository lboro ac ukconnector:repository_lboro_ac_uk@1.0.0
concepts[field].local:field:engineeringmapping · repository lboro ac ukconnector:repository_lboro_ac_uk@1.0.0
concepts[field].local:field:engineeringmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:humanitiesmapping · repository lboro ac ukconnector:repository_lboro_ac_uk@1.0.0
concepts[field].local:field:humanitiesmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:life-sciencesmapping · repository lboro ac ukconnector:repository_lboro_ac_uk@1.0.0
concepts[field].local:field:life-sciencesmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:materials-sciencemapping · repository lboro ac ukconnector:repository_lboro_ac_uk@1.0.0
concepts[field].local:field:mathematics-statisticsmapping · repository lboro ac ukconnector:repository_lboro_ac_uk@1.0.0
concepts[field].local:field:medicine-healthmapping · repository lboro ac ukconnector:repository_lboro_ac_uk@1.0.0
concepts[field].local:field:medicine-healthmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:ocean-atmosphericmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:psychology-behavioralmapping · repository lboro ac ukconnector:repository_lboro_ac_uk@1.0.0
concepts[field].local:field:social-sciencemapping · repository lboro ac ukconnector:repository_lboro_ac_uk@1.0.0
concepts[field].local:field:social-sciencemapping · figshare comconnector:figshare_com@1.0.0
concepts[modality].local:modality:tabularenrichment · figshare comkeyword-concept-rules@1.0.0title+description (65%)
descriptionsource · figshare comconnector:figshare_com@1.0.0/metadata/dc/description
licensesource · figshare comconnector:figshare_com@1.0.0/metadata/dc/rights
publication_datesource · figshare comconnector:figshare_com@1.0.0
titlesource · figshare comconnector:figshare_com@1.0.0/metadata/dc/title