Constarium
← Search

Table · dataset · 2026

<p>Median values by prompt type with effect sizes.</p>

Listed in figshare

<div><p>Large Language Models (LLMs) can generate text describing scientific concepts, but the characteristics of these outputs remain poorly understood.

Description

We present a multi-dimensional characterization framework analyzing 9,666 outputs from seven models (GPT-4.1, GPT-5.2, GPT-5.5, o4-mini, Claude Sonnet 4.5, Claude Opus 4.5, and the open-weight Gemma-3-27B) across five scientific domains. Rather than making claims about creativity or novelty, we measure five independent dimensions: coherence, domain relevance, lexical profile, structural properties, and semantic position using four sentence embedding models spanning 2020–2024.

After quality filtering (99.2% coherence, 99.9% domain relevance pass rates), we find that outputs exhibit graduate-level readability (median Flesch-Kincaid grade 16.3) and occupy semantic positions at the 83rd percentile of calibration distributions. All 39 metrics differ significantly across models (Kruskal-Wallis, <i>p</i> < 0.05), with structural properties showing the largest effects ( = 0.35–0.54) and semantic position showing small-to-medium effects ( 0.01–0.14).

Read the rest (2 more)

Dunn’s post-hoc tests with Bonferroni correction confirm that all model pairs differ on the top structural metrics (21/21 pairs for paragraph count, 20/21 for the next three). Centroid positions are robust to calibration sampling (bootstrap cosine similarity 0.99) and exceed a shuffled-domain baseline in 96.8–98.7% of cases; cross-model embedding consistency is moderate (Spearman = 0.38–0.67) and is not explained by embedding dimensionality.

Temperature effects replicate across two independent full-range models (GPT-4.1 and Gemma). This work provides calibrated measurements and validated methodology for future research without making interpretive claims about novelty.</p></div>

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Text 75%
Provenance · 1 source records, 19 field assertions
SourceKeyLast seenRaw
figshareoai:figshare.com:article/338139159 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · figshare comconnector:figshare_com@1.0.0
concepts[field].anzsrc:field:440710mapping · figshare comvocabulary-mapper@1.0.0keywords['Science Policy']
concepts[field].anzsrc:group:3101mapping · figshare comvocabulary-mapper@1.0.0keywords['Cell Biology']
concepts[field].local:field:astronomymapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:chemistrymapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:computer-science-aimapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:earth-environmentalmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:economics-financemapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:engineeringmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:humanitiesmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:life-sciencesmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:medicine-healthmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:ocean-atmosphericmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:social-sciencemapping · figshare comconnector:figshare_com@1.0.0
concepts[modality].local:modality:textenrichment · figshare comkeyword-concept-rules@1.0.0title+description (75%)
descriptionsource · figshare comconnector:figshare_com@1.0.0/metadata/dc/description
licensesource · figshare comconnector:figshare_com@1.0.0/metadata/dc/rights
publication_datesource · figshare comconnector:figshare_com@1.0.0
titlesource · figshare comconnector:figshare_com@1.0.0/metadata/dc/title