Table · dataset · 2026
Introspective Uncertainty Estimation for LLM-Based Code Generation
Listed in figshare
<p dir="ltr">Abstract:</p><p dir="ltr">Large Language Models (LLMs) are increasingly used for code generation but can produce fluent yet functionally incorrect outputs, which limits trust in their usage for practical software engineering workflows.
Description
This thesis investigates whether Introspective Uncertainty Estimation (IUE), based on internal hidden-state representations of LLMs, can reliably indicate correctness at the response and line levels for code generation tasks.
The objective is to determine the extent to which hidden states encode information about functional code correctness and how this can be leveraged for practical risk assessment and fault localization. Methodologically, this thesis combines response-level evaluation on LiveCodeBench (LCB) and BigCodeBench (BCB) with an augmentation pipeline that derives token- and line-level labels from incorrect programs. In this setup, it compares static and dynamic response-level features, evaluates generalization across tasks, programming domains, and token positions, and studies line-level fault localization.</p><p dir="ltr">The results show that hidden states contain a strong response-level correctness signal.
Read the rest (3 more)
Static single-token probes perform best, reaching 0.90 AUROC and 0.96 F1 on LCB in the best settings, generally surpassing the thresholds of previously reported static probe baselines for IUE. More elaborate dynamic token-selection and sequence-modeling strategies yield no consistent gains. While generalization across tasks, domains, and token positions is feasible, setting-dependent degradation largely remains for real-world software projects.
At a fine granularity, line-level prediction in mixed-program settings is substantially harder than response-level estimation. However, in a conditional localization setup with known-incorrect programs, Top-K point-of-failure ranking remains effective, achieving a Top-3 hit rate of 81% in the best setting. Overall, the findings suggest that hidden states are a robust and informative resource for estimating functional code correctness and localizing faults, supporting a two-stage workflow that combines response-level risk screening with targeted line-level prioritization, and motivate further research on IUE for LLM-based code generation.</p><p dir="ltr">This archive contains multiple dataset contributions from the thesis:</p><p dir="ltr">• Comprehensive Code Dataset: We curate a dataset of code generations across two state-of-the-art code generation benchmarks, produced with four distinct open-weight LLMs, and assign response-level correctness labels using benchmark-provided test suites.</p><p dir="ltr">• High-Granularity Augmented Dataset: We augment one benchmark for two models with automatically repaired program versions of each incorrect program.
This augmented dataset provides fine-grained, token- and line-level labels, which are derived by computing diffs between the original and repaired code.</p>
Links
Where it is published
- DOI doi.org/10.6084/m9.figshare.33696658.v1 ↗
DOI / persistent id · from figshare com
Catalogue records · 1
- OAI-PMH record api.figshare.com/v2/oai?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai%3Af… ↗
metadata API · from figshare com
Topics
Provenance · 1 source records, 18 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| figshare | oai:figshare.com:article/33696658 | 5 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · figshare com | connector:figshare_com@1.0.0 | |
| concepts[field].anzsrc:field:461104 | mapping · figshare com | vocabulary-mapper@1.0.0 | keywords['Neural networks'] |
| concepts[field].anzsrc:field:461207 | mapping · figshare com | vocabulary-mapper@1.0.0 | keywords['Software quality, processes and metrics'] |
| concepts[field].local:field:astronomy | mapping · figshare com | connector:figshare_com@1.0.0 | |
| concepts[field].local:field:chemistry | mapping · figshare com | connector:figshare_com@1.0.0 | |
| concepts[field].local:field:computer-science-ai | mapping · figshare com | connector:figshare_com@1.0.0 | |
| concepts[field].local:field:earth-environmental | mapping · figshare com | connector:figshare_com@1.0.0 | |
| concepts[field].local:field:economics-finance | mapping · figshare com | connector:figshare_com@1.0.0 | |
| concepts[field].local:field:engineering | mapping · figshare com | connector:figshare_com@1.0.0 | |
| concepts[field].local:field:humanities | mapping · figshare com | connector:figshare_com@1.0.0 | |
| concepts[field].local:field:life-sciences | mapping · figshare com | connector:figshare_com@1.0.0 | |
| concepts[field].local:field:medicine-health | mapping · figshare com | connector:figshare_com@1.0.0 | |
| concepts[field].local:field:ocean-atmospheric | mapping · figshare com | connector:figshare_com@1.0.0 | |
| concepts[field].local:field:social-science | mapping · figshare com | connector:figshare_com@1.0.0 | |
| description | source · figshare com | connector:figshare_com@1.0.0 | /metadata/dc/description |
| license | source · figshare com | connector:figshare_com@1.0.0 | /metadata/dc/rights |
| publication_date | source · figshare com | connector:figshare_com@1.0.0 | |
| title | source · figshare com | connector:figshare_com@1.0.0 | /metadata/dc/title |