Constarium
← Search

Table · dataset · 2026

Introspective Uncertainty Estimation for LLM-Based Code Generation

Listed in figshare

<p dir="ltr">Abstract:</p><p dir="ltr">Large Language Models (LLMs) are increasingly used for code generation but can produce fluent yet functionally incorrect outputs, which limits trust in their usage for practical software engineering workflows.

Description

This thesis investigates whether Introspective Uncertainty Estimation (IUE), based on internal hidden-state representations of LLMs, can reliably indicate correctness at the response and line levels for code generation tasks.

The objective is to determine the extent to which hidden states encode information about functional code correctness and how this can be leveraged for practical risk assessment and fault localization. Methodologically, this thesis combines response-level evaluation on LiveCodeBench (LCB) and BigCodeBench (BCB) with an augmentation pipeline that derives token- and line-level labels from incorrect programs. In this setup, it compares static and dynamic response-level features, evaluates generalization across tasks, programming domains, and token positions, and studies line-level fault localization.</p><p dir="ltr">The results show that hidden states contain a strong response-level correctness signal.

Read the rest (3 more)

Static single-token probes perform best, reaching 0.90 AUROC and 0.96 F1 on LCB in the best settings, generally surpassing the thresholds of previously reported static probe baselines for IUE. More elaborate dynamic token-selection and sequence-modeling strategies yield no consistent gains. While generalization across tasks, domains, and token positions is feasible, setting-dependent degradation largely remains for real-world software projects.

At a fine granularity, line-level prediction in mixed-program settings is substantially harder than response-level estimation. However, in a conditional localization setup with known-incorrect programs, Top-K point-of-failure ranking remains effective, achieving a Top-3 hit rate of 81% in the best setting. Overall, the findings suggest that hidden states are a robust and informative resource for estimating functional code correctness and localizing faults, supporting a two-stage workflow that combines response-level risk screening with targeted line-level prioritization, and motivate further research on IUE for LLM-based code generation.</p><p dir="ltr">This archive contains multiple dataset contributions from the thesis:</p><p dir="ltr">• Comprehensive Code Dataset: We curate a dataset of code generations across two state-of-the-art code generation benchmarks, produced with four distinct open-weight LLMs, and assign response-level correctness labels using benchmark-provided test suites.</p><p dir="ltr">• High-Granularity Augmented Dataset: We augment one benchmark for two models with automatically repaired program versions of each incorrect program.

This augmented dataset provides fine-grained, token- and line-level labels, which are derived by computing diffs between the original and repaired code.</p>

Links

Where it is published

Catalogue records · 1

Topics

Provenance · 1 source records, 18 field assertions
SourceKeyLast seenRaw
figshareoai:figshare.com:article/336966585 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · figshare comconnector:figshare_com@1.0.0
concepts[field].anzsrc:field:461104mapping · figshare comvocabulary-mapper@1.0.0keywords['Neural networks']
concepts[field].anzsrc:field:461207mapping · figshare comvocabulary-mapper@1.0.0keywords['Software quality, processes and metrics']
concepts[field].local:field:astronomymapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:chemistrymapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:computer-science-aimapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:earth-environmentalmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:economics-financemapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:engineeringmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:humanitiesmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:life-sciencesmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:medicine-healthmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:ocean-atmosphericmapping · figshare comconnector:figshare_com@1.0.0
concepts[field].local:field:social-sciencemapping · figshare comconnector:figshare_com@1.0.0
descriptionsource · figshare comconnector:figshare_com@1.0.0/metadata/dc/description
licensesource · figshare comconnector:figshare_com@1.0.0/metadata/dc/rights
publication_datesource · figshare comconnector:figshare_com@1.0.0
titlesource · figshare comconnector:figshare_com@1.0.0/metadata/dc/title