Data · dataset · 2026
Calibration-Aware Reinforcement Learning for Large Language Models: A Survey of Objectives, Optimization, and Decision-Making
Listed in ZivaHub
<p dir="ltr">Large language models increasingly emit confidence reports, predictive distributions, and typed decisions that determine whether a system answers, abstains, retrieves evidence, or spends more computation.
Description
We survey calibration-aware reinforcement learning (RL), in which a reported probability is scored by the reward, consumed by the policy’s actions, or both. Such a probability means something only relative to an event, the reporter’s information, and a population or policy.
Unlike post-hoc calibration, RL can change the answers being assessed, the incentive actually optimized, and the behavior that consumes the number. We organize the literature around three gaps. Between reporting and capability, proper-score geometry shows that a joint answer-confidence reward decomposes into terms for mean accuracy, the distribution of success across inputs, and reporting error, so it can prefer a less accurate policy even under truthful reporting.
Read the rest (2 more)
Between the stated objective and the implemented update, group standardization, the treatment of parameter-dependent rewards, and finite-ensemble scoring can change what training optimizes. Between ranking and decision value, a score can keep its ordering while losing the numerical meaning that a cost-derived threshold requires. Exact constructions make each gap concrete, and a method landscape traces representative protocols to primary sources.
We then give a comparator guide; an evaluation contract that separates reporting gains from changes in the answer policy, effects of the implemented update, and operating-point selection; and open questions. The practical conclusion is a comparator rule: post-hoc fitting is the comparator to beat when predictions can remain fixed, and direct supervision when the desired report has a tractable differentiable loss; RL is a distinct intervention only when reasoning, retrieval, answering, abstention, or information acquisition must adapt through outcome feedback.</p>
Links
Where it is published
- DOI doi.org/10.1184/r1/33989641.v2 ↗
DOI / persistent id · from zivahub uct ac za
Catalogue records · 1
- OAI-PMH record api.figshare.com/v2/oai?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai%3Af… ↗
metadata API · from zivahub uct ac za
Topics
Provenance · 1 source records, 10 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| ZivaHub | oai:figshare.com:article/33989641 | 6 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · zivahub uct ac za | connector:zivahub_uct_ac_za@1.0.0 | |
| concepts[field].anzsrc:field:461105 | mapping · zivahub uct ac za | vocabulary-mapper@1.0.0 | keywords['Reinforcement Learning'] |
| concepts[field].anzsrc:group:4602 | mapping · zivahub uct ac za | vocabulary-mapper@1.0.0 | keywords['Artificial intelligence'] |
| concepts[field].local:field:computer-science-ai | mapping · zivahub uct ac za | connector:zivahub_uct_ac_za@1.0.0 | |
| concepts[field].local:field:earth-environmental | mapping · zivahub uct ac za | connector:zivahub_uct_ac_za@1.0.0 | |
| concepts[field].local:field:humanities | mapping · zivahub uct ac za | connector:zivahub_uct_ac_za@1.0.0 | |
| description | source · zivahub uct ac za | connector:zivahub_uct_ac_za@1.0.0 | /metadata/dc/description |
| license | source · zivahub uct ac za | connector:zivahub_uct_ac_za@1.0.0 | /metadata/dc/rights |
| publication_date | source · zivahub uct ac za | connector:zivahub_uct_ac_za@1.0.0 | |
| title | source · zivahub uct ac za | connector:zivahub_uct_ac_za@1.0.0 | /metadata/dc/title |