Does In-IDE Calibration of Large Language Models work at Scale?

Conference Paper (2026)
Author(s)

Roham Koohestani (Student TU Delft, JetBrains Research)

Agnia Sergeyuk (JetBrains Research)

David Gros (University of California)

Claudio Spiess (University of California)

Sergey Titov (JetBrains Research)

Premkumar Devanbu (University of California)

Maliheh Izadi (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Research Group
Software Engineering
DOI related publication
https://doi.org/10.1145/3803437.3805234 Final published version
More Info
expand_more
Publication Year
2026
Language
English
Research Group
Software Engineering
Pages (from-to)
609-619
Publisher
ACM
ISBN (electronic)
9798400726361
Event
ACM International Conference on the Foundations of Software Engineering, FSE 2026 (2026-07-05 - 2026-07-09), Montreal, Canada
Page Views
78
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Code assistants powered by large language models are now embedded in integrated development environments, yet developers lack reliable signals for when to trust generated code. Model confidence could serve as a signal, but only if it accurately reflects the likelihood of acceptance. Post-hoc calibration aims to achieve this alignment, though its efficacy in production settings remains understudied. We investigate in-IDE confidence calibration from two perspectives: (1) scalable methods for calibrating confidence signals and (2) interface design for communicating reliability to developers. We introduce a flexible calibration framework for open-source models and evaluate calibration against developer acceptance behavior using over 24 million real-world IDE interactions across multiple languages. We find that a general Platt-scaling calibrator does not, consistently improve the usefulness of confidence as a reliability signal, while personalized calibration can help when sufficient user interaction data is available. Complementing this, a multi-phase design study with expert designers and 153 professional developers indicates a preference for non-numerical, color-coded reliability indicators embedded in the in-editor generation workflow.