Conformal Prediction for Complex Fact-Checking with Large Language Models

Master Thesis (2026)
Author(s)

A. Nechita (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Contributor(s)

P.K. Murukannaiah – Graduation committee member (TU Delft - Electrical Engineering, Mathematics and Computer Science)

S. Mukherjee – Graduation committee member (TU Delft - Electrical Engineering, Mathematics and Computer Science)

J. Yang – Graduation committee member (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Faculty
Electrical Engineering, Mathematics and Computer Science
More Info
expand_more
Publication Year
2026
Language
English
Graduation Date
03-07-2026
Awarding Institution
Delft University of Technology
Programme
Computer Science, Data Science and Artificial Intelligence Technology
Faculty
Electrical Engineering, Mathematics and Computer Science
Downloads counter
30
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Large Language Models (LLMs) are increasingly deployed in high-stakes applications, where unreliable outputs can have serious consequences. However, quantifying their uncertainty and providing verifiable guarantees on their predictions remains an open challenge. We investigated the use of conformal prediction (CP), a statistical framework for generating prediction sets with provable coverage guarantees, to address this gap in LLM-based fact-checking. Using Llama-3.1-8B-Instruct and Mistral-7B-Instruct-v0.3, we evaluated three non-conformity scores on more than 10,000 claims verified by PolitiFact, spanning statements made between the website's launch in 2007 and January 2026. By formulating PolitiFact verdict classification as a multiple-choice question answering (MCQA) problem, we assessed the LLMs' confidence in classifying claims into PolitiFact's six "Truth-O-Meter" veracity ratings. We found that access to the corresponding fact-check article sharply reduced prediction set sizes, yielding sets small enough to be useful to a human fact-checker. Without this evidence, prediction sets typically contained both true and false labels, indicating that the models could not reliably distinguish true from false claims on their own at the 90% marginal coverage level. In this way, this research advances the application of LLMs in high-stakes contexts requiring verifiable bounds on error rates.

Files

License info not available