The Effect of Evidence Distribution Across Conversational Memory and Documents on QA Agent Answer Correctness

Master Thesis (2026)
Author(s)

T. Mladenović (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Contributor(s)

J. Urbano Merino – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)

M. Skrodzki – Graduation committee member (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Faculty
Electrical Engineering, Mathematics and Computer Science
More Info
expand_more
Publication Year
2026
Language
English
Graduation Date
26-08-2026
Awarding Institution
Delft University of Technology
Programme
Computer Science, Data Science and Artificial Intelligence Technology
Faculty
Electrical Engineering, Mathematics and Computer Science
Downloads counter
12
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Personalized QA agents increasingly retrieve from both external documents and private conversational memory. This raises a basic design question: does answer correctness depend on where the supporting evidence is stored? Existing evaluations make this difficult to answer because memory and document conditions often differ in their questions, facts, or available sources. We construct a controlled evaluation setting from MultiHop-RAG in which the same questions and evidence are held fixed while golden evidence is placed entirely in documents, entirely in conversational memory, or split across both. We evaluate four retrieval architectures across three persona-specific memory corpora. The results show that evidence placement does not exert a simple, uniform penalty. Instead, its effect is significantly moderated by the topic match of the question to the persona's memory profile. In particular, questions whose domain matched the surrounding memory profile were answered more accurately, a surprising result that challenges the expectation that topically concentrated memory should harm retrieval and reasoning. This performance degradation during off-topic exploration suggests personalized agents may create soft algorithmic `filter bubbles' that penalize user curiosity.

Files

Thesis_Paper_updated.pdf
(pdf | 2.11 Mb)
License info not available