Seeking Information with RAG-Assistants

Does Model Size Matter in Human-AI Collaborations?

Conference Paper (2026)
Author(s)

Lennard C. Froma (Universiteit Leiden)

Tom Kouwenhoven (Universiteit Leiden)

Maaike H.T. De Boer (TNO)

Catholijn M. Jonker (Universiteit Leiden, TU Delft - Electrical Engineering, Mathematics and Computer Science)

Max J. Van Duijn (TNO)

Research Group
Interactive Intelligence
DOI related publication
https://doi.org/10.3233/FAIA260531 Final published version
More Info
expand_more
Publication Year
2026
Language
English
Research Group
Interactive Intelligence
Pages (from-to)
441-458
Publisher
IOS Press
ISBN (electronic)
9781643686707
Event
5th International Conference on Hybrid Human-Artificial Intelligence, HHAI 2026 (2026-07-06 - 2026-07-10), Brussels, Belgium
Page Views
60
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Much research on LLMs has focused on increasing benchmark performance. However, the evaluation of such models in real-world collaborative human-AI workflows has stayed behind. This work evaluates a chatbot-style assistant based on Retrieval-Augmented Generation (RAG) in a realistic multi-turn information-seeking scenario inspired by workplace settings where compliance with local legislation and secure handling of sensitive data are often key. Specifically, we examine the performance of humans (N=112) assisted by RAG-assistants compared to LLM-only or LLM+RAG baselines. In this setting, we investigate how underlying model size (3B, 8B, and 70B) shapes the human-AI collaborative dynamic and how it influences perceived usability and satisfaction. Results show that the performance gain of human-AI collaboration over the model-only baselines is significant, irrespective of model size, suggesting that hybrid systems are beneficial in information-seeking scenarios. Interestingly, however, perceived usability and satisfaction among participants showed little difference across model sizes. This demonstrates a nuanced trade-off between model size, performance, and user perception. Our work highlights the added value of evaluating AI applications in actual multi-turn interactions with human users, looking at usability and satisfaction besides accuracy, rather than focusing on benchmark performance only.