Information-Processing Alignment in Human-Robot Interaction
R. Zhang (TU Delft - Mechanical Engineering)
J.C.F. de Winter – Promotor (TU Delft - Mechanical Engineering)
D. Dodou – Promotor (TU Delft - Mechanical Engineering)
Y.B. Eisma – Copromotor (TU Delft - Mechanical Engineering)
H.C. Seyffert – Copromotor (TU Delft - Mechanical Engineering)
More Info
expand_more
Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.
Abstract
Machines that work alongside humans have progressed from mechanical devices to AI-driven robots that can perceive, reason, and act in environments shared with people. Multimodal large language models (MLLMs) now enable robots to interpret open-ended instructions, plan multi-step actions, and generate natural language explanations. However, technical capability alone does not guarantee effective collaboration, because humans and AI-driven robots perceive, reason, and respond through different information-processing mechanisms, and when these processes are misaligned, interaction can fail regardless of the robot’s capability.
This thesis investigated where alignment and misalignment occur between humans and AI-driven robots, and how human-robot interaction should be designed for complementary collaboration. This thesis adopts the vocabulary of Wickens’ Multiple Resource Theory, which distinguishes processing stages (perception, cognition, responding), processing modalities (visual and auditory), and processing codes (spatial, verbal). Four empirical studies targeted different processing stages and progressed in platform complexity and realism.
Across the four studies, a combined 2,292 participants generated evidence that misalignment can occur at every information-processing stage. Three overarching findings are extracted. First, outcome-level results alone are insufficient for detecting misalignment: high correlations in Chapter 2 masked different search mechanisms, and verbal fluency in Chapter 4 masked shallow reasoning. Second, as the AI system becomes more automated, the human role changes from direct operator to supervisor, from commanding every movement (Chapter 3), through evaluating automated perception and strategy (Chapter 4), to diagnosing complete LLM-generated plans (Chapter 5). Third, the quality of collaboration depends on how cognition is distributed, not only on the absolute capability of either partner: voice control freed spatial resources while gestures overloaded them (Chapter 3), and students learned to redistribute cognitive work based on model capability and task difficulty (Chapter 5).
This thesis also has limitations. The studies focused heavily on the visual modality; haptic and auditory modalities were not explicitly investigated. Each study concerned a specific domain and a relatively homogeneous participant population. Whether the documented forms of misalignment transfer to other human-robot-interaction domains, more complex environments, and more diverse populations requires further investigation. Nonetheless, the findings support a set of design principles: use modality switching to offload overloaded cognitive resources, expose AI uncertainty rather than only AI outcomes, provide calibration training before deployment, and design function allocation as dynamic instead of static.