Human-Like Attention? A Psychophysical Comparison of Visual Search in Humans and MLLMs

Journal Article (2026)
Author(s)

R. Zhang (TU Delft - Mechanical Engineering)

J. C.F. de Winter (TU Delft - Mechanical Engineering)

D. Dodou (TU Delft - Mechanical Engineering)

H. Seyffert (TU Delft - Mechanical Engineering)

Y. B. Eisma (TU Delft - Mechanical Engineering)

Research Group
Human-Robot Interaction
DOI related publication
https://doi.org/10.1007/s42113-026-00333-4 Final published version
More Info
expand_more
Publication Year
2026
Language
English
Research Group
Human-Robot Interaction
Journal title
Computational Brain and Behavior
Downloads counter
11
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Visual search is a fundamental cognitive ability. This study investigates whether Multimodal Large Language Models (MLLMs) exhibit human-like difficulty signatures in visual search tasks. We compared search performance of humans (n = 1,250) and MLLMs using identical 2D and 3D stimuli across different set sizes. Both groups showed efficient performance in feature searches, most clearly when the target had a unique color, but performance degradation in conjunction searches as set sizes increased. Additionally, we found strong correlations between human and MLLM error rates (ρ = 0.82), which suggests that MLLMs are sensitive to similar objective complexities, such as stimulus heterogeneity. However, differences were found as well: whereas humans invested extra search time to respond accurately on target-absent trials, MLLMs exhibited extreme present/absent response biases in complex searches. We conclude that MLLMs replicate high-level human performance signatures, yet their underlying computations differ significantly.