Circular Image

J.C.F. de Winter

info

Please Note

73 records found

Testing input format and prompt structure on the social robotics learning literature

Master thesis (2026) - S. Gulikers, J.C.F. de Winter, Y.B. Eisma, D. Dodou
This study evaluates recent large language models for extracting meta-analytic data from scientific articles in which relevant results are reported in text, tables, and figures. The study compares GPT-5.2, Claude Opus 4.6, Gemini 3.1 Pro Preview, and Gemini 3.1 Flash Lite Preview across three input representations: the original PDF, structured Markdown derived from the PDF, and a combined input consisting of the PDF together with cropped tables and figures. The final analyzed corpus contains 56 papers from a social robotics meta-analysis and includes 193 statistical data rows for which pre- and post-intervention values had to be extracted. Each row was scored on five numerical fields relevant for effect size calculation (Pre-Mean, Pre-SD, Post-Mean, Post-SD, and n), yielding 965 scored cells in total. Model outputs were compared against manually corrected ground-truth values from the original meta-analysis, which were verified against the source papers. Predicted rows were first matched to the corresponding ground-truth rows, after which the individual numerical values were scored using predefined numerical tolerances. This design allowed extraction errors to be interpreted not only as aggregate model failures, but also in relation to source format, input representation, target-row construction, and numerical reading. Across the full evaluation corpus, Markdown was often the strongest or near-strongest input representation overall, although the strongest representation differed across text-, table-, and figure-dominant papers. The highest overall performance was achieved by Claude Opus 4.6 with Markdown input. Papers in which the target values had to be extracted mainly from figures were the most difficult. A diagnostic follow-up on 17 papers that repeatedly showed errors in identifying the correct paper-condition combinations found that performance improved when the relevant condition was specified in advance and the model only had to extract the corresponding values. This suggests that many remaining errors were associated with identifying the correct extraction target row rather than with numerical reading alone. ...

The Impact of Spatial Representations and Frames of Reference

Current evaluations of Large Language Model (LLM) spatial reasoning focus on several isolated competencies rather than a unified task, and use an array of different input formats. As a result, the influence of spatial representation and output Frame of Reference (FoR) on performance in navigation tasks remains unclear. This study asks: how do spatial representations and frames of reference influence LLMs' spatial reasoning capabilities, and which combinations are conducive to it?
This research investigates the spatial reasoning and navigation capabilities of one reasoning and one non-reasoning LLM. Using perfect mazes as a controlled testbed, this thesis examines how various input spatial representations, including visual (JPG and ASCII), grid-based (JSON and Tagged per-cell), and graph-based (Adjacency List) formats, interact with different output FoRs to influence model performance.
The methodology involves an evaluation using Gemini 2.5 Pro (reasoning) and Gemini 2.5 Flash-Lite (non-reasoning) across 11 spatial representations and three output FoRs: allocentric using absolute coordinates ("coordinates"), allocentric using absolute directions ("absolute directions"), and egocentric (relative directions). Performance is measured using two metrics: a "completion score", defined as the percentage of the path navigated correctly before the first error, and the mean number of output tokens generated, used as a proxy for efficiency.
The findings of this research indicate that performance is highest when mazes are expressed using structured graph-based spatial representations, particularly Adjacency List JSON (a graph-based representation formatted as a JSON file), across model types, while the choice of output FoR strongly shapes outcomes, with absolute coordinate responses yielding substantially better results than egocentric ones that require continuous relational analysis and state tracking and therefore lead to markedly lower completion scores, especially for the non-reasoning model. In addition, inspection of internal reasoning traces suggests that the use of formal graph-solving algorithms is positively correlated with success, while exclusive reliance on heuristics and unfounded declarations of confidence are negatively correlated with completion scores.
By systematically varying input spatial representation and output FoR this work provides the first integrated evaluation of these factors, addressing the lack of unified benchmarks and clarifying how methodological choices shape observed LLM spatial reasoning performance. ...

The behavioural differences between a 1D merging agent controlled by a Large Language Model and human driving data

Human driver models are essential for the development and testing of Automated Driving Systems (ADS), yet current approaches often struggle to capture the complex, stochastic nature of human tactical decision-making. Large Language Models (LLMs) have emerged as potential reasoning agents capable of emulating human-like social behaviour, but their application as direct vehicle control agents remains largely underexplored.

This thesis investigates the extent to which a base LLM, guided by systematic prompt engineering, can replicate the tactical decisions and control of human drivers in a 1-D highway merging scenario. Using the OpenAI o3 model, an LLM-driven agent was developed and systematically benchmarked against a dataset of human driver behaviour recorded in a simulator experiment.
The study utilised Linear Mixed-Effects Regression (LMER) to analyse decision-making mechanisms and performed a sensitivity analysis using the Google Gemini-2.5-pro model to assess generalisability.

The results demonstrate that the LLM agent successfully replicated high-level tactical behaviours, satisfying qualitative criteria such as symmetrical yielding in neutral conditions and increased yield rates when the opposing vehicle held a headway advantage. However, a fundamental disparity was observed in operational control. While human drivers relied significantly on relative velocity to negotiate merges (p = 1.88 × 10−26), the LLM adopted a conservative, calculation-heavy gap-based strategy driven by absolute distance, resulting in average safety margins more than double the human benchmark (9.18 m vs. 3.85 m). Furthermore, a sensitivity analysis revealed severe model dependency. While the optimised prompt achieved a 0.0% collision rate with the o3 model, it resulted in a 25.5% collision rate with Gemini-2.5-pro.

This research concludes that while base LLMs possess the emergent reasoning capabilities to function as high-level strategic agents, their lack of continuous perceptual flow limits their validity as direct operational controllers. The findings suggest that future implementations should adopt hierarchical architectures, leveraging LLMs for tactical reasoning while relying on physics-based controllers for dynamic execution. ...
Machines that work alongside humans have progressed from mechanical devices to AI-driven robots that can perceive, reason, and act in environments shared with people. Multimodal large language models (MLLMs) now enable robots to interpret open-ended instructions, plan multi-step actions, and generate natural language explanations. However, technical capability alone does not guarantee effective collaboration, because humans and AI-driven robots perceive, reason, and respond through different information-processing mechanisms, and when these processes are misaligned, interaction can fail regardless of the robot’s capability.

This thesis investigated where alignment and misalignment occur between humans and AI-driven robots, and how human-robot interaction should be designed for complementary collaboration. This thesis adopts the vocabulary of Wickens’ Multiple Resource Theory, which distinguishes processing stages (perception, cognition, responding), processing modalities (visual and auditory), and processing codes (spatial, verbal). Four empirical studies targeted different processing stages and progressed in platform complexity and realism.

Across the four studies, a combined 2,292 participants generated evidence that misalignment can occur at every information-processing stage. Three overarching findings are extracted. First, outcome-level results alone are insufficient for detecting misalignment: high correlations in Chapter 2 masked different search mechanisms, and verbal fluency in Chapter 4 masked shallow reasoning. Second, as the AI system becomes more automated, the human role changes from direct operator to supervisor, from commanding every movement (Chapter 3), through evaluating automated perception and strategy (Chapter 4), to diagnosing complete LLM-generated plans (Chapter 5). Third, the quality of collaboration depends on how cognition is distributed, not only on the absolute capability of either partner: voice control freed spatial resources while gestures overloaded them (Chapter 3), and students learned to redistribute cognitive work based on model capability and task difficulty (Chapter 5).

This thesis also has limitations. The studies focused heavily on the visual modality; haptic and auditory modalities were not explicitly investigated. Each study concerned a specific domain and a relatively homogeneous participant population. Whether the documented forms of misalignment transfer to other human-robot-interaction domains, more complex environments, and more diverse populations requires further investigation. Nonetheless, the findings support a set of design principles: use modality switching to offload overloaded cognitive resources, expose AI uncertainty rather than only AI outcomes, provide calibration training before deployment, and design function allocation as dynamic instead of static. ...
Numerous techniques have been developed in order to explain the reasoning process of black-box models. Among them is a class of models that are designed to be inherently interpretable: select-then-predict models (a.k.a. rationale-based models). These models are meant to explain their prediction by highlighting part of the input as evidence. The evidence, called the rationale, should consist of the most salient parts of the input text that contribute the most to the model's decision.
However, according to some recent studies, these models are not truly interpretable, because they do not provide faithful explanations (i.e., explanations that accurately reflect the true reasoning process of the model).
In this thesis we give a formal definition of the degree of unfaithfulness to quantify unfaithful behavior. Then, we introduce an experiment to test the faithfulness of select-then-predict models and prove that select-then-predict models can provide unfaithful rationales. Lastly, we introduce a loss function, which we call the unfaithfulness loss, which minimizes the degree of unfaithfulness of select-then-predict models and teaches them to produce more faithful and plausible rationales. ...
Master thesis (2025) - A.C.G. Hutani, J.C.F. de Winter, D. Dodou, R. de Leeuw van Weenen
This thesis investigates how temporal design choices affect the real-time feasibility of human motion prediction models. Two state-of-the-art models were evaluated: GCNext, a data-driven graph convolutional model, and PhysMoP, a hybrid model combining a physics-based and data-driven branch. Controlled experiments showed the influence of input history length, temporal resolution, and the model architecture on prediction accuracy and latency. Results showed that longer observation windows do not necessarily improve accuracy, while increasing the latency. Both models were sensitive to changes in temporal resolution, as they implicitly assumed a fixed sampling rate. Real-time performance analysis indicated that single-pass architectures were favoured, while autoregressive models suffered from compounding delay. Retraining GCNext with shorter input histories and optimising autoregressive passes achieved substantial latency reduction with minimal accuracy loss. These results show that temporal configurations are critical design choices for achieving real-time feasibility of human motion prediction models. The code for this paper is available at https://github.com/AndrewHutani/HMP ...
Doctoral thesis (2025) - V. Onkhar, J.C.F. de Winter, D. Dodou
A large number of traffic accidents occur worldwide each year, of which a sizable portion involve pedestrians, making them a vulnerable group on the road. Many of these accidents occur due to visual distraction, meaning drivers and pedestrians fail to look where they should be looking. In addition to this tendency for distraction, the eyes are a means of exchanging information between road users, via behaviors such as eye contact. However, the role and importance of eye contact in traffic in connection with traffic safety and the decisions of road users is not yet entirely clear. Further, with the advent of automated vehicles, the role of eye contact in traffic may change or disappear altogether, due to the absence of drivers. One promising way to shed light on this matter is to use eye-tracking, a technology which can measure the eye movements of road users, and which might allow the engineering of solutions to mitigate the frequency and severity of accidents.

This dissertation aims to investigate the role of eye contact between drivers and pedestrians, as well as its influence on pedestrians’ road crossing intentions. Another aim of this dissertation is to assess the accuracy of eye-tracking devices and to objectively detect and operationalize driver-pedestrian eye contact using eye-tracking. Finally, this thesis aims to develop safety systems based on eye-tracking that can automatically analyze and contextualize gaze in traffic and warn vulnerable road users of danger. This thesis consists of four independently readable and empirical research papers.

The first study examines the effect of drivers’ eye contact on pedestrians’ crossing decisions using an online crowdsourced experiment. It shows that, although a car’s kinematics have a dominant effect, a driver’s eye contact also makes pedestrians feel safer and more likely to cross the road, and that the timing of the driver’s eye contact has an influence as well. The second study benchmarks the accuracies of mobile eye-trackers under static and dynamic conditions, finding that eccentricity worsens accuracy, but dynamicity does not necessarily worsen it. The third study presents a method to objectively detect and operationalize driver-pedestrian eye contact using two synchronized eye-trackers and computer vision, defining eye contact as mutual gaze within a 4° threshold. The fourth study explores the integration of mobile eye-tracking, object detection, and a vision-language model in an attempt to develop a real-time, context-aware safety system that can assess risk in traffic and enhance the situational awareness of road users.

This dissertation concludes that while eye contact is neither as powerful a cue as kinematics nor essential for crossing, it is still a “should-have” in driver-pedestrian interactions as it can increase perceived safety and willingness to cross. This thesis also concludes that certain types of external Human Machine Interfaces (eHMIs) – substitutes for the missing eye contact between pedestrians and automated vehicles – would be beneficial to maintain existing levels of comfort in interactions. Finally, this thesis also highlights the potential of using mobile eye-tracking in combination with computer vision and AI for applications in the traffic, manufacturing, medical, education, and other domains, and recommends topics for further research into eye contact and eye-tracking.
...

From Raw Data to Context-Aware Interpretations

Doctoral thesis (2025) - T. Driessen, J.C.F. de Winter, D. Dodou, Dick de Waard
Road traffic accidents remain a major public health concern worldwide. Technological advances in vehicle sensing, automation, and artificial intelligence present novel opportunities to assess and improve human driving. This dissertation explores these opportunities by developing and evaluating algorithms to assess the behavior of car and truck drivers.

Initial research establishes the perspectives of driving examiners and professional truck drivers on the acceptance of data-driven tools to assess driver behavior. The work then demonstrates that practical methods using readily available GPS and accelerometer data can successfully identify driving styles and predict negative outcomes like fines and damage incidents at a population level. However, these simple metrics prove insufficient for fair individual assessment due to the lack of situational context embedded in such data.

To address this limitation, the thesis explores modern AI-based approaches. It demonstrates how AI systems from automated driving can provide continuous behavioral references to evaluate human performance, and concludes by showing that vision-language models can establish a more holistic, "context-aware" risk assessment using images of typical traffic situations. ...

A Turing Test experiment using Think Aloud and Eye Tracking methods

Master thesis (2024) - R. Koerts, Y.B. Eisma, J.C.F. de Winter, D. Dodou
With the advancement of Artificial Intelligence leading to increasingly human-like outputs, assessing a machine’s ability to exhibit human-like intelligence has become more essential than ever. This study aims to investigate how human-like chess players perceive four conditions: one human opponent and three different types of algorithms. One of these algorithms, Maia, has been trained on human data and aims to play the most human-like move. In a custom-designed experiment similar to a Turing test, chess players faced off against Maia, Stockfish and a human without knowing their opponent’s nature. After each game, the chess player assessed how human-like the moves of the opponent were and estimated whether they played against an engine or a human opponent. During the game, participants were asked to think aloud about their next move and react towards the moves of the opponent. Additionally, the gaze of the player was captured with the SR EyeLink Portable Duo at 1000Hz, with the goal of finding differences within the player’s gaze while participants tried to discover the nature of their opponent. Results from the experiment revealed that, based on responses to a subjective questionnaire, the perceived humanness of Maia is statistically similar to a human and different from the other two chess engines. From the analysis of the voice recordings, categories of sentences were identified that could suggest recognition of the opponent, specifically: "expected", "unexpected", "human-like" and "engine-like". From the eye-tracking results, the average fixation duration and pupil diameter changes following the opponent’s move were compared for each condition, but showed no statistical differences between conditions. In summary, Maia was perceived more human-like compared with other chess engines. However, differences in underlying cognitive processes on how the human perceived this difference in a Turing Test experiment were not identified. ...
This thesis presents the design and evaluation of a comprehensive system for developing voice-based interfaces to support users in supermarkets. These interfaces enable customers to convey their needs across both generic and specific queries. While current state-of-the-art systems like GPTs by OpenAI are easily accessible and adaptable, featuring low-code deployment with options for functional integration, they still face challenges such as increased response times and limitations in strategic control for tailored use-case and cost optimisation. Motivated by the goal of crafting inclusive, personalised, and efficient conversational agents, this study advances on three fronts: 1) a comparative analysis of four popular off-the-shelf speech recognition technologies to identify the most accurate model for different genders (male/female) and languages (English/Dutch); 2) an assessment of the effects of personalised recommendations versus generic responses, using a blindfolded, counterbalanced within-subject experiment; and 3) the development and evaluation of a novel multi-LLM supermarket chatbot framework, comparing its performance with a specialized GPT model powered by the GPT-4 Turbo, using the Artificial Social Agent Questionnaire (ASAQ) in a counterbalanced within-subjects experiment and qualitative participant feedback. Our find-ings reveal that OpenAI’s Whisper leads in speech recognition accuracy across genders and languages, users significantly prefer personalised chatbots over the non-personalised counterparts and that our proposed multi-LLM chatbot architecture outperformed the benchmarked GPT model across all 13 measured criteria, including statistically significant improvements in four key areas: performance, user satisfaction, user-agent partnership, and self-image enhancement. The thesis concludes by presenting a simple method for supermarket robot navigation by mapping the final chatbot response to correct shelf numbers towards which the robot can plan sequential visits. This later enables effective use of low-level perception, motion planning, and control capabilities for product retrieval and collection. We hope this work encourages more efforts into using multiple, specialised smaller models instead of always relying on a single powerful model. ...
Immersive virtual reality (IVR) with head-mounted displays (HMDs) is expanding in various fields like training, but its effects on cognitive load from visual, auditory, and mental stimuli in virtual environments remain uncertain. This is particularly relevant in neurorehabilitation, where patients may suffer from training in overstimulating environments due to cognitive impairments. This study further explores how low and high levels of visual, auditory, and cognitive demands affect the cognitive load. Twenty-two participants used an HMD for a virtual shopping task, which involved selecting listed products and placing them in a cart, under baseline (the task without additional demands) and two stimulus complexity levels (low and high) for visual, auditory, and mental demands. The study assessed cognitive load using conventional (heart rate, variability, skin conductance, performance, self-reported questionnaire) and underexplored measures (head stillness, hand smoothness, gaze behavior), and explored behavior changes due to stimulus impact. Results showed that visual and auditory stimuli had minimal effects on cognitive load, with only specific measures showing any notable differences from the baseline. Mental stimuli significantly impacted cognitive load, with high mental tasks notably affecting the measures and behavior, whereas low mental tasks showed fewer changes. This research concluded that mental stimuli significantly increased the cognitive load in virtual environments, more than visual or auditory stimuli, suggesting future virtual reality designs should prioritize managing mental load. Furthermore, the study highlights the effectiveness of head stillness and gaze behavior as innovative measures for evaluating cognitive load. ...

Enhancing Cyclist Interaction with Automated Vehicles through Human-Machine Interfaces

Doctoral thesis (2024) - S.H. Berge, M.P. Hagenzieker, J.C.F. de Winter
This dissertation explores cyclist-automated vehicle interactions, emphasising developing and integrating human-machine interfaces (HMIs) to enhance cyclist safety and communication. Adopting a cyclist-centric perspective, it recognises cyclists' unique characteristics and communication strategies in shared traffic environments. Using semi-structured interviews, literature reviews, data triangulation, an eye-tracking field experiment, and a cycling simulator study, the research addresses five key research questions, providing qualitative and quantitative insights.

The main contributions of this dissertation include a thorough investigation of cyclists' expectations for future interactions with automated vehicles, highlighting the need for reliable detection by automated vehicles and placing the responsibility for safety on vehicle developers rather than cyclists. The research offers objective data and self-reported insights into cyclist-automated vehicle interactions and evaluates cyclists' ability to visually detect the presence or absence of a driver. Additionally, it introduces 20 scenarios of cyclist-automated vehicle interaction, serving as a resource for safety assessments and HMI research. A comprehensive literature review of existing HMIs for cyclists was conducted, identifying 92 concepts involving vehicles, bicycles, cyclists, and infrastructure.

The dissertation concludes with design recommendations for cyclist-centric HMIs, proposing an omnidirectional on-vehicle external HMI (eHMI) to communicate detection and automated driving mode. This dissertation provides valuable insights for researchers, policymakers, and automated vehicle developers, aiming for the safer, more inclusive, and sustainable urban traffic environments of tomorrow. ...
Doctoral thesis (2024) - W. Tabone, J.C.F. de Winter, R. Happee
This thesis explores how automated vehicles will interact with pedestrians in the urban environment through augmented reality technology. Nine distinct AR interfaces were designed, developed, and evaluated to assess how different design elements (symbols, text, colour) and distinct mappings of the AR (on the road, on the vehicle, or head-locked) would affect comprehension, and ultimately whether the pedestrian would trust and be convinced to cross in front of an automated vehicle displaying a safe message. Using increasing levels of ecological validity, from an online questionnaire to a CAVE simulator and an AR HMD experiment, the evaluation also explored how different AR anchoring (and mapping) positions affect pedestrians' crossing initiation times and the intuitiveness of the message. The thesis also explores the use of diminished reality (removal of information) to assist pedestrians in occluded scenarios, as well as the utilisation of Large Language Models in evaluating qualitative data in experiments. The outcomes of the thesis are a set of guidelines based on empirical evidence on how to design effective AR interfaces which promote safe and transparent interactions between pedestrians and automated vehicles. ...
Objective: The aim of this study was to investigate the effect of autonomous-vehicle-to-pedestrian (AV2P) communication through augmented reality (AR) interfaces on the road crossing behaviour of pedestrians, and research whether subjective results from a previous Cave Automatic Virtual Environment (CAVE) study replicated in a real world AR experiment.
Background: Previous studies investigating the effects of AV2P communication have mostly been conducted through virtual reality (VR) providing researchers with safe experimentation methods and high experimental control, but also resulting in a common limitation: the lack of ecological validity and realism, thereby affecting participants’ behaviour and causing distractions. This study therefore introduces AR experiments that have been conducted in a real world environment to increase ecological validity.
Methods: An AR experiment was conducted in which 28 participants were situated in the real world with the objective to cross the road. The virtual vehicle, that was projected through a Varjo XR-3 head mounted display, approached from the right at a speed of 30 km/h while 4 interfaces (2x world-locked, head-locked, and vehicle-locked) appeared to communicate the vehicle’s intention towards the participants, in addition to a no-interface baseline. Participants were tasked with indicating when they were willing to cross through the push of a remote button from which their Willingness to cross and Decision certainty could be derived. Subjective data was collected after the trials and after the experiment through interviews and a questionnaire respectively.
Results: Results suggest a positive effect of the AV2P interfaces on the Willingness to cross and Decision certainty, although statistically not significant. In other words, Willingness to cross increases when the vehicle indicates that it will yield, and decreases when the vehicle communicates that it will not yield. Decision certainty also increases when an interface is present compared to the no-interface baseline. Moreover, participants indicated using the interfaces as a tool to validate their own decisions. Compared to the CAVE study, subjective intuitiveness ratings replicate in terms of observing higher intuitiveness of the interfaces than the no-interface baseline. However, the intuitiveness ratings were higher in the CAVE study than the real world AR experiment. Furthermore, the order of the top 3 most preferred interfaces ranking is in the opposite order. Both differences suggest that the increased ecological validity of the real world AR experiment introduces new insights into participants’ perception of interfaces. The Van der Laan acceptance scale shows that participants believe interfaces to be useful and satisfying overall.
Conclusion: The experiments suggest that AV2P interfaces have a positive effect on the crossing behaviour of pedestrians. Furthermore, participants indicate using the interfaces as a tool to validate their own decision, which increases confidence in their decisions. Although results partially replicate a previous virtual environment study, there are differences that suggest that real world AR experiments provide valuable insights into participants’ perception of interfaces in a more realistic experiment. ...
Master thesis (2023) - F.C.J. Lijcklama à Nijeholt, Joost Broekens, J.C.F. de Winter, D. Dodou
As technology continues to evolve at a rapid pace, robots are becoming an increasingly common sight in our daily lives.
Robots that work with humans need to adapt to a variety of users and tasks, and learn to optimise their behaviour. For non-specialist users to interact with such robots, the robot's learning process needs to be transparent through its behaviour. Reinforcement Learning (RL) is a promising learning method to achieve this adaptability. However, the behaviour generated by RL is not inherently transparent because of the exploration/exploitation trade-off that is needed to optimise a policy for a specific task.

A RL algorithm is Temporal Difference (TD) learning. In TD learning, the algorithm updates a Q-table to keep track of Q-values. Q-values represent the expected future rewards that the agent (the actor that decides what action to take) can receive by taking a specific action in a certain state. Calculating the Q-values involves a value called the Temporal Difference, which is the difference between the current Q-value with the received reward added and the Q-value for the future state and chosen action.

Emotions are a natural way of communicating intent and situational appraisal for humans. In this study, emotional expressions based on Temporal Differences were implemented as a means to increase the transparency of a robot's learning progress. The effects on the robot's learning progress, learning result, and user experience were analysed.

A between-subject experiment with 61 participants on the following three robot modes was performed: no emotions, simulated emotions, and simulated emotions with matching attribution (see Table \ref{table:robotModes}). The simulated emotions are hope, fear, joy, and distress, which were expressed by a humanoid robot. The robot mode with simulated emotions and matching attributions would explain for what task it was feeling hope or fear. The task was a simple task where a human teacher had to help a humanoid robot to learn to express three different colours based on human commands.

The results demonstrate minimal differences between these three conditions. This means that for simple tasks, emotional expressions grounded in RL do not have a significant effect, and thus do not help nor hurt. The findings are discussed, and it is proposed that emotion simulation is beneficial for tasks that are more complex, afford some robot autonomy, and for which the emotion is informative about how the user should influence the robot's actions to the benefit of the robot's policy. ...
Master thesis (2023) - J.E. van Dorth tot Medler, J.C.F. de Winter, K. Elsendoorn, Y.B. Eisma, A.E. Zaidman
Background: For rigorous software testing, integration and end-to-end tests are essential to ensure the expected behavior of multiple interacting components of the system. When software is subjected to integration or end-to-end tests, it is often unfeasible to test every code change individually, as the runtime of these tests is usually significantly larger compared to unit tests. For this reason, batches of code changes from multiple authors are often tested simultaneously. Problem: An issue with testing multiple changes simultaneously is that it can be unclear which change form which author caused the failure when tests fail, as all changes from all authors included in the test can be at fault. Design: To solve this, a new automatic fault localization algorithm called GitFL is introduced, which combines state-of-the-art fault localization with version control history information for enhanced performance. GitFL was evaluated on a C++ repository at Adyen where tests are considered to be end-to-end. Findings: It showed that the addition of version control history information significantly increases the performance of fault localization for systems where multiple changes are tested simultaneously. Societal implications: This work provides insights on improved fault localization for these systems, which could enable organizations which develop these systems to speed up their testing and development processes. Originality: This work contributes by focusing on fault localization specifically for systems where multiple changes are tested simultaneously, which was not researched before.
...
Cardiovascular diseases (CVDs) are a group of disorders of the heart and blood vessels.
CVDs are the leading cause of death worldwide. To diagnose and treat CVDs, clinicians and cardiologists use multiple noninvasive imaging techniques. These scans are used to segment certain structures of the heart. Deep learning-based cardiac segmentation on short-axis cardiac magnetic resonance images (CMRI) has gained popularity over the past few years because of its generalisability and accuracy. This has exponentially reduced contouring times for clinicians. The development of such deep learning techniques has seen a common trend. In order to accommodate learning for larger cardiac datasets, the depth and size of segmentation networks have been increased. Unfortunately, the environmental impact of exploding such networks is not taken into account. One solution to mitigate having computationally expensive networks is to incorporate anatomical knowledge in the form of shape priors. The Gridnet and UNet with a shape prior are computationally efficient networks that are used to evaluate segmentation performance on a large and varied cardiac dataset (Combination of the Automated Cardiac Diagnosis Challenge - ACDC and Multi-Centre, Multi-Vendor and Multi-Disease Cardiac Segmentation challenge - M&M datasets). On average, these networks segment CMRIs with an average dice score of 0.87 and a Hausdorff distance of 11.7mm. In parallel, one of the major issues in cardiac technology is the under-representation of women in cardiac datasets. Purposefully curated cardiac datasets such as ACDC and M&M try and maintain equal representation. In real-world scenarios, this might not always be the case. Clinical trials to collect such data often report female representation as low as 25%. Evaluation of segmentation performance between a balanced and skewed dataset is conducted. This is to address if bias in such cardiac training datasets affects the performance of segmentation networks between male and female test patients.

...
An event-based camera enables capturing a video at a high temporal resolution, high dynamical range, reduced power consumption and minimal data bandwidth while the camera has minimal physical dimensions compared to a frame-based camera with the same vision properties. The limiting factor, however, of an event-based camera is the spatial resolution which ranges between 40 × 40 and 1280 × 960. To counter this deficiency, a method is researched to super resolve event-based vision in order to enhance spatial resolution. A selection of different neural network types and configurations are researched in a step-by-step fashion. Subsequent experiments tested the selected networks on their ability to process event-based data and extract features from it. Followed by experiments that exploited the limitations of the networks to super resolve at different ratios, lengths of eventstreams and more complex event-based data. Results of various experiments showed that a network configuration that utilizes a transformer architecture was best able to super resolve event-based vision. This type of network leverages the ability to extract features based on dependencies between events which aligns with the characteristics of event- based vision. Based on the obtained results from the exper- iments, a pipeline is proposed to super resolve event-based vision and consists of a combination of a transformer network, multilayer perceptrons and a k-nearest-neighbor algorithm. Using this pipeline, eventstreams can be super resolved in the spatial resolution at a scaling ratio of 4. Visually, these super resolved eventstreams resemble more detailed and enhanced version to the low-resolution input. This proposed pipeline can be considered as a starting point in further research toward the super-resolution of event-based data and thereby contributes to the extension of application possibilities of event-based vision. ...
Master thesis (2022) - A. Bakay, Y.B. Eisma, J.C.F. de Winter, D. Dodou
Human operators who are tasked with monitoring automation systems may experience a high visual demand to process the information streams from these systems. The visual sampling behavior of human operators can be described using mathematical models. These models can help designers improve environments where multiple signals are present for human operators to monitor, to a configuration that can be processed properly.

This study consisted of two parts. The first part investigated how peripheral vision plays a role in visual sampling behavior and task performance, specifically in the experimental eye-tracking setup presented in Eisma et al. (2018). In this setup, participants were instructed to monitor a bank of six dials, of which each dial pointer had a threshold indicator, and press a response key whenever a dial pointer crossed the threshold indicator. In the second part, the sampling models as presented by Senders (1983) are implemented to predict sampling trajectories. The sampling characteristics that resulted from the
predictions were then evaluated.

The results of the experiments show that peripheral vision plays a role in visual sampling and task performance. More specifically, sampling behavior is more evenly distributed among dials, and task performance is lower when peripheral vision is absent. The main attractor in the peripheral vision is shown to be the pointer speed. Moreover, the learning effect presented in Eisma et al. (2018) is not apparent when peripheral vision is absent.

The results of the predictions showcase the sampling behavior characteristics, some of which show similarities with the results from the experimental data. ...

A vulnerable road user study in a pedestrian crossing environment

Objective: In this thesis, we explore whether augmented and diminished reality interfaces, which, respectively, add and remove information from the environment, improve a pedestrian's feeling of road crossing safety, and how this information should be conveyed to the pedestrian.

Background: Literature shows that view occlusion is a prominent cause in pedestrian collisions. The research focus is currently on vehicle technology and pedestrian warning systems. Whether aiding pedestrians with camera views from unobstructed positions helps to overcome the view occlusion problem is unclear.

Methods: Twenty-eight participants engaged in a virtual reality urban road crossing scenario, in which they took on the role of a pedestrian. The pedestrian was situated on the curb and positioned such that the view on the road was largely obstructed. An autonomous vehicle approached and drove past from the left of the pedestrian. Through a head-mounted display, the participants experienced seven prototypes: baseline (i.e., no display), see-through display, transparent car, and both a head-locked and body-locked display with and without view guidance. The order in which participants encountered the prototypes was determined by a balanced Latin square, and each interface was tested by means of six trials with a non-yielding and a yielding scenario randomly selected such that in total three non-yielding and three yielding scenarios occurred in each block. The participants were instructed to continuously indicate whether they felt safe to cross by pressing a button. The interface's acceptance, workload and preference were measured with questionnaires.

Results: The participants' perceived feeling of safety revealed improved performance for all interfaces compared to the baseline condition. For the baseline condition, in which the vulnerable road user did not have access to occlusion-free information, the perceived feeling of safety was the lowest on average and decreased the earliest in the autonomous vehicle approaching phase, as well as scoring the lowest rating on acceptance. The see-through display and the transparent car interfaces, which used a combination of augmented and diminished reality properties to convey the information in a world-anchored manner to the pedestrians, achieved a higher acceptance and perceived feeling of safety than the head-locked and body-locked display interfaces.

Conclusion: A vulnerable road user's perceived feeling of safety can be increased by means of camera views from unobstructed positions to help overcome the view occlusion problem in common road crossing scenarios. This study's findings suggest a positive effect for diminished reality techniques for pedestrians, and future research could examine this technology further in more demanding scenarios. ...