Circular Image

U.K. Gadiraju

info

Please Note

93 records found

To explore high-dimensional Pareto frontiers calculated using multi-objective optimisation, analysts rely on visualisation techniques ranging from general-purpose charts to purpose-built dashboards. Prior empirical comparisons were not able to select a single chart type as best practice. Additionally, dashboards are evaluated as complete systems, and no study covers investment portfolio optimisation. We address this gap through a study in which domain experts helped characterise the analytical tasks and served as an initial filter for the techniques we evaluate. We then assessed six visualisation techniques arranged in an interactive dashboard with 57 non-expert and 10 expert participants, across a single-frontier analysis and a frontier comparison on 3D and 6D datasets. The results show that the view choices depend on the task being solved, while no single technique dominates overall in usage or performance. Specialised views stand out on specific tasks, while the tables' role as a robust and easy-to-learn baseline is reinforced by the findings in this study. Our framework is available at: \url{https://github.com/Gunterinos/master_thesis}
...
Master thesis (2026) - Z. LEI, U.K. Gadiraju, S.K. Freire, E. Niforatos
Large language model (LLM) agents can plan, call tools, and incorporate intermediate results, but errors may arise at different workflow stages and propagate without being apparent in the final outcome. Human-visible plan-then-execute workflows create opportunities for oversight, yet process visibility alone does not tell users when intervention is warranted or what response remains feasible. This thesis investigates how risk-aware human oversight support can be designed and evaluated for such workflows. The proposed framework identifies PLAN, ACTION/pre-execution, and OUTPUT/tool-output as three types of critical oversight junctures. At each juncture, it keeps task-impact assessment separate from a potential-error assessment based on stage-specific runtime evidence. A rule-based mapping translates these inputs into an Oversight Cue that presents a suggested response and assessment rationales alongside the controls available at that stage. Risk-aware refers to this use of task impact in the framework design; it was not manipulated as an independent mechanism. The framework was evaluated in a between-subjects study with 61 participants completing four simulated daily assistant tasks. Both conditions used the same staged workflow, artifacts, checkpoints, and intervention controls. Condition B additionally received the complete Oversight Cue, whereas Condition A used unguided staged oversight. Condition B achieved higher appropriate reliance (mean difference = .107, 95% confidence interval (CI) [.051, .163], Hedges' g = .95), primarily through a higher correct intervention rate (mean difference = .205, 95% CI [.106, .304], g = 1.03). Correct acceptance did not differ significantly between conditions. Condition B also produced higher-quality plans (mean difference = .25, 95% CI [.11, .40], g = .86). The primary analysis favored Condition B for action-sequence accuracy, but this result did not remain significant after covariate adjustment and familywise correction. No statistically detectable condition differences were found for final-outcome accuracy or overall perceived workload. The findings show that structured oversight support can improve intervention on unacceptable agent behavior and the quality of an intermediate workflow artifact without producing corresponding improvements in every downstream outcome. Because the cue elements were presented as a bundle and the LLM-based assessments were not independently validated, the results apply to the integrated implementation rather than to the independent effectiveness of each component or the diagnostic accuracy of the assessments.
...

An Exploratory Study into Trust, Perceived Trustworthiness, and Opinion Formation on Simulated Users

With the current rise of Large Language Models (LLMs), it also raises concerns that sycophantic responses may influence how users form opinions and trust in such models. This paper investigates how LLM sycophancy affects trust, perceived trustworthiness, and opinion formation among simulated users representing young adults. A 2 x 2 mixed experimental design was conducted in which simulated users between the ages of 18 and 25 interacted with either a neutral or sycophantic model across two topics: autonomous vehicles and AI in society. Users completed pre- and post-interaction questions for each topic. In addition, their open-ended reflection responses were qualitatively analyzed. Both neutral and sycophantic conditions were configured on the model Llama 3.1 8B. The results suggest that the sycophantic model increases perceived trustworthiness, while the effects on trust and opinion formation were insignificant. These results indicate that sycophantic behavior may make models appear more trustworthy even when it does not strongly influence users' opinions or trust. Results from manipulation checks show that there was only a significant difference in perceived validation between both conditions, suggesting that the perceived trustworthiness may have been influenced more by validation than by broader sycophantic behavior. Since the study uses simulated users, the results should be interpreted as exploratory rather than direct evidence of human behavior. The paper contributes an experimental setup for studying LLM sycophancy on simulated users and highlights the need for further validation with real human participants. ...

Using AI Personas to Evaluate Trustworthiness and Misinformation Detection

The increasing use of generative AI has had a significant impact on how people experience, interact, and interpret media. The widespread adoption of generative AI has raised concerns regarding the spread of AI-generated misinformation and its influence on the perceived trustworthiness of information. This study investigated how AI-personas representing young adults evaluated AI-generated and human-generated statements. A mixed factorial experimental design was used with three independent variables: statement truthfulness, statement source, and source label visibility. 124 AI-personas completed surveys where they were asked to evaluate short statements based on their truthfulness, confidence and trustworthiness. Mixed ANOVA was conducted to examine the effects of content source, truthfulness and labeling.

The results showed that AI-generated misinformation was not identified less accurately than human-generated misinformation. Source labeling did not significantly affect confidence in truthfulness judgments. Trustworthiness ratings were significantly influenced by both statement condition and label visibility. When source labels were hidden, AI-generated statements received higher trustworthiness ratings than human-generated statements. However, when the source labels were revealed, the trustworthiness ratings for AI-generated content were reduced, while human-made statements received higher trustworthiness scores. These findings suggest that knowledge of content origin influences the perceived trustworthiness. ...
AI-generated content has become very hard to distinguish and it has evolved into a challenge for users to judge whether the media was created by a human or by a machine. This study examines whether AI and media literacy interventions can improve the ability of AI-agent personas, prompted as young adults, to detect AI-generated texts and images. Twenty AI-agent personas completed pre- and post-intervention detection tasks across both modalities. Overall detection accuracy increased from 85.75% before the intervention to 94.25% after the intervention, with a larger improvement for image stimuli compared to text stimuli. Text detection accuracy was already high before the intervention, while image detection still showed room for improvement. The findings suggest that AI and media literacy guidance can produce measurable changes by using specific cues, but they should not be treated as direct evidence of how real young adults would respond. This study contributes an exploratory test of using AI-agent personas to evaluate intervention designs before human-participant research. ...

A Guiding Tool for the Design of Incentive Formulas in Crowdsourcing

Master thesis (2026) - V. Macsim, U.K. Gadiraju, M.L. Tielman
The rapid growth of artificial intelligence has driven demand for large volumes of real-world data, making crowdsourcing an essential practice. However, crowdsourcing remains largely unregulated, with minimal disclosure of compensation practices in academic literature or dataset documentation. This lack of transparency undermines two important goals: collecting high-quality, realistic data for AI systems, and ensuring fair treatment of workers. Without clear guidance on incentive design, it becomes difficult to distinguish between requesters' lack of knowledge and poor practices—a problem that affects both data quality and worker welfare.

To address this gap, a wizard tool was developed to guide requesters through the process of designing payment schemas for crowdsourcing tasks. A user study was conducted to investigate how structured guidance affects incentive design: first, by comparing designs created with and without the tool, and second, by examining whether the tool produces consistency in compensation decisions across different requesters. The study evaluated both the designs participants created and their feedback on the tool itself.

The analysis reveals three primary insights. First, the tool's primary strength lies in structuring the design process rather than fundamentally altering participants' compensation decisions. The extent to which structured guidance benefited participants depended significantly on their prior experience with crowdsourcing, suggesting that the tool's value is contingent on user expertise. Second, the tool produced convergence around a limited set of high-level design elements, though participants used varied implementation approaches within these patterns, such as specific bonus sums.

These findings indicate that the tool could serve a valuable function in documenting and contextualizing design rationales, capturing the constraints and considerations that shaped dataset creation decisions. However, realizing the tool's full potential as a design aid requires enhancements to customization options and user experience refinement. Despite these limitations, the tool shows promise as an educational resource for introducing beginners to crowdsourcing incentive design, offering a structured entry point into a complex domain.
...

How Does External Cognitive Load Affect Young Adults’ Ability to Evaluate AI-Generated Content?

In recent years, there has been a gradual increase in the use of generative artificial intelligence (AI) among young adults. At the same time, they tend to process textual information while under conditions of divided attention. As a result, young adults may encounter AI-generated misinformation when their cognitive resources are occupied, potentially affecting their ability to evaluate information critically. Previous research has linked external cognitive load (CL) to task performance, but less is known about its impact on the evaluation of AI-generated misinformation. To address this gap, this study used a simulated experiment in which AI personas representing young adults evaluated the veracity of AI-generated true and false statements under no-load, low-load, and high-load conditions, measuring accuracy, confidence, and sharing intention. High CL reduced personas' accuracy and confidence in evaluating veracity, whereas low CL did not differ significantly from the no-load condition. No statistically significant effect of CL was found for sharing intention. As the study is simulation-based, the results should not be interpreted as direct evidence of the behaviour of real young adults. ...

Fostering Responsible Opinion Formation Among Young Adults in the Age of Generative AI

The use of LLMs (Large Language Models) as "thinking partners", conversational partners actively partaking in user's reasoning, is on the rise. As young adults become the demographic that engages with LLMs the most, concerns over whether different AI "thinking partner" styles can help or hinder responsible opinion formation become more prevalent. This study investigates how three "thinking partner" styles, Steelman, Socratic, and Neutral, affect opinion change, epistemic trust, and epistemic autonomy in simulated young adult participants. A between subjects study was conducted, using simulated personas as participants. Each persona engages with a "thinking partner" condition for a five exchange session on the topic of individual versus systemic responsibility for climate action. Opinion change differed significantly across conditions, with the Steelman producing a shift away from individual climate action, while the Socratic and Neutral produced comparable positive shifts towards agreement. No significant change was noted for epistemic trust and autonomy, both of which were rated consistently high regardless of the condition. These findings suggest that an adversarial AI may provoke resistance rather than persuasion, while trust and sense of autonomy is preserved across interaction styles. This study serves as a preliminary methodological pilot, future work should replicate the experiment with human participants. ...
Master thesis (2026) - K. Yordanov, U.K. Gadiraju, P.K. Murukannaiah
The spread of online misinformation undermines responsible opinion formation, this risk being exacerbated for heavily debated topics where emotionally-driven reasoning and the formation of echo chambers weaken critical scrutiny. Visual misinformation labels which nudge users towards more mindful information-seeking behaviors are common-place, yet, despite its increasing popularity, the podcast medium remains overlooked in these efforts compared to Web search. This thesis examines the design of auditory misinformation warnings: brief, non-verbal sound cues embedded in podcast audio to indicate that a nearby statement is misleading. To verify the relevant criteria for these signals’ effectiveness and the factors influencing their reception by listeners, two exploratory crowdsourcing studies were conducted. The first assessed 15 auditory icons in terms of recognizability, consistency of conceptual mapping, and induced disruption, isolating the most favorably perceived candidates. The second embedded the latter as purposeful interventions within AI-generated podcast dialogues on three controversial scientific topics interspersed with myths and falsehoods, employing a between-subjects design across four warning-placement configurations (before, after, enclosing, and concurrent with a misinformation occurrence) and treating several psychometric traits as exploratory factors. The findings indicate salient auditory icons that alert listeners are a viable choice provided they do not severely depart from the podcast’s setting or listeners’ expectations of traditional podcast-editing effects; otherwise, they may trigger prolonged dissatisfaction and ultimately break immersion. No conclusive data on the cues’ role in misinformation recognition emerged, although some listeners correctly inferred their intended function when these followed or enclosed a misleading claim. Warnings placed after a falsehood were least disruptive while concurrent placement was associated with the lowest content recall. Regarding contextual factors, listeners’ prior stance on a debated topic and a podcast’s perceived density and pace exhibited significant correlations with several Likert-scale sound-cue evaluations. Participant’s reactions and follow-up attitudes towards the interventions diverged significantly – anticipation, habituation, and active resistance against perceived paternalism all being expressed. These results yield preliminary recommendations for auditory misinformation warnings and, given the many non-trivial trade-offs at play, outline a multitude of directions for future research. ...
Master thesis (2026) - R.Q. van Berkel, U.K. Gadiraju, B.J.W. Dudzik, Reem Younan
Root cause analysis (RCA) in large-scale industrial software systems requires engineers to correlate heterogeneous logs, source code, documentation, and domain knowledge, often under time pressure. This thesis studies whether agentic AI can support this work through human-agent collaboration. In a case study at ASML, we design and evaluate TalkToYieldStar, a multi-agent system for RCA on diagnostic log dumps from wafer metrology tools. The system decomposes RCA into bounded subtasks, including log inspection, source-code search, knowledge retrieval, hypothesis formation, and report generation. Evaluation through a case study with eleven professional engineers shows that agentic AI is useful primarily as an evidence-structuring investigation aid rather than an autonomous diagnostician: 8 of 11 participants reported the agent session as more efficient, and all rated willingness for daily use positively. Participants valued support for finding evidence and offloading tiring subtasks—including in unfamiliar domains and off-hours situations—but required steering, verification, and final human judgment. Gains were conditional on case type: the hardest cases, requiring deep domain-state reconstruction, resisted both human-guided and one-shot approaches, with hallucinated reasoning appearing in both baseline systems on those cases. The findings indicate that reliable industrial RCA agents need calibrated autonomy, inspectable evidence trails, domain-state representations, clearer planning checkpoints, and reduced computational overhead.  ...

The Interplay of Experience, Information Design, and Intervention Options

Master thesis (2026) - D. Viero, U.K. Gadiraju, T.A. Draws, D.S. Murray-Rust
Human oversight of artificial intelligence systems is increasingly mandated by regulation, yet empirical evidence on which interface design choices actually improve oversight quality remains scarce. This thesis investigates how two modifiable design factors, information signal granularity and intervention option range, affect the effectiveness and perceived workload of human overseers in an AI-assisted fraud detection task. A 3×3 between-subjects experiment was conducted online via Prolific (N = 144), in which participants reviewed 30 bank account applications flagged by a Gradient Boosting model under one of nine conditions, crossing three levels of information signal (plain case features, categorical risk level indicator, and continuous model confidence score) with three levels of intervention options (binary decision, decision with delegation, and decision with flagged delegation). Oversight effectiveness was operationalised as a composite score rewarding correct classifications and appropriate delegation decisions, and penalising both misclassifications and over-delegation. Perceived workload was measured using the NASA Task Load Index. Self-reported domain expertise was included as a covariate. Neither main effect reached the pre-registered Bonferroni-corrected significance threshold of α = 0.008, and no significant interaction was found on either dependent variable. A marginal effect of information signal granularity on oversight effectiveness was observed (F(2, 134) = 3.29, p = .040, η²p = .047), with a non-monotonic pattern in which the categorical risk level indicator outperformed both the plain information baseline and the continuous model confidence score. This reversal of the hypothesised ordering suggests that, for non-expert overseers, a well-designed categorical signal may be more actionable than a continuous probability score, as interpreting the latter requires complementary domain knowledge not uniformly present in a general population. The results are treated as preliminary, since the study was underpowered relative to the pre-registered target of N = 288, due to the planned expert-screened wave not being completed within the thesis timeline. Theoretical and practical implications for the design of human oversight interfaces are discussed, with particular attention to the relationship between signal granularity and overseer expertise. ...
Social engineering is a major threat in today's cybersecurity landscape. Unlike other types of cyberattacks, social engineering focuses on exploiting human psychology rather than technical vulnerabilities. By manipulating individuals into revealing sensitive information or taking actions against their own interests, malicious actors can cause significant harm, ranging from financial loss to compromised systems. As a socio-technical problem, it requires not only technical measures but also awareness on the human side to defend against these threats. With advancements in technology and, most recently, generative artificial intelligence, social engineering attacks are becoming increasingly sophisticated, highlighting the urgency to educate individuals about how to recognise and defend against them. While existing efforts in research and industry mostly focus on interventions for workplace settings, little attention has been given to approaches targeted at the general public or everyday life contexts. This thesis addresses this gap by designing and evaluating a serious game to raise awareness of social engineering, focusing on families as the target group. The resulting game, “Connected & Protected”, employs an asymmetric tabletop format in which one player takes on the role of a social engineer, while the others play as family members, trying to protect their personal data from the attacker. The aim of the game is to educate players about common influence principles used in social engineering, how to counter them, and the risks that come with sharing personal data. The project takes a human-centred approach and involves users in multiple phases, including feedback sessions and playtesting to support concept development and prototyping. The game was evaluated using a pre- and post-test study design with 15 participants across five family groups, including a delayed post-test two weeks later to measure knowledge retention. The evaluation focused on three main areas: social engineering awareness, self-perceptions, and game experience. Results indicated that the game had positive effects on social engineering awareness, with significant improvements after gameplay that were largely retained at the delayed post-test. Participants improved in their ability to recognise influence principles in social engineering scenarios and showed increased knowledge about defensive strategies and data-sharing risks. Self-efficacy and perceived preparedness to deal with social engineering attempts likewise increased significantly after gameplay, while perceived susceptibility did not change notably. The game experience was overall well-received, with particularly high scores for enjoyment and audiovisual appeal, along with a positively rated perceived learning effect and willingness to play the game again. With this project, we provide novel insights into designing serious games for social engineering education aimed at the general public, leveraging the family setting as a medium for shared learning and intergenerational knowledge exchange. The contributions can serve as a starting point for practitioners and policymakers to extend security awareness interventions beyond the workplace, making social engineering education more accessible by bringing it into people's homes. ...
Master thesis (2025) - A.R. Moraru, U.K. Gadiraju, S. Biswas, P. Pawelczak
Error messages are a primary feedback channel in programming environments, yet they often obstruct  progress, especially for novices. Although large language models (LLMs) are widely used for code  generation and debugging assistance, there is limited empirical evidence that LLM-rephrased error  messages consistently improve code correction capabilities, and skill-adaptive designs remain largely  unexplored. We introduce a framework that uses an LLM to rewrite Python standard interpreter errors  in two different styles, which are designed to be tailored to user expertise: the pragmatic style, which is  concise and action oriented, and the contingent style, which provides scaffolded, actionable guidance  organized by a clear argumentation model. To measure Python skill level reliably, we first ran a pilot  study that informed the design of a short 8 multiple-choice question assessment focused on debugging  and error-message interpretation. We then used this instrument in the main study to classify participants  as either novices or experts.

To gauge the effectiveness of our LLM-enhanced programming error messages (PEMs), we evaluated  the framework in a crowdsourced Prolific study with 103 participants. We measured objective outcomes  such as fix rate, time to fix, and number of attempts to fix, while also capturing subjective perceptions of  PEMs, including readability, cognitive load, and authoritativeness. Objectively, LLM-enhanced PEMs  showed favorable trends but did not produce statistically significant improvements over the standard  interpreter. Subjectively, novices and experts alike, rated the pragmatic messages as significantly  more readable and helpful, lower in intrinsic and extraneous cognitive load, and considerably less  authoritative. Contingent messages exceeded the baseline on average but did not consistently reach  statistical significance across all of our measurements, which points to a need for tighter control of error  message verbosity and granularity, particularly for beginners.   

These results show that LLMs, especially small-sized ones, are already capable of delivering targeted  text-rewriting interventions that improve the perceived quality of error feedback. Future work should  validate the effects at larger scale and across languages, expand coverage of real-world error contexts,  and pursue true adaptivity in which error message style and level of detail adjust dynamically to user  skill and task state.

...

Design, Implementation, and Evaluation of Adaptive XAI Feedback

Bachelor thesis (2025) - A.S. Kumar, D. Zhan, U.K. Gadiraju, M.A. Neerincx
Pose estimation models offer promising opportunities for automated feedback in cricket training, but their practical impact is limited by the lack of personalized and understandable explanations. This study investigates how explanation formats can be tailored to users’ expertise levels, focusing on beginner, intermediate, and expert levels, to improve the effectiveness of AI-generated feedback. Based on a literature review of explanation needs and generation methods, we propose a taxonomy linking expertise levels to suitable explanation modes: visual, comparative, and statistical. We implement a set of explanation prototypes aligned with this taxonomy and evaluate them through a user study involving 17 participants across the three expertise levels. Results show that participants rated explanations tailored to their skill level as more useful, trustworthy, and easier to interpret. Statistical validation using Kruskal-Wallis and Dunn’s tests confirmed significant differences in perception between user groups, especially between beginners and experts. These findings demonstrate the value of expertise-based explanation design in cricket analytics and offer design guidelines for future explainable pose estimation systems in sports
...

Addressing challenges and Inter-Keypoint dependencies in Cricket Pose Analysis

Bachelor thesis (2025) - A.M. Semov, U.K. Gadiraju, D. Zhan
Pose estimation models predict multiple interdependent body keypoints, making them a prototypical example of multi-target tasks in machine learning. While existing explainable AI (XAI) techniques have advanced our ability to interpret model outputs in single-target domains, their application to structured outputs remains underdeveloped. This work investigates how XAI methods can be adapted to explain pose estimation models, particularly in the context of cricket shot analysis. Guided by three research questions, we identify key challenges such as capturing inter-keypoint dependencies and providing interpretable explanations of structured outputs. We analyze both geometric and heatmap-level behavior of a pretrained pose estimation model to distinguish between two cricket shots - the pull and the cover drive. Through techniques like cosine similarity on heatmaps and polynomial trajectory modeling, we reveal how the model internally differentiates between similar motion patterns. Our framework introduces novel techniques for inter-keypoint explanation, contributes domain-specific insights into model behavior, and demonstrates the feasibility of interpretable structured predictions in high-dimensional, real-world tasks. ...

An Experimental Study on Human-like Chatbot Design and Question Sensitivity in Mental Health Contexts

AI-powered mental health chatbots offer scalable and accessible support, but their effectiveness hinges on users’ willingness to self-disclose—an outcome shaped by chatbot communication style and the sensitivity of the topic. While prior work has explored empathy and rapport, the role of conversational anthropomorphism remains underexamined, particularly in relation to question sensitivity as a potential moderator. This study addresses that gap through a mixed-design experiment (n = 30) in which participants interacted with either an anthropomorphic or a neutral chatbot and rated their willingness to respond to questions varying in sensitivity. Although no effects reached statistical significance, descriptive trends suggest that anthropomorphic cues—such as informal tone, emojis, and adaptive responses—may increase willingness to disclose, while higher question sensitivity slightly reduces it. No significant interaction effect was found, but anthropomorphic language appeared to promote disclosure regardless of sensitivity level. These findings offer tentative support for the use of calibrated human-like design in mental health chatbots. Future work should incorporate open-ended interactions, behavioral measures, and longitudinal designs to better capture disclosure dynamics and trust formation. ...

A User Study on the Importance of Privacy and Question Sensitivity in Mental Health Chatbots

Mental health chatbots are increasingly adopted to address shortage mental health services, by offering non-judgmental, always-available support. User self-disclosure is a critical factor which allows mental health chatbots to better understand users and provide more therapeutic experiences. Although prior work has explored how factors such as chatbot modality and tone affect self-disclosure, the role of privacy policies and how question sensitivity affects disclosure remains under examined. In this study, we investigate how privacy policies and the sensitivity of questions in voice-based mental health chatbots impacts user self-disclosure. Through a controlled user study, we explore whether the presence of a privacy policy leads to increased self-disclosure, whether question sensitivity influences self-disclosure willingness and whether there is any interaction effect between these two factors. Preliminary findings indicate that while providing a privacy policy did not significantly impact users' privacy understanding or willingness to self-disclose, question sensitivity notably influenced disclosure. Specifically, participants were more willing to disclose to low and medium sensitivity questions compared to high sensitivity. No interaction effect between the privacy policy and the question sensitivity was observed. Future research should expand participant pools, investigate self-disclosure in free-form interactions, and explore alternative methods of communicating privacy information for deeper insights into user perceptions regarding privacy, sensitivity and disclosure. ...

The Impact of Self-Disclosure Techniques on the User Disclosure

Bachelor thesis (2025) - Y. Shan, U.K. Gadiraju, E.C.S. de Groot, M.L. Tielman
As mental health issues continue to rise around the world, AI chatbots are becoming a promising way to provide accessible and scalable support. This study explores how different levels of chatbot self-disclosure affect users’ willingness to share personal information in a mental health context. A within-subjects experiment was conducted with 94 participants, each interacting with three versions of a chatbot: one with no self-disclosure, one with factual self-disclosure, and one with emotional self-disclosure. Participants engaged in a fictional role-play and rated their willingness to disclose across five personal topics, as well as their level of trust and comfort with the chatbot.

The chatbot using factual self-disclosure received the highest average scores for trust, comfort, and willingness to disclose. However, statistical tests (ANOVA) showed no significant differences between chatbot types on these measures, except for changes in willingness to disclose. Participants who interacted with the emotional chatbot were more likely to report a negative change in their willingness to share. This result was unexpected and suggests that emotional self-disclosure may reduce user openness during early interactions, possibly because it feels unnatural or too personal too soon.

These findings show that emotional expression is not always the best approach. Instead, it is important to match the chatbot’s disclosure style to the situation and the user's comfort level, especially in sensitive areas like mental health support. ...
Bachelor thesis (2025) - G. Vitner, U.K. Gadiraju, D. Zhan, M.A. Neerincx
Explainable Artificial Intelligence (XAI) has the potential to enhance user understanding and trust in AI systems, especially in domains where interpretability is crucial, such as cricket training. This study investigates the impact of different explanation formats on user experience within a cricket-specific context. Two prototypes were developed, each including four explanation formats: textual, visual, rule-based, and mixed. The second prototype introduced interactive features to examine their influence on user experience and explanation effectiveness. A small-scale user study evaluated the explanations based on satisfaction and trust. Results show that rule-based explanations were significantly less preferred in terms of satisfaction than the other explanation formats. Furthermore, the addition of interactive features led to a significant increase in user trust, though they did not enhance satisfaction levels. These findings highlight the importance of selecting appropriate explanation formats and the potential of interactive features to enhance trust in AI-generated explanations in a cricket-specific context. ...
This study investigates whether empathetic language in chatbot interactions influences users’ willingness to disclose mental health-related information. Using a two-by-two mixed factorial design, 114 participants were assigned to either an empathetic or neutral chatbot condition and responded to both emotional and behavioural health questions. While prior research suggests that empathy can pormote trust and openness, results from this study revealed no significant difference in disclosure willingness across chatbot styles or question types, although the manipulation check showed a well perceived empathy. The study highlights the importance of individual predispositions, such as prior readiness to disclose, in shaping interactions with digital mental health tools. Future work should explore longer-term interactions and real disclosure behaviour to better understand the role of empathy in chatbot design. ...