Circular Image

U.K. Gadiraju

info

Please Note

100 records found

Conference paper (2026) - Shreyan Biswas, Alexander Erlei, Ujwal Gadiraju
Large language models (LLMs) increasingly support heterogeneous tasks within a single interface, requiring users to form, update, and act upon beliefs about one system across domains with different reliability profiles. Understanding how such beliefs transfer across tasks and shape delegation is therefore critical for the design of multipurpose AI systems. We report a preregistered experiment (N = 240, 7,200 trials) in which participants interacted with a controlled AI simulation across grammar checking, travel planning, and visual question answering, each with fixed, domain-typical accuracy levels. Delegation was operationalized as a binary reliance decision - accepting the AI's output versus acting independently and belief dynamics were evaluated against Bayesian benchmarks. We find three main results. First, participants do not reset beliefs between tasks: priors in a new task depend on posteriors from the previous task, with a 10-point increase predicting a 3-4 point higher subsequent prior. Second, within tasks, belief updating follows the Bayesian direction but is substantially conservative, proceeding at roughly half the normative Bayesian rate. Third, delegation is driven primarily by subjective beliefs about AI accuracy rather than self-confidence, though confidence independently reduces reliance when beliefs are held constant. Together, these findings show that users form global, path-dependent expectations about multipurpose AI systems, update them conservatively, and rely on AI primarily based on subjective beliefs rather than objective performance. We discuss implications for expectation calibration, reliance design, and the risks of belief spillovers in deployed LLM-based interfaces. ...
Conference paper (2026) - Tim Schrills, Patricia Kahr, Markus Langer, Harmanpreet Kaur, Ujwal Gadiraju
As AI systems are increasingly adopted in high-stakes domains such as healthcare, autonomous driving, and criminal justice, their failures may threaten human safety and rights. Human oversight of AI systems is therefore critically important, as a potential safeguard to prevent harmful consequences in high-risk AI applications. Although regulations like the European AI Act mandate human oversight for high-risk AI, we lack methodologies and conceptual clarity to implement it effectively. Independent of policy and regulation, poorly designed oversight can create dangerous illusions of safety while obscuring accountability. This interdisciplinary workshop aims to bring together researchers from various disciplines, including AI, HCI, psychology, law, and policy, to address this critical gap. We will explore the following questions — How can we design AI systems that enable meaningful human oversight? What methods effectively communicate system states and risks to human overseers? How do we ensure scalable and effective interventions? Through papers, talks, and interactive group discussions, participants will identify oversight challenges, examine stakeholder roles, discuss supporting tools, methods, regulatory frameworks, and establish a collaborative research agenda. Our central goal is to further a roadmap that enables effective human oversight for the responsible deployment of AI in society. ...

Privacy Harms vs. Economic Risk in Personalized AI Adoption

Conference paper (2026) - Alexander Erlei, Tahir Abbas, Kilian Bizer, Ujwal Gadiraju
Privacy concerns significantly impact AI adoption, yet little is known about how information environments shape user responses to data leak threats. We conducted a 2 × 3 between-subjects experiment (N = 610) examining how risk versus ambiguity about privacy leaks affects the adoption of AI personalization. Participants chose between standard and AI-personalized product baskets, with personalization requiring data sharing that could leak to pricing algorithms. Under risk (30% leak probability), we found no difference in AI adoption between privacy-threatening and neutral conditions (ca. 50% adoption). Under ambiguity (10-50% range), privacy threats significantly reduced adoption compared to neutral conditions. This effect holds for sensitive demographic data as well as anonymized preference data. Users systematically over-bid for privacy disclosure labels, suggesting strong demand for transparency institutions. Notably, privacy leak threats did not affect subsequent bargaining behavior with algorithms. Our findings indicate that ambiguity over data leaks, rather than only privacy preferences per se, drives avoidance behavior among users towards personalized AI. ...
Conference paper (2026) - Saumya Pareek, Nattapat Boonprakong, Naja Kathrine Kollerup, Si Chen, Simo Hosio, Koji Yatani, Yi Chieh Lee, Ujwal Gadiraju, Niels van Berkel, Jorge Goncalves
Despite decades of advancements in Artificial Intelligence (AI), fostering appropriate trust in AI systems remains a challenge. Cognitive biases - systematic deviations from rational judgement - profoundly influence human decision-making, and reliance on such “mental shortcuts” can make AI systems appear more or less trustworthy than they really are, often undermining collaboration outcomes. As AI evolves with more sophisticated and persuasive natural language outputs, particularly through Generative AI (GenAI) and Large Language Models (LLMs), these biases may manifest in new and unpredictable ways, calling for their comprehensive examination. This workshop brings together diverse researchers from HCI, human-centred AI, cognitive psychology, interaction design, and related fields to collaboratively explore how cognitive biases influence trust calibration in human-AI interaction and establish a research agenda. We will explore how biases emerge across the human-AI interaction pipeline, what design strategies can mitigate or even harness these heuristics, and what methods are needed to study these dynamics effectively. Through a highly interactive 90-minute session, participants will map out open challenges, brainstorm tensions and solutions, chart future research directions, and share perspectives from their own diverse disciplinary lenses. Through this workshop, we aim to build a shared understanding of how cognitive biases influence trust in evolving AI systems, and derive a forward-looking, bias-aware research agenda that promotes appropriate trust in human-AI interaction. ...

Using LLMs to Support Illegal Content Reporting under the Digital Services Act

Conference paper (2026) - Marie Therese Sekwenz, Shreyan Biswas, Rita Hermann-Gsenger, Ujwal Gadiraju
Illegal content reporting mechanisms are a key technical and organizational measure through which online platforms address the dissemination of illegal content under European Union law. Under the Digital Services Act (DSA), user notices submitted pursuant to Article 16 must be sufficiently substantiated and provided in good faith, requiring users to interpret legal and procedural language and translate it into legally meaningful categories and reasons. In practice, however, reporting illegal content remains cumbersome across major social media platforms, placing substantial cognitive and legal demands on users. Without effective support at the reporting interface, operationalizing Article 16 in practice remains challenging. We investigate how large language model (LLM)-based assistants can support illegal content reporting. In a controlled user study (N = 450) using an interface modeled on a major platform's reporting workflow, we compare three conditions: (1) a conventional explainable AI assistant (XAI) that suggests a single legal category with a rationale, (2) an evaluative AI assistant (EvalAI) that presents balanced pro and con arguments across candidate legal provisions for user deliberation, and (3) a baseline reflecting unaided reporting (Baseline). We further examine these assistance forms under systematically varied AI error regimes. Our results show that EvalAI improves provision-level accuracy under AI error regimes and reduces misclassification distance relative to conventional XAI, particularly for near-miss and overbreadth errors. In contrast, conventional XAI does not improve - and can degrade - the quality of users' rationales relative to unaided reporting, despite enabling faster decisions when the AI output is correct. We discuss implications for the design of compliance-oriented reporting interfaces, highlighting trade-offs between accuracy, deliberation, and vulnerability to misleading AI output. ...

Reflexive Annotating for Situated AI Alignment

AI alignment relies on annotator judgments, yet annotation pipelines often treat annotators as interchangeable, obscuring how their social position shapes annotation. We introduce reflexive annotating as a probe that invites crowd workers to reflect on how their positionality informs subjective annotation judgments in a language model alignment context. Through a qualitative study with crowd workers (N = 30), including follow-up interviews (N = 5), we examine how our probe shapes annotators' behaviour, experience, and the situated metadata it elicits. We find that reflexive annotating captures epistemic metadata beyond static demographics by eliciting intersectional reasoning, surfacing positional humility, and nudging viewpoint change. Crucially, we also denote tensions between reflexive engagement and affective demands such as emotional exposure. We discuss the implications of our work for richer value elicitation and alignment practices that treat annotator judgments as situated and selectively integrate positional metadata. ...
Conference paper (2026) - A. Erlei, F.M. Cau, R. Georgiev, S. Chethan Kumar, K. Bizer, U. Gadiraju
AI consumer markets are characterized by severe buyer-supplier market asymmetries. Complex AI systems can appear highly accurate while making costly errors or embedding hidden defects. While there have been regulatory efforts surrounding different forms of disclosure, large information gaps remain. This paper provides the first experimental evidence on the important role of information asymmetries and disclosure designs in shaping user adoption of AI systems. We systematically vary the density of low-quality AI systems and the depth of disclosure requirements in a simulated AI product market to gauge how people react to the risk of accidentally relying on a low-quality AI system. Then, we compare participants' choices to a rational Bayesian model, analyzing the degree to which partial information disclosure can improve AI adoption. Our results underscore the deleterious effects of information asymmetries on AI adoption, but also highlight the potential of partial disclosure designs to improve the overall efficiency of human decision-making. ...
Journal article (2026) - Esra Cemre Su C. de Groot, Ujwal Gadiraju, Olya Kudina, Loes Keijsers, Manon H. J. H. Hillegers, Willem-Paul Brinkman
Digital technologies are on the rise to promote health. To improve the engagement and effectiveness of these technologies, there is a growing interest in algorithmic personalization. However, the user input data for these algorithms (e.g., data from wearables or self-reported data) can come with ethical and regulatory implications. Despite a growing amount of theoretical work, there is no practical precedent on how to consider these implications in the development of personalization algorithms. Therefore, our work aims to tackle this challenge by proposing a stepwise method for Responsible Data Selection (ReDS) for algorithmic personalization of mHealth. The ReDs method acts from a duty of care and promotes an active search for ethically less risky data. We demonstrate the six-step method through a real-world use case on an mHealth app promoting adolescents' mental well-being, using a dataset of 1181 adolescents (5199 interactions) who received coping strategy challenges based on cognitive behavioral therapy. First, we identified the personalization objective in the case study (step 1). The objective was to personalize the type of challenge to promote adherence while diversifying the coping strategy types within the completed challenges. Next, we identified the emotional state of the adolescent and prior completion rates as promising input data (step 2). However, personal emotion data can be considered sensitive, personal, and private, implying ethical implications (step 3). As a potential alternative, tiredness data can be perceived as less sensitive to share and collect (step 4). Subsequently, we analyzed the utility of all data features (step 5) using evaluative simulations with reinforcement learning models. This revealed that solely using the completion rates of the previous day could already benefit the personalization objective and that adding emotion data or tiredness data could similarly further increase the performance of the personalization algorithm. When determining the utility-risk trade-off (step 6), we conclude that tiredness data can be used as an alternative for emotion data if risk mitigation strategies are deployed. Through this case study, we demonstrate the practical utility of the ReDS method. We hope that our work will inspire future developers of personalization algorithms to explicitly incorporate ethical considerations in the algorithm development process. ...
Conference paper (2026) - Malik Khadar, Julia Cecil, Leon Van Der Neut, Nikola Banovic, Kevin Baum, Stevie Chancellor, Enrico Costanza, Ujwal Gadiraju, Harmanpreet Kaur, More Authors
As AI systems are increasingly adopted in high-stakes domains such as healthcare, autonomous driving, and criminal justice, their failures may threaten human safety and rights. Human oversight of AI systems is therefore critically important as a potential safeguard to prevent harmful consequences in high-risk AI applications. The global regulatory and policy landscape for AI governance remains understandably fragmented and diverse. While frameworks like the European AI Act require human oversight for high-risk AI systems, there is currently a lack of well-defined methodologies and conceptual clarity to operationalize such oversight effectively. Independent of policy and regulation, poorly designed oversight can create dangerous illusions of safety while obscuring accountability. This interdisciplinary workshop aims to bring together researchers from various disciplines, including AI, HCI, psychology, law, and policy, to address this critical gap. We will explore the following questions: (1) What are the greatest challenges to achieving effective human oversight of AI systems? (2) How can we design AI systems that enable meaningful human oversight? (3) How do we assign responsibilities to and support the various stakeholders involved in oversight? Through talks and interactive group discussions, participants will identify oversight challenges; examine stakeholder roles; discuss supporting tools, methods, and regulatory frameworks; and establish a collaborative research agenda. Our central goal is to further a roadmap that enables effective human oversight for the responsible deployment of AI in society. ...

Qualitative and quantitative insights from a user survey of a mental health promoting app

Journal article (2026) - Esra Cemre Su de Groot, Lianne P. de Vries, Ujwal Gadiraju, Olya Kudina, Loes Keijsers, Manon H.J. Hillegers, Willem Paul Brinkman
While mental health apps can help to promote adolescents’ mental health, prevent mental health problems, and reduce symptoms, maintaining sufficient user engagement with these apps remains challenging. This is often caused by a mismatch between the needs and preferences of adolescents and what the apps offer. Therefore, we need a better understanding of (i) adolescents’ needs and preferences and (ii) potential differences based on user characteristics. To this end, we qualitatively and quantitatively analyzed a dataset describing the user experience of 1312 Dutch adolescents (12–25 years) from the general population after they interacted for several weeks with a gamified mHealth app (the Grow It! app) that aims to promote momentary emotional awareness, reflection, and adaptive coping. A total of 4833 free-text survey responses spanning five user experience survey questions were analyzed using an inductive and iterative coding process, while accounting for intercoder reliability. We used (i) a thematic analysis to identify adolescents’ needs and preferences related to the app, and (ii) an exploratory quantitative analysis of the subthemes to investigate potential differences in which needs and preferences were mentioned by adolescents based on demographics. Through our thematic analysis, we identified three overarching themes related to the app’s design: usability , psychological impact , and meaningful interactive features . Furthermore, we identified two overarching themes that related to the adolescents’ motivation to use the app: intrinsic (de)motivators , and social–environmental factors impacting usage . Each of these themes consisted of four subthemes. Our exploratory statistical analysis shed light on several differences in how frequently these subthemes were mentioned based on age, sex, and educational level. By synthesizing our insights, we identify five design implications that can help tailor future mHealth apps to adolescents’ needs and preferences. These include concrete suggestions to personalize self-monitoring, include actionable insights, align content with personal needs, implement meaningful interactive features (e.g., competitions, gamification, and social communication), and make apps appealing to the entire target group. ...

How Search Engine Result Pages and AI-generated Podcasts Interact to Influence User Attitudes on Controversial Topics

Conference paper (2026) - Junjie Wang, Gaole He, Alisa Rieger, Ujwal Gadiraju
Compared to search engine result pages (SERPs), AI-generated podcasts represent a relatively new and relatively more passive modality of information consumption, delivering narratives in a naturally engaging format. As these two media increasingly converge in everyday information-seeking behavior, it is essential to explore how their interaction influences user attitudes, particularly in contexts involving controversial, value-laden, and often debated topics. Addressing this need, we aim to understand how information mediums of present-day SERPs and AI-generated podcasts interact to shape the opinions of users. To this end, through a controlled user study (N = 483), we investigated user attitudinal effects of consuming information via SERPs and AI-generated podcasts, focusing on how the sequence and modality of exposure shape user opinions. A majority of users in our study corresponded to attitude change outcomes, and we found an effect of sequence on attitude change. Our results further revealed a role of viewpoint bias and the degree of topic controversiality in shaping attitude change, although we found no effect of individual moderators. ...

Towards Intelligent Integration of Gestures As an Input Modality for Microtask Crowdsourcing

Conference paper (2025) - Garrett Allen, Ujwal Gadiraju
Human input is pivotal in building AI systems. Aiding the gathering of high-quality and representative human input on demand, microtask crowdsourcing platforms have thrived. Despite the benefits available, the lack of health provisions, safeguards, and existing practices threaten the sustainability of crowd work. Prior work investigated the usefulness of a dual-purpose input modality of ergonomically-informed gestures across different microtasks, finding that gestures as inputs offer a realistic trade-off between worker accuracy and potential short to long-term health benefits. However, little is understood about the effect of switching input modalities from one task to another on worker experiences and task-related outcomes. Addressing this research and empirical gap, we conducted a between-subjects study (N = 717) with varying sequences of input modalities across 16 experimental conditions to systematically understand the effect of switching input modalities. We found that the order of the input modality can influence the time it takes to complete tasks but does not affect accuracy. Further, the cognitive load perceived by workers was not significantly different between conditions. Our findings hint that ergonomically informed gestures can be effectively intertwined with conventional input modalities without a detrimental impact on worker experiences and quality-related outcomes. Our work has important implications for the design of human-centered crowdsourcing platforms that cater to worker health and wellbeing. ...
Book chapter (2025) - Ujwal Gadiraju, Agathe Balayn
The exponential advances in generative AI and agentic technologies have presented organizations with an unprecedented opportunity to seek and attain competitive advantages. Some aim to do so through product and service innovation, while others have begun to pursue operational efficiency and productivity. There is a pivotal role that human-AI collaboration can play in enterprise artificial intelligence design, development, and deployment today. In this chapter, we synthesize developments in human-AI collaboration over the last decade and explore how organizations can design AI systems that augment rather than replace human capabilities to achieve optimal experiential and performance-related outcomes for different stakeholders. Drawing from empirical work across different domains, we analyze the key factors that shape trust and reliance and determine successful human-AI collaboration. This includes various human factors, task factors, and AI system factors and the complex interplay between them in different configurations. This synthesis reflects the importance of broadening the spectrum of metrics for evaluating human-AI collaboration (e.g., by considering stakeholder values). We discuss promising ideas to address common challenges such as under-reliance or over-reliance on AI systems. We argue for broadening the lens of human-AI collaboration to consider the AI supply chain and the underlying value chains to ensure the responsible design of enterprise AI. Our insights collectively suggest that enterprise AI will benefit from the human-centered approach while creating collaborative workflows and human-AI configurations that can augment and complement humans with the computational power and scalability of AI. ...

An Online Conversational Survey for Understanding Worker Health in Crowdsourcing Platforms

Conference paper (2025) - Sihang Qiu, Ujwal Gadiraju, Xiaolong Zheng
Crowdsourcing marketplaces have gradually flourished over the last decade. With the growing landscape of online work in general, and the rise of paid microtask crowdsourcing in particular, the health and wellbeing of crowd workers has become an important concern. In this paper, we present an online conversational survey, named HealthInsights, for understanding the status quo of workers’ health-related background, physical health, mental health, and their needs. We carried out a study on two popular platforms - Mechanical Turk and Prolific. Results show that the survey has acceptable reliability and validity. We found that workers across these platforms reported similar health-related issues, but also exhibited certain differences. Based on our findings, we argue that crowdsourcing platforms, task requesters, and academic researchers need to take the collective responsibility of creating better work environments. Our work has important implications on task and workflow design that are centered around worker health on crowdsourcing platforms. ...
Conference paper (2025) - Ivica Kostric, Krisztian Balog, Ujwal Gadiraju
Conversational recommender systems (CRSs) provide users with an interactive means to express preferences and receive real-time personalized recommendations. The success of these systems is heavily influenced by the preference elicitation process. While existing research mainly focuses on what questions to ask during preference elicitation, there is a notable gap in understanding what role broader interaction patterns - including tone, pacing, and level of proactiveness - play in supporting users in completing a given task. This study investigates the impact of different conversational styles on preference elicitation, task performance, and user satisfaction with CRSs. We conducted a controlled experiment in the context of scientific literature recommendation, contrasting two distinct conversational styles - high involvement (fast-paced, direct, and proactive with frequent prompts) and high considerateness (polite and accommodating, prioritizing clarity and user comfort) - alongside a flexible experimental condition where users could switch between the two. Our results indicate that adapting conversational strategies based on user expertise and allowing flexibility between styles can enhance both user satisfaction and the effectiveness of recommendations in CRSs. Overall, our findings hold important implications for the design of future CRSs. ...

An Empirical Exploration to Foster Trustworthy LLM Production & Use

Conference paper (2025) - Agathe Balayn, Mireia Yurrita, Fanny Rancourt, Fabio Casati, Ujwal Gadiraju
Research on trust in AI is limited to several trustors (e.g., end-users) and trustees (especially AI systems), and empirical explorations remain in laboratory settings, overlooking factors that impact trust relations in the real world. Here, we broaden the scope of research by accounting for the supply chains that AI systems are part of. To this end, we present insights from an in-situ, empirical, study of LLM supply chains. We conducted interviews with 71 practitioners, and analyzed their (collaborative) practices using the lens of trust drawing from literature in organizational psychology. Our work reveals complex trust dynamics at the junctions of the chains, with interactions between diverse technical artifacts, individuals, or organizations. These junctions might constitute terrain for uncalibrated reliance when trustors lack supply chain knowledge or power dynamics are at play. Our findings bear implications for AI researchers and policymakers to promote AI governance that fosters calibrated trust. ...

Human-AI Decision Making With a Conversational XAI Assistant

Conference paper (2025) - Gaole He, Nilay Aishwarya, Ujwal Gadiraju
Explainable artificial intelligence (XAI) methods are being proposed to help interpret and understand how AI systems reach specific predictions. Inspired by prior work on conversational user interfaces, we argue that augmenting existing XAI methods with conversational user interfaces can increase user engagement and boost user understanding of the AI system. In this paper, we explored the impact of a conversational XAI interface on users’ understanding of the AI system, their trust, and reliance on the AI system. In comparison to an XAI dashboard, we found that the conversational XAI interface can bring about a better understanding of the AI system among users and higher user trust. However, users of both the XAI dashboard and conversational XAI interfaces showed clear over-reliance on the AI system. Enhanced conversations powered by large language model (LLM) agents amplified over-reliance. Based on our findings, we reason that the potential cause of such over-reliance is the illusion of explanatory depth that is concomitant with both XAI interfaces. Our findings have important implications for designing effective conversational XAI interfaces to facilitate appropriate reliance and improve human-AI collaboration. ...

An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant

Conference paper (2025) - Gaole He, Gianluca Demartini, Ujwal Gadiraju
Since the explosion in popularity of ChatGPT, large language models (LLMs) have continued to impact our everyday lives. Equipped with external tools that are designed for a specific purpose (e.g., for flight booking or an alarm clock), LLM agents exercise an increasing capability to assist humans in their daily work. Although LLM agents have shown a promising blueprint as daily assistants, there is a limited understanding of how they can provide daily assistance based on planning and sequential decision making capabilities. We draw inspiration from recent work that has highlighted the value of g'LLM-modulo' setups in conjunction with humans-in-the-loop for planning tasks. We conducted an empirical study (N = 248) of LLM agents as daily assistants in six commonly occurring tasks with different levels of risk typically associated with them (e.g., flight ticket booking and credit card payments). To ensure user agency and control over the LLM agent, we adopted LLM agents in a plan-then-execute manner, wherein the agents conducted step-wise planning and step-by-step execution in a simulation environment. We analyzed how user involvement at each stage affects their trust and collaborative team performance. Our findings demonstrate that LLM agents can be a double-edged sword - (1) they can work well when a high-quality plan and necessary user involvement in execution are available, and (2) users can easily mistrust the LLM agents with plans that seem plausible. We synthesized key insights for using LLM agents as daily assistants to calibrate user trust and achieve better overall task outcomes. Our work has important implications for the future design of daily assistants and human-AI collaboration with LLM agents. ...

Understanding the Effect of Decision-Makers' Configuration on Decision-Subjects' Fairness Perceptions

Human intervention is claimed to safeguard decision-subjects’ rights in algorithmic decision-making and contribute to their fairness perceptions. However, how decision-subjects perceive hybrid decision-maker configurations (i.e., combining humans and algorithms) is unclear. We address this gap through a mixed-methods study in an algorithmic policy enforcement context. Through qualitative interviews (Study 1; N1 = 21), we identify three characteristics (i.e., decision-maker’s profile, model type, input data provenance) that affect how decision-subjects perceive decision-makers’ ability, benevolence, and integrity (ABI). Through a quantitative study (Study 2; N2 = 223), we then systematically evaluate the individual and combined effects of these characteristics on decision-subjects’ perceptions towards decision-makers, and fairness perceptions. We found that only decision-maker’s profile contributes to perceived ability, benevolence, and integrity. Interestingly, the effect of decision-maker’s profile on fairness perceptions was mediated by perceived ability and integrity. Our findings have design implications for ensuring effective human intervention as a protection against harmful algorithmic decisions. ...
Contestability has been proposed as a key element in designing algorithmic decision-making processes that safeguard decision subjects' rights to dignity and autonomy. However, little is known about how contestability can be operationalized based on decision subjects' needs and preferences. We address this research gap by identifying decision subjects' information and procedural needs for enacting meaningful contestability. To this end, we chose an illegal holiday rental detection scenario as our case; a high-risk decision-making process in the public sector. We conducted 21 semi-structured interviews with citizens with experience renting their homes out and different levels of AI literacy. We found that decision subjects request interventions that facilitate (1) cooperation in sense-making, (2) support in contestation acts, and (3) appropriate responsibility attribution. Our results highlight the cooperative work behind contestability, and motivate future efforts to structure individual and collective action, to personalize explanations for contestability, and to open up sites of contestation in AI pipelines. ...