U.K. Gadiraju
Please Note
93 records found
1
An Analysis of Visualisation Techniques for High-Dimensional Pareto Frontiers
A Case Study in Investment Portfolio Optimisation
...
...
When AI Flatters Too Much
An Exploratory Study into Trust, Perceived Trustworthiness, and Opinion Formation on Simulated Users
Trust in Information in the Age of Generative AI
Using AI Personas to Evaluate Trustworthiness and Misinformation Detection
The results showed that AI-generated misinformation was not identified less accurately than human-generated misinformation. Source labeling did not significantly affect confidence in truthfulness judgments. Trustworthiness ratings were significantly influenced by both statement condition and label visibility. When source labels were hidden, AI-generated statements received higher trustworthiness ratings than human-generated statements. However, when the source labels were revealed, the trustworthiness ratings for AI-generated content were reduced, while human-made statements received higher trustworthiness scores. These findings suggest that knowledge of content origin influences the perceived trustworthiness. ...
The results showed that AI-generated misinformation was not identified less accurately than human-generated misinformation. Source labeling did not significantly affect confidence in truthfulness judgments. Trustworthiness ratings were significantly influenced by both statement condition and label visibility. When source labels were hidden, AI-generated statements received higher trustworthiness ratings than human-generated statements. However, when the source labels were revealed, the trustworthiness ratings for AI-generated content were reduced, while human-made statements received higher trustworthiness scores. These findings suggest that knowledge of content origin influences the perceived trustworthiness.
Can AI and Media Literacy Guidance Improve AI-Generated Content Detection?
An Intervention Study with Simulated Young Adults
The Wizard of Incentive
A Guiding Tool for the Design of Incentive Formulas in Crowdsourcing
To address this gap, a wizard tool was developed to guide requesters through the process of designing payment schemas for crowdsourcing tasks. A user study was conducted to investigate how structured guidance affects incentive design: first, by comparing designs created with and without the tool, and second, by examining whether the tool produces consistency in compensation decisions across different requesters. The study evaluated both the designs participants created and their feedback on the tool itself.
The analysis reveals three primary insights. First, the tool's primary strength lies in structuring the design process rather than fundamentally altering participants' compensation decisions. The extent to which structured guidance benefited participants depended significantly on their prior experience with crowdsourcing, suggesting that the tool's value is contingent on user expertise. Second, the tool produced convergence around a limited set of high-level design elements, though participants used varied implementation approaches within these patterns, such as specific bonus sums.
These findings indicate that the tool could serve a valuable function in documenting and contextualizing design rationales, capturing the constraints and considerations that shaped dataset creation decisions. However, realizing the tool's full potential as a design aid requires enhancements to customization options and user experience refinement. Despite these limitations, the tool shows promise as an educational resource for introducing beginners to crowdsourcing incentive design, offering a structured entry point into a complex domain.
...
To address this gap, a wizard tool was developed to guide requesters through the process of designing payment schemas for crowdsourcing tasks. A user study was conducted to investigate how structured guidance affects incentive design: first, by comparing designs created with and without the tool, and second, by examining whether the tool produces consistency in compensation decisions across different requesters. The study evaluated both the designs participants created and their feedback on the tool itself.
The analysis reveals three primary insights. First, the tool's primary strength lies in structuring the design process rather than fundamentally altering participants' compensation decisions. The extent to which structured guidance benefited participants depended significantly on their prior experience with crowdsourcing, suggesting that the tool's value is contingent on user expertise. Second, the tool produced convergence around a limited set of high-level design elements, though participants used varied implementation approaches within these patterns, such as specific bonus sums.
These findings indicate that the tool could serve a valuable function in documenting and contextualizing design rationales, capturing the constraints and considerations that shaped dataset creation decisions. However, realizing the tool's full potential as a design aid requires enhancements to customization options and user experience refinement. Despite these limitations, the tool shows promise as an educational resource for introducing beginners to crowdsourcing incentive design, offering a structured entry point into a complex domain.
Too Distracted to Think Straight?
How Does External Cognitive Load Affect Young Adults’ Ability to Evaluate AI-Generated Content?
Can AI Make a ”Thinking Partner” for Young Adults
Fostering Responsible Opinion Formation Among Young Adults in the Age of Generative AI
Effective Human Oversight of AI Systems
The Interplay of Experience, Information Design, and Intervention Options
To gauge the effectiveness of our LLM-enhanced programming error messages (PEMs), we evaluated the framework in a crowdsourced Prolific study with 103 participants. We measured objective outcomes such as fix rate, time to fix, and number of attempts to fix, while also capturing subjective perceptions of PEMs, including readability, cognitive load, and authoritativeness. Objectively, LLM-enhanced PEMs showed favorable trends but did not produce statistically significant improvements over the standard interpreter. Subjectively, novices and experts alike, rated the pragmatic messages as significantly more readable and helpful, lower in intrinsic and extraneous cognitive load, and considerably less authoritative. Contingent messages exceeded the baseline on average but did not consistently reach statistical significance across all of our measurements, which points to a need for tighter control of error message verbosity and granularity, particularly for beginners.
These results show that LLMs, especially small-sized ones, are already capable of delivering targeted text-rewriting interventions that improve the perceived quality of error feedback. Future work should validate the effects at larger scale and across languages, expand coverage of real-world error contexts, and pursue true adaptivity in which error message style and level of detail adjust dynamically to user skill and task state. ...
To gauge the effectiveness of our LLM-enhanced programming error messages (PEMs), we evaluated the framework in a crowdsourced Prolific study with 103 participants. We measured objective outcomes such as fix rate, time to fix, and number of attempts to fix, while also capturing subjective perceptions of PEMs, including readability, cognitive load, and authoritativeness. Objectively, LLM-enhanced PEMs showed favorable trends but did not produce statistically significant improvements over the standard interpreter. Subjectively, novices and experts alike, rated the pragmatic messages as significantly more readable and helpful, lower in intrinsic and extraneous cognitive load, and considerably less authoritative. Contingent messages exceeded the baseline on average but did not consistently reach statistical significance across all of our measurements, which points to a need for tighter control of error message verbosity and granularity, particularly for beginners.
These results show that LLMs, especially small-sized ones, are already capable of delivering targeted text-rewriting interventions that improve the perceived quality of error feedback. Future work should validate the effects at larger scale and across languages, expand coverage of real-world error contexts, and pursue true adaptivity in which error message style and level of detail adjust dynamically to user skill and task state.
Generating Expertise-Specific Explanations in Cricket Pose Estimation
Design, Implementation, and Evaluation of Adaptive XAI Feedback
...
Adapting Explainable AI methods for multi-target tasks
Addressing challenges and Inter-Keypoint dependencies in Cricket Pose Analysis
Talking Like a Human: How Conversational Anthropomorphism Affects Self-Disclosure to Mental Health Chatbots
An Experimental Study on Human-like Chatbot Design and Question Sensitivity in Mental Health Contexts
Do Privacy Policies Matter? Investigating Self-Disclosure in Mental Health Chatbots
A User Study on the Importance of Privacy and Question Sensitivity in Mental Health Chatbots
Designing Mental Health Chatbots
The Impact of Self-Disclosure Techniques on the User Disclosure
The chatbot using factual self-disclosure received the highest average scores for trust, comfort, and willingness to disclose. However, statistical tests (ANOVA) showed no significant differences between chatbot types on these measures, except for changes in willingness to disclose. Participants who interacted with the emotional chatbot were more likely to report a negative change in their willingness to share. This result was unexpected and suggests that emotional self-disclosure may reduce user openness during early interactions, possibly because it feels unnatural or too personal too soon.
These findings show that emotional expression is not always the best approach. Instead, it is important to match the chatbot’s disclosure style to the situation and the user's comfort level, especially in sensitive areas like mental health support. ...
The chatbot using factual self-disclosure received the highest average scores for trust, comfort, and willingness to disclose. However, statistical tests (ANOVA) showed no significant differences between chatbot types on these measures, except for changes in willingness to disclose. Participants who interacted with the emotional chatbot were more likely to report a negative change in their willingness to share. This result was unexpected and suggests that emotional self-disclosure may reduce user openness during early interactions, possibly because it feels unnatural or too personal too soon.
These findings show that emotional expression is not always the best approach. Instead, it is important to match the chatbot’s disclosure style to the situation and the user's comfort level, especially in sensitive areas like mental health support.