TD
T.A. Draws
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
2 records found
1
Effective Human Oversight of AI Systems
The Interplay of Experience, Information Design, and Intervention Options
Human oversight of artificial intelligence systems is increasingly mandated by regulation, yet empirical evidence on which interface design choices actually improve oversight quality remains scarce. This thesis investigates how two modifiable design factors, information signal granularity and intervention option range, affect the effectiveness and perceived workload of human overseers in an AI-assisted fraud detection task. A 3×3 between-subjects experiment was conducted online via Prolific (N = 144), in which participants reviewed 30 bank account applications flagged by a Gradient Boosting model under one of nine conditions, crossing three levels of information signal (plain case features, categorical risk level indicator, and continuous model confidence score) with three levels of intervention options (binary decision, decision with delegation, and decision with flagged delegation). Oversight effectiveness was operationalised as a composite score rewarding correct classifications and appropriate delegation decisions, and penalising both misclassifications and over-delegation. Perceived workload was measured using the NASA Task Load Index. Self-reported domain expertise was included as a covariate. Neither main effect reached the pre-registered Bonferroni-corrected significance threshold of α = 0.008, and no significant interaction was found on either dependent variable. A marginal effect of information signal granularity on oversight effectiveness was observed (F(2, 134) = 3.29, p = .040, η²p = .047), with a non-monotonic pattern in which the categorical risk level indicator outperformed both the plain information baseline and the continuous model confidence score. This reversal of the hypothesised ordering suggests that, for non-expert overseers, a well-designed categorical signal may be more actionable than a continuous probability score, as interpreting the latter requires complementary domain knowledge not uniformly present in a general population. The results are treated as preliminary, since the study was underpowered relative to the pre-registered target of N = 288, due to the planned expert-screened wave not being completed within the thesis timeline. Theoretical and practical implications for the design of human oversight interfaces are discussed, with particular attention to the relationship between signal granularity and overseer expertise.
...
Human oversight of artificial intelligence systems is increasingly mandated by regulation, yet empirical evidence on which interface design choices actually improve oversight quality remains scarce. This thesis investigates how two modifiable design factors, information signal granularity and intervention option range, affect the effectiveness and perceived workload of human overseers in an AI-assisted fraud detection task. A 3×3 between-subjects experiment was conducted online via Prolific (N = 144), in which participants reviewed 30 bank account applications flagged by a Gradient Boosting model under one of nine conditions, crossing three levels of information signal (plain case features, categorical risk level indicator, and continuous model confidence score) with three levels of intervention options (binary decision, decision with delegation, and decision with flagged delegation). Oversight effectiveness was operationalised as a composite score rewarding correct classifications and appropriate delegation decisions, and penalising both misclassifications and over-delegation. Perceived workload was measured using the NASA Task Load Index. Self-reported domain expertise was included as a covariate. Neither main effect reached the pre-registered Bonferroni-corrected significance threshold of α = 0.008, and no significant interaction was found on either dependent variable. A marginal effect of information signal granularity on oversight effectiveness was observed (F(2, 134) = 3.29, p = .040, η²p = .047), with a non-monotonic pattern in which the categorical risk level indicator outperformed both the plain information baseline and the continuous model confidence score. This reversal of the hypothesised ordering suggests that, for non-expert overseers, a well-designed categorical signal may be more actionable than a continuous probability score, as interpreting the latter requires complementary domain knowledge not uniformly present in a general population. The results are treated as preliminary, since the study was underpowered relative to the pre-registered target of N = 288, due to the planned expert-screened wave not being completed within the thesis timeline. Theoretical and practical implications for the design of human oversight interfaces are discussed, with particular attention to the relationship between signal granularity and overseer expertise.
Perspective Discovery in Controversial Debates
An exploration of unsupervised topic models
Since the introduction of the Web, online platforms have become a place to share opinions across various domains (e.g., social media platforms, discussion fora or webshops). Consequently, many researchers have seen a need to classify, summarise or categorise these large sets of unstructured user-generated content. A field related to this task is also known as opinion mining in which various applications have focused on sentiment analysis techniques to classify opinionated documents based on sentiment. More recent, researchers have focused on stance classification to classify opinionated documents based on stance in controversial debates. However, in the case of such controversial debates it would be equally interesting to know the underlying reasons behind a stance in order to truly understand a discussion. We can call these underlying reasons as perspectives. Few have focused on distilling such perspectives from text and in this research we aim to explore the use of an unsupervised model - called joint topic models - to perform the task of perspective discovery. We define perspective discovery on a controversial debate as the process of automatically finding and extracting a structured overview of perspectives from unstructured text. The aim is to quantify how well existing joint topic models can extract human understandable perspectives between and within stances for more fine-grained opinion mining on textual debates. To perform this evaluation we propose an evaluation setup with an extensive user study. This setup focuses on the topic model’s clustering ability of perspectives as well as the human understandability of the topic model’s output. Based on the results we may derive that topic models can discover some of the perspectives from text. Moreover, the results suggest that users are not influenced by their pre-existing stance when interpreting the output of topic models.
...
Since the introduction of the Web, online platforms have become a place to share opinions across various domains (e.g., social media platforms, discussion fora or webshops). Consequently, many researchers have seen a need to classify, summarise or categorise these large sets of unstructured user-generated content. A field related to this task is also known as opinion mining in which various applications have focused on sentiment analysis techniques to classify opinionated documents based on sentiment. More recent, researchers have focused on stance classification to classify opinionated documents based on stance in controversial debates. However, in the case of such controversial debates it would be equally interesting to know the underlying reasons behind a stance in order to truly understand a discussion. We can call these underlying reasons as perspectives. Few have focused on distilling such perspectives from text and in this research we aim to explore the use of an unsupervised model - called joint topic models - to perform the task of perspective discovery. We define perspective discovery on a controversial debate as the process of automatically finding and extracting a structured overview of perspectives from unstructured text. The aim is to quantify how well existing joint topic models can extract human understandable perspectives between and within stances for more fine-grained opinion mining on textual debates. To perform this evaluation we propose an evaluation setup with an extensive user study. This setup focuses on the topic model’s clustering ability of perspectives as well as the human understandability of the topic model’s output. Based on the results we may derive that topic models can discover some of the perspectives from text. Moreover, the results suggest that users are not influenced by their pre-existing stance when interpreting the output of topic models.