ZL

Z. LEI

info

Please Note

2 records found

Master thesis (2026) - Z. LEI, U.K. Gadiraju, S.K. Freire, E. Niforatos
Large language model (LLM) agents can plan, call tools, and incorporate intermediate results, but errors may arise at different workflow stages and propagate without being apparent in the final outcome. Human-visible plan-then-execute workflows create opportunities for oversight, yet process visibility alone does not tell users when intervention is warranted or what response remains feasible. This thesis investigates how risk-aware human oversight support can be designed and evaluated for such workflows. The proposed framework identifies PLAN, ACTION/pre-execution, and OUTPUT/tool-output as three types of critical oversight junctures. At each juncture, it keeps task-impact assessment separate from a potential-error assessment based on stage-specific runtime evidence. A rule-based mapping translates these inputs into an Oversight Cue that presents a suggested response and assessment rationales alongside the controls available at that stage. Risk-aware refers to this use of task impact in the framework design; it was not manipulated as an independent mechanism. The framework was evaluated in a between-subjects study with 61 participants completing four simulated daily assistant tasks. Both conditions used the same staged workflow, artifacts, checkpoints, and intervention controls. Condition B additionally received the complete Oversight Cue, whereas Condition A used unguided staged oversight. Condition B achieved higher appropriate reliance (mean difference = .107, 95% confidence interval (CI) [.051, .163], Hedges' g = .95), primarily through a higher correct intervention rate (mean difference = .205, 95% CI [.106, .304], g = 1.03). Correct acceptance did not differ significantly between conditions. Condition B also produced higher-quality plans (mean difference = .25, 95% CI [.11, .40], g = .86). The primary analysis favored Condition B for action-sequence accuracy, but this result did not remain significant after covariate adjustment and familywise correction. No statistically detectable condition differences were found for final-outcome accuracy or overall perceived workload. The findings show that structured oversight support can improve intervention on unacceptable agent behavior and the quality of an intermediate workflow artifact without producing corresponding improvements in every downstream outcome. Because the cue elements were presented as a bundle and the LLM-based assessments were not independently validated, the results apply to the integrated implementation rather than to the independent effectiveness of each component or the diagnostic accuracy of the assessments.
...
Bachelor thesis (2022) - Z. LEI, R.S. Verhagen, M.L. Tielman, A. Nadeem
Artificial intelligence systems assist humans in more and more cases. However, such systems' lack of explainability could lead to a bad performance of the teamwork, as humans might not cooperate with or trust systems with black-box algorithms opaque to them. This research attempts to improve the explainability of artificial intelligence systems by proposing a framework which models human workload in a value and tailors explanations to this value. Such explanations could provide agents' confidence, causes of making decisions and counterfactual parts to support their suggestions and are adjusted according to agents' knowledge of humans. Results show that adjusted explanations could improve participants' subjective trust in agents and make participants' take more suggestions, while no impact on collaboration fluency or teamwork performance is found. ...