J.M. Duran
Please Note
7 records found
1
Evaluating LLM-based decision support in safety-critical operational environments
An evaluation framework constructed from expert knowledge, applied to FSRU operations
This research addresses that gap by developing an evaluation framework for LLM-based decision support in safety-critical operational environments and applying it to EXMAR's FSRU in Eemshaven. The research delivers both the framework and an operationalised evaluation tool for EXMAR, which serves as the framework's test case.
The framework consists of five phases. First, a suitable operational scenario is selected using five requirements: improvement potential, frequent occurrence, safety involvement, sufficient structure, and data availability. Second, tacit operator knowledge is elicited through Cognitive Task Analysis and translated into two assessment components. Decision alignment measures how closely LLM recommendations reproduce expert decisions using the F1-score and Spearman correlation. Decision quality evaluates recommendations against expert-derived criteria, utility functions, and weights that reflect operators' judgments of good decisions. These criteria are elicited from reflective judgment rather than observed decisions, which may be satisficed, enabling a Multi-Criteria Decision Making approach complemented by qualitative reasoning analysis. Third, a benchmark of current operator decision-making is constructed, as no objectively optimal decisions exist in these contexts. Fourth, recommendations are generated by varying prompt design. Fifth, recommendations are evaluated through factual correctness and safety gates before assessing decision quality and alignment, with results interpreted using a 2×2 matrix distinguishing adoptable, investigable, and rejectable recommendations.
The framework is applied to EXMAR's FSRU in Eemshaven using the regasification configuration decision. Interviews with three operators identified four criteria: energy efficiency, operational robustness, operational effort, and safety as a non-negotiable gate. These were operationalised into indicators, utility functions, and weights. Benchmark scores of 0.83 and 0.90 indicate high-quality operator decisions consistent with satisficing behaviour predicted by Naturalistic Decision Making research. Evaluation of sixteen prompt configurations shows that prompt design strongly influences recommendation quality, with structured prompts using explicit performance criteria achieving results comparable to or exceeding the benchmark. Qualitative reasoning identified six failure modes, particularly stopping at the first feasible option, over-fitting to criteria at the expense of safety, and overly conservative equipment loading. Addressing these informed a second prompt iteration, eliminating safety violations across all configurations. Validation across four additional scenarios demonstrated consistent results, and in one case an operator revised their decision after reviewing an LLM recommendation.
This research demonstrates that evaluating LLM-based decision support in safety-critical expert environments is feasible while highlighting its limitations. Some tacit knowledge cannot be fully captured, limiting any evaluation framework. Nevertheless, several prompt configurations matched or exceeded expert performance, suggesting that LLMs and operators are better viewed as complementary. Rather than replacing experts, these systems are most valuable in supporting operational decision-making while leaving the final decision to the operator. ...
This research addresses that gap by developing an evaluation framework for LLM-based decision support in safety-critical operational environments and applying it to EXMAR's FSRU in Eemshaven. The research delivers both the framework and an operationalised evaluation tool for EXMAR, which serves as the framework's test case.
The framework consists of five phases. First, a suitable operational scenario is selected using five requirements: improvement potential, frequent occurrence, safety involvement, sufficient structure, and data availability. Second, tacit operator knowledge is elicited through Cognitive Task Analysis and translated into two assessment components. Decision alignment measures how closely LLM recommendations reproduce expert decisions using the F1-score and Spearman correlation. Decision quality evaluates recommendations against expert-derived criteria, utility functions, and weights that reflect operators' judgments of good decisions. These criteria are elicited from reflective judgment rather than observed decisions, which may be satisficed, enabling a Multi-Criteria Decision Making approach complemented by qualitative reasoning analysis. Third, a benchmark of current operator decision-making is constructed, as no objectively optimal decisions exist in these contexts. Fourth, recommendations are generated by varying prompt design. Fifth, recommendations are evaluated through factual correctness and safety gates before assessing decision quality and alignment, with results interpreted using a 2×2 matrix distinguishing adoptable, investigable, and rejectable recommendations.
The framework is applied to EXMAR's FSRU in Eemshaven using the regasification configuration decision. Interviews with three operators identified four criteria: energy efficiency, operational robustness, operational effort, and safety as a non-negotiable gate. These were operationalised into indicators, utility functions, and weights. Benchmark scores of 0.83 and 0.90 indicate high-quality operator decisions consistent with satisficing behaviour predicted by Naturalistic Decision Making research. Evaluation of sixteen prompt configurations shows that prompt design strongly influences recommendation quality, with structured prompts using explicit performance criteria achieving results comparable to or exceeding the benchmark. Qualitative reasoning identified six failure modes, particularly stopping at the first feasible option, over-fitting to criteria at the expense of safety, and overly conservative equipment loading. Addressing these informed a second prompt iteration, eliminating safety violations across all configurations. Validation across four additional scenarios demonstrated consistent results, and in one case an operator revised their decision after reviewing an LLM recommendation.
This research demonstrates that evaluating LLM-based decision support in safety-critical expert environments is feasible while highlighting its limitations. Some tacit knowledge cannot be fully captured, limiting any evaluation framework. Nevertheless, several prompt configurations matched or exceeded expert performance, suggesting that LLMs and operators are better viewed as complementary. Rather than replacing experts, these systems are most valuable in supporting operational decision-making while leaving the final decision to the operator.
Responsible AI for Criminal Investigations
A Governance Framework for the Royal Netherlands Marechaussee
Between Privacy and Protection
A proportionality based privacy framework for AML/CFT in the Netherlands
Freedom in the Digital Age
Designing for Non-Domination
Ethical behavior with artificial intelligence in the ICT-industry in the Netherlands
Towards an improved code of ethics
Towards a Responsible Implementation of Artificial Intelligence in Healthcare
The case of Royal Philips