JD

J.M. Duran

info

Please Note

7 records found

An evaluation framework constructed from expert knowledge, applied to FSRU operations

Master thesis (2026) - A.M.K. Brosens, M. Yang, P.H.A.J.M. van Gelder, J.M. Duran, F. Van Nuffel
Floating Storage and Regasification Units (FSRUs) play a critical role in energy security by providing flexible LNG import capacity. Operating these vessels efficiently requires daily decisions that balance multiple competing objectives and rely heavily on tacit expert knowledge that is difficult to articulate or codify. Recent advances in large language models (LLMs) offer opportunities to support this expert-driven decision-making. Before deployment, however, organisations need a structured method to evaluate whether LLM-generated recommendations are adequate. Existing evaluation approaches assess general language quality or assume that domain-specific performance criteria already exist. In expert-knowledge environments, they do not.

This research addresses that gap by developing an evaluation framework for LLM-based decision support in safety-critical operational environments and applying it to EXMAR's FSRU in Eemshaven. The research delivers both the framework and an operationalised evaluation tool for EXMAR, which serves as the framework's test case.

The framework consists of five phases. First, a suitable operational scenario is selected using five requirements: improvement potential, frequent occurrence, safety involvement, sufficient structure, and data availability. Second, tacit operator knowledge is elicited through Cognitive Task Analysis and translated into two assessment components. Decision alignment measures how closely LLM recommendations reproduce expert decisions using the F1-score and Spearman correlation. Decision quality evaluates recommendations against expert-derived criteria, utility functions, and weights that reflect operators' judgments of good decisions. These criteria are elicited from reflective judgment rather than observed decisions, which may be satisficed, enabling a Multi-Criteria Decision Making approach complemented by qualitative reasoning analysis. Third, a benchmark of current operator decision-making is constructed, as no objectively optimal decisions exist in these contexts. Fourth, recommendations are generated by varying prompt design. Fifth, recommendations are evaluated through factual correctness and safety gates before assessing decision quality and alignment, with results interpreted using a 2×2 matrix distinguishing adoptable, investigable, and rejectable recommendations.

The framework is applied to EXMAR's FSRU in Eemshaven using the regasification configuration decision. Interviews with three operators identified four criteria: energy efficiency, operational robustness, operational effort, and safety as a non-negotiable gate. These were operationalised into indicators, utility functions, and weights. Benchmark scores of 0.83 and 0.90 indicate high-quality operator decisions consistent with satisficing behaviour predicted by Naturalistic Decision Making research. Evaluation of sixteen prompt configurations shows that prompt design strongly influences recommendation quality, with structured prompts using explicit performance criteria achieving results comparable to or exceeding the benchmark. Qualitative reasoning identified six failure modes, particularly stopping at the first feasible option, over-fitting to criteria at the expense of safety, and overly conservative equipment loading. Addressing these informed a second prompt iteration, eliminating safety violations across all configurations. Validation across four additional scenarios demonstrated consistent results, and in one case an operator revised their decision after reviewing an LLM recommendation.

This research demonstrates that evaluating LLM-based decision support in safety-critical expert environments is feasible while highlighting its limitations. Some tacit knowledge cannot be fully captured, limiting any evaluation framework. Nevertheless, several prompt configurations matched or exceeded expert performance, suggesting that LLMs and operators are better viewed as complementary. Rather than replacing experts, these systems are most valuable in supporting operational decision-making while leaving the final decision to the operator. ...

A Governance Framework for the Royal Netherlands Marechaussee

Master thesis (2026) - N.A.M. ter Avest, M.E. Warnier, J.M. Duran

A proportionality based privacy framework for AML/CFT in the Netherlands

Master thesis (2025) - S. Uffing, J.M. Duran, Marcela Tuler de Oliveira, U. Pesch
This thesis investigates the interplay between privacy and safety within the context of Anti-Money Laundering and Countering the Financing of Terrorism (AML/CFT) practices in the Dutch banking sector. As financial institutions face increasing pressure to detect and report Financial Economic Crime (FEC), the demand for advanced surveillance techniques such as: Artificial Intelligence (AI)-driven monitoring, Public Private Partnerships (PPPs) and cross-bank data sharing, has grown. However, these innovations face barriers in their implementation due to concerns regarding financial privacy and data protection. By conducting a structural privacy assessment, this research identifies and categorizes the specific privacy harms that emerge from transaction monitoring. It analyses the tensions between key legal frameworks, including the General Data Protection Regulation (GDPR), the Dutch AML/CFT law (Wwft) and the recently introduced EU Anti Money Laundering Regulation (AMLR), highlighting the regulatory ambiguities and ethical dilemmas they present. Using expert interviews and a conceptual privacy framework grounded in academic theory, the study evaluates the proportionality of privacy and safety trade-offs. The key message of this thesis is that successful AML/CFT will remain politically and technically fragile until banks, regulators and developers adopt a structured, continuously-revised understanding of privacy harms and use that lens to decide which monitoring practices, data-sharing schemes and analytic tools are ethically and legally proportionate. The thesis therefore supplies both an analytic privacy framework tailored to transaction monitoring and a map of the legal, technical and governance tensions that must be resolved before developments such as AI, data-sharing and PPPs can be deployed responsibly. Keywords: Privacy, Anti-Money Laundering, Counter Terrorism Financing, GDPR, Data-sharing, PublicPrivate Partnership, Artificial Intelligence ...

Designing for Non-Domination

The effects of digital technologies on freedom and democracy have garnered increasing attention in recent years. Many have raised concerns about surveillance capitalism, technofeudalism, and general threats to constitutional democracies—with a special convergence on the worry that uncontrolled power of online platforms undermines people’s freedom. However, it remains unclear how ‘freedom’ should be understood, what the relation is between freedom and uncontrolled power, and to what extent these worries extend beyond online platforms. In this dissertation, I argue that these problems are best answered by appealing to a neo-republican account of freedom as non-domination, where ‘domination’ is understood as a condition of living under an agent’s uncontrolled power. In the context of AI systems used in core societal sectors such as healthcare, I show that domination of a system’s (in)direct end-users by the system’s developers occurs in at least three ways: (1) the distribution of decision-making power, (2) technical limitations of AI systems, and (3) underlying societal structures that empower developers and disempower end-users. To safeguard freedom in the digital age, I propose that AI development requires the explicit intention to 'design for non-domination'. This requires us to consider the broader societal contexts within which these systems operate, such as current regulatory initiatives and the political economy. ...
Master thesis (2020) - Arnoud Nederpel, Sabine Roeser, Juan Manuel Duran, Laurens Rook, Dirk van Roode
Artificial intelligence (AI) has the potential to revolutionize many industries across the world. However, artificial intelligence systems raise a series of ethical challenges that hamper their sustainable development. These challenges have often been discussed from a philosophical or theoretical perspective, but rarely from the more practical perspective of businesses. The lack of such a perspective for the ICT-sector in the Netherlands has formed a barrier to develop adequate ethical governance mechanisms. Specifically, the scarcity of practical data of ethical challenges and governance of business in the industry has made it difficult to establish the quality of the current code of ethics for the industry and the need for additional ethical governance measures. This research provides a descriptive study of ethical challenges and ethical governance of AI for ICT-businesses in the Netherlands. A form of mixed methodology (i.e., the exploratory sequential method) combines literature, interviews, and a questionnaire to acquire the relevant data. The results of these three methods are triangulated to identify room for improvement in the current code of ethics and identify the need for additional governance measures. Explainability, fairness, safety, and privacy were identified as the most pressing ethical challenges for businesses in the industry. This research also found that the first three of these challenges are currently insufficiently addressed in the code of ethics for the ICT-industry. Moreover, this study found a significantly low level of ethical governance of AI among small and medium-sized enterprises as compared to large companies. Future research should focus on further normative argumentation on principles for explainability, fairness, and safety to improve the code of ethics. Furthermore, this work lays the basis for researching and developing different concrete ethical governance measures of AI for the industry, especially for SMEs. ...
Master thesis (2020) - Aaron Khaleghi, Martijn Warnier, Hadi Asghari, Juan Manuel Duran, Benjamin Timmermans
The consumer lending domain has increasingly leveraged Artificial Intelligence (AI) to make loan approval processes more efficient and to make use of larger amount of information to predict their applicants’ repayment ability. Over time, however, valid concerns have been raised about whether decisions made about individuals using these data-driven technologies can lead to bias against women. In an attempt to assess the fairness of an algorithm, 21 prominent definitions of fairness have been proposed by the computer science community over the years. However, what remains absent is consensus on which definitions are suitable for assessing gender equality in consumer lending. There is also a lack of knowledge on how to appropriately implement these metrics in practice. To tackle the problems mentioned above, this research has investigated how automated loan approval processes can be assessed for gender equality. Two essential elements for assessing predictive tools were identified and investigated through a separate research question: What fairness metrics are suitable for assessing gender equality in consumer lending? How can the metrics be applied to observe gender bias in lending history data? Based on the questions above, the research was conducted in two stages: Stage 1 focused on analyzing the prominent definitions of group fairness, but before doing so, it conceptualizes gender equality in consumer lending by conducting an extensive literature review encompassing domains of philosophy, economics, gender studies, and history. In investigating the first research question, it is found that group fairness metrics are a measure of distributive justice. These metrics are based on three different statistical criteria commonly known as independence, sufficiency, and separation and each underlie different moral assumptions which should be verified based on the application scenario at hand. In Stage 2 of this work, the second research question was investigated by conducting an exploratory case study in which a logistic regression model is built to classify a sample of loan applicants in an open source dataset. Both the dataset and the model are then tested for bias using IBM’s open source Python AIF36 toolkit. After applying the group fairness metrics, it was found that the choice of separation and sufficiency can have different repercussions for each demographic group in the dataset. When false distribution of utility is under inspection, sufficiency advantaged male applicants more than the female applicants while separation advantages males more than females. Such inconsistency highlights the importance of realizing how relevant distribution of harm/benefit depends on the choice of fairness criteria made by decision makers. Lastly, the research provides an extensive discussion on possible root causes of bias and some recommendations to managers and data stewards on how to tackle bias issues that stakeholders may face in the context of consumer lending. ...
Master thesis (2019) - Juan Ruiz Reina, Ibo van de Poel, Geerten van de Kaa, Juan Manuel Duran, Andrea Franco
The introduction of artificial intelligence (AI) technologies in healthcare is expected to set a paradigm shift to medical practice because these systems will have a significant role in applications such as diagnosis-support and image analysis. However, this implementation does not come without risks. There are important ethical concerns that should be addressed beforehand to ensure public trust and acceptability. Privacy, safety, transparency, reliability and potential biases are some of the issues to consider. Responsible Research and Innovation (RRI) frameworks have been designed by academics to tackle this sort of problems but there is no application of these frameworks in the field of AI in healthcare. This problem is even more salient in the private sector, due to the unawareness of the RRI concept in industry. Consequently, the research objective of this project was to offer recommendations on how to implement RRI practices to avoid potential risks and improve the social acceptability of AI. For this, we studied the case of Philips and carried out interviews with the company’s experts in AI and corporate social responsibility (CSR). This information was complemented with a comprehensive study of the literature on topics related to AI in healthcare and RRI. The results from these activities were used to create a roadmap to introduce RRI practices in the AI innovation activities within Philips. The results showed that large companies should build upon their existing CSR practices to develop RRI. This will increase the acceptability of RRI within the research and development (R&D) teams. Based on that, we came up with 14 recommendations for the case of Philips. These actions range from current practices, such as continuing with the rigorous process of patient data selection and curation, to novel solutions such as including better interactive features in the design of telehealth platforms (i.e. virtual reality, video calls or social networks). Further research can be carried out in different companies to come up with common principles that contribute to the creation of a more comprehensive RRI framework for AI in healthcare. ...