NY

N. Yorke-Smith

info

Please Note

66 records found

Regime-Switching Reinforcement Learning for Portfolio Allocation in Pairs Trading

Bachelor thesis (2026) - T.B. Ilieva, F. Yu, N. Yorke-Smith, F.A. Oliehoek
Pairs trading is a well-studied strategy in statistical arbitrage. By using asset pairs with correlated changes in their historical prices, the strategy profits from exploiting the non-permanent divergence of their price relationship, assuming that this relationship will revert to its long-term equilibrium. However, the dynamics of this relationship may vary over time, as the spread, which measures the deviation between the prices of paired assets, can exhibit different levels of volatility and mean-reverting behavior under different market conditions. In this paper, we propose a regime-aware reinforcement learning framework for portfolio optimization in pairs trading. We model the spread between assets and characterize its behavior using statistical features capturing its relative position to historical equilibrium, its volatility, and the strength of its mean-reverting behavior. These features are used within a Hidden Markov Model to infer latent market regimes, which represent distinct states of spread dynamics over time. The inferred regimes are incorporated into the state representation of a reinforcement learning agent, which learns to dynamically allocate capital across pairs. We evaluate the proposed approach against a regime-agnostic reinforcement learning benchmark and a classical z-score threshold strategy. In a controlled simulation study, the regime-aware agent achieves a mean Sharpe ratio of 1.354 versus 0.738 for the baseline on V/MA (ΔSharpe = +0.616) and 1.183 versus 0.564 on V/JKHY (ΔSharpe = +0.619), consistent across 10 training seeds. On real out-of-sample data from 2023 to 2026, the regime agent achieves Sharpe ratios of 0.567 and 0.609 on V/MA and V/JKHY respectively, outperforming the baseline in both cases. ...

A Framework for Experimentally Relevant Materials Discovery in Well-Understood Chemical Spaces

Current inverse materials discovery methods face a trade-off between broad exploration of chemical space and control over chemical validity, synthesisability, and target properties. Here, we present the COMPosition Aware Search Strategy (COMPASS), a constrained multi-objective, multi-fidelity search framework for crystalline composition spaces. COMPASS introduces a discrete mixed-site encoding for material families with fixed site stoichiometries and up to two species mixed on each crystallographic site. This encoding preserves chemical identity, allowing empirical chemical rules and property constraints to be evaluated directly during optimisation. COMPASS combines fast composition-only screening with a constrained genetic algorithm, structure-based verification using machine-learning interatomic potentials, and active learning to improve the low-fidelity model. Applied to mixed-site ABX3 perovskites, COMPASS identifies 15,922 computationally promising candidates satisfying chemical, novelty, stability, and band-gap criteria. In the same constrained discovery task, COMPASS achieves an approximately two-orders-of-magnitude higher yield of desired candidates than the tested open-source MatterGen baselines [Zeni et al., Nature, 2025, 639, 624--632]. These results position COMPASS as a framework for chemically well-understood discovery problems where chemical constraints can guide search through large composition spaces. ...
Bachelor thesis (2026) - R.K. Georgiev, F. Yu, F.A. Oliehoek, N. Yorke-Smith
Pairs trading has grown increasingly popular over the past several decades, and its application has extended into the domain of portfolio optimization. Reinforcement learning (RL) strategies, particularly Proximal Policy Optimization (PPO), have been used to address this problem. However, while substantial research exists for the single-pair case, a systematic investigation of RL models for portfolio optimization across multiple pairs simultaneously has been lacking. To address this gap, we develop and compare two PPO models that trade on several cointegrated pairs identified within the energy sector of the S&P 500. The two models differ in their information set: one is given explicit knowledge of the asset pairs it trades, while the other operates without this information, learning to allocate capital from price and portfolio data alone. We find that the pair-aware model achieves an annual return of 20.1% and a Sharpe ratio of 0.877, and maintains consistent performance across varying numbers of traded pairs, though no clear relationship emerges between the number of pairs traded and performance. These results suggest that the multi-pair approach to portfolio optimization is promising and highlight the need for further investigation. ...

Reinforcement Learning for Regime-Dependent Optimal Stopping in Pairs Trading

Bachelor thesis (2026) - M.O. Bankov, F. Yu, F.A. Oliehoek, N. Yorke-Smith
Pairs trading is a strategy that utilises the mean-reverting spread between two correlated assets (stocks). An important factor in such strategies is the market regime, which captures characteristics like trend and volatility of the data, and can shift over time. This paper investigates whether incorporating regime awareness improves the performance of Reinforcement Learning agents for pairs trading entry and exit decisions. Three Double Deep-Q network variants are implemented and compared: a baseline DQN (Deep-Q network), a Recurrent DQN, and a Hidden Markov Model-based DQN that maintains a separate agent per inferred regime. The agents are evaluated on intraday Corn and Wheat futures data, as well as on single-regime generated daily data. Results show that the Recurrent DQN does not significantly improve over the baseline, suggesting it does not implicitly capture regime information. The Markov DQN outperforms the baseline on real data (p = 0.033), while performing worse on generated data, though between-run variance is high in both cases. This supports the hypothesis that explicit regime modelling can benefit pairs trading on real-world data. ...

Evaluating the Robustness of a Graph Neural Network Agent under Uncertainty of Estimated Time of Arrival

The Dynamic Berth Allocation Problem is a port
scheduling problem where vessels arrive dynamically over time and must be assigned to a berth. A pre-trained Graph Neural Network (GNN) based
reinforcement learning approach solves the problem efficiently but is dependent on estimated times of arrival [11]. However, vessels predominantly arrive later than estimated. this introduces unwanted uncertainty. We evaluate a pre-trained GNN agent under controlled ETA perturbations to create an information gap between the estimated and actual arrival times, using the original full information setting as a reference. The agent is evaluated using a factorial experimental setup over different deviation levels and instances to evaluate robustness under ETA uncertainty. The results show that the agent is robust to optimistic ETA deviations, with only limited performance degradation even at
larger introduced deviations. The robustness appears to mainly stem from the reactive nature of the agent and its associated scheduling process, which
limits the influence of ETA deviations on the decision making process. ...
Bachelor thesis (2026) - T. Pagu, F. Yu, F.A. Oliehoek, N. Yorke-Smith
Pairs trading is a type of algorithmic trading strategy that exploits temporary diver-
gences between assets that tend to follow each other, which we describe as cointegrated. As a special case of statistical arbitrage, it has long been studied by both practitioners and academics. We hypothesize that a common failure of existing pairs trading strategies is their behavior when the cointegration relationship is not constant over long periods of time, which is often the case in practice. We show that given future knowledge of the cointegration relation, a strategy can yield dramatically better returns, up to 16% annualized during cointegration periods. This finding motivates a data-driven approach for estimating the cointegration regime. To solve this problem, we propose a GRU model that tracks the cointegration regime better than chance, though its reliability varies substantially by pair. We trained an RL model to exploit cointegration periods using synthetic data, and experimented with limiting trading to only periods of predicted cointegration. We tested this on three commonly used pairs and found it outperformed the risk-free rate, with Sharpe ratios of 0.30–0.65. Our work shows the potential of cointegration-aware approaches through an oracle analysis, proposes a way to approximate it in a realistic strategy, and identifies current limitations of the model. ...
Bachelor thesis (2026) - C. Petre-Luca, F. Yu, F.A. Oliehoek, N. Yorke-Smith
Pairs trading exploits the mean reversion of a cointegrated spread of two stocks, classically traded with fixed z-score rules. We recast it as a continuous portfolio-optimisation problem and train reinforcement-learning (PPO) agents to size the two legs and a risk-free asset, comparing a constrained agent forced into the market-neutral hedge against a free agent that weights the legs independently. Agents are trained on a generative market model calibrated to each real pair, and evaluated out of sample, with and without transaction costs, against the classical z-rule. The constrained agent learns the spread’s direction but does not beat the z-rule: it over-trades when no arbitrage is available instead of stepping aside. The free agent earns higher but far more variable returns, mixing directional market exposure with some genuine arbitrage. Transaction costs push both toward smoother, more conservative policies. We outline a potential improvement to the constrained agent to leverage its sizing capabilities in a future work. ...
Master thesis (2026) - M. Vossen, N. Yorke-Smith, D.P. Peck
Trade classifications change over time, which makes detailed trade data difficult to compare consistently across years. This is especially important for critical raw materials (CRM) analysis, where product-specific mappings are often needed to link trade codes to material relevance. Revisions to systems such as the Harmonized System (HS) and the Combined Nomenclature (CN) make those mappings harder to maintain over time. This thesis examines whether the Lukaszuk–Torun (LT) method can be credibly applied to CN data to convert trade flows between classification vintages and express them on a common basis. The results support LT-based CN harmonisation for adjacent-year conversions and indicate that it can be suitable for longitudinal analysis on a common classification basis, especially at aggregate or group level. Longer conversion chains are more approximate and require greater caution, particularly for narrow code trajectories. The contribution is primarily methodological: the thesis shows that higher-detail CN data can be made more usable over time while clarifying the limits of LT-based harmonisation for CRM-motivated research. ...
Master thesis (2025) - D. Hogendoorn, F. Schulte, N. Yorke-Smith, Y. Pang, Bart van Riessen
Reliable container-tracking depends on the quality of estimated time-of-arrival (ETA) data, yet existing logistics platforms offer little guidance on how trustworthy those timestamps really are. This thesis proposes a fit-for-use data-quality (DQ) framework for Digital Container Shipping Association (DCSA)-compliant event logs that flags ETA records likely to deviate from actual time of arrival (ATA) by more than one calendar day.

Event logs from $\sim$90\,k transport legs were preprocessed into records capturing origin-destination pair, carrier, publisher type, and timing information. Four supervised models, namely Linear Regression (LR), Random Forest, XGBoost, and a Neural Network, were trained to predict leg duration. A prediction that placed ATA \(>1\) day from the published ETA labeled that record \textit{low-quality}. Model outputs were evaluated with a precision-oriented \(\mathrm{F}_{\beta}\)-score, where a false alarm is 50 times more costly than a missed detection (\(\beta \approx 0.141\)).

The simplest model prevailed: standard LR achieved the highest overall \(\mathrm{F}_{0.141}\)-score (68.5 \%), balancing few false positives with robust recall, while more-complex tree-based and neural models produced excessive false alarms. When the analysis was narrowed to early-stage ETAs published by carriers (arguably the least reliable yet most operationally valuable subset) LR’s score rose to 72.0 \%. These findings highlight that careful feature engineering and data curation outweigh algorithmic complexity for this task.

The study delivers the first systematic, event-data-only method to quantify DQ in container tracking, enabling near-real-time plausibility checks without AIS feeds. Limitations include a three-month observation window and absence of exogenous factors such as weather or port congestion. Future work should extend the temporal scope, integrate AIS-derived and environmental features, and explore meta-learning techniques to adapt to disruptions. It could also use process-mining to uncover anomalous event sequences to take a different approach in dataquality assessment within container-eventlogs.

By demonstrating that a transparent LR baseline can reliably surface dubious ETAs, the thesis provides a practical blueprint for logistics platforms seeking to bolster trust in their tracking data and to prioritise corrective action where it matters most. ...
Master thesis (2025) - M.L. Le Blansch, Périne Cunat, N. Yorke-Smith, J.L. Cremer
This work proposes a new Modelling-to-Generate Alternatives (MGA) method for Energy System Optimisation Models (ESOMs) using a Genetic Algorithm (GA).
Instead of generating each alternative one by one, the GA aims to optimise for a diverse set of alternatives, meaning they cover the space of possible alternatives as evenly as possible.
Such a diverse set of alternatives has the potential to improve the decision-making process by accelerating the extraction of stakeholder requirements and finding more agreeable compromises.
Before designing the algorithm, we investigate what diversity metric is most suitable to optimise.
The components of the GA are designed to exploit useful properties of ESOMs to increase efficiency.
The performance of the GA is tested in terms of output quality and scalability for increasingly large ESOMs, showing promising performance in terms of output quality for a similar computational burden as state-of-the-art MGA methods.
A potential issue caused by the curse of dimensionality is formulated, requiring further investigation on its impact on the quality of the method's output.
We show the generated output of applying the proposed method to the European power system, which encourages further testing of the method on increasingly large ESOMs. ...
A display map category, originally just called a class of display maps with a stability condition, can be used to model dependent type theory. There are several other constructions on categories that can serve a similar purpose, such as comprehension categories. In fact, the similarity of such concept has been well-known, and there even have been comparisons made using bicategories of such categorical notions.
In this thesis, I aim to formalise one such comparison, and implement it in a proof assistant. In order to do this, I needed to formalise display map categories, some related concepts, to then construct their bicategory, and show the comparison as a pseudofunctor into the bicategory of comprehension categories. The formalisation has been done using Univalent Foundations, while the implementation has been completed using Rocq, and more specifically the UniMath library.
...
This thesis explores the automated construction of Chemical Reaction Networks (CRNs) from incomplete experimental data, a task traditionally dependent on expert knowledge and manual effort. CRNs model the interactions between chemical species through a network of reactions and are essential in fields such as medicine and chemistry. However, many real-world systems include unobserved or unmeasurable species, making CRN construction challenging. To address this, this thesis frames CRN discovery as a program synthesis problem, using grammars and constraints to define the space of possible CRNs. A modular synthesis pipeline is developed that incrementally builds candidate molecules, reactions, and networks given a problem definition. Experimental results demonstrate that constraints effectively reduce the search space and that the solver is capable of identifying the correct reaction networks. Moreover, a scoring mechanism ranks the expected CRN highly among generated candidates. ...
Probabilistic programming offers an intuitive and expressive way to define statistical models, rendering it particularly effective in modeling problems where uncertainty plays a crucial role.
As adoption increases and models become more expressive, the challenge of effective inference becomes increasingly pronounced.
Effective inference often requires tailoring algorithms to the structure of the underlying model. While many probabilistic programming systems allow users to implement custom inference strategies via programmable inference, this process remains largely manual and heavily reliant on domain-specific expertise, particularly for sampling-based methods.
This paper investigates the use of Satisfiability Modulo Theories (SMT) to automate the generation of tailored, observation-aware proposals for guiding inference within Metropolis-Hastings in Gen, a probabilistic programming system. By reformulating the search for high-likelihood traces as a constraint optimization problem, this work explores whether SMT-based solutions can improve proposal quality and convergence. Empirical results indicate that SMT-derived traces offer a promising starting point for inference but are less effective as an active search heuristic. These findings suggest a new direction for automated, structure-aware proposal generation in probabilistic programming. ...

CertERoute: A Framework for Routing and Charging Scheduling under Time and Energy Consumption Uncertainty

The shift towards zero-emission transport has driven rapid adoption of Heavy Goods Electric Vehicles (HGEVs) across Europe. This trend is also evident in the Netherlands, where their numbers have grown exponentially over the past six years. The integration of HGEVs into fleet operations introduces new challenges for fleet planning due to limited battery ranges and significant operational uncertainties in both the time and energy domains.

To address these challenges, this thesis introduces CertERoute, an Adaptive Robust Optimization (ARO) framework that allows for joint optimization of routing and charging scheduling under time and energy consumption uncertainty. It incorporates both depot and en-route charging while accounting for charger availability and other realistic operational constraints. Time-related uncertainties in service, waiting, travel, and charging durations are modelled using uncertainty sets, which require minimal assumptions about the underlying probability distributions. Energy consumption is modelled as a function of time-domain uncertainty, with environmental and vehicle-specific energy uncertainty factors accounted for through Monte Carlo Simulation (MCS). The adaptive design of the framework supports a two-stage decision process where routing and charger visits are planned in advance to hedge against worst-case scenarios, while charging amount and timing decisions remain flexible during route execution to avoid overly conservative and costly solutions.

To ensure computational tractability, the framework employs a Column and Constraint Generation (CCG) approximation method enhanced by a novel One-Step Look-ahead Pessimization (OSLP) algorithm, which selectively integrates only provably infeasible scenarios into the optimization problem. Despite its theoretical vulnerability to premature convergence, this algorithm empirically produces highly robust solutions. To improve scalability to larger instances, a multi-scenario Adaptive Large Neighbourhood Search (ALNS) metaheuristic is developed, integrating charging and timing decisions into neighbourhood generation to enable more informed solution exploration.

The framework is evaluated using representative yet synthetic European HGEV planning scenarios. The results highlight the scalability of the proposed solution methods and demonstrate that the resulting plans are highly robust under operational uncertainty. As a result, this work offers a practical foundation for integrated fleet and energy management systems and supports Shell eMobility’s strategic goal of enabling sustainable, cost-efficient transport through offering intelligent charging solutions. ...
Bachelor thesis (2025) - A. Moreno, Neil Yorke-Smith, Pascal van der Vaart
This paper investigates how Random Network Distillation (RND), coupled with Boltzmann exploration, influences exploration behaviour and learning dynamics in value-based agents such as Deep Q-Learning (DQN) across a range of environments, from classic control tasks to behaviour suite benchmarks and contextual bandits. The study addresses the sensitivity of RND to key hyperparameters, the impact of exploration strategy design, and the transferability of settings across tasks. The results reveal that RND remains benefitial within DQN in both sequential and non-sequential tasks, but requires careful tuning of reward scaling, temperature, and network capacity to be effective. No universal hyperparameter configuration generalizes across environments, and inappropriate tuning can lead to unstable learning or suboptimal outcomes. These findings provide practical insights into the strengths and limitations of applying RND within value-based reinforcement learning frameworks. ...
Program synthesis aims to solve problems through coding by removing the need to write the programs yourself.
Given the grammar and problem specification, it aims to find a program that adheres to your problem specification.
This is done by iterating over many failing programs until a solution that adheres to the problem specification is found.
Conflict analysis automatically takes these failing solutions and learns new constraints to make the search more efficient.
Unfortunately, many conflict analysis techniques are heavily specialized.
They are difficult to apply to diverse problems or when trying to use a different search algorithm.
In this work, we present a modular framework in which these techniques can be implemented in a generalized way and applied independently to different problems and solvers, while having the ability to share generated constraints and parsed information.
We identify two distinct conflict types and show how to use semantics in conflict analysis effectively.
The framework is evaluated on two diverse domains; it prunes up to 96\% of the search space, where combining techniques can further improve its average effectiveness.
Domain applicability for the techniques has to be considered for optimal framework performance.
Compared to an enumeration solver, the framework shows marginal improvements on a real-world benchmark.
While the framework's overhead needs improvement, its modularity allows for comparison between solvers, problems, and conflict analysis techniques. ...

Solving the Equality Problem with Realistic Noise

Bachelor thesis (2025) - T.G. Jacobs, T.B. Propp, S.D.C. Wehner, N. Yorke-Smith
Quantum computers allow us to solve certain problems that are unsolvable using classical computers. In this study we focus on solving the equality problem by simulating a three quantum computer network and using the communication complexity to determine if our theoretical quantum advantage is still there in practice. We want to know how the noise from realistic quantum networks that already exist affect this communication complexity. We found that we can beat the classical solution when simulating a laboratory setup in which the quantum computers are in close proximity to each other and when using only a small bit strings. However, when moving to setups in which there are kilometres between quantum computers instead of metres or when using larger bit strings as input to our problem we see that the noise becomes too much to simulate. ...
Denotational semantics of type theories provide a framework for understanding and reasoning about type theories and the behaviour of programs and proofs. In particular, it is important to study what can and can not be proved within Martin-Löf Type Theory (MLTT) as it is the basis of proof assistants like Agda, Lean and Coq. Many models, including a certain class of comprehension categories, full and split comprehension categories, have been studied for the semantics of dependent type theories. The motivation for this work comes from the fact that not all comprehension categories are full and split, and one expects that type theories more general than MLTT can be interpreted in a comprehension category which is not full and split.

In this thesis, we first study how MLTT is interpreted in full split comprehension categories through concrete examples. Next, we investigate type theories that can be interpreted in comprehension cat- egories which are not necessarily full and split. For this, we propose a candidate type theory for the internal language of comprehension categories by extracting a type theory from the semantics given by a general comprehension category which is not full and split. We also give an interpretation of this type theory in every comprehension category. ...

Bayesian Network-based Fault Detection and Diagnosis of an Air-handling Unit in a Dutch University Building

A significant part of worldwide energy consumption is used to provide a comfortable climate inside buildings. 
This energy is mostly used by heating, ventilation and air-conditioning (HVAC) systems. 
A part of this energy is currently wasted due to faults, such as incorrect control signals or component failures. 
In the field of HVAC fault detection and diagnosis (FDD), methods are developed that aim to detect that a fault is present in the system, and consequently diagnose which fault this is.

However, most of the methods that are currently being developed are data-driven, which require large amounts of labelled data. 
In practice, this data often is not available, leading to a lack of adoption of FDD methods by building operators. 
Additionally, most of the FDD approaches proposed in the literature have not been tested in real-time, instead validation has been performed on already existing datasets.

In this thesis, an FDD method is proposed to diagnose part of the HVAC system, air-handling units (AHUs), in real-time.
Air-handling units are an important part of the HVAC system, responsible for a significant part of the energy consumption. 
The method applies the four symptoms three faults (4S3F) framework, to construct a diagnostic Bayesian network (DBN) without relying on historical data for symptom detection or fault diagnosis. 

This DBN is implemented in Python to diagnose a case study AHU in a Dutch university building.
The method was validated with fault experiments, which were diagnosed with diagnosis periods ranging from one 10-minute sample to three hours. 
All of the faults that were introduced and included in the DBN were accurately diagnosed for all diagnosis periods. 
For the AHU operation data of 2022 and 2023, mostly control and temperature faults were diagnosed. 
Incorrect control of the heating coil valve was found in 25\% and incorrect control of the ERW in 11\% of the days considered.
Additionally, a deviation between the measured intended supply temperature of more than 1 $\deg$ was found for 79\% of the days in the dataset. 

Finally, the application of the method for real-time fault diagnosis has also been demonstrated, by collecting sensor data through the data streaming platform Apache Kafka. 
This data was consequently stored locally in the time-series database InfluxDB, which could be queried for performing the diagnosis.
The processing time of performing the diagnosis was approximately ten seconds, demonstrating the potential for performing diagnoses per sample.

Recommendations for future work include extending the DBN to diagnose faults related to the cooling coil and faults occurring outside of operation hours. 
Additionally, the developed diagnosis method should be applied to other AHUs, to further research the generalisability of the DBN.  

...

Summarising patient experiences for healthcare professionals

Summarising patient interactions creates a huge workload for the healthcare professionals. This research finds that patient interactions contain a lot of noise that is subjective of nature. To explore the problem area interviews with a summarisation prototype have been conducted to extract system requirements and validate those by implementing them in the prototype. Filtering noise, reinforcement learning and numerical factual correctness are of paramount importance to a successful summarisation system. ...