F. Fang
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
13 records found
1
Structural Validation of Neural Network Calibration for Stochastic Volatility Models
An Explainable AI Perspective
This thesis investigates neural network-based approaches for the inverse calibration of stochastic volatility models, with a focus on both predictive performance and structural interpretability. Calibration is formulated as an inverse problem, where model parameters are inferred from implied volatility surfaces, a setting that is inherently non-linear and potentially ill-conditioned. While neural networks offer a computationally efficient alternative to classical optimisation-based methods, existing work has primarily focused on predictive accuracy, with comparatively limited attention given to the interpretability of the learned inverse mappings. This lack of transparency raises concerns in financial applications, where understanding model behaviour is essential for validation and risk management. Recent work has begun to address this issue, but a systematic framework for validating the learned inverse mappings remains limited. To address this, we evaluate several neural network architectures, including the multilayer perceptron (MLP), Highway networks, and Generalised Highway networks, trained on synthetically generated implied volatility surfaces from the Heston and rough Heston models. We introduce a framework for the structural validation of learned calibration mappings using eXplainable Artificial Intelligence (XAI) methods, specifically SHAP and νSHAP, which capture complementary notions of feature relevance.
The results show that neural networks achieve highly accurate and fast calibration within the training domain, with the Generalised Highway architecture consistently outperforming its counterparts in terms of validation error. The XAI analysis confirms that the learned mappings rely on structurally meaningful regions of the implied volatility surface, most notably short maturities and the wings of the smile, aligning with financial intuition. At the same time, the comparison between SHAP and νSHAP highlights non-homogeneous parameter identifiability: while some parameters are associated with concentrated and stable regions of importance, others are inferred from diffuse and overlapping information, indicating redundancy in the inverse mapping.
In addition, we demonstrate that XAI can be used to inform model design. In particular, a νSHAP-guided reduction of the input grid yields comparable or improved predictive performance despite a substantial reduction in input dimensionality.
Overall, the findings show that neural network-based calibration can be both accurate and structurally interpretable when combined with appropriate analysis tools, while also revealing important limitations related to generalisation and parameter identifiability. ...
The results show that neural networks achieve highly accurate and fast calibration within the training domain, with the Generalised Highway architecture consistently outperforming its counterparts in terms of validation error. The XAI analysis confirms that the learned mappings rely on structurally meaningful regions of the implied volatility surface, most notably short maturities and the wings of the smile, aligning with financial intuition. At the same time, the comparison between SHAP and νSHAP highlights non-homogeneous parameter identifiability: while some parameters are associated with concentrated and stable regions of importance, others are inferred from diffuse and overlapping information, indicating redundancy in the inverse mapping.
In addition, we demonstrate that XAI can be used to inform model design. In particular, a νSHAP-guided reduction of the input grid yields comparable or improved predictive performance despite a substantial reduction in input dimensionality.
Overall, the findings show that neural network-based calibration can be both accurate and structurally interpretable when combined with appropriate analysis tools, while also revealing important limitations related to generalisation and parameter identifiability. ...
This thesis investigates neural network-based approaches for the inverse calibration of stochastic volatility models, with a focus on both predictive performance and structural interpretability. Calibration is formulated as an inverse problem, where model parameters are inferred from implied volatility surfaces, a setting that is inherently non-linear and potentially ill-conditioned. While neural networks offer a computationally efficient alternative to classical optimisation-based methods, existing work has primarily focused on predictive accuracy, with comparatively limited attention given to the interpretability of the learned inverse mappings. This lack of transparency raises concerns in financial applications, where understanding model behaviour is essential for validation and risk management. Recent work has begun to address this issue, but a systematic framework for validating the learned inverse mappings remains limited. To address this, we evaluate several neural network architectures, including the multilayer perceptron (MLP), Highway networks, and Generalised Highway networks, trained on synthetically generated implied volatility surfaces from the Heston and rough Heston models. We introduce a framework for the structural validation of learned calibration mappings using eXplainable Artificial Intelligence (XAI) methods, specifically SHAP and νSHAP, which capture complementary notions of feature relevance.
The results show that neural networks achieve highly accurate and fast calibration within the training domain, with the Generalised Highway architecture consistently outperforming its counterparts in terms of validation error. The XAI analysis confirms that the learned mappings rely on structurally meaningful regions of the implied volatility surface, most notably short maturities and the wings of the smile, aligning with financial intuition. At the same time, the comparison between SHAP and νSHAP highlights non-homogeneous parameter identifiability: while some parameters are associated with concentrated and stable regions of importance, others are inferred from diffuse and overlapping information, indicating redundancy in the inverse mapping.
In addition, we demonstrate that XAI can be used to inform model design. In particular, a νSHAP-guided reduction of the input grid yields comparable or improved predictive performance despite a substantial reduction in input dimensionality.
Overall, the findings show that neural network-based calibration can be both accurate and structurally interpretable when combined with appropriate analysis tools, while also revealing important limitations related to generalisation and parameter identifiability.
The results show that neural networks achieve highly accurate and fast calibration within the training domain, with the Generalised Highway architecture consistently outperforming its counterparts in terms of validation error. The XAI analysis confirms that the learned mappings rely on structurally meaningful regions of the implied volatility surface, most notably short maturities and the wings of the smile, aligning with financial intuition. At the same time, the comparison between SHAP and νSHAP highlights non-homogeneous parameter identifiability: while some parameters are associated with concentrated and stable regions of importance, others are inferred from diffuse and overlapping information, indicating redundancy in the inverse mapping.
In addition, we demonstrate that XAI can be used to inform model design. In particular, a νSHAP-guided reduction of the input grid yields comparable or improved predictive performance despite a substantial reduction in input dimensionality.
Overall, the findings show that neural network-based calibration can be both accurate and structurally interpretable when combined with appropriate analysis tools, while also revealing important limitations related to generalisation and parameter identifiability.
Master thesis
(2025)
-
M.H. Chang, A. Papapantoleon, M.A. Sharifi Kolarijani, F. Fang, Cees Krijgsman
Options markets offer traders powerful leverage and sophisticated hedging tools, yet their nonlinear
payoffs and high volatility could cause catastrophic losses when not traded carefully. On the other
hand, when managed properly, options can yield substantial profits, underscoring their high-risk, highreward nature. Despite the high stakes and growing automation in finance, applying Reinforcement
Learning (RL) to intraday options trading remains underexplored.
This thesis aims to determine whether RL can learn robust and profitable trading strategies for S&P
500 options and to identify which reward designs best balance return and risk. To do so, we extend
the FinRL framework to simulate a multi-option trading environment rich in state information, including prices, Greeks, and implied volatility. FinRL is an open-source Deep Reinforcement Learning
library specifically designed for financial applications, providing standardized data pipelines, market
simulators, and agent interfaces. We chose FinRL because it offers backtesting capabilities, solid environment designs, and seamless integration with popular RL algorithms, which allowed us to focus
on customizing the option-specific state and reward structures rather than building infrastructure from
scratch.
We trained the Proximal Policy Optimization (PPO) agents under eight distinct reward formulations,
including (un)realized PnL, margin penalties, and normalized profit to margin ratios. Evaluation on out
of sample data shows that many agents fail to generalize the returns it has achieved during training.
Margin penalties enforce safety at the expense of profitability, and normalized rewards improve agent
behaviour slightly but suffer from unstable learning. A combined reward that integrates realized profit
and margin penalties achieves the best balance, producing small positive test returns with intended
margin control.
Shortcomings include sensitivity to hyperparameter choices, challenges in reward term scaling, choice
of data granularity in the simulated environment, limited state-action representation, and the inherent noisy data of the financial markets. These findings underscore the role of reward engineering in
learning viable options trading policies. Future work should focus on hyperparameters tuning, state
and action space engineering and alternative RL algorithms with diverse reward functions to further
enhance robustness and real-world applicability.
Overall, this thesis contributes to bridging the gap between RL and options trading by demonstrating
the role of reward engineering in shaping agent behavior.
...
Options markets offer traders powerful leverage and sophisticated hedging tools, yet their nonlinear
payoffs and high volatility could cause catastrophic losses when not traded carefully. On the other
hand, when managed properly, options can yield substantial profits, underscoring their high-risk, highreward nature. Despite the high stakes and growing automation in finance, applying Reinforcement
Learning (RL) to intraday options trading remains underexplored.
This thesis aims to determine whether RL can learn robust and profitable trading strategies for S&P
500 options and to identify which reward designs best balance return and risk. To do so, we extend
the FinRL framework to simulate a multi-option trading environment rich in state information, including prices, Greeks, and implied volatility. FinRL is an open-source Deep Reinforcement Learning
library specifically designed for financial applications, providing standardized data pipelines, market
simulators, and agent interfaces. We chose FinRL because it offers backtesting capabilities, solid environment designs, and seamless integration with popular RL algorithms, which allowed us to focus
on customizing the option-specific state and reward structures rather than building infrastructure from
scratch.
We trained the Proximal Policy Optimization (PPO) agents under eight distinct reward formulations,
including (un)realized PnL, margin penalties, and normalized profit to margin ratios. Evaluation on out
of sample data shows that many agents fail to generalize the returns it has achieved during training.
Margin penalties enforce safety at the expense of profitability, and normalized rewards improve agent
behaviour slightly but suffer from unstable learning. A combined reward that integrates realized profit
and margin penalties achieves the best balance, producing small positive test returns with intended
margin control.
Shortcomings include sensitivity to hyperparameter choices, challenges in reward term scaling, choice
of data granularity in the simulated environment, limited state-action representation, and the inherent noisy data of the financial markets. These findings underscore the role of reward engineering in
learning viable options trading policies. Future work should focus on hyperparameters tuning, state
and action space engineering and alternative RL algorithms with diverse reward functions to further
enhance robustness and real-world applicability.
Overall, this thesis contributes to bridging the gap between RL and options trading by demonstrating
the role of reward engineering in shaping agent behavior.
Can Context-Aware Incremental Nets Outperform GBDTs Over Time?
A Tabular Lifelong-Learning Study
Master thesis
(2025)
-
F.M. Gunnarsson, M.S. Pera, R. Hai, Y. Wang, F. Fang, J. Roeder, S. van Haren
Modern machine learning systems face unprecedented challenges in processing continuously arriving data streams while maintaining both computational efficiency and privacy compliance. Traditional batch learning approaches exhibit quadratic scaling in memory and computational requirements, making them unsuitable for long-term deployment in resource-constrained environments. Despite significant advances in continual learning for computer vision and natural language processing, tabular data represents the majority of industrial machine learning applications.
This thesis introduces IMLP (Incremental MLP), an attention-based architecture for energy-efficient continual learning on tabular data streams. IMLP augments a standard multilayer perceptron with attention-based feature rehearsal, maintaining a fixed-size buffer of learned 256-dimensional representations rather than raw historical samples. This design achieves constant computational complexity regardless of stream length while preserving task-relevant knowledge without storing personally identifiable information.
We conduct comprehensive evaluation across 36 diverse TabZilla classification tasks against 14 baseline methods spanning gradient boosting, classical machine learning, and neural architectures. Using calibrated power measurement equipment and rigorous statistical analysis via Friedman omnibus tests with post-hoc comparisons, we establish that IMLP achieves a $4.2\times$ median speedup and 79.6\% energy reduction compared to standard MLPs while maintaining competitive accuracy (80.6\% vs 82.9\% balanced accuracy).
Our key findings demonstrate that IMLP successfully trades a modest 2.3 percentage point accuracy reduction for substantial efficiency gains, achieving 97.5\% of cumulative learning performance using only current segment data. The approach proves robust across datasets spanning 5 to 2,000 features and diverse domains including medical diagnosis, sensor data, and financial applications. Moreover, we introduce NetScore-T, a composite metric for evaluating accuracy-efficiency trade-offs, positioning IMLP optimally on the neural network Pareto frontier.
Therefore, this work establishes the feasibility of practical continual learning for resource-constrained environments while contributing the first systematic study of energy consumption in neural continual learning for tabular data, enabling deployment scenarios previously considered computationally infeasible.
...
This thesis introduces IMLP (Incremental MLP), an attention-based architecture for energy-efficient continual learning on tabular data streams. IMLP augments a standard multilayer perceptron with attention-based feature rehearsal, maintaining a fixed-size buffer of learned 256-dimensional representations rather than raw historical samples. This design achieves constant computational complexity regardless of stream length while preserving task-relevant knowledge without storing personally identifiable information.
We conduct comprehensive evaluation across 36 diverse TabZilla classification tasks against 14 baseline methods spanning gradient boosting, classical machine learning, and neural architectures. Using calibrated power measurement equipment and rigorous statistical analysis via Friedman omnibus tests with post-hoc comparisons, we establish that IMLP achieves a $4.2\times$ median speedup and 79.6\% energy reduction compared to standard MLPs while maintaining competitive accuracy (80.6\% vs 82.9\% balanced accuracy).
Our key findings demonstrate that IMLP successfully trades a modest 2.3 percentage point accuracy reduction for substantial efficiency gains, achieving 97.5\% of cumulative learning performance using only current segment data. The approach proves robust across datasets spanning 5 to 2,000 features and diverse domains including medical diagnosis, sensor data, and financial applications. Moreover, we introduce NetScore-T, a composite metric for evaluating accuracy-efficiency trade-offs, positioning IMLP optimally on the neural network Pareto frontier.
Therefore, this work establishes the feasibility of practical continual learning for resource-constrained environments while contributing the first systematic study of energy consumption in neural continual learning for tabular data, enabling deployment scenarios previously considered computationally infeasible.
...
Modern machine learning systems face unprecedented challenges in processing continuously arriving data streams while maintaining both computational efficiency and privacy compliance. Traditional batch learning approaches exhibit quadratic scaling in memory and computational requirements, making them unsuitable for long-term deployment in resource-constrained environments. Despite significant advances in continual learning for computer vision and natural language processing, tabular data represents the majority of industrial machine learning applications.
This thesis introduces IMLP (Incremental MLP), an attention-based architecture for energy-efficient continual learning on tabular data streams. IMLP augments a standard multilayer perceptron with attention-based feature rehearsal, maintaining a fixed-size buffer of learned 256-dimensional representations rather than raw historical samples. This design achieves constant computational complexity regardless of stream length while preserving task-relevant knowledge without storing personally identifiable information.
We conduct comprehensive evaluation across 36 diverse TabZilla classification tasks against 14 baseline methods spanning gradient boosting, classical machine learning, and neural architectures. Using calibrated power measurement equipment and rigorous statistical analysis via Friedman omnibus tests with post-hoc comparisons, we establish that IMLP achieves a $4.2\times$ median speedup and 79.6\% energy reduction compared to standard MLPs while maintaining competitive accuracy (80.6\% vs 82.9\% balanced accuracy).
Our key findings demonstrate that IMLP successfully trades a modest 2.3 percentage point accuracy reduction for substantial efficiency gains, achieving 97.5\% of cumulative learning performance using only current segment data. The approach proves robust across datasets spanning 5 to 2,000 features and diverse domains including medical diagnosis, sensor data, and financial applications. Moreover, we introduce NetScore-T, a composite metric for evaluating accuracy-efficiency trade-offs, positioning IMLP optimally on the neural network Pareto frontier.
Therefore, this work establishes the feasibility of practical continual learning for resource-constrained environments while contributing the first systematic study of energy consumption in neural continual learning for tabular data, enabling deployment scenarios previously considered computationally infeasible.
This thesis introduces IMLP (Incremental MLP), an attention-based architecture for energy-efficient continual learning on tabular data streams. IMLP augments a standard multilayer perceptron with attention-based feature rehearsal, maintaining a fixed-size buffer of learned 256-dimensional representations rather than raw historical samples. This design achieves constant computational complexity regardless of stream length while preserving task-relevant knowledge without storing personally identifiable information.
We conduct comprehensive evaluation across 36 diverse TabZilla classification tasks against 14 baseline methods spanning gradient boosting, classical machine learning, and neural architectures. Using calibrated power measurement equipment and rigorous statistical analysis via Friedman omnibus tests with post-hoc comparisons, we establish that IMLP achieves a $4.2\times$ median speedup and 79.6\% energy reduction compared to standard MLPs while maintaining competitive accuracy (80.6\% vs 82.9\% balanced accuracy).
Our key findings demonstrate that IMLP successfully trades a modest 2.3 percentage point accuracy reduction for substantial efficiency gains, achieving 97.5\% of cumulative learning performance using only current segment data. The approach proves robust across datasets spanning 5 to 2,000 features and diverse domains including medical diagnosis, sensor data, and financial applications. Moreover, we introduce NetScore-T, a composite metric for evaluating accuracy-efficiency trade-offs, positioning IMLP optimally on the neural network Pareto frontier.
Therefore, this work establishes the feasibility of practical continual learning for resource-constrained environments while contributing the first systematic study of energy consumption in neural continual learning for tabular data, enabling deployment scenarios previously considered computationally infeasible.
A vital aspect of managing inflation risk is the use of inflation-indexed derivatives. Currently, inflation-indexed bonds and swaps are the primary instruments purchased by institutions. Inflation options (also known as inflation caps/floors) are also available in the market. Risk-neutral pricing of these derivatives is a difficult challenge due to the connection between inflation and interest rates.
In this thesis, the Heston model and its extensions to stochastic interest rates are investigated in the context of inflation-indexed derivatives. First, existing analytical pricing formulas and simulation methods are summarized. Then the multilevel Monte Carlo (MLMC) method is applied as a potent variance reduction technique. For the standard Heston model, the MLMC method reduces the computation costs by a factor of 10 to 50 for short maturities. The Python code implementing the applied methods is also published. ...
In this thesis, the Heston model and its extensions to stochastic interest rates are investigated in the context of inflation-indexed derivatives. First, existing analytical pricing formulas and simulation methods are summarized. Then the multilevel Monte Carlo (MLMC) method is applied as a potent variance reduction technique. For the standard Heston model, the MLMC method reduces the computation costs by a factor of 10 to 50 for short maturities. The Python code implementing the applied methods is also published. ...
A vital aspect of managing inflation risk is the use of inflation-indexed derivatives. Currently, inflation-indexed bonds and swaps are the primary instruments purchased by institutions. Inflation options (also known as inflation caps/floors) are also available in the market. Risk-neutral pricing of these derivatives is a difficult challenge due to the connection between inflation and interest rates.
In this thesis, the Heston model and its extensions to stochastic interest rates are investigated in the context of inflation-indexed derivatives. First, existing analytical pricing formulas and simulation methods are summarized. Then the multilevel Monte Carlo (MLMC) method is applied as a potent variance reduction technique. For the standard Heston model, the MLMC method reduces the computation costs by a factor of 10 to 50 for short maturities. The Python code implementing the applied methods is also published.
In this thesis, the Heston model and its extensions to stochastic interest rates are investigated in the context of inflation-indexed derivatives. First, existing analytical pricing formulas and simulation methods are summarized. Then the multilevel Monte Carlo (MLMC) method is applied as a potent variance reduction technique. For the standard Heston model, the MLMC method reduces the computation costs by a factor of 10 to 50 for short maturities. The Python code implementing the applied methods is also published.
This research project, conducted in collaboration between TU Delft and MN, a pension fund asset manager, focuses on the optimal venue selection in FX trading. The objective is to investigate how the venue selection affects trading performance and to improve MN trading execution algorithm, named ALGO. The research aims to propose a new approach for the venue selection problem by allocating weights to different venues instead of solely selecting the best one. It utilizes advanced statistical and machine learning techniques and develops a matching engine capable of reconstructing historical orderbooks for backtesting strategies. The outcomes of this thesis show clear ideas for improving the current venue selection model. The proposed models are anticipated to consistently outperform ALGO, leading to improved trade execution. The insights and methodologies developed in this research will contribute to further investigations in improving venue selection processes and optimizing execution strategies in the FX market. The thesis provides a comprehensive analysis of the problem, explores the mathematical framework, presents real-world data-driven approaches, and discusses the findings and conclusions, offering valuable insights and recommendations for future MN research.
...
This research project, conducted in collaboration between TU Delft and MN, a pension fund asset manager, focuses on the optimal venue selection in FX trading. The objective is to investigate how the venue selection affects trading performance and to improve MN trading execution algorithm, named ALGO. The research aims to propose a new approach for the venue selection problem by allocating weights to different venues instead of solely selecting the best one. It utilizes advanced statistical and machine learning techniques and develops a matching engine capable of reconstructing historical orderbooks for backtesting strategies. The outcomes of this thesis show clear ideas for improving the current venue selection model. The proposed models are anticipated to consistently outperform ALGO, leading to improved trade execution. The insights and methodologies developed in this research will contribute to further investigations in improving venue selection processes and optimizing execution strategies in the FX market. The thesis provides a comprehensive analysis of the problem, explores the mathematical framework, presents real-world data-driven approaches, and discusses the findings and conclusions, offering valuable insights and recommendations for future MN research.
In this research a new method for pricing continuous Arithmetic averaged Asian options is proposed. The computation is based on Fourier-cosine expansion, namely the COS method. Therefore, we derive the characteristic function of Integrated Geometric Brownian Motion based on Bougerol's identity.
Extensive numerical error analysis on the CDF recovery of IGBM and the option prices is performed. Via numerical tests, the convergence of errors using our new method has been proved. We are able to price continuous Arithmetic averaged Asian options with a minimal error of order 10-2, and a maximum precision of order 10-5 within seconds.
...
Extensive numerical error analysis on the CDF recovery of IGBM and the option prices is performed. Via numerical tests, the convergence of errors using our new method has been proved. We are able to price continuous Arithmetic averaged Asian options with a minimal error of order 10-2, and a maximum precision of order 10-5 within seconds.
...
In this research a new method for pricing continuous Arithmetic averaged Asian options is proposed. The computation is based on Fourier-cosine expansion, namely the COS method. Therefore, we derive the characteristic function of Integrated Geometric Brownian Motion based on Bougerol's identity.
Extensive numerical error analysis on the CDF recovery of IGBM and the option prices is performed. Via numerical tests, the convergence of errors using our new method has been proved. We are able to price continuous Arithmetic averaged Asian options with a minimal error of order 10-2, and a maximum precision of order 10-5 within seconds.
Extensive numerical error analysis on the CDF recovery of IGBM and the option prices is performed. Via numerical tests, the convergence of errors using our new method has been proved. We are able to price continuous Arithmetic averaged Asian options with a minimal error of order 10-2, and a maximum precision of order 10-5 within seconds.
This thesis investigates the estimation of option-implied probability density functions for inflation using inflation options, focusing not only on the expected value but the whole distribution. The aim is to identify the most effective method for measuring the market expectation of future inflation. The research explores both parametric and non-parametric approaches for deriving these density functions from inflation option prices. Methodologies include parametric models such as expansion, generalised distribution, and mixture methods, alongside non-parametric techniques using Breeden and Litzenberger’s result, such as curve-fitting and kernel methods. Implementing these methods involved analysing inflation option data sourced from the BVOL Bloomberg database, specifically for Harmonised Index of Consumer Prices excluding Tobacco (HICPxT) options from January 1, 2013. The study employed Shimko’s method, various spline methods, the Delta method, and Kernel method, assessing their effectiveness and challenges. Results reveal diverse implications for each method. Visual comparisons showcase the varying outcomes of the implemented methods, Likelihood-based assessments present a more numerical approach benefiting the Delta and Kernel methods due to higher scores and fewer negative likelihoods. Conclusions suggest that while multiple methods offer insights into inflation prediction, the Kernel method shows promise in its reliability while the Delta method scores highest in the numerical methods. However, challenges in accurately modelling extreme values and tail behaviours persist across methodologies. Recommendations for further research involve addressing these limitations and exploring enhancements to refine inflation prediction models.F
...
This thesis investigates the estimation of option-implied probability density functions for inflation using inflation options, focusing not only on the expected value but the whole distribution. The aim is to identify the most effective method for measuring the market expectation of future inflation. The research explores both parametric and non-parametric approaches for deriving these density functions from inflation option prices. Methodologies include parametric models such as expansion, generalised distribution, and mixture methods, alongside non-parametric techniques using Breeden and Litzenberger’s result, such as curve-fitting and kernel methods. Implementing these methods involved analysing inflation option data sourced from the BVOL Bloomberg database, specifically for Harmonised Index of Consumer Prices excluding Tobacco (HICPxT) options from January 1, 2013. The study employed Shimko’s method, various spline methods, the Delta method, and Kernel method, assessing their effectiveness and challenges. Results reveal diverse implications for each method. Visual comparisons showcase the varying outcomes of the implemented methods, Likelihood-based assessments present a more numerical approach benefiting the Delta and Kernel methods due to higher scores and fewer negative likelihoods. Conclusions suggest that while multiple methods offer insights into inflation prediction, the Kernel method shows promise in its reliability while the Delta method scores highest in the numerical methods. However, challenges in accurately modelling extreme values and tail behaviours persist across methodologies. Recommendations for further research involve addressing these limitations and exploring enhancements to refine inflation prediction models.F
In this research, we consider neural network-algorithms for option pricing. We use the Black-Scholes model and the lifted Heston model. We derive the option pricing partial differential equation (PDE), which we solve with a neural network, and the conditional characteristic function of the stock price which leads to the option price with the COS method. We consider two neural network-algorithms: the Deep Galerkin Method (DGM) and the Time Deep Nitsche Method (TDNM). We extend the TDNM to be able to solve the option pricing PDE by splitting the PDE operator in a symmetric part and an asymmetric part. The splitting method is more stable than transforming the Black-Scholes option pricing PDE to the symmetric heat equation. The DGM can predict options prices perfectly in the Black-Scholes model even when $r$ and $\sigma$ are added as variables to the neural network. In the lifted Heston model, the DGM can predict option prices perfectly for small dimensions, but has larger errors for larger dimensions. The TDNM is as good as DGM for small times to maturity and volatilities, but has larger errors for large times to maturity and volatilities.
...
In this research, we consider neural network-algorithms for option pricing. We use the Black-Scholes model and the lifted Heston model. We derive the option pricing partial differential equation (PDE), which we solve with a neural network, and the conditional characteristic function of the stock price which leads to the option price with the COS method. We consider two neural network-algorithms: the Deep Galerkin Method (DGM) and the Time Deep Nitsche Method (TDNM). We extend the TDNM to be able to solve the option pricing PDE by splitting the PDE operator in a symmetric part and an asymmetric part. The splitting method is more stable than transforming the Black-Scholes option pricing PDE to the symmetric heat equation. The DGM can predict options prices perfectly in the Black-Scholes model even when $r$ and $\sigma$ are added as variables to the neural network. In the lifted Heston model, the DGM can predict option prices perfectly for small dimensions, but has larger errors for larger dimensions. The TDNM is as good as DGM for small times to maturity and volatilities, but has larger errors for large times to maturity and volatilities.
Option Pricing Techniques
Using Neural Networks
With the emergence of more complex option pricing models, the demand for fast and accurate numerical pricing techniques is increasing. Due to a growing amount of accessible computational power, neural networks have become a feasible numerical method for approximating solutions to these pricing models. This work concentrates on analysing various neural network architectures on option pricing optimisation problems in a supervised and semi-supervised learning setting. We compare the mean-squared error (MSE) and computational training time of a multilayer perceptron (MLP), highway architecture and a recently developed DGM network (Sirignano et al., 2018) along with slight variations on the Black-Scholes and Heston European call option pricing problem as well as the implied volatility problem. We find that on nearly all the supervised learning problems, the generalised highway architecture outperforms its counterparts in terms of MSE relative to computation time. On the Black-Scholes problem, we noticed a reduction of 9.8% in MSE for the generalised highway network while containing 96.2% fewer parameters compared to the MLP considered in (Liu et al., 2019).
On the semi-supervised learning problem, where we directly optimise the neural network to fit the partial differential equation (PDE) and boundary/initial conditions, we concluded that the network architecture of the DGM allows for optimisation of both the interior condition as well as the non-smooth terminal condition. As this was not the case for the MLP and highway networks, the DGM network turned out to be the best performing network architecture on the semi-supervised learning problems. Additionally, we found indications that on the semi-supervised learning problem the performance of the DGM network remained consistent when increasing the dimensionality of the problem. ...
On the semi-supervised learning problem, where we directly optimise the neural network to fit the partial differential equation (PDE) and boundary/initial conditions, we concluded that the network architecture of the DGM allows for optimisation of both the interior condition as well as the non-smooth terminal condition. As this was not the case for the MLP and highway networks, the DGM network turned out to be the best performing network architecture on the semi-supervised learning problems. Additionally, we found indications that on the semi-supervised learning problem the performance of the DGM network remained consistent when increasing the dimensionality of the problem. ...
With the emergence of more complex option pricing models, the demand for fast and accurate numerical pricing techniques is increasing. Due to a growing amount of accessible computational power, neural networks have become a feasible numerical method for approximating solutions to these pricing models. This work concentrates on analysing various neural network architectures on option pricing optimisation problems in a supervised and semi-supervised learning setting. We compare the mean-squared error (MSE) and computational training time of a multilayer perceptron (MLP), highway architecture and a recently developed DGM network (Sirignano et al., 2018) along with slight variations on the Black-Scholes and Heston European call option pricing problem as well as the implied volatility problem. We find that on nearly all the supervised learning problems, the generalised highway architecture outperforms its counterparts in terms of MSE relative to computation time. On the Black-Scholes problem, we noticed a reduction of 9.8% in MSE for the generalised highway network while containing 96.2% fewer parameters compared to the MLP considered in (Liu et al., 2019).
On the semi-supervised learning problem, where we directly optimise the neural network to fit the partial differential equation (PDE) and boundary/initial conditions, we concluded that the network architecture of the DGM allows for optimisation of both the interior condition as well as the non-smooth terminal condition. As this was not the case for the MLP and highway networks, the DGM network turned out to be the best performing network architecture on the semi-supervised learning problems. Additionally, we found indications that on the semi-supervised learning problem the performance of the DGM network remained consistent when increasing the dimensionality of the problem.
On the semi-supervised learning problem, where we directly optimise the neural network to fit the partial differential equation (PDE) and boundary/initial conditions, we concluded that the network architecture of the DGM allows for optimisation of both the interior condition as well as the non-smooth terminal condition. As this was not the case for the MLP and highway networks, the DGM network turned out to be the best performing network architecture on the semi-supervised learning problems. Additionally, we found indications that on the semi-supervised learning problem the performance of the DGM network remained consistent when increasing the dimensionality of the problem.
This thesis investigates the application of machine learning models on foreign exchange data around the WM/R 4pm Closing Spot Rate (colloquially known as the WMR Fix). Due to the nature of the market dynamics around the WMR Fix, inefficiencies can occur and therefore some predictability might be expected. We aim to find these inefficiencies. This is done by applying machine learning models, specifically recurrent neural networks, on limit order book data of foreign exchange (FX). The focus will be on the Euro - US dollar exchange rate.
...
This thesis investigates the application of machine learning models on foreign exchange data around the WM/R 4pm Closing Spot Rate (colloquially known as the WMR Fix). Due to the nature of the market dynamics around the WMR Fix, inefficiencies can occur and therefore some predictability might be expected. We aim to find these inefficiencies. This is done by applying machine learning models, specifically recurrent neural networks, on limit order book data of foreign exchange (FX). The focus will be on the Euro - US dollar exchange rate.
Since the introduction of rough volatility there have been numerous attempts at combining it with existing models in order to better approximate the volatility surface with a low number of parameters. The drawback of rough volatility is usually the time needed to compute a volatility surface. We compare three major rough volatility models and compare their ability to fit to the market volatility surface. We implement the rough Bergomi, rough Heston and lifted Heston models
and introduce a fourth model by taking a Nelson-Siegel parameterization for the instantaneous forward variance curve in the rough Bergomi model. We then minimize the volatility surfaces of the models to that of the market and compare these results through minimization time, fit and hedging possibilities. We find that for both Bergomi models, a single volatility surface is seven times faster to compute than for the lifted Heston model, which in turn is five times faster to compute than for the rough Heston model. Minimization times for the rough Heston models are comparable, however significantly higher than for the rough Bergomi models. We also find that the rough Bergomi model is under-parameterised, this is however fixed by the Nelson-Siegel parameterization, which has a minimization error in line with that of both Heston models. Finally, hedging specifically against a parameter rarely improves the outcome, although a hedge against the instantaneous forward variance curve is consistently good and can improve the error between only a delta hedge by 25-50\% throughout all models. ...
and introduce a fourth model by taking a Nelson-Siegel parameterization for the instantaneous forward variance curve in the rough Bergomi model. We then minimize the volatility surfaces of the models to that of the market and compare these results through minimization time, fit and hedging possibilities. We find that for both Bergomi models, a single volatility surface is seven times faster to compute than for the lifted Heston model, which in turn is five times faster to compute than for the rough Heston model. Minimization times for the rough Heston models are comparable, however significantly higher than for the rough Bergomi models. We also find that the rough Bergomi model is under-parameterised, this is however fixed by the Nelson-Siegel parameterization, which has a minimization error in line with that of both Heston models. Finally, hedging specifically against a parameter rarely improves the outcome, although a hedge against the instantaneous forward variance curve is consistently good and can improve the error between only a delta hedge by 25-50\% throughout all models. ...
Since the introduction of rough volatility there have been numerous attempts at combining it with existing models in order to better approximate the volatility surface with a low number of parameters. The drawback of rough volatility is usually the time needed to compute a volatility surface. We compare three major rough volatility models and compare their ability to fit to the market volatility surface. We implement the rough Bergomi, rough Heston and lifted Heston models
and introduce a fourth model by taking a Nelson-Siegel parameterization for the instantaneous forward variance curve in the rough Bergomi model. We then minimize the volatility surfaces of the models to that of the market and compare these results through minimization time, fit and hedging possibilities. We find that for both Bergomi models, a single volatility surface is seven times faster to compute than for the lifted Heston model, which in turn is five times faster to compute than for the rough Heston model. Minimization times for the rough Heston models are comparable, however significantly higher than for the rough Bergomi models. We also find that the rough Bergomi model is under-parameterised, this is however fixed by the Nelson-Siegel parameterization, which has a minimization error in line with that of both Heston models. Finally, hedging specifically against a parameter rarely improves the outcome, although a hedge against the instantaneous forward variance curve is consistently good and can improve the error between only a delta hedge by 25-50\% throughout all models.
and introduce a fourth model by taking a Nelson-Siegel parameterization for the instantaneous forward variance curve in the rough Bergomi model. We then minimize the volatility surfaces of the models to that of the market and compare these results through minimization time, fit and hedging possibilities. We find that for both Bergomi models, a single volatility surface is seven times faster to compute than for the lifted Heston model, which in turn is five times faster to compute than for the rough Heston model. Minimization times for the rough Heston models are comparable, however significantly higher than for the rough Bergomi models. We also find that the rough Bergomi model is under-parameterised, this is however fixed by the Nelson-Siegel parameterization, which has a minimization error in line with that of both Heston models. Finally, hedging specifically against a parameter rarely improves the outcome, although a hedge against the instantaneous forward variance curve is consistently good and can improve the error between only a delta hedge by 25-50\% throughout all models.
Master thesis
(2021)
-
T.C.A.M. Hack, C.W. Oosterlee, F. Fang, A. Papapantoleon, Natalia Larina-Borovykh, George Inkoom
Interbank-offered-rates play a critical role in the hedging processes of banks, hedge funds or institutional investors. However, the financial stability board recommended to replace these rates by alternative risk-free-rates at the end of 2021. The new rates will be backward-looking rates and therefore, the payoff definitions of interest rate derivatives will change and the currently used Libor Market model to price exotic interest rate derivatives is no longer feasible. This thesis examines a new type of model, the forward market model, which is able to generate both the new backward-looking rates as the current forward-looking rates under the same stochastic process. Besides, contrary to the Libor Market Model, the dynamics under the risk-neutral measure can obtained. Consequently, the new forward market model should always be chosen over the Libor market model. Two issues regarding the forward market model are also considered in this thesis. First of all, the forward market model cannot deal with negative interest rate, this is solved by implementing a shifted version of the log-normal model. Second, a log-normal model is unable to reproduce the implied volatility smile which is present in the market. We solve this issue by combining the forward market model together with the SABR model. Under a few assumptions we derive the shifted SABR forward market model which hasn't been derived in the literature. The model is validated by pricing a new type of caplet that will be present in the post-Libor world, where the payoff won't be known until the payment date. We find that the implementation of this new shifted SABR-FMM can accurately price zero-coupon bonds and caplets in the market. Therefore, we conclude that this new type of model is a possible solution to price exotic interest rate derivatives in the post-Libor world.
...
Interbank-offered-rates play a critical role in the hedging processes of banks, hedge funds or institutional investors. However, the financial stability board recommended to replace these rates by alternative risk-free-rates at the end of 2021. The new rates will be backward-looking rates and therefore, the payoff definitions of interest rate derivatives will change and the currently used Libor Market model to price exotic interest rate derivatives is no longer feasible. This thesis examines a new type of model, the forward market model, which is able to generate both the new backward-looking rates as the current forward-looking rates under the same stochastic process. Besides, contrary to the Libor Market Model, the dynamics under the risk-neutral measure can obtained. Consequently, the new forward market model should always be chosen over the Libor market model. Two issues regarding the forward market model are also considered in this thesis. First of all, the forward market model cannot deal with negative interest rate, this is solved by implementing a shifted version of the log-normal model. Second, a log-normal model is unable to reproduce the implied volatility smile which is present in the market. We solve this issue by combining the forward market model together with the SABR model. Under a few assumptions we derive the shifted SABR forward market model which hasn't been derived in the literature. The model is validated by pricing a new type of caplet that will be present in the post-Libor world, where the payoff won't be known until the payment date. We find that the implementation of this new shifted SABR-FMM can accurately price zero-coupon bonds and caplets in the market. Therefore, we conclude that this new type of model is a possible solution to price exotic interest rate derivatives in the post-Libor world.