D. Kurowicka
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
21 records found
1
At the Netherlands Forensic Institute, additional replicate measurements of the same DNA trace, referred to as rework, can be performed to obtain more information from a DNA mixture profile. Rework may increase the evidential value, expressed by the likelihood ratio (LR), but it also costs laboratory time, resources, and DNA sample material. This thesis investigates whether the LR after rework can be predicted from the original DNA mixture profile.
Two main contributions were made. First, a simulation framework was developed to construct predictive distributions for the rework LR. Starting from the deconvolution of the original profile, plausible contributor genotypes are sampled, additional replicate profiles are simulated, and the LR of the combined profile is calculated. Second, a Bayesian MCMC implementation was developed for the EuroForMix/DNAStatistX peak-height model, making it possible to propagate uncertainty in the nuisance parameters when computing LR values.
The framework was evaluated on cleaned two-person NFI research data, focusing on minor contributors. The frequentist plug-in simulation was not sufficiently calibrated: nominal 95% prediction intervals covered only 69.0% of the observed minor true-donor rework LRs. Including Bayesian parameter uncertainty improved the empirical coverage to 81.6% and reduced the mean interval score from 50.5 to 21.6. However, the predicted distributions remained insufficiently calibrated for casework use.
Overall, this thesis shows that predicting rework LRs is possible in principle and that parameter uncertainty is important for such predictions. The current framework should be viewed as a mathematical proof of concept rather than an operational tool. Further work is needed on artefact modelling, computational scaling, full MCMC validation, extension to more complex mixtures, and validation on casework-like data. ...
Two main contributions were made. First, a simulation framework was developed to construct predictive distributions for the rework LR. Starting from the deconvolution of the original profile, plausible contributor genotypes are sampled, additional replicate profiles are simulated, and the LR of the combined profile is calculated. Second, a Bayesian MCMC implementation was developed for the EuroForMix/DNAStatistX peak-height model, making it possible to propagate uncertainty in the nuisance parameters when computing LR values.
The framework was evaluated on cleaned two-person NFI research data, focusing on minor contributors. The frequentist plug-in simulation was not sufficiently calibrated: nominal 95% prediction intervals covered only 69.0% of the observed minor true-donor rework LRs. Including Bayesian parameter uncertainty improved the empirical coverage to 81.6% and reduced the mean interval score from 50.5 to 21.6. However, the predicted distributions remained insufficiently calibrated for casework use.
Overall, this thesis shows that predicting rework LRs is possible in principle and that parameter uncertainty is important for such predictions. The current framework should be viewed as a mathematical proof of concept rather than an operational tool. Further work is needed on artefact modelling, computational scaling, full MCMC validation, extension to more complex mixtures, and validation on casework-like data. ...
At the Netherlands Forensic Institute, additional replicate measurements of the same DNA trace, referred to as rework, can be performed to obtain more information from a DNA mixture profile. Rework may increase the evidential value, expressed by the likelihood ratio (LR), but it also costs laboratory time, resources, and DNA sample material. This thesis investigates whether the LR after rework can be predicted from the original DNA mixture profile.
Two main contributions were made. First, a simulation framework was developed to construct predictive distributions for the rework LR. Starting from the deconvolution of the original profile, plausible contributor genotypes are sampled, additional replicate profiles are simulated, and the LR of the combined profile is calculated. Second, a Bayesian MCMC implementation was developed for the EuroForMix/DNAStatistX peak-height model, making it possible to propagate uncertainty in the nuisance parameters when computing LR values.
The framework was evaluated on cleaned two-person NFI research data, focusing on minor contributors. The frequentist plug-in simulation was not sufficiently calibrated: nominal 95% prediction intervals covered only 69.0% of the observed minor true-donor rework LRs. Including Bayesian parameter uncertainty improved the empirical coverage to 81.6% and reduced the mean interval score from 50.5 to 21.6. However, the predicted distributions remained insufficiently calibrated for casework use.
Overall, this thesis shows that predicting rework LRs is possible in principle and that parameter uncertainty is important for such predictions. The current framework should be viewed as a mathematical proof of concept rather than an operational tool. Further work is needed on artefact modelling, computational scaling, full MCMC validation, extension to more complex mixtures, and validation on casework-like data.
Two main contributions were made. First, a simulation framework was developed to construct predictive distributions for the rework LR. Starting from the deconvolution of the original profile, plausible contributor genotypes are sampled, additional replicate profiles are simulated, and the LR of the combined profile is calculated. Second, a Bayesian MCMC implementation was developed for the EuroForMix/DNAStatistX peak-height model, making it possible to propagate uncertainty in the nuisance parameters when computing LR values.
The framework was evaluated on cleaned two-person NFI research data, focusing on minor contributors. The frequentist plug-in simulation was not sufficiently calibrated: nominal 95% prediction intervals covered only 69.0% of the observed minor true-donor rework LRs. Including Bayesian parameter uncertainty improved the empirical coverage to 81.6% and reduced the mean interval score from 50.5 to 21.6. However, the predicted distributions remained insufficiently calibrated for casework use.
Overall, this thesis shows that predicting rework LRs is possible in principle and that parameter uncertainty is important for such predictions. The current framework should be viewed as a mathematical proof of concept rather than an operational tool. Further work is needed on artefact modelling, computational scaling, full MCMC validation, extension to more complex mixtures, and validation on casework-like data.
Functional Principal Component Analysis (FPCA) is a powerful method for identifying the dominant modes of variation within a large collection of curves, thereby revealing their underlying low-dimensional structure. However, the existing FPCA framework assumes that the functional data are directly observed. This thesis develops a new theoretical framework for applying FPCA when the functional objects are estimated rather than observed, focusing specifically on estimated conditional Kendall's tau (CKT) curves. The analysis is supported by additional theoretical work on conditional U-statistics, which form the basis for establishing the asymptotic properties of the proposed method. In the theoretical setting where the CKT curves are estimated over the entire domain, we prove the consistency of the associated Hilbert-Schmidt integral operator and under suitable regularity and asymptotic conditions, we show that the spectral decomposition of the estimated operator converges to the true one with an error rate of O(h² + (nhd)-½). We then extend the analysis to the practically relevant setting in which the curves are observed on a finite grid of N points and establish convergence of the spectral properties with a rate of O(N-1/d + h² + (nhd)-½). This illustrates the bias-variance trade-off with respect to the choice of the bandwidth h, as well as the numerical bias introduced by the discretization step, N-1/d.
We support our theoretical results with an empirical study of global financial assets, applying ARMA-GARCH filtering and conditioning on different markets to recover the most dominant modes of variation within a large dataset of CKT curves.
We find that all CKT curves can be approximately summarized by only two degrees of freedom: a constant baseline level, representing the average conditional Kendall's tau, and a parabolic component that contrasts extreme market conditions with more typical periods, capturing the contagion effects.
Ultimately, this framework provides a foundation for future applications, including evaluating the simplifying assumption. ...
We support our theoretical results with an empirical study of global financial assets, applying ARMA-GARCH filtering and conditioning on different markets to recover the most dominant modes of variation within a large dataset of CKT curves.
We find that all CKT curves can be approximately summarized by only two degrees of freedom: a constant baseline level, representing the average conditional Kendall's tau, and a parabolic component that contrasts extreme market conditions with more typical periods, capturing the contagion effects.
Ultimately, this framework provides a foundation for future applications, including evaluating the simplifying assumption. ...
Functional Principal Component Analysis (FPCA) is a powerful method for identifying the dominant modes of variation within a large collection of curves, thereby revealing their underlying low-dimensional structure. However, the existing FPCA framework assumes that the functional data are directly observed. This thesis develops a new theoretical framework for applying FPCA when the functional objects are estimated rather than observed, focusing specifically on estimated conditional Kendall's tau (CKT) curves. The analysis is supported by additional theoretical work on conditional U-statistics, which form the basis for establishing the asymptotic properties of the proposed method. In the theoretical setting where the CKT curves are estimated over the entire domain, we prove the consistency of the associated Hilbert-Schmidt integral operator and under suitable regularity and asymptotic conditions, we show that the spectral decomposition of the estimated operator converges to the true one with an error rate of O(h² + (nhd)-½). We then extend the analysis to the practically relevant setting in which the curves are observed on a finite grid of N points and establish convergence of the spectral properties with a rate of O(N-1/d + h² + (nhd)-½). This illustrates the bias-variance trade-off with respect to the choice of the bandwidth h, as well as the numerical bias introduced by the discretization step, N-1/d.
We support our theoretical results with an empirical study of global financial assets, applying ARMA-GARCH filtering and conditioning on different markets to recover the most dominant modes of variation within a large dataset of CKT curves.
We find that all CKT curves can be approximately summarized by only two degrees of freedom: a constant baseline level, representing the average conditional Kendall's tau, and a parabolic component that contrasts extreme market conditions with more typical periods, capturing the contagion effects.
Ultimately, this framework provides a foundation for future applications, including evaluating the simplifying assumption.
We support our theoretical results with an empirical study of global financial assets, applying ARMA-GARCH filtering and conditioning on different markets to recover the most dominant modes of variation within a large dataset of CKT curves.
We find that all CKT curves can be approximately summarized by only two degrees of freedom: a constant baseline level, representing the average conditional Kendall's tau, and a parabolic component that contrasts extreme market conditions with more typical periods, capturing the contagion effects.
Ultimately, this framework provides a foundation for future applications, including evaluating the simplifying assumption.
High-dimensional Pearson's chi-squared test
Hoge dimensionale Pearson chi-sqaured toets
This paper revisits Pearson's chi-square test and studies its properties, highlighting the behavior of the test when applied to large supports, i.e., the number of cells versus the sample size. First, we explore the general behavior through a controlled simulation, wherein we find that the test exhibits an increased number of type I errors. These errors occur when the sample size is small relative to the number of cells. This behavior will be explained using a generalized central limit theorem, showing that the support needs to be $o\left(\sqrt{\frac{n}{\log{n}}} \right)$.
...
This paper revisits Pearson's chi-square test and studies its properties, highlighting the behavior of the test when applied to large supports, i.e., the number of cells versus the sample size. First, we explore the general behavior through a controlled simulation, wherein we find that the test exhibits an increased number of type I errors. These errors occur when the sample size is small relative to the number of cells. This behavior will be explained using a generalized central limit theorem, showing that the support needs to be $o\left(\sqrt{\frac{n}{\log{n}}} \right)$.
Statistical inference of low-frequency time series is a challenge present in various fields, such as financial risk management and weather forecasting. Practical difficulties arise due to the scarcity of non-overlapping observations. The “direct method”, which directly uses the available low-frequency data to construct estimators, often results in inaccurate estimations.
In this thesis, we propose a novel “simulation-based method” for statistical inference of low-frequency time series that result from the aggregation of a higher-frequency time series over a period of time. We start by estimating the distribution of this higher-frequency process. We then simulate a large number of paths from this estimated distribution. By independently aggregating each simulated path, we generate corresponding low-frequency data. This provides us with a large simulated dataset of the low-frequency process, which enables us to apply estimation procedures and bypass the limitations posed by the shortage of original low-frequency data.
We also provide a theoretical framework and propose three families of estimators constructed from the estimated higher-frequency distribution, analyzing their properties under additional assumptions. Through a comprehensive simulation study, we compare the simulation-based method with the traditional direct method across different scenarios and objectives. While our study focuses on the marginal distributions of low-frequency processes, the simulation-based method’s applicability extends to joint distributions across multiple time points. This research offers a robust method for parameter estimation when faced with limited low-frequency data. ...
In this thesis, we propose a novel “simulation-based method” for statistical inference of low-frequency time series that result from the aggregation of a higher-frequency time series over a period of time. We start by estimating the distribution of this higher-frequency process. We then simulate a large number of paths from this estimated distribution. By independently aggregating each simulated path, we generate corresponding low-frequency data. This provides us with a large simulated dataset of the low-frequency process, which enables us to apply estimation procedures and bypass the limitations posed by the shortage of original low-frequency data.
We also provide a theoretical framework and propose three families of estimators constructed from the estimated higher-frequency distribution, analyzing their properties under additional assumptions. Through a comprehensive simulation study, we compare the simulation-based method with the traditional direct method across different scenarios and objectives. While our study focuses on the marginal distributions of low-frequency processes, the simulation-based method’s applicability extends to joint distributions across multiple time points. This research offers a robust method for parameter estimation when faced with limited low-frequency data. ...
Statistical inference of low-frequency time series is a challenge present in various fields, such as financial risk management and weather forecasting. Practical difficulties arise due to the scarcity of non-overlapping observations. The “direct method”, which directly uses the available low-frequency data to construct estimators, often results in inaccurate estimations.
In this thesis, we propose a novel “simulation-based method” for statistical inference of low-frequency time series that result from the aggregation of a higher-frequency time series over a period of time. We start by estimating the distribution of this higher-frequency process. We then simulate a large number of paths from this estimated distribution. By independently aggregating each simulated path, we generate corresponding low-frequency data. This provides us with a large simulated dataset of the low-frequency process, which enables us to apply estimation procedures and bypass the limitations posed by the shortage of original low-frequency data.
We also provide a theoretical framework and propose three families of estimators constructed from the estimated higher-frequency distribution, analyzing their properties under additional assumptions. Through a comprehensive simulation study, we compare the simulation-based method with the traditional direct method across different scenarios and objectives. While our study focuses on the marginal distributions of low-frequency processes, the simulation-based method’s applicability extends to joint distributions across multiple time points. This research offers a robust method for parameter estimation when faced with limited low-frequency data.
In this thesis, we propose a novel “simulation-based method” for statistical inference of low-frequency time series that result from the aggregation of a higher-frequency time series over a period of time. We start by estimating the distribution of this higher-frequency process. We then simulate a large number of paths from this estimated distribution. By independently aggregating each simulated path, we generate corresponding low-frequency data. This provides us with a large simulated dataset of the low-frequency process, which enables us to apply estimation procedures and bypass the limitations posed by the shortage of original low-frequency data.
We also provide a theoretical framework and propose three families of estimators constructed from the estimated higher-frequency distribution, analyzing their properties under additional assumptions. Through a comprehensive simulation study, we compare the simulation-based method with the traditional direct method across different scenarios and objectives. While our study focuses on the marginal distributions of low-frequency processes, the simulation-based method’s applicability extends to joint distributions across multiple time points. This research offers a robust method for parameter estimation when faced with limited low-frequency data.
This thesis concerns modeling residential real estate selling prices in a hedonic price model framework on a small spatial-temporal granularity. The research addresses the challenge of sparse spatial-temporal real estate data, i.e. many combinations of location and time with few or no transactions, by employing spatial dynamic factor models (SDFMs). Two types of SDFMs are employed: an SDFM with a 1D spatial structure based on the spatial random walk and an SDFM with a 2D spatial structure based on the Gaussian random field. To capture the information on the property characteristics, spatial dynamic factor models are combined with two different data-driven models, namely a neural network (NN) and an interpretable version of an NN, the local generalized linear model network (LGLMN). Both a Bayesian approach and an algorithmic approach are employed to estimate the models on both a PC and a high-performance computer (HPC). A simulation study is conducted to demonstrate the ability of an NN to capture linear and non-linear structures when combined with an SDFM and to show the ability of the LGLMN to replicate a linear structure. Furthermore, the models are evaluated on real transaction data from the municipality of Rotterdam. The findings demonstrate that the algorithmically estimated NN-adjusted SDFM based on the spatial random walk (NN-SRW-DFM) outperforms the other models in terms of accuracy with an out-of-sample MAPE of 0.128. Moreover, the results highlight a trade-off between accuracy, speed, and interpretability.
...
This thesis concerns modeling residential real estate selling prices in a hedonic price model framework on a small spatial-temporal granularity. The research addresses the challenge of sparse spatial-temporal real estate data, i.e. many combinations of location and time with few or no transactions, by employing spatial dynamic factor models (SDFMs). Two types of SDFMs are employed: an SDFM with a 1D spatial structure based on the spatial random walk and an SDFM with a 2D spatial structure based on the Gaussian random field. To capture the information on the property characteristics, spatial dynamic factor models are combined with two different data-driven models, namely a neural network (NN) and an interpretable version of an NN, the local generalized linear model network (LGLMN). Both a Bayesian approach and an algorithmic approach are employed to estimate the models on both a PC and a high-performance computer (HPC). A simulation study is conducted to demonstrate the ability of an NN to capture linear and non-linear structures when combined with an SDFM and to show the ability of the LGLMN to replicate a linear structure. Furthermore, the models are evaluated on real transaction data from the municipality of Rotterdam. The findings demonstrate that the algorithmically estimated NN-adjusted SDFM based on the spatial random walk (NN-SRW-DFM) outperforms the other models in terms of accuracy with an out-of-sample MAPE of 0.128. Moreover, the results highlight a trade-off between accuracy, speed, and interpretability.
Prisons serve as amplifiers of Tuberculosis (TB) transmission due to overpopulation, lack of hygiene, and bad ventilation [Mabud et al., 2019] [Baussano et al., 2010]. The risk of TB is elevated during and after incarceration, only returning to the general population’s risk seven years after liberation. This is an indication that formerly-incarcerated individuals are a highrisk group. By preventing TB among ex-prisoners, the burden and transmission to the general population can remain limited. This study aims to analyze the cost-effectiveness of various preventive measures among formerly-incarcerated individuals in Brazil. With the imprisonment rate growing in the country and the high TB incidence rates, Brazil can benefit greatly from proposed measures. The research mathematically models the effect of several interventions and quantifies associated health benefits and costs with the use of Disability Adjusted Life Years and Incremental Cost-Effectiveness Ratios. The parameters of the model are estimated with Bayesian statistics, using likelihood functions and prior distributions to calculate a posterior distribution. The results of the research indicate how the different interventions behave and which is most advisable for implementation in the Brazilian health system. Our conclusion is that shorter preventive therapies in combination with the skin test Tuberculin Skin Test perform better than longer therapies and/or the blood test Interferon Gamma Release Assays. This research was conducted in collaboration with the Harvard T.H. Chan School of Public Health and the TU Delft, with help from the Brazilian ministry of health.
...
Prisons serve as amplifiers of Tuberculosis (TB) transmission due to overpopulation, lack of hygiene, and bad ventilation [Mabud et al., 2019] [Baussano et al., 2010]. The risk of TB is elevated during and after incarceration, only returning to the general population’s risk seven years after liberation. This is an indication that formerly-incarcerated individuals are a highrisk group. By preventing TB among ex-prisoners, the burden and transmission to the general population can remain limited. This study aims to analyze the cost-effectiveness of various preventive measures among formerly-incarcerated individuals in Brazil. With the imprisonment rate growing in the country and the high TB incidence rates, Brazil can benefit greatly from proposed measures. The research mathematically models the effect of several interventions and quantifies associated health benefits and costs with the use of Disability Adjusted Life Years and Incremental Cost-Effectiveness Ratios. The parameters of the model are estimated with Bayesian statistics, using likelihood functions and prior distributions to calculate a posterior distribution. The results of the research indicate how the different interventions behave and which is most advisable for implementation in the Brazilian health system. Our conclusion is that shorter preventive therapies in combination with the skin test Tuberculin Skin Test perform better than longer therapies and/or the blood test Interferon Gamma Release Assays. This research was conducted in collaboration with the Harvard T.H. Chan School of Public Health and the TU Delft, with help from the Brazilian ministry of health.
Master thesis
(2023)
-
M. Galanis, A.F.F. Derumigny, Gianluca Finocchio, A.W. van der Vaart, D. Kurowicka
In this thesis, we explore the structure of consistent bootstrap statistics in hypothesis testing. Bootstrap, as a very useful technique when theoretical distributions are not available or when the sample size is small, enjoys a lot of interest from applied statisticians. Historically, guidelines for performing Bootstrap have been proposed. One of the guidelines proposed is to center the bootstrap statistic around the true statistic, calculated from the original sample. The second, is to perform resampling in a way such that the new sample reflects the hypothesis tested. However, both of the guidelines are proposed based mostly on an empirical point of view. In this project, we show that the calculation of the bootstrap statistic is directly related to the way the new sample is generated. We describe the specific conditions under which the Bootstrap statistic should or should not be centered around the true. As mentioned the resampling scheme that is picked directly influences this choice. The motivation is derived from the independence test and the same arguments apply to the regression slope test. Finally, we provide a generalized setting where a consistent bootstrap statistic is provided, based on the resampling scheme that is picked.
...
In this thesis, we explore the structure of consistent bootstrap statistics in hypothesis testing. Bootstrap, as a very useful technique when theoretical distributions are not available or when the sample size is small, enjoys a lot of interest from applied statisticians. Historically, guidelines for performing Bootstrap have been proposed. One of the guidelines proposed is to center the bootstrap statistic around the true statistic, calculated from the original sample. The second, is to perform resampling in a way such that the new sample reflects the hypothesis tested. However, both of the guidelines are proposed based mostly on an empirical point of view. In this project, we show that the calculation of the bootstrap statistic is directly related to the way the new sample is generated. We describe the specific conditions under which the Bootstrap statistic should or should not be centered around the true. As mentioned the resampling scheme that is picked directly influences this choice. The motivation is derived from the independence test and the same arguments apply to the regression slope test. Finally, we provide a generalized setting where a consistent bootstrap statistic is provided, based on the resampling scheme that is picked.
Dutch electricity spot price forecasting
Two study cases using structured expert judgement
Master thesis
(2022)
-
A.D.S. Bachasingh, G.F. Nane, D. Kurowicka, P. Chen, M. Koenen, J. Duivenvoorden
Electricity is a quite unique commodity. Due to the economically non-storable nature of the commodity that electricity is, the constant balance between consumption and production, weather effects, such as temperature, wind speed, solar intensity etc, and the intensity of everyday and business activities, e.g. holi- days, weekends, on- and off-peak hours etc., the price dynamics of this commodity are quite unique as they are not observed in any other market. This extreme price volatility has forced market participants to hedge volumes as well as price risks. Naturally, electricity price forecast models are of great interest to portfolio managers.
Current data-driven models are combined with expertise of traders due to the considerable uncertainty of these models. We aim to subject trader expertise to a transparent methodology using structured expert judgement. During this study, we conduct two elicitations whose variables of interest concern average day- ahead spot prices for 2025, 2030 and 2035. We distinguish between baseload and peakload price. The first elicitation uses assessments regarding current Dutch day-ahead electricity spot prices to forecast forecast the variables of interest. We forecast the average day-ahead electricity spot price for 2025, 2030 and 2035. The second elicitation uses assessments regarding data on past, current and future developments in the Dutch electricity markets to forecast forecast the variables of interest. The second elicitation results in a decision maker with a higher price forecast than the decision maker of the first elicitation. Uncertainty regarding the forecast is higher for 2025 and 2030 during the second study. Uncertainty is equally high regarding the peakload price forecast for 2035 in both studies, however the range shifts. ...
Current data-driven models are combined with expertise of traders due to the considerable uncertainty of these models. We aim to subject trader expertise to a transparent methodology using structured expert judgement. During this study, we conduct two elicitations whose variables of interest concern average day- ahead spot prices for 2025, 2030 and 2035. We distinguish between baseload and peakload price. The first elicitation uses assessments regarding current Dutch day-ahead electricity spot prices to forecast forecast the variables of interest. We forecast the average day-ahead electricity spot price for 2025, 2030 and 2035. The second elicitation uses assessments regarding data on past, current and future developments in the Dutch electricity markets to forecast forecast the variables of interest. The second elicitation results in a decision maker with a higher price forecast than the decision maker of the first elicitation. Uncertainty regarding the forecast is higher for 2025 and 2030 during the second study. Uncertainty is equally high regarding the peakload price forecast for 2035 in both studies, however the range shifts. ...
Electricity is a quite unique commodity. Due to the economically non-storable nature of the commodity that electricity is, the constant balance between consumption and production, weather effects, such as temperature, wind speed, solar intensity etc, and the intensity of everyday and business activities, e.g. holi- days, weekends, on- and off-peak hours etc., the price dynamics of this commodity are quite unique as they are not observed in any other market. This extreme price volatility has forced market participants to hedge volumes as well as price risks. Naturally, electricity price forecast models are of great interest to portfolio managers.
Current data-driven models are combined with expertise of traders due to the considerable uncertainty of these models. We aim to subject trader expertise to a transparent methodology using structured expert judgement. During this study, we conduct two elicitations whose variables of interest concern average day- ahead spot prices for 2025, 2030 and 2035. We distinguish between baseload and peakload price. The first elicitation uses assessments regarding current Dutch day-ahead electricity spot prices to forecast forecast the variables of interest. We forecast the average day-ahead electricity spot price for 2025, 2030 and 2035. The second elicitation uses assessments regarding data on past, current and future developments in the Dutch electricity markets to forecast forecast the variables of interest. The second elicitation results in a decision maker with a higher price forecast than the decision maker of the first elicitation. Uncertainty regarding the forecast is higher for 2025 and 2030 during the second study. Uncertainty is equally high regarding the peakload price forecast for 2035 in both studies, however the range shifts.
Current data-driven models are combined with expertise of traders due to the considerable uncertainty of these models. We aim to subject trader expertise to a transparent methodology using structured expert judgement. During this study, we conduct two elicitations whose variables of interest concern average day- ahead spot prices for 2025, 2030 and 2035. We distinguish between baseload and peakload price. The first elicitation uses assessments regarding current Dutch day-ahead electricity spot prices to forecast forecast the variables of interest. We forecast the average day-ahead electricity spot price for 2025, 2030 and 2035. The second elicitation uses assessments regarding data on past, current and future developments in the Dutch electricity markets to forecast forecast the variables of interest. The second elicitation results in a decision maker with a higher price forecast than the decision maker of the first elicitation. Uncertainty regarding the forecast is higher for 2025 and 2030 during the second study. Uncertainty is equally high regarding the peakload price forecast for 2035 in both studies, however the range shifts.
In this thesis, we present a study to obtain a clear and accurate overview of the progress and behaviour of COVID-19 in the Netherlands. We distinguish two parts for this study. The first part is to estimate the total number of infected people as a function of time by combining data from hospital admissions, daily reported cases and serological data. Using these data sets, we found that our estimation for the number of infected people was comparable to the estimations provided by the RIVM and Sanquin. Furthermore, we found that on average only 39.3% of the total number of cases were detected. 1.2% of the total number of infected people is admitted to the hospital and 18.6% of the hospitalized patients is admitted to the ICU. The second part is to develop a representative model that reproduces the estimated total number of infections using a modified SEIR model. These modifications include modelling the infection rate β(t) as a function of time using a simple linear ODE, a system of ODEs inspired by the Lotka-Volterra equations, the implementation of gamma distributed exposed and infected stages and lastly the incorporation of spatial heterogeneity. We found that our Lotka-Volterra inspired model was able to model multiple consecutive waves, which differs from the standard compartmental models. The other modifications however seemed to have only minor effects on the model and had some difficulties with matching historical data. We conclude that our Lotka-Volterra inspired model should be used to model consecutive waves for a longer period of time. The other modifications can be used to optimize the model.
...
In this thesis, we present a study to obtain a clear and accurate overview of the progress and behaviour of COVID-19 in the Netherlands. We distinguish two parts for this study. The first part is to estimate the total number of infected people as a function of time by combining data from hospital admissions, daily reported cases and serological data. Using these data sets, we found that our estimation for the number of infected people was comparable to the estimations provided by the RIVM and Sanquin. Furthermore, we found that on average only 39.3% of the total number of cases were detected. 1.2% of the total number of infected people is admitted to the hospital and 18.6% of the hospitalized patients is admitted to the ICU. The second part is to develop a representative model that reproduces the estimated total number of infections using a modified SEIR model. These modifications include modelling the infection rate β(t) as a function of time using a simple linear ODE, a system of ODEs inspired by the Lotka-Volterra equations, the implementation of gamma distributed exposed and infected stages and lastly the incorporation of spatial heterogeneity. We found that our Lotka-Volterra inspired model was able to model multiple consecutive waves, which differs from the standard compartmental models. The other modifications however seemed to have only minor effects on the model and had some difficulties with matching historical data. We conclude that our Lotka-Volterra inspired model should be used to model consecutive waves for a longer period of time. The other modifications can be used to optimize the model.
Measuring variable importance is often a difficult task: among others models can be complex and covariates can interact with each other and can be correlated. This study focuses on two questions: First, what should be the theoretical measure of variable importance under a given data-generating model? And second, what are the best estimates of these theoretical measures? Two theoretical measures and some corresponding estimates are presented of which one is the well-known random forests variable importance measure (Breiman, 2001). A simulation study is done for both linear and nonlinear models to find out what are the best estimates of variable importance measures for given data-generating models. Most measures struggle when covariates are correlated, but make an improvement in performance when the number of split variables is tuned.
...
Measuring variable importance is often a difficult task: among others models can be complex and covariates can interact with each other and can be correlated. This study focuses on two questions: First, what should be the theoretical measure of variable importance under a given data-generating model? And second, what are the best estimates of these theoretical measures? Two theoretical measures and some corresponding estimates are presented of which one is the well-known random forests variable importance measure (Breiman, 2001). A simulation study is done for both linear and nonlinear models to find out what are the best estimates of variable importance measures for given data-generating models. Most measures struggle when covariates are correlated, but make an improvement in performance when the number of split variables is tuned.
Aligning AI with Human Norms
Multi-Objective Deep Reinforcement Learning with Active Preference Elicitation
The field of deep reinforcement learning has seen major successes recently, achieving superhuman performance in discrete games such as Go and the Atari domain, as well as astounding results in continuous robot locomotion tasks. However, the correct specification of human intentions in a reward function is highly challenging, which is why state-of-the-art methods lack interpretability and may lead to unforeseen societal impacts when deployed in the real world. To tackle this, we propose multi-objective reinforced active learning (MORAL), a novel framework based on inverse reinforcement learning for combining a diverse set of human norms into a single Pareto optimal policy. We show that through the combination of active preference learning and multi-objective decision-making, one can interactively train an agent to trade off a variety of learned norms as well as primary reward functions, thus mitigating negative side effects. Furthermore, we introduce two toy environments called Burning Warehouse and Delivery, which allow for studying the scalability of our approach in both size of the state space and reward complexity. We find that through mixing expert demonstrations and preferences, we can achieve superior efficiency compared to employing a single type of expert feedback and, finally, suggest that unlike previous literature, MORAL is able to learn a deep reward model consisting of multiple expert utility functions.
...
The field of deep reinforcement learning has seen major successes recently, achieving superhuman performance in discrete games such as Go and the Atari domain, as well as astounding results in continuous robot locomotion tasks. However, the correct specification of human intentions in a reward function is highly challenging, which is why state-of-the-art methods lack interpretability and may lead to unforeseen societal impacts when deployed in the real world. To tackle this, we propose multi-objective reinforced active learning (MORAL), a novel framework based on inverse reinforcement learning for combining a diverse set of human norms into a single Pareto optimal policy. We show that through the combination of active preference learning and multi-objective decision-making, one can interactively train an agent to trade off a variety of learned norms as well as primary reward functions, thus mitigating negative side effects. Furthermore, we introduce two toy environments called Burning Warehouse and Delivery, which allow for studying the scalability of our approach in both size of the state space and reward complexity. We find that through mixing expert demonstrations and preferences, we can achieve superior efficiency compared to employing a single type of expert feedback and, finally, suggest that unlike previous literature, MORAL is able to learn a deep reward model consisting of multiple expert utility functions.
Sensitivity analysis for hydrodynamic model of the North Sea
Considering the correlations and dependencies between parameters
In this thesis, sensitivity analysis is used to study the influences of parameters on specific outputs in the hydrodynamic model 3D DCSM-FM of the North Sea. The sensitivity analysis is the study of how uncertainty in the outputs of a model can be divided and allocated to different sources of uncertainty in its inputs. The software Delft3D FM is used to simulate the model, which is developed by Deltares. This thesis is supported by German pilot of the UNITED project, which studies the possibilities of combing blue mussel and sugar kelp cultivation with wind energy. The investigation is conducted at FINO3 platform, 80 km off Sylt.
In this thesis project, temperature and current velocity are selected as outputs in sensitivity analysis, which are influential factors of blue mussel and sugar kelp growth. Specific parameters in the hydrodynamic model are selected as inputs respectively. Three sensitivity methods are conducted: Morris, copula-based and variance-based method. Among them, copula-based and variance-based consider the independency information, while parameters are assumed dependent in Morris method. To generate samples, parameters are transformed into a unit hypercube in Morris method and run on the contour of the grid cell, composing paths. Between two steps within the paths, only one parameter changes at a time (OAT method). The input domain is scanned with a better strategy to separate the paths maximizing the dispersion. Each parameter’s elementary effects are calculated within paths, evaluating the changes in outputs contributed by the single parameter. Then (absolute) mean and variance of elementary effects are used as sensitivity indices. Copula-based method uses similar sampling and the same measurements, but it gathers parameters into copulas before sampling to include the dependency information. In variance-based method, the variance of the outputs’ conditional expectations is used to measure the sensitivity. Only random samples are needed in this method.
Morris and copula-based method prove the temporal and spatial similarities of parameters’ sensitivity behavior. Convective and evaporative heat flux are the most influential parameters of temperature, and they also show correlations with other parameters. Air density influences current velocity the most, while Smagorinsky factor is the most correlated parameter of current velocity. Variance-based method gives similar results about rankings of influences. Correlations are proved to exist among parameters. Uniform horizontal eddy viscosity/viscosity in definition files have no impacts, as they are overwritten. Three methods are compared. Except the differences in independency, sampling and measurement, there are some other differences. Much more samples are required in variance-based method and it is used mainly to decide the existence of significant correlations and whether a parameter can be neglected. While Morris and copula-based method ranks the influences and correlations. ...
In this thesis project, temperature and current velocity are selected as outputs in sensitivity analysis, which are influential factors of blue mussel and sugar kelp growth. Specific parameters in the hydrodynamic model are selected as inputs respectively. Three sensitivity methods are conducted: Morris, copula-based and variance-based method. Among them, copula-based and variance-based consider the independency information, while parameters are assumed dependent in Morris method. To generate samples, parameters are transformed into a unit hypercube in Morris method and run on the contour of the grid cell, composing paths. Between two steps within the paths, only one parameter changes at a time (OAT method). The input domain is scanned with a better strategy to separate the paths maximizing the dispersion. Each parameter’s elementary effects are calculated within paths, evaluating the changes in outputs contributed by the single parameter. Then (absolute) mean and variance of elementary effects are used as sensitivity indices. Copula-based method uses similar sampling and the same measurements, but it gathers parameters into copulas before sampling to include the dependency information. In variance-based method, the variance of the outputs’ conditional expectations is used to measure the sensitivity. Only random samples are needed in this method.
Morris and copula-based method prove the temporal and spatial similarities of parameters’ sensitivity behavior. Convective and evaporative heat flux are the most influential parameters of temperature, and they also show correlations with other parameters. Air density influences current velocity the most, while Smagorinsky factor is the most correlated parameter of current velocity. Variance-based method gives similar results about rankings of influences. Correlations are proved to exist among parameters. Uniform horizontal eddy viscosity/viscosity in definition files have no impacts, as they are overwritten. Three methods are compared. Except the differences in independency, sampling and measurement, there are some other differences. Much more samples are required in variance-based method and it is used mainly to decide the existence of significant correlations and whether a parameter can be neglected. While Morris and copula-based method ranks the influences and correlations. ...
In this thesis, sensitivity analysis is used to study the influences of parameters on specific outputs in the hydrodynamic model 3D DCSM-FM of the North Sea. The sensitivity analysis is the study of how uncertainty in the outputs of a model can be divided and allocated to different sources of uncertainty in its inputs. The software Delft3D FM is used to simulate the model, which is developed by Deltares. This thesis is supported by German pilot of the UNITED project, which studies the possibilities of combing blue mussel and sugar kelp cultivation with wind energy. The investigation is conducted at FINO3 platform, 80 km off Sylt.
In this thesis project, temperature and current velocity are selected as outputs in sensitivity analysis, which are influential factors of blue mussel and sugar kelp growth. Specific parameters in the hydrodynamic model are selected as inputs respectively. Three sensitivity methods are conducted: Morris, copula-based and variance-based method. Among them, copula-based and variance-based consider the independency information, while parameters are assumed dependent in Morris method. To generate samples, parameters are transformed into a unit hypercube in Morris method and run on the contour of the grid cell, composing paths. Between two steps within the paths, only one parameter changes at a time (OAT method). The input domain is scanned with a better strategy to separate the paths maximizing the dispersion. Each parameter’s elementary effects are calculated within paths, evaluating the changes in outputs contributed by the single parameter. Then (absolute) mean and variance of elementary effects are used as sensitivity indices. Copula-based method uses similar sampling and the same measurements, but it gathers parameters into copulas before sampling to include the dependency information. In variance-based method, the variance of the outputs’ conditional expectations is used to measure the sensitivity. Only random samples are needed in this method.
Morris and copula-based method prove the temporal and spatial similarities of parameters’ sensitivity behavior. Convective and evaporative heat flux are the most influential parameters of temperature, and they also show correlations with other parameters. Air density influences current velocity the most, while Smagorinsky factor is the most correlated parameter of current velocity. Variance-based method gives similar results about rankings of influences. Correlations are proved to exist among parameters. Uniform horizontal eddy viscosity/viscosity in definition files have no impacts, as they are overwritten. Three methods are compared. Except the differences in independency, sampling and measurement, there are some other differences. Much more samples are required in variance-based method and it is used mainly to decide the existence of significant correlations and whether a parameter can be neglected. While Morris and copula-based method ranks the influences and correlations.
In this thesis project, temperature and current velocity are selected as outputs in sensitivity analysis, which are influential factors of blue mussel and sugar kelp growth. Specific parameters in the hydrodynamic model are selected as inputs respectively. Three sensitivity methods are conducted: Morris, copula-based and variance-based method. Among them, copula-based and variance-based consider the independency information, while parameters are assumed dependent in Morris method. To generate samples, parameters are transformed into a unit hypercube in Morris method and run on the contour of the grid cell, composing paths. Between two steps within the paths, only one parameter changes at a time (OAT method). The input domain is scanned with a better strategy to separate the paths maximizing the dispersion. Each parameter’s elementary effects are calculated within paths, evaluating the changes in outputs contributed by the single parameter. Then (absolute) mean and variance of elementary effects are used as sensitivity indices. Copula-based method uses similar sampling and the same measurements, but it gathers parameters into copulas before sampling to include the dependency information. In variance-based method, the variance of the outputs’ conditional expectations is used to measure the sensitivity. Only random samples are needed in this method.
Morris and copula-based method prove the temporal and spatial similarities of parameters’ sensitivity behavior. Convective and evaporative heat flux are the most influential parameters of temperature, and they also show correlations with other parameters. Air density influences current velocity the most, while Smagorinsky factor is the most correlated parameter of current velocity. Variance-based method gives similar results about rankings of influences. Correlations are proved to exist among parameters. Uniform horizontal eddy viscosity/viscosity in definition files have no impacts, as they are overwritten. Three methods are compared. Except the differences in independency, sampling and measurement, there are some other differences. Much more samples are required in variance-based method and it is used mainly to decide the existence of significant correlations and whether a parameter can be neglected. While Morris and copula-based method ranks the influences and correlations.
Safe hypothesis tests are tests that are robust under accumulation bias, namely when there are dependencies between the results of previous studies and the decision whether to conduct further studies. We construct two types of safe test for the 2 × 2 contingency table, the conditional and unconditional safe tests. In general safe tests are given by an information projection that may be difficult to compute. The conditional tests we construct however are given either in explicit form or implicitly via a defining equation. The same can be said of many of the unconditional tests we construct, for which we prove a number of theoretical results enabling their quick calculation when not given explicitly. The method we develop to accomplish this may perhaps be used to identify optimal safe tests in many other scenarios.
...
Safe hypothesis tests are tests that are robust under accumulation bias, namely when there are dependencies between the results of previous studies and the decision whether to conduct further studies. We construct two types of safe test for the 2 × 2 contingency table, the conditional and unconditional safe tests. In general safe tests are given by an information projection that may be difficult to compute. The conditional tests we construct however are given either in explicit form or implicitly via a defining equation. The same can be said of many of the unconditional tests we construct, for which we prove a number of theoretical results enabling their quick calculation when not given explicitly. The method we develop to accomplish this may perhaps be used to identify optimal safe tests in many other scenarios.
In this thesis we shall consider sample covariance matrices Sn in the case when the dimension of the data increases with the sample size to infinity ,while the ratio approaches a fixed constant. We will derive a new statistic based on the general linear shrinkage estimator by Bodnar et al. (2014)[1] We will show that the new statistic is normally distributed under the null hypothesis that the true covariance matrix is the identity, where we assume the existence of the fourth moment of our data.
Furthermore, we will do simulation study that compares our new statistic to tests from finite dimensional statistics that have been altered to work in high dimensional statistics by Wang and Yao [3]. We will look at three different hypothesis, the equicorrelation case, the auto-regressive case and a fixed ratio case. After that, we will look at the non-linear shrinkage estimator based on the work by Ledoit and Peche (2011) [11], and show that, under the null hypothesis, constructing a test is not directly possible like it is in the linear case. ...
Furthermore, we will do simulation study that compares our new statistic to tests from finite dimensional statistics that have been altered to work in high dimensional statistics by Wang and Yao [3]. We will look at three different hypothesis, the equicorrelation case, the auto-regressive case and a fixed ratio case. After that, we will look at the non-linear shrinkage estimator based on the work by Ledoit and Peche (2011) [11], and show that, under the null hypothesis, constructing a test is not directly possible like it is in the linear case. ...
In this thesis we shall consider sample covariance matrices Sn in the case when the dimension of the data increases with the sample size to infinity ,while the ratio approaches a fixed constant. We will derive a new statistic based on the general linear shrinkage estimator by Bodnar et al. (2014)[1] We will show that the new statistic is normally distributed under the null hypothesis that the true covariance matrix is the identity, where we assume the existence of the fourth moment of our data.
Furthermore, we will do simulation study that compares our new statistic to tests from finite dimensional statistics that have been altered to work in high dimensional statistics by Wang and Yao [3]. We will look at three different hypothesis, the equicorrelation case, the auto-regressive case and a fixed ratio case. After that, we will look at the non-linear shrinkage estimator based on the work by Ledoit and Peche (2011) [11], and show that, under the null hypothesis, constructing a test is not directly possible like it is in the linear case.
Furthermore, we will do simulation study that compares our new statistic to tests from finite dimensional statistics that have been altered to work in high dimensional statistics by Wang and Yao [3]. We will look at three different hypothesis, the equicorrelation case, the auto-regressive case and a fixed ratio case. After that, we will look at the non-linear shrinkage estimator based on the work by Ledoit and Peche (2011) [11], and show that, under the null hypothesis, constructing a test is not directly possible like it is in the linear case.
In this modern age, data is being generated constantly and data is being saved for analysis everywhere. In the maritime industry, interest in the analysis of ship data has grown over the years. In this thesis, we will take a look at AIS data coupled with sea state data. AIS data is data generated from the ship, concerning the ship locations, speed, and heading, among others. When coupled with data such as the wave height and wave directions at these locations, we can analyse the ship operations in different sea conditions. We analyzed 46 Damen ships of the same type, that operate in different regions of the world. The aim was to make interpretable groups of ships that have similar operation profiles, and to investigate the effect of different sea states on the ship operations. We first enrich the data with port labels, from which we can define trips as sequences of points away from port. We also estimate path lengths between points using Bézier curves. From this we get a relevant set of variables that can use for an unsupervised learning task. We clustered the ships using three methods: principal components analysis, Kmeans, and hierarchical clustering. Principal components analysis showed variation in the ships, but interpretation and definitive clusters were not clear. We then used the K-means method to make 12 clusters of ships, of which six clusters proved to be stable. Hierarchical clustering showed similar results. Interpretation of these clusters was possible, mainly by looking at separate trips. We therefore also clustered the trips, to get classes of trips. We used the K-means method and obtained six clusters of trips, of which five were stable. We also look at ship availability in different regions during different sea conditions. We use an isotonic regression method to test whether ships ships stay in port more often during heavy weather. We found regions where availability decreases during high waves and regions where availability seemed independent from wave height. This most likely has to do with the function of the ship. And finally we look at sailing speeds during different sea states and find that sea state data alone is not sufficient to adequately estimate sailing speeds of a ship. The conclusion is that using the variables that we created, stable clusters can be obtained. These clusters are interpretable and can lead to a better understanding of customer needs. Coupling the AIS data with more data sources however would be a recommendation, since that can lead to more informative clusters, and might lead to more insight into sailing speeds.
...
In this modern age, data is being generated constantly and data is being saved for analysis everywhere. In the maritime industry, interest in the analysis of ship data has grown over the years. In this thesis, we will take a look at AIS data coupled with sea state data. AIS data is data generated from the ship, concerning the ship locations, speed, and heading, among others. When coupled with data such as the wave height and wave directions at these locations, we can analyse the ship operations in different sea conditions. We analyzed 46 Damen ships of the same type, that operate in different regions of the world. The aim was to make interpretable groups of ships that have similar operation profiles, and to investigate the effect of different sea states on the ship operations. We first enrich the data with port labels, from which we can define trips as sequences of points away from port. We also estimate path lengths between points using Bézier curves. From this we get a relevant set of variables that can use for an unsupervised learning task. We clustered the ships using three methods: principal components analysis, Kmeans, and hierarchical clustering. Principal components analysis showed variation in the ships, but interpretation and definitive clusters were not clear. We then used the K-means method to make 12 clusters of ships, of which six clusters proved to be stable. Hierarchical clustering showed similar results. Interpretation of these clusters was possible, mainly by looking at separate trips. We therefore also clustered the trips, to get classes of trips. We used the K-means method and obtained six clusters of trips, of which five were stable. We also look at ship availability in different regions during different sea conditions. We use an isotonic regression method to test whether ships ships stay in port more often during heavy weather. We found regions where availability decreases during high waves and regions where availability seemed independent from wave height. This most likely has to do with the function of the ship. And finally we look at sailing speeds during different sea states and find that sea state data alone is not sufficient to adequately estimate sailing speeds of a ship. The conclusion is that using the variables that we created, stable clusters can be obtained. These clusters are interpretable and can lead to a better understanding of customer needs. Coupling the AIS data with more data sources however would be a recommendation, since that can lead to more informative clusters, and might lead to more insight into sailing speeds.
In de statistiek zijn er verschillende methodes voor het uitvoeren van model selectie. Het verschil in deze methodes komt voort uit het verschil in stromingen. Voor niet-geneste model selectie zijn de meest ganbare stromingen de Bayes Factor en de likelihood ratio. D. M. Ommen en C. P. Saunders presenteerden theoretische resultaten voor de relatie tussen de Bayes Factor en de likelihood ratio, waardoor de resultaten van beide paradigma's met elkaar kunnen worden vergeleken [5]. De Bayes Factor wordt uitgedrukt in de verwachting van likelihoodratio functie met betrekking tot de posterior verdeling van de parameters. In de bewijzen van deze theoriën ontbraken een aantal belangrijke aspecten. Om deze resultaten volledig te kunnen bewijzen, wordt in dit verslag aangetoond dat eerdere theoretische resultaten moeten worden uitgebreid met extra aannames of volledig moeten worden aangepast. Aan de hand van deze aannames en aanpassingen worden de theoretische resultaten bewezen. De theoretische resultaten zijn belangrijk in de toepassing van niet-geneste model selectie, omdat er nu gecommuniceerd kan worden tussen experts vanuit een ander paradigma.
...
In de statistiek zijn er verschillende methodes voor het uitvoeren van model selectie. Het verschil in deze methodes komt voort uit het verschil in stromingen. Voor niet-geneste model selectie zijn de meest ganbare stromingen de Bayes Factor en de likelihood ratio. D. M. Ommen en C. P. Saunders presenteerden theoretische resultaten voor de relatie tussen de Bayes Factor en de likelihood ratio, waardoor de resultaten van beide paradigma's met elkaar kunnen worden vergeleken [5]. De Bayes Factor wordt uitgedrukt in de verwachting van likelihoodratio functie met betrekking tot de posterior verdeling van de parameters. In de bewijzen van deze theoriën ontbraken een aantal belangrijke aspecten. Om deze resultaten volledig te kunnen bewijzen, wordt in dit verslag aangetoond dat eerdere theoretische resultaten moeten worden uitgebreid met extra aannames of volledig moeten worden aangepast. Aan de hand van deze aannames en aanpassingen worden de theoretische resultaten bewezen. De theoretische resultaten zijn belangrijk in de toepassing van niet-geneste model selectie, omdat er nu gecommuniceerd kan worden tussen experts vanuit een ander paradigma.
In dit onderzoek zijn twee manieren van A/B-testen met elkaar vergeleken. A/B-testen is het vergelijken van verschillende website versies om te achterhalen welke versie voor een hogere opbrengt zorgt. Consumenten krijgen afzonderlijk meerdere versies van een website te zien: versie A, versie B, versie C, et cetera. De websites verschillen op basis van één onderscheidend kenmerk van elkaar.
De twee manieren van A/B-testen zijn de klassieke methode en de meer recent ontwikkelde bandit-methode. In de klassieke methode van A/B-testen krijgen meerdere groepen van consumenten de verschillende versies van de website te zien en wordt achteraf bepaald welke website versie de betere versie is op basis van het aantal conversies (conversies zijn bijvoorbeeld aankopen, clicks, et cetera). Het nadeel van deze test is dat je pas achteraf weet welke website de hoogste opbrengst oplevert, terwijl deze uitkomst misschien al tijdens de steekproef duidelijk wordt. De bandit-methode heeft meerdere varianten, waarvan vijf varianten in dit onderzoek zijn onderzocht en vergeleken. De bandit-methode van A/B-testen blijft gedurende de steekproef de groep consumenten die naar de verschillende versies van de website gestuurd worden aanpassen zodat al tijdens het testen zo min mogelijk consumenten naar de slechter presterende website gestuurd worden. Dit betekent namelijk misgelopen opbrengsten, ook wel spijt in A/B-testen. In dit onderzoek wordt daarom de klassieke methode vergeleken met een vijftal bandit-methoden.
Aan de hand van een gesimuleerde steekproef en drie fictieve versies van een website (A,B en C) worden de uitkomsten van de zes methoden op basis van statistische analyses vergeleken. De analyses zijn in RStudio geprogrammeerd en uitgevoerd. Hieruit blijkt dat, alhoewel alle zes de methoden uiteindelijk de best presterende versie kunnen aanwijzen, er duidelijke verschillen zichtbaar zijn in de totale conversies en de spijt. De bandit-methode verbetert op die onderdelen de klassieke methode. Daarnaast zijn er aanwijzingen dat de bandit-methoden de verschillende conversieratio’s van de website versies eerder statistisch significant kunnen aantonen. ...
De twee manieren van A/B-testen zijn de klassieke methode en de meer recent ontwikkelde bandit-methode. In de klassieke methode van A/B-testen krijgen meerdere groepen van consumenten de verschillende versies van de website te zien en wordt achteraf bepaald welke website versie de betere versie is op basis van het aantal conversies (conversies zijn bijvoorbeeld aankopen, clicks, et cetera). Het nadeel van deze test is dat je pas achteraf weet welke website de hoogste opbrengst oplevert, terwijl deze uitkomst misschien al tijdens de steekproef duidelijk wordt. De bandit-methode heeft meerdere varianten, waarvan vijf varianten in dit onderzoek zijn onderzocht en vergeleken. De bandit-methode van A/B-testen blijft gedurende de steekproef de groep consumenten die naar de verschillende versies van de website gestuurd worden aanpassen zodat al tijdens het testen zo min mogelijk consumenten naar de slechter presterende website gestuurd worden. Dit betekent namelijk misgelopen opbrengsten, ook wel spijt in A/B-testen. In dit onderzoek wordt daarom de klassieke methode vergeleken met een vijftal bandit-methoden.
Aan de hand van een gesimuleerde steekproef en drie fictieve versies van een website (A,B en C) worden de uitkomsten van de zes methoden op basis van statistische analyses vergeleken. De analyses zijn in RStudio geprogrammeerd en uitgevoerd. Hieruit blijkt dat, alhoewel alle zes de methoden uiteindelijk de best presterende versie kunnen aanwijzen, er duidelijke verschillen zichtbaar zijn in de totale conversies en de spijt. De bandit-methode verbetert op die onderdelen de klassieke methode. Daarnaast zijn er aanwijzingen dat de bandit-methoden de verschillende conversieratio’s van de website versies eerder statistisch significant kunnen aantonen. ...
In dit onderzoek zijn twee manieren van A/B-testen met elkaar vergeleken. A/B-testen is het vergelijken van verschillende website versies om te achterhalen welke versie voor een hogere opbrengt zorgt. Consumenten krijgen afzonderlijk meerdere versies van een website te zien: versie A, versie B, versie C, et cetera. De websites verschillen op basis van één onderscheidend kenmerk van elkaar.
De twee manieren van A/B-testen zijn de klassieke methode en de meer recent ontwikkelde bandit-methode. In de klassieke methode van A/B-testen krijgen meerdere groepen van consumenten de verschillende versies van de website te zien en wordt achteraf bepaald welke website versie de betere versie is op basis van het aantal conversies (conversies zijn bijvoorbeeld aankopen, clicks, et cetera). Het nadeel van deze test is dat je pas achteraf weet welke website de hoogste opbrengst oplevert, terwijl deze uitkomst misschien al tijdens de steekproef duidelijk wordt. De bandit-methode heeft meerdere varianten, waarvan vijf varianten in dit onderzoek zijn onderzocht en vergeleken. De bandit-methode van A/B-testen blijft gedurende de steekproef de groep consumenten die naar de verschillende versies van de website gestuurd worden aanpassen zodat al tijdens het testen zo min mogelijk consumenten naar de slechter presterende website gestuurd worden. Dit betekent namelijk misgelopen opbrengsten, ook wel spijt in A/B-testen. In dit onderzoek wordt daarom de klassieke methode vergeleken met een vijftal bandit-methoden.
Aan de hand van een gesimuleerde steekproef en drie fictieve versies van een website (A,B en C) worden de uitkomsten van de zes methoden op basis van statistische analyses vergeleken. De analyses zijn in RStudio geprogrammeerd en uitgevoerd. Hieruit blijkt dat, alhoewel alle zes de methoden uiteindelijk de best presterende versie kunnen aanwijzen, er duidelijke verschillen zichtbaar zijn in de totale conversies en de spijt. De bandit-methode verbetert op die onderdelen de klassieke methode. Daarnaast zijn er aanwijzingen dat de bandit-methoden de verschillende conversieratio’s van de website versies eerder statistisch significant kunnen aantonen.
De twee manieren van A/B-testen zijn de klassieke methode en de meer recent ontwikkelde bandit-methode. In de klassieke methode van A/B-testen krijgen meerdere groepen van consumenten de verschillende versies van de website te zien en wordt achteraf bepaald welke website versie de betere versie is op basis van het aantal conversies (conversies zijn bijvoorbeeld aankopen, clicks, et cetera). Het nadeel van deze test is dat je pas achteraf weet welke website de hoogste opbrengst oplevert, terwijl deze uitkomst misschien al tijdens de steekproef duidelijk wordt. De bandit-methode heeft meerdere varianten, waarvan vijf varianten in dit onderzoek zijn onderzocht en vergeleken. De bandit-methode van A/B-testen blijft gedurende de steekproef de groep consumenten die naar de verschillende versies van de website gestuurd worden aanpassen zodat al tijdens het testen zo min mogelijk consumenten naar de slechter presterende website gestuurd worden. Dit betekent namelijk misgelopen opbrengsten, ook wel spijt in A/B-testen. In dit onderzoek wordt daarom de klassieke methode vergeleken met een vijftal bandit-methoden.
Aan de hand van een gesimuleerde steekproef en drie fictieve versies van een website (A,B en C) worden de uitkomsten van de zes methoden op basis van statistische analyses vergeleken. De analyses zijn in RStudio geprogrammeerd en uitgevoerd. Hieruit blijkt dat, alhoewel alle zes de methoden uiteindelijk de best presterende versie kunnen aanwijzen, er duidelijke verschillen zichtbaar zijn in de totale conversies en de spijt. De bandit-methode verbetert op die onderdelen de klassieke methode. Daarnaast zijn er aanwijzingen dat de bandit-methoden de verschillende conversieratio’s van de website versies eerder statistisch significant kunnen aantonen.
Efficient Inference with Panel Data
On the pass-through of the Dutch 2001 and 2012 VAT increases to consumer prices
This thesis evaluates the pass-through of the 2001 and 2012 Dutch Value Added Tax (VAT) increases to customer prices using a difference-in-differences model. To this end, the first difference and feasible generalised least squares estimators are introduced. Contrary to the conventional pooled OLS estimator, these estimators always show significant causal effects for both VAT hikes. These also dramatically improve the accuracy of the estimates compared to earlier research on the incidence of VAT. For the 2012 tax increase, the null hypothesis of full pass through is even rejected. This result is a novelty in the econometric literature. Even in more general settings, the estimators used in this thesis prove far superior over conventional causal estimation techniques of difference-in-differences models.
...
...
This thesis evaluates the pass-through of the 2001 and 2012 Dutch Value Added Tax (VAT) increases to customer prices using a difference-in-differences model. To this end, the first difference and feasible generalised least squares estimators are introduced. Contrary to the conventional pooled OLS estimator, these estimators always show significant causal effects for both VAT hikes. These also dramatically improve the accuracy of the estimates compared to earlier research on the incidence of VAT. For the 2012 tax increase, the null hypothesis of full pass through is even rejected. This result is a novelty in the econometric literature. Even in more general settings, the estimators used in this thesis prove far superior over conventional causal estimation techniques of difference-in-differences models.
Master thesis
(2017)
-
Max Roberto Ortega Del Vecchyo, Dion Gijswijt, Dorota Kurowicka, Karen Aardal, F. Phillipson, A. Sangers
In this report we present an interactive multi-objective optimization tool that was developed as part of this graduation project. This appliance is meant to be used as a decision support tool for transportation planners working on a synchromodal transportation network on the container-to-mode assignment, where different attributes are considered important. The tool offers the planner a range of solutions according to her/his preferences, and offers the opportunity to seek for new ones if the planner is not satisfied with the solutions found so far. Before presenting the tool, a framework for synchromodal transportation problems is introduced (which was developed as part of a collaborative work with two other students and two supervisors). Then an analysis is done on mathematical modeling approaches of container-to-mode assignment, with special emphasis on computational time, given the time-sensitivity of this problem on synchromodal transport networks. From this analysis a model is chosen on which to built upon the multi-objective optimization tool, for which a thorough analysis on the attributes to be considered is carried out.
...
In this report we present an interactive multi-objective optimization tool that was developed as part of this graduation project. This appliance is meant to be used as a decision support tool for transportation planners working on a synchromodal transportation network on the container-to-mode assignment, where different attributes are considered important. The tool offers the planner a range of solutions according to her/his preferences, and offers the opportunity to seek for new ones if the planner is not satisfied with the solutions found so far. Before presenting the tool, a framework for synchromodal transportation problems is introduced (which was developed as part of a collaborative work with two other students and two supervisors). Then an analysis is done on mathematical modeling approaches of container-to-mode assignment, with special emphasis on computational time, given the time-sensitivity of this problem on synchromodal transport networks. From this analysis a model is chosen on which to built upon the multi-objective optimization tool, for which a thorough analysis on the attributes to be considered is carried out.
Master thesis
(2017)
-
Veerle Berghuis, Jakob Söhl, Bram Van Hoof, Geurt Jongbloed, Dorota Kurowicka, Maarten Voncken
ASML produces TwinScan NXT machines that are used for the production of microchips. The machines ensure that an accurate pattern of DUV-light passes a lens and that it is projected as accurate as possible on the wafer. To ensure that the focal point of the converged DUV-light falls exactly onto the wafer, the leveling functionality is of great importance. That is, placing the wafer in the correct depth of focus by rotating the wafer and moving the wafer up or down during exposure.
In order to meet the imaging requirements, the performance is investigated by analyzing errors of the machines at customers' site, considering one-year data. The most important errors are A, B, C and D. To reduce the total unscheduled down (USD) time of those errors, we should focus on reducing the mean USD time for errors A and C; and focus on reducing the frequency for errors B andD where these last two errors are likely to be solved together.
Different nominal customer-related variables are considered as possible causes of USD times such as location, system type or type of sensors. After applying hierarchical clustering and multidimensional scaling, the variable set is reduced. This set is used to model the USD time of one error: B. Significant differences in USD times are found, showed by the robust and distribution free rank tests: Wilcoxon and Kruskal-Wallis test.
To discover interesting patterns among variables, regression models are applied. The linear regression model and generalized linear model not seem to be the right model to the data. The zero adjusted exponential model seems to be the correct model and show that AG type, location and FSM flex package are the most important explanatory variables. This directs to a potential root cause where ASML is working further upon. ...
In order to meet the imaging requirements, the performance is investigated by analyzing errors of the machines at customers' site, considering one-year data. The most important errors are A, B, C and D. To reduce the total unscheduled down (USD) time of those errors, we should focus on reducing the mean USD time for errors A and C; and focus on reducing the frequency for errors B andD where these last two errors are likely to be solved together.
Different nominal customer-related variables are considered as possible causes of USD times such as location, system type or type of sensors. After applying hierarchical clustering and multidimensional scaling, the variable set is reduced. This set is used to model the USD time of one error: B. Significant differences in USD times are found, showed by the robust and distribution free rank tests: Wilcoxon and Kruskal-Wallis test.
To discover interesting patterns among variables, regression models are applied. The linear regression model and generalized linear model not seem to be the right model to the data. The zero adjusted exponential model seems to be the correct model and show that AG type, location and FSM flex package are the most important explanatory variables. This directs to a potential root cause where ASML is working further upon. ...
ASML produces TwinScan NXT machines that are used for the production of microchips. The machines ensure that an accurate pattern of DUV-light passes a lens and that it is projected as accurate as possible on the wafer. To ensure that the focal point of the converged DUV-light falls exactly onto the wafer, the leveling functionality is of great importance. That is, placing the wafer in the correct depth of focus by rotating the wafer and moving the wafer up or down during exposure.
In order to meet the imaging requirements, the performance is investigated by analyzing errors of the machines at customers' site, considering one-year data. The most important errors are A, B, C and D. To reduce the total unscheduled down (USD) time of those errors, we should focus on reducing the mean USD time for errors A and C; and focus on reducing the frequency for errors B andD where these last two errors are likely to be solved together.
Different nominal customer-related variables are considered as possible causes of USD times such as location, system type or type of sensors. After applying hierarchical clustering and multidimensional scaling, the variable set is reduced. This set is used to model the USD time of one error: B. Significant differences in USD times are found, showed by the robust and distribution free rank tests: Wilcoxon and Kruskal-Wallis test.
To discover interesting patterns among variables, regression models are applied. The linear regression model and generalized linear model not seem to be the right model to the data. The zero adjusted exponential model seems to be the correct model and show that AG type, location and FSM flex package are the most important explanatory variables. This directs to a potential root cause where ASML is working further upon.
In order to meet the imaging requirements, the performance is investigated by analyzing errors of the machines at customers' site, considering one-year data. The most important errors are A, B, C and D. To reduce the total unscheduled down (USD) time of those errors, we should focus on reducing the mean USD time for errors A and C; and focus on reducing the frequency for errors B andD where these last two errors are likely to be solved together.
Different nominal customer-related variables are considered as possible causes of USD times such as location, system type or type of sensors. After applying hierarchical clustering and multidimensional scaling, the variable set is reduced. This set is used to model the USD time of one error: B. Significant differences in USD times are found, showed by the robust and distribution free rank tests: Wilcoxon and Kruskal-Wallis test.
To discover interesting patterns among variables, regression models are applied. The linear regression model and generalized linear model not seem to be the right model to the data. The zero adjusted exponential model seems to be the correct model and show that AG type, location and FSM flex package are the most important explanatory variables. This directs to a potential root cause where ASML is working further upon.