A.F.F. Derumigny
Please Note
20 records found
1
The first part introduces the theoretical background on time series, copulas, maximum likelihood estimation, Generalized Autoregressive Score (GAS) updating, and GARCH marginal models. The second part uses simulation studies to explain why the observed behaviour of a time series can be affected by marginal mean and volatility dynamics and why copula parameters should be estimated using filtered pseudo-observations rather than the raw return series. The third part performs a Monte Carlo study to assess the finite sample properties of selected specifications of time-varying dependence. It concludes that estimation precision increases with sample size, while the complexity of recursive and score-driven models makes them computationally more demanding.
The empirical analysis uses data from five financial markets: stocks, silver, gold, Bitcoin, and an energy market. The results show that cross-market dependence is not constant over time. Different market pairs are characterized by different dependence specifications. This thesis uses Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC) to select the dependence specifications. Using AIC, GAS-type models are frequently selected, while BIC more often favours simpler constant or recursive specifications.
Finally, this thesis studies the minimum-variance portfolio allocation and CoVaR-based downside risk analysis by using the estimated dependence paths. The dynamic minimum-variance portfolio results show that time-varying dependence estimates can reduce portfolio volatility compared with an equal-weight portfolio. The ΔCoVaR analysis shows that downside risk transmission changes over time, differs by direction, and varies substantially across market pairs.
Overall, the thesis provides a framework for time-varying Gaussian copula models to analyse dynamic financial market dependence and its implications for portfolio and risk management.
...
The first part introduces the theoretical background on time series, copulas, maximum likelihood estimation, Generalized Autoregressive Score (GAS) updating, and GARCH marginal models. The second part uses simulation studies to explain why the observed behaviour of a time series can be affected by marginal mean and volatility dynamics and why copula parameters should be estimated using filtered pseudo-observations rather than the raw return series. The third part performs a Monte Carlo study to assess the finite sample properties of selected specifications of time-varying dependence. It concludes that estimation precision increases with sample size, while the complexity of recursive and score-driven models makes them computationally more demanding.
The empirical analysis uses data from five financial markets: stocks, silver, gold, Bitcoin, and an energy market. The results show that cross-market dependence is not constant over time. Different market pairs are characterized by different dependence specifications. This thesis uses Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC) to select the dependence specifications. Using AIC, GAS-type models are frequently selected, while BIC more often favours simpler constant or recursive specifications.
Finally, this thesis studies the minimum-variance portfolio allocation and CoVaR-based downside risk analysis by using the estimated dependence paths. The dynamic minimum-variance portfolio results show that time-varying dependence estimates can reduce portfolio volatility compared with an equal-weight portfolio. The ΔCoVaR analysis shows that downside risk transmission changes over time, differs by direction, and varies substantially across market pairs.
Overall, the thesis provides a framework for time-varying Gaussian copula models to analyse dynamic financial market dependence and its implications for portfolio and risk management.
I first consider the ideal case, where the marginal distributions are known. In this setting, the true copula observations are available directly. The main difficulty is that the test does not use only one fixed split point, but searches over many possible locations. Therefore, the likelihood approximations need to hold uniformly over all candidate change-points. To do this, I first control the score process and the observed information matrices uniformly over the candidate splits. These results are then used to describe the estimation errors of the maximum likelihood estimators. From there, the likelihood ratio statistic can be rewritten as a quadratic form involving a centered score process. This process converges to a standard multivariate Brownian bridge, which gives the limiting distribution of the maximized test statistic.
The feasible case is more difficult because the marginal distributions are unknown. The true copula observations are then replaced by rank-based pseudo-observations. For every possible change-point, the ranks are recomputed separately in the two resulting segments. Because of this, the score terms are no longer simple partial sums of independent observations. To handle this, I use the sequential localrank empirical copula process and combine it with a function-indexed integration-by-parts argument. This makes it possible to derive the weak limit of the local-rank score process and, from this, the limiting distribution of the feasible likelihood ratio statistic. The resulting limit still has a Brownian-bridge-type structure, but its covariance also reflects the additional effect of estimating the marginal distributions through ranks.
Finally, I study the finite-sample behaviour of the test through simulations using Gaussian and Student 𝑡 copulas. The results show a clear general pattern: the test becomes more powerful when the sample size increases or when the change in dependence becomes larger. Small changes are harder to detect, especially for the feasible procedure based on local ranks. At the same time, the difference between the ideal and feasible procedures becomes smaller for larger samples and stronger changes in dependence. Overall, the thesis develops the theoretical justification for a likelihood-ratio change-point test in copula models and studies how the procedure behaves in finite samples.
...
I first consider the ideal case, where the marginal distributions are known. In this setting, the true copula observations are available directly. The main difficulty is that the test does not use only one fixed split point, but searches over many possible locations. Therefore, the likelihood approximations need to hold uniformly over all candidate change-points. To do this, I first control the score process and the observed information matrices uniformly over the candidate splits. These results are then used to describe the estimation errors of the maximum likelihood estimators. From there, the likelihood ratio statistic can be rewritten as a quadratic form involving a centered score process. This process converges to a standard multivariate Brownian bridge, which gives the limiting distribution of the maximized test statistic.
The feasible case is more difficult because the marginal distributions are unknown. The true copula observations are then replaced by rank-based pseudo-observations. For every possible change-point, the ranks are recomputed separately in the two resulting segments. Because of this, the score terms are no longer simple partial sums of independent observations. To handle this, I use the sequential localrank empirical copula process and combine it with a function-indexed integration-by-parts argument. This makes it possible to derive the weak limit of the local-rank score process and, from this, the limiting distribution of the feasible likelihood ratio statistic. The resulting limit still has a Brownian-bridge-type structure, but its covariance also reflects the additional effect of estimating the marginal distributions through ranks.
Finally, I study the finite-sample behaviour of the test through simulations using Gaussian and Student 𝑡 copulas. The results show a clear general pattern: the test becomes more powerful when the sample size increases or when the change in dependence becomes larger. Small changes are harder to detect, especially for the feasible procedure based on local ranks. At the same time, the difference between the ideal and feasible procedures becomes smaller for larger samples and stronger changes in dependence. Overall, the thesis develops the theoretical justification for a likelihood-ratio change-point test in copula models and studies how the procedure behaves in finite samples.
We support our theoretical results with an empirical study of global financial assets, applying ARMA-GARCH filtering and conditioning on different markets to recover the most dominant modes of variation within a large dataset of CKT curves.
We find that all CKT curves can be approximately summarized by only two degrees of freedom: a constant baseline level, representing the average conditional Kendall's tau, and a parabolic component that contrasts extreme market conditions with more typical periods, capturing the contagion effects.
Ultimately, this framework provides a foundation for future applications, including evaluating the simplifying assumption. ...
We support our theoretical results with an empirical study of global financial assets, applying ARMA-GARCH filtering and conditioning on different markets to recover the most dominant modes of variation within a large dataset of CKT curves.
We find that all CKT curves can be approximately summarized by only two degrees of freedom: a constant baseline level, representing the average conditional Kendall's tau, and a parabolic component that contrasts extreme market conditions with more typical periods, capturing the contagion effects.
Ultimately, this framework provides a foundation for future applications, including evaluating the simplifying assumption.
This thesis investigates how such bias can be detected and corrected. When the magnitude of the bias is known, it can be adjusted for directly, restoring the reliability of the results. If the bias is not known exactly but is bounded within a certain range, two alternative strategies can still guarantee reliable conclusions, albeit at the cost of reduced sensitivity to small effects. However, if no information about the distortion is available, the situation becomes hopeless: no method can reliably distinguish genuine findings from random fluctuations.
The thesis therefore demonstrates that some knowledge about the degree of sampling imbalance is not merely beneficial, but essential for drawing meaningful statistical conclusions. ...
This thesis investigates how such bias can be detected and corrected. When the magnitude of the bias is known, it can be adjusted for directly, restoring the reliability of the results. If the bias is not known exactly but is bounded within a certain range, two alternative strategies can still guarantee reliable conclusions, albeit at the cost of reduced sensitivity to small effects. However, if no information about the distortion is available, the situation becomes hopeless: no method can reliably distinguish genuine findings from random fluctuations.
The thesis therefore demonstrates that some knowledge about the degree of sampling imbalance is not merely beneficial, but essential for drawing meaningful statistical conclusions.
Semiparametric dependence modeling in illiquid financial markets using multivariate zero‐inflated GARCH‐X models
Case study on Voluntary Carbon Credits in collaboration with Rabobank
The multivariate extension incorporates two different types of dependence. First, cross-dependence in trading activity is modeled using Markov networks applied to binary trading indicators. Second, cross-dependence in returns is analyzed using a copula-GARCH framework applied to residuals, with residuals corresponding to zero returns treated as undefined values. To make copula methods applicable to zero-inflated data, we introduce a joint probability integral transform approach. In this construction, the univariate marginals are defined conditional on the simultaneous trading activity of each asset, rather than conditioning only on each asset’s own trading activity. We prove that this method yields a consistent copula estimator when applied to the subset of observations where all assets have simultaneous non-zero trading activity. Dependence is quantified using both unconditional and conditional Kendall’s tau, estimated via kernel-based methods.
Theoretical results include consistency and asymptotic normality of the quasi-maximum likelihood estimator for the multivariate model under stationary covariates. The empirical study covers seven voluntary carbon credits and six conventional financial assets. We find no significant dependence between carbon credits and conventional assets, but observe strong correlation within the carbon market, especially among nature-based credits. Results suggest that voluntary carbon markets may operate independently of more liquid assets and are influenced by peer pricing due to a lack of standardization. ...
The multivariate extension incorporates two different types of dependence. First, cross-dependence in trading activity is modeled using Markov networks applied to binary trading indicators. Second, cross-dependence in returns is analyzed using a copula-GARCH framework applied to residuals, with residuals corresponding to zero returns treated as undefined values. To make copula methods applicable to zero-inflated data, we introduce a joint probability integral transform approach. In this construction, the univariate marginals are defined conditional on the simultaneous trading activity of each asset, rather than conditioning only on each asset’s own trading activity. We prove that this method yields a consistent copula estimator when applied to the subset of observations where all assets have simultaneous non-zero trading activity. Dependence is quantified using both unconditional and conditional Kendall’s tau, estimated via kernel-based methods.
Theoretical results include consistency and asymptotic normality of the quasi-maximum likelihood estimator for the multivariate model under stationary covariates. The empirical study covers seven voluntary carbon credits and six conventional financial assets. We find no significant dependence between carbon credits and conventional assets, but observe strong correlation within the carbon market, especially among nature-based credits. Results suggest that voluntary carbon markets may operate independently of more liquid assets and are influenced by peer pricing due to a lack of standardization.
A Vine Copula Approach for Portfolio Optimisation
Exploring the Effect of Copulas and Vine Models on Optimal Investment Allocation of Stock Index Returns
In this thesis, we propose a novel “simulation-based method” for statistical inference of low-frequency time series that result from the aggregation of a higher-frequency time series over a period of time. We start by estimating the distribution of this higher-frequency process. We then simulate a large number of paths from this estimated distribution. By independently aggregating each simulated path, we generate corresponding low-frequency data. This provides us with a large simulated dataset of the low-frequency process, which enables us to apply estimation procedures and bypass the limitations posed by the shortage of original low-frequency data.
We also provide a theoretical framework and propose three families of estimators constructed from the estimated higher-frequency distribution, analyzing their properties under additional assumptions. Through a comprehensive simulation study, we compare the simulation-based method with the traditional direct method across different scenarios and objectives. While our study focuses on the marginal distributions of low-frequency processes, the simulation-based method’s applicability extends to joint distributions across multiple time points. This research offers a robust method for parameter estimation when faced with limited low-frequency data. ...
In this thesis, we propose a novel “simulation-based method” for statistical inference of low-frequency time series that result from the aggregation of a higher-frequency time series over a period of time. We start by estimating the distribution of this higher-frequency process. We then simulate a large number of paths from this estimated distribution. By independently aggregating each simulated path, we generate corresponding low-frequency data. This provides us with a large simulated dataset of the low-frequency process, which enables us to apply estimation procedures and bypass the limitations posed by the shortage of original low-frequency data.
We also provide a theoretical framework and propose three families of estimators constructed from the estimated higher-frequency distribution, analyzing their properties under additional assumptions. Through a comprehensive simulation study, we compare the simulation-based method with the traditional direct method across different scenarios and objectives. While our study focuses on the marginal distributions of low-frequency processes, the simulation-based method’s applicability extends to joint distributions across multiple time points. This research offers a robust method for parameter estimation when faced with limited low-frequency data.
Estimators for the population mean and variance for stratified sampling
The search for unbiased estimators in a suboptimal sample
In this thesis, different estimators for the true population mean and variance are defined and examined in terms of bias and variance in the case of two subgroups. Weighing the measurements according to the true proportions creates unbiased estimators for both the mean and the variance. These unbiased estimators are compared with other, biased, estimators, including naive ones in which the influence of different subgroups is not taken into account. The naive estimators are not only biased, they also have a variance of the same order as the unbiased ones. When the true proportions are not available, one can only take a guess. A guess lying close to the true proportions leads to a smaller bias and therefore a better estimator. This underlines the importance of obtaining sufficient knowledge about the population. ...
In this thesis, different estimators for the true population mean and variance are defined and examined in terms of bias and variance in the case of two subgroups. Weighing the measurements according to the true proportions creates unbiased estimators for both the mean and the variance. These unbiased estimators are compared with other, biased, estimators, including naive ones in which the influence of different subgroups is not taken into account. The naive estimators are not only biased, they also have a variance of the same order as the unbiased ones. When the true proportions are not available, one can only take a guess. A guess lying close to the true proportions leads to a smaller bias and therefore a better estimator. This underlines the importance of obtaining sufficient knowledge about the population.
Walking on Powered VR Shoes to Virtual Reality Motion
A User Experience Evaluation
Hawkes Processes in Large-Scale Service Systems
Improving service management at ING
In more detail, the data from the monitoring stream consists (among other things) of a message and a time stamp. Moreover, the monitoring data stream of this bank consists of two natures of information. These natures are either automatically generated warnings in the form of events or unplanned outages, referred to as incidents. The events and incidents are referred to as arrivals. As a first requirement to obtain better granularity, both event and incident messages with similar semantics should be grouped together. To this extent, the message component from each arrival is transformed into a numerical vector, the dimension of the obtained vector is reduced, and the collection of vectors is clustered. Once the individual arrival from the IT monitoring data stream is attached to a cluster based on their message component, the arrival is assigned a mark. This mark consists of a combination of the assigned cluster, the nature, and three different levels of service from the IT architecture on which the arrival occurred.
From a mathematical point of view, we can now view the monitoring data stream from different levels of service as a marked point process. Our primary focus centers on a specific category of marked point processes, known as marked Hawkes processes. Given the marked Hawkes process, we assume that each arrival from the IT monitoring data stream results in an instantaneous increase in the probability of some other arrivals in the near future. From here, we estimate the excitation matrix, representing the instantaneous increases among all assigned marks. Once the estimated excitation matrix is obtained, we decompose it into the different levels of service as defined within the mark. In particular, the decomposition has been performed through means of hierarchical linear models. Finally, the decomposition resulted in a comprehensive overview of the excitation behavior in large-scale service systems. This overview can directly be incorporated into the field of Software Architecture in order to uncover associations within complex IT infrastructures. ...
In more detail, the data from the monitoring stream consists (among other things) of a message and a time stamp. Moreover, the monitoring data stream of this bank consists of two natures of information. These natures are either automatically generated warnings in the form of events or unplanned outages, referred to as incidents. The events and incidents are referred to as arrivals. As a first requirement to obtain better granularity, both event and incident messages with similar semantics should be grouped together. To this extent, the message component from each arrival is transformed into a numerical vector, the dimension of the obtained vector is reduced, and the collection of vectors is clustered. Once the individual arrival from the IT monitoring data stream is attached to a cluster based on their message component, the arrival is assigned a mark. This mark consists of a combination of the assigned cluster, the nature, and three different levels of service from the IT architecture on which the arrival occurred.
From a mathematical point of view, we can now view the monitoring data stream from different levels of service as a marked point process. Our primary focus centers on a specific category of marked point processes, known as marked Hawkes processes. Given the marked Hawkes process, we assume that each arrival from the IT monitoring data stream results in an instantaneous increase in the probability of some other arrivals in the near future. From here, we estimate the excitation matrix, representing the instantaneous increases among all assigned marks. Once the estimated excitation matrix is obtained, we decompose it into the different levels of service as defined within the mark. In particular, the decomposition has been performed through means of hierarchical linear models. Finally, the decomposition resulted in a comprehensive overview of the excitation behavior in large-scale service systems. This overview can directly be incorporated into the field of Software Architecture in order to uncover associations within complex IT infrastructures.
We have also developed an R package that facilitates the application of machine learning methods. This package leverages a range of machine learning models, including decision trees, neural networks, random forests, and bagging neural networks. Through systematic learning of the intricate relationships between the covariates and the binary variables, we effectively estimate conditional CDFs.
To enhance the accuracy and reliability of the estimated CDFs, we incorporate a rearrangement technique which transforms the estimated functions into monotonic representations, aligning them more closely with the target CDFs and mitigating potential inconsistencies [6].
Through simulations, we evaluate the performance of the estimation approach under various scenarios and assess the impact of sample size and correlation on estimation accuracy, using Mean Integrated Squared Error as a key performance metric. The results demonstrate the effectiveness and robustness of the methodology in estimating conditional CDFs, providing a valuable tool for capturing complex dependencies in multivariate data, with potential applications in risk assessment, finance, and environmental modeling.
...
We have also developed an R package that facilitates the application of machine learning methods. This package leverages a range of machine learning models, including decision trees, neural networks, random forests, and bagging neural networks. Through systematic learning of the intricate relationships between the covariates and the binary variables, we effectively estimate conditional CDFs.
To enhance the accuracy and reliability of the estimated CDFs, we incorporate a rearrangement technique which transforms the estimated functions into monotonic representations, aligning them more closely with the target CDFs and mitigating potential inconsistencies [6].
Through simulations, we evaluate the performance of the estimation approach under various scenarios and assess the impact of sample size and correlation on estimation accuracy, using Mean Integrated Squared Error as a key performance metric. The results demonstrate the effectiveness and robustness of the methodology in estimating conditional CDFs, providing a valuable tool for capturing complex dependencies in multivariate data, with potential applications in risk assessment, finance, and environmental modeling.
of the conditional dependence relates to characteristics of an asset such as geographical properties and type of asset. ...
of the conditional dependence relates to characteristics of an asset such as geographical properties and type of asset.
...
Financial Stock Market Modeling and the COVID-19 crisis
Has COVID-19 structurally changed the dynamics of the stock market?