JC
J. Cai
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
5 records found
1
In dit onderzoek is een Monte Carlo algoritme beschreven voor het schatten van de dominante eigenwaarde en bijbehorende eigenvector van een niet-negatieve, irreducibele matrix. Naast de werking van het Monte Carlo algoritme is ook de convergentiesnelheid onderzocht. Tevens zijn een aantal simulaties in de praktijk uitgevoerd voor zowel het Monte Carlo algoritme als de power methode en aan de hand hiervan bekeken of het Monte Carlo algoritme een goede concurrent is van de power methode.
...
In dit onderzoek is een Monte Carlo algoritme beschreven voor het schatten van de dominante eigenwaarde en bijbehorende eigenvector van een niet-negatieve, irreducibele matrix. Naast de werking van het Monte Carlo algoritme is ook de convergentiesnelheid onderzocht. Tevens zijn een aantal simulaties in de praktijk uitgevoerd voor zowel het Monte Carlo algoritme als de power methode en aan de hand hiervan bekeken of het Monte Carlo algoritme een goede concurrent is van de power methode.
Time series analysis is used to predict future behaviour of processes and is widely used in the finance sector. In this paper we will analyse the modelling of multivariate time series of financial data using vector autoregressive processes. The goal is that the reader will understand the presented models and could theoretically perform time series analysis by himself. Two specific models will be explained: the Vector Autoregressive model (VAR model) and the Vector Error Correction Model (VECM). We will describe various methods to analyse multivariate time series using these models, such as forecasting the process, variance decomposition of the forecast error, causality analysis and impulse response analysis. Examples of these models and analysis methods will be presented and investigated. Finally, we will perform a time series analysis with these models on Dutch indices and stock data. We conclude that real-world data often does not fit the VAR model and VECM requirements and that further improved models should be considered as well.
...
Time series analysis is used to predict future behaviour of processes and is widely used in the finance sector. In this paper we will analyse the modelling of multivariate time series of financial data using vector autoregressive processes. The goal is that the reader will understand the presented models and could theoretically perform time series analysis by himself. Two specific models will be explained: the Vector Autoregressive model (VAR model) and the Vector Error Correction Model (VECM). We will describe various methods to analyse multivariate time series using these models, such as forecasting the process, variance decomposition of the forecast error, causality analysis and impulse response analysis. Examples of these models and analysis methods will be presented and investigated. Finally, we will perform a time series analysis with these models on Dutch indices and stock data. We conclude that real-world data often does not fit the VAR model and VECM requirements and that further improved models should be considered as well.
The primary goal of this report is to provide a general overview of offline change-point literature as it is known today. Change-point methods are important statistical problems, where we are interested in determining whenever a certain data-set changes in structure. Furthermore, the term "off-line" is meant to indicate that the data itself is already known, whereas on the other hand we have "on-line" methods which deal with situations where new data is yet being received during localisation of the change-points. In this report we mainly consider off-line methods, as we feel off-line methods provide a more friendly introduction into change-point analysis and on-line methods are in principle just an extension of their off-line counterparts. First off, these off-line change-point methods are considered under different assumptions (parametric, non-parametric). In each case, we treat a solution to the change-point problem under different models(normal and gamma model, mean or varianche change etc.). Eventually we shall also treat some widely used algorithms, meant to extend the problem into the localisation of multiple change-points.indent Aside from theoretical considerations, an equally important part of this report will be focused on empirical results. Both for the statistics as algorithms will there be a performance study where the different methods will be empirically assessed and compared under different models, namely the robustness against "outliers"(extreme data-values) will be investigated. So to complement the primary goal, we will also focus on the following two subgoals:
1) Assessment and comparison of different change-point models
2) Evaluating and improving robustness against outliers
...
1) Assessment and comparison of different change-point models
2) Evaluating and improving robustness against outliers
...
The primary goal of this report is to provide a general overview of offline change-point literature as it is known today. Change-point methods are important statistical problems, where we are interested in determining whenever a certain data-set changes in structure. Furthermore, the term "off-line" is meant to indicate that the data itself is already known, whereas on the other hand we have "on-line" methods which deal with situations where new data is yet being received during localisation of the change-points. In this report we mainly consider off-line methods, as we feel off-line methods provide a more friendly introduction into change-point analysis and on-line methods are in principle just an extension of their off-line counterparts. First off, these off-line change-point methods are considered under different assumptions (parametric, non-parametric). In each case, we treat a solution to the change-point problem under different models(normal and gamma model, mean or varianche change etc.). Eventually we shall also treat some widely used algorithms, meant to extend the problem into the localisation of multiple change-points.indent Aside from theoretical considerations, an equally important part of this report will be focused on empirical results. Both for the statistics as algorithms will there be a performance study where the different methods will be empirically assessed and compared under different models, namely the robustness against "outliers"(extreme data-values) will be investigated. So to complement the primary goal, we will also focus on the following two subgoals:
1) Assessment and comparison of different change-point models
2) Evaluating and improving robustness against outliers
1) Assessment and comparison of different change-point models
2) Evaluating and improving robustness against outliers
Master thesis
(2019)
-
Attila Borsos, Marjan Hagenzieker, Haneen Farah, Juanjuan Cai, Aliaksei Laureshyn
The most common way to evaluate traffic safety is investigating the occurrence and severity of crashes using historical data. This approach however has a number of limitations, the most important of which is probably its reactive nature. An alternative method using non-crash events has gained a lot of attention recently, especially thanks to the rapid improvement of sensing technologies. By gathering trajectory data and calculating various Surrogate Measures of Safety it has become possible to analyse safety without waiting for accidents to happen. Using these indicators combined with Extreme Value Theory (EVT) one can estimate the probability of crashes as extreme (unobserved) events. The primary goal of this thesis is to contribute to the research that has been done so far on the application of Extreme Value Theory to Surrogate Measures for traffic safety analysis. Research questions seek for answers to what we can learn from applying univariate EVT using indicators describing collision course and crossing course interactions, and how we can predict nearness to collision and severity using bivariate EVT models.
...
The most common way to evaluate traffic safety is investigating the occurrence and severity of crashes using historical data. This approach however has a number of limitations, the most important of which is probably its reactive nature. An alternative method using non-crash events has gained a lot of attention recently, especially thanks to the rapid improvement of sensing technologies. By gathering trajectory data and calculating various Surrogate Measures of Safety it has become possible to analyse safety without waiting for accidents to happen. Using these indicators combined with Extreme Value Theory (EVT) one can estimate the probability of crashes as extreme (unobserved) events. The primary goal of this thesis is to contribute to the research that has been done so far on the application of Extreme Value Theory to Surrogate Measures for traffic safety analysis. Research questions seek for answers to what we can learn from applying univariate EVT using indicators describing collision course and crossing course interactions, and how we can predict nearness to collision and severity using bivariate EVT models.
Over the past three decades, the singular value decomposition has been increasingly used for various big data applications. As it allows for rank reduction of the input data matrix, it is not only able to compress the information contained, but can even reveal underlying patterns in the data through feature identification. This thesis explores algorithms for large-scale SVD calculation, and uses these to demonstrate how the SVD can be applied to a variety of fields, including information retrieval, recommender systems and image processing.
The algorithms discussed are Golub-Kahan-Lanczos bidiagonalization, randomized SVD and block power SVD. Each algorithm is implemented in Matlab and both error and time taken by each algorithm are compared. We find that the block power SVD is very effective, especially when only the truncated SVD is required. Due to its simplicity, speed and relatively small error for low-rank matrix approximation, it is an ideal method for the applications discussed in this thesis.
We show how the SVD can be used for information retrieval, through Latent Semantic Indexing. The method is tested on the Time collection and we find that the SVD removes much of the noise present in the data and solves the issues of synonymy and polysemy. Then, SVD-based algorithms for recommender systems are presented. We implement a basic SVD algorithm called Average Rating Filling, and a (biased) stochastic gradient descent algorithm, which was developed for the Netflix recommender-system prize. These are tested on the Movielens 100k dataset, resulting in the best performance by biased stochastic gradient descent. Finally, the SVD is used for image compression and we find that, while not very useful for face recognition, the SVD could provide a time- and space-efficient method for searching through an image database for similar images. ...
The algorithms discussed are Golub-Kahan-Lanczos bidiagonalization, randomized SVD and block power SVD. Each algorithm is implemented in Matlab and both error and time taken by each algorithm are compared. We find that the block power SVD is very effective, especially when only the truncated SVD is required. Due to its simplicity, speed and relatively small error for low-rank matrix approximation, it is an ideal method for the applications discussed in this thesis.
We show how the SVD can be used for information retrieval, through Latent Semantic Indexing. The method is tested on the Time collection and we find that the SVD removes much of the noise present in the data and solves the issues of synonymy and polysemy. Then, SVD-based algorithms for recommender systems are presented. We implement a basic SVD algorithm called Average Rating Filling, and a (biased) stochastic gradient descent algorithm, which was developed for the Netflix recommender-system prize. These are tested on the Movielens 100k dataset, resulting in the best performance by biased stochastic gradient descent. Finally, the SVD is used for image compression and we find that, while not very useful for face recognition, the SVD could provide a time- and space-efficient method for searching through an image database for similar images. ...
Over the past three decades, the singular value decomposition has been increasingly used for various big data applications. As it allows for rank reduction of the input data matrix, it is not only able to compress the information contained, but can even reveal underlying patterns in the data through feature identification. This thesis explores algorithms for large-scale SVD calculation, and uses these to demonstrate how the SVD can be applied to a variety of fields, including information retrieval, recommender systems and image processing.
The algorithms discussed are Golub-Kahan-Lanczos bidiagonalization, randomized SVD and block power SVD. Each algorithm is implemented in Matlab and both error and time taken by each algorithm are compared. We find that the block power SVD is very effective, especially when only the truncated SVD is required. Due to its simplicity, speed and relatively small error for low-rank matrix approximation, it is an ideal method for the applications discussed in this thesis.
We show how the SVD can be used for information retrieval, through Latent Semantic Indexing. The method is tested on the Time collection and we find that the SVD removes much of the noise present in the data and solves the issues of synonymy and polysemy. Then, SVD-based algorithms for recommender systems are presented. We implement a basic SVD algorithm called Average Rating Filling, and a (biased) stochastic gradient descent algorithm, which was developed for the Netflix recommender-system prize. These are tested on the Movielens 100k dataset, resulting in the best performance by biased stochastic gradient descent. Finally, the SVD is used for image compression and we find that, while not very useful for face recognition, the SVD could provide a time- and space-efficient method for searching through an image database for similar images.
The algorithms discussed are Golub-Kahan-Lanczos bidiagonalization, randomized SVD and block power SVD. Each algorithm is implemented in Matlab and both error and time taken by each algorithm are compared. We find that the block power SVD is very effective, especially when only the truncated SVD is required. Due to its simplicity, speed and relatively small error for low-rank matrix approximation, it is an ideal method for the applications discussed in this thesis.
We show how the SVD can be used for information retrieval, through Latent Semantic Indexing. The method is tested on the Time collection and we find that the SVD removes much of the noise present in the data and solves the issues of synonymy and polysemy. Then, SVD-based algorithms for recommender systems are presented. We implement a basic SVD algorithm called Average Rating Filling, and a (biased) stochastic gradient descent algorithm, which was developed for the Netflix recommender-system prize. These are tested on the Movielens 100k dataset, resulting in the best performance by biased stochastic gradient descent. Finally, the SVD is used for image compression and we find that, while not very useful for face recognition, the SVD could provide a time- and space-efficient method for searching through an image database for similar images.