Circular Image

J. Söhl

info

Please Note

17 records found

Preprint (2025) - Koen van Arem, Floris Goes-Smit, Jakob Söhl
Transfers in professional football (soccer) are risky investments because of the large transfer fees and high risks involved. Although data-driven models can be used to improve transfer decisions, existing models focus on describing players' historical progress, leaving their future performance unknown. Moreover, recent developments have called for the use of explainable models combined with uncertainty quantification of predictions. This paper assesses explainable machine learning models based on predictive accuracy and uncertainty quantification methods for the prediction of the future development in quality and transfer value of professional football players. Using a historical data set of data-driven indicators describing player quality and the transfer value of a football player, the models are trained to forecast player quality and player value one year ahead. These two prediction problems demonstrate the efficacy of tree-based models, particularly random forest and XGBoost, in making accurate predictions. In general, the random forest model is found to be the most suitable model because it provides accurate predictions as well as an uncertainty quantification method that naturally arises from the bagging procedure of the random forest model. Additionally, our research shows that the development of player performance contains nonlinear patterns and interactions between variables, and that time series information can provide useful information for the modeling of player performance metrics. Our research provides models to help football clubs make more informed, data-driven transfer decisions by forecasting player quality and transfer value. ...
Book chapter (2025) - K.W. van Arem, Jakob Söhl, Mirjam Bruinsma, Geurt Jongbloed
With an average football (soccer) match recording over 3,000 on-ball events, effective use of this event data is essential for practitioners at football clubs to obtain meaningful insights. Models can extract more information from this data, and explainable methods can make them more accessible to practitioners. The Expected Threat model has been praised for its explainability and offers an accessible option. However, selecting the grid size is a challenging key design choice that has to be made when applying the Expected Threat model. Using a finer grid leads to a more flexible model that can better distinguish between different situations, but the accuracy of the estimates deteriorates with a more flexible model. Consequently, practitioners face challenges in balancing the trade-off between model flexibility and model accuracy.
In this study, the Expected Threat model \added{is analyzed} from a theoretical perspective and simulations are performed based on the Markov chain of the model to examine its behavior in practice. Our theoretical results establish an upper bound on the error of the Expected Threat model for different flexibilities. Based on the simulations, a more accurate characterization of the model’s error is provided, improving over the theoretical bound. Finally, these insights are converted into a practical rule of thumb to help practitioners choose the right balance between the model flexibility and the desired accuracy of the Expected Threat model. ...
Journal article (2025) - K.W. van Arem, Floris Goes-Smit, J. Söhl
Featured Application: This paper studies what models are most suitable for forecasting future values of player performance metrics in association football (soccer). The resulting forecast statistics find applications in team management and player scouting at football clubs. As transfer decisions concern whether a player should play for a club in the future, the predictions of future performance metrics offer a forward-looking improvement over the traditional backward-looking assessments. Transfers in professional football (soccer) are risky investments because of the large transfer fees and high risks involved. Although data-driven models can be used to improve transfer decisions, existing models focus on describing players’ historical progress, leaving their future performance unknown. Moreover, recent developments have called for the use of explainable models combined with methods for uncertainty quantification of predictions to improve applicability for practitioners. This paper assesses explainable machine learning models in a practitioner-oriented way for the prediction of the future development in quality and transfer value of professional football players. To this end, the methods for uncertainty quantification are studied through the literature. The predictive accuracy is studied by training the models to predict the quality and value of players one year ahead, equivalent to one season. This is carried out by training them on two data sets containing data-driven indicators describing the player quality and player value in historical settings. In this paper, the random forest model is found to be the most suitable model because it provides accurate predictions as well as an uncertainty quantification method that naturally arises from the bagging procedure of the random forest model. Additionally, this research shows that the development of player performance contains nonlinear patterns and interactions between variables, and that time series information can provide useful information for the modeling of player performance metrics. The resulting models can help football clubs make more informed, data-driven transfer decisions by forecasting player quality and transfer value. ...
Journal article (2024) - Michał G. Ciszewski, Jakob Söhl, Ton Leenen, Bart van Trigt, Geurt Jongbloed
Often the question arises whether (Formula presented.) can be predicted based on (Formula presented.) using a certain model. Especially for highly flexible models such as neural networks one may ask whether a seemingly good prediction is actually better than fitting pure noise or whether it has to be attributed to the flexibility of the model. This paper proposes a rigorous permutation test to assess whether the prediction is better than the prediction of pure noise. The test avoids any sample splitting and is based instead on generating new pairings of (Formula presented.). It introduces a new formulation of the null hypothesis and rigorous justification for the test, which distinguishes it from the previous literature. The theoretical findings are applied both to simulated data and to sensor data of tennis serves in an experimental context. The simulation study underscores how the available information affects the test. It shows that the less informative the predictors, the lower the probability of rejecting the null hypothesis of fitting pure noise and emphasizes that detecting weaker dependence between variables requires a sufficient sample size. ...
Journal article (2023) - Michał Ciszewski, Jakob Söhl, Geurt Jongbloed
The past decade has seen an increased interest in human activity recognition based on sensor data. Most often, the sensor data come unannotated, creating the need for fast labelling methods. For assessing the quality of the labelling, an appropriate performance measure has to be chosen. Our main contribution is a novel post-processing method for activity recognition. It improves the accuracy of the classification methods by correcting for unrealistic short activities in the estimate. We also propose a new performance measure, the Locally Time-Shifted Measure (LTS measure), which addresses uncertainty in the times of state changes. The effectiveness of the post-processing method is evaluated, using the novel LTS measure, on the basis of a simulated dataset and a real application on sensor data from football. The simulation study is also used to discuss the choice of the parameters of the post-processing method and the LTS measure. ...
Journal article (2020) - Dirk Paulsen, Jakob Söhl
When the in-sample Sharpe ratio is obtained by optimizing over a k-dimensional parameter space, it is a biased estimator for what can be expected on unseen data (out-of-sample). We derive (1) an unbiased estimator adjusting for both sources of bias: noise fit and estimation error. We then show (2) how to use the adjusted Sharpe ratio as model selection criterion analogously to the Akaike Information Criterion (AIC). Selecting a model with the highest adjusted Sharpe ratio selects the model with the highest estimated out-of-sample Sharpe ratio in the same way as selection by AIC does for the log-likelihood as a measure of fit. ...
Journal article (2019) - Richard Nickl, Jakob Söhl
We study nonparametric Bayesian statistical inference for the parameters governing a pure jump process of the form (Formula Presented) where N(t) is a standard Poisson process of intensity λ, and Z k are drawn i.i.d. from jump measure μ. A high-dimensional wavelet series prior for the Lévy measure ν = λμ is devised and the posterior distribution arises from observing discrete samples Y Δ, Y , …, Y at fixed observation distance Δ, giving rise to a nonlinear inverse inference problem. We derive contraction rates in uniform norm for the posterior distribution around the true Lévy density that are optimal up to logarithmic factors over Hölder classes, as sample size n increases. We prove a functional Bernstein–von Mises theorem for the distribution functions of both μ and ν, as well as for the intensity λ, establishing the fact that the posterior distribution is approximated by an infinite-dimensional Gaussian measure whose covariance structure is shown to attain the information lower bound for this inverse problem. As a consequence posterior based inferences, such as nonparametric credible sets, are asymptotically valid and optimal from a frequentist point of view. ...
Journal article (2018) - T. van Groeningen, H. Driessen, J. Söhl, Robert Voûte
When the fire brigade arrives at a burning building, it is of vital importance that people who are still inside can quickly be found. Smart buildings should be able to expose this location data to the fire brigade working in a smart city. In this paper the feasibility is researched of using ultrasonic sound sensors for human presence detection in smoke-filled spaces. This type of sensor could assist the fire brigade when evacuating a large building by directing them to the places where their help is most needed. The advantage of ultrasonic sound over other sensors or cameras is that its signal is able to pierce through smoke, does not require badges or other wearable devices and introduces little privacy and security risks. In addition, ultrasonic sensors are very inexpensive making it possible to equip every room of a building with an ultrasonic presence detector. In this research both a preliminary ultrasound measuring device and signal processing algorithm have been designed. Testing results show that the walking movement of a person in an indoor area can be detected with the combination of the sensor and the algorithms. In addition, tests of the signal strength in smoke have shown that ultrasound is capable of “looking through” the smoke. The algorithm based on a particle filter allows for more information to be extracted from the relatively simple sensor signal by detecting human walking movement specifically and opens up the way for an ultrasound based indoor positioning system that can be used in emergency situations. ...
A method is proposed for calculating the shear viscosity of a liquid from finite-size effects of self-diffusion coefficients in Molecular Dynamics simulations. This method uses the difference in the self-diffusivities, computed from at least two system sizes, and an analytic equation to calculate the shear viscosity. To enable the efficient use of this method, a set of guidelines is developed. The most efficient number of system sizes is two and the large system is at least four times the small system. The number of independent simulations for each system size should be assigned in such a way that 50%-70% of the total available computational resources are allocated to the large system. We verified the method for 250 binary and 26 ternary Lennard-Jones systems, pure water, and an ionic liquid ([Bmim][Tf2N]). The computed shear viscosities are in good agreement with viscosities obtained from equilibrium Molecular Dynamics simulations for all liquid systems far from the critical point. Our results indicate that the proposed method is suitable for multicomponent mixtures and highly viscous liquids. This may enable the systematic screening of the viscosities of ionic liquids and deep eutectic solvents. ...
Journal article (2017) - Richard Nickl, Jakob Söhl
We consider nonparametric Bayesian inference in a reflected diffusionmodel dXt = b(Xt)dt + σ(Xt)dWt , with discretely sampled observationsX0,X, . . . , Xn. We analyse the nonlinear inverse problem correspondingto the “low frequency sampling” regime where >0 is fixed and n→∞.A general theorem is proved that gives conditions for prior distributions on the diffusion coefficient σ and the drift function b that ensure minimaxoptimal contraction rates of the posterior distribution over Hölder–Sobolevsmoothness classes. These conditions are verified for natural examples ofnonparametric random wavelet series priors. For the proofs, we derive newconcentration inequalities for empirical processes arising from discretely observeddiffusions that are of independent interest. ...
Journal article (2016) - Richard Nickl, Markus Reiß, Jakob Söhl, Mathias Trabs
Donsker-type functional limit theorems are proved for empirical processes arising from discretely sampled increments of a univariate Lévy process. In the asymptotic regime the sampling frequencies increase to infinity and the limiting object is a Gaussian process that can be obtained from the composition of a Brownian motion with a covariance operator determined by the Lévy measure. The results are applied to derive the asymptotic distribution of natural estimators for the distribution function of the Lévy jump measure. As an application we deduce Kolmogorov–Smirnov type tests and confidence bands. ...

Estimating the invariant measure and the drift

Journal article (2016) - Jakob Söhl, Mathias Trabs
As a starting point we prove a functional central limit theorem for estimators of the invariant measure of a geometrically ergodic Harris-recurrent Markov chain in a multi-scale space. This allows to construct confidence bands for the invariant density with optimal (up to undersmoothing) L-diameter by using wavelet projection estimators. In addition our setting applies to the drift estimation of diffusions observed discretely with fixed observation distance. We prove a functional central limit theorem for estimators of the drift function and finally construct adaptive confidence bands for the drift by using a completely data-driven estimator. ...
Journal article (2015) - Jakob Söhl
We consider the Grenander estimator that is the maximum likelihood estimator for non-increasing densities. We prove uniform central limit theorems for certain subclasses of bounded variation functions and for Hölder balls of smoothness s >1/2. We do not assume that the density is differentiable or continuous. The proof can be seen as an adaptation of the method for the parametric maximum likelihood estimator to the nonparametric setting. Since nonparametric maximum likelihood estimators lie on the boundary, the derivative of the likelihood cannot be expected to equal zero as in the parametric case. Nevertheless, our proofs rely on the fact that the derivative of the likelihood can be shown to be small at the maximum likelihood estimator. ...

Confidence intervals and empirical results

Journal article (2014) - Jakob Söhl, Mathias Trabs
Observing prices of European put and call options, we calibrate exponential Lévy models nonparametrically. We discuss the efficient implementation of the spectral estimation procedures for Lévy models of finite jump activity as well as for self-decomposable Lévy models. Based on finite sample variances, confidence intervals are constructed for the volatility, for the drift and, pointwise, for the jump density. As demonstrated by simulations, these intervals perform well in terms of size and coverage probabilities. We compare the performance of the procedures for finite and infinite jump activity based on options on the German DAX index and find that both methods achieve good calibration results. The stability of the finite activity model is studied when the option prices are observed in a sequence of trading days. ...
Journal article (2014) - Jakob Söhl
Confidence intervals and joint confidence sets are constructed for the nonparametric calibration of exponential Lévy models based on prices of European options. To this end, we show joint asymptotic normality in the spectral calibration method for the estimators of the volatility, the drift, the jump intensity and the Lévy density at finitely many points. ...
Journal article (2012) - Jakob Söhl, Mathias Trabs
We estimate linear functionals in the classical deconvolution problem by kernel estimators. We obtain a uniform central limit theorem with √n-rate on the assumption that the smoothness of the functionals is larger than the ill-posedness of the problem, which is given by the polynomial decay rate of the characteristic function of the error. The limit distribution is a generalized Brownian bridge with a covariance structure that depends on the characteristic function of the error and on the functionals. The proposed estimators are optimal in the sense of semiparametric efficiency. The class of linear functionals is wide enough to incorporate the estimation of distribution functions. The proofs are based on smoothed empirical processes and mapping properties of the deconvolution operator. ...
Journal article (2010) - Jakob Söhl
This paper studies polar sets for anisotropic Gaussian random fields, i.e. sets which a Gaussian random field does not hit almost surely. The main assumptions are that the eigenvalues of the covariance matrix are bounded from below and that the canonical metric associated with the Gaussian random field is dominated by an anisotropic metric. We deduce an upper bound for the hitting probabilities and conclude that sets with small Hausdorff dimension are polar. Moreover, the results allow for a translation of the Gaussian random field by a random field, that is independent of the Gaussian random field and whose sample functions are of bounded Hölder norm. ...