Leave-Multiple-Out Informal Benchmarking

A Simulation Study on the Impact of Sample Size

Bachelor Thesis (2026)
Author(s)

M.P. Czerwinska (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Contributor(s)

J.H. Krijthe – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)

M. Havelka – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)

A. Anand – Graduation committee member (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Faculty
Electrical Engineering, Mathematics and Computer Science
More Info
expand_more
Publication Year
2026
Language
English
Graduation Date
23-06-2026
Awarding Institution
Delft University of Technology
Project
CSE3000 Research Project
Programme
Computer Science and Engineering
Faculty
Electrical Engineering, Mathematics and Computer Science
Downloads counter
62
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

In observational studies, establishing robust causal relationships is frequently challenged by the pres- ence of hidden variables that distort the true rela- tionships within the data. While sensitivity analysis offers a formal framework to assess how vulnera- ble a study’s conclusions are to such hidden biases, its widespread adoption is often hindered by com- plex mathematical assumptions and interpretation hurdles. To make these tools more intuitive, re- searchers frequently employ informal benchmark- ing—a technique that uses the measured explana- tory power of observed features as a baseline to calibrate the hypothetical strength a hidden variable would need to overturn a conclusion. While widely used, the behavior of these benchmarks across varying sample scales lacks a systematic evaluation under different structural conditions. This thesis addresses this gap by investigating the sample size dynamics within the Leave-Multiple-Out (LMO) informal benchmarking framework. Leveraging an extended simulation grid across 18 structural sce- narios and sample sizes ranging from N = 20 to N = 50, 000, we demonstrate that increasing data volume does not produce a uniform trajec- tory. Instead, the benchmark follows two entirely different paths depending on the predictive power of the available features. In weak or uncorrelated settings, increasing the sample size acts as a vari- ance reducer, eliminating small-sample overfitting (N ≤ 200) and stabilizing scores at a conserva- tive baseline. In contrast, in strong or highly cor- related settings where true propensity scores span a wide range, scaling the sample size grants well- specified models the statistical power to accurately map the extreme tails of the distribution. Be- cause the non-linear mechanics of odds ratio cal- culations are highly sensitive to probabilities near 0 or 1, these captured tail cases drive the bench- mark scores steadily upward. While the framework successfully reflects model quality as the sample size increases, these dynamics prompt a theoreti- cal warning: the benchmark inherently rewards the generation of extreme probabilities, rendering it po- tentially vulnerable to overparameterized or uncal- ibrated estimators.

Files

License info not available