Md
M. de Koning
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
2 records found
1
Master thesis
(2025)
-
M. de Koning, A. Panichella, Pouria Derakhshanfar, M.J.G. Olsthoorn, S.E. Verwer
Effective LLM-based automated program repair (APR) methods can lead to massive cost reductions and have improved significantly in recent times. However, the validity of many APR evaluations as they are conducted at this point is at risk due to data leakage: Prior research has shown that LLMs can memorize solutions to problems if the evaluation benchmark overlaps with the training set, leading to overinflated results.
In this study, we examine the potential of using metamorphic transformations to mitigate the effects of data leakage. For this, we create a variant benchmark for two popular, well-established benchmarks Defects4J and GitBug-Java, and evaluate the APR performance of several LLMs on these benchmarks and their transformed counterparts. In addition, we investigate to what extent our results align with data leakage metrics from other studies.
Our results show that state-of-the-art LLMs for code repair exhibit significant performance degradation (Up to 4.1% for Claude-3.7-Sonnet) on a metamorphically transformed Defecsts4J benchmark. Moreover, we find a significant correlation between our results and the negative log-likelihood as a metric of data leakage. Our results demonstrate the potential of using metamorphic transformations to mitigate the overinflation of evaluation results due to data leakage. We recommend that researchers report results on both original and metamorphically transformed benchmarks in future evaluations.
...
In this study, we examine the potential of using metamorphic transformations to mitigate the effects of data leakage. For this, we create a variant benchmark for two popular, well-established benchmarks Defects4J and GitBug-Java, and evaluate the APR performance of several LLMs on these benchmarks and their transformed counterparts. In addition, we investigate to what extent our results align with data leakage metrics from other studies.
Our results show that state-of-the-art LLMs for code repair exhibit significant performance degradation (Up to 4.1% for Claude-3.7-Sonnet) on a metamorphically transformed Defecsts4J benchmark. Moreover, we find a significant correlation between our results and the negative log-likelihood as a metric of data leakage. Our results demonstrate the potential of using metamorphic transformations to mitigate the overinflation of evaluation results due to data leakage. We recommend that researchers report results on both original and metamorphically transformed benchmarks in future evaluations.
...
Effective LLM-based automated program repair (APR) methods can lead to massive cost reductions and have improved significantly in recent times. However, the validity of many APR evaluations as they are conducted at this point is at risk due to data leakage: Prior research has shown that LLMs can memorize solutions to problems if the evaluation benchmark overlaps with the training set, leading to overinflated results.
In this study, we examine the potential of using metamorphic transformations to mitigate the effects of data leakage. For this, we create a variant benchmark for two popular, well-established benchmarks Defects4J and GitBug-Java, and evaluate the APR performance of several LLMs on these benchmarks and their transformed counterparts. In addition, we investigate to what extent our results align with data leakage metrics from other studies.
Our results show that state-of-the-art LLMs for code repair exhibit significant performance degradation (Up to 4.1% for Claude-3.7-Sonnet) on a metamorphically transformed Defecsts4J benchmark. Moreover, we find a significant correlation between our results and the negative log-likelihood as a metric of data leakage. Our results demonstrate the potential of using metamorphic transformations to mitigate the overinflation of evaluation results due to data leakage. We recommend that researchers report results on both original and metamorphically transformed benchmarks in future evaluations.
In this study, we examine the potential of using metamorphic transformations to mitigate the effects of data leakage. For this, we create a variant benchmark for two popular, well-established benchmarks Defects4J and GitBug-Java, and evaluate the APR performance of several LLMs on these benchmarks and their transformed counterparts. In addition, we investigate to what extent our results align with data leakage metrics from other studies.
Our results show that state-of-the-art LLMs for code repair exhibit significant performance degradation (Up to 4.1% for Claude-3.7-Sonnet) on a metamorphically transformed Defecsts4J benchmark. Moreover, we find a significant correlation between our results and the negative log-likelihood as a metric of data leakage. Our results demonstrate the potential of using metamorphic transformations to mitigate the overinflation of evaluation results due to data leakage. We recommend that researchers report results on both original and metamorphically transformed benchmarks in future evaluations.
As single-cell RNA sequencing techniques improve and more cells are measured in individual experiments, cell clustering procedures become increasingly more computationally intensive. This paper studies the runtime performance impact of a specialized clustering algorithm for data converted to a binary format, in order to reduce computational burden. We experimentally show that our specialized algorithm runs faster than the Seurat library on small datasets, and that with proper dimensionality reduction and approximation techniques, the algorithm could be more scalable than current methods. Optimizations for cluster quality and memory efficiency are not considered in this paper.
...
As single-cell RNA sequencing techniques improve and more cells are measured in individual experiments, cell clustering procedures become increasingly more computationally intensive. This paper studies the runtime performance impact of a specialized clustering algorithm for data converted to a binary format, in order to reduce computational burden. We experimentally show that our specialized algorithm runs faster than the Seurat library on small datasets, and that with proper dimensionality reduction and approximation techniques, the algorithm could be more scalable than current methods. Optimizations for cluster quality and memory efficiency are not considered in this paper.