K.A. Hildebrandt
Please Note
51 records found
1
Analysis of results in the ML research field
LLM assistance in peer reviewing ML papers
fication and evaluation of responsible research checklists impose a significant burden on
reviewers. This study investigates the ability of Large Language Models (LLMs) to au-
tomatically classify research papers as empirical, theoretical, or hybrid, and to extract
checklist compliance data. Using a dataset of publicly available NeurIPS papers, we
designed an automated pipeline and evaluated its outputs against a human-annotated
ground truth. Our results demonstrate that the LLM achieves high accuracy in the
core classification task, reliably distinguishing the papers core methodology by iden-
tifying clear structural indicators like mathematical proofs and benchmark datasets.
Furthermore, the model excels at extracting objective checklist elements, performing
well on close-ended extraction tasks that rely on clear structural indicators. However,
performance noticeably decreased on structurally scattered or subjective criteria, such
as broader impacts and the declaration of AI usage. This drop highlights a limitation in
the model’s broader reading comprehension, as it struggles to merge contextual infor-
mation without explicit headers. Notably, this automated failure closely mirrors human
task ambiguity, as these exact subjective items also generated the lower inter-annotator
agreement among human annotators. Conclusively, while LLMs provide a highly con-
sistent baseline for classifying paper typologies and extracting explicit methodological
data, their reliance on structural cues indicates they should serve as assistive screening
tools rather than autonomous evaluators in academic peer review. ...
fication and evaluation of responsible research checklists impose a significant burden on
reviewers. This study investigates the ability of Large Language Models (LLMs) to au-
tomatically classify research papers as empirical, theoretical, or hybrid, and to extract
checklist compliance data. Using a dataset of publicly available NeurIPS papers, we
designed an automated pipeline and evaluated its outputs against a human-annotated
ground truth. Our results demonstrate that the LLM achieves high accuracy in the
core classification task, reliably distinguishing the papers core methodology by iden-
tifying clear structural indicators like mathematical proofs and benchmark datasets.
Furthermore, the model excels at extracting objective checklist elements, performing
well on close-ended extraction tasks that rely on clear structural indicators. However,
performance noticeably decreased on structurally scattered or subjective criteria, such
as broader impacts and the declaration of AI usage. This drop highlights a limitation in
the model’s broader reading comprehension, as it struggles to merge contextual infor-
mation without explicit headers. Notably, this automated failure closely mirrors human
task ambiguity, as these exact subjective items also generated the lower inter-annotator
agreement among human annotators. Conclusively, while LLMs provide a highly con-
sistent baseline for classifying paper typologies and extracting explicit methodological
data, their reliance on structural cues indicates they should serve as assistive screening
tools rather than autonomous evaluators in academic peer review.
Large Language Models for Reviewing Research Papers
Evaluating Claim-Level Completeness in Machine Learning Research
Analysis of results in the ML research field
Investigating the Efficacy of LLMs in Extracting Stated Research Limitations
Analysis of Results in the ML Research Field
How well can an LLM decide the reproducibility of a paper?
We evaluate different neural architectures, comparing a multilayer perceptron to DeltaConv, a graph convolutional model, and find that the MLP provides superior performance. In addition, we assess multiple segmentation strategies and identify watershed as the most effective, followed by hierarchical segmentation. We also find that the segmentation algorithms do not achieve real-time performance for large meshes.
These results highlight the potential of machine learning-based fracture simulations, but also indicate that distance field segmentation is not capable of real-time performance using our tested algorithms. This suggests that future work should focus on directly learning the labels rather than relying on distance fields as an intermediary representation in real-time scenarios.
...
We evaluate different neural architectures, comparing a multilayer perceptron to DeltaConv, a graph convolutional model, and find that the MLP provides superior performance. In addition, we assess multiple segmentation strategies and identify watershed as the most effective, followed by hierarchical segmentation. We also find that the segmentation algorithms do not achieve real-time performance for large meshes.
These results highlight the potential of machine learning-based fracture simulations, but also indicate that distance field segmentation is not capable of real-time performance using our tested algorithms. This suggests that future work should focus on directly learning the labels rather than relying on distance fields as an intermediary representation in real-time scenarios.
We use the user–user covariance (or its inverse, the precision matrix) as a graph shift operator (GSO) and train SelectionGNN-based VNNs on the MovieLens-100K dataset.
Two training regimes are evaluated: (i) RMSE-only (No-BAM-SVNN) and (ii) a compound loss that also includes novelty and diversity terms (BAM-SVNN).
For each regime we sweep six graph configurations: covariance/precision crossed with {dense, hard-threshold, soft-threshold} sparsification, under five random seeds, yielding 30 runs per regime.
Baseline comparisons include PCA, a naive mean–std model, and a random predictor.
The best SVNN configuration increases recommendation Novelty by 2.8 percentage points and matches PCA’s Diversity while incurring only a 0.03 RMSE penalty.
Hard-thresholded precision graphs provide the lowest SVNN RMSE (0.952), whereas dense covariance graphs maximise diversity (0.868).
Integrating novelty/diversity directly into the loss offers no additional benefit yet multiplies runtime by x33.
One-way ANOVA indicates that model family explains 97.6% of RMSE variance (\(\eta^2=0.976\)) and 77.8% of novelty variance.
This work is the first to benchmark (sparsified) VNNs on beyond-accuracy metrics, demonstrating a favourable accuracy–novelty trade-off and clarifying when sparsification and BAM-weighted training pay off.
All code, data splits and statistical notebooks are released for full reproducibility. ...
We use the user–user covariance (or its inverse, the precision matrix) as a graph shift operator (GSO) and train SelectionGNN-based VNNs on the MovieLens-100K dataset.
Two training regimes are evaluated: (i) RMSE-only (No-BAM-SVNN) and (ii) a compound loss that also includes novelty and diversity terms (BAM-SVNN).
For each regime we sweep six graph configurations: covariance/precision crossed with {dense, hard-threshold, soft-threshold} sparsification, under five random seeds, yielding 30 runs per regime.
Baseline comparisons include PCA, a naive mean–std model, and a random predictor.
The best SVNN configuration increases recommendation Novelty by 2.8 percentage points and matches PCA’s Diversity while incurring only a 0.03 RMSE penalty.
Hard-thresholded precision graphs provide the lowest SVNN RMSE (0.952), whereas dense covariance graphs maximise diversity (0.868).
Integrating novelty/diversity directly into the loss offers no additional benefit yet multiplies runtime by x33.
One-way ANOVA indicates that model family explains 97.6% of RMSE variance (\(\eta^2=0.976\)) and 77.8% of novelty variance.
This work is the first to benchmark (sparsified) VNNs on beyond-accuracy metrics, demonstrating a favourable accuracy–novelty trade-off and clarifying when sparsification and BAM-weighted training pay off.
All code, data splits and statistical notebooks are released for full reproducibility.
Recommender systems via Covariance Neural Networks
Using precision matrices as Graph Collaborative Filter
Representing CNN Feature Maps with Implicit Neural Representations
A Proof-of-Concept Study Using SIRENs
served to filter out irregular plaques in the unlabeled Aβ data. After performing 5-fold cross-validation, the fine-tuned models demonstrated an average accuracy of 85.71% and precision of 89.47% on the primary types. The final classifier used on the unlabeled Aβ dataset was an ensemble model that incorporated majority voting, combining the predictions of the five models trained during cross-validation. Aβ loads were calculated for each primary Aβ type based on the classifier’s predictions. It was observed that across all considered primary Aβ types, the centenarians’ Aβ loads were statistically significantly lower compared to the AD cohort. The lower Aβ load in centenarians also held true across the frontal, temporal, parietal, and occipital cerebral regions for each primary Aβ type. To partially validate the
model’s performance, correlations were computed between the predicted Aβ loads of the primary Aβ types and neuropathological assessments collected from 75 centenarians. These assessments are common in related works and serve as a benchmark for the ensemble model. They include the Thal Aβ phase, which categorizes the distribution of general Aβ in the brain; the Thal CAA stage, which measures the severity of CAA; and CERAD NP scores, which evaluates the spread of neuritic plaques in the brain, which are a subset of cored plaques. Consequently, statistically significant positive correlations were revealed between: the Thal Aβ phase and the Aβ load of all primary types (r ranging 0.59-0.73); the Thal CAA stage and Aβ load of predicted CAA deposits (r = 0.66); and the CERAD NP scores and Aβ load of cored plaques (r = 0.68). Since the correlated types coincide with what the benchmark staging schemes measure, the model’s predictions seems to align with existing literature. ...
served to filter out irregular plaques in the unlabeled Aβ data. After performing 5-fold cross-validation, the fine-tuned models demonstrated an average accuracy of 85.71% and precision of 89.47% on the primary types. The final classifier used on the unlabeled Aβ dataset was an ensemble model that incorporated majority voting, combining the predictions of the five models trained during cross-validation. Aβ loads were calculated for each primary Aβ type based on the classifier’s predictions. It was observed that across all considered primary Aβ types, the centenarians’ Aβ loads were statistically significantly lower compared to the AD cohort. The lower Aβ load in centenarians also held true across the frontal, temporal, parietal, and occipital cerebral regions for each primary Aβ type. To partially validate the
model’s performance, correlations were computed between the predicted Aβ loads of the primary Aβ types and neuropathological assessments collected from 75 centenarians. These assessments are common in related works and serve as a benchmark for the ensemble model. They include the Thal Aβ phase, which categorizes the distribution of general Aβ in the brain; the Thal CAA stage, which measures the severity of CAA; and CERAD NP scores, which evaluates the spread of neuritic plaques in the brain, which are a subset of cored plaques. Consequently, statistically significant positive correlations were revealed between: the Thal Aβ phase and the Aβ load of all primary types (r ranging 0.59-0.73); the Thal CAA stage and Aβ load of predicted CAA deposits (r = 0.66); and the CERAD NP scores and Aβ load of cored plaques (r = 0.68). Since the correlated types coincide with what the benchmark staging schemes measure, the model’s predictions seems to align with existing literature.
An Experimental Look at the Stability of Graph Neural Networks against Topological Perturbations
The Relationship Between Graph Properties and Stability
Beyond Spectral Graph Theory
An Explainability-Driven Approach to Analyzing the Stability of Graph Neural Networks to Topology Perturbations
Urban Change Detection Based on Remote Sensing Data
How are Recurrent Neural Networks applied in the context of urban change detection?