Gennady Roshchupkin
Please Note
5 records found
1
Traditional statistical approaches have advanced our understanding of the genetics of complex diseases, yet are limited to linear additive models. Here we applied machine learning (ML) to genome-wide data from 41,686 individuals in the largest European consortium on Alzheimer’s disease (AD) to investigate the effectiveness of various ML algorithms in replicating known findings, discovering novel loci, and predicting individuals at risk. We utilised Gradient Boosting Machines (GBMs), biological pathway-informed Neural Networks (NNs), and Model-based Multifactor Dimensionality Reduction (MB-MDR) models. ML approaches successfully captured all genome-wide significant genetic variants identified in the training set and 22% of associations from larger meta-analyses. They highlight 6 novel loci which replicate in an external dataset, including variants which map to ARHGAP25, LY6H, COG7, SOD1 and ZNF597. They further identify novel association in AP4E1, refining the genetic landscape of the known SPPL2A locus. Our results demonstrate that machine learning methods can achieve predictive performance comparable to classical approaches in genetic epidemiology and have the potential to uncover novel loci that remain undetected by traditional GWAS. These insights provide a complementary avenue for advancing the understanding of AD genetics.
cortical measures and brain regions, and 160 genome-wide significant associations pointing to wnt/β-catenin, TGF-β and sonic hedgehog pathways. There is enrichment for genes involved in anthropometric traits, hindbrain development, vascular and neurodegenerative disease and psychiatric conditions. These data are a rich resource for studies of the biological mechanisms behind cortical development and aging. ...
cortical measures and brain regions, and 160 genome-wide significant associations pointing to wnt/β-catenin, TGF-β and sonic hedgehog pathways. There is enrichment for genes involved in anthropometric traits, hindbrain development, vascular and neurodegenerative disease and psychiatric conditions. These data are a rich resource for studies of the biological mechanisms behind cortical development and aging.
We present the largest population-based heritability study of the human brain structural connectome, including a pathology-sensitive extension, the disconnectome. The disconnectome maps the effect of white matter lesions throughout the brain. The connectome and disconnectome were generated from diffusion-weighted images of 3255 unrelated subjects from the Rotterdam Study aged between 45 and 99 years. Graph theory measures were derived for both the connectome and disconnectome. Genotypes were used to derive genetic relationship matrices between individuals for heritability analyses. High measures of heritability, from 33% to 51%, were found across all connectivity measures. The disconnectome showed more significantly heritable connectivity measures than the connectome, suggesting that the new proposed measure may reveal additional or complementary information about the genetic architecture of the human brain.
Large-scale distributed analyses of over 30,000 magnetic resonance imaging scans recently detected common genetic variants associated with the volumes of subcortical brain structures. Scaling up these efforts, still greater computational challenges arise in screening the genome for statistical associations at each voxel in the brain, localizing effects using "image-wide genome-wide" testing (voxelwise genome-wide association studies, vGWASs). Here we benefit from distributed computations at multiple sites to metaanalyze genome-wide image-wide data, allowing private genomic data to stay at the site where it was collected. Site-specific tensor-based morphometry is performed with a custom template for each site, using a multichannel registration. A single vGWAS testing 107 variants against 2million voxels can yield hundreds of terabytes (TB) of summary statistics, which would need to be transferred and pooled for metaanalysis. We propose a two-step method, which reduces data transfer for each site to a subset of single-nucleotide polymorphisms and voxels guaranteed to contain all significant hits.