GB

G.A. Bouland

info

Please Note

12 records found

Aging is the biological process that changes the body over time. When we age our bodies become more prone to disease and other health risks. But not everyone experiences these changes at the same age. This is because the age of our cells (biological age) does not always match our chronological age (time since birth). Being able to predict someone’s biological age and comparing it to their chronological age can be used to infer if someone is indeed more prone to diseases or other health risks.

Other studies have been able to predict the age of cells by using gene expressions. They explore the number of expressions in young and old individuals to identify genes that are affected by age. What has not yet been explored is how the correlation of gene pairs are affected by age. How genes cooperate can change with age, this can be captured by looking at how genes correlate and how that correlation changes with age. This paper will explore these correlations and answer the following question. By performing a correlation analysis between features of young individuals, and on the same features for old individuals, can we interpret any differences and use those to improve current age prediction models?

During this study we found a lot of gene pairs that have a significant difference in correlation from younger to older individuals. We also identified hub genes that change correlation with many other genes. Using these genes to train a linear regression model we were able to predict the age of cells with a Mean Absolute Error of 9.7835.

Using the hub genes we were not able to improve the current existing linear regression model. But we did identify genes that have earlier been linked to aging. Like LIMD2, but also a lot of ribosomal genes and mitochondrial genes, both of which lose functionality with aging. ...
Single-cell RNA sequencing (scRNAseq) is a measuring technique of gene expressions in single cells that has allowed researchers to tackle Alzheimer’s disease (AD) in many ways. Single-cell data has been joined with machine learning to classify brain cells as affected by AD. However, not much is known regarding the usage of such classification models in a spatial setting. This paper analyzes how models trained on scRNAseq data can be used to find AD properties of single cells when measuring them with spatially resolved transcriptomics. With that we study the hypothesis that cells labeled as affected by the disease should appear closer to amyloid plaques, than those that are unaffected. To find out if this holds, three models are used to classify single cells spatially and their predictions are analyzed. Two single-cell datasets are used for training, each giving a drastically different classification outcome. The models do not come to a consensus on the hypothesis’ validity either, as the analysis finds no significant correlation between the variables. ...
Alzheimer's Disease (AD) is a complex heterogeneous disease and is the leading cause of dementia around the world. Treatment options remain limited and the underlying mechanisms are not yet fully understood. To get more insight on this celular level, single-cell gene expression data can be used. It has proven to be effective with machine learning for tasks like cell type classification. While prior studies have explored AD classification using scRNA-seq, this has only been a binary classification. Severity of AD is classified using multiple measures, ranging from cognitive ability scores, to neuro pathological measures. This research explores the possibility of expanding the binary prediction of AD by including these measures for AD severity. In addition, given that these measures are associated, we also investigate if Multi Task Learning (MTL) models can improve the predictions by learning multiple AD related data points. If successful, this approach can give additional analysis into key tasks, genes and/or cells (sub)types that drive the models, which would lead to more possibilities for personalized treatment options, alongside more insight into the development of AD in the brain. We used a three-layer neural network architecture alongside a translation from cellular level to individual level to make individual-level predictions. Results show that Cognitive Ability can be classified best, but overal performance is only slightly above Naive Bayes. Furthermore, MTL does not appear to have any measurable positive effect on scores compared to single task models. A link to the github repository is available at \url{https://github.com/WillemDieleman/ADseverityCSE3000}. ...
The aim of this research is to investigate whether physical gene characteristics can predict age-related changes in gene expression. Specifically, we analyze gene length, GC content, distance to the ends of the chromosome, and similar features to determine their connection with differential expression between young and old individuals. Among these features, gene length consistently shows a strong correlation with age-related expression patterns. However, when combined, the selected features do not provide sufficient predictive power to train a classifier capable of exceeding a modest 66% accuracy. These findings highlight the limitations of the current feature set and point toward the need for more complex feature preprocessing steps or biologically relevant features in future predictive models. ...

Enhancing Accuracy and Biological Interpretability

Biological aging clocks estimate age from molecular data and provide insights into age-related functional decline. While aging clocks based on bulk transcriptomic data are well-studied, their single-cell counterparts remain limited and underexplored. In this study, we replicate and enhance a recent single-cell RNA-seq aging clock for human immune cells using ElasticNet, improving its performance through refined preprocessing, feature selection, and regularization. We also explore LightGBM to assess nonlinear modeling potential. Our enhanced models reduce prediction error, generalize better across external datasets, and identify biologically relevant genes through SHAP analysis. These findings support the development of accurate, interpretable, cell-type-specific aging clocks using single-cell data. ...
As single-cell RNA sequencing techniques improve and more cells are measured in individual experiments, cell clustering procedures become increasingly more computationally intensive. This paper studies the runtime performance impact of a specialized clustering algorithm for data converted to a binary format, in order to reduce computational burden. We experimentally show that our specialized algorithm runs faster than the Seurat library on small datasets, and that with proper dimensionality reduction and approximation techniques, the algorithm could be more scalable than current methods. Optimizations for cluster quality and memory efficiency are not considered in this paper. ...

What is the gain in peak memory usage of the binary clustering algorithm compared to current state-of-the-art clustering methods?

The rapid increase in the size of single-cell RNAseq datasets presents significant performance challenges when conducting evaluations and extracting information. We research an alternative input data format that utilizes binarization. Our main focus is an analysis of peak memory usage. An in-depth exploration of the solution’s design and implementation is provided, specifically emphasizing the strategies used to minimize memory usage. We analyzed and validated memory usage patterns and asymptotes using memory profiling tools. However, our findings suggest that gains in reducing memory usage on big modern datasets are attributable only to binarized data format rather than workflow interaction with the new format, which we found to be independent of the input format. ...

How close can we get to state-of-the-art ?

Analysing single-cell RNA sequencing data is becoming an increasingly tedious task as the size of data sets grows. As a proposed solution, recent discoveries suggest that these data sets can be binarized without losing much information. This in turn should allow for memory and time efficient methods of storage and computation. Numerous analyses techniques require cell clustering as a preliminary procedure, which suggests the need to evaluate binary representation performance under that context. In this work we present a comparison between binary clustering results and the state-of-the-art, with a focus on similarity metric choice and the impact on intermediate steps of the procedure (i.e. similarity matrices and kNN graphs). The method was evaluated on single-cell transcriptomic data sets, utilizing a combination of R and C++ as an evaluation framework. Through these means we found that some of the similarity metrics operating on continuous input can possibly be reproduced with similarity metrics operating on binary input. ...

The impact of binarized scRNA-seq data on clustering through community detection algorithms

Single-cell RNA sequencing data clustering is a valuable technique for demonstrating cell-to-cell heterogeneity and revealing cell dynamics within and amongst groups. Large up-scaling of scRNA-seq datasets in recent years pose computational challenges for existing state-of-the-art clustering techniques. A possible solution to tackle these challenges is to binarize the scRNA-seq data and perform clustering using optimized binary methods. Using a binary clustering pipeline we demonstrate that binary clustering solutions resemble conventional clustering solutions for large clusters, but show less resemblance for smaller clusters. We also show that the Leiden community detection algorithm can achieve higher cluster quality compared to the Louvain algorithm for the binarized data. ...
Understanding the role of genes and genetic variants is a key challenge in unraveling the driving mechanisms of Alzheimer's disease (AD). Single-cell RNA sequencing is a technique that quantifies gene expression at the cell (type) level enabling investigation of the roles of different cell types in disease. We analyzed changes in gene (co-)expression associated with genetic variants using single-cell RNA sequencing data (>1.3 million cells) from the dorsolateral prefrontal cortex (DLPFC) of 379 individuals of the ROSMAP cohort. Our single cell expression quantitative trait loci (sc-eQTL) analysis determined 3,337,065 sc-eQTLs, linking 1,882,645 SNPs to changes in expression of 8,057 genes in 7 major cell types. Next, we investigated the association of genetic variants with changes in co-expression for gene pairs (co-eQTLs), focusing on a set of variants and genes relevant to AD. Our novel non-parametric method for co-eQTL analysis compares gene co-expression distributions between SNP genotypes. We found 6,878 cell type specific co-eQTLs (variant-gene-gene combinations) relating to 18 AD variants. Although a substantial proportion of the findings is driven by eQTL effects, our method identified co-eQTLs that would not have been discovered in a correlation-based analysis. Most notable, we found variant rs13237518 (located in the TMEM106B gene) to associate with expression changes in a subset of 25 genes in excitatory neurons which is possibly indicative of higher-level disruptions related to the variant. Overall, we show that exploring genetic variant-associated changes in gene (co-)expression is a promising approach in finding cell type specific mechanisms that may be altered in AD. ...