Circular Image

A. Urhan

info

Please Note

11 records found

Journal article (2026) - Bruna F. Sgardioli, Matthew C. Phillips, Arjun M. Miklos, Kailey Fleiszig-Evans, Abigail L. Manson, Suelen Scarpa de Mello, Terrance Shea, A. Urhan, Ryan Whipple, More authors...
Enterococci appear to have originated in the guts of early terrestrializing arthropods and invertebrates over 425 million years ago-hosts that are now highly diverse and widespread in nature today. Yet most knowledge of the genus comes from human infection-associated lineages with genomes swollen by the recent accretion of foreign DNA conveyed by mobile elements. Because invertebrates dominate terrestrial animal diversity and biomass, they would be predicted to constitute a major but little-explored reservoir of enterococcal diversity. We therefore systematically examined Enterococcus association and species diversification in invertebrate hosts of the comparatively natural, isolated, but well-characterized environment of the Azorean island of Terceira. Over 100 invertebrate specimens were examined for associated enterococci, which were taxonomically classified by whole-genome sequencing. Supporting the existence of a large pool of uncharacterized enterococci and Enterococcus-adapted genes, 40% (eight of 20) of the Enterococcus species identified were either undescribed, including four candidate new species described here, or very recently discovered. In contrast, control isolates from vertebrates were exclusively of known species typical of sampling elsewhere, discounting geographic isolation as a main driver of the novelty observed. Further, because of the abundance of E. casseliflavus and E. flavescens in this collection, we obtained the resolution necessary to quantify the divergence and decipher the drivers of speciation in the controversial division between these naturally vancomycin-resistant species. These findings provide robust support for the existence of a large pool of new species and unexplored adaptive traits in invertebrate-associated enterococci-diverse environmental survival traits optimized for expression in an enterococcal background, and well positioned for transmission into human-associated enterococcal strains.IMPORTANCEEnterococci are auxotrophic gut-associated bacteria that co-evolved with their terrestrial hosts over many eons. In the last 75 years-the "antibiotic era"-E. faecalis and E. faecium gained genes for antibiotic resistance and enhanced virulence, emerging as leading causes of multidrug-resistant infection. Little is known about the source of those genes or the pathway by which they entered human-associated strains. A recent global survey suggested a potentially large repository of uncharacterized genetic diversity in the enterococci of invertebrates. We directly tested this prospect by examining enterococci of invertebrate hosts in a largely natural and pastoral environment. Our findings provide clear evidence that invertebrates naturally harbor vast unexplored enterococcal diversity. Moreover, associations are likely driven by intrinsic host selection factors rather than geographic isolation. This expands our knowledge of Enterococcus biodiversity, including the identification of four novel species, identifying a vast reservoir of enterococcal genes available to species that colonize and infect humans. ...

Synteny-aware gene function prediction for bacteria using protein embeddings

Journal article (2024) - Aysun Urhan, Bianca-Maria Cosma, Ashlee M. Earl, Abigail L. Manson, Thomas Abeel
Motivation: Today, we know the function of only a small fraction of the protein sequences predicted from genomic data. This problem is even more salient for bacteria, which represent some of the most phylogenetically and metabolically diverse taxa on Earth. This low rate of bacterial gene annotation is compounded by the fact that most function prediction algorithms have focused on eukaryotes, and conventional annotation approaches rely on the presence of similar sequences in existing databases. However, often there are no such sequences for novel bacterial proteins. Thus, we need improved gene function prediction methods tailored for bacteria. Recently, transformer-based language models - adopted from the natural language processing field - have been used to obtain new representations of proteins, to replace amino acid sequences. These representations, referred to as protein embeddings, have shown promise for improving annotation of eukaryotes, but there have been only limited applications on bacterial genomes. Results: To predict gene functions in bacteria, we developed SAFPred, a novel synteny-aware gene function prediction tool based on protein embeddings from state-of-the-art protein language models. SAFpred also leverages the unique operon structure of bacteria through conserved synteny. SAFPred outperformed both conventional sequence-based annotation methods and state-of-the-art methods on multiple bacterial species, including for distant homolog detection, where the sequence similarity to the proteins in the training set was as low as 40%. Using SAFPred to identify gene functions across diverse enterococci, of which some species are major clinical threats, we identified 11 previously unrecognized putative novel toxins, with potential significance to human and animal health. ...
Journal article (2024) - Julia A. Schwartzman, Francois Lebreton, Rauf Salamzade, Terrance Shea, Melissa J. Martin, Katharina Schaufler, Aysun Urhan, Thomas Abeel, Ilana L.B.C. Camargo, More authors...
Enterococci are gut microbes of most land animals. Likely appearing first in the guts of arthropods as they moved onto land, they diversified over hundreds of millions of years adapting to evolving hosts and host diets. Over 60 enterococcal species are now known. Two species, Enterococcus faecalis and Enterococcus faecium, are common constituents of the human microbiome. They are also now leading causes of multidrug-resistant hospital-associated infection. The basis for host association of enterococcal species is unknown. To begin identifying traits that drive host association, we collected 886 enterococcal strains from widely diverse hosts, ecologies, and geographies. This identified 18 previously undescribed species expanding genus diversity by >25%. These species harbor diverse genes including toxins and systems for detoxification and resource acquisition. Enterococcus faecalis and E. faecium were isolated from diverse hosts highlighting their generalist properties. Most other species showed a more restricted distribution indicative of specialized host association. The expanded species diversity permitted the Enterococcus genus phylogeny to be viewed with unprecedented resolution, allowing features to be identified that distinguish its four deeply rooted clades, and the entry of genes associated with range expansion such as B-vitamin biosynthesis and flagellar motility to be mapped to the phylogeny. This work provides an unprecedentedly broad and deep view of the genus Enterococcus, including insights into its evolution, potential new threats to human health, and where substantial additional enterococcal diversity is likely to be found. ...
Doctoral thesis (2024) - A. Urhan, M.J.T. Reinders, Thomas Abeel
We are witnessing an era of rapid technological advancements, which led to an explosion in the amount of genomic data collected. The field of comparative genomics, in parallel, is expanding at an unrepentant rate. Comparative genomics explores the similarities and differences in the genomes of various organisms, species or strains, and it is one of our most useful tools today for unraveling the complexities of microbial biology. However, despite growing interest in microbial genomics, there remains a significant gap in our understanding of microbial diversity and function. The microbial dark matter remains elusive, and we have a lot more to uncover.
This dissertation aims to leverage comparative genomics, and develop novel algorithms tailored for microbial genomes to enhance our understanding of microbial biology and address existing knowledge gaps. More specifically, it focuses on the representation of microbial diversity and the functional annotation of poorly characterized taxa. By harnessing large-scale genomic datasets, novel approaches and algorithms are designed to uncover hidden traits in microorganisms.
We begin our journey at the smallest scale with viruses; our study of SARS-CoV-2 genomes in the Netherlands during the COVID-19 pandemic showcases the power of genomic data to understand disease dynamics. The remainder of the dissertation concerns bacteria. We explore pangenome graphs to represent bacterial populations. As I discuss the limitations of current methods, I propose an ensemble approach to exploit graph representations for structural variant calling. This work sets the stage for future developments in pangenome graphs as a powerful framework to model bacterial populations and analyze their genetic makeup.
Following recent developments in algorithms for eukaryotes, I draw inspiration from natural language processing to predict gene functions in bacteria. I present SAFPred, a novel tool in which I integrate bacterial synteny into the predictive model, and demonstrate its use to identify variants of toxin genes in Enterococcus. The novelty of my approach lies partly in how I incorporated bacterial synteny into the function prediction algorithm. Thus, I also release our synteny database, SAFPredDB, that can facilitate various comparative genomic analyses in the future. Our journey comes to an end in our study of the Enterococcus genus through the largest collection of genome assemblies. Here, I emphasize the importance of understanding microbial diversity and antibiotic resistance mechanisms once again, and note the power of large scale genomic analyses.
Overall, my main goal with this dissertation is to showcase the potential of comparative genomics in unraveling the mysteries of microbial life and addressing pressing global challenges in health, agriculture, and biotechnology. Through innovative methods and large-scale data analysis, my work, first and foremost, offers valuable insights into microbial biology and evolution, paving the way for future research in the field. And I hope it also encourages further exploration and appreciation of the mighty world of microbes. ...

Identifying antimicrobial resistance gene transfer between plasmids

Motivation: Plasmids are carriers for antimicrobial resistance (AMR) genes and can exchange genetic material with other structures, contributing to the spread of AMR. There is no reliable approach to identify the transfer of AMR genes across plasmids. This is mainly due to the absence of a method to assess the phylogenetic distance of plasmids, as they show large DNA sequence variability. Identifying and quantifying such transfer can provide novel insight into the role of small mobile elements and resistant plasmid regions in the spread of AMR. Results: We developed SHIP, a novel method to quantify plasmid similarity based on the dynamics of plasmid evolution. This allowed us to find conserved fragments containing AMR genes in structurally different and phylogenetically distant plasmids, which is evidence for lateral transfer. Our results show that regions carrying AMR genes are highly mobilizable between plasmids through transposons, integrons, and recombination events, and contribute to the spread of AMR. Identified transferred fragments include a multi-resistant complex class 1 integron in Escherichia coli and Klebsiella pneumoniae, and a region encoding tetracycline resistance transferred through recombination in Enterococcus faecalis. ...
The success of antibiotics as a therapeutic agent has led to their ineffectiveness. The continuous use and misuse in clinical and non-clinical areas have led to the emergence and spread of antibiotic-resistant bacteria and its genetic determinants. This is a multi-dimensional problem that has now become a global health crisis. Antibiotic resistance research has primarily focused on the clinical healthcare sectors while overlooking the non-clinical sectors. The increasing antibiotic usage in the environment – including animals, plants, soil, and water – are drivers of antibiotic resistance and function as a transmission route for antibiotic resistant pathogens and is a source for resistance genes. These natural compartments are interconnected with each other and humans, allowing the spread of antibiotic resistance via horizontal gene transfer between commensal and pathogenic bacteria. Identifying and understanding genetic exchange within and between natural compartments can provide insight into the transmission, dissemination, and emergence mechanisms. The development of high-throughput DNA sequencing technologies has made antibiotic resistance research more accessible and feasible. In particular, the combination of metagenomics and powerful bioinformatic tools and platforms have facilitated the identification of microbial communities and has allowed access to genomic data by bypassing the need for isolating and culturing microorganisms. This review aimed to reflect on the different sequencing techniques, metagenomic approaches, and bioinformatics tools and pipelines with their respective advantages and limitations for antibiotic resistance research. These approaches can provide insight into resistance mechanisms, the microbial population, emerging pathogens, resistance genes, and their dissemination. This information can influence policies, develop preventative measures and alleviate the burden caused by antibiotic resistance. ...

Haplotype assembly tool using short and error-prone long reads

Motivation: Haplotypes are the set of alleles co-occurring on a single chromosome and inherited together to the next generation. Because a monoploid reference genome loses this co-occurrence information, it has limited use in associating phenotypes with allelic combinations of genotypes. Therefore, methods to reconstruct the complete haplotypes from DNA sequencing data are crucial. Recently, several attempts have been made at haplotype reconstructions, but significant limitations remain. High-quality continuous haplotypes cannot be created reliably, particularly when there are few differences between the homologous chromosomes. Results: Here, we introduce HAT, a haplotype assembly tool that exploits short and long reads along with a reference genome to reconstruct haplotypes. HAT tries to take advantage of the accuracy of short reads and the length of the long reads to reconstruct haplotypes. We tested HAT on the aneuploid yeast strain Saccharomyces pastorianus CBS1483 and multiple simulated polyploid datasets of the same strain, showing that it outperforms existing tools. ...

Acinetobacter baumannii pan-genome reveals structural variation in antimicrobial resistance-carrying plasmids

Journal article (2021) - A. Urhan, T.E.P.M.F. Abeel
Microbial organisms have diverse populations, where using a single linear reference sequence in comparative studies introduces reference-bias in downstream analyses, and leads to a failure to account for variability in the population. Recently, pan-genome graphs have emerged as an alternative to the traditional linear reference with many successful applications and a rapid increase in the number of methods available in the literature. Despite this enthusiasm, there has been no attempt at exploring these graph construction methods in depth, demonstrating their practical use. In this study, we aim to develop a general guide to help researchers who may want to incorporate pan-genomes in their analyses of microbial organisms. We evaluated the state-of- the art pan-genome construction tools to model a collection of 70 Acinetobacter baumannii strains. Our results suggest that all tools produced pan-genome graphs conforming to our expectations based on previous literature, and that their approach to homologue detection is likely to be the most influential in determining the final size and complexity of the pan-genome. The graphs overlapped most in the core pan-genome content while the cloud genes varied significantly among tools. We propose an alternative approach for pan-genome construction by combining two of the tools, Panaroo and Ptolemy, to further exploit them in downstream analyses, and demonstrate the effectiveness of our pipeline for structural variant calling in beta-lactam resistance genes in the same set of A. baumannii isolates, identifying various transposon structures for carbapenem resistance in chromosome, as well as plasmids. We identify a novel plasmid structure in two multidrug-resistant clinical isolates that had previously been studied, and which could be important for their resistance phenotypes ...
Journal article (2021) - Aysun Urhan, Thomas Abeel
Coronavirus disease 2019 (COVID-19) has emerged in December 2019 when the first case was reported in Wuhan, China and turned into a pandemic with 27 million (September 9th) cases. Currently, there are over 95,000 complete genome sequences of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the virus causing COVID-19, in public databases, accompanying a growing number of studies. Nevertheless, there is still much to learn about the viral population variation when the virus is evolving as it continues to spread. We have analyzed SARS-CoV-2 genomes to identify the most variant sites, as well as the stable, conserved ones in samples collected in the Netherlands until June 2020. We identified the most frequent mutations in different geographies. We also performed a phylogenetic study focused on the Netherlands to detect novel variants emerging in the late stages of the pandemic and forming local clusters. We investigated the S and N proteins on SARS-CoV-2 genomes in the Netherlands and found the most variant and stable sites to guide development of diagnostics assays and vaccines. We observed that while the SARS-CoV-2 genome has accumulated mutations, diverging from reference sequence, the variation landscape is dominated by four mutations globally, suggesting the current reference does not represent the virus samples circulating currently. In addition, we detected novel variants of SARS-CoV-2 almost unique to the Netherlands that form localized clusters and region-specific sub-populations indicating community spread. We explored SARS-CoV-2 variants in the Netherlands until June 2020 within a global context; our results provide insight into the viral population diversity for localized efforts in tracking the transmission of COVID-19, as well as sequenced-based approaches in diagnostics and therapeutics. We emphasize that little diversity is observed globally in recent samples despite the increased number of mutations relative to the established reference sequence. We suggest sequence-based analyses should opt for a consensus representation to adequately cover the genomic variation observed to speed up diagnostics and vaccine design. ...
Journal article (2020) - Aysun Urhan, Burak Alakent
Most applications of soft sensors in process industries require learning from a stream of data, which may exhibit nonstationary dynamics, or concept drift. In this study, we develop a relevance vector machine (RVM) based novel adaptive learning algorithm called MWAdp-JITL, to meet the demands of continuous processes. The resulting algorithm combines active and passive learning: A moving window (MW) algorithm, which adapts the window size against virtual/real concept drifts, is coupled with a just-in-time learning (JITL) model, constructed using an appropriate region of historical data, and the ensemble weights of the MW and JITL models are adjusted for each query point. Tests on four real industrial datasets and a synthetic data, comprising various concept drift scenarios, show that MWAdp-JITL yields superior prediction accuracy and is generally more robust to changes in algorithm parameters compared to conventional adaptive learning methods and state-of-the-art algorithms from the literature. MWAdp-JITL complies with time limits of online prediction, and is applicable for high dimensional processes under various types of concept drifts. It is seen that MWAdp-JITL can successfully achieve a good balance in bias-variance tradeoff, justifying the use of only two exquisitely selected learners in ensemble learning. ...
Preprint (2019) - A. Urhan, Burak Alakent
Data driven soft sensor design has recently gained immense popularity, due to advances in sensory devices, and a growing interest in data mining. While partial least squares (PLS) is traditionally used in the process literature for designing soft sensors, the statistical literature has focused on sparse learners, such as Lasso and relevance vector machine (RVM), to solve the high dimensional data problem. In the current study, predictive performances of three regression techniques, PLS, Lasso and RVM were assessed and compared under various offline and online soft sensing scenarios applied on datasets from five real industrial plants, and a simulated process. In offline learning, predictions of RVM and Lasso were found to be superior to those of PLS when a large number of time-lagged predictors were used. Online prediction results gave a slightly more complicated picture. It was found that the minimum prediction error achieved by PLS under moving window (MW), or just-in-time learning scheme was decreased up to ~5-10% using Lasso, or RVM. However, when a small MW si ze was used, or the optimum number of PLS components was as low as ~1, prediction performance of PLS surpassed RVM, which was found to yield occasional unstable predictions. PLS and Lasso models constructed via online parameter tuning generally did not yield better predictions compared to those constructed via offline tuning. We present evidence to suggest that retaining a large portion of the available process measurement data in the predictor matrix, instead of preselecting variables, would be more advantageous for sparse learners in increasing prediction accuracy. As a result, Lasso is recommended as a better substitute for PLS in soft sensors; while performance of RVM should be validated before online application. ...