<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
A Spark Framework for Cost Effective, Fast and Accurate DNA Analysis at Scale
Conference paper(2017)
-
Hamid Mushtaq, Frank Liu, Carlos Costa, Gang Liu, Peter Hofstee, Zaid Al-Ars
In recent years, the cost of NGS (Next Generation Sequencing) technology has dramatically reduced, making it a viable method for
diagnosing genetic diseases. The large amount of data generated by NGS technology, usually in the order of hundreds of gigabytes per experiment, have to be analyzed quickly to generate meaningful variant results. The GATK best practices pipeline from the Broad
Institute is one of the most popular computational pipelines for DNA analysis. Many components of the GATK pipeline are not very
parallelizable though. In this paper, we present SparkGA, a parallel implementation of a DNA analysis pipeline based on the big data
Apache Spark framework. This implementation is highly scalable and capable of parallelizing computation by utilizing data-level
parallelism as well as load balancing techniques. In order to reduce the analysis cost, SparkGA can run on nodes with as little memory as 16GB. For whole genome sequencing experiments, we show that the runtime can be reduced to about 1.5 hours on a 20-node cluster with an accuracy of up to 99.9981%. Moreover, SparkGA is about 71% faster than other state-of-the-art solutions while also being more accurate. The source code of SparkGA is publicly available at ttps://github.com/HamidMushtaq/SparkGA1.git.
...
In recent years, the cost of NGS (Next Generation Sequencing) technology has dramatically reduced, making it a viable method for
diagnosing genetic diseases. The large amount of data generated by NGS technology, usually in the order of hundreds of gigabytes per experiment, have to be analyzed quickly to generate meaningful variant results. The GATK best practices pipeline from the Broad
Institute is one of the most popular computational pipelines for DNA analysis. Many components of the GATK pipeline are not very
parallelizable though. In this paper, we present SparkGA, a parallel implementation of a DNA analysis pipeline based on the big data
Apache Spark framework. This implementation is highly scalable and capable of parallelizing computation by utilizing data-level
parallelism as well as load balancing techniques. In order to reduce the analysis cost, SparkGA can run on nodes with as little memory as 16GB. For whole genome sequencing experiments, we show that the runtime can be reduced to about 1.5 hours on a 20-node cluster with an accuracy of up to 99.9981%. Moreover, SparkGA is about 71% faster than other state-of-the-art solutions while also being more accurate. The source code of SparkGA is publicly available at ttps://github.com/HamidMushtaq/SparkGA1.git.
Journal article(2016)
-
Gang Liu, Liping Li, Shaopeng Wu, Martin van de Ven
A rapidly increasing source of reclaimed asphalt (RA) containing polymer-modified bitumen (PMB) offers a potential premium material contribution when performing recycling. This manuscript studies the influence of soft virgin paving grade binder and PMB binder on the rheological properties of three PMB-containing RA binders from different ‘old’ surface-layer asphalt mixtures in Europe, by using the dynamic shear rheometer (DSR). The results indicated that the soft virgin binder and PMB binder can change the flow abilities of RA binder. After the normalisation, the Cole–Cole curve for the modulus can be used to describe the influence of soft fresh binder on the rheological properties of reclaimed binders. The Cole–Cole curve in terms of the viscosity can be used to describe the colloidal shifting of reclaimed binder when blending with fresh binder, and the fresh binder dominated the structure of blended binders. Mixing with soft PMB binder made it possible to rejuvenate the rheological properties of the reclaimed PMB binder. Then, the polymer from the reclaimed binder can still function in a new mixture when recycled.
...
A rapidly increasing source of reclaimed asphalt (RA) containing polymer-modified bitumen (PMB) offers a potential premium material contribution when performing recycling. This manuscript studies the influence of soft virgin paving grade binder and PMB binder on the rheological properties of three PMB-containing RA binders from different ‘old’ surface-layer asphalt mixtures in Europe, by using the dynamic shear rheometer (DSR). The results indicated that the soft virgin binder and PMB binder can change the flow abilities of RA binder. After the normalisation, the Cole–Cole curve for the modulus can be used to describe the influence of soft fresh binder on the rheological properties of reclaimed binders. The Cole–Cole curve in terms of the viscosity can be used to describe the colloidal shifting of reclaimed binder when blending with fresh binder, and the fresh binder dominated the structure of blended binders. Mixing with soft PMB binder made it possible to rejuvenate the rheological properties of the reclaimed PMB binder. Then, the polymer from the reclaimed binder can still function in a new mixture when recycled.
Cookie settings
We use necessary cookies to make the TU Delft Repository work.
Help us improve the Repository
With your permission, we use privacy-friendly Matomo analytics to understand how people use the
Repository — for example, which features are used and where we can improve the search experience. The analytics are managed by TU Delft and are not used for advertising or commercial tracking. Your IP
address is anonymized, and analytics data is not shared with third parties.
Choosing “Accept all” helps the Library improve the Repository for researchers, students, and other
users. You can change your choice at any time using the cookie settings icon in the footer. For more information, read our
privacy statement.