Search PubMedSearch

Biomedical subjects

Aaron R Quinlan

Publications and source records attributed to Aaron R Quinlan.

3 recordsLinked to original sources

AVITI sequencing of a four-generation CEPH/Utah pedigree confirms low mutation rates at homopolymer loci despite their low sequence complexity.

BACKGROUND: Short tandem repeats (STRs) and homopolymers are among the most mutable loci in the human genome. Despite their presumed mutability owing to replication slippage, homopolymer loci exhibit lower mutation rates and minimal paternal age effects compared to other STRs. This paradox questions if technical limitations, rather than biological mechanisms, explain these observations. RESULTS: We used the Element Biosciences AVITI platform to sequence the genomes of a 48-member, four-generation CEPH/Utah pedigree. As the AVITI platform reduces error rates at repetitive sequences compared to Illumina, this design enabled accurate mutation discovery at 90% of assayed homopolymers and a 1.7-fold increase in discoverable mutations compared to Illumina. We identified a median of 35 de novo homopolymer mutations per trio and a mutation rate of 5.28 &#xd7; 10-5 DNMs per locus per generation, confirming a lower rate than dinucleotides (1.94 &#xd7; 10-4). Most DNMs were single base-pair expansions or contractions. Despite comprising <1% of homopolymer loci, G/C homopolymers showed 18-fold higher mutation rates than A/T homopolymers; in contrast, the high dinucleotide mutation rate is not driven by a particular motif class. Parent-of-origin analysis revealed 78% of homopolymer mutations are paternal in origin, but no significant paternal age effect was observed. CONCLUSIONS: This study confirms that homopolymers exhibit lower mutation rates and lack strong paternal age effects compared to other STRs, likely owing to the combination of a lower propensity to form slippage-causing secondary structures and more efficient mismatch repair. Our set of high-quality mutations suggest these phenomena are biological rather than technical in nature. Finally, we demonstrate that AVITI sequencing unlocks previously intractable regions of the genome and will be a powerful tool for continued investigation of repeat mutation.

AVITI

A genome-wide approach for the discovery of novel repeat expansion disorders in the Undiagnosed Diseases Network cohort.

PURPOSE: The Undiagnosed Diseases Network is a National Institutes of Health funded research study that aims to solve a broad clinical spectrum of challenging rare disease cases. Participants receive care from multiple clinical specialists, who collaborate to perform deep phenotyping and state-of-the-art multiomics analyses. As bioinformatics of short-read sequencing has matured, the discovery of repeat expansion disorders (REDs) is accelerating. REDs comprise approximately 60 characterized disorders, which exhibit a broad spectrum of phenotypes. Thus, a largely unbiased genome-wide approach in a phenotypically diverse sample will add to the diagnostic depth, explore the limits of short-read genome analysis, and establish novel candidate RED loci. METHODS: Here, we present a genome-wide analysis of repeat expansions conducted on 1018 genomes from the Undiagnosed Diseases Network. By leveraging 2 distinct bioinformatics tools, ExpansionHunter Denovo and STRling, we showed that repeat expansions can be accurately detected in short-read genomes. RESULTS: We demonstrated that a genotype-first approach can diagnose atypical cases of known REDs and provide valuable clinical insights. We present clinical details on participants with expansions in ATXN7, DMPK, FMR1, GLS, HTT, RFC1, AFF3, and MARCH6. Importantly, we highlight 2 cases of juvenile Huntington disease that were discovered through our analysis. Finally, we present a list of novel candidate short tandem repeats (TR) that could potentially be pathogenic if expanded. CONCLUSION: Importantly, our approach showcases the bioinformatic advancements in genome analysis for RED detection and highlights its practical applications.

Humans

Vcfexpress: flexible, rapid user-expressions to filter and format VCFs.

MOTIVATION: Variant call format (VCF) files are the standard output format for various software tools that identify genetic variation from DNA sequencing experiments. Downstream analyses require the ability to query, filter, and modify them simply and efficiently. Several tools are available to perform these operations from the command line, including BCFTools, vembrane, slivar, and others. RESULTS: Here, we introduce vcfexpress, a new, high-performance toolset for the analysis of VCF files, written in the Rust programming language. It is nearly as fast as BCFTools, but adds functionality to execute user expressions in the lua programming language for precise filtering and reporting of variants from a VCF or BCF file. We demonstrate performance and flexibility by comparing vcfexpress to other tools using the vembrane benchmark. AVAILABILITY AND IMPLEMENTATION: vcfexpress is available under the MIT license at https://github.com/brentp/vcfexpress with code used for the manuscript deposited in https://doi.org/10.5281/zenodo.14756838.

Software