Search PubMed⌕ Search

Biomedical subjects

Donna M Muzny

Publications and source records attributed to Donna M Muzny.

15 recordsLinked to original sources

Development and extensive sequencing of a broadly-consented Genome in a Bottle matched tumor-normal pair.

The Genome in a Bottle Consortium (GIAB), hosted by the National Institute of Standards and Technology (NIST), is developing new matched tumor-normal samples, the first explicitly consented for public dissemination of genomic data and cell lines. Here, we describe a comprehensive genomic dataset from the first individual, HG008, including DNA from an adherent, epithelial-like pancreatic ductal adenocarcinoma (PDAC) tumor cell line and matched normal cells from duodenal and pancreatic tissues. Data for the tumor-normal matched samples comes from seventeen distinct state-of-the-art whole genome measurement technologies, including high depth short and long-read bulk whole genome sequencing (WGS), single cell WGS, Hi-C, and karyotyping. These data will be used by the GIAB Consortium to develop matched tumor-normal benchmarks for somatic variant detection. We expect these data to facilitate innovation for whole genome measurement technologies, de novo assembly of tumor and normal genomes, and bioinformatic tools to identify small and structural somatic variants. This first-of-its-kind broadly consented open-access resource will facilitate further understanding of sequencing methods used for cancer biology.

Humans↗

Development and extensive sequencing of a broadly-consented Genome in a Bottle matched tumor-normal pair.

The Genome in a Bottle Consortium (GIAB), hosted by the National Institute of Standards and Technology (NIST), is developing new matched tumor-normal samples, the first to be explicitly consented for public dissemination of genomic data and cell lines. Here, we describe a comprehensive genomic dataset from the first individual, HG008, including DNA from an adherent, epithelial-like pancreatic ductal adenocarcinoma (PDAC) tumor cell line and matched normal cells from duodenal and pancreatic tissues. Data for the tumor-normal matched samples comes from seventeen distinct state-of-the-art whole genome measurement technologies, including high depth short and long-read bulk whole genome sequencing (WGS), single cell WGS, and Hi-C, and karyotyping. In future publications, these data will be used by the GIAB Consortium to develop matched tumor-normal benchmarks for somatic variant detection. We expect these data to facilitate innovation for whole genome measurement technologies, de novo assembly of tumor and normal genomes, and bioinformatic tools to identify small and structural somatic mutations. This first-of-its-kind broadly consented open-access resource will facilitate further understanding of sequencing methods used for cancer biology.

Journal Article↗

Inverted triplications formed by iterative template switches generate structural variant diversity at genomic disorder loci.

The duplication-triplication/inverted-duplication (DUP-TRP/INV-DUP) structure is a complex genomic rearrangement (CGR). Although it has been identified as an important pathogenic DNA mutation signature in genomic disorders and cancer genomes, its architecture remains unresolved. Here, we studied the genomic architecture of DUP-TRP/INV-DUP by investigating the DNA of 24 patients identified by array comparative genomic hybridization (aCGH) on whom we found evidence for the existence of 4 out of 4 predicted structural variant (SV) haplotypes. Using a combination of short-read genome sequencing (GS), long-read GS, optical genome mapping, and single-cell DNA template strand sequencing (strand-seq), the haplotype structure was resolved in 18 samples. The point of template switching in 4 samples was shown to be a segment of ∼2.2-5.5 kb of 100% nucleotide similarity within inverted repeat pairs. These data provide experimental evidence that inverted low-copy repeats act as recombinant substrates. This type of CGR can result in multiple conformers generating diverse SV haplotypes in susceptible dosage-sensitive loci.

Humans↗

Dominant negative variants in KIF5B cause osteogenesis imperfecta via down regulation of mTOR signaling.

BACKGROUND: Kinesin motor proteins transport intracellular cargo, including mRNA, proteins, and organelles. Pathogenic variants in kinesin-related genes have been implicated in neurodevelopmental disorders and skeletal dysplasias. We identified de novo, heterozygous variants in KIF5B, encoding a kinesin-1 subunit, in four individuals with osteogenesis imperfecta. The variants cluster within the highly conserved kinesin motor domain and are predicted to interfere with nucleotide binding, although the mechanistic consequences on cell signaling and function are unknown. METHODS: To understand the in vivo genetic mechanism of KIF5B variants, we modeled the p.Thr87Ile variant that was found in two patients in the C. elegans ortholog, unc-116, at the corresponding position (Thr90Ile) by CRISPR/Cas9 editing and performed functional analysis. Next, we studied the cellular and molecular consequences of the recurrent p.Thr87Ile variant by microscopy, RNA and protein analysis in NIH3T3 cells, primary human fibroblasts and bone biopsy. RESULTS: C. elegans heterozygous for the unc-116 Thr90Ile variant displayed abnormal body length and motility phenotypes that were suppressed by additional copies of the wild type allele, consistent with a dominant negative mechanism. Time-lapse imaging of GFP-tagged mitochondria showed defective mitochondria transport in unc-116 Thr90Ile neurons providing strong evidence for disrupted kinesin motor function. Microscopy studies in human cells showed dilated endoplasmic reticulum, multiple intracellular vacuoles, and abnormal distribution of the Golgi complex, supporting an intracellular trafficking defect. RNA sequencing, proteomic analysis, and bone immunohistochemistry demonstrated down regulation of the mTOR signaling pathway that was partially rescued with leucine supplementation in patient cells. CONCLUSION: We report dominant negative variants in the KIF5B kinesin motor domain in individuals with osteogenesis imperfecta. This study expands the spectrum of kinesin-related disorders and identifies dysregulated signaling targets for KIF5B in skeletal development.

Animals↗

Human NK cell deficiency as a result of biallelic mutations in MCM10.

Human natural killer cell deficiency (NKD) arises from inborn errors of immunity that lead to impaired NK cell development, function, or both. Through the understanding of the biological perturbations in individuals with NKD, requirements for the generation of terminally mature functional innate effector cells can be elucidated. Here, we report a cause of NKD resulting from compound heterozygous mutations in minichromosomal maintenance complex member 10 (MCM10) that impaired NK cell maturation in a child with fatal susceptibility to CMV. MCM10 has not been previously associated with monogenic disease and plays a critical role in the activation and function of the eukaryotic DNA replisome. Through evaluation of patient primary fibroblasts, modeling patient mutations in fibroblast cell lines, and MCM10 knockdown in human NK cell lines, we have shown that loss of MCM10 function leads to impaired cell cycle progression and induction of DNA damage-response pathways. By modeling MCM10 deficiency in primary NK cell precursors, including patient-derived induced pluripotent stem cells, we further demonstrated that MCM10 is required for NK cell terminal maturation and acquisition of immunological system function. Together, these data define MCM10 as an NKD gene and provide biological insight into the requirement for the DNA replisome in human NK cell maturation and function.

Alleles↗

The DNA sequence, annotation and analysis of human chromosome 3.

After the completion of a draft human genome sequence, the International Human Genome Sequencing Consortium has proceeded to finish and annotate each of the 24 chromosomes comprising the human genome. Here we describe the sequencing and analysis of human chromosome 3, one of the largest human chromosomes. Chromosome 3 comprises just four contigs, one of which currently represents the longest unbroken stretch of finished DNA sequence known so far. The chromosome is remarkable in having the lowest rate of segmental duplication in the genome. It also includes a chemokine receptor gene cluster as well as numerous loci involved in multiple human cancers such as the gene encoding FHIT, which contains the most common constitutive fragile site in the genome, FRA3B. Using genomic sequence from chimpanzee and rhesus macaque, we were able to characterize the breakpoints defining a large pericentric inversion that occurred some time after the split of Homininae from Ponginae, and propose an evolutionary history of the inversion.

Animals↗

The finished DNA sequence of human chromosome 12.

Human chromosome 12 contains more than 1,400 coding genes and 487 loci that have been directly implicated in human disease. The q arm of chromosome 12 contains one of the largest blocks of linkage disequilibrium found in the human genome. Here we present the finished sequence of human chromosome 12, which has been finished to high quality and spans approximately 132 megabases, representing approximately 4.5% of the human genome. Alignment of the human chromosome 12 sequence across vertebrates reveals the origin of individual segments in chicken, and a unique history of rearrangement through rodent and primate lineages. The rate of base substitutions in recent evolutionary history shows an overall slowing in hominids compared with primates and rodents.

Animals↗

The genome sequence of Mannheimia haemolytica A1: insights into virulence, natural competence, and Pasteurellaceae phylogeny.

The draft genome sequence of Mannheimia haemolytica A1, the causative agent of bovine respiratory disease complex (BRDC), is presented. Strain ATCC BAA-410, isolated from the lung of a calf with BRDC, was the DNA source. The annotated genome includes 2,839 coding sequences, 1,966 of which were assigned a function and 436 of which are unique to M. haemolytica. Through genome annotation many features of interest were identified, including bacteriophages and genes related to virulence, natural competence, and transcriptional regulation. In addition to previously described virulence factors, M. haemolytica encodes adhesins, including the filamentous hemagglutinin FhaB and two trimeric autotransporter adhesins. Two dual-function immunoglobulin-protease/adhesins are also present, as is a third immunoglobulin protease. Genes related to iron acquisition and drug resistance were identified and are likely important for survival in the host and virulence. Analysis of the genome indicates that M. haemolytica is naturally competent, as genes for natural competence and DNA uptake signal sequences (USS) are present. Comparison of competence loci and USS in other species in the family Pasteurellaceae indicates that M. haemolytica, Actinobacillus pleuropneumoniae, and Haemophilus ducreyi form a lineage distinct from other Pasteurellaceae. This observation was supported by a phylogenetic analysis using sequences of predicted housekeeping genes.

Actinobacillus pleuropneumoniae↗

Comparative genome sequencing of Drosophila pseudoobscura: chromosomal, gene, and cis-element evolution.

We have sequenced the genome of a second Drosophila species, Drosophila pseudoobscura, and compared this to the genome sequence of Drosophila melanogaster, a primary model organism. Throughout evolution the vast majority of Drosophila genes have remained on the same chromosome arm, but within each arm gene order has been extensively reshuffled, leading to a minimum of 921 syntenic blocks shared between the species. A repetitive sequence is found in the D. pseudoobscura genome at many junctions between adjacent syntenic blocks. Analysis of this novel repetitive element family suggests that recombination between offset elements may have given rise to many paracentric inversions, thereby contributing to the shuffling of gene order in the D. pseudoobscura lineage. Based on sequence similarity and synteny, 10,516 putative orthologs have been identified as a core gene set conserved over 25-55 million years (Myr) since the pseudoobscura/melanogaster divergence. Genes expressed in the testes had higher amino acid sequence divergence than the genome-wide average, consistent with the rapid evolution of sex-specific proteins. Cis-regulatory sequences are more conserved than random and nearby sequences between the species--but the difference is slight, suggesting that the evolution of cis-regulatory elements is flexible. Overall, a pattern of repeat-mediated chromosomal rearrangement, and high coadaptation of both male genes and cis-regulatory sequences emerges as important themes of genome divergence between these species of Drosophila.

Animals↗

Genome sequence of the Brown Norway rat yields insights into mammalian evolution.

The laboratory rat (Rattus norvegicus) is an indispensable tool in experimental medicine and drug development, having made inestimable contributions to human health. We report here the genome sequence of the Brown Norway (BN) rat strain. The sequence represents a high-quality 'draft' covering over 90% of the genome. The BN rat sequence is the third complete mammalian genome to be deciphered, and three-way comparisons with the human and mouse genomes resolve details of mammalian evolution. This first comprehensive analysis includes genes and proteins and their relation to human disease, repeated sequences, comparative genome-wide studies of mammalian orthologous chromosomal regions and rearrangement breakpoints, reconstruction of ancestral karyotypes and the events leading to existing species, rates of variation, and lineage-specific and lineage-independent evolutionary events such as expansion of gene families, orthology relations and protein evolution.

Animals↗

The status, quality, and expansion of the NIH full-length cDNA project: the Mammalian Gene Collection (MGC).

The National Institutes of Health's Mammalian Gene Collection (MGC) project was designed to generate and sequence a publicly accessible cDNA resource containing a complete open reading frame (ORF) for every human and mouse gene. The project initially used a random strategy to select clones from a large number of cDNA libraries from diverse tissues. Candidate clones were chosen based on 5'-EST sequences, and then fully sequenced to high accuracy and analyzed by algorithms developed for this project. Currently, more than 11,000 human and 10,000 mouse genes are represented in MGC by at least one clone with a full ORF. The random selection approach is now reaching a saturation point, and a transition to protocols targeted at the missing transcripts is now required to complete the mouse and human collections. Comparison of the sequence of the MGC clones to reference genome sequences reveals that most cDNA clones are of very high sequence quality, although it is likely that some cDNAs may carry missense variants as a consequence of experimental artifact, such as PCR, cloning, or reverse transcriptase errors. Recently, a rat cDNA component was added to the project, and ongoing frog (Xenopus) and zebrafish (Danio) cDNA projects were expanded to take advantage of the high-throughput MGC pipeline.

Animals↗

Finishing a whole-genome shotgun: release 3 of the Drosophila melanogaster euchromatic genome sequence.

BACKGROUND: The Drosophila melanogaster genome was the first metazoan genome to have been sequenced by the whole-genome shotgun (WGS) method. Two issues relating to this achievement were widely debated in the genomics community: how correct is the sequence with respect to base-pair (bp) accuracy and frequency of assembly errors? And, how difficult is it to bring a WGS sequence to the accepted standard for finished sequence? We are now in a position to answer these questions. RESULTS: Our finishing process was designed to close gaps, improve sequence quality and validate the assembly. Sequence traces derived from the WGS and draft sequencing of individual bacterial artificial chromosomes (BACs) were assembled into BAC-sized segments. These segments were brought to high quality, and then joined to constitute the sequence of each chromosome arm. Overall assembly was verified by comparison to a physical map of fingerprinted BAC clones. In the current version of the 116.9 Mb euchromatic genome, called Release 3, the six euchromatic chromosome arms are represented by 13 scaffolds with a total of 37 sequence gaps. We compared Release 3 to Release 2; in autosomal regions of unique sequence, the error rate of Release 2 was one in 20,000 bp. CONCLUSIONS: The WGS strategy can efficiently produce a high-quality sequence of a metazoan genome while generating the reagents required for sequence finishing. However, the initial method of repeat assembly was flawed. The sequence we report here, Release 3, is a reliable resource for molecular genetic experimentation and computational analysis.

Animals↗

Generation and initial analysis of more than 15,000 full-length human and mouse cDNA sequences.

The National Institutes of Health Mammalian Gene Collection (MGC) Program is a multiinstitutional effort to identify and sequence a cDNA clone containing a complete ORF for each human and mouse gene. ESTs were generated from libraries enriched for full-length cDNAs and analyzed to identify candidate full-ORF clones, which then were sequenced to high accuracy. The MGC has currently sequenced and verified the full ORF for a nonredundant set of >9,000 human and >6,000 mouse genes. Candidate full-ORF clones for an additional 7,800 human and 3,500 mouse genes also have been identified. All MGC sequences and clones are available without restriction through public databases and clone distribution networks (see http:mgc.nci.nih.gov).

Algorithms↗

Glass bead purification of plasmid template DNA for high throughput sequencing of mammalian genomes.

To meet the new challenge of generating the draft sequences of mammalian genomes, we describe the development of a novel high throughput 96-well method for the purification of plasmid DNA template using size-fractionated, acid-washed glass beads. Unlike most previously described approaches, the current method has been designed and optimized to facilitate the direct binding of alcohol-precipitated plasmid DNA to glass beads from alkaline lysed bacterial cells containing the insoluble cellular aggregate material. Eliminating the tedious step of separating the cleared lysate significantly simplifies the method and improves throughput and reliability. During a 4 month period of 96-capillary DNA sequencing of the Rattus norvegicus genome at the Baylor College of Medicine Human Genome Sequencing Center, the average success rate and read length derived from >1 800 000 plasmid DNA templates prepared by the direct lysis/glass bead method were 82.2% and 516 bases, respectively. The cost of this direct lysis/glass bead method in September 2001 was approximately 10 cents per clone, which is a significant cost saving in high throughput genomic sequencing efforts.

Animals↗