Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

A complete diploid human genome benchmark for personalized genomics.

Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.

Journal Article↗

Fine-scale structural variation of the human genome.

Inversions, deletions and insertions are important mediators of disease and disease susceptibility. We systematically compared the human genome reference sequence with a second genome (represented by fosmid paired-end sequences) to detect intermediate-sized structural variants >8 kb in length. We identified 297 sites of structural variation: 139 insertions, 102 deletions and 56 inversion breakpoints. Using combined literature, sequence and experimental analyses, we validated 112 of the structural variants, including several that are of biomedical relevance. These data provide a fine-scale structural variation map of the human genome and the requisite sequence precision for subsequent genetic studies of human disease.

Base Pairing↗

Colora: a Snakemake workflow for complete chromosome-scale de novo genome assembly.

MOTIVATION: De novo assembly creates reference genomes that underpin many modern biodiversity and conservation studies. Large numbers of new genomes are being assembled by labs around the world. To avoid duplication of efforts and variable data quality, we desire a best-practice assembly process, implemented as an automated portable workflow. RESULTS: Here, we present Colora, a Snakemake workflow that produces chromosome-scale de novo primary or phased genome assemblies complete with organelles using Pacific Biosciences HiFi, Hi-C, and optionally Oxford Nanopore Technologies reads as input. Colora is a user-friendly, versatile, and reproducible pipeline that is ready to use by researchers looking for an automated way to obtain high-quality de novo genome assemblies. AVAILABILITY AND IMPLEMENTATION: The source code of Colora is available on GitHub (https://github.com/LiaOb21/colora) and has been deposited in Zenodo under DOI https://doi.org/10.5281/zenodo.13321576. Colora is also available at the Snakemake Workflow Catalog (https://snakemake.github.io/snakemake-workflow-catalog/? usage=LiaOb21%2Fcolora).

Software↗

The oryza map alignment project: the golden path to unlocking the genetic potential of wild rice species.

The wild species of the genus Oryza offer enormous potential to make a significant impact on agricultural productivity of the cultivated rice species Oryza sativa and Oryza glaberrima. To unlock the genetic potential of wild rice we have initiated a project entitled the 'Oryza Map Alignment Project' (OMAP) with the ultimate goal of constructing and aligning BAC/STC based physical maps of 11 wild and one cultivated rice species to the International Rice Genome Sequencing Project's finished reference genome--O. sativa ssp. japonica c. v. Nipponbare. The 11 wild rice species comprise nine different genome types and include six diploid genomes (AA, BB, CC, EE, FF and GG) and four tetrapliod genomes (BBCC, CCDD, HHKK and HHJJ) with broad geographical distribution and ecological adaptation. In this paper we describe our strategy to construct robust physical maps of all 12 rice species with an emphasis on the AA diploid O. nivara--thought to be the progenitor of modern cultivated rice.

Chromosome Mapping↗

Bacterial species determination from DNA-DNA hybridization by using genome fragments and DNA microarrays.

Whole genomic DNA-DNA hybridization has been a cornerstone of bacterial species determination but is not widely used because it is not easily implemented. We have developed a method based on random genome fragments and DNA microarray technology that overcomes the disadvantages of whole-genome DNA-DNA hybridization. Reference genomes of four fluorescent Pseudomonas species were fragmented, and 60 to 96 genome fragments of approximately 1 kb from each strain were spotted on microarrays. Genomes from 12 well-characterized fluorescent Pseudomonas strains were labeled with Cy dyes and hybridized to the arrays. Cluster analysis of the hybridization profiles revealed taxonomic relationships between bacterial strains tested at species to strain level resolution, suggesting that this approach is useful for the identification of bacteria as well as determining the genetic distance among bacteria. Since arrays can contain thousands of DNA spots, a single array has the potential for broad identification capacity. In addition, the method does not require laborious cross-hybridizations and can provide an open database of hybridization profiles, avoiding the limitations of traditional DNA-DNA hybridization.

DNA, Bacterial↗

HPRC2: A human pangenome reference with near-complete coverage of common genetic variation.

A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium's (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.

Journal Article↗

Features of Arabidopsis genes and genome discovered using full-length cDNAs.

Arabidopsis is currently the reference genome for higher plants. A new, more detailed statistical analysis of Arabidopsis gene structure is presented including intron and exon lengths, intergenic distances, features of promoters, and variant 5'-ends of mRNAs transcribed from the same transcription unit. We also provide a statistical characterization of Arabidopsis transcripts in terms of their size, UTR lengths, 3'-end cleavage sites, splicing variants, and coding potential. These analyses were facilitated by scrutiny of our collection of sequenced full-length cDNAs and much larger collection of 5'-ESTs, together with another set of full-length cDNAs from Salk/Stanford/Plant Gene Expression Center/RIKEN. Examples of alternative splicing are observed for transcripts from 7% of the genes and many of these genes display multiple spliced isoforms. Most splicing variants lie in non-coding regions of the transcripts. Non-canonical splice sites constitute less than 1% of all splice sites. Genes with fewer than four introns display reduced average mRNA levels. Putative alternative transcription start sites were observed in 30% of highly expressed genes and in more than 50% of the genes with low expression. Transcription start sites correlate remarkably well with a CG skew peak in the DNA sequences. The intergenic distances vary considerably, those where genes are transcribed towards one another being significantly shorter. New transcripts, missing in the current TIGR genome annotation and ESTs that are non-coding, including those antisense to known genes, are derived and cataloged in the Supplementary Material. They identify 148 new loci in the Arabidopsis genome. The conclusions drawn provide a better understanding of the Arabidopsis genome and how the gene transcripts are processed. The results also allow better predictions to be made for, as yet, poorly defined genes and provide a reference for comparisons with other plant genomes whose complete sequences are currently being determined. Some comparisons with rice are included in this paper.

Alternative Splicing↗

Biomedical applications and studies of molecular evolution: a proposal for a primate genomic library resource.

The anticipated completion of two of the most biomedically relevant genomes, mouse and human, within the next three years provides an unparalleled opportunity for the large-scale exploration of genome evolution. Targeted sequencing of genomic regions in a panel of primate species and comparison to reference genomes will provide critical insight into the nature of single-base pair variation, mechanisms of chromosomal rearrangement, patterns of selection, and species adaptation. Although not recognized as model "genetic organisms" because of their longevity and low fecundity, 30 of the approximately 300 primate species are targets of biomedical research. The existence of a human reference sequence and genomic primate BAC libraries greatly facilitates the recovery of genes/genomic regions of high biological interest because of an estimated maximum neutral nucleotide sequence divergence of 25%. Primate species, therefore, may be regarded as the ideal model "genomic organisms". Based on existing BAC library resources, we propose the construction of a panel of primate BAC libraries from phylogenetic anchor species for the purpose of comparative medicine as well as studies of genome evolution.

Animals↗

Genomic and bioinformatics analyses of HAdV-4vac and HAdV-7vac, two human adenovirus (HAdV) strains that constituted original prophylaxis against HAdV-related acute respiratory disease, a reemerging epidemic disease.

Vaccine strains of human adenovirus serotypes 4 and 7 (HAdV-4vac and HAdV-7vac) have been used successfully to prevent adenovirus-related acute respiratory disease outbreaks. The genomes of these two vaccine strains have been sequenced, annotated, and compared with their prototype equivalents with the goals of understanding their genomes for molecular diagnostics applications, vaccine redevelopment, and HAdV pathoepidemiology. These reference genomes are archived in GenBank as HAdV-4vac (35,994 bp; AY594254) and HAdV-7vac (35,240 bp; AY594256). Bioinformatics and comparative whole-genome analyses with their recently reported and archived prototype genomes reveal six mismatches and four insertions-deletions (indels) between the HAdV-4 prototype and vaccine strains, in contrast to the 611 mismatches and 130 indels between the HAdV-7 prototype and vaccine strains. Annotation reveals that the HAdV-4vac and HAdV-7vac genomes contain 51 and 50 coding units, respectively. Neither vaccine strain appears to be attenuated for virulence based on bioinformatics analyses. There is evidence of genome recombination, as the inverted terminal repeat of HAdV-4vac is initially identical to that of species C whereas the prototype is identical to species B1. These vaccine reference sequences yield unique genome signatures for molecular diagnostics. As a molecular forensics application, these references identify the circulating and problematic 1950s era field strains as the original HAdV-4 prototype and the Greider prototype, from which the vaccines are derived. Thus, they are useful for genomic comparisons to current epidemic and reemerging field strains, as well as leading to an understanding of pathoepidemiology among the human adenoviruses.

Acute Disease↗

Genomic technologies and the interrogation of the transcriptome.

Functional genomics refers to the study of whole genomes and the function of its constituent parts to explain biological processes. Though these investigations may involve whole proteome analysis, the primary focus is on the transcriptome and how it is regulated. Recent advances in technologies that can interrogate cellular transcripts on a genome-wide scale seek the complete disclosure of the transcriptome over time-intervals and across many different cellular states. This massively complex data when viewed as a whole can provide surprisingly precise assessment of cellular conditions. Moreover, these data can define hierarchies of importance and have shown us new transcriptional elements. Herein, we describe the technologies and the experimental strategies to study the transcriptome that would be pertinent to cancer and ageing research.

Animals↗

Comparative genome assembly.

One of the most complex and computationally intensive tasks of genome sequence analysis is genome assembly. Even today, few centres have the resources, in both software and hardware, to assemble a genome from the thousands or millions of individual sequences generated in a whole-genome shotgun sequencing project. With the rapid growth in the number of sequenced genomes has come an increase in the number of organisms for which two or more closely related species have been sequenced. This has created the possibility of building a comparative genome assembly algorithm, which can assemble a newly sequenced genome by mapping it onto a reference genome. We describe here a novel algorithm for comparative genome assembly that can accurately assemble a typical bacterial genome in less than four minutes on a standard desktop computer. The software is available as part of the open-source AMOS project.

Algorithms↗

An intermediate grade of finished genomic sequence suitable for comparative analyses.

Although the cost of generating draft-quality genomic sequence continues to decline, refining that sequence by the process of "sequence finishing" remains expensive. Near-perfect finished sequence is an appropriate goal for the human genome and a small set of reference genomes; however, such a high-quality product cannot be cost-justified for large numbers of additional genomes, at least for the foreseeable future. Here we describe the generation and quality of an intermediate grade of finished genomic sequence (termed comparative-grade finished sequence), which is tailored for use in multispecies sequence comparisons. Our analyses indicate that this sequence is very high quality (with the residual gaps and errors mostly falling within repetitive elements) and reflects 99% of the total sequence. Importantly, comparative-grade sequence finishing requires approximately 40-fold less reagents and approximately 10-fold less personnel effort compared to the generation of near-perfect finished sequence, such as that produced for the human genome. Although applied here to finishing sequence derived from individual bacterial artificial chromosome (BAC) clones, one could envision establishing routines for refining sequences emanating from whole-genome shotgun sequencing projects to a similar quality level. Our experience to date demonstrates that comparative-grade sequence finishing represents a practical and affordable option for sequence refinement en route to comparative analyses.

Animals↗

The transcriptome of the intraerythrocytic developmental cycle of Plasmodium falciparum.

Plasmodium falciparum is the causative agent of the most burdensome form of human malaria, affecting 200-300 million individuals per year worldwide. The recently sequenced genome of P. falciparum revealed over 5,400 genes, of which 60% encode proteins of unknown function. Insights into the biochemical function and regulation of these genes will provide the foundation for future drug and vaccine development efforts toward eradication of this disease. By analyzing the complete asexual intraerythrocytic developmental cycle (IDC) transcriptome of the HB3 strain of P. falciparum, we demonstrate that at least 60% of the genome is transcriptionally active during this stage. Our data demonstrate that this parasite has evolved an extremely specialized mode of transcriptional regulation that produces a continuous cascade of gene expression, beginning with genes corresponding to general cellular processes, such as protein synthesis, and ending with Plasmodium-specific functionalities, such as genes involved in erythrocyte invasion. The data reveal that genes contiguous along the chromosomes are rarely coregulated, while transcription from the plastid genome is highly coregulated and likely polycistronic. Comparative genomic hybridization between HB3 and the reference genome strain (3D7) was used to distinguish between genes not expressed during the IDC and genes not detected because of possible sequence variations. Genomic differences between these strains were found almost exclusively in the highly antigenic subtelomeric regions of chromosomes. The simple cascade of gene regulation that directs the asexual development of P. falciparum is unprecedented in eukaryotic biology. The transcriptome of the IDC resembles a "just-in-time" manufacturing process whereby induction of any given gene occurs once per cycle and only at a time when it is required. These data provide to our knowledge the first comprehensive view of the timing of transcription throughout the intraerythrocytic development of P. falciparum and provide a resource for the identification of new chemotherapeutic and vaccine candidates.

Animals↗

SwinePan for pig graph-based pangenome and multiomics data mining.

Pigs are one of the most important livestock species worldwide. Although multiple high-quality reference genomes exist, reliance on a single linear reference limits the detection of structural variants (SVs) and the characterization of population-specific genetic diversity. To address this limitation, we developed SwinePan, a comprehensive and integrated multiomics database for pigs built on a graph-based pangenome framework. SwinePan incorporates a variome derived from the graph-based pangenome, covering 2,598 individuals across 35 breeds, including 185,759 SVs, 117 million SNPs, and 6.8 million indels. The database also integrates transcriptomic data from liver, loin muscle, abdominal fat, and backfat, along with over 150,000 phenotypic records. The online toolkit deployed in SwinePan enables genome-wide association studies (GWAS), expression quantitative trait locus (eQTL) mapping, and colocalization, while interactive modules visualize population structure and multiomics associations, streamlining candidate gene and variant exploration. Additionally, two proof-of-concept analyses demonstrate how SwinePan pinpoints trait-associated loci and deciphers their potential regulatory mechanisms.

Journal Article↗

Genome-wide analysis of DNA copy-number changes using cDNA microarrays.

Gene amplifications and deletions frequently contribute to tumorigenesis. Characterization of these DNA copy-number changes is important for both the basic understanding of cancer and its diagnosis. Comparative genomic hybridization (CGH) was developed to survey DNA copy-number variations across a whole genome. With CGH, differentially labelled test and reference genomic DNAs are co-hybridized to normal metaphase chromosomes, and fluorescence ratios along the length of chromosomes provide a cytogenetic representation of DNA copy-number variation. CGH, however, has a limited ( approximately 20 Mb) mapping resolution, and higher-resolution techniques, such as fluorescence in situ hybridization (FISH), are prohibitively labour-intensive on a genomic scale. Array-based CGH, in which fluorescence ratios at arrayed DNA elements provide a locus-by-locus measure of DNA copy-number variation, represents another means of achieving increased mapping resolution. Published array CGH methods have relied on large genomic clone (for example BAC) array targets and have covered only a small fraction of the human genome. cDNAs representing over 30,000 radiation-hybrid (RH)-mapped human genes provide an alternative and readily available genomic resource for mapping DNA copy-number changes. Although cDNA microarrays have been used extensively to characterize variation in human gene expression, human genomic DNA is a far more complex mixture than the mRNA representation of human cells. Therefore, analysis of DNA copy-number variation using cDNA microarrays would require a sensitivity of detection an order of magnitude greater than has been routinely reported. We describe here a cDNA microarray-based CGH method, and its application to DNA copy-number variation analysis in breast cancer cell lines and tumours. Using this assay, we were able to identify gene amplifications and deletions genome-wide and with high resolution, and compare alterations in DNA copy number and gene expression.

Chromosomes, Human, Pair 17↗

High resolution analysis of DNA copy number variation using comparative genomic hybridization to microarrays.

Gene dosage variations occur in many diseases. In cancer, deletions and copy number increases contribute to alterations in the expression of tumour-suppressor genes and oncogenes, respectively. Developmental abnormalities, such as Down, Prader Willi, Angelman and Cri du Chat syndromes, result from gain or loss of one copy of a chromosome or chromosomal region. Thus, detection and mapping of copy number abnormalities provide an approach for associating aberrations with disease phenotype and for localizing critical genes. Comparative genomic hybridization (CGH) was developed for genome-wide analysis of DNA sequence copy number in a single experiment. In CGH, differentially labelled total genomic DNA from a 'test' and a 'reference' cell population are cohybridized to normal metaphase chromosomes, using blocking DNA to suppress signals from repetitive sequences. The resulting ratio of the fluorescence intensities at a location on the 'cytogenetic map', provided by the chromosomes, is approximately proportional to the ratio of the copy numbers of the corresponding DNA sequences in the test and reference genomes. CGH has been broadly applied to human and mouse malignancies. The use of metaphase chromosomes, however, limits detection of events involving small regions (of less than 20 Mb) of the genome, resolution of closely spaced aberrations and linking ratio changes to genomic/genetic markers. Therefore, more laborious locus-by-locus techniques have been required for higher resolution studies. Hybridization to an array of mapped sequences instead of metaphase chromosomes could overcome the limitations of conventional CGH (ref. 6) if adequate performance could be achieved. Copy number would be related to the test/reference fluorescence ratio on the array targets, and genomic resolution could be determined by the map distance between the targets, or by the length of the cloned DNA segments. We describe here our implementation of array CGH. We demonstrate its ability to measure copy number with high precision in the human genome, and to analyse clinical specimens by obtaining new information on chromosome 20 aberrations in breast cancer.

Animals↗

A genomic catalog of Earth's bacterial and archaeal symbionts.

Microbial symbiosis drives the functional and phylogenomic diversification of life on Earth yet remains underexplored because of culturing challenges. This study used machine learning (ML) to predict symbiotic lifestyles in more than a hundred thousand microbial genomes from diverse environmental metagenome samples and reference genomes. Predictions were performed using symclatron, an ML framework developed to identify genomic signatures of symbionts. Predictions were deposited in a catalog we established called Symbiont Genomes (SymGs). The results indicate that 15-23% of uncultivated microorganisms likely engage in symbiotic relationships with other organisms, categorized as host-associated or obligate intracellular lifestyles, and are present in half of all known bacterial and archaeal phyla. We also identify genomic signatures of symbiotic lifestyles, including the loss of certain metabolic functions and the differential presence of metabolic modules that may enable host-dependent living. The symclatron software and the SymGs catalog represent valuable resources for studying symbioses, potentially facilitating future mechanistic investigations and engineering of host-microorganism associations.

Journal Article↗

A chromosomal level genome assembly of Nguni Sheep, Ovis aries.

Nguni sheep (Ovis aries) are indigenous to the Southern Africa region and common within the smallholder and poor resources farming systems. They are well adapted to different agroecological regions. However, limited genomic resources such as high-quality reference genomes have hindered our understanding of its adaptation and establishment of an effective breeding program. To address this, we assembled a chromosomal-level genome of Nguni sheep using a combination of PacBio HiFi reads and Omni-C reads. The genome size was estimated to be 2.9 Gb with a contig/scaffold N50 74 Mb and 99.6 Mb and a genome completeness of 96.1%, as estimated by the Benchmarking Universal Single-Copy Orthologs (BUSCO) program. The final genome encompassed a total of 25,926 protein-coding genes. The findings of this study provide a valuable genomic resource for understanding the adaptability of the Nguni sheep and the establishment of effective breeding programs.

Animals↗