Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Detecting disease-causing mutations in the human genome by haplotype matching.

Comparisons between haplotypes from affected patients and the human reference genome are frequently used to identify candidates for disease-causing mutations, even though these alignments are expected to reveal a high level of background neutral polymorphism. This limits the scope of genetic studies to relatively small genomic intervals, because current methods for distinguishing potential causal mutations from neutral variation are inefficient. Here we describe a new strategy for detecting mutations that is based on comparing affected haplotypes with closely matched control sequences from healthy individuals, rather than with the human reference genome. We use theory, simulation, and a real data set to show that this approach is expected to reduce the number of sequence variants that must be subjected to follow-up analysis by at least a factor of 20 when closely matched control sequences are selected from a reference panel with as few as 100 control genomes. We also define a reference data resource that would allow efficient application of this strategy to large critical intervals across the genome.

Alleles↗

Use of RNA and genomic DNA references for inferred comparisons in DNA microarray analyses.

In most microarray assays, labeled cDNA molecules derived from reference and query RNA samples are co-hybridized to probes arrayed on a glass surface. Gene expression profiles are then calculated for each gene based on the relative hybridization intensities measured between the two samples. The most commonly used reference samples are typically isolates from a single representative RNA source (RNA-0) or pooled mixtures of RNA derived from a plurality of sources (RNA-p). Genomic DNA offers an alternative reference nucleic acid with a number of potential advantages, including stability, reproducibility, and a potentially uniform representation of all genes, as each unique gene should have equal representation in a haploid genome. Using hydrogen peroxide-treated Arabidopsis thaliana plants as a model, we evaluated genomic DNA and RNA-p as reference samples and compared expression levels inferred through the reference relative to unexposed plants with expression levels measured directly using an RNA-0 reference. Our analysis demonstrates that while genomic DNA can serve as a reasonable reference source for microarray assays, a much greater correlation with direct measurements can be achieved using an RNA-based reference sample.

Arabidopsis↗

Improving long-read somatic structural variant calling with pangenome and de novo personal genome assembly.

Accurate detection of mosaic and somatic structural variants (SVs) provides early diagnostic and therapeutic evidence for cancers. While long-read whole-genome sequencing leads to more accurate SV detection than short read sequencing, existing long-read SV callers only look at alignment against a single reference genome and are susceptible to systematic false discovery caused by germline differences between the individual genome and the reference genome. Here we develop a new SV filtering method that jointly considers the alignment against a pangenome and the de novo assembly of the germline genome. It dramatically reduces false positive mosaic and somatic SVs in cancer cell lines with little loss in sensitivity for existing long read SV callers. Our study highlights the essential need for pangenome or personal genome assembly to integrate SV calls for both SV discoveries and clinical diagnostics.

Journal Article↗

Detecting uber-operons in prokaryotic genomes.

We present a study on computational identification of uber-operons in a prokaryotic genome, each of which represents a group of operons that are evolutionarily or functionally associated through operons in other (reference) genomes. Uber-operons represent a rich set of footprints of operon evolution, whose full utilization could lead to new and more powerful tools for elucidation of biological pathways and networks than what operons have provided, and a better understanding of prokaryotic genome structures and evolution. Our prediction algorithm predicts uber-operons through identifying groups of functionally or transcriptionally related operons, whose gene sets are conserved across the target and multiple reference genomes. Using this algorithm, we have predicted uber-operons for each of a group of 91 genomes, using the other 90 genomes as references. In particular, we predicted 158 uber-operons in Escherichia coli K12 covering 1830 genes, and found that many of the uber-operons correspond to parts of known regulons or biological pathways or are involved in highly related biological processes based on their Gene Ontology (GO) assignments. For some of the predicted uber-operons that are not parts of known regulons or pathways, our analyses indicate that their genes are highly likely to work together in the same biological processes, suggesting the possibility of new regulons and pathways. We believe that our uber-operon prediction provides a highly useful capability and a rich information source for elucidation of complex biological processes, such as pathways in microbes. All the prediction results are available at our Uber-Operon Database: http://csbl.bmb.uga.edu/uber, the first of its kind.

Algorithms↗

The chromosome-level genome assembly and annotation of the silver-lipped pearl oyster, Pinctada maxima.

The silver-lipped pearl oyster (Pinctada maxima) is a valuable tropical aquaculture species, playing a crucial economic role in the global pearl industry. However, the lack of genomic reference limits our in-depth understanding of this species in genome-based breeding, conservation, evolution and adaptation. Here, annotated chromosome-level reference genome for P. maxima was generated by integrating PacBio long-read sequencing, Illumina short-read sequencing, and Hi-C sequencing data. The total genome size is 1,264.93&#x2009;Mb, with contig N50 and scaffold N50 of 649&#x2009;kb and 89.19&#x2009;Mb, respectively. The majority (97.94%) of the assembled genome was anchored to the 14 chromosomes by Hi-C analysis. The relatively high genome completeness was observed, with 97.38% (metazoa_odb10 database) and 95.26% (mollusca_odb10 database) in BUSCO analysis. Genome annotation revealed approximately 65.46% of the repeat sequences and 26,315 protein-coding genes. Comparative genome analysis revealed 28 expanded and 48 contracted families (p&#x2009;<&#x2009;0.05) in P. maxima, with 3.2% of genes (894) being species-specific. This chromosome-level genome serves as an essential resource for research in evolutionary genomics, phylogenetics, and biomineralization.

Animals↗

nf-core/magmap: Map metatranscriptomes to large collections of genomes.

SUMMARY: The lack of publicly available reference genomes has forced annotation of metatranscriptomes to either use direct alignment of sequence reads to reference databases or de novo assembly. As more and more natural environments are covered by metagenomic surveys, this is rapidly changing. This opens up the possibility of genome-resolved studies of prokaryotic metatranscriptomes by mapping to genomes from public repositories or metagenome-assembled genomes derived from the same environment. Here, we present the nf-core/magmap pipeline that provides a reproducible, easy-to-access, and well-documented workflow for selecting reference genomes, mapping to them, and quantifying features. Genomes can be drawn from public sources or originate from private collections. The pipeline is primarily aimed at prokaryotic communities but can, together with collections of reference mature gene sequences, also be applied to eukaryotes. AVAILABILITY AND IMPLEMENTATION: The nf-core/magmap pipeline is implemented in Nextflow and part of the nf-core collaboration. The pipeline is available at the nf-core website (https://nf-co.re/magmap) and GitHub (https://github.com/nf-core/magmap).

Software↗

Complete telomere-to-telomere genome assembly of Guazuma ulmifolia uncovers evolutionary mechanisms, drought adaptation, and flavonoid biosynthesis.

The first T2T reference genome of Guazuma ulmifolia is reported, which serves as a core genomic resource for stress adaptation research and stress-tolerant breeding in cacao wild relatives. Climate change, particularly increased incidence of drought, poses a major threat to food security. Understanding the genomic basis of environmental adaptation in crop wild relatives can provide valuable resources for improving stress resilience. Guazuma ulmifolia, a wild relative of Theobroma cacao with important ecological and medicinal value, lacks high-quality reference genomic resources. Here, we report the first telomere-to-telomere (T2T) chromosome-level genome assembly of G. ulmifolia, with a genome size of 311.31&#xa0;Mb, contig N50 of 35.19&#xa0;Mb, and 98.70% BUSCO completeness. Repetitive sequences constitute 27.43% of the G. ulmifolia genome, with LTR retrotransposons as the predominant class. Comparative genomic analyses revealed that genome-size variation among Malvaceae species is associated with differences in polyploidization history and TE dynamics. Ancestral karyotype reconstruction identified five lineage-specific chromosome fusion events distinguishing G. ulmifolia from T. cacao. Comparative analyses further identified tandem duplication-associated expansion of stress-related LEA and GST gene families, suggesting potential genomic features associated with stress responses. Flavonoid biosynthesis genes were largely conserved in copy number but showed tissue-specific expression patterns, providing candidate genes for investigating secondary metabolism. Together, this study establishes a high-quality T2T genome resource for exploring genome evolution, chromosome organization, and stress-related genomic features in Malvaceae.

Genome, Plant↗

ERGA-BGE chromosome-level genome assembly of the giant stream lacewing&#xa0; Osmylus fulvicephalus (Scopoli, 1763).

The giant stream lacewing, Osmylus fulvicephalus (Scopoli, 1763), is a widespread European species belonging to the insect order Neuroptera. Its cryptic larvae are predators found at the banks of streams and smaller rivers where they use their piercing, lance-shaped stylets to inject venom into their arthropod prey. Here, we present the reference genome of the giant stream lacewing as a crucial resource for uncovering the genetic basis of venom evolution in Neuroptera. The chromosome-level genome encompasses 674.7 Mb and is composed of 60 contigs and 24 scaffolds where 99.2% of the assembly is distributed among the 6 contiguous chromosomal pseudomolecules and two sex chromosomes (X and Y). Contig and scaffold N50 have a value of 51.5&#xa0;Mb and 116.2&#xa0;Mb, respectively. This reference genome is the first genomic resource from the family of lance lacewings, providing valuable data for clarifying the phylogenetic placement of the family Osmylidae within Neuroptera.

Biodiversity Genomics Europe↗

European ash pangenome reveals widespread structural variation and genetic basis of low ash dieback susceptibility.

European Ash (Fraxinus excelsior) is a keystone tree species, whose populations are being decimated by ash dieback disease (ADB) - better characterisation of genetic variants associated with low susceptibility to the disease is needed. Here, we develop a F. excelsior pangenome to more fully capture sequence variability within this species compared with a linear reference genome, using a geographically diverse set of fifty F. excelsior samples. We identify 362,965 structural variants (SVs), including 174&#x2009;Mb of sequence absent from the linear reference genome (22% of the linear reference size), and identify 3,412 high-confidence dispensable genes (those present only in some individuals). We use the pangenome to analyse existing genomic data from over 1,200 individuals, revealing 220 single nucleotide polymorphisms (SNPs) showing consistent allele frequency shifts between healthy individuals and those highly damaged by ADB, across UK seed sources, explicitly demonstrating the existence of a shared genetic component to low ADB susceptibility.

Polymorphism, Single Nucleotide↗

Evolution of toxicology for risk assessment.

The science of toxicology has served society well in protecting public health and the environment. Governments, the industrial sector, and the public have relied on toxicology as the foundation to assess risks to both human and ecological populations from environmental factors, including chemicals, biologic agents, physical agents, and other stressors. To maintain its prominence, the science and practice of toxicology will need to embrace the revolution underway in biology. Systems biology and biotechnologies derived from sequencing of the human genome, referred to as "genomics," have created exciting possibilities for application to human health and environmental risk assessment. Yet this rapid advance of science and technology can be overshadowed by inconsistency in study design and sampling strategies; by the lack of quantitative or qualitative correlations of exposure, dose, or adverse effects; and by the lack of bioinformatics tools and analytical methods necessary to manage the volume of research findings. These limitations may render results uninterpretable and difficult, if not impossible, to use in risk assessment. Recommendations will be discussed to improve integrating systems biology and genomics into risk assessment so that the inherent promise of these new approaches can be realized.

Biology↗

GiantHunter: accurate detection of giant virus in metagenomic data using reinforcement-learning and Monte Carlo tree search.

MOTIVATION: Nucleocytoplasmic large DNA viruses (NCLDVs) are notable for their large genomes and extensive gene repertoires, which contribute to their widespread environmental presence and critical roles in processes such as host metabolic reprogramming and nutrient cycling. Metagenomic sequencing has emerged as a powerful tool for uncovering novel NCLDVs in environmental samples. However, identifying NCLDV sequences in metagenomic data remains challenging due to their high genomic diversity, limited reference genomes, and shared regions with other microbes. Existing alignment-based and machine learning methods struggle with achieving optimal trade-offs between sensitivity and precision. RESULTS: In this work, we present GiantHunter, a reinforcement learning-based tool for identifying NCLDVs from metagenomic data. By employing a Monte Carlo tree search strategy, GiantHunter dynamically selects representative non-NCLDV sequences as the negative training data, enabling the model to establish a robust decision boundary. Benchmarking on rigorously designed experiments shows that GiantHunter achieves high precision while maintaining competitive sensitivity, improving the F1-score by 10% and reducing computational cost by 90% compared to the second-best method. To demonstrate its real-world utility, we applied GiantHunter to 60 metagenomic datasets collected from six cities along the Yangtze River, located both upstream and downstream of the Three Gorges Dam. The results reveal significant differences in NCLDV diversity correlated with proximity to the dam, likely influenced by reduced flow velocity caused by the dam. These findings highlight GiantHunter's potential to advance our understanding of NCLDVs and their ecological roles in diverse environments. AVAILABILITY AND IMPLEMENTATION: The source code of GiantHunter is available via: https://github.com/FuchuanQu/GiantHunter.

Metagenomics↗

Defining and cataloging variants in pangenome graphs.

Structural variation causes some human haplotypes to align poorly with the linear reference genome, leading to 'reference bias'. A pangenome reference graph could ameliorate this bias by relating a sample to multiple reference assemblies. However, this approach requires a new definition of a 'genetic variant.' We introduce a definition of pangenome variants and a method, pantree, to identify them. Our approach involves a pangenome reference tree which includes all nodes (sequences) of the pangenome graph, but only a subset of its edges; non-reference edges are variant edges. Our variants are biallelic and have well-defined positions. Analyzing the Minigraph-Cactus draft human pangenome reference graph, we identified 29.6 million genetic variants. Most variants (99.2%) are small, and most small variants (73.9%) are SNPs. 3.5 million variants (11.7%) have a reference allele which is not on GRCh38; these variants are difficult to detect without a pangenome reference, or with existing pangenome-based approaches. They tend to be embedded within tangled, multiallelic regions. We analyze two medically relevant regions, around the HLA-A and RHD genes, identifying thousands of small variants embedded within several large insertions, deletions, and inversions. We release an open-source software tool together with a VCF variant catalogue.

Journal Article↗

Genomic divergences among cattle, dog and human estimated from large-scale alignments of genomic sequences.

BACKGROUND: Approximately 11 Mb of finished high quality genomic sequences were sampled from cattle, dog and human to estimate genomic divergences and their regional variation among these lineages. RESULTS: Optimal three-way multi-species global sequence alignments for 84 cattle clones or loci (each >50 kb of genomic sequence) were constructed using the human and dog genome assemblies as references. Genomic divergences and substitution rates were examined for each clone and for various sequence classes under different functional constraints. Analysis of these alignments revealed that the overall genomic divergences are relatively constant (0.32-0.37 change/site) for pairwise comparisons among cattle, dog and human; however substitution rates vary across genomic regions and among different sequence classes. A neutral mutation rate (2.0-2.2 x 10(-9) change/site/year) was derived from ancestral repetitive sequences, whereas the substitution rate in coding sequences (1.1 x 10(-9) change/site/year) was approximately half of the overall rate (1.9-2.0 x 10(-9) change/site/year). Relative rate tests also indicated that cattle have a significantly faster rate of substitution as compared to dog and that this difference is about 6%. CONCLUSION: This analysis provides a large-scale and unbiased assessment of genomic divergences and regional variation of substitution rates among cattle, dog and human. It is expected that these data will serve as a baseline for future mammalian molecular evolution studies.

Animals↗

Image analysis for comparative genomic hybridization based on a karyotyping program for windows.

OBJECTIVE: To describe the image processing and analysis techniques we developed for the quantitative analysis of comparative genomic hybridization (CGH). STUDY DESIGN: A system for CGH cytometry is based on a semiautomated karyotyping program using the Windows graphic user interface. RESULTS: After alignment and normalization of the fluorescence images of the test genome (FITC) and reference genome (TRITC), the chromosomes are segmented and arranged as a CGH karyogram. The CGH karyograms from different metaphases of one tumor sample are represented as a CGH sum karyogram. Mean DAPI, FITC, TRITC and RATIO images, as well as ratio profiles with the 95% confidence interval, can be displayed. CONCLUSION: Representation of CGH results in the form of pseudocolored mean ratio chromosomes enhances the visibility of the method. The ratio profile plus confidence interval facilitates the identification of preparation artifacts and the classification of chromosomal imbalances. The sum karyograms of several tumors can be combined into a CGH superkaryogram of a tumor subgroup; that helps identify recurrent DNA changes in a tumor and might lead to genetic tumor classification based on CGH.

DNA, Neoplasm↗

Comparing Neanderthal introgression maps reveals core agreement but substantial heterogeneity.

Statistical methods to identify Neanderthal ancestry in modern human genomes rest on varying assumptions and inputs. Nonetheless, most studies of introgression use only a single method to define Neanderthal ancestry. Due to a lack of "ground truth," we have a limited understanding of the accuracy, comparative strengths and weaknesses, and the sensitivity of downstream conclusions for these methods. Here, we performed large-scale comparisons of genome-wide introgression maps from 12 representative Neanderthal introgression detection algorithms. These span methods that consider archaic and human reference genomes not from Africa (ArchaicSeeker2, CRF, DICAL-ADMIX), only archaic genomes (S*, Sprime, HMM, SARGE, ARGWeaver-D), only human reference genomes, including from Africa (IBDmix), or simulated data (ArchIE). Our results highlight a core set of regions predicted by nearly all methods, as well as substantial heterogeneity in commonly used Neanderthal introgression maps. Furthermore, we find that downstream analyses may result in different conclusions depending on the map used. Thus, we recommend careful consideration of map(s) chosen for an analysis and support the use of multiple maps to ensure robustness of conclusions. We make integrated prediction sets available, enabling further understanding of Neanderthal introgression's legacy on modern humans.

Journal Article↗

Web-based visualization tools for bacterial genome alignments.

With the increase in the flow of sequence data, both in contigs and whole genomes, visual aids for comparison and analysis studies are becoming imperative. We describe three web-based tools for visualizing alignments of bacterial genomes. The first, called Enteric, produces a graphical, hypertext view of pairwise alignments between a reference genome and sequences from each of several related organisms, covering 20 kb around a user-specified position. Insertions, deletions and rearrangements relative to the reference genome are color-coded, which reveals many intriguing differences among genomes. The second, Menteric, computes and displays nucleotide-level multiple alignments of the same sequences, together with annotations of ORFs and regulatory sites, in a 1 kb region surrounding a given address. The third, a Java-based viewer called Maj, combines some features of the previous tools, and adds a zoom-in mechanism. We compare the Escherichia coli K-12 genome with the partially sequenced genomes of Klebsiella pneumoniae, Yersinia pestis, Vibrio cholerae, and the Salmonella enterica serovars Typhimurium, Typhi and Paratyphi A. Examination of the pairwise and multiple alignments in a region allows one to draw inferences about regulatory patterns and functional assignments. For example, these tools revealed that rffH, a gene involved in enterobacterial common antigen (ECA) biosynthesis, is partly deleted in one of the genomes. We used PCR to show that this deletion occurs sporadically in some strains of some serovars of S.enterica subspecies I but not in any strains tested from six other subspecies. The resulting cell surface diversity may be associated with selection by the host immune response.

DNA, Bacterial↗

High-quality chromosome-level genome of three Meretrix species using Nanopore and Hi-C technologies.

Meretrix is a commercially valuable bivalve genus in Asia, but only one reference genome has hindered comprehensive genetic studies and germplasm resource evaluation. In this study, we present three reference genomes of Meretrix species: Meretrix sp. MF1, Meretrix sp. MT1, and Meretrix lamarckii JML1. Meretrix sp. MF1 was assembled at the chromosome level using Nanopore sequencing and Hi-C technologies, whereas Meretrix sp. MT1 and Meretrix lamarckii were assembled as scaffold-level assemblies. The chromosome-level genome of Meretrix sp. MF1 consists of 36 contigs, including 19 chromosomes and 17 scaffolds, with a total length of 883.3&#x2009;Mb and a scaffold N50 of 46.87&#x2009;Mb. Notably, the genome of Meretrix sp. MF1, a putative novel species, exhibits an Average Nucleotide Identity (ANI) of 94.33% with its closest relative, Meretrix lamarckii. These genomic resources not only provide a crucial foundation for genetic research on Meretrix but also contribute to the development of effective conservation strategies for its sustainable management.

Animals↗

Chromosome-Level Genome Assembly of Solanum carolinense.

Horsenettle (Solanum carolinense L.) is a noxious weed widely distributed across North America and increasingly invasive in other regions. Its strong environmental adaptability, complex defense strategies, and distinctive reproductive traits make it an important model for studying plant-herbivore coevolution. However, the absence of high-quality genomic resources has limited deeper investigation into its adaptive evolutionary mechanisms. In this study, we generated a chromosome-level reference genome assembly for S. carolinense using an integrated approach combining PacBio HiFi long-read sequencing, Illumina second-generation sequencing, and Hi-C chromatin interaction scaffolding. The final genome assembly had a total length of 915.40 Mb, with a contig N50 of 51.06 Mb and a scaffold N50 of 73.17 Mb; 96.05% of the sequences were successfully anchored onto 12 pseudochromosomes. The genome was characterized by a high proportion of repetitive sequences (73.64%) and substantial heterozygosity (1.13%), consistent with a highly repetitive and moderately high heterozygous genome. BUSCO analysis indicated that the chromosome-level genome assembly of S. carolinense reached a completeness score of 94.8%. A total of 32,206 protein-coding genes were annotated, of which 97.95% received functional annotations. The evaluation of the annotated protein-coding gene set returned a completeness value of 94.9%. This reference genome provides a valuable resource for advancing research on the adaptive evolution of weedy Solanaceae species, supports the development of more effective management strategies for this troublesome species, and offers a technical reference for assembling other highly heterozygous weed genomes.

Solanum carolinense↗