Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

ERGA-BGE reference genome of the Mediterranean monk seal ( Monachus monachus), an IUCN Vulnerable species.

The Mediterranean monk seal, Monachus monachus, is the only pinniped that lives in the Mediterranean Sea and one of the rarest marine mammals in the world. The species was recently classified as "vulnerable" by the IUCN, considering an improvement in its overall status. However, the species' populations have undergone severe bottlenecks due to systematic persecution by humans over the past centuries. Today, the global population of M. monachus is estimated to be no more than 1,000 individuals. The Mediterranean monk seal is mainly using marine caves as resting and pupping sites. It is an opportunistic apex predator and as a result its role is considered important for maintaining the structure and function of marine ecosystems. Nowadays, the species is threatened mainly by the destruction of its habitat (due to coastal development, mass tourism, and pollution) and by the depletion of its prey due to overfishing. The Mediterranean monk seal is an emblematic species; its ecological importance and its vulnerable status render its protection and effective management necessary. The entirety of the genome sequence of a female specimen was assembled into 16 contiguous chromosomal pseudomolecules, one sex chromosome (X), and one mitochondrial genome. This chromosome-level assembly encompasses 2.4 Gb, composed of 316 contigs and 275 scaffolds, with contig and scaffold N50 values of 90.9 Mb and 157 Mb, respectively.

Biodiversity Genomics Europe↗

The reference genome of the human diploid cell line RPE-1.

Recent technological advances have facilitated the assembly of telomere-to-telomere (T2T) genomes. The current T2T CHM13 showcases the complete architecture of the human genome, yet its use in functional experiments is limited by discrepancies with the actual genome of the specific biological system under study. Access to reference assemblies for experimentally relevant cell lines is therefore essential in advancing sequencing-based analyses and precise manipulation, particularly in highly variable regions such as centromeres. Here, we present RPE1v1.1, the near-complete diploid genome assembly of the hTERT RPE-1 cell line, a non-cancerous human retinal epithelial model with a stable karyotype. Using high-coverage Pacific Biosciences and Oxford Nanopore Technologies long-read sequencing, we generate a high-quality de novo assembly, validate it through multiple methods, and phase it by integrating high-throughput chromosome conformation capture (Hi-C) data. Our assembly includes chromosome-level scaffolds that span centromeres for all chromosomes. Comparing both haplotypes with the CHM13 genome, we detect haplotype-specific genomic variations, including the translocation between chromosome 10 and chromosome X t(X;10)(Xq28;10q21.2) characteristic of RPE-1 cells, and divergence peaking at centromeres. Altogether, the RPE1v1.1 genome provides a reference-quality diploid assembly of a widely used cell line, supporting high-precision genetic and epigenetic studies in this model system.

Humans↗

Analysis of human mRNAs with the reference genome sequence reveals potential errors, polymorphisms, and RNA editing.

The NCBI Reference Sequence (RefSeq) project and the NIH Mammalian Gene Collection (MGC) together define a set of approximately 30,000 nonredundant human mRNA sequences with identified coding regions representing 17,000 distinct loci. These high-quality mRNA sequences allow for the identification of transcribed regions in the human genome sequence, and many researchers accept them as the correct representation of each defined gene sequence. Computational comparison of these mRNA sequences and the recently published essentially finished human genome sequence reveals several thousand undocumented nonsynonymous substitution and frame shift discrepancies between the two resources. Additional analysis is undertaken to verify that the euchromatic human genome is sufficiently complete--containing nearly the whole mRNA collection, thus allowing for a comprehensive analysis to be undertaken. Many of the discrepancies will prove to be genuine polymorphisms in the human population, somatic cell genomic variants, or examples of RNA editing. It is observed that the genome sequence variant has significant additional support from other mRNAs and ESTs, almost four times more often than does the mRNA variant, suggesting that the genome sequence is more accurate. In approximately 15% of these cases, there is substantial support for both variants, suggestive of an undocumented polymorphism. An initial screening against a 24-individual genomic DNA diversity panel verified 60% of a small set of potential single nucleotide polymorphisms from which successful results could be obtained. We also find statistical evidence that a few of these discrepancies are due to RNA editing. Overall, these results suggest that the mRNA collections may contain a substantial number of errors. For current and future mRNA collections, it may be prudent to fully reconcile each genome sequence discrepancy, classifying each as a polymorphism, site of RNA editing or somatic cell variation, or genome sequence error.

Computational Biology↗

Rapid identification of gene sequences for transcriptional map assembly by direct cDNA screening of genomic reference libraries.

We have used the direct cDNA screening protocol to identify sequences transcribed in cerebral cortex from a reference library of human Xq28. To derive coding sequences from these genomic clones, we first identified fragments containing transcribed sequences and subjected these to exon trapping or to partial sequencing and analysis by Grail. In a preliminary analysis of three clones, coding sequences from two novel genes expressed in brain were identified. This method allows the rapid identification of coding sequences of genes expressed in specific tissues without recourse to cDNA libraries. The approach is amenable to large scale applications and should be useful for isolating candidate disease genes and in particular for assembling integrated transcriptional maps from large genomic regions.

Amino Acid Sequence↗

HIV-1 subtype H near-full length genome reference strains and analysis of subtype-H-containing inter-subtype recombinants.

OBJECTIVE: To characterize near-full-length genomes of two HIV-1 subtype H strains. To extend sequence data to include full env and gag, and analyse and redefine, previously documented subtype H strains. DESIGN: Near-full-length genomes of HIV-1 env subtype H strains VI991 and VI997 were amplified, cloned, sequenced, phylogenetically analysed and compared with a panel of 23 HIV-1 group M reference isolates. The mosaic nature of previously published subtype H strains VI557 and CA13 was reanalysed. MATERIALS AND METHODS: Peripheral blood mononuclear cells (PBMC) from individuals harbouring strains VI991 and VI997 were co-cultivated with PHA stimulated donor PBMC. Near-full-length genomes of VI991 and VI997, and gag and env genes of CA13 and VI557, were amplified by polymerase chain reaction, cloned and sequenced. Intersubtype recombination analyses were performed by similarity plot, bootscanning and phylogenetic analysis. RESULTS: Near-full-length clones of HIV-1 VI991 and VI997 are representative of subtype H. They form a phylogenetic cluster with the only previously described subtype H representative HIV-1 90CF056.1, regardless of the genome region analysed. VI557 is redefined as a gag and env subtype H mosaic virus containing unclassified fragments. CA13 is a complex intersubtype recombinant between subtypes A, H and unclassified strains CONCLUSION: Near-full-length genome analysis identified HIV-1 VI991 and VI997 as two new subtype H representatives. These reagents will allow defining and classifying non-recombinant as well as recombinant HIV-1, eventually helping to solve the puzzle of HIV-1 subtypes.

Base Sequence↗

A chromosome-level reference genome assembly of the Small snakehead (Channa asiatica).

The Small snakehead (Channa asiatica) is an economically important species in both aquaculture and ornamental trade, mainly distributed in South China and Southeast Asia. Despite its significance, limited genomic resources have impeded in-depth genetic studies and breeding programs. In this study, we used PacBio HiFi long-read sequencing, Illumina short-read sequencing, and Hi-C technologies to generate a high-quality chromosome-level genome of the C. asiatica. The final genome spans 659.44 Mb, with an impressive 98.18% anchored to 23 chromosomes. Notably, the contig N50 and scaffold N50 are 23.92 Mb and 29.61 Mb, validated by a BUSCO completeness score of 98.93%. Genome annotation identified 26,603 protein-coding genes, 99.29% of which were confirmed by BUSCO analysis, and 93.68% were functionally annotated. Approximately 27.72% of the genome sequences were classified as repeat elements. This high-fidelity genome assembly provides a robust foundation for advancing molecular breeding, comparative genomics, and evolutionary studies of C. asiatica and related species.

Animals↗

A telomere-to-telomere reference genome assembly of the red silk cotton tree (Bombax ceiba).

Bombax ceiba, an important ornamental tree and potential fiber resource in the textile industry, is widely distributed in tropical and subtropical regions. In this study, we assembled a nearly gap-free telomere-to-telomere (T2T) genome of B. ceiba using Illumina, PacBio High-fidelity (HiFi), ONT ultra-long, and Hi-C sequencing technologies. The genome spanned approximately 807.89 Mb, with a scaffold N50 of 16.58 Mb, and 754.68 Mb (93.41%) of genomic sequences were anchored onto 48 pseudo-chromosomes. Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis revealed a completeness of 99.40%, identifying 1,378 single-copy and 213 duplicated genes out of 1,614. The genome contained 67.72% (547.11 Mb) repeat regions, with 39,708 predicted protein-coding genes. Collectively, our study provides valuable genomic data for investigating the evolutionary history of the Malvaceae family.

Genome, Plant↗

Defining three dimensional chromatin structures of pediatric and adolescent B cells using primary B cell and EBV-immortalized B cell reference genomes.

BACKGROUND/PURPOSE: Knowledge of the 3D genome is essential to elucidate genetic mechanisms driving autoimmune diseases. The 3D genome is distinct for each cell type, and it is uncertain whether cell lines faithfully recapitulate the 3D architecture of primary human cells or whether developmental aspects of the pediatric immune system require use of pediatric samples. We undertook a systematic analysis of B cells and B cell lines to compare 3D genomic features encompassing risk loci for juvenile idiopathic arthritis (JIA), systemic lupus (SLE), and type 1 diabetes (T1D). METHODS: We isolated B cells from four healthy individuals, ages 9-17. HiChIP was performed using a CTCF antibody, and CTCF peaks were called within each sample separately. Peaks observed in all four samples were identified. CTCF loops were called within the pediatric samples using three CTCF peak datasets: 1) self-called CTCF consensus peaks called within the pediatric samples, 2) ENCODE's publicly available GM12878 CTCF ChIP-seq peaks, and 3) ENCODE's primary B cell CTCF ChIP-seq peaks from two adult females. Differential looping was assessed within the pediatric samples and each of the three peak datasets. RESULTS: The number of consensus peaks called in the pediatric samples was similar to that identified in ENCODE's GM12878 and primary B cell datasets. We observed&#x2009;<&#x2009;1% of loops that demonstrated significantly differential looping between peaks called within the pediatric samples themselves and when called using ENCODE GM12878 peaks. Significant looping differences were even fewer when comparing loops of the pediatric called peaks to those of the ENCODE primary B cell peaks. When querying loops found in juvenile idiopathic arthritis, type 1 diabetes, or systemic lupus erythematosus risk haplotypes, we observed significant differences in only 2.2%, 1.0%, and 1.3% loops, respectively, when comparing peaks called within the pediatric samples and ENCODE GM12878 dataset. The differences were even less apparent when comparing loops called with the pediatric vs ENCODE adult primary B cell peak datasets. CONCLUSION: The 3D chromatin architecture in B cells is similar across pediatric, adult, and EBV-transformed cell lines. This conservation of 3D structure includes regions encompassing autoimmune risk haplotypes. Thus, even for pediatric autoimmune diseases, publicly available adult B cell and cell line datasets may be sufficient for assessing effects exerted in the 3D genomic space.

Humans↗

An analysis of the use of genomic DNA as a universal reference in two channel DNA microarrays.

BACKGROUND: DNA microarray is an invaluable tool for gene expression explorations. In the two-dye microarray, fluorescence intensities of two samples, each labeled with a different dye, are compared after hybridization. To compare a large number of samples, the 'reference design' is widely used, in which all RNA samples are hybridized to a common reference. Genomic DNA is an attractive candidate for use as a universal reference, especially for bacterial systems with a low percentage of non-coding sequences. However, genomic DNA, comprising of both the sense and anti-sense strands, is unlike the single stranded cDNA usually used in microarray hybridizations. The presence of the antisense strand in the 'reference' leads to reactions between complementary labeled strands in solution and may cause the assay result to deviate from true values. RESULTS: We have developed a mathematical model to predict the validity of using genomic DNA as a reference in the microarray assay. The model predicts that the assay can accurately estimate relative concentrations for a wide range of initial cDNA concentrations. Experimental results of DNA microarray assay using genomic DNA as a reference correlated well to those obtained by a direct hybridization between two cDNA samples. The model predicts that the initial concentrations of labeled genomic DNA strands and immobilized strands, and the hybridization time do not significantly affect the assay performance. At low values of the rate constant for hybridization between immobilized and mobile strands, the assay performance varies with the hybridization time and initial cDNA concentrations. For the case where a microarray with immobilized single strands is used, results from hybridizations using genomic DNA as a reference will correspond to true ratios under all conditions. CONCLUSION: Simulation using the mathematical model, and the experimental study presented here show the potential utility of microarray assays using genomic DNA as a reference. We conclude that the use of genomic DNA as reference DNA should greatly facilitate comparative transcriptome analysis.

DNA↗

The First Highly Contiguous Genome Assembly for the Western Bluebird (Sialia mexicana).

The western bluebird (Sialia mexicana) is a secondary cavity-nesting thrush that has experienced historical population declines, local extirpations, and more recent recoveries associated with nest box programs. Despite these regional successes, recent eBird estimates suggest continued range-wide declines and substantial geographic variation in population trajectories, making this species a useful system for future studies of demographic change, connectivity, and conservation genomics. However, genomic resources for western bluebirds remain limited, and no reference genome currently exists for any species in the genus Sialia. Here, we present the first high-quality de novo reference genome for S. mexicana. Using PacBio HiFi long-read sequencing from an adult female, we generated a highly contiguous, phased 1.3&#x2005;Gb nuclear assembly with a contig N50 of 24.8&#x2005;Mb and high BUSCO completeness of 98.3%. We annotated the nuclear genome using transcriptomic and protein evidence, identifying 16,656 protein-coding genes and 26,060 transcripts/protein isoforms. We also assembled a complete &#x223c;16&#x2005;kb mitochondrial genome from Illumina short-read data. This reference genome provides a foundational resource for future studies of population structure, genetic diversity, connectivity, demographic history, and adaptation in western bluebirds and related taxa.

Animals↗

Bacteriocin (mutacin) production by Streptococcus mutans genome sequence reference strain UA159: elucidation of the antimicrobial repertoire by genetic dissection.

Streptococcus mutans UA159, the genome sequence reference strain, exhibits nonlantibiotic mutacin activity. In this study, bioinformatic and mutational analyses were employed to demonstrate that the antimicrobial repertoire of strain UA159 includes mutacin IV (specified by the nlm locus) and a newly identified bacteriocin, mutacin V (encoded by SMU.1914c).

Bacterial Proteins↗

Genome evolution of the ancient hexaploid Platanus &#xd7; acerifolia (London planetree).

Whole-genome duplication (WGD; i.e., polyploidy) and chromosomal rearrangement (i.e., genome shuffling) significantly influence genome structure and organization. Many polyploids show extensive genome shuffling relative to their pre-WGD ancestors. No reference genome is currently available for Platanaceae (Proteales), one of the sister groups to the core eudicots. Moreover, Platanus &#xd7; acerifolia (London planetree; Platanaceae) is a widely used street tree. Given the pivotal phylogenetic position of Platanus and its 2-y flowering transition, understanding its flowering-time regulatory mechanism has significant evolutionary implications; however, the impact of Platanus genome evolution on flowering-time genes remains unknown. Here, we assembled a high-quality, chromosome-level reference genome for P. &#xd7; acerifolia using a phylogeny-based subgenome phasing method. Comparative genomic analyses revealed that P. &#xd7; acerifolia (2n = 42) is an ancient hexaploid with three subgenomes resulting from two sequential WGD events; Platanus does not seem to share any WGD with other Proteales or with core eudicots. Each P. &#xd7; acerifolia subgenome is highly similar in structure and content to the reconstructed pre-WGD ancestral eudicot genome without chromosomal rearrangements. The P. &#xd7; acerifolia genome exhibits karyotypic stasis and gene sub-/neo-functionalization and lacks subgenome dominance. The copy number of flowering-time genes in P. &#xd7; acerifolia has undergone an expansion compared to other noncore eudicots, mainly via the WGD events. Sub-/neo-functionalization of duplicated genes provided the genetic basis underlying the unique flowering-time regulation in P. &#xd7; acerifolia. The P. &#xd7; acerifolia reference genome will greatly expand understanding of the evolution of genome organization, genetic diversity, and flowering-time regulation in angiosperms.

Polyploidy↗

KCFtools: rapid alignment-free method for introgression screening and GWAS using k-mer profiles.

MOTIVATION: In the era of multiple genome references, researchers often align sequencing reads against distinct assemblies or even multiple references simultaneously. This enables applications such as the detection of introgressed segments or highly variable genomic regions, which are especially prevalent in large-genome crop species such as lettuce or wheat. However, these applications come at the cost of increased computational burden, inconsistencies in mapping methods, and reduced reproducibility across studies. To address these limitations, we developed KCFtools, a Java-based toolkit that identifies the presence and absence of k-mers in nonoverlapping genomic or transcriptomic windows by comparing query and reference genomes. This alignment-free approach enables the efficient computation of an identity score for each window, thereby facilitating robust detection of introgressed or variable regions across genomes. RESULTS: We systematically evaluated the performance and accuracy of the k-mer-based method implemented in KCFtools, benchmarking it against conventional single nucleotide variation-based introgression detection pipelines. Our results demonstrate that KCFtools effectively captures introgressed segments and structurally diverse regions, even in species with fragmented or highly divergent reference genomes. In addition, we extended KCFtools to generate genotype matrices from k-mer variation tables. These matrices are compatible with genome-wide association studies software and allow the identification of loci associated with phenotypic traits. We showcase the utility of this approach by detecting known and novel associations for downy mildew resistance in lettuce, underscoring the pipeline's potential for high-resolution, reference-agnostic population genetic analysis. AVAILABILITY AND IMPLEMENTATION: https://github.com/sivasubramanics/kcftools.

Software↗

Chromosome-Level Assembly and Annotation of the Grey Reef Shark (Carcharhinus amblyrhynchos) Genome.

To date less than 5% of shark species have nuclear reference genomes, despite next-generation sequencing advances. Particularly for threatened shark species, there is a lack of reliable genomes which are crucial in facilitating research and conservation applications. We assembled the first nuclear reference genome of the endangered grey reef shark (Carcharhinus amblyrhynchos) using long-read PacBio HiFi and Omni-C sequencing to reach chromosome-level contiguity (36 pseudochromosomes; 2.9&#x2005;Gbp) and high completeness (94% complete BUSCOs). BRAKER3 annotated 16,505 protein-coding genes after masking repetitive elements which accounted for 59% of the genome. We identified potential X and Y sex chromosomes on pseudochromosomes 36 and 57, respectively. The quality and completeness of the draft genome of C. amblyrhynchos will enable researchers to investigate genetic variations and adaptations specific to this species as well as across other Carcharhinus spp., opening new venues for comparative genomics and advancing conservation genetic applications.

Animals↗

Genetic Differentiation is Constrained to Chromosomal Inversions and Putative Centromeres in Locally Adapted Populations With Higher Gene Flow.

The impact of genome structure on adaptation is a growing focus in evolutionary biology, revealing an important role for structural variation and recombination landscapes in shaping genetic diversity across genomes and among populations. This is particularly relevant when local adaptation occurs despite gene flow, where clustering of differentiated loci can maintain locally adapted variants by reducing recombination between them. However, the limited genomic resources for nonmodel species, including reference genomes and recombination maps, have constrained our understanding of these patterns. In this study, we leverage the Atlantic silverside-a nonmodel fish with extensive local adaptation across a steep latitudinal gradient-as an ideal system to explore how genome structure influences adaptation under varying levels of gene flow, using a newly available reference genome and multiple recombination maps. Analyzing 168 genomes from four populations, we found a continuum of genome-wide differentiation increasing from south to north, reflecting higher connectivity among southern populations and reduced gene flow at northern latitudes. With increasing gene flow, the number and clustering of FST outlier loci also increased, with differentiated loci found exclusively within large haploblocks harboring inversions and smaller peaks overlapping putative centromeric regions. Notably, sequence divergence was only evident in inversions, supporting their role in adaptive divergence with gene flow, whereas centromeric regions appeared differentiated because of low recombination and diversity, with no indication of elevated divergence. Our results support the hypothesis that clustered genomic architectures evolve with high gene flow and enhance our understanding of how inversions and centromeres are linked to different evolutionary processes.

Gene Flow↗

Data rotation improves genomotyping efficiency.

Unsequenced bacterial strains can be characterized by comparing their genomic DNA to a sequenced reference genome of the same species. This comparative genomic approach, also called genomotyping, is leading to an increased understanding of bacterial evolution and pathogenesis. It is efficiently accomplished by comparative genomic hybridization on custom-designed cDNA microarrays. The microarray experiment results in fluorescence intensities for reference and sample genome for each gene. The log-ratio of these intensities is usually compared to a cut-off, classifying each gene of the sample genome as a candidate for an absent or present gene with respect to the reference genome. Reducing the usually high rate of false positives in the list of candidates for absent genes is decisive for both time and costs of the experiment. We propose a novel method to improve efficiency of genomotyping experiments in this sense, by rotating the normalized intensity data before setting up the list of candidate genes. We analyze simulated genomotyping data and also re-analyze an experimental data set for comparison and illustration. We approximately halve the proportion of false positives in the list of candidate absent genes for the example comparative genomic hybridization experiment as well as for the simulation experiments.

Algorithms↗

Metagenomics indicates new taxa in Candidatus Saccharimonadia and proposal of Parviradicicola hetaonensis gen. nov. sp. nov. and Parviputeicola dengkouensis gen. nov. sp. nov. following the rules of the SeqCode.

Candidatus Saccharimonadia is a core lineage within the phylum Patescibacteriota (formerly the bacterial candidate phyla radiation, CPR), yet the class has long lacked a standardized, complete taxonomic framework. This nomenclatural gap severely hinders consistent academic exchange and global research into its diversity, evolutionary history, and ecological roles. Here, we recovered 29 medium- to high-quality Ca. Saccharimonadia metagenome-assembled genomes (MAGs) from groundwater, rhizosphere soil, and saline-alkali soil in the Hetao Irrigation District, Inner Mongolia, China, and performed integrated phylogenomic, genome size evolution, and metabolic analyses alongside reference genomes from the GTDB r220 database. Based on robust polyphasic taxonomic evidence (multi-dimensional phylogenetic analyses, widely accepted genome-wide ANI/AAI thresholds) and SeqCode rules, we formally propose two novel taxa: Parviradicicola hetaonensis gen. nov., sp. nov. (type material: txb011_bin.8.strictTS) and Parviputeicola dengkouensis gen. nov., sp. nov. (type material: sgl022_bin.19.origTS), plus two novel families and one novel order. We further identified potential drivers and important associations related to Ca. Saccharimonadia genome size evolution and adaptive metabolic traits. This work refines the Ca. Saccharimonadia taxonomic framework, providing critical genomic references for follow-up research.

Phylogeny↗