Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

An alignment-free strategy for circulating tumor DNA detection and tumor fraction estimation from whole-genome sequencing data.

Circulating tumor DNA (ctDNA) is emerging as a promising biomarker for postoperative monitoring of cancer patients. Precise estimation of circulating tumor fraction is crucial for evaluating treatment effects and timely detection of disease recurrence. All current ctDNA detection methods that utilize whole-genome sequencing (WGS) data rely on the reference genome alignment of sequencing reads and often apply separate tools for detecting different variant types. However, various bioinformatic analysis confounders and the application of external variant calling tools could be avoided by analyzing k-mers from unaligned sequencing reads. While k-mer-based methods have successfully been applied for somatic variant validation and detection, the potential of k-mer-based ctDNA detection is unexplored. We have developed a tumor-informed alignment-free ctDNA detection tool called ctDNAmer that detects tumor-specific somatic variation directly from unaligned sequencing data by identifying k-mers unique to the tumor DNA. ctDNAmer detects variant information across the genome by comparing the primary tumor and germline WGS data and accounts for sample-specific germline variability and technical noise in the same framework. We tested the utility of ctDNAmer for tumor fraction estimation on postoperative plasma cfDNA WGS data (mean sequencing depth ~ 28x) from 90 stage III colorectal cancer patients with three years of follow-up. The tumor fraction (TF) estimates agreed with the available clinical information and ctDNA was detected in 77% (17/22) of recurring patients with a median lead time of 8 months compared to radiological imaging. We further validated ctDNAmer's tumor fraction estimates based on a comparison with the mean cfDNA allele frequencies of somatic clonal SNVs identified from aligned primary tumor sequencing data. The TF estimates showed a strong Pearson correlation of 0.897 with the mean allele frequencies and improved ctDNA detection results across samples with an AUC of 0.79 compared to 0.75 if the mean allele frequency of clonal mutations is used.

Circulating Tumor DNA↗

Characterisation of the growth and differentiation in vivo and in vitro-of bloodstream-form Trypanosoma brucei strain TREU 927.

Trypanosoma brucei TREU 927/4 has been chosen as the reference strain targeted for complete sequencing of the genome of the African trypanosome. This line is pleomorphic in mammalian hosts and is fly transmissible; however it is relatively unstable with respect to variable surface glycoprotein (VSG) expression. Therefore, we subjected TREU 927/4 to 27 rapid syringe passages through mice, and derived a cloned line which expressed Glasgow University Trypanozoon antigen type (GUTat) 10.1 with relative stability. This line also retained pleomorphism in the bloodstream, being able to generate homogeneous populations of stumpy forms in mice. Furthermore, these parasites remain able to transform to procyclic forms synchronously in vitro and can complete their life cycle in tsetse flies. The passaged cell line was also adapted to in vitro bloodstream-form culture and transfected with a construct encoding the tetracycline repressor (TETR) protein. The resulting TETR subline no longer expressed the GUTat 10.1 VSG but remained able to generate uniform populations of stumpy form cells in mice immunocompromised with cyclophosphamide. They could also differentiate to procyclic forms synchronously in vitro. The generated lines and analyses of their growth and differentiation will provide a basic resource for the analysis and interpretation of gene function in the T. brucei genome reference strain.

Animals↗

Evaluation of sequencing reads at scale using rdeval.

MOTIVATION: Large sequencing datasets are being produced and deposited into public archives at unprecedented rates. The availability of tools that can reliably and efficiently generate and store sequencing read summary statistics has become critical. RESULTS: As part of the effort by the Vertebrate Genomes Project (VGP) to generate high-quality reference genomes at scale, we sought to address the community's need for efficient sequence data evaluation by developing rdeval, a standalone tool to quickly compute and interactively display sequencing read metrics. Rdeval can either run on the fly or store key sequence data metrics in tiny read 'snapshot' files. Statistics can then be efficiently recalled from snapshots for additional processing. Rdeval can convert fa*[.gz] files to and from other popular formats including BAM and CRAM for better compression. Overall, while CRAM achieves the best compression, the gain compared to BAM is marginal, and BAM achieves the best compromise between data compression and access speed. Rdeval also generates a detailed visual report with multiple data analytics that can be exported in various formats. We showcase rdeval's functionalities using long-read data from different sequencing platforms and species, including human. For PacBio long-read sequencing, our analysis shows dramatic improvements in both read length and quality over time, as well as the benefit of increased coverage for genome assembly, though the magnitude varies by taxa. AVAILABILITY AND IMPLEMENTATION: Rdeval is implemented in C++ for data processing and in R for data visualization. Precompiled releases (Linux, MacOS, Windows) and commented source code for rdeval are available under MIT license at https://github.com/vgl-hub/rdeval. Documentation is available on ReadTheDocs (https://rdeval-documentation.readthedocs.io). Rdeval is also available in Bioconda and in Galaxy (https://usegalaxy.org). An automated test workflow ensures the consistency of software updates.

Software↗

Persistent Genomic Erosion in Whooping Cranes Despite Demographic Recovery.

Integrating in-situ (wild) and ex-situ (captive) conservation efforts can mitigate genetic diversity loss and help prevent extinction of endangered wild populations. The whooping crane (Grus americana) experienced severe population declines in the 18th century, culminating in a collapse to ~20 individuals by 1944. Legal protections and conservation actions have since increased the census population from a stock of 16 individuals to approximately 840 individuals, yet the impact on genomic diversity remains unclear. We analysed the temporal dynamics of genomic erosion by sequencing a high-quality reference genome, and re-sequencing 16 historical (years 1867-1893) and 37 modern (2007-2020) genomes, including wild individuals and four generations of captive-bred individuals. Genomic demographic reconstructions reveal a steady decline, accelerating over the past 300 years with the European settlement of North America. Temporal genomic analyses show that despite demographic recovery, the species has lost 70% of its historical genetic diversity and has increased its inbreeding. Although the modern population bottleneck reduced the ancestral genetic load, modern populations possess more realised load than masked load, possibly resulting in a chronic loss of fitness. Integrating pedigree and genomic data, we underscore the role of breeding management in reducing recent inbreeding. Yet ongoing heterozygosity loss, load accumulation, and persistent effects of historical inbreeding (i.e., background inbreeding) argue against the species' downlisting from its current Endangered status on the IUCN Red List and the Endangered Species Act. The presence of private genetic variation in wild and captive populations suggests that wild-captive crosses could enhance genetic diversity and reduce the realised load. Our findings emphasise the role of genomics in informing conservation management and policy.

Animals↗

Advances in Cytoplasmic Male Sterility in Sugar Beet from Mitochondrial Genome Structural Dynamics and Nuclear-Cytoplasmic Coordination.

Sugar beet (Beta vulgaris L.) is a globally important sugar crop whose hybrid breeding system relies heavily on cytoplasmic male sterility (CMS) lines. Recent advances in sugar beet genomics, particularly the release of high-quality reference genomes and the characterization of organellar genomes, have provided a foundation for elucidating the molecular genetic mechanisms of CMS. Furthermore, innovations in gene editing technologies are enabling transformative functional studies in this field. The precise targeting of CMS-associated mitochondrial genes and nuclear restorer-of-fertility genes not only allows for direct investigation of theoretical models governing fertility regulation through nuclear-cytoplasmic interactions but also holds promise for the targeted development of sterile and restorer lines. This review systematically summarizes progresses in sugar beet genomics, the development of gene editing tools, and the current understanding of the molecular genetics of CMS and fertility restoration in sugar beet. Although challenges remain-such as efficient delivery of editing tools into mitochondria and coordinated editing of multiple genes-the integration of genomic and gene editing technologies is expected to accelerate multi-omics-guided dissection of CMS mechanisms. These advances will facilitate the precise design of high-yield, high-sugar, and stress-resistant sugar beet hybrids, thereby providing core scientific and technological support for the sustainable development of the global sugar industry.

Beta vulgaris↗

Chromosome-scale assembly with improved annotation provides insights into breed-wide genomic structure and diversity in domestic cats.

INTRODUCTION: Comprehensive genomic resources offer insights into biological features, including traits/disease-related genetic loci. The current reference genome assembly for the domestic cat (Felis catus), Felis_Catus_9.0 (felCat9), derived from sequences of the Abyssinian cat, may inadequately represent the general cat population, limiting the extent of deducible genetic variations. OBJECTIVES: The goal was to develop Anicom American Shorthair 1.0 (AnAms1.0), a reference-grade chromosome-scale cat genome assembly. METHODS: In contrast to prior assemblies relying on Abyssinian cat sequences, AnAms1.0 was constructed from the sequences of more popular American Shorthair breed, which is related to more breeds than the Abyssinian cat. By combining advanced genomics technologies, including PacBio long-read sequencing and Hi-C- and optical mapping data-based sequence scaffolding, we compared AnAms1.0 to existing Felidae genome assemblies (20 scaffolds, scaffolds N50 > 150 Mbp). Homology-based and ab initio gene annotation through Iso-Seq and RNA-Seq was used to identify new coding genes and splice variants. RESULTS: AnAms1.0 demonstrated superior contiguity and accuracy than existing Felidae genome assemblies. Using AnAms1.0, we identified over 1.5 thousand structural variants and 29 million repetitions compared to felCat9. Additionally, we identified > 1,600 novel protein-coding genes. Notably, olfactory receptor structural variants and cardiomyopathy-related variants were identified. CONCLUSION: AnAms1.0 facilitates the discovery of novel genes related to normal and disease phenotypes in domestic cats. The analyzed data are publicly accessible on Cats-I (https://cat.annotation.jp/), which we established as a platform for accumulating and sharing genomic resources to discover novel genetic traits and advance veterinary medicine.

Animals↗

Plant genome databases: from references to inference tools.

Plant genome databases play an important role in the archiving and dissemination of data arising from the international genome projects. Recent developments in bioinformatics, such as new software tools, programming languages and standards, have produced better access across the Internet to the data held within them. An increasing emphasis is placed on data analysis and indeed many resources now provide tools allied to the databases, to aid in the analysis and interpretation of the data. However, a considerable wealth of information lies untapped by considering the databases as single entities and will only be exploited by linking them with a wide range of data sources. Data from research programs such as comparative mapping and germplasm studies may be used as tools, to gain additional knowledge but without additional experimentation. To date, the current plant genome databases are not yet linked comprehensively with each other or with these additional resources, although they are clearly moving toward this. Here, the current wealth of public plant genome databases is reviewed, together with an overview of initiatives underway to bind them to form a single plant genome infrastructure.

Computational Biology↗

Deletions below 10 megabasepairs are detected in comparative genomic hybridization by standard reference intervals.

Comparative genomic hybridization (CGH) is a widely used technique for studying chromosomal imbalances. The sensitivity of the technique is, however, relatively low. Deletions down to a size of 10-12 Mbp have been detected by the use of fixed diagnostic thresholds. In this study, we applied standard reference intervals as detection criteria on a number of deletions in the range of 3 Mbp to 14-18 Mbp. All deletions were detected. Thus, detection by standard reference intervals confers a considerably higher sensitivity to CGH analysis compared to fixed diagnostic thresholds. Genes Chromosomes Cancer 25:410-413, 1999.

Chromosome Aberrations↗

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals↗

Detection of aneuploidy in single cells using comparative genomic hybridization.

The ability of comparative genomic hybridization (CGH) to detect aneuploidy following universal amplification of DNA from a single cell, or a small number of cells, was investigated with a view to preimplantation diagnosis following in vitro fertilization, and prenatal diagnosis using fetal erythroblasts obtained from maternal blood. The DNA obtained from lysed single cells was amplified using degenerate oligonucleotide-primed PCR (DOP-PCR). This product was labelled using nick translation and hybridized together with normal reference genomic DNA. The CGH fluorescent ratio profiles obtained could be used to determine aneuploidy with cut-off thresholds of 0.75 and 1.25. Deviation in the profiles in the heterochromatic regions was reduced by using, as a reference sample, normal genomic DNA that had also undergone DOP-PCR. Single cells known to be trisomic for chromosomes 13, 18 or 21 were analysed using this technique. The resolution of CGH with amplified DNA from a single cell is of the order of 40 Mb, sufficient for the diagnosis of trisomy 21, and possibly segmental aneuploidy of equivalent size. These results, and those of others, demonstrate that diagnosis of chromosomal aneuploidy in single cells is possible using CGH with DOP-PCR amplified DNA.

Aneuploidy↗

Analysis of the compositional biases in Plasmodium falciparum genome and proteome using Arabidopsis thaliana as a reference.

Comparative genomic analysis of the malaria causative agent, Plasmodium falciparum, with other eukaryotes for which the complete genome is available, revealed that the genome from P. falciparum was more similar to the genome of a plant, Arabidopsis thaliana, than to other non-apicomplexan taxa. Plant-like sequences are thought to result from horizontal gene transfers after a secondary endosymbiosis involving an algal ancestor. The use of the A. thaliana genome and proteome as a reference gives an opportunity to refine our understanding of the extreme compositional bias in the P. falciparum genome that leads to a proteome-wide amino acid bias. A set of pairs of non-redundant protein homologues was selected owing to rigorous genome-wide sequence comparison methods. The introduction of A. thaliana as a reference was a mean to weight the magnitude of the protein evolutionary divergence in P. falciparum. The correlation of the amino acid proportions with evolutionary time supports the hypothesis that amino acids encoded by GC-rich codons are directionally substituted into amino acids encoded by AT-rich codons in the P. falciparum proteome. The long-term deviation of codons in malarial sequences appears as a possible consequence of a genome-wide tri-nucleotidic signature imprinting. Additionally, this study suggests possible working guidelines to improve the accuracy of P. falciparum sequence comparisons, for homology searches and phylogenetic studies.

AT Rich Sequence↗

Rapid recent growth and divergence of rice nuclear genomes.

By employing the nuclear DNA of the African rice Oryza glaberrima as a reference genome, the timing, natures, mechanisms, and specificities of recent sequence evolution in the indica and japonica subspecies of Oryza sativa were identified. The data indicate that the genome sizes of both indica and japonica have increased substantially, >2% and >6%, respectively, since their divergence from a common ancestor, mainly because of the amplification of LTR-retrotransposons. However, losses of all classes of DNA sequence through unequal homologous recombination and illegitimate recombination have attenuated the growth of the rice genome. Small deletions have been particularly frequent throughout the genome. In >1 Mb of orthologous regions that we analyzed, no cases of complete gene acquisition or loss from either indica or japonica were found, nor was any example of precise transposon excision detected. The sequences between genes were observed to have a very high rate of divergence, indicating a molecular clock for transposable elements that is at least 2-fold more rapid than synonymous base substitutions within genes. We found that regions prone to frequent insertions and deletions also exhibit higher levels of point mutation. These results indicate a highly dynamic rice genome with competing processes for the generation and removal of genetic variation.

Base Sequence↗

Analysis of primate genomic variation reveals a repeat-driven expansion of the human genome.

We performed a detailed analysis of both single-nucleotide and large insertion/deletion events based on large-scale comparison of 10.6 Mb of genomic sequence from lemur, baboon, and chimpanzee to human. Using a human genomic reference, optimal global alignments were constructed from large (>50-kb) genomic sequence clones. These alignments were examined for the pattern, frequency, and nature of mutational events. Whereas rates of single-nucleotide substitution remain relatively constant (1-2 x 10(-9) substitutions/site/year), rates of retrotransposition vary radically among different primate lineages. These differences have lead to a 15%-20% expansion of human genome size over the last 50 million years of primate evolution, 90% of it due to new retroposon insertions. Orthologous comparisons with the chimpanzee suggest that the human genome continues to significantly expand due to shifts in retrotransposition activity. Assuming that the primate genome sequence we have sampled is representative, we estimate that human euchromatin has expanded 30 Mb and 550 Mb compared to the primate genomes of chimpanzee and lemur, respectively.

Animals↗

Draft genome sequence of Breoghania corrubedonensis DSM 23382T.

We report the genome sequence of Breoghania corrubedonensis DSM 23382T isolated from oil-spill contaminated beach sand. The 5,331,589-bp genome with 63.62% G + C encodes 4,746 genes. This reference genome will facilitate experimental studies investigating B. corrubedonensis's role in oil-contaminated marine environments, particularly oil degradation or resistivity.

computational biology↗

The cost and cost trajectory of genome sequencing and bioinformatics analysis for Indigenous children with suspected rare diseases.

PURPOSE: Indigenous peoples are underrepresented in reference genome libraries. Consequently, rare disease diagnosis may require bespoke bioinformatics analyses of genome sequences. Establishing diagnostic cost is crucial to support policy development for equitable diagnosis of rare diseases. We estimated the cost and cost trajectory of diagnostic genome sequencing and bioinformatics for Indigenous participants with suspected rare diseases. METHODS: We conducted a microcosting study of Indigenous children and their families receiving genome sequencing through Canada's Silent Genomes Project. Invoice data informed the costs of genome sequencing. We conducted a time-and-motion study for bioinformatics analyses, including labor, computing, and data storage costs. RESULTS: With standard bioinformatics, costs ranged from C$3645 (SD: 455) for singletons to C$7402 (SD: 566) for trios. With advanced, bespoke bioinformatics, costs ranged from C$5344 (SD: 634) for singletons to C$9760 (SD: 822) for trios. Genome sequencing was a primary cost driver; however, sequencing costs decreased by 61% over 4 years. Bioinformatics costs ranged from 21.3% to 58.3% of the total costs. The time required for bioinformatics ranged from 71 hours to 215 hours for standard and advanced analyses, respectively. CONCLUSION: Genome sequencing costs decreased over time. Bioinformatics is a significant cost driver, particularly for bespoke analyses arising from nonrepresentative reference libraries.

Humans↗

What is the future of electrophoresis in large-scale genomic sequencing?

Although a finished human genome reference sequence is now available, the ability to sequence large, complex genomes remains critically important for researchers in the biological sciences, and in particular, continued human genomic sequence determination will ultimately help to realize the promise of medical care tailored to an individual's unique genetic identity. Many new technologies are being developed to decrease the costs and to dramatically increase the data acquisition rate of such sequencing projects. These new sequencing approaches include Sanger reaction-based technologies that have electrophoresis as the final separation step as well as those that use completely novel, nonelectrophoretic methods to generate sequence data. In this review, we discuss the various advances in sequencing technologies and evaluate the current limitations of novel methods that currently preclude their complete acceptance in large-scale sequencing projects. Our primary goal is to analyze and predict the continuing role of electrophoresis in large-scale DNA sequencing, both in the near and longer term.

Animals↗

Fugu and human sequence comparison identifies novel human genes and conserved non-coding sequences.

The compact genome of the pufferfish, Fugu rubripes, has been proposed as a 'reference' genome to aid in annotating and analysing the human genome. We have annotated and compared 85 kb of Fugu sequence containing 17 genes with its homologous loci in the human draft genome and identified three 'novel' human genes that were missed or incompletely predicted by the previous gene prediction methods. Two of the novel genes contain zinc finger domains and are designated ZNF366 and ZNF367. They map to human chromosomes 5q13.2 and 9q22.32, respectively. The third novel gene, designated C9orf21, maps to chromosome 9q22.32. This gene is unique to vertebrates, and the protein encoded by it does not contain any known domains. We could not find human homologs for two Fugu genes, a novel chemokine gene and a kinase gene. These genes are either specific to teleosts or lost in the human lineage. The Fugu-human comparison identified several conserved non-coding sequences in the promoter and intronic regions. These sequences, conserved during 450 million years of vertebrate evolution, are likely to be involved in gene regulation. The 85 kb Fugu locus is dispersed over four human loci, occupying about 1.5 Mb. Contiguity is conserved in the human genome between six out of 16 Fugu gene pairs. These contiguous chromosomal segments should share a common evolutionary history dating back to the common ancestor of mammals and teleosts. We propose contiguity as strong evidence to identify orthologous genes in distant organisms. This study confirms the utility of the Fugu as a supplementary tool to uncover and confirm novel genes and putative gene regulatory regions in the human genome.

Amino Acid Sequence↗

Forty new genomes shed light on sexual reproduction and the origin of tetraploidy in Microsporidia.

Microsporidia are single-celled, obligately intracellular parasites with growing public health, agricultural, and economic importance. Despite this, Microsporidia remain relatively enigmatic, with many aspects of their biology and evolution unexplored. Key questions include whether Microsporidia undergo sexual reproduction, and the nature of the relationship between tetraploid and diploid lineages. While few high-quality microsporidian genomes currently exist to help answer such questions, large-scale biodiversity genomics initiatives, such as the Darwin Tree of Life project, can generate high-quality genome assemblies for microsporidian parasites when sequencing infected host species. Here, we present 40 new microsporidian genome assemblies from infected arthropod hosts that were sequenced to create reference genomes. Out of the 40, 32 are complete genomes, eight of which are chromosome-level, and eight are partial microsporidian genomes. We characterized 14 of these as polyploid and five as diploid. We found that tetraploid genome haplotypes are consistent with autopolyploidy, in that they coalesce more recently than species, and that they likely recombine. Within some genomes, we found large-scale rearrangements between the homeologous genomes. We also observed a high rate of rearrangement between genomes from different microsporidian groups, and a striking tolerance for segmental duplications. Analysis of chromatin conformation capture (Hi-C) data indicated that tetraploid genomes are likely organized into two diploid units, similar to dikaryotic cells in fungi, with evidence of recombination within and between units. Together, our results provide evidence for the existence of a sexual cycle in Microsporidia, and suggest a model for the microsporidian lifecycle that mirrors fungal reproduction.

Genome, Fungal↗