Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.

NCBI's reference sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) is a curated non-redundant collection of sequences representing genomes, transcripts and proteins. The database includes 3774 organisms spanning prokaryotes, eukaryotes and viruses, and has records for 2,879,860 proteins (RefSeq release 19). RefSeq records integrate information from multiple sources, when additional data are available from those sources and therefore represent a current description of the sequence and its features. Annotations include coding regions, conserved domains, tRNAs, sequence tagged sites (STS), variation, references, gene and protein product names, and database cross-references. Sequence is reviewed and features are added using a combined approach of collaboration and other input from the scientific community, prediction, propagation from GenBank and curation by NCBI staff. The format of all RefSeq records is validated, and an increasing number of tests are being applied to evaluate the quality of sequence and annotation, especially in the context of complete genomic sequence.

Amino Acid Sequence↗

A method for genome comparisons and hybridization studies using known megabase-scale DNA sequences as a reference.

We present a method for genome comparisons and high-resolution hybridization analyses using megabase stretches of known DNA sequences as a reference. The method employs two-dimensional gel electrophoresis, separating genomic segments cut with different restriction endonucleases in the first and second dimensions, to generate filters suitable for image analysis and repeated nucleic acid hybridizations. The corresponding two-dimensional pattern is computed from the reference nucleotide sequence and matched to the observed pattern, thereby identifying each fragment on the filter; at the same time the technique uncovers discrepancies from the reference sequence. This permits genome comparisons as well as automated identification and quantification of hybridization patterns with various probes. The technique is illustrated by an analysis of Saccharomyces cerevisiae chromosome IX.

Chromosome Walking↗

Genomic biomarkers for cancer assessment: implementation challenges for laboratory practice.

Genomic biomarkers are an emerging class of laboratory tests, which present special implementation challenges for clinical laboratory services, compared to conventional laboratory tests. These challenges, which include analytical, bioinformatics, bioethical, interpretation and commercialization issues, represent real obstacles to widespread implementation of these tests. Technical challenges include the capacity to detect and identify many different kinds of markers for different diseases in a short time period, capacity to identify simultaneously gene rearrangements, amplification, inhibition, deletions and replications. Bioinformatics challenges include rapid analysis of genomic data, as well as the cross reference to other genomic data, and to other laboratory tests. Bioethical issues relate to consent to retain and use genetic data, which may be obtained inadvertently during analysis for genomic markers. Interpretation challenges include observations that the particular genomic markers may not be independent variables, as other undetected genomic alterations could invalidate or alter genomic marker interpretation. Further, as early experience with predictive genetic markers for cancer has shown, proprietary commercial interests may conflict with public health values of identifying genomic markers in subject populations. Based on our 10 years of experience with genomic biomarkers, important implementation strategies for genomic markers include development of:Standard high throughput analyzers capable of detecting any alteration of any genomic variant at any time. Bioinformatics analysis online, coupled to stored patient data. Laboratory service framework that preserves confidentiality but integrates genomic data with other laboratory tests. Laboratory service framework, which links consents, genomic analysis, reports to both specimen and data repositories. Overall, the laboratory service challenges for genomic markers are to manage very large analytical sets and very large data sets in finite time with responsible interpretation, all within finite funding. To meet these challenges, implementation strategies beyond the one disease, one diagnosis, one genomic marker concept must begin now.

Biomarkers, Tumor↗

Whole-genome sequence variation among multiple isolates of Pseudomonas aeruginosa.

Whole-genome shotgun sequencing was used to study the sequence variation of three Pseudomonas aeruginosa isolates, two from clonal infections of cystic fibrosis patients and one from an aquatic environment, relative to the genomic sequence of reference strain PAO1. The majority of the PAO1 genome is represented in these strains; however, at least three prominent islands of PAO1-specific sequence are apparent. Conversely, approximately 10% of the sequencing reads derived from each isolate fail to align with the PAO1 backbone. While average sequence variation among all strains is roughly 0.5%, regions of pronounced differences were evident in whole-genome scans of nucleotide diversity. We analyzed two such divergent loci, the pyoverdine and O-antigen biosynthesis regions, by complete resequencing. A thorough analysis of isolates collected over time from one of the cystic fibrosis patients revealed independent mutations resulting in the loss of O-antigen synthesis alternating with a mucoid phenotype. Overall, we conclude that most of the PAO1 genome represents a core P. aeruginosa backbone sequence while the strains addressed in this study possess additional genetic material that accounts for at least 10% of their genomes. Approximately half of these additional sequences are novel.

Adolescent↗

CoreGenes: a computational tool for identifying and cataloging "core" genes in a set of small genomes.

BACKGROUND: Improvements in DNA sequencing technology and methodology have led to the rapid expansion of databases comprising DNA sequence, gene and genome data. Lower operational costs and heightened interest resulting from initial intriguing novel discoveries from genomics are also contributing to the accumulation of these data sets. A major challenge is to analyze and to mine data from these databases, especially whole genomes. There is a need for computational tools that look globally at genomes for data mining. RESULTS: CoreGenes is a global JAVA-based interactive data mining tool that identifies and catalogs a "core" set of genes from two to five small whole genomes simultaneously. CoreGenes performs hierarchical and iterative BLASTP analyses using one genome as a reference and another as a query. Subsequent query genomes are compared against each newly generated "consensus." These iterations lead to a matrix comprising related genes from this set of genomes, e. g., viruses, mitochondria and chloroplasts. Currently the software is limited to small genomes on the order of 330 kilobases or less. CONCLUSION: A computational tool CoreGenes has been developed to analyze small whole genomes globally. BLAST score-related and putatively essential "core" gene data are displayed as a table with links to GenBank for further data on the genes of interest. This web resource is available at http://pumpkins.ib3.gmu.edu:8080/CoreGenes or http://www.bif.atcc.org/CoreGenes.

Algorithms↗

Use of reference libraries and hybridisation fingerprinting for relational genome analysis.

The concept of relational genome analysis by hybridisation has been developed into a working system. Various genomic and cDNA libraries have been generated and are distributed via a reference system. Analysis procedures have been tested successfully in the mapping of the entire Schizosaccharomyces pombe genome. In another test-case for their refinement, analyses on the Drosophila genome are well under way. Human and mouse libraries are being studied on all levels, from generating YAC maps to partially sequencing representative cDNA libraries. The automation of the involved processes and the development of improved image detection and analysis are well advanced.

Animals↗

Using Mapping-Profiles to Refine Strain-Level Metagenomic Classification.

Metagenomic classification at the strain level remains challenging due to high sequence similarity among closely related genomes, which leads to ambiguous read mappings and frequent false-positive strain detections. Reducing such errors improves the reliability of strain-level analyses, which is critical for applications such as pathogen detection. We introduce StrainRefine, a post-mapping refinement method that analyzes read-reference mapping profiles to resolve ambiguous assignments among highly similar genomes. The method represents candidate reference genomes using binary profiles that capture read-support patterns and measures similarity between references based on profile overlap. The method clusters references based on similar mapping profiles, filters weakly supported genomes, and reassigns reads to representative references, reducing redundant reporting of near-identical strains. StrainRefine substantially reduces false-positive strain detections while preserving recall and improving agreement between predicted and true abundance profiles. On large-scale metagenomic datasets, it achieves a substantially improved precision-recall balance compared with existing mapping-based approaches, with the standalone method obtaining the highest read-level classification accuracy on the most complex evaluated dataset. Unlike many strain-level tools designed for individual species, StrainRefine operates without prior assumptions about sample composition or curated species-specific reference collections, while still achieving comparable performance in single-species settings on species-specific reference databases. These results highlight mapping-profile similarity as an effective signal for improving strain-level metagenomic classification.

false-positive reduction↗

HinCyc: a knowledge base of the complete genome and metabolic pathways of H. influenzae.

We present a methodology for predicting the metabolic pathways of an organism from its genomic sequence by reference to a knowledge base of known metabolic pathways. We applied these techniques to the genome of H. influenzae by reference to the EcoCyc knowledge base to predict which of 81 metabolic pathways of E. coli are found in H. influenzae. The resulting prediction is a complex hypothesis that is presented in computer form as HinCyc: an electronic encyclopedia of the genes and metabolic pathways of H. influenzae. HinCyc connects the predicted genes, enzymes, enzyme-catalyzed reactions, and biochemical pathways in a WWW-accessible knowledge base to allow scientists to explore this complex hypothesis.

Computer Communication Networks↗

Experimental analysis of the annotation of promoters in the public database.

The ability to identify and examine promoter elements is important to researchers who wish to understand how gene expression is regulated in normal and pathological states. Unfortunately, the number of human promoters that have been directly experimentally defined is small. In order to determine if promoter sequences can be identified by simply aligning mRNA and genomic sequences, we have used a reporter gene assay to assess the promoter activity of the immediate 5' region flanking 38 mRNAs mapping to chromosome 21. For comparison, we have measured the activities of 19 sequences not thought to be promoters and 39 sequences taken from the Eukaryotic Promoter Database. Our results suggest that alignment of reference mRNAs to genomic sequence allows promoters to be identified for at least 75% of genes. These data provide the first empirical evidence that the current state of annotation of the genome is sufficient to allow molecular geneticists to correctly identify promoter sequences for most genes for which reference mRNA and genomic sequences are available.

Cell Line↗

Statistical methods for detecting genomic alterations through array-based comparative genomic hybridization (CGH).

Array-based comparative genomic hybridization (ABCGH) is an emerging high-resolution and high-throughput molecular genetic technique that allows genome-wide screening for chromosome alterations associated with tumorigenesis. Like the cDNA microarrays, ABCGH uses two differentially labeled test and reference DNAs which are cohybridized to cloned genomic fragments immobilized on glass slides. The hybridized DNAs are then detected in two different fluorochromes, and the significant deviation from unity in the ratios of the digitized intensity values is indicative of copy-number differences between the test and reference genomes. Proper statistical analyses need to account for many sources of variation besides genuine differences between the two genomes. In particular, spatial correlations, the variable nature of the ratio variance and non-Normal distribution call for careful statistical modeling. We propose two new statistics, the standard t-statistic and its modification with variances smoothed along the genome, and two tests for each statistic, the standard t-test and a test based on the hybrid adaptive spline (HAS). Simulations indicate that the smoothed t-statistic always improves the performance over the standard t-statistic. The t-tests are more powerful in detecting isolated alterations while those based on HAS are more powerful in detecting a cluster of alterations. We apply the proposed methods to the identification of genomic alterations in endometrium in women with endometriosis.

Chromosome Aberrations↗

The complete telomere-to-telomere sequence of a mouse Y chromosome.

The mouse Y chromosome is essential for male reproduction, yet the GRCm39 reference contains 25 gaps, particularly in repetitive and complex regions. Here, we assembled a telomere-to-telomere Y chromosome (mT2T Y) of 95.21 Mb from a C57BL/6 mouse incorporating parental genomes. This assembly fills all gaps, corrects structural errors, and adds over 8.70 Mb of previously unassembled sequence to the reference genome. We annotated 142 previously unidentified genes, identified Y specific satellite arrays, and mapped homologous recombination loci in the pseudoautosomal region (PAR). Analysis of X Y homologous gene expression revealed a Y chromosome dosage compensation mechanism. By combining mT2T Y with T2T mhaESC, we completed the T2T assembly of all C57BL/6 chromosomes, designated T2T mhaESC+Y, providing a complete C57BL/6 reference genome.

Animals↗

DNA-DNA hybridization study of Burkholderia species using genomic DNA macro-array analysis coupled to reverse genome probing.

The present study was aimed at simplifying procedures to delineate species and identify isolates based on DNA-DNA reassociation. DNA macro-arrays harbouring genomic DNA of reference strains of several Burkholderia species were produced. Labelled genomic DNA, hybridized to such an array, allowed multiple relative pairwise comparisons. Based on the relative DNA-DNA relatedness values, a complete data matrix was constructed and the ability of the method to discriminate strains belonging to different species was assessed. This simple approach led successfully to the discrimination of Burkholderia mallei from Burkholderia pseudomallei, but also discriminated Burkholderia cepacia genomovars I and III, Burkholderia multivorans, Burkholderia pyrrocinia, Burkholderia stabilis and Burkholderia vietnamiensis. Present data showed a sufficient degree of congruence with previous DNA-DNA reassociation techniques. As part of a polyphasic taxonomic scheme, this straightforward approach is proposed to improve species definition, especially for application in the rapid screening necessary for large numbers of clinical or environmental isolates.

Bacterial Typing Techniques↗

Genes, genomes and identity. Projections on matter.

This paper aims to show that references to genes and genomes are counterproductive in legal and political understandings of what it is to be human and a unique individual. To support this claim, I will give a brief overview of the many incompatible meanings the term 'identity' has gathered in reference to genes or genome in the contexts of biology and family ancestry, personal identity, species identity. One finds various and incompatible understandings of these expressions. While genetics is usually considered to deliver definitive knowledge about history and the future, genomics seems to work with more complicated relations between DNA, inheritance and phenotype. In genomics, 'identity' is no longer about identification and status markers but about individualization. Regulatory and legal documents project from traits to genomes, implying that individuality is at least represented, if not created, in a unique genome. Boundaries between humans and other animals, between different 'kinds' of humans, and between all individual humans are re-established via reference to the chemical matter of DNA. My analysis will show how this trend is a reactionary response to modern understandings of identities as social products and that it ignores new biomedical understandings of human bodies.

Animals↗

Multiplexed discovery of sequence polymorphisms using base-specific cleavage and MALDI-TOF MS.

The completion of the Human Genome Project provides researchers with a reference sequence that covers about 99% of the gene-containing regions and is more than 99.9% accurate. Sequence drafts and completed sequences for several other species are also available to researchers worldwide. The ongoing effort to provide more and more genomic reference information now enables the detection of deviations from this 'genetic blueprint'. Comparative sequencing projects will play a major role in elucidating the meaning of the genetic code and in establishing a correlation between genotype and phenotype. As part of this effort, a number of projects will focus on distinct functional aspects, like resequencing of exons or HLA determining regions. Typically these target regions are short in length and their analysis does not require long read length. To find an efficient solution for these applications, we developed a novel method that allows simultaneous analysis of multiple independent target regions (Multiplexed Comparative Sequence Analysis) by employing base-specific cleavage biochemistry and MALDI TOF-MS analysis.

Humans↗

Proceedings of the SMBE Tri-National Young Investigators' Workshop 2005. Insight into the diversity and evolution of the cryptomonad nucleomorph genome.

The cryptomonads are an enigmatic group of marine and freshwater unicellular algae that acquired their plastids through the engulfment and retention of a eukaryotic ("secondary") endosymbiont. Together with the chlorarachniophyte algae, the cryptomonads are unusual in that they have retained the nucleus of their endosymbiont in a miniaturized form called a nucleomorph. The nucleomorph genome of the cryptomonad Guillardia theta has been completely sequenced and with only three chromosomes and a total size of 551 kb, is a model of nuclear genome compaction. Using this genome as a reference, we have investigated the structure and content of nucleomorph genomes in a wide range of cryptomonad algae. In this study, we have sequenced nine new cryptomonad nucleomorph 18S ribosomal DNA (rDNA) genes and four heat shock protein 90 (hsp90) gene fragments, and using pulsed-field gel electrophoresis and Southern hybridizations, have obtained nucleomorph genome size estimates for nine different species. We also used long-range polymerase chain reaction to obtain nucleomorph genomic fragments from Hanusia phi CCMP325 and Proteomonas sulcata CCMP704 that are syntenic with the subtelomeric region of nucleomorph chromosome I in G. theta. Our results indicate that (1) the presence of three chromosomes is a common feature of the nucleomorph genomes of these organisms, (2) nucleomorph genome size varies dramatically in the cryptomonads examined, (3) unidentified cryptomonad species CCMP1178 has the largest nucleomorph genome identified to date at approximately 845 kb, (4) nucleomorph genome size reductions appear to have occurred multiple times independently during cryptomonad evolution, (5) the relative positions of the 18S rDNA, ubc4, and hsp90 genes are conserved in three different cryptomonad genera, and (6) interchromosomal recombination appears to be rapidly changing the size and sequence of a repetitive subtelomeric region of the nucleomorph genome between the 18S rDNA and ubc4 loci. These results provide a glimpse into the genetic diversity of nucleomorph genomes in cryptomonads and set the stage for more comprehensive sequence-based studies in closely and distantly related taxa.

Bayes Theorem↗

B-lymphoblastoid cell lines as a source of reference DNA for human platelet and neutrophil antigen genotyping.

BACKGROUND: Human platelet and neutrophil antigens (HPAs, HNAs) are targets for platelet or granulocyte antibodies causing immune thrombocytopenia or neutropenia, respectively. Currently, genotyping is replacing phenotyping as the preferred method of diagnosis of immune cytopenia. To establish a reliable genotyping analysis, however, the availability as reference DNA of genomic DNA from persons of known genotype is essential. STUDY DESIGN AND METHODS: By the use of Epstein-Barr virus transformation, panels of B-lympho-blastoid cell lines (B-LCLs) from HPA- and HNA-phenotyped individuals were developed. Genomic DNA was isolated from these cell lines and tested as reference DNA for genotyping of persons for HPAs and HNAs. RESULTS: DNA derived from these B-LCLs was typed by polymerase chain reaction-restriction fragment length polymorphism and -sequence-specific primers. The results were in accordance with the genotyping from peripheral blood cells. These results were confirmed by 24 laboratories in Germany in a blind study. CONCLUSION: The inexhaustible source of reference DNA derived from B-LCLs allowed the evaluation of reliable HPA and HNA genotyping for quality control purposes. It should facilitate the development of DNA typing in blood centers and clinical laboratories.

Antigens↗

Reference-Free Variant Calling with Local Graph Construction with ska lo (SKA).

The study of genomic variants is increasingly important for public health surveillance of pathogens. Traditional variant-calling methods from whole-genome sequencing data rely on reference-based alignment, which can introduce biases and require significant computational resources. Alignment- and reference-free approaches offer an alternative by leveraging k-mer-based methods, but existing implementations often suffer from sensitivity limitations, particularly in high mutation density genomic regions. Here, we present ska lo, a graph-based algorithm that aims to identify within-strain variants in pathogen whole-genome sequencing data by traversing a colored De Bruijn graph and building variant groups (i.e. sets of variant combinations). Through in silico benchmarking and real-world dataset analyses, we demonstrate that ska lo achieves high sensitivity in single-nucleotide polymorphism (SNP) calls while also enabling the detection of insertions and deletions, as well as SNP positioning on a reference genome for recombination analyses. These findings highlight ska lo as a simple, fast, and effective tool for pathogen genomic epidemiology, extending the range of reference-free variant-calling approaches. ska lo is freely available as part of the SKA program (https://github.com/bacpop/ska.rust).

Polymorphism, Single Nucleotide↗