Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

International Rice Genome Sequencing Project: the effort to completely sequence the rice genome.

The International Rice Genome Sequencing Project (IRGSP) involves researchers from ten countries who are working to completely and accurately sequence the rice genome within a short period. Sequencing uses a map-based clone-by-clone shotgun strategy; shared bacterial artificial chromosome/P1-derived artificial chromosome libraries have been constructed from Oryza sativa ssp. japonica variety 'Nipponbare'. End-sequencing, fingerprinting and marker-aided PCR screening are being used to make sequence-ready contigs. Annotated sequences are immediately released for public use and are made available with supplemental information at each IRGSP member's website. The IRGSP works to promote the development of rice and cereal genomics in addition to producing genome sequence data.

Chromosome Mapping↗

Complete genome sequence and comparative genomic analysis of an emerging human pathogen, serotype V Streptococcus agalactiae.

The 2,160,267 bp genome sequence of Streptococcus agalactiae, the leading cause of bacterial sepsis, pneumonia, and meningitis in neonates in the U.S. and Europe, is predicted to encode 2,175 genes. Genome comparisons among S. agalactiae, Streptococcus pneumoniae, Streptococcus pyogenes, and the other completely sequenced genomes identified genes specific to the streptococci and to S. agalactiae. These in silico analyses, combined with comparative genome hybridization experiments between the sequenced serotype V strain 2603 V/R and 19 S. agalactiae strains from several serotypes using whole-genome microarrays, revealed the genetic heterogeneity among S. agalactiae strains, even of the same serotype, and provided insights into the evolution of virulence mechanisms.

Amino Acid Sequence↗

Genome sequencing and population genomics provide insights into the demographic history, genetic load, and local adaptation of an endangered Tertiary relict.

Endangered Tertiary relict trees represent an exceptional evolutionary heritage with small and isolated populations, yet little is known about how demographic history, local adaptation, and genetic load have affected their long-term survival and extinction risk. We performed whole-genome sequencing and population genomic analyses on Ulmus elongata L. K. Fu & C. S. Ding, an endangered Tertiary relict tree endemic to East Asia. By integrating genomes from U. elongata and seven other endangered trees from public databases, we identified rate-decelerated genes across endangered trees and genes under positive selection of U. elongata associated with tissue development, detoxification, and immune response, and signal transduction and regulation mechanisms potentially leading to endangered status. Demographic analyses revealed continuous population decline from the late Miocene to present, especially during the last glacial maximum (LGM) and last 10&#x2009;000&#x2009;years. Spearman correlation indicated a strong negative relationship between effective population size and human population density (rpopulation density&#x2009;=&#x2009;-0.90, P&#x2009;<&#x2009;0.001) as well as cropland use (rcropland use&#x2009;=&#x2009;-0.89, P&#x2009;<&#x2009;0.001). Genotype-environment association (GEA) analyses identified a set of candidate genes associated with temperature and precipitation, supporting a polygenic adaptation model in U. elongata. Overall, our findings underscore the severe population bottlenecks that have led to the fixation of strongly deleterious mutations and inbreeding, further compromising the adaptive potential and long-term viability of U. elongata. Furthermore, assessments of genomic vulnerability under future climate scenarios revealed higher genetic offsets in northern region of Fujian and Jiangxi populations, suggesting these regions require prioritized conservation efforts due to reduced adaptive capacity.

Endangered Species↗

Fast and sensitive multiple alignment of large genomic sequences.

BACKGROUND: Genomic sequence alignment is a powerful method for genome analysis and annotation, as alignments are routinely used to identify functional sites such as genes or regulatory elements. With a growing number of partially or completely sequenced genomes, multiple alignment is playing an increasingly important role in these studies. In recent years, various tools for pair-wise and multiple genomic alignment have been proposed. Some of them are extremely fast, but often efficiency is achieved at the expense of sensitivity. One way of combining speed and sensitivity is to use an anchored-alignment approach. In a first step, a fast search program identifies a chain of strong local sequence similarities. In a second step, regions between these anchor points are aligned using a slower but more accurate method. RESULTS: Herein, we present CHAOS, a novel algorithm for rapid identification of chains of local pair-wise sequence similarities. Local alignments calculated by CHAOS are used as anchor points to improve the running time of DIALIGN, a slow but sensitive multiple-alignment tool. We show that this way, the running time of DIALIGN can be reduced by more than 95% for BAC-sized and longer sequences, without affecting the quality of the resulting alignments. We apply our approach to a set of five genomic sequences around the stem-cell-leukemia (SCL) gene and demonstrate that exons and small regulatory elements can be identified by our multiple-alignment procedure. CONCLUSION: We conclude that the novel CHAOS local alignment tool is an effective way to significantly speed up global alignment tools such as DIALIGN without reducing the alignment quality. We likewise demonstrate that the DIALIGN/CHAOS combination is able to accurately align short regulatory sequences in distant orthologues.

Algorithms↗

Complete genome sequence and comparative genomics of Shigella flexneri serotype 2a strain 2457T.

We determined the complete genome sequence of Shigella flexneri serotype 2a strain 2457T (4,599,354 bp). Shigella species cause >1 million deaths per year from dysentery and diarrhea and have a lifestyle that is markedly different from those of closely related bacteria, including Escherichia coli. The genome exhibits the backbone and island mosaic structure of E. coli pathogens, albeit with much less horizontally transferred DNA and lacking 357 genes present in E. coli. The strain is distinctive in its large complement of insertion sequences, with several genomic rearrangements mediated by insertion sequences, 12 cryptic prophages, 372 pseudogenes, and 195 S. flexneri-specific genes. The 2457T genome was also compared with that of a recently sequenced S. flexneri 2a strain, 301. Our data are consistent with Shigella being phylogenetically indistinguishable from E. coli. The S. flexneri-specific regions contain many genes that could encode proteins with roles in virulence. Analysis of these will reveal the genetic basis for aspects of this pathogenic organism's distinctive lifestyle that have yet to be explained.

Base Sequence↗

Progress in Arabidopsis genome sequencing and functional genomics.

Arabidopsis thaliana has a relatively small genome of approximately 130 Mb containing about 10% repetitive DNA. Genome sequencing studies reveal a gene-rich genome, predicted to contain approximately 25000 genes spaced on average every 4.5 kb. Between 10 to 20% of the predicted genes occur as clusters of related genes, indicating that local sequence duplication and subsequent divergence generates a significant proportion of gene families. In addition to gene families, repetitive sequences comprise individual and small clusters of two to three retroelements and other classes of smaller repeats. The clustering of highly repetitive elements is a striking feature of the A. thaliana genome emerging from sequence and other analyses.

Agriculture↗

Genome BLAST distance phylogenies inferred from whole plastid and whole mitochondrion genome sequences.

BACKGROUND: Phylogenetic methods which do not rely on multiple sequence alignments are important tools in inferring trees directly from completely sequenced genomes. Here, we extend the recently described Genome BLAST Distance Phylogeny (GBDP) strategy to compute phylogenetic trees from all completely sequenced plastid genomes currently available and from a selection of mitochondrial genomes representing the major eukaryotic lineages. BLASTN, TBLASTX, or combinations of both are used to locate high-scoring segment pairs (HSPs) between two sequences from which pairwise similarities and distances are computed in different ways resulting in a total of 96 GBDP variants. The suitability of these distance formulae for phylogeny reconstruction is directly estimated by computing a recently described measure of "treelikeness", the so-called delta value, from the respective distance matrices. Additionally, we compare the trees inferred from these matrices using UPGMA, NJ, BIONJ, FastME, or STC, respectively, with the NCBI taxonomy tree of the taxa under study. RESULTS: Our results indicate that, at this taxonomic level, plastid genomes are much more valuable for inferring phylogenies than are mitochondrial genomes, and that distances based on breakpoints are of little use. Distances based on the proportion of "matched" HSP length to average genome length were best for tree estimation. Additionally we found that using TBLASTX instead of BLASTN and, particularly, combining TBLASTX and BLASTN leads to a small but significant increase in accuracy. Other factors do not significantly affect the phylogenetic outcome. The BIONJ algorithm results in phylogenies most in accordance with the current NCBI taxonomy, with NJ and FastME performing insignificantly worse, and STC performing as well if applied to high quality distance matrices. delta values are found to be a reliable predictor of phylogenetic accuracy. CONCLUSION: Using the most treelike distance matrices, as judged by their delta values, distance methods are able to recover all major plant lineages, and are more in accordance with Apicomplexa organelles being derived from "green" plastids than from plastids of the "red" type. GBDP-like methods can be used to reliably infer phylogenies from different kinds of genomic data. A framework is established to further develop and improve such methods. delta values are a topology-independent tool of general use for the development and assessment of distance methods for phylogenetic inference.

Algorithms↗

Virulence Searcher: a tool for searching raw genome sequences from bacterial genomes for putative virulence factors.

There is often a delay between completion of a genome sequence and its publication, mainly because of the lengthy process of annotation. For most researchers, the raw sequence alone does not easily yield the rich information it contains. An online tool (Virulence Searcher) has been designed that enables scientists interested in bacterial pathogenesis to search sequences from unannotated bacterial genomes for putative genes encoding virulence factors. This will facilitate an immediate start on important research into bacterial disease without having to wait for the annotated sequence to be published.

Genes, Bacterial↗

WIT: integrated system for high-throughput genome sequence analysis and metabolic reconstruction.

The WIT (What Is There) (http://wit.mcs.anl.gov/WIT2/) system has been designed to support comparative analysis of sequenced genomes and to generate metabolic reconstructions based on chromosomal sequences and metabolic modules from the EMP/MPW family of databases. This system contains data derived from about 40 completed or nearly completed genomes. Sequence homologies, various ORF-clustering algorithms, relative gene positions on the chromosome and placement of gene products in metabolic pathways (metabolic reconstruction) can be used for the assignment of gene functions and for development of overviews of genomes within WIT. The integration of a large number of phylogenetically diverse genomes in WIT facilitates the understanding of the physiology of different organisms.

Databases, Factual↗

Genome sequencing and comparative genomics of tropical disease pathogens.

The sequencing of eukaryotic genomes has lagged behind sequencing of organisms in the other domains of life, archae and bacteria, primarily due to their greater size and complexity. With recent advances in high-throughput technologies such as robotics and improved computational resources, the number of eukaryotic genome sequencing projects has increased significantly. Among these are a number of sequencing projects of tropical pathogens of medical and veterinary importance, many of which are responsible for causing widespread morbidity and mortality in peoples of developing countries. Uncovering the complete gene complement of these organisms is proving to be of immense value in the development of novel methods of parasite control, such as antiparasitic drugs and vaccines, as well as the development of new diagnostic tools. Combining pathogen genome sequences with the host and vector genome sequences is promising to be a robust method for the identification of host-pathogen interactions. Finally, comparative sequencing of related species, especially of organisms used as model systems in the study of the disease, is beginning to realize its potential in the identification of genes, and the evolutionary forces that shape the genes, that are involved in evasion of the host immune response.

Animals↗

Vibrio cholerae phage K139: complete genome sequence and comparative genomics of related phages.

In this report, we characterize the complete genome sequence of the temperate phage K139, which morphologically belongs to the Myoviridae phage family (P2 and 186). The prophage genome consists of 33,106 bp, and the overall GC content is 48.9%. Forty-four open reading frames were identified. Homology analysis and motif search were used to assign possible functions for the genes, revealing a close relationship to P2-like phages. By Southern blot screening of a Vibrio cholerae strain collection, two highly K139-related phage sequences were detected in non-O1, non-O139 strains. Combinatorial PCR analysis revealed almost identical genome organizations. One region of variable gene content was identified and sequenced. Additionally, the tail fiber genes were analyzed, leading to the identification of putative host-specific sequence variations. Furthermore, a K139-encoded Dam methyltransferase was characterized.

Bacteriophages↗

Reconstruction of ancient genome and gene order from complete microbial genome sequences.

Microbial genome sequences provide us with the fossil records for inferring their origination and evolution. Assuming that current microbial genomes are the evolutionary results of ancient genomes or fragments and the neighboring genes in ancient genomes are more likely neighbors in current genomes, in this paper we proposed a paleontological algorithm and assembled the orthologous gene groups from 66 complete and current microbial genome sequences into a pseudo-ancient genome, which consists of continuous fragments of various sizes. We performed bootstrap resampling and correlation analyses and the results showed that the assembled ancient genome and fragments are statistically significant and the genes of the same fragment are inherently related and likely derived from common ancestors. This method provides a new computational tool for studying microbial genome structure and evolution.

Algorithms↗

The complete genome sequence and comparative genome analysis of the high pathogenicity Yersinia enterocolitica strain 8081.

The human enteropathogen, Yersinia enterocolitica, is a significant link in the range of Yersinia pathologies extending from mild gastroenteritis to bubonic plague. Comparison at the genomic level is a key step in our understanding of the genetic basis for this pathogenicity spectrum. Here we report the genome of Y. enterocolitica strain 8081 (serotype 0:8; biotype 1B) and extensive microarray data relating to the genetic diversity of the Y. enterocolitica species. Our analysis reveals that the genome of Y. enterocolitica strain 8081 is a patchwork of horizontally acquired genetic loci, including a plasticity zone of 199 kb containing an extraordinarily high density of virulence genes. Microarray analysis has provided insights into species-specific Y. enterocolitica gene functions and the intraspecies differences between the high, low, and nonpathogenic Y. enterocolitica biotypes. Through comparative genome sequence analysis we provide new information on the evolution of the Yersinia. We identify numerous loci that represent ancestral clusters of genes potentially important in enteric survival and pathogenesis, which have been lost or are in the process of being lost, in the other sequenced Yersinia lineages. Our analysis also highlights large metabolic operons in Y. enterocolitica that are absent in the related enteropathogen, Yersinia pseudotuberculosis, indicating major differences in niche and nutrients used within the mammalian gut. These include clusters directing, the production of hydrogenases, tetrathionate respiration, cobalamin synthesis, and propanediol utilisation. Along with ancestral gene clusters, the genome of Y. enterocolitica has revealed species-specific and enteropathogen-specific loci. This has provided important insights into the pathology of this bacterium and, more broadly, into the evolution of the genus. Moreover, wider investigations looking at the patterns of gene loss and gain in the Yersinia have highlighted common themes in the genome evolution of other human enteropathogens.

Evolution, Molecular↗

The Genome Sequence DataBase version 1.0 (GSDB): from low pass sequences to complete genomes.

The Genome Sequence DataBase (GSDB) has completed its conversion to an improved relational database. The new database, GSDB 1.0, is fully operational and publicly available. Data contributions, including both original sequence submissions and community annotation, are being accomplished through the use of a graphical client-server interface tool, the GSDB Annotator, and via GIO (GSDB Input/Output) files. Data retrieval services are being provided through a new Web Query Tool and direct SQL. All methods of data contribution and data retrieval fully support the new data types that have been incorporated into GSDB, including discontiguous sequences, multiple sequence alignments, and community annotation.

Animals↗

Achieving congruency of phylogenetic trees generated by W-curves of genomic sequences.

Comparative genomic analysis at its most fundamental level involves alignment and analysis of linear strings of DNA. Many useful and powerful tools, such as BlastN and ClustalW are able to respectively, search for, and align similar strings of DNA from a variety of species. However, interesting genomic patterns cannot be immediately visualized within the information contact embedded in long genomic strings without extensive a priori knowledge. More problematic is the question of whether we will be able to crystallize long genomic sequences and analyze their true secondary and tertiary structures. It is, of course, these putative motifs that are binding to the three-dimensional structures of proteins and inducing replication and transcription events. The W-curve is a numerical mapping algorithm that allows one to geometrically visualize the information content of genomic motifs. Patterns of ALU, LINES, SINEs, and duplication sequences may be easily visualized with the W-curve. It is our hope that this pattern recognition algorithm will lead to visualization tools to track the evolutionary history of motif patterns. The combinatorics of DNA motif crossover-recombination events will be more easily followed as we continue to sequence more and more genomes. In our laboratory we are currently collaborating with mathematicians and computer scientists to develop and test tools, such as the W-curve, for analyzing patterns of long genomic sequences. In this paper, we examine the limitations of using the W-curve to infer the phylogenetic history of species.

Bacteria↗

Genome sequencing networks.

Genome sequencing projects have been undertaken in one of three ways: in a purpose-built and professionally staffed genome centre, by a small number of traditional research laboratories or by an extensive network of traditional research laboratories that are linked by the Internet. Sequencing networks are an attractive option in many circumstances as they are easy to create, bring together diverse types of expertise, integrate the eventual users of a genome sequence with its determination and generally foster a collaborative spirit.

Genome↗

A model of the statistical power of comparative genome sequence analysis.

Comparative genome sequence analysis is powerful, but sequencing genomes is expensive. It is desirable to be able to predict how many genomes are needed for comparative genomics, and at what evolutionary distances. Here I describe a simple mathematical model for the common problem of identifying conserved sequences. The model leads to some useful rules of thumb. For a given evolutionary distance, the number of comparative genomes needed for a constant level of statistical stringency in identifying conserved regions scales inversely with the size of the conserved feature to be detected. At short evolutionary distances, the number of comparative genomes required also scales inversely with distance. These scaling behaviors provide some intuition for future comparative genome sequencing needs, such as the proposed use of "phylogenetic shadowing" methods using closely related comparative genomes, and the feasibility of high-resolution detection of small conserved features.

Animals↗

The Z curve database: a graphic representation of genome sequences.

MOTIVATION: Genome projects for many prokaryotic and eukaryotic species have been completed and more new genome projects are being underway currently. The availability of a large number of genomic sequences for researchers creates a need to find graphic tools to study genomes in a perceivable form. The Z curve is one of such tools available for visualizing genomes. The Z curve is a unique three-dimensional curve representation for a given DNA sequence in the sense that each can be uniquely reconstructed given the other. The Z curve database for more than 1000 genomes have been established here. RESULTS: The database contains the Z curves for archaea, bacteria, eukaryota, organelles, phages, plasmids, viroids and viruses, whose genomic sequences are currently available. All the 3-dimensional Z curves and their three component curves are stored in the database. The applications of the Z curve database on comparative genomics, gene prediction, computation of G+C content with a windowless technique, prediction of replication origins and terminations of bacterial and archaeal genomes and study of local deviations from the Chargaff Parity Rule 2 etc. are presented in detail. The Z curve database reported here is a treasure trove in which biologists could find useful biological knowledge.

Animals↗