Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

The complete genome sequence of J virus reveals a unique genome structure in the family Paramyxoviridae.

J virus (J-V) was isolated from feral mice (Mus musculus) trapped in Queensland, Australia, during the early 1970s. Although studies undertaken at the time revealed that J-V was a new paramyxovirus, it remained unclassified beyond the family level. The complete genome sequence of J-V has now been determined, revealing a genome structure unique within the family Paramyxoviridae. At 18,954 nucleotides (nt), the J-V genome is the largest paramyxovirus genome sequenced to date, containing eight genes in the order 3'-N-P/V/C-M-F-SH-TM-G-L-5'. The two genes located between the fusion (F) and attachment (G) protein genes, which have been named the small hydrophobic (SH) protein gene and the transmembrane (TM) protein gene, encode putative proteins of 69 and 258 amino acids, respectively. The 4,401-nt J-V G gene, much larger than other paramyxovirus attachment protein genes sequenced to date, encodes a putative attachment protein of 709 amino acids and distally contains a second open reading frame (ORF) of 2,115 nt, referred to as ORF-X. Taken together, these novel features represent the most significant divergence to date from the common six-gene genome structure of Paramyxovirinae. Although genome analysis has confirmed that J-V can be classified as a member of the subfamily Paramyxovirinae, it cannot be assigned to any of the five existing genera within this subfamily. Interestingly, a recently isolated paramyxovirus appears to be closely related to J-V, and preliminary phylogenetic analyses based on putative matrix protein sequences indicate that these two viruses will likely represent a new genus within the subfamily Paramyxovirinae.

Amino Acid Sequence↗

Alfresco--a workbench for comparative genomic sequence analysis.

Comparative analysis of genomic sequences provides a powerful tool for identifying regions of potential biologic function; by comparing corresponding regions of genomes from suitable species, protein coding or regulatory regions can be identified by their homology. This requires the use of several specific types of computational analysis tools. Many programs exist for these types of analysis; not many exist for overall view/control of the results, which is necessary for large-scale genomic sequence analysis. Using Java, we have developed a new visualization tool that allows effective comparative genome sequence analysis. The program handles a pair of sequences from putatively homologous regions in different species. Results from various different existing external analysis programs, such as database searching, gene prediction, repeat masking, and alignment programs, are visualized and used to find corresponding functional sequence domains in the two sequences. The user interacts with the program through a graphic display of the genome regions, in which an independently scrollable and zoomable symbolic representation of the sequences is shown. As an example, the analysis of two unannotated orthologous genomic sequences from human and mouse containing parts of the UTY locus is presented.

Algorithms↗

Differential distribution of simple sequence repeats in eukaryotic genome sequences.

Complete chromosome/genome sequences available from humans, Drosophila melanogaster, Caenorhabditis elegans, Arabidopsis thaliana, and Saccharomyces cerevisiae were analyzed for the occurrence of mono-, di-, tri-, and tetranucleotide repeats. In all of the genomes studied, dinucleotide repeat stretches tended to be longer than other repeats. Additionally, tetranucleotide repeats in humans and trinucleotide repeats in Drosophila also seemed to be longer. Although the trends for different repeats are similar between different chromosomes within a genome, the density of repeats may vary between different chromosomes of the same species. The abundance or rarity of various di- and trinucleotide repeats in different genomes cannot be explained by nucleotide composition of a sequence or potential of repeated motifs to form alternative DNA structures. This suggests that in addition to nucleotide composition of repeat motifs, characteristic DNA replication/repair/recombination machinery might play an important role in the genesis of repeats. Moreover, analysis of complete genome coding DNA sequences of Drosophila, C. elegans, and yeast indicated that expansions of codon repeats corresponding to small hydrophilic amino acids are tolerated more, while strong selection pressures probably eliminate codon repeats encoding hydrophobic and basic amino acids. The locations and sequences of all of the repeat loci detected in genome sequences and coding DNA sequences are available at http://www.ncl-india.org/ssr and could be useful for further studies.

Animals↗

Putative virulence-related genes in Vibrio anguillarum identified by random genome sequencing.

The genome of Vibrio anguillarum strain H775-3 was partially determined by a random sequencing procedure. A total of 2,300 clones, 2,100 from a plasmid library and 200 from a cosmid library, were sequenced and subjected to homology search by the BLAST algorithm. The total length of the sequenced clones is 1.5 Mbp. The nucleotide sequences were classified into 17 broad functional categories. Forty putative virulence-related genes were identified, 36 of which are novel in V. anguillarum, including a repeat in toxin gene cluster, haemolysin genes, enterobactin gene, protease genes, lipopolysaccharide biosynthesis genes, capsule biosynthesis gene, flagellar genes and pilus genes.

Bacterial Capsules↗

Complete genome sequence of USA300, an epidemic clone of community-acquired meticillin-resistant Staphylococcus aureus.

BACKGROUND: USA300, a clone of meticillin-resistant Staphylococcus aureus, is a major source of community-acquired infections in the USA, Canada, and Europe. Our aim was to sequence its genome and compare it with those of other strains of S aureus to try to identify genes responsible for its distinctive epidemiological and virulence properties. METHODS: We ascertained the genome sequence of FPR3757, a multidrug resistant USA300 strain, by random shotgun sequencing, then compared it with the sequences of ten other staphylococcal strains. FINDINGS: Compared with closely related S aureus, we noted that almost all of the unique genes in USA300 clustered in novel allotypes of mobile genetic elements. Some of the unique genes are involved in pathogenesis, including Panton-Valentine leucocidin and molecular variants of enterotoxin Q and K. The most striking feature of the USA300 genome is the horizontal acquisition of a novel mobile genetic element that encodes an arginine deiminase pathway and an oligopeptide permease system that could contribute to growth and survival of USA300. We did not detect this element, termed arginine catabolic mobile element (ACME), in other S aureus strains. We noted a high prevalence of ACME in S epidermidis, suggesting not only that ACME transfers into USA300 from S epidermidis, but also that this element confers a selective advantage to this ubiquitous commensal of the human skin. INTERPRETATION: USA300 has acquired mobile genetic elements that encode resistance and virulence determinants that could enhance fitness and pathogenicity.

Community-Acquired Infections↗

Fast and sensitive alignment of large genomic sequences.

Comparative analysis of syntenic genome sequences can be used to identify functional sites such as exons and regulatory elements. Here, the first step is to align two or several evolutionary related sequences and, in recent years, a number of computer programs have been developed for alignment of large genomic sequences. Some of these programs are extremely fast but often time-efficiency is achieved at the expense of sensitivity. One way of combining speed and sensitivity is to use an anchored-alignment approach. In a first step, a fast heuristic identifies a chain of strong sequence similarities that serve as anchor points. In a second step, regions between these anchor points are aligned using a slower but more sensitive method. We present CHAOS, a novel algorithm for rapid identification of chains of local sequence similarities among large genomic sequences. Similarities identified by CHAOS are used as anchor points to improve the running time of the DIALIGN alignment program. Systematic test runs show that this method can reduce the running time of DIALIGN by more than 93% while affecting the quality of the resulting alignments by only 1%. The source code for CHAOS is available at http://www.stanford.edu/~brudno/chaos/ An integrated program package containing CHAOS and DIALIGN is available at http://bibiserv.techfak.uni-bielefeld.de/dialign/

Algorithms↗

Molecular cloning, sequencing and restriction mapping of the genomic sequence encoding human proacrosin.

In the present study, molecular cloning, sequencing and restriction mapping of the genomic sequence encoding human proacrosin is described. The full-length cDNA encoding human proacrosin was utilized to recover a 17-kb human genomic clone which was sequenced without further subcloning. The nucleotide sequences of the exons agree with the sequence of the cDNA reported previously. More than 500 bases of the promoter region were sequenced and found to be highly GC rich but devoid of an identifiable TATA box. These findings are generally consistent with a recently published report [Keime, S., Adham, I. M. & Engel, W. (1990) Eur. J. Biochem. 190, 195-200]. However, further sequence analysis revealed discrepancies between our clone and that previously reported. Sequencing of the first intron showed similarity with the published data for 54 bases of the 5' region, beginning with the donor splice site, and for 114 bases at the 3' end. However, 500 bases sequenced distal to the initial 54 bases at the 5' end of intron 1 showed no similarity with the published sequence. In addition, the boundaries of intron 3 differed such that a cytosine residue previously reported to be in exon 3 was found to be the first base of exon 4. Detailed studies were undertaken to confirm that our clone constitutes the authentic sequence of human proacrosin. Cloning and characterization of the human proacrosin gene may allow for informative studies of its regulation, and for a more detailed examination of its role in fertilization.

Acrosin↗

Hypermethylation of calcitonin gene regulatory sequences in human breast cancer as revealed by genomic sequencing.

DNA methylation has been studied intensively during the past years in order to elucidate its role in the regulation of gene expression, gene imprinting and cancer progression. Earlier studies have shown that a general genomic under-methylation is associated with chronic lymphocytic leukemia and metastatic prostate cancer. Site-specific methylation changes, as revealed by the use of methylation-sensitive restriction enzymes, have been reported to occur in the promotor region of the calcitonin gene in chronic myeloid leukemia as it progresses from the chronic phase to blast crisis, in non-Hodgkin's lymphoid neoplasms and in non-lymphocytic leukemia. We have now explored possible methylation changes associated with benign and malignant breast tumors. Two approaches were employed: (i) chemical determination of general genomic methylation status and (ii) base-specific analysis of the methylation changes in the promoter of the calcitonin gene with the aid of genomic sequencing. The results did not reveal any changes of total DNA 5-methylcytosine content in ductal carcinoma of breast in comparison with benign tumors. There was a small, yet significant, increase in 5-methylcytosine content in lobular carcinoma. Genomic sequencing of the promoter region of the calcitonin gene, however, revealed a striking hypermethylation at or around the transcription start site of the gene in ductal carcinomas. In benign tumors and lobular carcinomas, this region was either entirely unmethylated or only slightly methylated. The latter changes may reflect a regional hypermethylation of the short arm of chromosome 11, which harbors, in addition to the calcitonin gene, a number of putative or established tumor-suppressor genes. Our results demonstrate that genomic sequencing in its present form can be used for a reliable and precise DNA methylation analysis of primary human tumors.

5-Methylcytosine↗

IdentiCS--identification of coding sequence and in silico reconstruction of the metabolic network directly from unannotated low-coverage bacterial genome sequence.

BACKGROUND: A necessary step for a genome level analysis of the cellular metabolism is the in silico reconstruction of the metabolic network from genome sequences. The available methods are mainly based on the annotation of genome sequences including two successive steps, the prediction of coding sequences (CDS) and their function assignment. The annotation process takes time. The available methods often encounter difficulties when dealing with unfinished error-containing genomic sequence. RESULTS: In this work a fast method is proposed to use unannotated genome sequence for predicting CDSs and for an in silico reconstruction of metabolic networks. Instead of using predicted genes or CDSs to query public databases, entries from public DNA or protein databases are used as queries to search a local database of the unannotated genome sequence to predict CDSs. Functions are assigned to the predicted CDSs simultaneously. The well-annotated genome of Salmonella typhimurium LT2 is used as an example to demonstrate the applicability of the method. 97.7% of the CDSs in the original annotation are correctly identified. The use of SWISS-PROT-TrEMBL databases resulted in an identification of 98.9% of CDSs that have EC-numbers in the published annotation. Furthermore, two versions of sequences of the bacterium Klebsiella pneumoniae with different genome coverage (3.9 and 7.9 fold, respectively) are examined. The results suggest that a 3.9-fold coverage of the bacterial genome could be sufficiently used for the in silico reconstruction of the metabolic network. Compared to other gene finding methods such as CRITICA our method is more suitable for exploiting sequences of low genome coverage. Based on the new method, a program called IdentiCS (Identification of Coding Sequences from Unfinished Genome Sequences) is delivered that combines the identification of CDSs with the reconstruction, comparison and visualization of metabolic networks (free to download at http://genome.gbf.de/bioinformatics/index.html). CONCLUSIONS: The reversed querying process and the program IdentiCS allow a fast and adequate prediction protein coding sequences and reconstruction of the potential metabolic network from low coverage genome sequences of bacteria. The new method can accelerate the use of genomic data for studying cellular metabolism.

Base Sequence↗

Gene-associated CpG islands in plants as revealed by analyses of genomic sequences.

We screened plant genome sequences, primarily from rice and Arabidopsis thaliana, for CpG islands, and identified DNA segments rich in CpG dinucleotides within these sequences. These CpG-rich clusters appeared in the analysed sequences as discrete peaks and occurred at the frequencies of one per 4.7 kb in rice and one per 4.0 kb in A. thaliana. In rice and A. thaliana, most of the CpG-rich clusters were associated with genes, which suggests that these clusters are useful landmarks in genome sequences for identifying genes in plants with small genomes. In contrast, in plants with larger genomes, only a few of the clusters were associated with genes. These plant CpG-rich clusters satisfied the criteria used for identifying human CpG islands, which suggests that these CpG clusters may be regarded as plant CpG islands. The position of each island relative to the 5'-end of its associated gene varied considerably. Genes in the analysed sequences were grouped into five classes according to the position of the CpG islands within their associated genes. A large proportion of the genes belonged to one of two classes, in which a CpG island occurred near the 5'-end of the gene or covered the whole gene region. The position of a plant CpG island within its associated gene appeared to be related to the extent of tissue-specific expression of the gene; the CpG islands of most of the widely expressed rice genes occurred near the 5'-end of the genes.

Arabidopsis↗

Visualizing associations between genome sequences and gene expression data using genome-mean expression profiles.

The combination of genome-wide expression patterns and full genome sequences offers a great opportunity to further our understanding of the mechanisms and logic of transcriptional regulation. Many methods have been described that identify sequence motifs enriched in transcription control regions of genes that share similar gene expression patterns. Here we present an alternative approach that evaluates the transcriptional information contained by specific sequence motifs by computing for each motif the mean expression profile of all genes that contain the motif in their transcription control regions. These genome-mean expression profiles (GMEP's) are valuable for visualizing the relationship between genome sequences and gene expression data, and for characterizing the transcriptional importance of specific sequence motifs. Analysis of GMEP's calculated from a dataset of 519 whole-genome microarray experiments in Saccharomyces cerevisiae show a significant correlation between GMEP's of motifs that are reverse complements, a result that supports the relationship between GMEP's and transcriptional regulation. Hierarchical clustering of GMEP's identifies clusters of motifs that correspond to binding sites of well-characterized transcription factors. The GMEP's of these clustered motifs have patterns of variation across conditions that reflect the known activities of these transcription factors. Software that computed GMEP's from sequence and gene expression data is available under the terms of the Gnu Public License from http://rana.lbl.gov/.

Algorithms↗

Sequence complexity profiles of prokaryotic genomic sequences: a fast algorithm for calculating linguistic complexity.

MOTIVATION: One of the major features of genomic DNA sequences, distinguishing them from texts in most spoken or artificial languages, is their high repetitiveness. Variation in the repetitiveness of genomic texts reflects the presence and density of different biologically important messages. Thus, deviation from an expected number of repeats in both directions indicates a possible presence of a biological signal. Linguistic complexity corresponds to repetitiveness of a genomic text, and potential regulatory sites may be discovered through construction of typical patterns of complexity distribution. RESULTS: We developed software for fast calculation of linguistic sequence complexity of DNA sequences. Our program utilizes suffix trees to compute the number of subwords present in genomic sequences, thereby allowing calculation of linguistic complexity in time linear in genome size. The measure of linguistic complexity was applied to the complete genome of Haemophilus influenzae. Maps of complexity along the entire genome were obtained using sliding windows of 40, 100, and 2000 nucleotides. This approach provided an efficient way to detect simple sequence repeats in this genome. In addition, local profiles of complexity distribution around the starts of translation were constructed for 21 complete prokaryotic genomes. We hypothesize that complexity profiles correspond to evolutionary relationships between organisms. We found principal differences in profiles of the GC-rich and other (non-GC-rich) genomes. We also found characteristic differences in profiles of AT genomes, which probably reflect individual species variations in translational regulation. AVAILABILITY: The program is available upon request from Alexander Bolshoy or at http://csweb.haifa.ac.il/library/#complex.

Algorithms↗

The genome sequence DataBase.

The Genome Sequence DataBase (GSDB) is a database of publicly available nucleotide sequences and their associated biological and bibliographic information. Several notable changes have occurred in the past year: GSDB stopped accepting data submissions from researchers; ownership of data submitted to GSDB was transferred to GenBank; sequence analysis capabilities were expanded to include Smith-Waterman and Frame Search; and Sequence Viewer became available to Mac users. The content of GSDB remains up-to-date because publicly available data is acquired from the International Nucleotide Sequence Database Collaboration databases (IC) on a nightly basis. This allows GSDB to continue providing researchers with the ability to analyze, query and retrieve nucleotide sequences in the database. GSDB and its related tools are freely accessible from the URL: http://www.ncgr.org

Databases, Factual↗

Reconstruction of amino acid biosynthesis pathways from the complete genome sequence.

The complete genome sequence of an organism contains information that has not been fully utilized in the current prediction methods of gene functions, which are based on piece-by-piece similarity searches of individual genes. We present here a method that utilizes a higher level information of molecular pathways to reconstruct a complete functional unit from a set of genes. Specifically, a genome-by-genome comparison is first made for identifying enzyme genes and assigning EC numbers, which is followed by the reconstruction of selected portions of the metabolic pathways by use of the reference biochemical knowledge. The completeness of the reconstructed pathway is an indicator of the correctness of the initial gene function assignment. This feature has become possible because of our efforts to computerize the current knowledge of metabolic pathways under the KEGG project. We found that the biosynthesis pathways of all 20 amino acids were completely reconstructed in Escherichia coli, Haemophilus influenzae, and Bacillus subtilis, and probably in Synechocystis and Saccharomyces cerevisiae as well, although it was necessary to assume wider substrate specificity for aspartate aminotransferases.

Amino Acids↗

Comparing low coverage random shotgun sequence data from Brassica oleracea and Oryza sativa genome sequence for their ability to add to the annotation of Arabidopsis thaliana.

Since the completion of the Arabidopsis thaliana genome sequence, there is an ongoing effort to annotate the genome as accurately as possible. Comparing genome sequences of related species complements the current annotation strategies by identifying genes and improving gene structure. A total of 595,321 Brassica oleracea shotgun reads were sequenced by TIGR (The Institute for Genome Research) and the collaboration of Washington University and Cold Spring Harbor. Vicogenta (a genome viewer based on GMOD and GBrowse) was created to view the current annotation and sequence alignments for Arabidopsis. Brassica reads were compared with the Arabidopsis genome and proteome databases using BLAST. Hypothetical genes and conserved unannotated regions on the short arm of chromosome 4 from Arabidopsis were experimentally verified using RT-PCR. We were able to improve the Arabidopsis annotation by identifying 25 genes that were missed, and confirming expression of 43 hypothetical genes in Arabidopsis. We were also able to detect conservation in genes whose transcription is normally suppressed due to methylation. We also examined how useful the O. sativa genome and ESTs from other species are, compared with Brassica, in improving the Arabidopsis annotation.

Amino Acid Sequence↗

Genome sequences and evolutionary biology, a two-way interaction.

Complete genome sequences are accumulating rapidly, culminating with the announcement of the human genome sequence in February 2001. In addition to cataloguing the diversity of genes and other sequences, genome sequences will provide the first detailed and complete data on gene families and genome organization, including data on evolutionary changes. Reciprocally, evolutionary biology will make important contributions to the efforts to understand functions of genes and other sequences in genomes. Large-scale, detailed and unbiased comparisons between species will illuminate the evolution of genes and genomes, and population genetics methods will enable detection of functionally important genes or sequences, including sequences that have been involved in adaptive changes.

Journal Article↗

Genomic sequencing reveals gene content, genomic organization, and recombination relationships in barley.

Barley (Hordeum vulgare L.) is one of the most important large-genome cereals with extensive genetic resources available in the public sector. Studies of genome organization in barley have been limited primarily to genetic markers and sparse sequence data. Here we report sequence analysis of 417.5 kb DNA from four BAC clones from different genomic locations. Sequences were analyzed with respect to gene content, the arrangement of repetitive sequences and the relationship of gene density to recombination frequencies. Gene densities ranged from 1 gene per 12 kb to 1 gene per 103 kb with an average of 1 gene per 21 kb. In general, genes were organized into islands separated by large blocks of nested retrotransposons. Single genes in apparent isolation were also found. Genes occupied 11% of the total sequence, LTR retrotransposons and other repeated elements accounted for 51.9% and the remaining 37.1% could not be annotated.

Chromosomes, Artificial, Bacterial↗