Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Molecular Sequence Annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

GPCR-GRAPA-LIB--a refined library of hidden Markov Models for annotating GPCRs.

GPCR-GRAPA-LIB is a library of HMMs describing G protein coupled receptor families. These families are initially defined by class of receptor ligand, with divergent families divided into subfamilies using phylogenic analysis and knowledge of GPCR function. Protein sequences are applied to the models with the GRAPA curve-based selection criteria. RefSeq sequences for Homo sapiens, Drosophila melanogaster, and Caenorhabditis elegans have been annotated using this approach.

Algorithms↗

Automatic annotation of eukaryotic genes, pseudogenes and promoters.

BACKGROUND: The ENCODE gene prediction workshop (EGASP) has been organized to evaluate how well state-of-the-art automatic gene finding methods are able to reproduce the manual and experimental gene annotation of the human genome. We have used Softberry gene finding software to predict genes, pseudogenes and promoters in 44 selected ENCODE sequences representing approximately 1% (30 Mb) of the human genome. Predictions of gene finding programs were evaluated in terms of their ability to reproduce the ENCODE-HAVANA annotation. RESULTS: The Fgenesh++ gene prediction pipeline can identify 91% of coding nucleotides with a specificity of 90%. Our automatic pseudogene finder (PSF program) found 90% of the manually annotated pseudogenes and some new ones. The Fprom promoter prediction program identifies 80% of TATA promoters sequences with one false positive prediction per 2,000 base-pairs (bp) and 50% of TATA-less promoters with one false positive prediction per 650 bp. It can be used to identify transcription start sites upstream of annotated coding parts of genes found by gene prediction software. CONCLUSION: We review our software and underlying methods for identifying these three important structural and functional genome components and discuss the accuracy of predictions, recent advances and open problems in annotating genomic sequences. We have demonstrated that our methods can be effectively used for initial automatic annotation of the eukaryotic genome.

Animals↗

Gene structure prediction by spliced alignment of genomic DNA with protein sequences: increased accuracy by differential splice site scoring.

Gene identification in genomic DNA from eukaryotes is complicated by the vast combinatorial possibilities of potential exon assemblies. If the gene encodes a protein that is closely related to known proteins, gene identification is aided by matching similarity of potential translation products to those target proteins. The genomic DNA and protein sequences can be aligned directly by scoring the implied residues of in-frame nucleotide triplets against the protein residues in conventional ways, while allowing for long gaps in the alignment corresponding to introns in the genomic DNA. We describe a novel method for such spliced alignment. The method derives an optimal alignment based on scoring for both sequence similarity of the predicted gene product to the protein sequence and intrinsic splice site strength of the predicted introns. Application of the method to a representative set of 50 known genes from Arabidopsis thaliana showed significant improvement in prediction accuracy compared to previous spliced alignment methods. The method is also more accurate than ab initio gene prediction methods, provided sufficiently close target proteins are available. In view of the fast growth of public sequence repositories, we argue that close targets will be available for the majority of novel genes, making spliced alignment an excellent practical tool for high-throughput automated genome annotation.

Algorithms↗

Prediction of signal recognition particle RNA genes.

We describe a method for prediction of genes that encode the RNA component of the signal recognition particle (SRP). A heuristic search for the strongly conserved helix 8 motif of SRP RNA is combined with covariance models that are based on previously known SRP RNA sequences. By screening available genomic sequences we have identified a large number of novel SRP RNA genes and we can account for at least one gene in every genome that has been completely sequenced. Novel bacterial RNAs include that of Thermotoga maritima, which, unlike all other non-gram-positive eubacteria, is predicted to have an Alu domain. We have also found the RNAs of Lactococcus lactis and Staphylococcus to have an unusual UGAC tetraloop in helix 8 instead of the normal GNRA sequence. An investigation of yeast RNAs reveals conserved sequence elements of the Alu domain that aid in the analysis of these RNAs. Analysis of the human genome reveals only two likely genes, both on chromosome 14. Our method for SRP RNA gene prediction is the first convenient tool for this task and should be useful in genome annotation.

Alu Elements↗

Characterization of a normalized cDNA library from bovine intestinal muscle and epithelial tissues.

Tissue-specific cDNA library sequences (expressed sequence tags, or EST) yield a detailed snapshot of gene expression and are useful in developing second-generation molecular resources (i.e., microarrays) for gene expression profiling. The objective of this study was to develop and characterize an intestine-specific cDNA library to examine the transcriptome of the bovine gut and identify expressed genes that influence ruminant nutrition and health. We describe BARC-8BOV, a normalized cDNA library developed from mRNA isolated from four distinct intestinal locations (duodenal, jejunal and ileal small intestine, colon) of Holstein dairy cattle resulting in 19,110 5'-EST deposited into the NCBI GenBank EST database. Assembly and clustering of these 19,110 clone sequences yielded 11,208 unique elements (3,419 contigs and 7,789 singletons) with an average length of 695 base pairs. Analysis strongly suggests normalization and tissue pooling were effective at increasing the discovery rate of new bovine sequence. A total of 1,123 sequence elements not previously identified in cattle, but with similarity to known genes in other animal species, were identified and shown to be involved in numerous critical biological processes. An additional 745 transcripts were not previously represented as EST in nucleotide or protein databases, and further analysis of these could lead to the identification of gut-specific transcript variants of known genes or potentially the discovery of novel bovine genes. Of the 11,208 assembled sequences, 11,034, or 98.4%, match sequences present in the bovine DNA trace archive at NCBI, and add to a bovine EST database previously lacking significant gut tissue representation. Ultimately, these data will also contribute in efforts to annotate the bovine genome.

Animals↗

Expressed sequence tags (ESTs) and simple sequence repeat (SSR) markers from octoploid strawberry (Fragaria x ananassa).

BACKGROUND: Cultivated strawberry (Fragaria x ananassa) represents one of the most valued fruit crops in the United States. Despite its economic importance, the octoploid genome presents a formidable barrier to efficient study of genome structure and molecular mechanisms that underlie agriculturally-relevant traits. Many potentially fruitful research avenues, especially large-scale gene expression surveys and development of molecular genetic markers have been limited by a lack of sequence information in public databases. As a first step to remedy this discrepancy a cDNA library has been developed from salicylate-treated, whole-plant tissues and over 1800 expressed sequence tags (EST's) have been sequenced and analyzed. RESULTS: A putative unigene set of 1304 sequences--133 contigs and 1171 singlets--has been developed, and the transcripts have been functionally annotated. Homology searches indicate that 89.5% of sequences share significant similarity to known/putative proteins or Rosaceae ESTs. The ESTs have been functionally characterized and genes relevant to specific physiological processes of economic importance have been identified. A set of tools useful for SSR development and mapping is presented. CONCLUSION: Sequences derived from this effort may be used to speed gene discovery efforts in Fragaria and the Rosaceae in general and also open avenues of comparative mapping. This report represents a first step in expanding molecular-genetic analyses in strawberry and demonstrates how computational tools can be used to optimally mine a large body of useful information from a relatively small data set.

Chromosome Mapping↗

Automated protein structure homology modeling: a progress report.

Understanding the molecular function of proteins is greatly enhanced by insights gained from their three-dimensional structures. Since experimental structures are only available for a small fraction of proteins, computational methods for protein structure modeling play an increasingly important role. Comparative protein structure modeling is currently the most accurate method, yielding models suitable for a wide spectrum of applications, such as structure-guided drug development or virtual screening. Stable and reliable automated prediction pipelines have been developed to apply large-scale comparative modeling to whole genomes or entire sequence databases. Model repositories give access to these annotated and evaluated models. In this review, we will discuss recent developments in automated comparative modeling and provide selected examples illustrating the use of homology models.

Animals↗

Genome sequencing and annotation of two rhizobacteria with antifungal activity: Pseudomonas tolaasii strain A46 and Pseudomonas palleroriana P61.

Here, we report the draft genome sequences of P. tolaasii A46 and P. palleroriana P61, two rhizobacteria previously shown to inhibit Rhizoctonia solani. These genomic resources will support future efforts to elucidate the molecular basis of fungal suppression and to assess the biocontrol potential of these Pseudomonas strains.

antagonistic rhizobacteria↗

Protein model determination from crystallographic data.

Crystallographic studies play a major role in current efforts towards protein structure determination. However, despite recent advances in computational tools for molecular modeling and graphics, the task of constructing a model of the tertiary structure of a protein from experimental data remains complex and time-consuming, requiring extensive expert intervention. This paper describes an approach to protein model determination that incorporates crystallographic data, along with sequence data. A model is represented as an annotated graph that traces the backbone and side chains for a protein. The proposed approach incorporates numerical techniques that are applied to construct and analyze an electron density map for a unit cell of a crystal. The purpose of this work is to advance the ability to discern meaningful features of protein structure through the use of topological analysis of the relative density. Experimental results, which demonstrate the viability of the approach, are reported.

Computer Graphics↗

LegumeDB1 bioinformatics resource: comparative genomic analysis and novel cross-genera marker identification in lupin and pasture legume species.

The identification of markers in legume pasture crops, which can be associated with traits such as protein and lipid production, disease resistance, and reduced pod shattering, is generally accepted as an important strategy for improving the agronomic performance of these crops. It has been demonstrated that many quantitative trait loci (QTLs) identified in one species can be found in other plant species. Detailed legume comparative genomic analyses can characterize the genome organization between model legume species (e.g., Medicago truncatula, Lotus japonicus) and economically important crops such as soybean (Glycine max), pea (Pisum sativum), chickpea (Cicer arietinum), and lupin (Lupinus angustifolius), thereby identifying candidate gene markers that can be used to track QTLs in lupin and pasture legume breeding. LegumeDB is a Web-based bioinformatics resource for legume researchers. LegumeDB analysis of Medicago truncatula expressed sequence tags (ESTs) has identified novel simple sequence repeat (SSR) markers (16 tested), some of which have been putatively linked to symbiosome membrane proteins in root nodules and cell-wall proteins important in plant-pathogen defence mechanisms. These novel markers by preliminary PCR assays have been detected in Medicago truncatula and detected in at least one other legume species, Lotus japonicus, Glycine max, Cicer arietinum, and (or) Lupinus angustifolius (15/16 tested). Ongoing research has validated some of these markers to map them in a range of legume species that can then be used to compile composite genetic and physical maps. In this paper, we outline the features and capabilities of LegumeDB as an interactive application that provides legume genetic and physical comparative maps, and the efficient feature identification and annotation of the vast tracks of model legume sequences for convenient data integration and visualization. LegumeDB has been used to identify potential novel cross-genera polymorphic legume markers that map to agronomic traits, supporting the accelerated identification of molecular genetic factors underpinning important agronomic attributes in lupin.

Chromosome Mapping↗

A new cancer genome anatomy project web resource for the community.

The National Cancer Institute's Cancer Genome Anatomy Project (CGAP) is developing publicly accessible information, technology, and material resources that provide a platform for the interface of cancer research and genomics. CGAP's efforts have focused toward (1) building and annotating catalogues of genes expressed during cancer development, (2) identifying polymorphisms in those genes, and (3) developing resources for the molecular characterization of cancer-related chromosomal aberrations. To date, CGAP has produced more than 1,000,000 expressed sequence tags, approximately 3,300,000 serial analysis of gene expression tags, and identified more than 10,000 human gene-based single-nucleotide polymorphisms. To enhance access to these datasets by the research community, a new Cancer Genome Project web site (http://cgap.nci.nih.gov/) is being introduced. The web site includes genomic data for humans and mice, including transcript sequence, gene expression patterns, single-nucleotide polymorphisms, clone resources, and cytogenetic information. Descriptions of the methods and reagents used in deriving the CGAP datasets are also provided. An extensive suite of informatics tools facilitates queries and analysis of the CGAP data by the community. One of the newest features of the CGAP web site is an electronic version of the Mitelman Database of Chromosome Aberrations in Cancer.

Chromosome Aberrations↗

From genes to proteins: in vitro expression of rickettsial proteins.

The availability of the complete genome sequences of several organisms allows the comparative analysis of genomes, a branch of bioinformatics known as genomics. With this approach, much can be learned about the biology of organisms that are difficult to culture, even when few, if any, of their proteins have been isolated and studied directly. We have focused our interest on Rickettsia conorii, an obligate intracellular bacterium responsible for Mediterranean spotted fever, a disease endemic in southern Europe. While bioinformatic annotation of the complete genome of this bacteria has allowed identification of 1,374 genes, a large number of them remain functionally uncharacterized. The final goal of many experiments in molecular biology is to use biological systems to synthesize the protein encoded by the gene being studied. Because three-dimensional structures are more resilient to evolution and change than amino acid sequences, structure determination of some open reading frames should also exhibit structural similarity to previously described protein families. We have thus initiated a systematic expression and structure determination program for the proteins encoded by rickettsial genes of interest. We have cloned different genes of R. conorii by recombinational cloning (GATEWAY), Invitrogen) a method that uses in vitro site-specific recombination to accomplish a directional cloning of PCR products and the subsequent automatic subcloning of the DNA segment into new vector backbones at high efficiency. The constructions in p-Dest17 yielded several clones able to express recombinant proteins with a C-terminal histidine tag. Expression of corresponding proteins was then performed using a cell-free protein expression system (Rapid Translation System, RTS, Roche Diagnostics). The recombinational cloning approach coupled to RTS provides an approach to rapid optimization of protein expression and is very useful to express rickettsial proteins. Moreover, this system is able to overcome some of the limitations encountered with rickettsial proteins highly toxic for E. coli or insect cells.

Bacterial Proteins↗

New local potential useful for genome annotation and 3D modeling.

A new potential energy function representing the conformational preferences of sequentially local regions of a protein backbone is presented. This potential is derived from secondary structure probabilities such as those produced by neural network-based prediction methods. The potential is applied to the problem of remote homolog identification, in combination with a distance-dependent inter-residue potential and position-based scoring matrices. This fold recognition jury is implemented in a Java application called JThread. These methods are benchmarked on several test sets, including one released entirely after development and parameterization of JThread. In benchmark tests to identify known folds structurally similar to (but not identical with) the native structure of a sequence, JThread performs significantly better than PSI-BLAST, with 10% more structures identified correctly as the most likely structural match in a fold library, and 20% more structures correctly narrowed down to a set of five possible candidates. JThread also improves the average sequence alignment accuracy significantly, from 53% to 62% of residues aligned correctly. Reliable fold assignments and alignments are identified, making the method useful for genome annotation. JThread is applied to predicted open reading frames (ORFs) from the genomes of Mycoplasma genitalium and Drosophila melanogaster, identifying 20 new structural annotations in the former and 801 in the latter.

Animals↗

Genome-wide protein interaction maps using two-hybrid systems.

Automated sequence technology has rendered functional biology amenable to genomic scale analysis. Among genome-wide exploratory approaches, the two-hybrid system in yeast (Y2H) has outranked other techniques because it is the system of choice to detect protein-protein interactions. Deciphering the cascade of binding events in a whole cell helps define signal transduction and metabolic pathways or enzymatic complexes. The function of proteins is eventually attributed through whole cell protein interaction maps where totally unknown proteins are partnered with fully annotated proteins belonging to the same functional category. Since its first description in the late 1980's, several versions of the Y2H have been developed in order to overcome the major limitations of the system, namely false positives and false negatives. Optimized versions have been recently applied at multi-molecular and genomic scale. These genome-wide surveys can be methodologically divided into two types of approaches: one either tests combinations of predefined polypeptides (the so-called matrix approach) using various short-cuts to speed up the process, or one screens with a given polypeptide (bait) for potential partners (preys) present in complex libraries of genomic or complementary DNA (library screening). In the former strategy, one tests what one knows, for example pair-wise interactions between full-length open reading frames from recently sequenced and annotated genomes. Although based on a one-by-one scheme, this method is reported to be amenable to large-scale genomics thanks to multicloning strategies and to the use of small robotics workstations. In the latter, highly complex cDNA or genomic libraries of protein domains can be screened to saturation with high-throughput screening systems allowing the discovery of yet unidentified proteins. Both approaches have strengths and drawbacks that will be discussed here. None yields a full proteome-wide screening since certain proteins (e.g. some transcription factors) are not usable in Y2H. Novel two-hybrid assays have been recently described in bacteria. Applications of these time- and cost-effective assays to genomic screening will be discussed and compared to the Y2H technology.

Animals↗

Protein sequencing by mass analysis of polypeptide ladders after controlled protein hydrolysis.

The characterization of protein modifications is essential for the study of protein function using functional genomic and proteomic approaches. However, current techniques are not efficient in determining protein modifications. We report an approach for sequencing proteins and determining modifications with high speed, sensitivity and specificity. We discovered that a protein could be readily acid-hydrolyzed within 1 min by exposure to microwave irradiation to form, predominantly, two series of polypeptide ladders containing either the N- or C-terminal amino acid of the protein, respectively. Mass spectrometric analysis of the hydrolysate produced a simple mass spectrum consisting of peaks exclusively from these polypeptide ladders, allowing direct reading of amino acid sequence and modifications of the protein. As examples, we applied this technique to determine protein phosphorylation sites as well as the sequences and several previously unknown modifications of 28 small proteins isolated from Escherichia coli K12 cells. This technique can potentially be automated for large-scale protein annotation.

Algorithms↗

An evaluation of new criteria for CpG islands in the human genome as gene markers.

MOTIVATION: Recently, more stringent criteria for CpG islands have been introduced to exclude Alu repeats, thereby enabling a higher proportion of CpG islands associating with genes to be identified. Using these new criteria, several types of associations between CpG islands and genes were investigated to further establish the importance of CpG islands as gene markers. RESULTS: The CpG islands were searched by CpGIE, a java software program developed for CpG island identification. CpGIE was advanced in identification accuracy compared with other tools. According to our results, about 70% of the identified CpG islands were associating with the human genes and over half of them are in the promoters. Furthermore, the investigation of genes in the confirmed gene model showed that 56% of them had a CpG island overlapping the transcription start sites. In comparison, the new criteria were found capable of filtering a large fraction of Alu repeats that was identified as CpG islands by the generally accepted criteria within the genes, but very few CpG islands associating with the promoters were affected. The genes in the predicted gene model were not obviously associated with CpG islands, suggesting that CpG islands can be used to evaluate the accuracy of gene annotation. AVAILABILITY: http://bioinfo.hku.hk/cpgieintro

Algorithms↗

Genepi: a blackboard framework for genome annotation.

BACKGROUND: Genome annotation can be viewed as an incremental, cooperative, data-driven, knowledge-based process that involves multiple methods to predict gene locations and structures. This process might have to be executed more than once and might be subjected to several revisions as the biological (new data) or methodological (new methods) knowledge evolves. In this context, although a lot of annotation platforms already exist, there is still a strong need for computer systems which take in charge, not only the primary annotation, but also the update and advance of the associated knowledge. In this paper, we propose to adopt a blackboard architecture for designing such a system RESULTS: We have implemented a blackboard framework (called Genepi) for developing automatic annotation systems. The system is not bound to any specific annotation strategy. Instead, the user will specify a blackboard structure in a configuration file and the system will instantiate and run this particular annotation strategy. The characteristics of this framework are presented and discussed. Specific adaptations to the classical blackboard architecture have been required, such as the description of the activation patterns of the knowledge sources by using an extended set of Allen's temporal relations. Although the system is robust enough to be used on real-size applications, it is of primary use to bioinformatics researchers who want to experiment with blackboard architectures. CONCLUSION: In the context of genome annotation, blackboards have several interesting features related to the way methodological and biological knowledge can be updated. They can readily handle the cooperative (several methods are implied) and opportunistic (the flow of execution depends on the state of our knowledge) aspects of the annotation process.

Algorithms↗

Evolutionary relationships among G protein-coupled receptors using a clustered database approach.

Guanine nucleotide-binding protein-coupled receptors (GPCRs) comprise large and diverse gene families in fungi, plants, and the animal kingdom. GPCRs appear to share a common structure with 7 transmembrane segments, but sequence similarity is minimal among the most distant GPCRs. To reevaluate the question of evolutionary relationships among the disparate GPCR families, this study takes advantage of the dramatically increased number of cloned GPCRs. Sequences were selected from the National Center for Biotechnology Information (NCBI) nonredundant peptide database using iterative BLAST (Basic Local Alignment Search Tool) searches to yield a database of approximately 1700 GPCRs and unrelated membrane proteins as controls, divided into 34 distinct clusters. For each cluster, separate position-specific matrices were established to optimize sequence comparisons among GPCRs. This approach resulted in significant alignments between distant GPCR families, including receptors for the biogenic amine/peptide, VIP/secretin, cAMP, STE3/MAP3 fungal pheromones, latrophilin, developmental receptors frizzled and smoothened, as well as the more distant metabotrobic glutamate receptors, the STE2/MAM2 fungal pheromone receptors, and GPR1, a fungal glucose receptor. On the other hand, alignment scores between these recognized GPCR clades with p40 (putative GPCR) and pm1 (putative GPCR), as well as bacteriorhodopsins, failed to support a finding of homology. This study provides a refined view of GPCR ancestry and serves as a reference database with hyperlinks to other sources. Moreover, it may facilitate database annotation and the assignment of orphan receptors to GPCR families.

Amino Acid Sequence↗