Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

iMOTdb--a comprehensive collection of spatially interacting motifs in proteins.

Realization of conserved residues that represent a protein family is crucial for clearer understanding of biological function as well as for the better recognition of additional members in sequence databases. Functionally important residues are recognized well due to their high degree of conservation in closely related sequences and are annotated in functional motif databases. Structural motifs are central to the integrity of the fold and require careful analysis for their identification. We report the availability of a database of spatially interacting motifs in single protein structures as well as those among distantly related protein structures that belong to a superfamily. Spatial interactions amongst conserved motifs are automatically measured using sequence similarity scores and distance calculations. Interactions between pairs of conserved motifs are described in the form of pseudoenergies. iMOTdb database provides information for 854,488 motifs corresponding to 60,849 protein structural domains and 22,648 protein structural entries.

Amino Acid Motifs↗

Database verification studies of SWISS-PROT and GenBank.

PROBLEM STATEMENT: We have studied the relationships among SWISS-PROT, TrEMBL, and GenBank with two goals. First is to determine whether users can reliably identify those proteins in SWISS-PROT whose functions were determined experimentally, as opposed to proteins whose functions were predicted computationally. If this information was present in reasonable quantities, it would allow researchers to decrease the propagation of incorrect function predictions during sequence annotation, and to assemble training sets for developing the next generation of sequence-analysis algorithms. Second is to assess the consistency between translated GenBank sequences and sequences in SWISS-PROT and TrEMBL. RESULTS: (1) Contrary to claims by the SWISS-PROT authors, we conclude that SWISS-PROT does not identify a significant number of experimentally characterized proteins. (2) SWISS-PROT is more incomplete than we expected in that version 38.0 from July 1999 lacks many proteins from the full genomes of important organisms that were sequenced years earlier. (3) Even if we combine SWISS-PROT and TrEMBL, some sequences from the full genomes are missing from the combined dataset. (4) In many cases, translated GenBank genes do not exactly match the corresponding SWISS-PROT sequences, for reasons that include missing or removed methionines, differing translation start positions, individual amino-acid differences, and inclusion of sequence data from multiple sequencing projects. For example, results show that for Escherichia coli, 80.6% of the proteins in the GenBank entry for the complete genome have identical sequence matches with SWISS-PROT/TrEMBL sequences, 13.4% have exact substring matches, and matches for 4.1% can be found using BLAST search; the remaining 2.0% of E.coli protein sequences (most of which are ORFs) have no clear matches to SWISS-PROT/TrEMBL. Although many of these differences can be explained by the complexity of the DB, and by the curation processes used to create it, the scale of the differences is notable.

Algorithms↗

Whole genome shotgun sequencing of Brassica oleracea and its application to gene discovery and annotation in Arabidopsis.

Through comparative studies of the model organism Arabidopsis thaliana and its close relative Brassica oleracea, we have identified conserved regions that represent potentially functional sequences overlooked by previous Arabidopsis genome annotation methods. A total of 454,274 whole genome shotgun sequences covering 283 Mb (0.44 x) of the estimated 650 Mb Brassica genome were searched against the Arabidopsis genome, and conserved Arabidopsis genome sequences (CAGSs) were identified. Of these 229,735 conserved regions, 167,357 fell within or intersected existing gene models, while 60,378 were located in previously unannotated regions. After removal of sequences matching known proteins, CAGSs that were close to one another were chained together as potentially comprising portions of the same functional unit. This resulted in 27,347 chains of which 15,686 were sufficiently distant from existing gene annotations to be considered a novel conserved unit. Of 192 conserved regions examined, 58 were found to be expressed in our cDNA populations. Rapid amplification of cDNA ends (RACE) was used to obtain potentially full-length transcripts from these 58 regions. The resulting sequences led to the creation of 21 gene models at 17 new Arabidopsis loci and the addition of splice variants or updates to another 19 gene structures. In addition, CAGSs overlapping already annotated genes in Arabidopsis can provide guidance for manual improvement of existing gene models. Published genome-wide expression data based on whole genome tiling arrays and massively parallel signature sequencing were overlaid on the Brassica-Arabidopsis conserved sequences, and 1399 regions of intersection were identified. Collectively our results and these data sets suggest that several thousand new Arabidopsis genes remain to be identified and annotated.

Arabidopsis↗

A novel genetic locus outside the symbiotic island is required for effective symbiosis of Bradyrhizobium japonicum with soybean Glycine max.

In order to investigate the symbiotic interaction between soybean and Bradyrhizobium japonicum, TnphoA mutagenesis of the microsymbiont was performed. Mutant strain 2-10 was found to induce a strongly reduced number of ineffective nodules. Ultrastructural analysis of the soybean nodule central tissue revealed the presence of numerous starch granules and vacuoles in the infected cells. In addition, the number of symbiosomes was extremely low, indicating an impaired interaction between the plant and invading bacteria. Cloning and sequencing of the mutated DNA region uncovered four open reading frames (ORFs) lacking any data base similarities. ORFs srrA1 and srrA2, the 2-10 TnphoA insertion site, are encoded in the same reading frame. A 35-kDa expression product in Escherichia coli indicated the presence of a common protein, called SrrA (symbiotically relevant region) in B. japonicum 110spc4, encoded by combined srrA1 and srrA2 genes. The analysis of gene disruption mutants revealed that srrB and srrC were also required for effective symbiosis with soybeans. Further downstream the gene for a putative inner membrane protein (pipA) of unknown function was encoded on the opposite strand. Primer extension studies led to the conclusion that the organization of genes differed from the RhizoBase annotation in this particular region of B. japonicum USDA110.

Amino Acid Sequence↗

Drosophila genomic sequence annotation using the BLOCKS+ database.

A simple and general homology-based method for gene finding was applied to the 2.9-Mb Drosophila melanogaster Adh region, the target sequence of the Genome Annotation Assessment Project (GASP). Each strand of the entire sequence was used as query of the BLOCKS+ database of conserved regions of proteins. This led to functional assignments for more than one-third of the genes and two-thirds of the transposons. Considering the enormous size of the query, the fact that only two false-positive matches were reported emphasizes the high selectivity of protein family-based methods for gene finding. We used the search results to improve BLOCKS+ by identifying compositionally biased blocks. Our results confirm that protein family databases can be used effectively in automated sequence annotation efforts.

Alcohol Dehydrogenase↗

Malaria and the red blood cell membrane.

Malaria is the most serious and widespread parasitic disease of humans and is arguably the commonest disease of red blood cells (RBCs). Malaria has exerted a powerful effect on human evolution and selection for resistance has led to the appearance and persistence of a number of inherited diseases. After parasite invasion, RBCs are progressively and dramatically modified. New structures appear inside the RBC and novel parasite proteins are exported to the erythrocyte cytoplasm and membrane skeleton. Radical biochemical, morphological, and rheological alterations manifest as increased membrane rigidity, reduced cell deformability, and greater adhesiveness for the vascular endothelium and other blood cells. Numerous protein-protein interactions between the malaria-parasite and the host RBC are important for many aspects of parasite biology and the pathogenesis of malaria. In addition, there are many other parasite proteins located within the infected red cell and at the membrane skeleton, for which no precise functional roles have yet been elucidated. Sequencing and annotation of the complete genome of Plasmodium falciparum, the production of proteomic and transcriptomic profiles of parasites, and the development of a transfection system for the asexual stage of the parasite are all recent achievements that should advance understanding of the molecular mechanisms that underlie the parasite-induced functional alterations in red cells.

Animals↗

Mining protein function from text using term-based support vector machines.

BACKGROUND: Text mining has spurred huge interest in the domain of biology. The goal of the BioCreAtIvE exercise was to evaluate the performance of current text mining systems. We participated in Task 2, which addressed assigning Gene Ontology terms to human proteins and selecting relevant evidence from full-text documents. We approached it as a modified form of the document classification task. We used a supervised machine-learning approach (based on support vector machines) to assign protein function and select passages that support the assignments. As classification features, we used a protein's co-occurring terms that were automatically extracted from documents. RESULTS: The results evaluated by curators were modest, and quite variable for different problems: in many cases we have relatively good assignment of GO terms to proteins, but the selected supporting text was typically non-relevant (precision spanning from 3% to 50%). The method appears to work best when a substantial set of relevant documents is obtained, while it works poorly on single documents and/or short passages. The initial results suggest that our approach can also mine annotations from text even when an explicit statement relating a protein to a GO term is absent. CONCLUSION: A machine learning approach to mining protein function predictions from text can yield good performance only if sufficient training data is available, and significant amount of supporting data is used for prediction. The most promising results are for combined document retrieval and GO term assignment, which calls for the integration of methods developed in BioCreAtIvE Task 1 and Task 2.

Computational Biology↗

Protein binding microarrays (PBMs) for rapid, high-throughput characterization of the sequence specificities of DNA binding proteins.

DNA binding proteins play a number of key roles in cells, in processes including transcriptional regulation, recombination, genome rearrangements, and DNA replication, repair, and modification. Of particular interest are the interactions between transcription factors and their DNA binding sites, as they are an integral part of the transcriptional regulatory networks that control gene expression. Despite their importance, the DNA binding specificities of most DNA binding proteins remain unknown, as earlier technologies aimed at characterizing DNA-protein interactions have been time consuming and not highly scalable. We have developed a new DNA microarray-based technology, termed protein binding microarrays (PBMs), that allows rapid, high-throughput characterization of the in vitro DNA binding site sequence specificities of transcription factors in a single day. The resulting DNA binding site data can be used in a number of ways, including for the prediction of the genes regulated by a given transcription factor, annotation of transcription factor function, and functional annotation of the predicted target genes.

Base Sequence↗

Prediction of peroxisomal targeting signal 1 containing proteins from amino acid sequence.

Peroxisomal matrix proteins have to be imported into their target organelle post-translationally. The major translocation pathway depends on a C-terminal targeting signal, termed PTS1. Our previous analysis of sequence variability in the PTS1 motif revealed that, in addition to the known C-terminal tripeptide, at least nine residues directly upstream are important for signal recognition in the PTS1-Pex5 receptor complex. The refined PTS1 motif description was implemented in a prediction tool composed of taxon-specific functions (metazoa, fungi, remaining taxa), capable of recognising potential PTS1s in query sequences. The composite score function consists of classical profile terms and additional terms penalising deviations from the derived physical property pattern over sequence segments. The prediction algorithm has been validated with a self-consistency and three different cross-validation tests. Additionally, we tested the tool on a large set of non-peroxisomal negatives, on mutation data, and compared the prediction rate to the PTS1 component of the PSORT2 program. The sensitivity of our predictor in recognising documented PTS1 signal containing proteins is close to 90% for reliable prediction. The predictor distinguishes even SKL-appended non-peroxisomally targeted proteins such as a mouse dihydrofolate reductase-SKL construct. The corresponding rate of false positives is not worse than 0.8%; thus, the tool can be applied for large-scale unsupervised sequence database annotation. A scan of public protein databases uncovered a number of yet uncharacterised proteins for which the PTS1 signal might be critical for biological function. The predicted presence of a PTS1 signal implies peroxisomal localisation in the absence of N-terminal targeting sequences such as the mitochondrial import signal.

Algorithms↗

The TRIPLES database: a community resource for yeast molecular biology.

TRIPLES is a web-accessible database of TRansposon-Insertion Phenotypes, Localization and Expression in Saccharomyces cerevisiae-a relational database housing nearly half a million data points generated from an ongoing study using large-scale transposon mutagenesis to characterize gene function in yeast. At present, TRIPLES contains three principal data sets (i.e. phenotypic data, protein localization data and expression data) for over 3500 annotated yeast genes as well as several hundred non-annotated open reading frames. In addition, the TRIPLES web site provides online order forms linked to each data set so that users may request any strain or reagent generated from this project free of charge. In response to user requests, the TRIPLES web site has undergone several recent modifications. Our localization data have been supplemented with approximately 500 fluorescent micrographs depicting actual staining patterns observed upon indirect immunofluorescence analysis of indicated epitope-tagged proteins. These localization data, as well as all other data sets within TRIPLES, are now available in full as tab-delimited text. To accommodate increased reagent requests, all orders are now cataloged in a separate database, and users are notified immediately of order receipt and shipment. Also, TRIPLES is one of five sites incorporated into the new functional analysis tool Function Junction provided by the Saccharomyces Genome Database. TRIPLES may be accessed from the Yale Genome Analysis Center (YGAC) homepage at http://ygac.med.yale.edu.

Computer Graphics↗

Prediction of functional sites by analysis of sequence and structure conservation.

We present a method for prediction of functional sites in a set of aligned protein sequences. The method selects sites which are both well conserved and clustered together in space, as inferred from the 3D structures of proteins included in the alignment. We tested the method using 86 alignments from the NCBI CDD database, where the sites of experimentally determined ligand and/or macromolecular interactions are annotated. In agreement with earlier investigations, we found that functional site predictions are most successful when overall background sequence conservation is low, such that sites under evolutionary constraint become apparent. In addition, we found that averaging of conservation values across spatially clustered sites improves predictions under certain conditions: that is, when overall conservation is relatively high and when the site in question involves a large macromolecular binding interface. Under these conditions it is better to look for clusters of conserved sites than to look for particular conserved sites.

Algorithms↗

An atlas of differential gene expression during early Xenopus embryogenesis.

We have carried out a large-scale, semi-automated whole-mount in situ hybridization screen of 8369 cDNA clones in Xenopus laevis embryos. We confirm that differential gene expression is prevalent during embryogenesis since 24% of the clones are expressed non-ubiquitously and 8% are organ or cell type specific marker genes. Sequence analysis and clustering yielded 723 unique genes displaying a differential expression pattern. Of these, 18% were already described in Xenopus, 47% have homologs and 35% are lacking significant sequence similarity in databases. Many of them encode known developmental regulators. We classified 363 of the 723 genes for which a Gene Ontology annotation for molecular function could be attributed and found 'DNA binding' and 'enzyme' the most represented terms. The most common protein domains encoded in these embryonic, differentially expressed genes are the homeobox and RNA Recognition Motif (RRM). Fifty-nine putative orthologs of human disease genes, and 254 organ or cell specific marker genes were identified. Markers were found for nasal placode and archenteron roof, organs for which a specific marker was previously unavailable. Markers were also found for novel subdomains of various other organs. The tissues for which most markers were found are muscle and epidermis. Expression of cell cycle regulators fell in two classes, containing proliferation-promoting and anti-proliferative genes, respectively. We identified 66 new members of the BMP4, chromatin, endoplasmic reticulum, and karyopherin synexpression groups, thus providing a first glimpse of their probable cellular roles. Cluster analysis of tissues to measure tissue relatedness yielded some unorthodox affinities besides expectable lineage relationships. In conclusion, this study represents an atlas of gene expression patterns, which reveals embryonic regionalization, provides novel marker genes, and makes predictions about the functional role of unknown genes.

Animals↗

A novel type of RNase III family proteins in eukaryotes.

The RNase III family of double-stranded RNA-specific endonucleases is characterized by the presence of a highly conserved 9 amino acid stretch in their catalytic center known as the RNase III signature motif. We isolated the drosha gene, a new member of this family in Drosophila melanogaster. Characterization of this gene revealed the presence of two RNase III signature motifs in its sequence that may indicate that it is capable of forming an active catalytic center as a monomer. The drosha protein also contains an 825 amino acid N-terminus with an unknown function. A search for the known homologues of the drosha protein revealed that it has a similarity to two adjacent annotated genes identified during C. elegans genome sequencing. Analysis of the genomic region of these genes by the Fgenesh program and sequencing of the EST cDNA clone derived from it revealed that this region encodes only one gene. This newly identified gene in nematode genome shares a high similarity to Drosophila drosha throughout its entire protein sequence. A potential drosha homologue is also found among the deposited human cDNA sequences. A comparison of these drosha proteins to other members of the RNase III family indicates that they form a new group of proteins within this family.

Amino Acid Sequence↗

Functional and structural genomics using PEDANT.

MOTIVATION: Enormous demand for fast and accurate analysis of biological sequences is fuelled by the pace of genome analysis efforts. There is also an acute need in reliable up-to-date genomic databases integrating both functional and structural information. Here we describe the current status of the PEDANT software system for high-throughput analysis of large biological sequence sets and the genome analysis server associated with it. RESULTS: The principal features of PEDANT are: (i) completely automatic processing of data using a wide range of bioinformatics methods, (ii) manual refinement of annotation, (iii) automatic and manual assignment of gene products to a number of functional and structural categories, (iv) extensive hyperlinked protein reports, and (v) advanced DNA and protein viewers. The system is easily extensible and allows to include custom methods, databases, and categories with minimal or no programming effort. PEDANT is actively used as a collaborative environment to support several on-going genome sequencing projects. The main purpose of the PEDANT genome database is to quickly disseminate well-organized information on completely sequenced and unfinished genomes. It currently includes 80 genomic sequences and in many cases serves as the only source of exhaustive information on a given genome. The database also acts as a vehicle for a number of research projects in bioinformatics. Using SQL queries, it is possible to correlate a large variety of pre-computed properties of gene products encoded in complete genomes with each other and compare them with data sets of special scientific interest. In particular, the availability of structural predictions for over 300 000 genomic proteins makes PEDANT the most extensive structural genomics resource available on the web.

Arabidopsis↗

The yeast proteome database (YPD) and Caenorhabditis elegans proteome database (WormPD): comprehensive resources for the organization and comparison of model organism protein information.

The Yeast Proteome Database (YPDtrade mark) has been for several years a resource for organized and accessible information about the proteins of Saccharomyces cerevisiae. We have now extended the YPD format to create a database containing complete proteome information about the model organism Caenorhabditis elegans (WormPDtrade mark). YPD and WormPD are designed for use not only by their respective research communities but also by the broader scientific community. In both databases, information gleaned from the literature is presented in a consistent, user-friendly Protein Report format: a single Web page presenting all available knowledge about a particular protein. Each Protein Report begins with a Title Line, a concise description of the function of that protein that is continually updated as curators review new literature. Properties and functions of the protein are presented in tabular form in the upper part of the Report, and free-text annotations organized by topic are presented in the lower part. Each Protein Report ends with a comprehensive reference list whose entries are linked to their MEDLINE s. YPD and WormPD are seamlessly integrated, with extensive links between the species. They are freely accessible to academic users on the WWW at http://www. proteome.com/databases/index.html, and are available by subscription to corporate users.

Animals↗

AraC-XylS database: a family of positive transcriptional regulators in bacteria.

The AraC-XylS database contains information about a family of positive transcriptional regulators broadly distributed in bacteria. This specific database focuses on protein sequences and on the biological and functional features of each of the proteins that belong to this family. Each entry provides information on the protein itself, the annotated protein sequence and, when the crystal is available, a comprehensive representation of its three-dimensional structure. The organization of the database is based on an exhaustive analysis of the scientific literature. The data are interconnected and linked with other databases. Multiple alignments of the members of the family, an extensive collection of references and a tutorial about the family provide additional information. The AraC-XylS database is accessible on the World Wide Web at http://www.AraC-XylS.org.

Amino Acid Sequence↗

CancerGenes: a gene selection resource for cancer genome projects.

The genome sequence framework provided by the human genome project allows us to precisely map human genetic variations in order to study their association with disease and their direct effects on gene function. Since the description of tumor suppressor genes and oncogenes several decades ago, both germ-line variations and somatic mutations have been established to be important in cancer-in terms of risk, oncogenesis, prognosis and response to therapy. The Cancer Genome Atlas initiative proposed by the NIH is poised to elucidate the contribution of somatic mutations to cancer development and progression through the re-sequencing of a substantial fraction of the total collection of human genes-in hundreds of individual tumors and spanning several tumor types. We have developed the CancerGenes resource to simplify the process of gene selection and prioritization in large collaborative projects. CancerGenes combines gene lists annotated by experts with information from key public databases. Each gene is annotated with gene name(s), functional description, organism, chromosome number, location, Entrez Gene ID, GO terms, InterPro descriptions, gene structure, protein length, transcript count, and experimentally determined transcript control regions, as well as links to Entrez Gene, COSMIC, and iHOP gene pages and the UCSC and Ensembl genome browsers. The user-friendly interface provides for searching, sorting and intersection of gene lists. Users may view tabulated results through a web browser or may dynamically download them as a spreadsheet table. CancerGenes is available at http://cbio.mskcc.org/cancergenes.

Databases, Genetic↗

Evaluation of methods for determination of a reconstructed history of gene sequence evolution.

With whole-genome sequences being completed at an increasing rate, it is important to develop and assess tools to analyze them. Following annotation of the protein content of a genome, one can compare sequences with previously characterized homologous genes to detect novel functions within specific proteins in the evolution of the newly sequenced genome. One common statistical method to detect such changes is to compare the ratios of nonsynonymous (K(a)) to synonymous (K(s)) nucleotide substitution rates. Here, the effects of several parameters that can influence this calculation (sequence reconstruction method, phylogenetic tree branch length weighting, GC content, and codon bias) are examined. Also, two new alternative measures of adaptive evolution, the point accepted mutations (PAM)/neutral evolutionary distance (NED) ratio and the sequence space assessment (SSA) statistic are presented. All of these methods are compared using two sequence families: the recent divergence of leptin orthologs in primates, and the more ancient divergence of the deoxyribonucleoside kinase family. The examination of these and other measures to detect changes of gene function along branches of a phylogenetic tree will become increasingly important in the postgenomic era.

Algorithms↗