Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Validation of Shewanella oneidensis MR-1 small proteins by AMT tag-based proteome analysis.

Using stringent criteria for protein identification by accurate mass and time (AMT) tag mass spectrometric methodology, we detected 36 proteins of <101 amino acids in length, including 10 that were annotated as hypothetical proteins, in 172 global tryptic digests of Shewanella oneidensis MR-1 proteins. Peptides that map to the conserved, but functionally uncharacterized proteins SO4134 and SO2787, were the most frequently detected peptides in these samples, while those that map to hypotheticals SO2669 and SO2063, conserved hypotheticals SO0335 and SO2176, and the SlyX protein (SO1063) were observed at frequencies similar to those from essential small proteins (ribosomal proteins and translation initiation factor IF-1), suggesting that they may function in similarly important cellular functions. In addition, peptides were detected that map to 30 genes predicted to encode frameshifts, point mutations, or recoding signals. Of these 30 genes, peptides that map to positions beyond internal stop codons were detected in 13 genes (SO0101, SO0419, SO0590, SO0738, SO1113, SO1211, SO3079, SO3130, SO3240, SO4231, SO4328, SO4422, and SO4657). While expression of the full-length formate dehydrogenase encoded by SO0101 can be explained by incorporation of selenocysteine at the internal stop codon, the mechanism of translating downstream sequences in the remaining genes remains unknown.

Amino Acid Sequence↗

The Schistosoma mansoni gene index: gene discovery and biology by reconstruction and analysis of expressed gene sequences.

Expressed sequence tag (EST) sequencing and analysis is a primary research tool to identify and characterize the Schistosoma mansoni transcriptome. As part of our gene discovery effort, a total of 5,793 ESTs have been generated from clones selected randomly from complementary DNA (cDNA) libraries constructed from male and female adult worms. Assembly analysis of all the 16,813 public S. mansoni ESTs has identified 1,920 distinct tentative consensus sequences (TCs) and 5,571 nonoverlapping ESTs (singletons). Of these, 376 TCs (20%) and 1,449 singletons (26%) are unique to the SUNY/TIGR sequencing effort. Tentative consensus sequences and singletons were distributed into various categories of biological roles associated with cell structure, metabolism, protein fate, signal transduction, transcription, protein synthesis, transporters, and cell growth. The TCs and singletons represent transcripts that can be used as a resource for functional annotation of genomic sequence data, comparative sequence analysis, and cDNA clone selection for microarray projects. The utility of EST analysis is demonstrated by identifying new protease genes, which may be involved in hemoglobin degradation.

Amino Acid Sequence↗

iProClass: an integrated database of protein family, function and structure information.

The iProClass database provides comprehensive, value-added descriptions of proteins and serves as a framework for data integration in a distributed networking environment. The protein information in iProClass includes family relationships as well as structural and functional classifications and features. The current version consists of about 830 000 non-redundant PIR-PSD, SWISS-PROT, and TrEMBL proteins organized with more than 36 000 PIR superfamilies, 145 000 families, 4000 domains, 1300 motifs and 550 000 FASTA similarity clusters. It provides rich links to over 50 database of protein sequences, families, functions and pathways, protein-protein interactions, post-translational modifications, protein expressions, structures and structural classifications, genes and genomes, ontologies, literature and taxonomy. Protein and superfamily summary reports present extensive annotation information and include membership statistics and graphical display of domains and motifs. iProClass employs an open and modular architecture for interoperability and scalability. It is implemented in the Oracle object-relational database system and is updated biweekly. The database is freely accessible from the web site at http://pir.georgetown.edu/iproclass/ and searchable by sequence or text string. The data integration in iProClass supports exploration of protein relationships. Such knowledge is fundamental to the understanding of protein evolution, structure and function and crucial to functional genomic and proteomic research.

Amino Acid Motifs↗

Characterization and comparative analysis of the EGLN gene family.

Rat Sm-20 is a homologue of the Caenorhabditis elegans gene egl-9 and has been implicated in the regulation of growth, differentiation and apoptosis in muscle and nerve cells. Null mutants in egl-9 result in a complete tolerance to an otherwise lethal toxin produced by Pseudomonas aeruginosa. This study describes the conserved Egl-Nine (EGLN) gene family of which rat SM-20 and C. elegans Egl-9 are members and characterizes the mouse and human homologues. Each of the human genes (EGLN1, EGLN2 and EGLN3) are of a conserved genomic structure consisting of five coding exons. Phylogenetic analysis and domain organization show that EGLN1 represents the ancestral form of the gene family and that EGLN3 is the human orthologue of rat Sm-20. The previously observed mitochondrial targeting of rat SM-20 is unlikely to be a general feature of the protein family and may be a feature specific to rats. An EGLN gene is unexpectedly found in the genome of P. aeruginosa, a bacterium known to produce a toxin that acts through the Egl-9 protein. The pathogenic bacterium Vibrio cholerae is also shown to have an EGLN gene suggesting that it is an important pathogenicity factor. These results provide new insights into host-pathogen interactions and a basis for further functional characterization of the gene family and resolve discrepancies in annotation between gene family members.

Amino Acid Sequence↗

Non-EST based prediction of exon skipping and intron retention events using Pfam information.

Most of the known alternative splice events have been detected by the comparison of expressed sequence tags (ESTs) and cDNAs. However, not all splice events are represented in EST databases since ESTs have several biases. Therefore, non-EST based approaches are needed to extend our view of a transcriptome. Here, we describe a novel method for the ab initio prediction of alternative splice events that is solely based on the annotation of Pfam domains. Furthermore, we applied this approach in a genome-wide manner to all human RefSeq transcripts and predicted a total of 321 exon skipping and intron retention events. We show that this method is very reliable as 78% (250 of 321) of our predictions are confirmed by ESTs or cDNAs. Subsequent analyses of splice events within Pfam domains revealed a significant preference of alternative exon junctions to be located at the protein surface and to avoid secondary structure elements. Thus, splice events within Pfams are probable to alter the structure and function of a domain which makes them highly interesting for detailed biological investigation. As Pfam domains are annotated in many other species, our strategy to predict exon skipping and intron retention events might be important for species with a lower number of ESTs.

Algorithms↗

The 'permeome' of the malaria parasite: an overview of the membrane transport proteins of Plasmodium falciparum.

BACKGROUND: The uptake of nutrients, expulsion of metabolic wastes and maintenance of ion homeostasis by the intraerythrocytic malaria parasite is mediated by membrane transport proteins. Proteins of this type are also implicated in the phenomenon of antimalarial drug resistance. However, the initial annotation of the genome of the human malaria parasite Plasmodium falciparum identified only a limited number of transporters, and no channels. In this study we have used a combination of bioinformatic approaches to identify and attribute putative functions to transporters and channels encoded by the malaria parasite, as well as comparing expression patterns for a subset of these. RESULTS: A computer program that searches a genome database on the basis of the hydropathy plots of the corresponding proteins was used to identify more than 100 transport proteins encoded by P. falciparum. These include all the transporters previously annotated as such, as well as a similar number of candidate transport proteins that had escaped detection. Detailed sequence analysis enabled the assignment of putative substrate specificities and/or transport mechanisms to all those putative transport proteins previously without. The newly-identified transport proteins include candidate transporters for a range of organic and inorganic nutrients (including sugars, amino acids, nucleosides and vitamins), and several putative ion channels. The stage-dependent expression of RNAs for 34 candidate transport proteins of particular interest are compared. CONCLUSION: The malaria parasite possesses substantially more membrane transport proteins than was originally thought, and the analyses presented here provide a range of novel insights into the physiology of this important human pathogen.

Amino Acid Sequence↗

Consistency checks for characterizing protein forms.

Proteomics enforces the reverse chronological order on the gene to protein dogma and imposes amino acid sequences as a starting point of an investigation relative to function. By this approach, proteomics data can confirm the presence of multiple forms of a protein. Notwithstanding variations attributed specific individual features of organisms and tissues, from two to over ten protein forms can be identified in a given sample. The present work describes some guidelines for tracking the origin of alternative protein forms and attempts to tag the details of sequence data in the literature. Working via these guidelines we have uncovered a third alternative form of the Pim subfamily of oncogenes. The term form is here combined with the qualification alternative to describe any product of a given gene including closely related paralogs. This paper also emphasizes the need for consistency checks in annotation processes, such as gene clustering, to avoid losing important details describing protein alternative forms. By identifying alternative protein forms, we illustrate the fact that rationalizing of protein function via the identification of protein-protein interactions should in reality be that of identifying (alternative) form-form interactions.

Amino Acid Sequence↗

A simple algorithm to infer gene duplication and speciation events on a gene tree.

MOTIVATION: When analyzing protein sequences using sequence similarity searches, orthologous sequences (that diverged by speciation) are more reliable predictors of a new protein's function than paralogous sequences (that diverged by gene duplication), because duplication enables functional diversification. The utility of phylogenetic information in high-throughput genome annotation ('phylogenomics') is widely recognized, but existing approaches are either manual or indirect (e.g. not based on phylogenetic trees). Our goal is to automate phylogenomics using explicit phylogenetic inference. A necessary component is an algorithm to infer speciation and duplication events in a given gene tree. RESULTS: We give an algorithm to infer speciation and duplication events on a gene tree by comparison to a trusted species tree. This algorithm has a worst-case running time of O(n(2)) which is inferior to two previous algorithms that are approximately O(n) for a gene tree of sequences. However, our algorithm is extremely simple, and its asymptotic worst case behavior is only realized on pathological data sets. We show empirically, using 1750 gene trees constructed from the Pfam protein family database, that it appears to be a practical (and often superior) algorithm for analyzing real gene trees. AVAILABILITY: http://www.genetics.wustl.edu/eddy/forester.

Algorithms↗

Structural characterization of proteins using residue environments.

A primary challenge for structural genomics is the automated functional characterization of protein structures. We have developed a sequence-independent method called S-BLEST (Structure-Based Local Environment Search Tool) for the annotation of previously uncharacterized protein structures. S-BLEST encodes the local environment of an amino acid as a vector of structural property values. It has been applied to all amino acids in a nonredundant database of protein structures to generate a searchable structural resource. Given a query amino acid from an experimentally determined or modeled structure, S-BLEST quickly identifies similar amino acid environments using a K-nearest neighbor search. In addition, the method gives an estimation of the statistical significance of each result. We validated S-BLEST on X-ray crystal structures from the ASTRAL 40 nonredundant dataset. We then applied it to 86 crystallographically determined proteins in the protein data bank (PDB) with unknown function and with no significant sequence neighbors in the PDB. S-BLEST was able to associate 20 proteins with at least one local structural neighbor and identify the amino acid environments that are most similar between those neighbors.

Amino Acids↗

Integrative analysis of cancer-related data using CAP.

The development of human cancer is a highly complex process and can be considered the result of several combined events, such as genetic alterations, disturbance of signal transduction, or failure of immunological surveillance. Cancer-related databases usually focus on specific fields of research, e.g., cancer genetics or cancer immunology, whereas the complexity of cancer genesis requires an integrated analysis of heterogeneous data from several sources. Here we present the cancer-associated protein database (CAP), a novel analysis system for cancer-related data. CAP integrates data from multiple external databases, augments these data with functional annotations, and offers tools for statistical analysis of these data. We have employed CAP to analyze genes that have been found to cause an autoimmune response in cancer. In particular, we explored the connection between the autoimmune response, mutations, and overexpression of these genes. Our preliminary results suggest that mutations are not significant contributors to raising an antibody response against tumor antigens, whereas overexpression seems to play a more important role. We hereby demonstrate how different types of data can be integrated and analyzed successfully, providing interesting results. As the amount of available data is growing rapidly, a combined analysis will play an important role in exploring the genetic and immunological basis of cancer. CAP is freely available at the following web site: http://www.bioinf.uni-sb.de/CAP/.

Autoimmunity↗

The SNF2 domain protein family in higher vertebrates displays dynamic expression patterns in Xenopus laevis embryos.

All eukaryotes share a common nuclear infrastructure, in which DNA is packaged into nucleosomal chromatin. Its functional states, in particular the accessibility of the chromatin fiber to trans-acting factors, are determined by two classes of evolutionarily conserved enzymes, i.e. histone modifying enzymes and ATP-driven nucleosome remodeling machines. Browsing the annotated human genome database, we establish here a family of SNF2-like nuclear ATPases, which are the core enzymatic subunits of chromatin remodeling protein complexes. Homologues of those human genes are also to a large extent found in the Xenopus laevis genome, indicating a high degree of sequence conservation of this family among vertebrates. Expression analyses of the ATPase family of proteins reveal stage- and tissue-specific domains of peak RNA expression during early frog embryogenesis. These dynamic expression profiles suggest specific functional requirements for individual members of this family throughout early stages of vertebrate development.

Adenosine Triphosphatases↗

Search for Bacillus anthracis potential vaccine candidates by a functional genomic-serologic screen.

Bacillus anthracis proteins that possess antigenic properties and are able to evoke an immune response were identified by a reductive genomic-serologic screen of a set of in silico-preselected open reading frames (ORFs). The screen included in vitro expression of the selected ORFs by coupled transcription and translation of linear PCR-generated DNA fragments, followed by immunoprecipitation with antisera from B. anthracis-infected animals. Of the 197 selected ORFs, 161 were chromosomal and 36 were on plasmids pXO1 and pXO2, and 138 of the 197 ORFs had putative functional annotations (known ORFs) and 59 had no assigned functions (unknown ORFs). A total of 129 of the known ORFs (93%) could be expressed, whereas only 38 (64%) of the unknown ORFs were successfully expressed. All 167 expressed polypeptides were subjected to immunoprecipitation with the anti-B. anthracis antisera, which revealed 52 seroreactive immunogens, only 1 of which was encoded by an unknown ORF. The high percentage of seroreactive ORFs among the functionally annotated ORFs (37%; 51/129) attests to the predictive value of the bioinformatic strategy used for vaccine candidate selection. Furthermore, the experimental findings suggest that surface-anchored proteins and adhesins or transporters, such as cell wall hydrolases, proteins involved in iron acquisition, and amino acid and oligopeptide transporters, have great potential to be immunogenic. Most of the seroreactive ORFs that were tested as DNA vaccines indeed appeared to induce a humoral response in mice. We list more than 30 novel B. anthracis immunoreactive virulence-related proteins which could be useful in diagnosis, pathogenesis studies, and future anthrax vaccine development.

Animals↗

ZCURVE_CoV: a new system to recognize protein coding genes in coronavirus genomes, and its applications in analyzing SARS-CoV genomes.

A new system to recognize protein coding genes in the coronavirus genomes, specially suitable for the SARS-CoV genomes, has been proposed in this paper. Compared with some existing systems, the new program package has the merits of simplicity, high accuracy, reliability, and quickness. The system ZCURVE_CoV has been run for each of the 11 newly sequenced SARS-CoV genomes. Consequently, six genomes not annotated previously have been annotated, and some problems of previous annotations in the remaining five genomes have been pointed out and discussed. In addition to the polyprotein chain ORFs 1a and 1b and the four genes coding for the major structural proteins, spike (S), small envelop (E), membrane (M), and nuleocaspid (N), respectively, ZCURVE_CoV also predicts 5-6 putative proteins in length between 39 and 274 amino acids with unknown functions. Some single nucleotide mutations within these putative coding sequences have been detected and their biological implications are discussed. A web service is provided, by which a user can obtain the annotated result immediately by pasting the SARS-CoV genome sequences into the input window on the web site (http://tubic.tju.edu.cn/sars/). The software ZCURVE_CoV can also be downloaded freely from the web address mentioned above and run in computers under the platforms of Windows or Linux.

Algorithms↗

DBD: a transcription factor prediction database.

Regulation of gene expression influences almost all biological processes in an organism; sequence-specific DNA-binding transcription factors are critical to this control. For most genomes, the repertoire of transcription factors is only partially known. Hitherto transcription factor identification has been largely based on genome annotation pipelines that use pairwise sequence comparisons, which detect only those factors similar to known genes, or on functional classification schemes that amalgamate many types of proteins into the category of 'transcription factor'. Using a novel transcription factor identification method, the DBD transcription factor database fills this void, providing genome-wide transcription factor predictions for organisms from across the tree of life. The prediction method behind DBD identifies sequence-specific DNA-binding transcription factors through homology using profile hidden Markov models (HMMs) of domains. Thus, it is limited to factors that are homologus to those HMMs. The collection of HMMs is taken from two existing databases (Pfam and SUPERFAMILY), and is limited to models that exclusively detect transcription factors that specifically recognize DNA sequences. It does not include basal transcription factors or chromatin-associated proteins, for instance. Based on comparison with experimentally verified annotation, the prediction procedure is between 95% and 99% accurate. Between one quarter and one-half of our genome-wide predicted transcription factors represent previously uncharacterized proteins. The DBD (www.transcriptionfactor.org) consists of predicted transcription factor repertoires for 150 completely sequenced genomes, their domain assignments and the hand curated list of DNA-binding domain HMMs. Users can browse, search or download the predictions by genome, domain family or sequence identifier, view families of transcription factors based on domain architecture and receive predictions for a protein sequence.

Animals↗

Caenorhabditis elegans has two genes encoding functional d-aspartate oxidases.

Four cDNA clones that were annotated in the database as encoding d-amino acid oxidase (DAAO) or d-aspartate oxidase (DASPO) were isolated by RT-PCR from Caenorhabditis elegans RNA. The proteins (Y69Ap, C47Ap, F18Ep, and F20Hp) encoded by the cloned cDNAs were expressed in Escherichia coli as recombinant proteins with an N-terminal His-tag. All proteins except F20Hp were recovered in the soluble fractions. The recombinant Y69Ap has functional DAAO activity, as it can deaminate neutral and basic d-amino acids, whereas the recombinants C47Ap and F18Ep have functional DASPO activities, as they can deaminate acidic d-amino acids. Additional experiments using purified recombinant proteins revealed that Y69Ap deaminates d-Arg more efficiently than d-Ala and d-Met, and that C47Ap and F18Ep show distinct kinetic properties against d-Asp, d-Glu, and N-methyl-d-Asp. This is the first time that cDNA cloning of invertebrate DAAO and DASPO genes has been reported. In addition, our study reveals for the first time that C. elegans has at least two genes encoding functional DASPOs and one gene encoding DAAO, although it had previously been thought that organisms only bear one copy each of these genes. The two C. elegans DASPOs differ in their substrate specificities and possibly also in their subcellular localization.

Amino Acid Sequence↗

sc-PDB: an annotated database of druggable binding sites from the Protein Data Bank.

The sc-PDB is a collection of 6 415 three-dimensional structures of binding sites found in the Protein Data Bank (PDB). Binding sites were extracted from all high-resolution crystal structures in which a complex between a protein cavity and a small-molecular-weight ligand could be identified. Importantly, ligands are considered from a pharmacological and not a structural point of view. Therefore, solvents, detergents, and most metal ions are not stored in the sc-PDB. Ligands are classified into four main categories: nucleotides (< 4-mer), peptides (< 9-mer), cofactors, and organic compounds. The corresponding binding site is formed by all protein residues (including amino acids, cofactors, and important metal ions) with at least one atom within 6.5 angstroms of any ligand atom. The database was carefully annotated by browsing several protein databases (PDB, UniProt, and GO) and storing, for every sc-PDB entry, the following features: protein name, function, source, domain and mutations, ligand name, and structure. The repository of ligands has also been archived by diversity analysis of molecular scaffolds, and several chemoinformatics descriptors were computed to better understand the chemical space covered by stored ligands. The sc-PDB may be used for several purposes: (i) screening a collection of binding sites for predicting the most likely target(s) of any ligand, (ii) analyzing the molecular similarity between different cavities, and (iii) deriving rules that describe the relationship between ligand pharmacophoric points and active-site properties. The database is periodically updated and accessible on the web at http://bioinfo-pharma.u-strasbg.fr/scPDB/.

Algorithms↗

Predicting phenotype from patterns of annotation.

MOTIVATION: Predicting the outcome of specific experiments (such as the growth of a particular mutant strain in a particular medium) has the potential to allow researchers to devote resources to experiments with higher expected numbers of 'hits'. RESULTS: We use decision trees to predict phenotypes associated with Saccharomyces cerevisiae genes on the basis of Gene Ontology (GO) functional annotations from the Saccharomyces Genome Database (SGD) and other phenotypic annotations from the Yeast Phenotype Catalog at the Munich Information Center for Protein Sequences (MIPS). We assess the methodology in three ways: (1) we use cross-validation on the phenotypic annotations listed in MIPS, and show ROC curves indicating the tradeoff between true-positive rate and false-positive rate; (2) we do a literature-search for 100 of the predicted gene-phenotype associations that are not listed in MIPS, and find evidence for 43 of them; (3) we use deletion strains to experimentally assess 61 predicted gene-phenotype associations not listed in MIPS; significantly more of these deletion strains show abnormal growth than would be expected by chance.

Algorithms↗

Large-scale identification and characterization of alternative splicing variants of human gene transcripts using 56,419 completely sequenced and manually annotated full-length cDNAs.

We report the first genome-wide identification and characterization of alternative splicing in human gene transcripts based on analysis of the full-length cDNAs. Applying both manual and computational analyses for 56,419 completely sequenced and precisely annotated full-length cDNAs selected for the H-Invitational human transcriptome annotation meetings, we identified 6877 alternative splicing genes with 18 297 different alternative splicing variants. A total of 37,670 exons were involved in these alternative splicing events. The encoded protein sequences were affected in 6005 of the 6877 genes. Notably, alternative splicing affected protein motifs in 3015 genes, subcellular localizations in 2982 genes and transmembrane domains in 1348 genes. We also identified interesting patterns of alternative splicing, in which two distinct genes seemed to be bridged, nested or having overlapping protein coding sequences (CDSs) of different reading frames (multiple CDS). In these cases, completely unrelated proteins are encoded by a single locus. Genome-wide annotations of alternative splicing, relying on full-length cDNAs, should lay firm groundwork for exploring in detail the diversification of protein function, which is mediated by the fast expanding universe of alternative splicing variants.

Alternative Splicing↗