Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Predicting ligand-binding function in families of bacterial receptors.

The three-dimensional fold of a new protein sequence can often be inferred directly from sequence homology to a protein of known structure. The function of a new protein sequence is more difficult to predict, however, since homologues can have different molecular and cellular functions. To develop and automate computational methods for determining molecular function, we have analyzed ligand-binding specificity in two related families of binding proteins. One of these families includes Escherichia coli lactose repressor and ribose-binding protein, and the other includes E. coli sulfate- and phosphate-binding proteins. These proteins have similar folds but varying specificity, binding many different small molecules, including mono- and disaccharides, purines, oxyanions, ferric iron, and polyamines. Starting from template structural alignments, alignments of over 90 sequences per family were generated by iterative database searches with hidden Markov models. Phylogenetic trees were made of full-length sequences and of subsets of residues lining the binding cleft, to determine whether subbranches of the trees correlate with ligand-binding preference. Automated analyses of residues in the binding pocket were also used to predict ligand-binding function for many uncharacterized database sequences and to identify specific side chain-ligand contacts in proteins without solved structures. Our results demonstrate the utility of anchoring functional annotation within a protein family context.

Amino Acid Sequence↗

MitoP2: the mitochondrial proteome database--now including mouse data.

The MitoP2 database (http://www.mitop.de) integrates information on mitochondrial proteins, their molecular functions and associated diseases. The central database features are manually annotated reference proteins localized or functionally associated with mitochondria supplied for yeast, human and mouse. MitoP2 enables (i) the identification of putative orthologous proteins between these species to study evolutionarily conserved functions and pathways; (ii) the integration of data from systematic genome-wide studies such as proteomics and deletion phenotype screening; (iii) the prediction of novel mitochondrial proteins using data integration and the assignment of evidence scores; and (iv) systematic searches that aim to find the genes that underlie common and rare mitochondrial diseases. The data and analysis files are referenced to data sources in PubMed and other online databases and can be easily downloaded. MitoP2 users can explore the relationship between mitochondrial dysfunctions and disease and utilize this information to conduct systems biology approaches on mitochondria.

Animals↗

Design and characterization of a functional library for NMR screening against novel protein targets.

In the past few years, NMR has been extensively utilized as a screening tool for drug discovery using various types of compound libraries. The designs of NMR specific chemical libraries that utilize a fragment-based approach based on drug-like characteristics have been previously reported. In this article, a new type of compound library will be described that focuses on aiding in the functional annotation of novel proteins that have been identified from various ongoing genomics efforts. The NMR functional chemical library is comprised of small molecules with known biological activity such as: co-factors, inhibitors, metabolites and substrates. This functional library was developed through an extensive manual effort of mining several databases based on known ligand interactions with protein systems. In order to increase the efficiency of screening the NMR functional library, the compounds are screened as mixtures of 3-4 compounds that avoids the need to deconvolute positive hits by maintaining a unique NMR resonance and function for each compound in the mixture. The functional library has been used in the identification of general biological function of hypothetical proteins identified from the Protein Structure Initiative.

Binding Sites↗

MultiLoc: prediction of protein subcellular localization using N-terminal targeting sequences, sequence motifs and amino acid composition.

MOTIVATION: Functional annotation of unknown proteins is a major goal in proteomics. A key annotation is the prediction of a protein's subcellular localization. Numerous prediction techniques have been developed, typically focusing on a single underlying biological aspect or predicting a subset of all possible localizations. An important step is taken towards emulating the protein sorting process by capturing and bringing together biologically relevant information, and addressing the clear need to improve prediction accuracy and localization coverage. RESULTS: Here we present a novel SVM-based approach for predicting subcellular localization, which integrates N-terminal targeting sequences, amino acid composition and protein sequence motifs. We show how this approach improves the prediction based on N-terminal targeting sequences, by comparing our method TargetLoc against existing methods. Furthermore, MultiLoc performs considerably better than comparable methods predicting all major eukaryotic subcellular localizations, and shows better or comparable results to methods that are specialized on fewer localizations or for one organism. AVAILABILITY: http://www-bs.informatik.uni-tuebingen.de/Services/MultiLoc/

Algorithms↗

Benchmarking ortholog identification methods using functional genomics data.

BACKGROUND: The transfer of functional annotations from model organism proteins to human proteins is one of the main applications of comparative genomics. Various methods are used to analyze cross-species orthologous relationships according to an operational definition of orthology. Often the definition of orthology is incorrectly interpreted as a prediction of proteins that are functionally equivalent across species, while in fact it only defines the existence of a common ancestor for a gene in different species. However, it has been demonstrated that orthologs often reveal significant functional similarity. Therefore, the quality of the orthology prediction is an important factor in the transfer of functional annotations (and other related information). To identify protein pairs with the highest possible functional similarity, it is important to qualify ortholog identification methods. RESULTS: To measure the similarity in function of proteins from different species we used functional genomics data, such as expression data and protein interaction data. We tested several of the most popular ortholog identification methods. In general, we observed a sensitivity/selectivity trade-off: the functional similarity scores per orthologous pair of sequences become higher when the number of proteins included in the ortholog groups decreases. CONCLUSION: By combining the sensitivity and the selectivity into an overall score, we show that the InParanoid program is the best ortholog identification method in terms of identifying functionally equivalent proteins.

Algorithms↗

High-throughput yeast two-hybrid assays for large-scale protein interaction mapping.

Protein-protein interactions play fundamental roles in many biological processes. Hence, protein interaction mapping is becoming a well-established functional genomics approach to generate functional annotations for predicted proteins that so far have remained uncharacterized. The yeast two-hybrid system is currently one of the most standardized protein interaction mapping techniques. Here, we describe the protocols for a semiautomated, high-throughput, Gal4-based yeast two-hybrid system.

Culture Media↗

Inferring sub-cellular localization through automated lexical analysis.

MOTIVATION: The SWISS-PROT sequence database contains keywords of functional annotations for many proteins. In contrast, information about the sub-cellular localization is available for only a few proteins. Experts can often infer localization from keywords describing protein function. We developed LOCkey, a fully automated method for lexical analysis of SWISS-PROT keywords that assigns sub-cellular localization. With the rapid growth in sequence data, the biochemical characterisation of sequences has been falling behind. Our method may be a useful tool for supplementing functional information already automatically available. RESULTS: The method reached a level of more than 82% accuracy in a full cross-validation test. Due to a lack of functional annotations, we could infer localization for fewer than half of all proteins in SWISS-PROT. We applied LOCkey to annotate five entirely sequenced proteomes, namely Saccharomyces cerevisiae (yeast), Caenorhabditis elegans (worm), Drosophila melanogaster (fly), Arabidopsis thaliana (plant) and a subset of all human proteins. LOCkey found about 8000 new annotations of sub-cellular localization for these eukaryotes.

Abstracting and Indexing↗

Gene discovery in Plasmodium vivax through sequencing of ESTs from mixed blood stages.

Despite the significance of Plasmodium vivax as the most widespread human malaria parasite and a major public health problem, gene expression in this parasite is poorly understood. To accelerate gene discovery and facilitate the annotation phase of the P. vivax genome project, we have undertaken a transcriptome approach to study gene expression in the mixed blood stages of a P. vivax field isolate. Using a cDNA library constructed from purified blood stages, we have obtained single-pass sequences for approximately 21,500 expressed sequence tags (ESTs), the largest number of transcript tags obtained so far for this species. Cluster analysis revealed that the library is highly redundant, resulting in 5407 clusters. Clustered ESTs were searched against public protein databases for functional annotation, and more than one-third showed a significant match, the majority of these to Plasmodium falciparum proteins. The most abundant clusters were to genes encoding ribosomal proteins and proteins involved in metabolism, consistent with the predominance of trophozoites in the field isolate sample. In spite of the scarcity of other parasite stages in the field isolate, we could identify genes that are expressed in rings, schizonts and gametocytes. This study should facilitate our understanding of the gene expression in P. vivax asexual stages and provide valuable data for gene prediction and annotation of the P. vivax genome sequence.

Animals↗

Discovering hidden viral piracy.

MOTIVATION: Viruses and developers of anti-inflammatory therapies share a common interest in proteins that manipulate the immune response. Large double-stranded DNA viruses acquire host proteins to evade host defense mechanisms. Hence, viral pirated proteins may have a therapeutic potential. Although dozens of viral piracy events have already been identified, we hypothesized that sequence divergence impedes the discovery of many others. RESULTS: We developed a method to assess the number of viral/human homologs and discovered that at least 917 highly diverged homologs are hidden in low-similarity alignment hits that are usually ignored. However, these low-similarity homologs are masked by many false alignment hits. We therefore applied a filtering method to increase the proportion of viral/human homologous proteins. The homologous proteins we found may facilitate functional annotation of viral and human proteins. Furthermore, some of these proteins play a key role in immune modulation and are therefore therapeutic protein candidates.

Computational Biology↗

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa↗

Automatic evaluation of protein sequence functional patterns.

A procedure that automatically provides an evaluation of the diagnostic ability of a protein sequence functional pattern is described. The procedure relies on the identification of the closest definable set in terms of a (protein sequence) database functional annotation to the set of database instances containing a given pattern. Assuming annotation correctness and completeness in the protein sequence database, the degree of statistical association between these sets provides an appropriate measure of the diagnostic ability of the pattern. An experimental implementation of the procedure, using the NBRF/PIR protein database, has been applied to a diverse collection of published sequence patterns. Results obtained reveal that frequently it is not possible to define (in NBRF/PIR database terminology) the set of database instances containing a given pattern, suggesting either lack of pattern diagnostic ability or protein database annotation incompleteness and/or inconsistencies.

Algorithms↗

A cross-genomic approach for systematic mapping of phenotypic traits to genes.

We present a computational method for de novo identification of gene function using only cross-organismal distribution of phenotypic traits. Our approach assumes that proteins necessary for a set of phenotypic traits are preferentially conserved among organisms that share those traits. This method combines organism-to-phenotype associations,along with phylogenetic profiles,to identify proteins that have high propensities for the query phenotype; it does not require the use of any functional annotations for any proteins. We first present the statistical foundations of this approach and then apply it to a range of phenotypes to assess how its performance depends on the frequency and specificity of the phenotype. Our analysis shows that statistically significant associations are possible as long as the phenotype is neither extremely rare nor extremely common; results on the flagella,pili, thermophily,and respiratory tract tropism phenotypes suggest that reliable associations can be inferred when the phenotype does not arise from many alternate mechanisms.

Bacterial Proteins↗

The Universal Protein Resource (UniProt).

The ability to store and interconnect all available information on proteins is crucial to modern biological research. Accordingly, the Universal Protein Resource (UniProt) plays an increasingly important role by providing a stable, comprehensive, freely accessible central resource on protein sequences and functional annotation. UniProt is produced by the UniProt Consortium, formed in 2002 by the European Bioinformatics Institute (EBI), the Protein Information Resource (PIR) and the Swiss Institute of Bioinformatics (SIB). The core activities include manual curation of protein sequences assisted by computational analysis, sequence archiving, development of a user-friendly UniProt web site and the provision of additional value-added information through cross-references to other databases. UniProt is comprised of three major components, each optimized for different uses: the UniProt Archive, the UniProt Knowledgebase and the UniProt Reference Clusters. An additional component consisting of metagenomic and environmental sequences has recently been added to UniProt to ensure availability of such sequences in a timely fashion. UniProt is updated and distributed on a bi-weekly basis and can be accessed online for searches or download at http://www.uniprot.org.

Amino Acid Sequence↗

Protein molecular function prediction by Bayesian phylogenomics.

We present a statistical graphical model to infer specific molecular function for unannotated protein sequences using homology. Based on phylogenomic principles, SIFTER (Statistical Inference of Function Through Evolutionary Relationships) accurately predicts molecular function for members of a protein family given a reconciled phylogeny and available function annotations, even when the data are sparse or noisy. Our method produced specific and consistent molecular function predictions across 100 Pfam families in comparison to the Gene Ontology annotation database, BLAST, GOtcha, and Orthostrapper. We performed a more detailed exploration of functional predictions on the adenosine-5'-monophosphate/adenosine deaminase family and the lactate/malate dehydrogenase family, in the former case comparing the predictions against a gold standard set of published functional characterizations. Given function annotations for 3% of the proteins in the deaminase family, SIFTER achieves 96% accuracy in predicting molecular function for experimentally characterized proteins as reported in the literature. The accuracy of SIFTER on this dataset is a significant improvement over other currently available methods such as BLAST (75%), GeneQuiz (64%), GOtcha (89%), and Orthostrapper (11%). We also experimentally characterized the adenosine deaminase from Plasmodium falciparum, confirming SIFTER's prediction. The results illustrate the predictive power of exploiting a statistical model of function evolution in phylogenomic problems. A software implementation of SIFTER is available from the authors.

Adenosine Deaminase↗

NIFAS: visual analysis of domain evolution in proteins.

MOTIVATION: Multi-domain proteins have evolved by insertions or deletions of distinct protein domains. Tracing the history of a certain domain combination can be important for functional annotation of multi-domain proteins, and for understanding the function of individual domains. In order to analyze the evolutionary history of the domains in modular proteins it is desirable to inspect a phylogenetic tree based on sequence divergence with the modular architecture of the sequences superimposed on the tree. RESULT: A Java applet, NIFAS, that integrates graphical domain schematics for each sequence in an evolutionary tree was developed. NIFAS retrieves domain information from the Pfam database and uses CLUSTAL W to calculate a tree for a given Pfam domain. The tree can be displayed with symbolic bootstrap values, and to allow the user to focus on a part of the tree, the layout can be altered by swapping nodes, changing the outgroup, and showing/collapsing subtrees. NIFAS is integrated with the Pfam database and is accessible over the internet (http://www.cgr.ki.se/Pfam). As an example, we use NIFAS to analyze the evolution of domains in Protein Kinases C.

Computer Graphics↗

Functional replacement of the FabA and FabB proteins of Escherichia coli fatty acid synthesis by Enterococcus faecalis FabZ and FabF homologues.

The anaerobic unsaturated fatty acid synthetic pathway of Escherichia coli requires two specialized proteins, FabA and FabB. However, the fabA and fabB genes are found only in the Gram-negative alpha- and gamma-proteobacteria, and thus other anaerobic bacteria must synthesize these acids using different enzymes. We report that the Gram-positive bacterium Enterococcus faecalis encodes a protein, annotated as FabZ1, that functionally replaces the E. coli FabA protein, although the sequence of this protein aligns much more closely with E. coli FabZ, a protein that plays no specific role in unsaturated fatty acid synthesis. Therefore E. faecalis FabZ1 is a bifunctional dehydratase/isomerase, an enzyme activity heretofore confined to a group of Gram-negative bacteria. The FabZ2 protein is unable to replace the function of E. coli FabZ, although FabZ2, a second E. faecalis FabZ homologue, has this ability. Moreover, an E. faecalis FabF homologue (FabF1) was found to replace the function of E. coli FabB, whereas a second FabF homologue was inactive. From these data it is clear that bacterial fatty acid biosynthetic pathways cannot be deduced solely by sequence comparisons.

3-Oxoacyl-(Acyl-Carrier-Protein) Synthase↗

Proteomics-Based Identification of the Pyroptosis-Related Biomarker PCSK9 and Its Association With the Pathogenesis of Rheumatoid Arthritis.

Rheumatoid arthritis (RA) is a common autoimmune disease, and early diagnosis is critical for effective treatment. This study aims to identify potential biomarkers related to pyroptosis through serum proteomics analysis, offering new insights for the early diagnosis of RA. We enrolled 100 participants, including 50 patients with RA and 50 healthy controls. Serum samples were collected and analyzed using high-resolution liquid chromatography-tandem mass spectrometry (LC-MS/MS) for proteomics profiling. Differential protein expression analysis and functional annotation revealed significant upregulation of pyroptosis-related proteins in the serum of patients with RA. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses, along with protein-protein interaction (PPI) network analysis, showed that these proteins are involved in inflammation and immune pathways, particularly the activation of the NOD-like receptor protein 3 (NLRP3) inflammasome. Enzyme-linked immunosorbent assay (ELISA) validation confirmed a significant increase in PCSK9 levels in patients with RA, suggesting that PCSK9 may play a key role in the pathogenesis of RA. This study provides new directions for biomarker research in RA, particularly regarding the potential involvement of the pyroptosis pathway, with significant clinical application prospects.

Humans↗

Using evolutionary and structural information to predict DNA-binding sites on DNA-binding proteins.

Proteins that interact with DNA are involved in a number of fundamental biological activities such as DNA replication, transcription, and repair. A reliable identification of DNA-binding sites in DNA-binding proteins is important for functional annotation, site-directed mutagenesis, and modeling protein-DNA interactions. We apply Support Vector Machine (SVM), a supervised pattern recognition method, to predict DNA-binding sites in DNA-binding proteins using the following features: amino acid sequence, profile of evolutionary conservation of sequence positions, and low-resolution structural information. We use a rigorous statistical approach to study the performance of predictors that utilize different combinations of features and how this performance is affected by structural and sequence properties of proteins. Our results indicate that an SVM predictor based on a properly scaled profile of evolutionary conservation in the form of a position specific scoring matrix (PSSM) significantly outperforms a PSSM-based neural network predictor. The highest accuracy is achieved by SVM predictor that combines the profile of evolutionary conservation with low-resolution structural information. Our results also show that knowledge-based predictors of DNA-binding sites perform significantly better on proteins from mainly-alpha structural class and that the performance of these predictors is significantly correlated with certain structural and sequence properties of proteins. These observations suggest that it may be possible to assign a reliability index to the overall accuracy of the prediction of DNA-binding sites in any given protein using its sequence and structural properties. A web-server implementation of the predictors is freely available online at http://lcg.rit.albany.edu/dp-bind/.

Amino Acid Sequence↗