Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Bioinformatics of large-scale protein interaction networks.

We survey recent techniques for construction and prediction of large-scale protein interaction networks, focusing on computational processing steps. Special emphasis is placed on critical assessment of data completeness and reliability of the various approaches. Once built, protein interaction networks can be used for functional annotation or to generate higher-level biological hypotheses on pathways.

Bacterial Proteins↗

High-resolution functional proteomics by active-site peptide profiling.

Characterization and functional annotation of the large number of proteins predicted from genome sequencing projects poses a major scientific challenge. Whereas several proteomics techniques have been developed to quantify the abundance of proteins, these methods provide little information regarding protein function. Here, we present a gel-free platform that permits ultrasensitive, quantitative, and high-resolution analyses of protein activities in proteomes, including highly problematic samples such as undiluted plasma. We demonstrate the value of this platform for the discovery of both disease-related enzyme activities and specific inhibitors that target these proteins.

Animals↗

Prediction of the coding sequences of unidentified human genes. XVI. The complete sequences of 150 new cDNA clones from brain which code for large proteins in vitro.

We have carried out a human cDNA sequencing project to accumulate information regarding the coding sequences of unidentified human genes. As an extension of the preceding reports, we herein present the entire sequences of 150 cDNA clones of unknown human genes, named KIAA1294 to KIAA1443, from two sets of size-fractionated human adult and fetal brain cDNA libraries. The average sizes of the inserts and corresponding open reading frames of cDNA clones analyzed here reached 4.8 kb and 2.7 kb (910 amino acid residues), respectively. From sequence similarities and protein motifs, 73 predicted gene products were functionally annotated and 97% of them were classified into the following four functional categories: cell signaling/communication, nucleic acid management, cell structure/motility and protein management. Additionally, the chromosomal loci of the genes were assigned by using human-rodent hybrid panels for those genes whose mapping data were not available in the public databases. The expression profiles of the genes were also studied in 10 human tissues, 8 brain regions, spinal cord, fetal brain and fetal liver by reverse transcription-coupled polymerase chain reaction, products of which were quantified by enzyme-linked immunosorbent assay.

Adult↗

A novel representation of protein sequences for prediction of subcellular location using support vector machines.

As the number of complete genomes rapidly increases, accurate methods to automatically predict the subcellular location of proteins are increasingly useful to help their functional annotation. In order to improve the predictive accuracy of the many prediction methods developed to date, a novel representation of protein sequences is proposed. This representation involves local compositions of amino acids and twin amino acids, and local frequencies of distance between successive (basic, hydrophobic, and other) amino acids. For calculating the local features, each sequence is split into three parts: N-terminal, middle, and C-terminal. The N-terminal part is further divided into four regions to consider ambiguity in the length and position of signal sequences. We tested this representation with support vector machines on two data sets extracted from the SWISS-PROT database. Through fivefold cross-validation tests, overall accuracies of more than 87% and 91% were obtained for eukaryotic and prokaryotic proteins, respectively. It is concluded that considering the respective features in the N-terminal, middle, and C-terminal parts is helpful to predict the subcellular location.

Amino Acids, Basic↗

Zinc coordination environments in proteins determine zinc functions.

Estimates of the number of zinc proteins in humans are now possible and a functional annotation of the zinc proteome can begin. The catalytic and structural roles of zinc in hundreds of enzymes and thousands of so-called "zinc finger" protein domains have provided a molecular basis for the numerous biological functions of this essential element. Additional, regulatory functions of zinc/protein interactions are being recognized. They include roles of the zinc ion in signal transduction, in controlling the architecture of protein complexes, and in redox-active zinc sites, where the binding and release of zinc is under redox control. Moreover, a considerable number of proteins participate in cellular zinc homeostasis, e.g. membrane transporters, and cellular storage, sensor, and trafficking proteins. These proteins have evolved with mechanisms to handle zinc ions rather specifically and selectively. They perform their functions with a remarkably modest set: One redox state of the zinc ion and nitrogen, oxygen, and sulfur ligands from the side chains of histidine, glutamate/aspartate, and cysteine, respectively. By permutation of the ligands in this set, the functional potential of the zinc ion has been fully explored. Different coordination environments modulate the chemical characteristics of the zinc ion, control the kinetics of its binding, and allow it to be either metabolically active or inert. Insights into all these functions are building an understanding of why zinc is so critical for such a multitude of life processes.

Gene Expression Regulation↗

Mapping Gene Ontology to proteins based on protein-protein interaction data.

MOTIVATION: Gene Ontology (GO) consortium provides structural description of protein function that is used as a common language for gene annotation in many organisms. Large-scale techniques have generated many valuable protein-protein interaction datasets that are useful for the study of protein function. Combining both GO and protein-protein interaction data allows the prediction of function for unknown proteins. RESULT: We apply a Markov random field method to the prediction of yeast protein function based on multiple protein-protein interaction datasets. We assign function to unknown proteins with a probability representing the confidence of this prediction. The functions are based on three general categories of cellular component, molecular function and biological process defined in GO. The yeast proteins are defined in the Saccharomyces Genome Database (SGD). The protein-protein interaction datasets are obtained from the Munich Information Center for Protein Sequences (MIPS), including physical interactions and genetic interactions. The efficiency of our prediction is measured by applying the leave-one-out validation procedure to a functional path matching scheme, which compares the prediction with the GO description of a protein's function from the abstract level to the detailed level along the GO structure. For biological process, the leave-one-out validation procedure shows 52% precision and recall of our method, much better than that of the simple guilty-by-association methods.

Chromosome Mapping↗

Domain-based small molecule binding site annotation.

BACKGROUND: Accurate small molecule binding site information for a protein can facilitate studies in drug docking, drug discovery and function prediction, but small molecule binding site protein sequence annotation is sparse. The Small Molecule Interaction Database (SMID), a database of protein domain-small molecule interactions, was created using structural data from the Protein Data Bank (PDB). More importantly it provides a means to predict small molecule binding sites on proteins with a known or unknown structure and unlike prior approaches, removes large numbers of false positive hits arising from transitive alignment errors, non-biologically significant small molecules and crystallographic conditions that overpredict ion binding sites. DESCRIPTION: Using a set of co-crystallized protein-small molecule structures as a starting point, SMID interactions were generated by identifying protein domains that bind to small molecules, using NCBI's Reverse Position Specific BLAST (RPS-BLAST) algorithm. SMID records are available for viewing at http://smid.blueprint.org. The SMID-BLAST tool provides accurate transitive annotation of small-molecule binding sites for proteins not found in the PDB. Given a protein sequence, SMID-BLAST identifies domains using RPS-BLAST and then lists potential small molecule ligands based on SMID records, as well as their aligned binding sites. A heuristic ligand score is calculated based on E-value, ligand residue identity and domain entropy to assign a level of confidence to hits found. SMID-BLAST predictions were validated against a set of 793 experimental small molecule interactions from the PDB, of which 472 (60%) of predicted interactions identically matched the experimental small molecule and of these, 344 had greater than 80% of the binding site residues correctly identified. Further, we estimate that 45% of predictions which were not observed in the PDB validation set may be true positives. CONCLUSION: By focusing on protein domain-small molecule interactions, SMID is able to cluster similar interactions and detect subtle binding patterns that would not otherwise be obvious. Using SMID-BLAST, small molecule targets can be predicted for any protein sequence, with the only limitation being that the small molecule must exist in the PDB. Validation results and specific examples within illustrate that SMID-BLAST has a high degree of accuracy in terms of predicting both the small molecule ligand and binding site residue positions for a query protein.

Binding Sites↗

Survey of current protein family databases and their application in comparative, structural and functional genomics.

The last two decades have witnessed significant expansions in the databases storing information on the sequences and structures of proteins. This has led to the creation of many excellent protein family resources, which classify proteins according to their evolutionary relationship. These have allowed extensive insights into evolution and particularly how protein function mutates and evolves over time. Such analyses have greatly assisted the inheritance of functional annotations between experimentally characterised and uncharacterised genes. Moreover, the development of bioinformatics tools acts as a companion to the new technologies emerging in biology, such as transcriptomics and proteomics. The latter enable researchers to analyse gene expression profiles and interactions on a genome-wide scale, generating vast datasets of proteins, many of which include experimentally uncharacterised proteins. Protein family/function databases can be used to help interpret this data and allow us to benefit more fully from these technologies. This review aims to summarise the most popular sequence- and structure-based protein family databases. We also cover their application to comparative genomics and the functional annotation of the genomes.

Biological Evolution↗

Clustering proteins from interaction networks for the prediction of cellular functions.

BACKGROUND: Developing reliable and efficient strategies allowing to infer a function to yet uncharacterized proteins based on interaction networks is of crucial interest in the current context of high-throughput data generation. In this paper, we develop a new algorithm for clustering vertices of a protein-protein interaction network using a density function, providing disjoint classes. RESULTS: Applied to the yeast interaction network, the classes obtained appear to be biological significant. The partitions are then used to make functional predictions for uncharacterized yeast proteins, using an annotation procedure that takes into account the binary interactions between proteins inside the classes. We show that this procedure is able to enhance the performances with respect to previous approaches. Finally, we propose a new annotation for 37 previously uncharacterized yeast proteins. CONCLUSION: We believe that our results represent a significant improvement for the inference of cellular functions, that can be applied to other organism as well as to other type of interaction graph, such as genetic interactions.

Cluster Analysis↗

TransportDB: a comprehensive database resource for cytoplasmic membrane transport systems and outer membrane channels.

TransportDB (http://www.membranetransport.org/) is a comprehensive database resource of information on cytoplasmic membrane transporters and outer membrane channels in organisms whose complete genome sequences are available. The complete set of membrane transport systems and outer membrane channels of each organism are annotated based on a series of experimental and bioinformatic evidence and classified into different types and families according to their mode of transport, bioenergetics, molecular phylogeny and substrate specificities. User-friendly web interfaces are designed for easy access, query and download of the data. Features of the TransportDB website include text-based and BLAST search tools against known transporter and outer membrane channel proteins; comparison of transporter and outer membrane channel contents from different organisms; known 3D structures of transporters, and phylogenetic trees of transporter families. On individual protein pages, users can find detailed functional annotation, supporting bioinformatic evidence, protein/DNA sequences, publications and cross-referenced external online resource links. TransportDB has now been in existence for over 10 years and continues to be regularly updated with new evidence and data from newly sequenced genomes, as well as having new features added periodically.

Bacterial Outer Membrane Proteins↗

A Network Pharmacology and Molecular Docking Study of TongBi Formula for Osteoarthritis.

This study applied network pharmacology combined with molecular docking to predict the potential therapeutic targets and molecular mechanisms of TongBi Formula (TBF) in osteoarthritis (OA). Active components and corresponding targets of TBF were retrieved from the traditional Chinese medicine Systems Pharmacology Database and Analysis Platform, while OA-related targets were collected from Online Mendelian Inheritance in Man, GeneCards, DrugBank, and Therapeutic Target Database. A network visualization and analysis software was used to construct compound-target and protein-protein interaction (PPI) networks. Gene Ontology functional annotation and Kyoto Encyclopedia of Genes and Genomes pathway enrichment analyses were performed using the Database for Annotation, Visualization and Integrated Discovery platform. Molecular docking analysis was conducted using a molecular docking software to evaluate the predicted binding affinity between key active compounds and core target proteins. A total of 47 overlapping targets between TBF and OA were identified. PPI network analysis highlighted JUN, RELA, IL6, MAPK1, and IL10 as potential hub targets. Enrichment analysis suggested that TBF may regulate inflammation, lipid metabolism, and multiple intracellular signaling pathways associated with OA progression. Molecular docking results demonstrated favorable predicted binding affinities between core active compounds and key OA-related protein targets. These findings provide a computational framework for understanding the potential mechanisms of TBF against OA and support further experimental validation.

Molecular Docking Simulation↗

Annotated proteome of a human T-cell lymphoma.

As the reliable identification of proteins by tandem mass spectrometry becomes increasingly common, the full characterization of large data sets of proteins remains a difficult challenge. Our goal was to survey the proteome of a human T-cell lymphoma-derived cell line in a single set of experiments and present an automated method for the annotation of lists of proteins. A downstream application of these data includes the identification of novel pathogenetic and candidate diagnostic markers of T-cell lymphoma. Total protein isolated from cytoplasmic, membrane, and nuclear fractions of the SUDHL-1 T-cell lymphoma cell line was resolved by SDS-PAGE, and the entire gel lanes digested and analyzed by tandem mass spectrometry. Acquired data files were searched against the UniProt protein database using the SEQUEST algorithm. Search results for each subcellular fraction were analyzed using INTERACT and ProteinProphet. All protein identifications with an error rate of less than 10% were directly exported into excel and analyzed using GOMiner (NIH/NCI). The Gene ontology molecular function and cell location data were summarized for the identified proteins and results exported as user-interactive directed acyclic graphs. A total of 1105 unique proteins were identified and fully annotated, including numerous proteins that had not been previously characterized in lymphoma, in functional categories such as cell adhesion, migration, signaling, and stress response. This study demonstrates the utility of currently available bioinformatics tools for the robust identification and annotation of large numbers of proteins in a batchwise fashion.

Algorithms↗

From fold to function predictions: an apoptosis regulator protein BID.

With the rapidly increasing pace of genome sequencing projects and the resulting flood of predicted amino acid sequences of uncharacterized proteins, protein sequence analysis, and in particular, protein structure prediction is quickly gaining in importance. Prediction algorithms can be used for preliminary annotation of newly sequenced proteins and, at least in some cases, provide insights into their function and specific mode of action. Such annotations for several microbial genomes were performed by several groups and placed in public domain for evaluation. An example presented in this work comes from a related project of structural and functional predictions for proteins involved in the process of controlled cell death (apoptosis). The BID protein belongs to an important class of regulators of apoptosis identified by short sequence motifs. Here, several fold prediction methods are used to build a series of three-dimensional models. Structure analysis of the models with reference to the biological data available allows selection of the most appropriate model. It is found that the most likely structural model of BID is built on the structure of Bcl-X(L). The model is discussed in terms of experimental data on specific proteolytic cleavage of BID and its effect on BID interactions with other proteins and membranes.

Algorithms↗

Create and assess protein networks through molecular characteristics of individual proteins.

MOTIVATION: The study of biological systems, pathways and processes relies increasingly on analyses of networks. Most often, such analyses focus on network topology, thereby treating all proteins or genes as identical, featureless nodes. Integrating molecular data and insights about the qualities of individual proteins into the analysis may enhance our ability to decipher biological pathways and processes. RESULTS: Here, we introduce a novel platform for data integration that generates networks on the macro system-level, analyzes the molecular characteristics of each protein on the micro level, and then combines the two levels by using the molecular characteristics to assess networks. It also annotates the function and subcellular localization of each protein and displays the process on an image of a cell, rendering each protein in its respective cellular compartment. By thus visualizing the network in a cellular context we are able to analyze pathways and processes in a novel way. As an example, we use the system to analyze proteins implicated with Alzheimers disease and show how the integrated view corroborates previous observations and how it helps in the formulation of new hypotheses regarding the molecular underpinnings of the disease. AVAILABILITY: http://www.rostlab.org/services/pinat.

Cell Physiological Phenomena↗

ELISA: structure-function inferences based on statistically significant and evolutionarily inspired observations.

UNLABELLED: The problem of functional annotation based on homology modeling is primary to current bioinformatics research. Researchers have noted regularities in sequence, structure and even chromosome organization that allow valid functional cross-annotation. However, these methods provide a lot of false negatives due to limited specificity inherent in the system. We want to create an evolutionarily inspired organization of data that would approach the issue of structure-function correlation from a new, probabilistic perspective. Such organization has possible applications in phylogeny, modeling of functional evolution and structural determination. ELISA (Evolutionary Lineage Inferred from Structural Analysis, http://romi.bu.edu/elisa) is an online database that combines functional annotation with structure and sequence homology modeling to place proteins into sequence-structure-function "neighborhoods". The atomic unit of the database is a set of sequences and structural templates that those sequences encode. A graph that is built from the structural comparison of these templates is called PDUG (protein domain universe graph). We introduce a method of functional inference through a probabilistic calculation done on an arbitrary set of PDUG nodes. Further, all PDUG structures are mapped onto all fully sequenced proteomes allowing an easy interface for evolutionary analysis and research into comparative proteomics. ELISA is the first database with applicability to evolutionary structural genomics explicitly in mind. AVAILABILITY: The database is available at http://romi.bu.edu/elisa.

Amino Acid Sequence↗

Large-scale identification of proteins expressed in mouse embryonic stem cells.

A protein subset expressed in the mouse embryonic stem (ES) cell line, E14-1, was characterized by mass spectrometry-based protein identification technology and data analysis. In total, 1790 proteins including 365 potential nuclear and 260 membrane proteins were identified from tryptic digests of total cell lysates. The subset contained a variety of proteins in terms of physicochemical characteristics, subcellular localization, and biological function as defined by Gene Ontology annotation groups. In addition to many housekeeping proteins found in common with other cell types, the subset contained a group of regulatory proteins that may determine unique ES cell functions. We identified 39 transcription factors including Oct-3/4, Sox-2, and undifferentiated embryonic cell transcription factor I, which are characteristic of ES cells, 88 plasma membrane proteins including cell surface markers such as CD9 and CD81, 44 potential proteinaceous ligands for cell surface receptors including growth factors, cytokines, and hormones, and 100 cell signaling molecules. The subset also contained the products of 60 ES-specific and 41 stemness genes defined previously by the DNA microarray analysis of Ramalho-Santos et al. (Ramalho-Santos et al., Science 2002, 298, 597-600), as well as a number of components characteristic of differentiated cell types such as hematopoietic and neural cells. We also identified potential post-translational modifications in a number of ES cell proteins including five Lys acetylation sites and a single phosphorylation site. To our knowledge, this study provides the largest proteomic dataset characterized to date for a single mammalian cell species, and serves as a basic catalogue of a major proteomic subset that is expressed in mouse ES cells.

Animals↗

ESLpred: SVM-based method for subcellular localization of eukaryotic proteins using dipeptide composition and PSI-BLAST.

Automated prediction of subcellular localization of proteins is an important step in the functional annotation of genomes. The existing subcellular localization prediction methods are based on either amino acid composition or N-terminal characteristics of the proteins. In this paper, support vector machine (SVM) has been used to predict the subcellular location of eukaryotic proteins from their different features such as amino acid composition, dipeptide composition and physico-chemical properties. The SVM module based on dipeptide composition performed better than the SVM modules based on amino acid composition or physico-chemical properties. In addition, PSI-BLAST was also used to search the query sequence against the dataset of proteins (experimentally annotated proteins) to predict its subcellular location. In order to improve the prediction accuracy, we developed a hybrid module using all features of a protein, which consisted of an input vector of 458 dimensions (400 dipeptide compositions, 33 properties, 20 amino acid compositions of the protein and 5 from PSI-BLAST output). Using this hybrid approach, the prediction accuracies of nuclear, cytoplasmic, mitochondrial and extracellular proteins reached 95.3, 85.2, 68.2 and 88.9%, respectively. The overall prediction accuracy of SVM modules based on amino acid composition, physico-chemical properties, dipeptide composition and the hybrid approach was 78.1, 77.8, 82.9 and 88.0%, respectively. The accuracy of all the modules was evaluated using a 5-fold cross-validation technique. Assigning a reliability index (reliability index > or =3), 73.5% of prediction can be made with an accuracy of 96.4%. Based on the above approach, an online web server ESLpred was developed, which is available at http://www.imtech.res.in/raghava/eslpred/.

Artificial Intelligence↗

Visualizing the genome: techniques for presenting human genome data and annotations.

BACKGROUND: In order to take full advantage of the newly available public human genome sequence data and associated annotations, biologists require visualization tools ("genome browsers") that can accommodate the high frequency of alternative splicing in human genes and other complexities. RESULTS: In this article, we describe visualization techniques for presenting human genomic sequence data and annotations in an interactive, graphical format. These techniques include: one-dimensional, semantic zooming to show sequence data alongside gene structures; color-coding exons to indicate frame of translation; adjustable, moveable tiers to permit easier inspection of a genomic scene; and display of protein annotations alongside gene structures to show how alternative splicing impacts protein structure and function. These techniques are illustrated using examples from two genome browser applications: the Neomorphic GeneViewer annotation tool and ProtAnnot, a prototype viewer which shows protein annotations in the context of genomic sequence. CONCLUSION: By presenting techniques for visualizing genomic data, we hope to provide interested software developers with a guide to what features are most likely to meet the needs of biologists as they seek to make sense of the rapidly expanding body of public genomic data and annotations.

Alternative Splicing↗