Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Effects of growth phase and the developmentally significant bldA-specified tRNA on the membrane-associated proteome of Streptomyces coelicolor.

Previous proteomic analyses of Streptomyces coelicolor by two-dimensional electrophoresis and protein mass fingerprinting focused on extracts from total cellular material. Here, the membrane-associated proteome of cultures grown in a liquid minimal medium was partially characterized. The products of some 120 genes were characterized from the membrane fraction, with 70 predicted to possess at least one transmembrane helix. A notably high proportion of ABC transporter systems was represented; the specific types detected provided a snapshot of the nutritional requirements of the mycelium. The membrane-associated proteins did not change very much in abundance in different phases of growth in liquid minimal medium. Identification of gene products not expected to be present in membrane protein extracts led to a reconsideration of the genome annotation in two cases, and supplemented scarce information on 11 hypothetical/conserved hypothetical proteins of unknown function. The wild-type membrane proteome was compared with that of a bldA mutant lacking the only tRNA capable of efficient translation of the rare UUA (leucine) codon. Such mutants are unaffected in vegetative growth but are defective in many aspects of secondary metabolism and morphological differentiation. There were a few clear changes in the membrane proteome of the mutant. In particular, two hypothetical proteins (SCO4244 and SCO4252) were completely absent from the bldA mutant, and this was associated with the TTA-containing regulatory gene SCO4263. Evidence for the control of a cluster of function-unknown genes by the SCO4263 regulator revealed a new aspect of the pleiotropic bldA phenotype.

Base Sequence↗

Functional annotation prediction: all for one and one for all.

In an era of rapid genome sequencing and high-throughput technology, automatic function prediction for a novel sequence is of utter importance in bioinformatics. While automatic annotation methods based on local alignment searches can be simple and straightforward, they suffer from several drawbacks, including relatively low sensitivity and assignment of incorrect annotations that are not associated with the region of similarity. ProtoNet is a hierarchical organization of the protein sequences in the UniProt database. Although the hierarchy is constructed in an unsupervised automatic manner, it has been shown to be coherent with several biological data sources. We extend the ProtoNet system in order to assign functional annotations automatically. By leveraging on the scaffold of the hierarchical classification, the method is able to overcome some frequent annotation pitfalls.

Algorithms↗

Predicting metal-binding site residues in low-resolution structural models.

The accurate prediction of the biochemical function of a protein is becoming increasingly important, given the unprecedented growth of both structural and sequence databanks. Consequently, computational methods are required to analyse such data in an automated manner to ensure genomes are annotated accurately. Protein structure prediction methods, for example, are capable of generating approximate structural models on a genome-wide scale. However, the detection of functionally important regions in such crude models, as well as structural genomics targets, remains an extremely important problem. The method described in the current study, MetSite, represents a fully automatic approach for the detection of metal-binding residue clusters applicable to protein models of moderate quality. The method involves using sequence profile information in combination with approximate structural data. Several neural network classifiers are shown to be able to distinguish metal sites from non-sites with a mean accuracy of 94.5%. The method was demonstrated to identify metal-binding sites correctly in LiveBench targets where no obvious metal-binding sequence motifs were detectable using InterPro. Accurate detection of metal sites was shown to be feasible for low-resolution predicted structures generated using mGenTHREADER where no side-chain information was available. High-scoring predictions were observed for a recently solved hypothetical protein from Haemophilus influenzae, indicating a putative metal-binding site.

Binding Sites↗

Accurate and scalable identification of functional sites by evolutionary tracing.

A common difficulty in post genomics biology is that large-scale techniques of data collection often strip away information on the biological context of these data. The result is a massive number of disconnected observations on sequence, structure, and function from which underlying patterns and biological meaning are obscured. One solution is to build computational filters that pick out sufficiently few facts, relevant to a query, that their relationship is immediately apparent and experimentally testable. Typically, these filters rely on mathematics and statistics, and on first principles from physics and chemistry. We show here that evolution itself can be used to filter sequence and structure data in order to identify evolutionarily important amino acids. A general property of these residues is that they form clusters in native protein structures and point to regions where mutations have the greatest biological impact. The result is an accurate method of functional site annotation that is scalable for structural proteomics.

Amino Acid Sequence↗

ProtoBee: hierarchical classification and annotation of the honey bee proteome.

The recently sequenced genome of the honey bee (Apis mellifera) has produced 10,157 predicted protein sequences, calling for a computational effort to extract biological insights from them. We have applied an unsupervised hierarchical protein-clustering method, which was previously used in the ProtoNet system, to nearly 200,000 proteins consisting of the predicted honey bee proteins, the SWISS-PROT protein database, and the complete set of proteins of the mouse (Mus musculus) and the fruit fly (Drosophila melanogaster). The hierarchy produced by this method has been entitled ProtoBee. In ProtoBee, the proteins are hierarchically organized into 18,936 separate tree hierarchies, each representing a protein functional family. By using the mouse and Drosophila complete proteomes as reference, we are able to highlight functional groups of putative gene-loss events, putative novel proteins of unique functionality, and bee-specific paralogs. We have studied some of the ProtoBee findings and suggest their biological relevance. Examples include novel opsin genes and intriguing nuclear matches of mitochondrial genes. The organization of bee sequences into functional clusters suggests a natural way of automatically inferring functional annotation. Following this notion, we were able to assign functional annotation to about 70% of the sequences. ProtoBee is available at http://www.protobee.cs.huji.ac.il.

Animals↗

CDART: protein homology by domain architecture.

The Conserved Domain Architecture Retrieval Tool (CDART) performs similarity searches of the NCBI Entrez Protein Database based on domain architecture, defined as the sequential order of conserved domains in proteins. The algorithm finds protein similarities across significant evolutionary distances using sensitive protein domain profiles rather than by direct sequence similarity. Proteins similar to a query protein are grouped and scored by architecture. Relying on domain profiles allows CDART to be fast, and, because it relies on annotated functional domains, informative. Domain profiles are derived from several collections of domain definitions that include functional annotation. Searches can be further refined by taxonomy and by selecting domains of interest. CDART is available at http://www.ncbi.nlm.nih.gov/Structure/lexington/lexington.cgi.

BRCA1 Protein↗

Identification of functional modules in protein complexes via hyperclique pattern discovery.

Proteins usually do not act isolated in a cell but function within complicated cellular pathways, interacting with other proteins either in pairs or as components of larger complexes. While many protein complexes have been identified by large-scale experimental studies, due to a large number of false-positive interactions existing in current protein complexes 10, it is still difficult to obtain an accurate understanding of functional modules, which encompass groups of proteins involved in common elementary biological function. In this paper, we present a hyperclique pattern discovery approach for extracting functional modules (hyperclique patterns) from protein complexes. A hyperclique pattern is a type of association pattern containing proteins that are highly affiliated with each other. The analysis of hyperclique patterns shows that proteins within the same pattern tend to present in the protein complex together. Also, statistically significant annotations of proteins in a pattern using the Gene Ontology suggest that proteins within the same hyperclique pattern more likely perform the same function and participate in the same biological process. More interestingly, the 3-D structural view of proteins within a hyperclique pattern reveals that these proteins physically interactwith each other. In addition, we show that several hyperclique patterns corresponding to different functions can participate in the same protein complex as independent modules. Finally, we demonstrate that a hyperclique pattern can be involved in different complexes performing different higher-order biological functions, although the pattern corresponds to a specific elementary biological function.

Algorithms↗

Preferential duplication in the sparse part of yeast protein interaction network.

Gene duplication is an important mechanism driving the evolution of biomolecular network. Thus, it is expected that there should be a strong relationship between a gene's duplicability and the interactions of its protein product with other proteins in the network. We studied this question in the context of the protein interaction network (PIN) of Saccharomyces cerevisiae. We found that duplicates have, on average, significantly lower clustering coefficient (CC) than singletons, and the proportion of duplicates (PD) decreases steadily with CC. Furthermore, using functional annotation data, we observed a strong negative correlation between PD and the mean CC for functional categories. By partitioning the network into modules and assigning each protein a modularity measure Q(n), we found that CC of a protein is a reflection of its modularity. Moreover, the core components of complexes identified in a recent high-throughput experiment, characterized by high CC, have lower PD than that of the attachments. Subsequently, 2 types of hub were identified by their degree, CC and Q(n). Although PD of intramodular hubs is much less than the network average, PD of intermodular hubs is comparable to, or even higher than, the network average. Our results suggest that high CC, and thus high modularity, pose strong evolutionary constraints on gene duplicability, and gene duplication prefers to happen in the sparse part of PINs.

Cluster Analysis↗

Extensive analysis of the human platelet proteome by two-dimensional gel electrophoresis and mass spectrometry.

Platelets play a key role in the control of bleeding and wound healing, contributing to the formation of vascular plugs. Under pathologic circumstances, they are involved in thrombotic disorders, including heart disease. Since platelets do not have a nucleus, proteomics offers a powerful alternative approach to provide data on protein expression in these cells, helping to address their biology. In this publication we extend the previously reported analysis of the pI 4-5 region of the human platelet proteome to the pI 5-11 region. By using narrow pI range two-dimensional electrophoresis (2-DE) for protein separation followed by high-throughput tandem mass spectrometry (MS/MS) for protein identification, we were able to identify 760 protein features, corresponding to 311 different genes, resulting in the annotation of 54% of the pI 5-11 range 2-DE proteome map. We evaluated the physicochemical properties and functions of the identified platelet proteome. Importantly, the main group of proteins identified is involved in intracellular signalling and regulation of the cytoskeleton. In addition, 11 hypothetical proteins are reported. In conclusion, this study provides a unique inventory of the platelet proteome, contributing to our understanding of platelet function and building the basis for the identification of new drug targets.

Blood Platelets↗

BaCelLo: a balanced subcellular localization predictor.

MOTIVATION: The knowledge of the subcellular localization of a protein is fundamental for elucidating its function. It is difficult to determine the subcellular location for eukaryotic cells with experimental high-throughput procedures. Computational procedures are then needed for annotating the subcellular location of proteins in large scale genomic projects. RESULTS: BaCelLo is a predictor for five classes of subcellular localization (secretory pathway, cytoplasm, nucleus, mitochondrion and chloroplast) and it is based on different SVMs organized in a decision tree. The system exploits the information derived from the residue sequence and from the evolutionary information contained in alignment profiles. It analyzes the whole sequence composition and the compositions of both the N- and C-termini. The training set is curated in order to avoid redundancy. For the first time a balancing procedure is introduced in order to mitigate the effect of biased training sets. Three kingdom-specific predictors are implemented: for animals, plants and fungi, respectively. When distributing the proteins from animals and fungi into four classes, accuracy of BaCelLo reach 74% and 76%, respectively; a score of 67% is obtained when proteins from plants are distributed into five classes. BaCelLo outperforms the other presently available methods for the same task and gives more balanced accuracy and coverage values for each class. We also predict the subcellular localization of five whole proteomes, Homo sapiens, Mus musculus, Caenorhabditis elegans, Saccharomyces cerevisiae and Arabidopsis thaliana, comparing the protein content in each different compartment. AVAILABILITY: BaCelLo can be accessed at http://www.biocomp.unibo.it/bacello/.

Algorithms↗

MotD of Sinorhizobium meliloti and related alpha-proteobacteria is the flagellar-hook-length regulator and therefore reassigned as FliK.

The flagella of the soil bacterium Sinorhizobium meliloti differ from the enterobacterial paradigm in the complex filament structure and modulation of the flagellar rotary speed. The mode of motility control in S. meliloti has a molecular corollary in two novel periplasmic motility proteins, MotC and MotE, that are present in addition to the ubiquitous MotA/MotB energizing proton channel. A fifth motility gene is located in the mot operon downstream of the motB and motC genes. Its gene product was originally designated MotD, a cytoplasmic motility protein having an unknown function. We report here reassignment of MotD as FliK, the regulator of flagellar hook length. The FliK gene is one of the few flagellar genes not annotated in the contiguous flagellar regulon of S. meliloti. Characteristic for its class, the 475-residue FliK protein contains a conserved, compactly folded Flg hook domain in its carboxy-terminal region. Deletion of fliK leads to formation of prolonged flagellar hooks (polyhooks) with missing filament structures. Extragenic suppressor mutations all mapped in the cytoplasmic region of the transmembrane export protein FlhB and restored assembly of a flagellar filament, and thus motility, in the presence of polyhooks. The structural properties of FliK are consistent with its function as a substrate specificity switch of the flagellar export apparatus for switching from rod/hook-type substrates to filament-type substrates.

Amino Acid Sequence↗

Fishing for a reelGene: evaluating gene models with evolution and machine learning.

Assembled genomes and their associated annotations have transformed our study of gene function. However, each new annotated assembly generates new gene models. Inconsistencies between annotations likely arise from biological and technical causes, including pseudogene misclassification, transposon activity, and intron retention from sequencing of unspliced transcripts. To evaluate gene model predictions, we developed reelGene, a pipeline of machine learning models focused on (1) transcription boundaries, (2) mRNA integrity, and (3) protein structure. The first two models leverage sequence characteristics and evolutionary conservation across related taxa to learn the grammar of conserved transcription boundaries and mRNA sequences, while the third uses the conserved evolutionary grammar of protein sequences to predict whether a gene can produce a protein. Evaluating 1.8 million transcript models in Zea mays ssp. mays (maize), reelGene classified 28% as incorrectly annotated or non-functional. We find that reelGene classifies 92.2% of genes in the maize proteome and 99.2% of genes within the maize classical gene list as functional. reelGene also provides a way to further investigate genome biology- for instance, reelGene indicates that 10.3% of dispensable genes in B73 are functional, and within retained duplicate genes, reelGene identifies a 30% bias toward the retention of the M1 subgenome when one copy is functional and the other is non-functional. As an annotation-evaluating tool, reelGene is directly applicable to species of the Andropogoneae tribe, including other important crops like sorghum and miscanthus. As a community resource, reelGene has been integrated onto MaizeGDB both as a browser track and as an individual Shiny App, allowing researchers to evaluate gene model accuracy and further investigate genome biology.

Machine Learning↗

Local modeling of global interactome networks.

MOTIVATION: Systems biology requires accurate models of protein complexes, including physical interactions that assemble and regulate these molecular machines. Yeast two-hybrid (Y2H) and affinity-purification/mass-spectrometry (AP-MS) technologies measure different protein-protein relationships, and issues of completeness, sensitivity and specificity fuel debate over which is best for high-throughput 'interactome' data collection. Static graphs currently used to model Y2H and AP-MS data neglect dynamic and spatial aspects of macromolecular complexes and pleiotropic protein function. RESULTS: We apply the local modeling methodology proposed by Scholtens and Gentleman (2004) to two publicly available datasets and demonstrate its uses, interpretation and limitations. Specifically, we use this technology to address four major issues pertaining to protein-protein networks. (1) We motivate the need to move from static global interactome graphs to local protein complex models. (2) We formally show that accurate local interactome models require both Y2H and AP-MS data, even in idealized situations. (3) We briefly discuss experimental design issues and how bait selection affects interpretability of results. (4) We point to the implications of local modeling for systems biology including functional annotation, new complex prediction, pathway interactivity and coordination with gene-expression data. AVAILABILITY: The local modeling algorithm and all protein complex estimates reported here can be found in the R package apComplex, available at http://www.bioconductor.org CONTACT: dscholtens@northwestern.edu SUPPLEMENTARY INFORMATION: http://daisy.prevmed.northwestern.edu/~denise/pubs/LocalModeling

Algorithms↗

GPCR-GRAPA-LIB--a refined library of hidden Markov Models for annotating GPCRs.

GPCR-GRAPA-LIB is a library of HMMs describing G protein coupled receptor families. These families are initially defined by class of receptor ligand, with divergent families divided into subfamilies using phylogenic analysis and knowledge of GPCR function. Protein sequences are applied to the models with the GRAPA curve-based selection criteria. RefSeq sequences for Homo sapiens, Drosophila melanogaster, and Caenorhabditis elegans have been annotated using this approach.

Algorithms↗

Gene recognition based on nucleotide distribution of ORFs in a hyper-thermophilic crenarchaeon, Aeropyrum pernix K1.

The 2694 ORFs originally annotated as potential genes in the genome of Aeropyrum pernix can be categorized into three clusters (A, B, C), according to their nucleotide composition at three codon positions. Coding potential was found to be responsible for the phenomenon of three clusters in a 9-dimensional space derived from the nucleotide composition of ORFs: ORFs assigned to cluster A are coding ones, while those assigned to clusters B and C are non-coding ORFs. A "codingness" index called the AZ score is defined based on a clustering method used to recognize protein-coding genes in the A. pernix genome. The criterion for a coding or non-coding ORF is based on the AZ score. ORFs with AZ > 0 or AZ < 0 are coding or non-coding, respectively. Consequently, 620 out of 632 ORFs with putative functions based on the original annotation are contained in cluster A, which have positive AZ scores. In addition, all 29 ORFs encoding putative or conserved proteins newly added in RefSeq annotation also have positive AZ scores. Accordingly, the number of re-recognized protein-coding genes in the A. pernix genome is 1610, which is significantly less than 2694 in the original annotation and also much less than 1841 in the RefSeq annotation curated by NCBI staff. Annotation information of re-recognized genes and their AZ scores are available at: http://tubic.tju.edu.cn/Aper/.

Aeropyrum↗

Structural characterization of the human proteome.

This paper reports an analysis of the encoded proteins (the proteome) of the genomes of human, fly, worm, yeast, and representatives of bacteria and archaea in terms of the three-dimensional structures of their globular domains together with a general sequence-based study. We show that 39% of the human proteome can be assigned to known structures. We estimate that for 77% of the proteome, there is some functional annotation, but only 26% of the proteome can be assigned to standard sequence motifs that characterize function. Of the human protein sequences, 13% are transmembrane proteins, but only 3% of the residues in the proteome form membrane-spanning regions. There are substantial differences in the composition of globular domains of transmembrane proteins between the proteomes we have analyzed. Commonly occurring structural superfamilies are identified within the proteome. The frequencies of these superfamilies enable us to estimate that 98% of the human proteome evolved by domain duplication, with four of the 10 most duplicated superfamilies specific for multicellular organisms. The zinc-finger superfamily is massively duplicated in human compared to fly and worm, and occurrence of domains in repeats is more common in metazoa than in single cellular organisms. Structural superfamilies over- and underrepresented in human disease genes have been identified. Data and results can be downloaded and analyzed via web-based applications at http://www.sbg.bio.ic.ac.uk.

Algorithms↗

IdentiCS--identification of coding sequence and in silico reconstruction of the metabolic network directly from unannotated low-coverage bacterial genome sequence.

BACKGROUND: A necessary step for a genome level analysis of the cellular metabolism is the in silico reconstruction of the metabolic network from genome sequences. The available methods are mainly based on the annotation of genome sequences including two successive steps, the prediction of coding sequences (CDS) and their function assignment. The annotation process takes time. The available methods often encounter difficulties when dealing with unfinished error-containing genomic sequence. RESULTS: In this work a fast method is proposed to use unannotated genome sequence for predicting CDSs and for an in silico reconstruction of metabolic networks. Instead of using predicted genes or CDSs to query public databases, entries from public DNA or protein databases are used as queries to search a local database of the unannotated genome sequence to predict CDSs. Functions are assigned to the predicted CDSs simultaneously. The well-annotated genome of Salmonella typhimurium LT2 is used as an example to demonstrate the applicability of the method. 97.7% of the CDSs in the original annotation are correctly identified. The use of SWISS-PROT-TrEMBL databases resulted in an identification of 98.9% of CDSs that have EC-numbers in the published annotation. Furthermore, two versions of sequences of the bacterium Klebsiella pneumoniae with different genome coverage (3.9 and 7.9 fold, respectively) are examined. The results suggest that a 3.9-fold coverage of the bacterial genome could be sufficiently used for the in silico reconstruction of the metabolic network. Compared to other gene finding methods such as CRITICA our method is more suitable for exploiting sequences of low genome coverage. Based on the new method, a program called IdentiCS (Identification of Coding Sequences from Unfinished Genome Sequences) is delivered that combines the identification of CDSs with the reconstruction, comparison and visualization of metabolic networks (free to download at http://genome.gbf.de/bioinformatics/index.html). CONCLUSIONS: The reversed querying process and the program IdentiCS allow a fast and adequate prediction protein coding sequences and reconstruction of the potential metabolic network from low coverage genome sequences of bacteria. The new method can accelerate the use of genomic data for studying cellular metabolism.

Base Sequence↗

Functional grouping based on signatures in protein termini.

The two ends of each protein are known as the amino (N-) and carboxyl (C-) termini. Short signatures in a protein's termini often carry vital cellular function. No systematic research has been conducted to address the importance of short signatures (3 to 10 amino acids) in protein termini at the proteomic level. Specifically, it is unknown whether such signatures are evolutionarily conserved, and if so, whether this conservation confers shared biological functions. Current signature detection methods fail to detect such short signatures due to inadequate statistical scores. The findings presented in this study strongly support the notion that functional significance of protein sets may be captured by short signatures at their termini. A positional search method was applied to over one million proteins from the UniProt database. The result is a collection of about a thousand significant signature groups (SIGs) that include previously identified as well as many novel signatures in protein termini. These SIGs represent protein sets with minimal or no overall sequence similarity excepting the similarity at their termini. The most significant SIGs are assigned by their strong correspondence to functional annotations derived from external databases such as Gene Ontology. Each of the SIGs is associated with the statistical significance of its functional association. These SIGs provide a valuable source for testing previously overlooked signatures in protein termini and allow for the investigation of the role played by such signatures throughout evolution. The SIGs archive and advanced search options are available at http://www.proteus.cs.huji.ac.il.

Amino Acid Sequence↗