Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Functional annotation of the putative orphan Caenorhabditis elegans G-protein-coupled receptor C10C6.2 as a FLP15 peptide receptor.

This report describes the cloning and functional annotation of a Caenorhabditis elegans orphan G-protein-coupled receptor (GPCR) (C10C6.2) as a receptor for the FMRFamide-related peptides (FaRPs) encoded on the flp15 precursor gene, leading to the receptor designation FLP15-R. A cDNA encoding C10C6.2 was obtained using PCR techniques, confirmed identical to the Worm-pep-predicted sequence, and cloned into a vector appropriate for eucaryotic expression. A [35S]guanosine 5'-O-(thiotriphosphate) (GTPgammaS) assay with membranes prepared from Chinese hamster ovary (CHO) cells transiently transfected with FLP15-R was used as a read-out for receptor activation. FLP15-R was activated by putative FLP15 peptides, GGPQGPLRF-NH2 (FLP15-1), RGPSGPLRF-NH2 (FLP15-2A), its des-Arg1 counterpart, GPSGPLRF-NH2 (FLP15-2B), and to a lesser extent, by a tobacco hornworm Manduca sexta FaRP, GNSFLRFNH2 (F7G) (potency ranking FLP15-2A > FLP15-1 > FLP15-2B >> F7G). FLP15-R activation was abolished in the transfected cells pretreated with pertussis toxin, suggesting a preferential receptor coupling to Gi/Go proteins. The functional expression of FLP15-R in mammalian cells was temperature-dependent. Either no stimulation or significantly lower ligand-evoked [35S]GTPgammaS binding was observed in membranes prepared from transfected FLP15-R/CHO cells cultured at 37 degrees C. However, a 37 to 28 degrees C temperature shift implemented 24 h post-transfection consistently resulted in an improved activation signal and was essential for detectable functional expression of FLP15-R in CHO cells. To our knowledge, the FLP15 receptor is only the second deorphanized C. elegans neuropeptide GPCR reported to date.

Amino Acid Sequence↗

Cardiovascular-related proteins identified in human plasma by the HUPO Plasma Proteome Project pilot phase.

Proteomic profiling of accessible bodily fluids, such as plasma, has the potential to accelerate biomarker/biosignature development for human diseases. The HUPO Plasma Proteome Project pilot phase examined human plasma with distinct proteomic approaches across multiple laboratories worldwide. Through this effort, we confidently identified 3020 proteins, each requiring a minimum of two high-scoring MS/MS spectra. A critical step subsequent to protein identification is functional annotation, in particular with regard to organ systems and disease. Performing exhaustive literature searches, we have manually annotated a subset of these 3020 proteins that have cardiovascular-related functions on the basis of an existing body of published information. These cardiovascular-related proteins can be organized into eight groups: markers of inflammation and/or cardiovascular disease, vascular and coagulation, signaling, growth and differentiation, cytoskeletal, transcription factors, channels/receptors and heart failure and remodeling. In addition, analysis of the peptide per protein ratio for MS/MS identification reveals group-specific trends. These findings serve as a resource to interrogate the functions of plasma proteins, and moreover, the list of cardiovascular-related proteins in plasma constitutes a baseline proteomic blueprint for the future development of biosignatures for diseases such as myocardial ischemia and atherosclerosis.

Arteriosclerosis↗

Functional annotation of putative aminoglycoside antibiotic modifying proteins in Mycobacterium tuberculosis H37Rv.

The growing availability of sequences of bacterial genomes has revealed a number of open reading frames predicted by sequence alignment to encode antibiotic resistance proteins. The presence of these putative resistance genes within bacterial genomes raises important questions regarding potential reservoirs of resistance elements and their evolution. Here we examine four gene products encoding predicted aminoglycoside-aminocyclitol antibiotic modifying enzymes, two phosphotransferases and two acetyltransferases, derived from analysis of the genome sequence of Mycobacterium tuberculosis strain H37Rv with the goal of assigning biochemical function by purification of each protein and characterization of their ability to modify aminoglycoside antibiotics. Only one of these enzymes, the previously characterized aminoglycoside acetyltransferase AAC(2')-Ic, displayed compelling aminoglycoside modifying activity. While the putative phosphotransferase encoded by the Rv3225c gene did display low levels of aminoglycoside kinase activity, the predicted kinase encoded by the Rv3817 gene lacked any such activity. A potential aminoglycoside 6'-acetyltransferase, encoded by the Rv1347c gene, did not show antibiotic acylation activity but did demonstrate selective thioesterase activity with numerous acyl-CoAs. This activity, together with the genomic environment of the Rv1347c gene in a likely polyketide synthesis cluster, suggests a role for this protein in secondary metabolism and not in antibiotic modification. It was thus shown that only one of four putative aminoglycosides modifying enzymes derived from the whole genome sequencing of M. tuberculosis H37Rv showed sufficient predicted enzyme activity to be annotated as an aminoglycoside resistance element. This study demonstrates the necessity of biochemical annotation methods as a follow up to in silico sequence alignment-based methods of assigning gene product function.

Acetyltransferases↗

Accelerated discovery of novel protein function in cultured human cells.

Experimental approaches that enable direct investigation of human protein function are necessary for comprehensive annotation of the human proteome. We introduce a cell-based platform for rapid and unbiased functional annotation of undercharacterized human proteins. Utilizing a library of antibody biomarkers, the full-length proteins are investigated by tracking phenotypic changes caused by overexpression in human cell lines. We combine reverse transfection and immunodetection by fluorescence microscopy to facilitate this procedure at high resolution. Demonstrating the advantage of this approach, new annotations are provided for two novel proteins: 1) a membrane-bound O-acyltransferase protein (C3F) that, when overexpressed, disrupts Golgi and endosome integrity due likely to an endoplasmic reticulum-Golgi transport block and 2) a tumor marker (BC-2) that prompts a redistribution of a transcriptional silencing protein (BMI1) and a mitogen-activated protein kinase mediator (Rac1) to distinct nuclear regions that undergo chromatin compaction. Our strategy is an immediate application for directly addressing those proteins whose molecular function remains unknown.

Biomarkers↗

Biological function made crystal clear - annotation of hypothetical proteins via structural genomics.

Many of the gene products of completely sequenced organisms are 'hypothetical' - they cannot be related to any previously characterized proteins - and so are of completely unknown function. Structural studies provide one means of obtaining functional information in these cases. A 'structural genomics' project has been initiated aimed at determining the structures of 50 hypothetical proteins from Haemophilus influenzae to gain an understanding of their function. Each stage of the project - target selection, protein production, crystallization, structure determination, and structure analysis - makes use of recent advances to streamline procedures. Early results from this and similar projects are encouraging in that some level of functional understanding can be deduced from experimentally solved structures.

Bacterial Proteins↗

Metaproteomic Analysis to Assess the Impact of Storage Media on Human Gut Microbiome in Fecal Samples.

The human gut microbiome is a diverse community of microorganisms residing in the gastrointestinal tract. The storage condition of fecal samples may impact the taxonomic and protein compositions of microbiomes in these samples. Here, we performed a mass spectrometry-based metaproteomic study to assess the impact of storage media on human gut microbiome in fecal samples. We evaluated FDA-authorized OMNIgene·GUT (OG), phosphate-buffered saline (PBS), and RNALater (RNAL) buffers and identified 38,185 microbial peptides corresponding to 7348 microbial proteins, which matched 16 phyla, 20 classes, 50 orders, 104 families, 332 genera, and 453 species. We found a high similarity among the fecal microbiomes preserved in OG, PBS, and RNAL in terms of the identification of proteins, taxa, and functional annotations. Both alpha and beta diversity suggested the high similarity among samples stored in the three media. Nonetheless, we also found some notable differences among buffers regarding the abundances of a few taxon groups. A partial human proteome (over 400 proteins) was identified in the fecal samples, with most of these proteins associated with the membrane and extracellular regions. The findings indicate the similarity among microbiomes in the fecal samples stored in OG, PBS, and RNAL regarding proteome profile, taxa, and functional capacity. SUMMARY: This study thoroughly analyzed and compared the metaproteomes of fecal samples preserved at -80°C in PBS, RNALater, and OMNIgene·GUT Dx buffers, offering novel insights into the effectiveness of these buffers in maintaining the stability and composition of the human gut microbiome. We found a high similarity in the identification and quantification of proteins, taxa, and functional annotations across the three buffers, with notable quantitative differences highlighting subtle yet important variations in preservation efficacy. The unique datasets and findings could offer valuable revelations into the impact of fecal sample preservation on translational and clinical analyses of the human gut microbiome.

Humans↗

EST sequencing and time course microarray hybridizations identify more than 700 Medicago truncatula genes with developmental expression regulation in flowers and pods.

To evaluate the molecular mechanisms during pod and seed formation in legumes, starting with the development of reproductive organs, we constructed two cDNA libraries from developing flowers (MtFLOW) and pods including seeds (MtPOSE) of the model plant Medicago truncatula Gaertner. A total of 2,516 expressed sequence tags (ESTs) clustered into 1,776 nonredundant sequences (2k-set), which were annotated and assigned to functional classes. While about 30% of the ESTs encoded proteins of yet unknown function, typical annotations pointed to seed storage proteins, LTPs and lipoxygenases. The 2k-set was used to upgrade Mt6k-RIT microarrays (Küster et al. in J Biotechnol 108: 95, 2004) to Mt8k versions representing approximately 6,300 nonredundant M. truncatula genes. These were used to perform time course expression profiling studies based on hybridizations of samples that covered eight different developmental stages from flower buds to almost mature pods versus leaves as a common reference. About 180 up- and 70 downregulated genes were typically found for each stage and in total, 782 genes were either twofold up- or downregulated in at least one of the eight stages investigated. Based on this set, a combination of self-organizing map and hierarchical clustering revealed genes displaying expression regulation during characteristic stages of M. truncatula flower and pod development. Amongst those, several genes encoded proteins related to seed metabolism and development including novel regulators and proteins involved in signaling.

Expressed Sequence Tags↗

Phydbac "Gene Function Predictor": a gene annotation tool based on genomic context analysis.

BACKGROUND: The large amount of completely sequenced genomes allows genomic context analysis to predict reliable functional associations between prokaryotic proteins. Major methods rely on the fact that genes encoding physically interacting partners or members of shared metabolic pathways tend to be proximate on the genome, to evolve in a correlated manner and to be fused as a single sequence in another organism. RESULTS: The new "Gene Function Predictor", linked to the web server Phydbac proposes putative associations between Escherichia coli K-12 proteins derived from a combination of these methods. We show that associations made by this tool are more accurate than linkages found in the other established databases. Predicted assignments to GO categories, based on pre-existing functional annotations of associated proteins are also available. This new database currently holds 9,379 pairwise links at an expected success rate of at least 80%, the 6,466 functional predictions to GO terms derived from these links having a level of accuracy higher than 70%. CONCLUSION: The "Gene Function Predictor" is an automatic tool that aims to help biologists by providing them hypothetical functional predictions out of genomic context characteristics. The "Gene Function predictor" is available at http://www.igs.cnrs-mrs.fr/phydbac/indexPS.html.

Algorithms↗

Dividing the large glycoside hydrolase family 13 into subfamilies: towards improved functional annotations of alpha-amylase-related proteins.

Family GH13, also known as the alpha-amylase family, is the largest sequence-based family of glycoside hydrolases and groups together a number of different enzyme activities and substrate specificities acting on alpha-glycosidic bonds. This polyspecificity results in the fact that the simple membership of this family cannot be used for the prediction of gene function based on sequence alone. In order to establish robust groups that show an improved correlation between sequence and enzymatic specificity, we have performed a large-scale analysis of 1691 family GH13 sequences by combining clustering, similarity search and phylogenetic methods. About 80% of the sequences could be reliably classified into 35 subfamilies. Most subfamilies appear monofunctional (i.e. contain enzymes with the same substrate and the same product). The close examination of the other, apparently polyspecific, subfamilies revealed that they actually group together enzymes with strongly related (or even sometimes virtually identical) activities. Overall our subfamily assignment allows to set the limits for genomic function prediction on this large family of biologically and industrially important enzymes.

Amino Acid Sequence↗

Functional annotation of two orphan G-protein-coupled receptors, Drostar1 and -2, from Drosophila melanogaster and their ligands by reverse pharmacology.

By combining a Drosophila genome data base search and reverse transcriptase-PCR-based cDNA isolation, two G-protein-coupled receptors were cloned, which are the closest known invertebrate homologs of the mammalian opioid/somatostatin receptors. However, when functionally expressed in Xenopus oocytes by injection of Drosophila orphan receptor RNAs together with a coexpressed potassium channel, neither receptor was activated by known mammalian agonists. By applying a reverse pharmacological approach, the physiological ligands were isolated from peptide extracts from adult flies and larvae. Edman sequencing and mass spectrometry of the purified ligands revealed two decapentapeptides, which differ only by an N-terminal pyroglutamate/glutamine. The peptides align to a hormone precursor sequence of the Drosophila genome data base and are almost identical to allatostatin C from Manduca sexta. Both receptors were activated by the synthetic peptides irrespective of the N-terminal modification. Site-directed mutagenesis of a residue in transmembrane region 3 and the loop between transmembrane regions 6 and 7 affect ligand binding, as previously described for somatostatin receptors. The two receptor genes each containing three exons and transcribed in opposite directions are separated by 80 kb with no other genes predicted between. Localization of receptor transcripts identifies a role of the new transmitter system in visual information processing as well as endocrine regulation.

Amino Acid Sequence↗

The Arabidopsis thaliana chloroplast proteome reveals pathway abundance and novel protein functions.

BACKGROUND: Chloroplasts are plant cell organelles of cyanobacterial origin. They perform essential metabolic and biosynthetic functions of global significance, including photosynthesis and amino acid biosynthesis. Most of the proteins that constitute the functional chloroplast are encoded in the nuclear genome and imported into the chloroplast after translation in the cytosol. Since protein targeting is difficult to predict, many nuclear-encoded plastid proteins are still to be discovered. RESULTS: By tandem mass spectrometry, we identified 690 different proteins from purified Arabidopsis chloroplasts. Most proteins could be assigned to known protein complexes and metabolic pathways, but more than 30% of the proteins have unknown functions, and many are not predicted to localize to the chloroplast. Novel structure and function prediction methods provided more informative annotations for proteins of unknown functions. While near-complete protein coverage was accomplished for key chloroplast pathways such as carbon fixation and photosynthesis, fewer proteins were identified from pathways that are downregulated in the light. Parallel RNA profiling revealed a pathway-dependent correlation between transcript and relative protein abundance, suggesting gene regulation at different levels. CONCLUSIONS: The chloroplast proteome contains many proteins that are of unknown function and not predicted to localize to the chloroplast. Expression of nuclear-encoded chloroplast genes is regulated at multiple levels in a pathway-dependent context. The combined shotgun proteomics and RNA profiling approach is of high potential value to predict metabolic pathway prevalence and to define regulatory levels of gene expression on a pathway scale.

Arabidopsis↗

Inference of protein function from protein structure.

Structural genomics has brought us three-dimensional structures of proteins with unknown functions. To shed light on such structures, we have developed ProKnow (http://www.doe-mbi.ucla.edu/Services/ProKnow/), which annotates proteins with Gene Ontology functional terms. The method extracts features from the protein such as 3D fold, sequence, motif, and functional linkages and relates them to function via the ProKnow knowledgebase of features, which links features to annotated functions via annotation profiles. Bayes' theorem is used to compute weights of the functions assigned, using likelihoods based on the extracted features. The description level of the assigned function is quantified by the ontology depth (from 1 = general to 9 = specific). Jackknife tests show approximately 89% correct assignments at ontology depth 1 and 40% at depth 9, with 93% coverage of 1507 distinct folded proteins. Overall, about 70% of the assignments were inferred correctly. This level of performance suggests that ProKnow is a useful resource in functional assessments of novel proteins.

Amino Acid Motifs↗

Whole-proteome prediction of protein function via graph-theoretic analysis of interaction maps.

MOTIVATION: Determining protein function is one of the most important problems in the post-genomic era. For the typical proteome, there are no functional annotations for one-third or more of its proteins. Recent high-throughput experiments have determined proteome-scale protein physical interaction maps for several organisms. These physical interactions are complemented by an abundance of data about other types of functional relationships between proteins, including genetic interactions, knowledge about co-expression and shared evolutionary history. Taken together, these pairwise linkages can be used to build whole-proteome protein interaction maps. RESULTS: We develop a network-flow based algorithm, FunctionalFlow, that exploits the underlying structure of protein interaction maps in order to predict protein function. In cross-validation testing on the yeast proteome, we show that FunctionalFlow has improved performance over previous methods in predicting the function of proteins with few (or no) annotated protein neighbors. By comparing several methods that use protein interaction maps to predict protein function, we demonstrate that FunctionalFlow performs well because it takes advantage of both network topology and some measure of locality. Finally, we show that performance can be improved substantially as we consider multiple data sources and use them to create weighted interaction networks. AVAILABILITY: http://compbio.cs.princeton.edu/function

Algorithms↗

A proteomic analysis of salivary glands of female Anopheles gambiae mosquito.

Understanding the development of the malaria parasite within the mosquito vector at the molecular level should provide novel targets for interrupting parasitic life cycle and subsequent transmission. Availability of the complete genomic sequence of the major African malaria vector, Anopheles gambiae, allows discovery of such targets through experimental as well as computational methods. In the female mosquito, the salivary gland tissue plays an important role in the maturation of the infective form of the malaria parasite. Therefore, we carried out a proteomic analysis of salivary glands from female An. gambiae mosquitoes. Salivary gland extracts were digested with trypsin using two complementary approaches and analyzed by LC-MS/MS. This led to identification of 69 unique proteins, 57 of which were novel. We carried out a functional annotation of all proteins identified in this study through a detailed bioinformatics analysis. Even though a number of cDNA and Edman degradation-based approaches to catalog transcripts and proteins from salivary glands of mosquitoes have been published previously, this is the first report describing the application of MS for characterization of the salivary gland proteome. Our approach should prove valuable for characterizing proteomes of parasites and vectors with sequenced genomes as well as those whose genomes are yet to be fully sequenced.

Animals↗

Enhanced functional information from predicted protein networks.

Experimentally derived genome-wide protein interaction networks have been useful in the elucidation of functional information that is not evident from examining individual proteins but determination of these networks is complex and time consuming. To address this problem, several computational methods for predicting protein networks in novel genomes have been developed. A recent publication by Date and Marcotte describes the use of phylogenetic profiling for elucidating novel pathways in proteomes that have not been experimentally characterized. This method, in combination with other computational methods for generating protein-interaction networks, might help identify novel functional pathways and enhance functional annotation of individual proteins.

Energy Metabolism↗

TOPOFIT-DB, a database of protein structural alignments based on the TOPOFIT method.

TOPOFIT-DB (T-DB) is a public web-based database of protein structural alignments based on the TOPOFIT method, providing a comprehensive resource for comparative analysis of protein structure families. The TOPOFIT method is based on the discovery of a saturation point on the alignment curve (topomax point) which presents an ability to objectively identify a border between common and variable parts in a protein structural family, providing additional insight into protein comparison and functional annotation. TOPOFIT also effectively detects non-sequential relations between protein structures. T-DB provides users with the convenient ability to retrieve and analyze structural neighbors for a protein; do one-to-all calculation of a user provided structure against the entire current PDB release with T-Server, and pair-wise comparison using the TOPOFIT method through the T-Pair web page. All outputs are reported in various web-based tables and graphics, with automated viewing of the structure-sequence alignments in the Friend software package for complete, detailed analysis. T-DB presents researchers with the opportunity for comprehensive studies of the variability in proteins and is publicly available at http://mozart.bio.neu.edu/topofit/index.php.

Databases, Protein↗

TIGRFAMs and Genome Properties: tools for the assignment of molecular function and biological process in prokaryotic genomes.

TIGRFAMs is a collection of protein family definitions built to aid in high-throughput annotation of specific protein functions. Each family is based on a hidden Markov model (HMM), where both cutoff scores and membership in the seed alignment are chosen so that the HMMs can classify numerous proteins according to their specific molecular functions. Most TIGRFAMs models describe 'equivalog' families, where both orthology and lateral gene transfer may be part of the evolutionary history, but where a single molecular function has been conserved. The Genome Properties system contains a queriable set of metabolic reconstructions, genome metrics and extractions of information from the scientific literature. Its genome-by-genome assertions of whether or not specific structures, pathways or systems are present provide high-level conceptual descriptions of genomic content. These assertions enable comparative genomics, provide a meaningful biological context to aid in manual annotation, support assignments of Gene Ontology (GO) biological process terms and help validate HMM-based predictions of protein function. The Genome Properties system is particularly useful as a generator of phylogenetic profiles, through which new protein family functions may be discovered. The TIGRFAMs and Genome Properties systems can be accessed at http://www.tigr.org/TIGRFAMs and http://www.tigr.org/Genome_Properties.

Archaeal Proteins↗

EUCLID: automatic classification of proteins in functional classes by their database annotations.

UNLABELLED: A tool is described for the automatic classification of sequences in functional classes using their database annotations. The Euclid system is based on a simple learning procedure from examples provided by human experts. AVAILABILITY: Euclid is freely available for academics at http://www.gredos.cnb.uam.es/EUCLID, with the corresponding dictionaries for the generation of three, eight and 14 functional classes. CONTACT: E-mail: valencia@cnb.uam.es SUPPLEMENTARY INFORMATION: The results of the EUCLID classification of different genomes are available at http://www.sander.ebi.ac. uk/genequiz/. A detailed description of the different applications mentioned in the text is available at http://www.gredos.cnb.uam. es/EUCLID/Full_Paper

Computational Biology↗