Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Nanodroplet profiling of enzymatic activities in a microarray.

We describe a generic method for the large-scale functional characterization of enzymes in a microarray. Poly-l-lysine and amine reactive slides were coated with fluorogenic substrates sensitive to proteases and phosphatases. Patterning enzymes on the slides by robotic printing produced spatially addressable, segregated droplets that were simultaneously exposed to the on-chip sensors. Multiple enzymes were profiled using this system that provided fluorescence readouts across temporal and stoichiometric dimensions concurrently on a single microarray substrate. This integrated microarray platform is applicable not only for the functional annotation of proteins, but also for the rapid agonist and antagonist discovery and in performing on-chip kinetics.

Microarray Analysis↗

Exploring biosynthetic potential of the endophytic Penicillium turbatum BLH34 using whole-genome sequence analysis and molecular networking.

An in-depth genomic and metabolomic investigation was conducted on the endophytic fungus Penicillium turbatum BLH34, isolated from Macleaya cordata. Hybrid sequencing (Illumina-Nanopore) generated a high-quality 27.9 Mb genome (GC 48.6%) encoding 9798 proteins, with functional annotation linking 5350 genes to the NCBI non-redundant database and 3404 to KEGG pathways. AntiSMASH analysis uncovered 35 biosynthetic gene clusters (BGCs), 23 of which lacked homology to known pathways, highlighting BLH34's potential for novel metabolite discovery. Molecular networking (GNPS) and LC-MS/MS identified 19 specialised metabolites, including antimicrobial polyketides. Bioassays demonstrated potent inhibition against Staphylococcus aureus (36 mm), Bacillus subtilis (28 mm) and Escherichia coli (24 mm), underscoring its pharmaceutical relevance.

Penicillium↗

A tool to assist the study of specific features at protein binding sites.

The Protein Data Bank contains a large amount of proteins that have been solved with small ligands bound to them. This constitutes a rich source of information for the study of the specific requirements of protein sites to bind small molecules with a favorable free energy. The specific atomic composition and three-dimensional geometric restraints of protein binding sites for different ligands could be easily obtained from there. The development of accurate binding site descriptors in proteins constitutes a valuable tool to assist in the large-scale prediction and annotation of protein function in whole genomes. In this work, an integrated database containing some processed and calculated protein/ligand information is described. It is expected that this database will constitute a useful tool for people working in the prediction of protein function from its structure. The database is accessible from the Internet through a web server located at: http://protein.bio.puc.cl

Binding Sites↗

Mutation accumulation in a hybrid parthenogenetic vertebrate.

Asexual lineages are thought to experience elevated extinction rates compared with sexual species, yet direct evidence for the underlying genetic causes remains scarce. Muller's ratchet predicts that the absence of recombination in asexual organisms facilitates the accumulation of deleterious mutations, thereby reducing long-term fitness. Here, we test this hypothesis in the hybrid-origin, parthenogenetic whiptail lizard Aspidoscelis tesselatus by integrating short-read RNAseq and long-read IsoSeq data from both the asexual lineage and its parental sexual species. We reconstructed phased transcripts for A. tesselatus to quantify mutation accumulation relative to the parental sexual species. Comparative analyses revealed elevated ω ratios in both parental genomic complements (subgenomes) of the parthenogenetic lineage, consistent with accelerated accumulation of nonsynonymous mutations. Structural variant analyses identified multiple indels in expressed transcripts predicted to disrupt protein domains. Functional annotation indicated that genes affected by both single-nucleotide variants and indels were enriched for roles in chromatin organization, apoptosis regulation, and transcriptional control. While both parental subgenomes showed similar evolutionary patterns, the maternal complement exhibited more structural and missense mutations than the paternal complement. Together, these results provide evidence that mutations accumulate in asexual A. tesselatus in genes involved in core cellular functions, supporting theoretical predictions that Muller's ratchet contributes to mutation accumulation in asexual lineages.

Animals↗

ArchDB: automated protein loop classification as a tool for structural genomics.

The annotation of protein function has become a crucial problem with the advent of sequence and structural genomics initiatives. A large body of evidence suggests that protein structural information is frequently encoded in local sequences, and that folds are mainly made up of a number of simple local units of super-secondary structural motifs, consisting of a few secondary structures and their connecting loops. Moreover, protein loops play an important role in protein function. Here we present ArchDB, a classification database of structural motifs, consisting of one loop plus its bracing secondary structures. ArchDB currently contains 12,665 super-secondary elements classified into 1496 motif subclasses. The database provides an easy way to retrieve functional information from protein structures sharing a common motif, to search motifs found in a given SCOP family, superfamily or fold, or to search by keywords on proteins with classified loops. The ArchDB database of loops is located at http://sbi.imim.es/archdb.

Amino Acid Motifs↗

Deciphering the genetic background of an industrial 2-ketogluconic acid-producing strain Pseudomonas plecoglossicida JUIM01 using whole-genome sequencing.

2-Ketogluconic acid (2KGA) is an important precursor for the food antioxidant erythorbic acid, currently produced via microbial fermentation using Pseudomonas species. To facilitate the genetic improvement of production strains, the complete genome of an industrial 2KGA producer P. plecoglossicida JUIM01 was sequenced and analyzed. The genome consists of a 5.13-Mb circular chromosome with a GC content of 63.58%, encoding 4,517 predicted proteins. Comprehensive functional annotation identified a putative global regulatory network comprising 75 core regulators, which were classified into six functionally cooperative modules, potentially governing the strain's metabolism and environmental adaptability. We further delineated the genetic determinants hypothetically linked to efficient 2KGA synthesis, including glucose metabolism, fatty acid metabolism, and the oxidative phosphorylation system. These outputs could provide the genomic resource for elucidating high productivity and robustness, and rationally engineering the high-performance chassis cells toward robust 2KGA production.

P. plecoglossicida↗

Who tangos with GOA?-Use of Gene Ontology Annotation (GOA) for biological interpretation of '-omics' data and for validation of automatic annotation tools.

The number of large-scale experimental datasets generated from high-throughput technologies has grown rapidly. Biological knowledge resources such as the Gene Ontology Annotation (GOA) database, which provides high-quality functional annotation to proteins within the UniProt Knowledgebase, can play an important role in the analysis of such data. The integration of GOA with analytical tools has proved to aid the clustering, annotation and biological interpretation of such large expression datasets. GOA is also useful in the development and validation of automated annotation tools, in particular text-mining systems. The increasing interest in GOA highlights the great potential of this freely available resource to assist both the biological research and bioinformatics communities.

Animals↗

High-throughput functional affinity purification of mannose binding proteins from Oryza sativa.

We have used affinity chromatography in combination with mass spectrometry to isolate, identify, and assign a preliminary functional annotation to a large number of both known and novel proteins from rice. Rice (Oryza sativa) leaf, root, and seed tissue extracts were fractionated by column affinity chromatography using alpha-D-mannose as the ligand. Bound fractions were eluted and subjected to one-dimensional electrophoresis, followed by high-performance liquid chromatography-tandem mass spectrometric analysis of separated proteins. This multiplexed technology resulted in the isolation and identification of 136 distinct mannose binding proteins from rice. A comparative analysis demonstrates very little overlap of identified proteins between the respective tissues, and confirms the correctly compartmentalized presence of a significant number of proteins from largely tissue-specific biochemical pathways. Over 30% of the identified proteins with a previously annotated function are directly involved in sugar metabolism, including several highly expressed known rice lectins. Direct comparison of the peptide sequences identified in this study to those peptides identified in the most comprehensive survey of the rice proteome to date indicates that our current data represents a significant enrichment of proteins unique to this dataset. Nearly 15% of the identified proteins, identified on the basis of exact peptide matching to sequences in the rice genomic database, represent proteins without a previously known functional annotation, indicating the potential of this combined chromatographic approach to assign a preliminary function to novel proteins in a high-throughput fashion.

Binding, Competitive↗

Protein sequence analysis in silico: application of structure-based bioinformatics to genomic initiatives.

The current pace of high-throughput genome sequencing programs coupled with high-throughput functional genomic screens has provided researchers with a bewildering array of sequence and biological data to contend with. Identification of proteins of interest from a particular biological study requires the application of bioinformatic tools to process and prioritise the data. From a protein function standpoint, transfer of annotation from known proteins to a novel target is currently the only practical way to convert vast quantities of raw sequence data into meaningful information. New bioinformatics tools now provide more sophisticated methods to transfer functional annotation, integrating sequence, family profile and structural search methodology. The importance of these approaches to medical research is increasing as we move to annotate the proteome through functional and structural genomic efforts.

Animals↗

A two-dimensional electrophoresis proteomic reference map and systematic identification of 1367 proteins from a cell suspension culture of the model legume Medicago truncatula.

The proteome of a Medicago truncatula cell suspension culture was analyzed using two-dimensional electrophoresis and nanoscale HPLC coupled to a tandem Q-TOF mass spectrometer (QSTAR Pulsar i) to yield an extensive protein reference map. Coomassie Brilliant Blue R-250 was used to visualize more than 1661 proteins, which were excised, subjected to in-gel trypsin digestion, and analyzed using nanoscale HPLC/MS/MS. The resulting spectral data were queried against a custom legume protein database using the MASCOT search engine. A total of 1367 of the 1661 proteins were identified with high rigor, yielding an identification success rate of 83% and 907 unique protein accession numbers. Functional annotation of the M. truncatula suspension cell proteins revealed a complete tricarboxylic acid cycle, a nearly complete glycolytic pathway, a significant portion of the ubiquitin pathway with the associated proteolytic and regulatory complexes, and many enzymes involved in secondary metabolism such as flavonoid/isoflavonoid, chalcone, and lignin biosynthesis. Proteins were also identified from most other functional classes including primary metabolism, energy production, disease/defense, protein destination/storage, protein synthesis, transcription, cell growth/division, and signal transduction. This work represents the most extensive proteomic description of M. truncatula suspension cells to date and provides a reference map for future comparative proteomic and functional genomic studies of the response of these cells to biotic and abiotic stress.

Amino Acid Sequence↗

Correlation between gene expression profiles and protein-protein interactions within and across genomes.

MOTIVATION: Function annotation of an unclassified protein on the basis of its interaction partners is well documented in the literature. Reliable predictions of interactions from other data sources such as gene expression measurements would provide a useful route to function annotation. We investigate the global relationship of protein-protein interactions with gene expression. This relationship is studied in four evolutionarily diverse species, for which substantial information regarding their interactions and expression is available: human, mouse, yeast and Escherichia coli. RESULTS: In E.coli the expression of interacting pairs is highly correlated in comparison to random pairs, while in the other three species, the correlation of expression of interacting pairs is only slightly stronger than that of random pairs. To strengthen the correlation, we developed a protocol to integrate ortholog information into the interaction and expression datasets. In all four genomes, the likelihood of predicting protein interactions from highly correlated expression data is increased using our protocol. In yeast, for example, the likelihood of predicting a true interaction, when the correlation is > 0.9, increases from 1.4 to 9.4. The improvement demonstrates that protein interactions are reflected in gene expression and the correlation between the two is strengthened by evolution information. The results establish that co-expression of interacting protein pairs is more conserved than that of random ones.

Animals↗

Integrative bioinformatics for functional genome annotation: trawling for G protein-coupled receptors.

G protein-coupled receptors (GPCR) are amongst the best studied and most functionally diverse types of cell-surface protein. The importance of GPCRs as mediates or cell function and organismal developmental underlies their involvement in key physiological roles and their prominence as targets for pharmacological therapeutics. In this review, we highlight the requirement for integrated protocols which underline the different perspectives offered by different sequence analysis methods. BLAST and FastA offer broad brush strokes. Motif-based search methods add the fine detail. Structural modelling offers another perspective which allows us to elucidate the physicochemical properties that underlie ligand binding. Together, these different views provide a more informative and a more detailed picture of GPCR structure and function. Many GPCRs remain orphan receptors with no identified ligand, yet as computer-driven functional genomics starts to elaborate their functions, a new understanding of their roles in cell and developmental biology will follow.

Algorithms↗

Identification and characterization of protein subcomplexes in yeast.

Protein complexes are major components of cellular organization. Based on large-scale protein complex data, we present the first statistical procedure to find insightful substructures in protein complexes: we identify protein subcomplexes (SCs), i.e., multiprotein assemblies residing in different protein complexes. Four protein complex datasets with different origins and variable reliability are separately analyzed. Our method identifies well-characterized protein assemblies with known functions, thereby confirming the utility of the procedure. In addition, we also identify hitherto unknown functional entities consisting of either functionally unknown proteins or proteins with different functional annotation. We show that SCs represent more reliable protein assemblies than the original complexes. Finally, we demonstrate unique properties of subcomplex proteins that underline the distinct roles of SCs: (i) SCs are functionally and spatially more homogeneous than complete protein complexes (this fact is utilized to predict functional roles and subcellular localizations for so far unannotated proteins); (ii) the abundance of subcomplex proteins is less variable than the abundance of other proteins; (iii) SCs are enriched with essential and synthetic lethal proteins; and (iv) mutations in SC-proteins have higher fitness effects than mutations in other proteins.

Gene Deletion↗

In silico discovery of enzyme-substrate specificity-determining residue clusters.

The binding between an enzyme and its substrate is highly specific, despite the fact that many different enzymes show significant sequence and structure similarity. There must be, then, substrate specificity-determining residues that enable different enzymes to recognize their unique substrates. We reason that a coordinated, not independent, action of both conserved and non-conserved residues determine enzymatic activity and specificity. Here, we present a surface patch ranking (SPR) method for in silico discovery of substrate specificity-determining residue clusters by exploring both sequence conservation and correlated mutations. As case studies we apply SPR to several highly homologous enzymatic protein pairs, such as guanylyl versus adenylyl cyclases, lactate versus malate dehydrogenases, and trypsin versus chymotrypsin. Without using experimental data, we predict several single and multi-residue clusters that are consistent with previous mutagenesis experimental results. Most single-residue clusters are directly involved in enzyme-substrate interactions, whereas multi-residue clusters are vital for domain-domain and regulator-enzyme interactions, indicating their complementary role in specificity determination. These results demonstrate that SPR may help the selection of target residues for mutagenesis experiments and, thus, focus rational drug design, protein engineering, and functional annotation to the relevant regions of a protein.

Adenylyl Cyclases↗

An accurate, sensitive, and scalable method to identify functional sites in protein structures.

Functional sites determine the activity and interactions of proteins and as such constitute the targets of most drugs. However, the exponential growth of sequence and structure data far exceeds the ability of experimental techniques to identify their locations and key amino acids. To fill this gap we developed a computational Evolutionary Trace method that ranks the evolutionary importance of amino acids in protein sequences. Studies show that the best-ranked residues form fewer and larger structural clusters than expected by chance and overlap with functional sites, but until now the significance of this overlap has remained qualitative. Here, we use 86 diverse protein structures, including 20 determined by the structural genomics initiative, to show that this overlap is a recurrent and statistically significant feature. An automated ET correctly identifies seven of ten functional sites by the least favorable statistical measure, and nine of ten by the most favorable one. These results quantitatively demonstrate that a large fraction of functional sites in the proteome may be accurately identified from sequence and structure. This should help focus structure-function studies, rational drug design, protein engineering, and functional annotation to the relevant regions of a protein.

Amino Acid Motifs↗

Physiological genomics of Escherichia coli protein families.

The well-researched Escherichia coli genome offers the opportunity to explore the value of using protein families within a single organism to enrich functional annotation procedures and to study mechanisms of protein evolution. Having identified multimodular proteins resulting from gene fusion, and treated each module as a separate protein, nonoverlapping sequence-similar families in E. coli could be assembled. Of 3,902 proteins of length 100 residues or more, 2,415 clustered into 609 protein families. The relatedness of function among members of each family was dissected in detail. Data on paralogous protein families provides valuable information in attributing putative function to unknown genes, supplementing existing function annotation. Enzymes, transporters, and regulators represent the three major types of proteins in E. coli. They are shown to have distinctive patterns in gene duplication and divergence and gene fusion, suggesting that details of protein evolution have been different for genes in these categories. Data for the complete list of paralogous protein families and updated functional annotation for E. coli K-12 are accessible in GenProtEC (http://genprotec.mbl.edu).

Carrier Proteins↗

Functional annotation of the putative orphan Caenorhabditis elegans G-protein-coupled receptor C10C6.2 as a FLP15 peptide receptor.

This report describes the cloning and functional annotation of a Caenorhabditis elegans orphan G-protein-coupled receptor (GPCR) (C10C6.2) as a receptor for the FMRFamide-related peptides (FaRPs) encoded on the flp15 precursor gene, leading to the receptor designation FLP15-R. A cDNA encoding C10C6.2 was obtained using PCR techniques, confirmed identical to the Worm-pep-predicted sequence, and cloned into a vector appropriate for eucaryotic expression. A [35S]guanosine 5'-O-(thiotriphosphate) (GTPgammaS) assay with membranes prepared from Chinese hamster ovary (CHO) cells transiently transfected with FLP15-R was used as a read-out for receptor activation. FLP15-R was activated by putative FLP15 peptides, GGPQGPLRF-NH2 (FLP15-1), RGPSGPLRF-NH2 (FLP15-2A), its des-Arg1 counterpart, GPSGPLRF-NH2 (FLP15-2B), and to a lesser extent, by a tobacco hornworm Manduca sexta FaRP, GNSFLRFNH2 (F7G) (potency ranking FLP15-2A > FLP15-1 > FLP15-2B >> F7G). FLP15-R activation was abolished in the transfected cells pretreated with pertussis toxin, suggesting a preferential receptor coupling to Gi/Go proteins. The functional expression of FLP15-R in mammalian cells was temperature-dependent. Either no stimulation or significantly lower ligand-evoked [35S]GTPgammaS binding was observed in membranes prepared from transfected FLP15-R/CHO cells cultured at 37 degrees C. However, a 37 to 28 degrees C temperature shift implemented 24 h post-transfection consistently resulted in an improved activation signal and was essential for detectable functional expression of FLP15-R in CHO cells. To our knowledge, the FLP15 receptor is only the second deorphanized C. elegans neuropeptide GPCR reported to date.

Amino Acid Sequence↗