Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Uncovering conserved patterns in bioactive peptides in Metazoa.

Bioactive (neuro)peptides play critical roles in regulating most biological processes in animals. Peptides belonging to the same family are characterized by a typical sequence pattern that is conserved among the family's peptide members. Such a conserved pattern or motif usually corresponds to the functionally important part of the biologically active peptide. In this paper, all known bioactive (neuro)peptides annotated in Swiss-Prot and TrEMBL protein databases are collected, and the pattern searching program Pratt is used to search these unaligned peptide sequences for conserved patterns. The obtained patterns are then refined by combining the information on amino acids at important functional sites collected from the literature. All the identified patterns are further tested by scanning them against Swiss-Prot and TrEMBL protein databases. The diagnostic power of each pattern is validated by the fact that any annotated protein from Swiss-Prot and TrEMBL that contains one of the established patterns, is indeed a known (neuro)peptide precursor. We discovered 155 novel peptide patterns in addition to the 56 established ones in the PROSITE database. All the patterns cover 110 peptide families. Fifty-five of these families are not characterized by the PROSITE signatures, and 12 are also not identified by other existing motif databases, such as Pfam and SMART. Using the newly identified peptide signatures as a search tool, we predicted 95 hypothetical proteins as putative peptide precursors.

Amino Acid Motifs↗

From structure to function: YrbI from Haemophilus influenzae (HI1679) is a phosphatase.

The crystal structure of the YrbI protein from Haemophilus influenzae (HI1679) was determined at a 1.67-A resolution. The function of the protein had not been assigned previously, and it is annotated as hypothetical in sequence databases. The protein exhibits the alpha/beta-hydrolase fold (also termed the Rossmann fold) and resembles most closely the fold of the L-2-haloacid dehalogenase (HAD) superfamily. Following this observation, a detailed sequence analysis revealed remote homology to two members of the HAD superfamily, the P-domain of Ca(2+) ATPase and phosphoserine phosphatase. The 19-kDa chains of HI1679 form a tetramer both in solution and in the crystalline form. The four monomers are arranged in a ring such that four beta-hairpin loops, each inserted after the first beta-strand of the core alpha/beta-fold, form an eight-stranded barrel at the center of the assembly. Four active sites are located at the subunit interfaces. Each active site is occupied by a cobalt ion, a metal used for crystallization. The cobalt is octahedrally coordinated to two aspartate side-chains, a backbone oxygen, and three solvent molecules, indicating that the physiological metal may be magnesium. HI1679 hydrolyzes a number of phosphates, including 6-phosphogluconate and phosphotyrosine, suggesting that it functions as a phosphatase in vivo. The physiological substrate is yet to be identified; however the location of the gene on the yrb operon suggests involvement in sugar metabolism.

Amino Acid Sequence↗

Automated protein function prediction--the genomic challenge.

Overwhelmed with genomic data, biologists are facing the first big post-genomic question--what do all genes do? First, not only is the volume of pure sequence and structure data growing, but its diversity is growing as well, leading to a disproportionate growth in the number of uncharacterized gene products. Consequently, established methods of gene and protein annotation, such as homology-based transfer, are annotating less data and in many cases are amplifying existing erroneous annotation. Second, there is a need for a functional annotation which is standardized and machine readable so that function prediction programs could be incorporated into larger workflows. This is problematic due to the subjective and contextual definition of protein function. Third, there is a need to assess the quality of function predictors. Again, the subjectivity of the term 'function' and the various aspects of biological function make this a challenging effort. This article briefly outlines the history of automated protein function prediction and surveys the latest innovations in all three topics.

Algorithms↗

Computational identification and systematic analysis of the ACR gene family in Oryza sativa.

Based on sequence similarity search and domain detection, nine ACT domain repeat protein-coding genes (the "ACR" genes) in rice were identified, which were mainly distributed on the chromosomes 2, 3, 4, and 8. An InterPro database search indicated that four copies of the ACT domain linearly occupied the entire polypeptide. The first three ACT domains were linked by two different sequences. However, the fourth ACT domain was extremely close to ACT3. Gene structure comparisons showed large differences in exon numbers, from three to eight, among members of the rice ACR gene family. In addition, it appeared that gene duplication might be operative when the compositions of exons and introns were analyzed. Phylogenetic analysis divided the ACR gene family into five distinct groups, and this division was generally according to the expression patterns of the ACR genes. The Arabidopsis and rice ACR proteins were clustered across together, suggesting that these ACR genes might originate from an ancient common ancestor. Notably, the identification of orthologues and paralogues would be useful for rice gene functional annotation.

Chromosome Mapping↗

Systematic learning of gene functional classes from DNA array expression data by using multilayer perceptrons.

Recent advances in microarray technology have opened new ways for functional annotation of previously uncharacterised genes on a genomic scale. This has been demonstrated by unsupervised clustering of co-expressed genes and, more importantly, by supervised learning algorithms. Using prior knowledge, these algorithms can assign functional annotations based on more complex expression signatures found in existing functional classes. Previously, support vector machines (SVMs) and other machine-learning methods have been applied to a limited number of functional classes for this purpose. Here we present, for the first time, the comprehensive application of supervised neural networks (SNNs) for functional annotation. Our study is novel in that we report systematic results for ~100 classes in the Munich Information Center for Protein Sequences (MIPS) functional catalog. We found that only ~10% of these are learnable (based on the rate of false negatives). A closer analysis reveals that false positives (and negatives) in a machine-learning context are not necessarily "false" in a biological sense. We show that the high degree of interconnections among functional classes confounds the signatures that ought to be learned for a unique class. We term this the "Borges effect" and introduce two new numerical indices for its quantification. Our analysis indicates that classification systems with a lower Borges effect are better suitable for machine learning. Furthermore, we introduce a learning procedure for combining false positives with the original class. We show that in a few iterations this process converges to a gene set that is learnable with considerably low rates of false positives and negatives and contains genes that are biologically related to the original class, allowing for a coarse reconstruction of the interactions between associated biological pathways. We exemplify this methodology using the well-studied tricarboxylic acid cycle.

Algorithms↗

AgBase: a unified resource for functional analysis in agriculture.

Analysis of functional genomics (transcriptomics and proteomics) datasets is hindered in agricultural species because agricultural genome sequences have relatively poor structural and functional annotation. To facilitate systems biology in these species we have established the curated, web-accessible, public resource 'AgBase' (www.agbase.msstate.edu). We have improved the structural annotation of agriculturally important genomes by experimentally confirming the in vivo expression of electronically predicted proteins and by proteogenomic mapping. Proteogenomic data are available from the AgBase proteogenomics link. We contribute Gene Ontology (GO) annotations and we provide a two tier system of GO annotations for users. The 'GO Consortium' gene association file contains the most rigorous GO annotations based solely on experimental data. The 'Community' gene association file contains GO annotations based on expert community knowledge (annotations based directly from author statements and submitted annotations from the community) and annotations for predicted proteins. We have developed two tools for proteomics analysis and these are freely available on request. A suite of tools for analyzing functional genomics datasets using the GO is available online at the AgBase site. We encourage and publicly acknowledge GO annotations from researchers and provide an online mechanism for agricultural researchers to submit requests for GO annotations.

Agriculture↗

A revised annotation and comparative analysis of Helicobacter pylori genomes.

Huge amounts of genomic information are currently being generated. Therefore, biologists require structured, exhaustive and comparative databases. The PyloriGene database (http://genolist.pasteur.fr/PyloriGene) was developed to respond to these needs, by integrating and connecting the information generated during the sequencing of two distinct strains of Helicobacter pylori. This led to the need for a general annotation consensus, as the physical and functional annotations of the two strains differed significantly in some cases. A revised functional classification system was created to accommodate the existing data and to make it possible to classify coding sequences (CDS) into several functional categories to harmonize CDS classification. The annotation of the two complete genomes was revised in the light of new data, allowing us to reduce the percentage of hypothetical proteins from approximately 40 to 33%. This resulted in the reassignment of functions for 108 CDS (approximately 7% of all CDS). Interestingly, the functions of only approximately 13% of CDS (222 out of 1658 CDS) were annotated as a result of work done directly on H.pylori genes. Finally, comparison of the two published genomes revealed a significant amount of size variation between corresponding (orthologous) CDS. Most of these size variations were due to natural polymorphisms, although other sources of variation were identified, such as pseudogenes, new genes potentially regulated by slipped-strand mispairing mechanism, or frame-shifts. 113 of these differences were due to different start codon assignments, a common problem when constructing physical annotations.

Databases, Nucleic Acid↗

Genomic survey of cAMP and cGMP signalling components in the cyanobacterium Synechocystis PCC 6803.

Cyanobacteria modulate intracellular levels of cAMP and cGMP in response to environmental conditions (light, nutrients and pH). In an attempt to identify components of the cAMP and cGMP signalling pathways in Synechocystis PCC 6803, the authors screened its complete genome sequence by using bioinformatic tools and data from sequence-function studies performed on both eukaryotic and prokaryotic cAMP/cGMP-dependent proteins. Sll1624 and Slr2100 were tentatively assigned as being two putative cyclic nucleotide phosphodiesterases. Five proteins were identified as having all the determinants required to be cyclic nucleotide receptors, two of them being probably more specific for cGMP (an element of two-component regulatory systems - Slr2104 - and a putative cyclic-nucleotide-gated cation channel - Slr1575), the three others being probably more specific for cAMP: (i) a protein of unidentified function (Slr0842); (ii) a putative cyclic-nucleotide-modulated permease (Slr0593), previously annotated as a kinase A regulatory subunit; and (iii) a putative transcription factor (CRP-SYN: =Sll1371), which possesses cAMP- and DNA-binding determinants homologous to those of the cAMP receptor protein of Escherichia coli (CRP-EC:). This homology, together with the presence in Synechocystis of CRP-EC:-like binding sites upstream of crp, cya1, slr1575, and several genes encoding enzymes involved in transport and metabolism, strongly suggests that CRP-SYN: is a global regulator.

3',5'-Cyclic-AMP Phosphodiesterases↗

metaExpertPro: A Computational Workflow for Metaproteomics Spectral Library Construction and Data-Independent Acquisition Mass Spectrometry Data Analysis.

Analysis of large-scale data-independent acquisition mass spectrometry metaproteomics data remains a computational challenge. Here, we present a computational pipeline called metaExpertPro for metaproteomics data analysis. This pipeline encompasses spectral library generation using data-dependent acquisition MS, protein identification and quantification using data-independent acquisition mass spectrometry, functional and taxonomic annotation, as well as quantitative matrix generation for both microbiota and hosts. By integrating FragPipe and DIA-NN, metaExpertPro offers compatibility with both Orbitrap and timsTOF MS instruments. To evaluate the depth and accuracy of identification and quantification, we conducted extensive assessments using human fecal samples and benchmark tests. Performance tests conducted on human fecal samples indicated that metaExpertPro quantified an average of 45,000 peptides in a 60-min diaPASEF injection. Notably, metaExpertPro outperformed three existing software tools by characterizing a higher number of peptides and proteins. Importantly, metaExpertPro maintained a low factual false discovery rate of approximately 5% for protein groups across four benchmark tests. Applying a filter of five peptides per genus, metaExpertPro achieved relatively high accuracy (F-score = 0.67-0.90) in genus diversity and showed a high correlation (rSpearman = 0.73-0.82) between the measured and true genus relative abundance in benchmark tests. Additionally, the quantitative results at the protein, taxonomy, and function levels exhibited high reproducibility and consistency across the commonly adopted public human gut microbial protein databases IGC and UHGP. In a metaproteomic analysis of dyslipidemia patients, metaExpertPro revealed characteristic alterations in microbial functions and potential interactions between the microbiota and the host.

Proteomics↗

Yeast Protein database (YPD): a database for the complete proteome of Saccharomyces cerevisiae.

The Yeast Protein Database (YPD) is a database for the proteins of the budding yeast,Saccharomyces cerevisiae. YPD is the first annotated database for the complete proteome of any organism. Now that the complete genome sequence of yeast is available, YPD contains entries for each of the characterized proteins and for each of the uncharacterized proteins predicted from the sequence. Contained in YPD are the calculated properties of each protein such as molecular weight and isoelectric point, experimentally determined properties such as subcellular localization and post-translational modifications, and extensive annotations from the yeast literature. YPD contains 25 000 lines of textual annotation that describe the known functions, mutant phenotypes, interactions, and other properties for the approximately 6000 proteins in the yeast proteome. The information in YPD is updated daily, and it is available on the World Wide Web at http://www.proteome.com/YPDhome.html .

Amino Acid Sequence↗

dbPTM: an information repository of protein post-translational modification.

dbPTM is a database that compiles information on protein post-translational modifications (PTMs), such as the catalytic sites, solvent accessibility of amino acid residues, protein secondary and tertiary structures, protein domains and protein variations. The database includes all of the experimentally validated PTM sites from Swiss-Prot, PhosphoELM and O-GLYCBASE. Only a small fraction of Swiss-Prot proteins are annotated with experimentally verified PTM. Although the Swiss-Prot provides rich information about the PTM, other structural properties and functional information of proteins are also essential for elucidating protein mechanisms. The dbPTM systematically identifies three major types of protein PTM (phosphorylation, glycosylation and sulfation) sites against Swiss-Prot proteins by refining our previously developed prediction tool, KinasePhos (http://kinasephos.mbc.nctu.edu.tw/). Solvent accessibility and secondary structure of residues are also computationally predicted and are mapped to the PTM sites. The resource is now freely available at http://dbPTM.mbc.nctu.edu.tw/.

Amino Acids↗

Prediction and functional analysis of native disorder in proteins from the three kingdoms of life.

An automatic method for recognizing natively disordered regions from amino acid sequence is described and benchmarked against predictors that were assessed at the latest critical assessment of techniques for protein structure prediction (CASP) experiment. The method attains a Wilcoxon score of 90.0, which represents a statistically significant improvement on the methods evaluated on the same targets at CASP. The classifier, DISOPRED2, was used to estimate the frequency of native disorder in several representative genomes from the three kingdoms of life. Putative, long (>30 residue) disordered segments are found to occur in 2.0% of archaean, 4.2% of eubacterial and 33.0% of eukaryotic proteins. The function of proteins with long predicted regions of disorder was investigated using the gene ontology annotations supplied with the Saccharomyces genome database. The analysis of the yeast proteome suggests that proteins containing disorder are often located in the cell nucleus and are involved in the regulation of transcription and cell signalling. The results also indicate that native disorder is associated with the molecular functions of kinase activity and nucleic acid binding.

Databases, Genetic↗

Organelle DB: an updated resource of eukaryotic protein localization and function.

Organelle DB (http://organelledb.lsi.umich.edu) is a web-accessible relational database presenting a supplemented catalog of organelle-localized proteins and major protein complexes. Since its release in 2004, Organelle DB has grown by 20% to encompass over 30,000 proteins from 138 eukaryotic organisms. Each protein in Organelle DB is presented with its subcellular localization, primary sequence and a detailed description of its function, as available. All records in Organelle DB have been annotated using controlled vocabulary from the Gene Ontology consortium. Protein localization data are inherently visual, and Organelle DB is a significant repository of biological images, housing 1500 micrographs of yeast cells carrying stained proteins. Furthermore, we report here the development of Organelle View, an extension of Organelle DB for the interactive visualization of organelles and subcellular structures in the budding yeast Saccharomyces cerevisiae. Organelle View offers a dimensional representation of a yeast cell; users can search Organelle View for proteins of interest, and the organelles housing these proteins will be highlighted in the cell image. Among other applications, Organelle View may serve as an educational aid engaging introductory biology students through a visually 'fun' interface. Organelle View can be accessed from the Organelle DB home page or directly at http://organelleview.lsi.umich.edu.

Animals↗

Real spherical harmonic expansion coefficients as 3D shape descriptors for protein binding pocket and ligand comparisons.

MOTIVATION: An increasing number of protein structures are being determined for which no biochemical characterization is available. The analysis of protein structure and function assignment is becoming an unexpected challenge and a major bottleneck towards the goal of well-annotated genomes. As shape plays a crucial role in biomolecular recognition and function, the examination and development of shape description and comparison techniques is likely to be of prime importance for understanding protein structure-function relationships. RESULTS: A novel technique is presented for the comparison of protein binding pockets. The method uses the coefficients of a real spherical harmonics expansion to describe the shape of a protein's binding pocket. Shape similarity is computed as the L2 distance in coefficient space. Such comparisons in several thousands per second can be carried out on a standard linux PC. Other properties such as the electrostatic potential fit seamlessly into the same framework. The method can also be used directly for describing the shape of proteins and other molecules. AVAILABILITY: A limited version of the software for the real spherical harmonics expansion of a set of points in PDB format is freely available upon request from the authors. Binding pocket comparisons and ligand prediction will be made available through the protein structure annotation pipeline Profunc (written by Roman Laskowski) which will be accessible from the EBI website shortly.

Algorithms↗

Integrating machine learning and GWAS for variant prioritization in the INCIPE cohort highlights ABC transporter genes in chronic kidney disease.

INTRODUCTION: Chronic kidney disease (CKD) is a major public health challenge, affecting approximately 674 million people worldwide and representing one of the fastest-growing causes of mortality. Since CKD is frequently asymptomatic in its early stages, the identification of novel genetic biomarkers may improve early detection and risk stratification. Genome-Wide Association Studies (GWAS) have identified numerous genetic loci associated with CKD and related traits; however, their performance is often limited in small and imbalanced cohorts, where reduced statistical power increases both false-positive and false-negative findings. Machine learning (ML) approaches can complement conventional GWAS by prioritizing biologically relevant genetic signals from high-dimensional genomic data. METHODS: In this study, we implemented a nested ensemble (NCBC) model composed of an undersampler and a CatBoostClassifier (CBC) to prioritize candidate genetic variants associated with CKD in the INCIPE cohort. Prioritized variants were functionally annotated and evaluated through enrichment analyses, GTEx gene expression profiling, and protein-protein interaction network analyses. Genes identified by the CKDGen Consortium were analysed as an external reference set and used to validate the biological relevance of the prioritized results. RESULTS: The NCBC model outperformed conventional ML classifiers, achieving a ROC AUC score of 87.77%, compared to 50%-53% for the other evaluated models. Among the prioritized genes, 56.25% showed protein-protein interactions with genes previously reported by the CKDGen Consortium, whereas only 1.9% of randomly generated gene sets showed interactions. DISCUSSION: Our study demonstrates that the NCBC model improves the prioritization of biologically plausible candidate variants in a small and imbalanced CKD cohort. Functional analyses suggested ABC transporter-related genes, including ABCA13, ABCA4, and ABCC4 genes, as promising candidate for future validation, with ABCA4 showing substantial expression in kidney tissues. Overall, these findings support the integration of ML with GWAS to prioritize candidate genes and investigate the genetic architecture of complex diseases.

SNP prioritization↗

The role of protein structure in genomics.

The genome projects produce an enormous amount of sequence data that needs to be annotated in terms of molecular structure and biological function. These tasks have triggered additional initiatives like structural genomics. The intention is to determine as many protein structures as possible, in the most efficient way, and to exploit the solved structures for the assignment of biological function to hypothetical proteins. We discuss the impact of these developments on protein classification, gene function prediction, and protein structure prediction.

Databases, Factual↗

SPD--a web-based secreted protein database.

With the improved secreted protein prediction approach and comprehensive data sources, including Swiss-Prot, TrEMBL, RefSeq, Ensembl and CBI-Gene, we have constructed secretomes of human, mouse and rat, with a total of 18 152 secreted proteins. All the entries are ranked according to the prediction confidence. They were further annotated via a proteome annotation pipeline that we developed. We also set up a secreted protein classification pipeline and classified our predicted secreted proteins into different functional categories. To make the dataset more convincing and comprehensive, nine reference datasets are also integrated, such as the secreted proteins from the Gene Ontology Annotation (GOA) system at the European Bioinformatics Institute, and the vertebrate secreted proteins from Swiss-Prot. All these entries were grouped via a TribeMCL based clustering pipeline. We have constructed a web-based secreted protein database, which has been publicly available at http://spd.cbi.pku.edu.cn. Users can browse the database via a GO assignment or chromosomal-location-based interface. Moreover, text query and sequence similarity search are also provided, and the sequence and annotation data can be downloaded freely from the SPD website.

Animals↗

Gene and protein profiling of the response of MA-10 Leydig tumor cells to human chorionic gonadotropin.

Activation of the steroidogenic machinery by peptide hormones involves a number of steps for transmitting signals from the plasma membrane to mitochondria in a spatially and temporally coordinated manner. Although key proteins mediating the hormonal signal have been identified, recent data suggest that the pathway might involve more complex protein-protein and protein-lipid interactions. Genomic and proteomic methods of analysis, namely the Affymetrix Murine Genome U74A v2 GeneChip and the BD PowerBlot Western Array, were used to identify human chorionic gonadotropin (hCG)-induced changes in mRNA and protein of MA-10 Leydig tumor cells that parallel the increase seen in progesterone synthesis. To analyze the massive amount of data that was generated, a comprehensive protein information matrix summarizing the features of each gene or protein, including its known properties, as well as annotations derived by homology-based functional inference, was developed. Of the genes examined by Affymetrix array, approximately 79 were differentially expressed and of gene products examined by PowerBlot, 9 were differentially expressed (above twofold). Changes in the expression of selected transcripts of interest were confirmed using real-time quantitative polymerase chain reaction and immunoblot analyses. Collectively, these results indicate that hormonal regulation of steroidogenesis is a complex phenomenon, involving proteins that participate in various known and novel pathways, which are implicated in transmitting signals from the plasma membrane to mitochondria and nucleus.

Blotting, Western↗