Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Proteomics-Based Identification of the Pyroptosis-Related Biomarker PCSK9 and Its Association With the Pathogenesis of Rheumatoid Arthritis.

Rheumatoid arthritis (RA) is a common autoimmune disease, and early diagnosis is critical for effective treatment. This study aims to identify potential biomarkers related to pyroptosis through serum proteomics analysis, offering new insights for the early diagnosis of RA. We enrolled 100 participants, including 50 patients with RA and 50 healthy controls. Serum samples were collected and analyzed using high-resolution liquid chromatography-tandem mass spectrometry (LC-MS/MS) for proteomics profiling. Differential protein expression analysis and functional annotation revealed significant upregulation of pyroptosis-related proteins in the serum of patients with RA. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses, along with protein-protein interaction (PPI) network analysis, showed that these proteins are involved in inflammation and immune pathways, particularly the activation of the NOD-like receptor protein 3 (NLRP3) inflammasome. Enzyme-linked immunosorbent assay (ELISA) validation confirmed a significant increase in PCSK9 levels in patients with RA, suggesting that PCSK9 may play a key role in the pathogenesis of RA. This study provides new directions for biomarker research in RA, particularly regarding the potential involvement of the pyroptosis pathway, with significant clinical application prospects.

Humans↗

Using evolutionary and structural information to predict DNA-binding sites on DNA-binding proteins.

Proteins that interact with DNA are involved in a number of fundamental biological activities such as DNA replication, transcription, and repair. A reliable identification of DNA-binding sites in DNA-binding proteins is important for functional annotation, site-directed mutagenesis, and modeling protein-DNA interactions. We apply Support Vector Machine (SVM), a supervised pattern recognition method, to predict DNA-binding sites in DNA-binding proteins using the following features: amino acid sequence, profile of evolutionary conservation of sequence positions, and low-resolution structural information. We use a rigorous statistical approach to study the performance of predictors that utilize different combinations of features and how this performance is affected by structural and sequence properties of proteins. Our results indicate that an SVM predictor based on a properly scaled profile of evolutionary conservation in the form of a position specific scoring matrix (PSSM) significantly outperforms a PSSM-based neural network predictor. The highest accuracy is achieved by SVM predictor that combines the profile of evolutionary conservation with low-resolution structural information. Our results also show that knowledge-based predictors of DNA-binding sites perform significantly better on proteins from mainly-alpha structural class and that the performance of these predictors is significantly correlated with certain structural and sequence properties of proteins. These observations suggest that it may be possible to assign a reliability index to the overall accuracy of the prediction of DNA-binding sites in any given protein using its sequence and structural properties. A web-server implementation of the predictors is freely available online at http://lcg.rit.albany.edu/dp-bind/.

Amino Acid Sequence↗

Identifying cysteines and histidines in transition-metal-binding sites using support vector machines and neural networks.

Accurate predictions of metal-binding sites in proteins by using sequence as the only source of information can significantly help in the prediction of protein structure and function, genome annotation, and in the experimental determination of protein structure. Here, we introduce a method for identifying histidines and cysteines that participate in binding of several transition metals and iron complexes. The method predicts histidines as being in either of two states (free or metal bound) and cysteines in either of three states (free, metal bound, or in disulfide bridges). The method uses only sequence information by utilizing position-specific evolutionary profiles as well as more global descriptors such as protein length and amino acid composition. Our solution is based on a two-stage machine-learning approach. The first stage consists of a support vector machine trained to locally classify the binding state of single histidines and cysteines. The second stage consists of a bidirectional recurrent neural network trained to refine local predictions by taking into account dependencies among residues within the same protein. A simple finite state automaton is employed as a postprocessing in the second stage in order to enforce an even number of disulfide-bonded cysteines. We predict histidines and cysteines in transition-metal-binding sites at 73% precision and 61% recall. We observe significant differences in performance depending on the ligand (histidine or cysteine) and on the metal bound. We also predict cysteines participating in disulfide bridges at 86% precision and 87% recall. Results are compared to those that would be obtained by using expert information as represented by PROSITE motifs and, for disulfide bonds, to state-of-the-art methods.

Amino Acid Sequence↗

Comparison of protein interaction networks reveals species conservation and divergence.

BACKGROUND: Recent progresses in high-throughput proteomics have provided us with a first chance to characterize protein interaction networks (PINs), but also raised new challenges in interpreting the accumulating data. RESULTS: Motivated by the need of analyzing and interpreting the fast-growing data in the field of proteomics, we propose a comparative strategy to carry out global analysis of PINs. We compare two PINs by combining interaction topology and sequence similarity to identify conserved network substructures (CoNSs). Using this approach we perform twenty-one pairwise comparisons among the seven recently available PINs of E.coli, H.pylori, S.cerevisiae, C.elegans, D.melanogaster, M.musculus and H.sapiens. In spite of the incompleteness of data, PIN comparison discloses species conservation at the network level and the identified CoNSs are also functionally conserved and involve in basic cellular functions. We investigate the yeast CoNSs and find that many of them correspond to known complexes. We also find that different species harbor many conserved interaction regions that are topologically identical and these regions can constitute larger interaction regions that are topologically different but similar in framework. Based on the species-to-species difference in CoNSs, we infer potential species divergence. It seems that different species organize orthologs in similar but not necessarily the same topology to achieve similar or the same function. This attributes much to duplication and divergence of genes and their associated interactions. Finally, as the application of CoNSs, we predict 101 protein-protein interactions (PPIs), annotate 339 new protein functions and deduce 170 pairs of orthologs. CONCLUSION: Our result demonstrates that the cross-species comparison strategy we adopt is powerful for the exploration of biological problems from the perspective of networks.

Animals↗

The institute for genomic research Osa1 rice genome annotation database.

We have developed a rice (Oryza sativa) genome annotation database (Osa1) that provides structural and functional annotation for this emerging model species. Using the sequence of O. sativa subsp. japonica cv Nipponbare from the International Rice Genome Sequencing Project, pseudomolecules, or virtual contigs, of the 12 rice chromosomes were constructed. Our most recent release, version 3, represents our third build of the pseudomolecules and is composed of 98% finished sequence. Genes were identified using a series of computational methods developed for Arabidopsis (Arabidopsis thaliana) that were modified for use with the rice genome. In release 3 of our annotation, we identified 57,915 genes, of which 14,196 are related to transposable elements. Of these 43,719 non-transposable element-related genes, 18,545 (42.4%) were annotated with a putative function, 5,777 (13.2%) were annotated as encoding an expressed protein with no known function, and the remaining 19,397 (44.4%) were annotated as encoding a hypothetical protein. Multiple splice forms (5,873) were detected for 2,538 genes, resulting in a total of 61,250 gene models in the rice genome. We incorporated experimental evidence into 18,252 gene models to improve the quality of the structural annotation. A series of functional data types has been annotated for the rice genome that includes alignment with genetic markers, assignment of gene ontologies, identification of flanking sequence tags, alignment with homologs from related species, and syntenic mapping with other cereal species. All structural and functional annotation data are available through interactive search and display windows as well as through download of flat files. To integrate the data with other genome projects, the annotation data are available through a Distributed Annotation System and a Genome Browser. All data can be obtained through the project Web pages at http://rice.tigr.org.

Computational Biology↗

The predicted secretome of Lactobacillus plantarum WCFS1 sheds light on interactions with its environment.

The predicted extracellular proteins of the bacterium Lactobacillus plantarum were analysed to gain insight into the mechanisms underlying interactions of this bacterium with its environment. Extracellular proteins play important roles in processes ranging from probiotic effects in the gastrointestinal tract to degradation of complex extracellular carbon sources such as those found in plant materials, and they have a primary role in the adaptation of a bacterium to changing environmental conditions. The functional annotation of extracellular proteins was improved using a wide variety of bioinformatics methods, including domain analysis and phylogenetic profiling. At least 12 proteins are predicted to be directly involved in adherence to host components such as collagen and mucin, and about 30 extracellular enzymes, mainly hydrolases and transglycosylases, might play a role in the degradation of substrates by L. plantarum to sustain its growth in different environmental niches. A comprehensive overview of all predicted extracellular proteins, their domains composition and their predicted function is provided through a database at http://www.cmbi.ru.nl/secretome which could serve as a basis for targeted experimental studies into the function of extracellular proteins.

Bacterial Adhesion↗

Efficient recognition of protein fold at low sequence identity by conservative application of Psi-BLAST: validation.

A substantial fraction of protein sequences derived from genomic analyses is currently classified as representing 'hypothetical proteins of unknown function'. In part, this reflects the limitations of methods for comparison of sequences with very low identity. We evaluated the effectiveness of a Psi-BLAST search strategy to identify proteins of similar fold at low sequence identity. Psi-BLAST searches for structurally characterized low-sequence-identity matches were carried out on a set of over 300 proteins of known structure. Searches were conducted in NCBI's non-redundant database and were limited to three rounds. Some 614 potential homologs with 25% or lower sequence identity to 166 members of the search set were obtained. Disregarding the expect value, level of sequence identity and span of alignment, correspondence of fold between the target and potential homolog was found in more than 95% of the Psi-BLAST matches. Restrictions on expect value or span of alignment improved the false positive rate at the expense of eliminating many true homologs. Approximately three-quarters of the putative homologs obtained by three rounds of Psi-BLAST revealed no significant sequence similarity to the target protein upon direct sequence comparison by BLAST, and therefore could not be found by a conventional search. Although three rounds of Psi-BLAST identified many more homologs than a standard BLAST search, most homologs were undetected. It appears that more than 80% of all homologs to a target protein may be characterized by a lack of significant sequence similarity. We suggest that conservative use of Psi-BLAST has the potential to propose experimentally testable functions for the majority of proteins currently annotated as 'hypothetical proteins of unknown function'.

Algorithms↗

Automatic annotation for biological sequences by extraction of keywords from MEDLINE abstracts. Development of a prototype system.

We have developed a prototype for the automatic annotation of functional characteristics in protein families. The system is able to extract biological information directly from scientific literature in the form of MEDLINE abstracts. The criterion for selecting relevant keywords is the difference between their frequency in the abstracts associated with the protein family under study and its frequency in other unrelated protein families. The concept of functional information associated to protein families is the key feature of our system and gathers evolutionary information into the problem of functional annotation of biological sequences. The system has been tested in two different scenarios: first, a large set of protein families with a small number of abstract per family and second, selected protein families with large number of abstracts attached to each one. In both cases the performances are compared with annotations provided by human experts showing a clear relation between the amount of information provided to the system and the quality of the annotations. The automatic annotations are in many cases of similar quality to the ones contained in current data bases. The possibilities and difficulties to be encountered during the development of a full system for automatic annotation are discussed.

Abstracting and Indexing↗

Effective function annotation through catalytic residue conservation.

Because of the extreme impact of genome sequencing projects, protein sequences without accompanying experimental data now dominate public databases. Homology searches, by providing an opportunity to transfer functional information between related proteins, have become the de facto way to address this. Although a single, well annotated, close relationship will often facilitate sufficient annotation, this situation is not always the case, particularly if mutations are present in important functional residues. When only distant relationships are available, the transfer of function information is more tenuous, and the likelihood of encountering several well annotated proteins with different functions is increased. The consequence for a researcher is a range of candidate functions with little way of knowing which, if any, are correct. Here, we address the problem directly by introducing a computational approach to accurately identify and segregate related proteins into those with a functional similarity and those where function differs. This approach should find a wide range of applications, including the interpretation of genomics/proteomics data and the prioritization of targets for high-throughput structure determination. The method is generic, but here we concentrate on enzymes and apply high-quality catalytic site data. In addition to providing a series of comprehensive benchmarks to show the overall performance of our approach, we illustrate its utility with specific examples that include the correct identification of haptoglobin as a nonenzymatic relative of trypsin, discrimination of acid-d-amino acid ligases from a much larger ligase pool, and the successful annotation of BioH, a structural genomics target.

Amino Acid Sequence↗

Modeling a whole organ using proteomics: the avian bursa of Fabricius.

While advances in proteomics have improved proteome coverage and enhanced biological modeling, modeling function in multicellular organisms requires understanding how cells interact. Here we used the chicken bursa of Fabricius, a common experimental system for B cell function, to model organ function from proteomics data. The bursa has two major functional cell types: B cells and the supporting stromal cells. We used differential detergent fractionation-multidimensional protein identification technology (DDF-MudPIT) to identify 5198 proteins from all cellular compartments. Of these, 1753 were B cell specific, 1972 were stroma specific and 1473 were shared between the two. By modeling programmed cell death (PCD), cell differentiation and proliferation, and transcriptional activation, we have improved functional annotation of chicken proteins and placed chicken-specific death receptors into the PCD process using phylogenetics. We have identified 114 transcription factors (TFs); 42 of the bursal B cell TFs have not been reported before in any B cells. We have also improved the structural annotation of a newly sequenced genome by confirming the in vivo expression of 4006 "predicted", and 6623 ab initio, ORFs. Finally, we have developed a novel method for facilitating structural annotation, "expressed peptide sequence tags" (ePSTs) and demonstrate its utility by identifying 521 potential novel proteins from the chicken "unassigned chromosome".

Amino Acid Sequence↗

Annotating nucleic acid-binding function based on protein structure.

Many of the targets of structural genomics will be proteins with little or no structural similarity to those currently in the database. Therefore, novel function prediction methods that do not rely on sequence or fold similarity to other known proteins are needed. We present an automated approach to predict nucleic-acid-binding (NA-binding) proteins, specifically DNA-binding proteins. The method is based on characterizing the structural and sequence properties of large, positively charged electrostatic patches on DNA-binding protein surfaces, which typically coincide with the DNA-binding-sites. Using an ensemble of features extracted from these electrostatic patches, we predict DNA-binding proteins with high accuracy. We show that our method does not rely on sequence or structure homology and is capable of predicting proteins of novel-binding motifs and protein structures solved in an unbound state. Our method can also distinguish NA-binding proteins from other proteins that have similar, large positive electrostatic patches on their surfaces, but that do not bind nucleic acids.

Amino Acid Motifs↗

Prediction of protein function from protein sequence and structure.

The sequence of a genome contains the plans of the possible life of an organism, but implementation of genetic information depends on the functions of the proteins and nucleic acids that it encodes. Many individual proteins of known sequence and structure present challenges to the understanding of their function. In particular, a number of genes responsible for diseases have been identified but their specific functions are unknown. Whole-genome sequencing projects are a major source of proteins of unknown function. Annotation of a genome involves assignment of functions to gene products, in most cases on the basis of amino-acid sequence alone. 3D structure can aid the assignment of function, motivating the challenge of structural genomics projects to make structural information available for novel uncharacterized proteins. Structure-based identification of homologues often succeeds where sequence-alone-based methods fail, because in many cases evolution retains the folding pattern long after sequence similarity becomes undetectable. Nevertheless, prediction of protein function from sequence and structure is a difficult problem, because homologous proteins often have different functions. Many methods of function prediction rely on identifying similarity in sequence and/or structure between a protein of unknown function and one or more well-understood proteins. Alternative methods include inferring conservation patterns in members of a functionally uncharacterized family for which many sequences and structures are known. However, these inferences are tenuous. Such methods provide reasonable guesses at function, but are far from foolproof. It is therefore fortunate that the development of whole-organism approaches and comparative genomics permits other approaches to function prediction when the data are available. These include the use of protein-protein interaction patterns, and correlations between occurrences of related proteins in different organisms, as indicators of functional properties. Even if it is possible to ascribe a particular function to a gene product, the protein may have multiple functions. A fundamental problem is that function is in many cases an ill-defined concept. In this article we review the state of the art in function prediction and describe some of the underlying difficulties and successes.

Amino Acid Sequence↗

Investigation of protein functions through data-mining on integrated human transcriptome database, H-Invitational database (H-InvDB).

H-Invitational Database (H-InvDB; ) is a human transcriptome database, containing integrative annotation of 41,118 full-length cDNA clones originated from 21,037 loci. H-InvDB is a product of the H-Invitational project, an international collaboration to systematically and functionally validate human genes by analysis of a unique set of high quality full-length cDNA clones using automatic annotation and human curation under unified criteria. Here, 19,574 proteins encoded by these cDNAs were classified into 11,709 function-known and 7865 function-unknown hypothetical proteins by similarity with protein databases and motif prediction (InterProScan). The proportion of "hypothetical proteins" in H-InvDB was as high as 40.4%. In this study, we thus conducted data-mining in H-InvDB with the aim of assigning advanced functional annotations to those hypothetical proteins. First, by data-mining in the H-InvDB version of GTOP, we identified 337 SCOP domains within 7865 H-Inv hypothetical proteins. Second, by data-mining of predicted subcellular localization by SOSUI and TMHMM in H-InvDB, we found 1032 transmembrane proteins within H-Inv hypothetical proteins. These results clearly demonstrate that structural prediction is effective for functional annotation of proteins with unknown functions. All the data in H-InvDB are shown in two main views, the cDNA view and the Locus view, and five auxiliary databases with web-based viewers; DiseaseInfo Viewer, H-ANGEL, Clustering Viewer, G-integra and TOPO Viewer; the data also are provided as flat files and XML files. The data consists of descriptions of their gene structures, novel alternative splicing isoforms, functional RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein 3D structure, mapping of SNPs and microsatellite repeat motifs in relation with orphan diseases, gene expression profiling, and comparisons with mouse full-length cDNAs in the context of molecular evolution. This unique integrative platform for conducting in silico data-mining represents a substantial contribution to resources required for the exploration of human biology and pathology.

Amino Acid Sequence↗

Identification of transmembrane protein functions by binary topology patterns.

We propose a novel method for identifying and classifying the functions of transmembrane (TM) proteins based on their TM topology [the number of TM segments (tms), the loop length and the N-terminus location]. In this method, the TM topology is expressed as a string of '0' and '1', and this is designated the binary topology pattern (BTP). We focused on TM proteins with up to 12 tms, with the exception of 1 and 9 tms, and classified them into 37 functional groups by the number of tms and the functional annotation. These grouped TM protein sequences were used to determine BTPs which are specific to the individual functional groups. Since the evaluated accuracies (sensitivity, specificity and self-consistency) of these patterns in functional identification were quite high overall, i.e. 0.940, 0.934 and 0.935, respectively, as averaged over the 37 functional groups, we confirmed that TM protein function can be identified by the number of tms and the characteristics of loop lengths, i.e. BTPs.

Data Interpretation, Statistical↗

WILMA-automated annotation of protein sequences.

Large-scale annotation of sets of proteins is a frequently occurring task in association with genome sequencing projects. Here, we present an automated platform for the functional annotation of large sets of protein sequences. Various bioinformatics tools are used to achieve a comprehensive description of protein sequences and to link these results to standard Gene Ontology descriptors for molecular function, biological processes and cellular components. Access to the annotation is provided via a web-interface and database queries. These interfaces allow to formulate proteome wide queries as well as the investigation of details of individual results. WILMA annotations of the proteomes of Homo sapiens, Mus musculus, Arabidopsis thaliana and Caenorhabditis elegans are accessible at http://www.came.sbg.ac.at/wilma/

Amino Acid Sequence↗

Alternate transcription of the Toll-like receptor signaling cascade.

BACKGROUND: Alternate splicing of key signaling molecules in the Toll-like receptor (Tlr) cascade has been shown to dramatically alter the signaling capacity of inflammatory cells, but it is not known how common this mechanism is. We provide transcriptional evidence of widespread alternate splicing in the Toll-like receptor signaling pathway, derived from a systematic analysis of the FANTOM3 mouse data set. Functional annotation of variant proteins was assessed in light of inflammatory signaling in mouse primary macrophages, and the expression of each variant transcript was assessed by splicing arrays. RESULTS: A total of 256 variant transcripts were identified, including novel variants of Tlr4, Ticam1, Tollip, Rac1, Irak1, 2 and 4, Mapk14/p38, Atf2 and Stat1. The expression of variant transcripts was assessed using custom-designed splicing arrays. We functionally tested the expression of Tlr4 transcripts under a range of cytokine conditions via northern and quantitative real-time polymerase chain reaction. The effects of variant Mapk14/p38 protein expression on macrophage survival were demonstrated. CONCLUSION: Members of the Toll-like receptor signaling pathway are highly alternatively spliced, producing a large number of novel proteins with the potential to functionally alter inflammatory outcomes. These variants are expressed in primary mouse macrophages in response to inflammatory mediators such as interferon-gamma and lipopolysaccharide. Our data suggest a surprisingly common role for variant proteins in diversification/repression of inflammatory signaling.

Alternative Splicing↗

Structural genomics sheds light on protein functions and remote homologs across the insect tree of life.

Protein structure bridges the sequence-function relationship, enabling deep exploration of biological processes across diverse organisms. Insects, the most diverse animal lineage, accounting for over 50% of all described animal species, provide an exceptional system for exploring sequence-structure-function relationships. Here, we reconstructed a comprehensive and well-resolved phylogeny of 4854 insects, spanning all orders. Leveraging this framework, we created an atlas of 13.29 million predicted protein structures from 824 representative species, including 11.63 million newly predicted structures. Structural clustering revealed that proteins with divergent sequences but similar structures could be effectively grouped together. Structural similarity searches against proteins with well-characterized functions yielded annotations for 7.61 million insect proteins, including up to 14% of previously unannotated proteins. We further identified 750 million remote homologs between insect proteins, many of which trace back to ancient branches of the insect phylogeny. Remarkably, despite extensive sequence divergence, cGAS-like receptors (cGLRs) were structurally conserved across all 824 insects. Experimental assays demonstrated that these structurally identified cGLRs play a crucial role in antiviral defense in the yellow fever mosquito. Our findings highlight the significance of structural genomics for understanding protein function and evolution across the tree of life.

Animals↗

The wheat (Triticum aestivum L.) leaf proteome.

The wheat leaf proteome was mapped and partially characterized to function as a comparative template for future wheat research. In total, 404 proteins were visualized, and 277 of these were selected for analysis based on reproducibility and relative quantity. Using a combination of protein and expressed sequence tag database searching, 142 proteins were putatively identified with an identification success rate of 51%. The identified proteins were grouped according to their functional annotations with the majority (40%) being involved in energy production, primary, or secondary metabolism. Only 8% of the protein identifications lacked ascertainable functional annotation. The 51% ratio of successful identification and the 8% unclear functional annotation rate are major improvements over most previous plant proteomic studies. This clearly indicates the advancement of the plant protein and nucleic acid sequence and annotation data available in the databases, and shows the enhanced feasibility of future wheat leaf proteome research.

Computational Biology↗