Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa↗

Automatic evaluation of protein sequence functional patterns.

A procedure that automatically provides an evaluation of the diagnostic ability of a protein sequence functional pattern is described. The procedure relies on the identification of the closest definable set in terms of a (protein sequence) database functional annotation to the set of database instances containing a given pattern. Assuming annotation correctness and completeness in the protein sequence database, the degree of statistical association between these sets provides an appropriate measure of the diagnostic ability of the pattern. An experimental implementation of the procedure, using the NBRF/PIR protein database, has been applied to a diverse collection of published sequence patterns. Results obtained reveal that frequently it is not possible to define (in NBRF/PIR database terminology) the set of database instances containing a given pattern, suggesting either lack of pattern diagnostic ability or protein database annotation incompleteness and/or inconsistencies.

Algorithms↗

Assessing annotation transfer for genomics: quantifying the relations between protein sequence, structure and function through traditional and probabilistic scores.

Measuring in a quantitative, statistical sense the degree to which structural and functional information can be "transferred" between pairs of related protein sequences at various levels of similarity is an essential prerequisite for robust genome annotation. To this end, we performed pairwise sequence, structure and function comparisons on approximately 30,000 pairs of protein domains with known structure and function. Our domain pairs, which are constructed according to the SCOP fold classification, range in similarity from just sharing a fold, to being nearly identical. Our results show that traditional scores for sequence and structure similarity have the same basic exponential relationship as observed previously, with structural divergence, measured in RMS, being exponentially related to sequence divergence, measured in percent identity. However, as the scale of our survey is much larger than any previous investigations, our results have greater statistical weight and precision. We have been able to express the relationship of sequence and structure similarity using more "modern scores," such as Smith-Waterman alignment scores and probabilistic P-values for both sequence and structure comparison. These modern scores address some of the problems with traditional scores, such as determining a conserved core and correcting for length dependency; they enable us to phrase the sequence-structure relationship in more precise and accurate terms. We found that the basic exponential sequence-structure relationship is very general: the same essential relationship is found in the different secondary-structure classes and is evident in all the scoring schemes. To relate function to sequence and structure we assigned various levels of functional similarity to the domain pairs, based on a simple functional classification scheme. This scheme was constructed by combining and augmenting annotations in the enzyme and fly functional classifications and comparing subsets of these to the Escherichia coli and yeast classifications. We found sigmoidal relationships between similarity in function and sequence, with clear thresholds for different levels of functional conservation. For pairs of domains that share the same fold, precise function appears to be conserved down to approximately 40 % sequence identity, whereas broad functional class is conserved to approximately 25 %. Interestingly, percent identity is more effective at quantifying functional conservation than the more modern scores (e.g. P-values). Results of all the pairwise comparisons and our combined functional classification scheme for protein structures can be accessed from a web database at http://bioinfo.mbb.yale.edu/alignCopyright 2000 Academic Press.

Animals↗

NIFAS: visual analysis of domain evolution in proteins.

MOTIVATION: Multi-domain proteins have evolved by insertions or deletions of distinct protein domains. Tracing the history of a certain domain combination can be important for functional annotation of multi-domain proteins, and for understanding the function of individual domains. In order to analyze the evolutionary history of the domains in modular proteins it is desirable to inspect a phylogenetic tree based on sequence divergence with the modular architecture of the sequences superimposed on the tree. RESULT: A Java applet, NIFAS, that integrates graphical domain schematics for each sequence in an evolutionary tree was developed. NIFAS retrieves domain information from the Pfam database and uses CLUSTAL W to calculate a tree for a given Pfam domain. The tree can be displayed with symbolic bootstrap values, and to allow the user to focus on a part of the tree, the layout can be altered by swapping nodes, changing the outgroup, and showing/collapsing subtrees. NIFAS is integrated with the Pfam database and is accessible over the internet (http://www.cgr.ki.se/Pfam). As an example, we use NIFAS to analyze the evolution of domains in Protein Kinases C.

Computer Graphics↗

High-quality protein knowledge resource: SWISS-PROT and TrEMBL.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation (such as the description of the function of a protein, its domain structure, post-translational modifications, variants, etc.), a minimal level of redundancy and a high level of integration with other databases. Together with its automatically annotated supplement TrEMBL, it provides a comprehensive and high-quality view of the current state of knowledge about proteins. Ongoing developments include the further improvement of functional and automatic annotation in the databases including evidence attribution with particular emphasis on the human, archaeal and bacterial proteomes and the provision of additional resources such as the International Protein Index (IPI) and XML format of SWISS-PROT and TrEMBL to the user community.

Amino Acid Sequence↗

Proteomics-Based Identification of the Pyroptosis-Related Biomarker PCSK9 and Its Association With the Pathogenesis of Rheumatoid Arthritis.

Rheumatoid arthritis (RA) is a common autoimmune disease, and early diagnosis is critical for effective treatment. This study aims to identify potential biomarkers related to pyroptosis through serum proteomics analysis, offering new insights for the early diagnosis of RA. We enrolled 100 participants, including 50 patients with RA and 50 healthy controls. Serum samples were collected and analyzed using high-resolution liquid chromatography-tandem mass spectrometry (LC-MS/MS) for proteomics profiling. Differential protein expression analysis and functional annotation revealed significant upregulation of pyroptosis-related proteins in the serum of patients with RA. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses, along with protein-protein interaction (PPI) network analysis, showed that these proteins are involved in inflammation and immune pathways, particularly the activation of the NOD-like receptor protein 3 (NLRP3) inflammasome. Enzyme-linked immunosorbent assay (ELISA) validation confirmed a significant increase in PCSK9 levels in patients with RA, suggesting that PCSK9 may play a key role in the pathogenesis of RA. This study provides new directions for biomarker research in RA, particularly regarding the potential involvement of the pyroptosis pathway, with significant clinical application prospects.

Humans↗

Automatic annotation for biological sequences by extraction of keywords from MEDLINE abstracts. Development of a prototype system.

We have developed a prototype for the automatic annotation of functional characteristics in protein families. The system is able to extract biological information directly from scientific literature in the form of MEDLINE abstracts. The criterion for selecting relevant keywords is the difference between their frequency in the abstracts associated with the protein family under study and its frequency in other unrelated protein families. The concept of functional information associated to protein families is the key feature of our system and gathers evolutionary information into the problem of functional annotation of biological sequences. The system has been tested in two different scenarios: first, a large set of protein families with a small number of abstract per family and second, selected protein families with large number of abstracts attached to each one. In both cases the performances are compared with annotations provided by human experts showing a clear relation between the amount of information provided to the system and the quality of the annotations. The automatic annotations are in many cases of similar quality to the ones contained in current data bases. The possibilities and difficulties to be encountered during the development of a full system for automatic annotation are discussed.

Abstracting and Indexing↗

Structural genomics sheds light on protein functions and remote homologs across the insect tree of life.

Protein structure bridges the sequence-function relationship, enabling deep exploration of biological processes across diverse organisms. Insects, the most diverse animal lineage, accounting for over 50% of all described animal species, provide an exceptional system for exploring sequence-structure-function relationships. Here, we reconstructed a comprehensive and well-resolved phylogeny of 4854 insects, spanning all orders. Leveraging this framework, we created an atlas of 13.29 million predicted protein structures from 824 representative species, including 11.63 million newly predicted structures. Structural clustering revealed that proteins with divergent sequences but similar structures could be effectively grouped together. Structural similarity searches against proteins with well-characterized functions yielded annotations for 7.61 million insect proteins, including up to 14% of previously unannotated proteins. We further identified 750 million remote homologs between insect proteins, many of which trace back to ancient branches of the insect phylogeny. Remarkably, despite extensive sequence divergence, cGAS-like receptors (cGLRs) were structurally conserved across all 824 insects. Experimental assays demonstrated that these structurally identified cGLRs play a crucial role in antiviral defense in the yellow fever mosquito. Our findings highlight the significance of structural genomics for understanding protein function and evolution across the tree of life.

Animals↗

Dictionary-driven protein annotation.

Computational methods seeking to automatically determine the properties (functional, structural, physicochemical, etc.) of a protein directly from the sequence have long been the focus of numerous research groups. With the advent of advanced sequencing methods and systems, the number of amino acid sequences that are being deposited in the public databases has been increasing steadily. This has in turn generated a renewed demand for automated approaches that can annotate individual sequences and complete genomes quickly, exhaustively and objectively. In this paper, we present one such approach that is centered around and exploits the Bio-Dictionary, a collection of amino acid patterns that completely covers the natural sequence space and can capture functional and structural signals that have been reused during evolution, within and across protein families. Our annotation approach also makes use of a weighted, position-specific scoring scheme that is unaffected by the over-representation of well-conserved proteins and protein fragments in the databases used. For a given query sequence, the method permits one to determine, in a single pass, the following: local and global similarities between the query and any protein already present in a public database; the likeness of the query to all available archaeal/ bacterial/eukaryotic/viral sequences in the database as a function of amino acid position within the query; the character of secondary structure of the query as a function of amino acid position within the query; the cytoplasmic, transmembrane or extracellular behavior of the query; the nature and position of binding domains, active sites, post-translationally modified sites, signal peptides, etc. In terms of performance, the proposed method is exhaustive, objective and allows for the rapid annotation of individual sequences and full genomes. Annotation examples are presented and discussed in Results, including individual queries and complete genomes that were released publicly after we built the Bio-Dictionary that is used in our experiments. Finally, we have computed the annotations of more than 70 complete genomes and made them available on the World Wide Web at http://cbcsrv.watson.ibm.com/Annotations/.

Algorithms↗

Functional information in SWISS-PROT: the basis for large-scale characterisation of protein sequences.

With the rapid growth of sequence databases, there is an increasing need for reliable functional characterisation and annotation of newly predicted proteins. To cope with such large data volumes, faster and more effective means of protein sequence characterisation and annotation are required. One promising approach is automatic large-scale functional characterisation and annotation, which is generated with limited human interaction. However, such an approach is heavily dependent on reliable data sources. The SWISS-PROT protein sequence database plays an essential role here owing to its high level of functional information.

Animals↗

A protein-protein interaction map of the Caenorhabditis elegans 26S proteasome.

The ubiquitin-proteasome proteolytic pathway is pivotal in most biological processes. Despite a great level of information available for the eukaryotic 26S proteasome-the protease responsible for the degradation of ubiquitylated proteins-several structural and functional questions remain unanswered. To gain more insight into the assembly and function of the metazoan 26S proteasome, a two-hybrid-based protein interaction map was generated using 30 Caenorhabditis elegans proteasome subunits. The results recapitulate interactions reported for other organisms and reveal new potential interactions both within the 19S regulatory complex and between the 19S and 20S subcomplexes. Moreover, novel potential proteasome interactors were identified, including an E3 ubiquitin ligase, transcription factors, chaperone proteins and other proteins not yet functionally annotated. By providing a wealth of novel biological hypotheses, this interaction map constitutes a framework for further analysis of the ubiquitin-proteasome pathway in a multicellular organism amenable to both classical genetics and functional genomics.

Animals↗

SUPFAM--a database of potential protein superfamily relationships derived by comparing sequence-based and structure-based families: implications for structural genomics and function annotation in genomes.

Members of a superfamily of proteins could result from divergent evolution of homologues with insignificant similarity in the amino acid sequences. A superfamily relationship is detected commonly after the three-dimensional structures of the proteins are determined using X-ray analysis or NMR. The SUPFAM database described here relates two homologous protein families in a multiple sequence alignment database of either known or unknown structure. The present release (1.1), which is the first version of the SUPFAM database, has been derived by analysing Pfam, which is one of the commonly used databases of multiple sequence alignments of homologous proteins. The first step in establishing SUPFAM is to relate Pfam families with the families in PALI, which is an alignment database of homologous proteins of known structure that is derived largely from SCOP. The second step involves relating Pfam families which could not be associated reliably with a protein superfamily of known structure. The profile matching procedure, IMPALA, has been used in these steps. The first step resulted in identification of 1280 Pfam families (out of 2697, i.e. 47%) which are related, either by close homologous connection to a SCOP family or by distant relationship to a SCOP family, potentially forming new superfamily connections. Using the profiles of 1417 Pfam families with apparently no structural information, an all-against-all comparison involving a sequence-profile match using IMPALA resulted in clustering of 67 homologous protein families of Pfam into 28 potential new superfamilies. Expansion of groups of related proteins of yet unknown structural information, as proposed in SUPFAM, should help in identifying 'priority proteins' for structure determination in structural genomics initiatives to expand the coverage of structural information in the protein sequence space. For example, we could assign 858 distinct Pfam domains in 2203 of the gene products in the genome of Mycobacterium tubercolosis. Fifty-one of these Pfam families of unknown structure could be clustered into 17 potentially new superfamilies forming good targets for structural genomics. SUPFAM database can be accessed at http://pauling.mbu.iisc.ernet.in/~supfam.

Animals↗

GXXXG and GXXXA motifs stabilize FAD and NAD(P)-binding Rossmann folds through C(alpha)-H... O hydrogen bonds and van der waals interactions.

Here we present evidence that domains in soluble proteins containing either the GXXXG or GXXXA motif are stabilized by the interaction of a beta-strand with the following alpha-helix. As an example, we characterized a beta-strand-helix interaction from the FAD or NAD(P)-binding Rossmann fold. The Rossmann fold is one of the three most highly represented folds in the Protein Data Bank (PDB). A subset of the proteins that adopt the Rossmann fold also bind to nucleotide cofactors such as FAD and NAD(P) and function as oxidoreductases. These Rossmann folds can often be identified by the short amino acid sequence motif, GX(1-2)GXXG. Here, we present evidence that in addition to this sequence motif, Rossmann folds that bind FAD and NAD(P) also typically contain either GXXXG or GXXXA motifs, where the first glycyl residue of these motifs and the third glycyl residue of the GX(1-2)GXXG motif are the same residue. These two motifs appear to stabilize the Rossmann fold: the first glycyl residue of either the GXXXG or GXXXA motif contacts the carbonyl oxygen atom from the first glycyl residue of the GX(1-2)GXXG motif consistent with the formation of a C(alpha)-H cdots, three dots, centered O hydrogen bond. In addition, both the glycyl and alanyl residues of the GXXXG or GXXXA motifs form van der Waals interactions with either a valine or isoleucine residue located either seven or eight residues further back along the polypeptide chain from the first glycine of the GXXXG or GXXXA motifs. Therefore, we combine both the GX(1-2)GXXG and GXXXG/A motifs into an extended motif, V/IXGX(1-2)GXXGXXXG/A, that is more strongly indicative than previously described motifs of Rossmann folds that bind FAD or NAD(P). The V/IXGX(1-2)GXXGXXXG/A motif can be used to search genomic sequence data and to annotate the function of proteins containing the motif as oxidoreductases, including proteins of previously unknown function.

Amino Acid Motifs↗

Whole-Genome Analysis of Bacillus Licheniformis Ali5 and Synthesis of Lichenysin via Genome Shuffling.

Whole-genome sequencing of Bacillus licheniformis Ali5 was performed via MGI-seq PE150 and Nanopore single-molecule real-time sequencing. The strain has a 4,114,664 bp circular genome encoding 4030 protein-coding genes. Functional annotation across NR, COG, GO, KEGG, CARD, BacMet, and CAZy databases identified 4025, 2812, 988, 1242, 72, 69, and 94 corresponding genes, respectively, and antiSMASH 6.0 revealed multiple antimicrobial biosynthetic gene clusters, including intact lichenysin and lichenicidin VK21 A1/A2 gene clusters. Three rounds of recursive protoplast fusion-based genome shuffling, paired with a dual-index screening system, significantly improved strain growth and lichenysin biosynthesis. Recombinants exhibited shortened lag phase, enhanced proliferation, improved stationary-phase stability, and higher diauxic peak biomass. PP3-176 and PP3-186 showed 4.6%-8.1% higher 12-h shake-flask titer and 3.1%-4.0% higher maximum titer than the parental average, with excellent fermentation stability. 1-L bioreactor validation confirmed strong scale-up potential. PP3-186 achieved 27.2% and 31.6% titer increases at 12 h and 20 h, while PP3-176 yielded 20.4% and 14.6% improvements with robust metabolic performance. This study validates genome shuffling as an effective strategy for enhancing lichenysin production, providing candidate strains and technical support for industrial application.

Bacillus licheniformis↗

Prediction of unidentified human genes on the basis of sequence similarity to novel cDNAs from cynomolgus monkey brain.

BACKGROUND: The complete assignment of the protein-coding regions of the human genome is a major challenge for genome biology today. We have already isolated many hitherto unknown full-length cDNAs as orthologs of unidentified human genes from cDNA libraries of the cynomolgus monkey (Macaca fascicularis) brain (parietal lobe and cerebellum). In this study, we used cDNA libraries of three other parts of the brain (frontal lobe, temporal lobe and medulla oblongata) to isolate novel full-length cDNAs. RESULTS: The entire sequences of novel cDNAs of the cynomolgus monkey were determined, and the orthologous human cDNA sequences were predicted from the human genome sequence. We predicted 29 novel human genes with putative coding regions sharing an open reading frame with the cynomolgus monkey, and we confirmed the expression of 21 pairs of genes by the reverse transcription-coupled polymerase chain reaction method. The hypothetical proteins were also functionally annotated by computer analysis. CONCLUSIONS: The 29 new genes had not been discovered in recent explorations for novel genes in humans, and the ab initio method failed to predict all exons. Thus, monkey cDNA is a valuable resource for the preparation of a complete human gene catalog, which will facilitate post-genomic studies.

Animals↗

Protein interaction mapping in C. elegans using proteins involved in vulval development.

Protein interaction mapping using large-scale two-hybrid analysis has been proposed as a way to functionally annotate large numbers of uncharacterized proteins predicted by complete genome sequences. This approach was examined in Caenorhabditis elegans, starting with 27 proteins involved in vulval development. The resulting map reveals both known and new potential interactions and provides a functional annotation for approximately 100 uncharacterized gene products. A protein interaction mapping project is now feasible for C. elegans on a genome-wide scale and should contribute to the understanding of molecular mechanisms in this organism and in human diseases.

Animals↗

Functional organization of the yeast proteome by systematic analysis of protein complexes.

Most cellular processes are carried out by multiprotein complexes. The identification and analysis of their components provides insight into how the ensemble of expressed proteins (proteome) is organized into functional units. We used tandem-affinity purification (TAP) and mass spectrometry in a large-scale approach to characterize multiprotein complexes in Saccharomyces cerevisiae. We processed 1,739 genes, including 1,143 human orthologues of relevance to human biology, and purified 589 protein assemblies. Bioinformatic analysis of these assemblies defined 232 distinct multiprotein complexes and proposed new cellular roles for 344 proteins, including 231 proteins with no previous functional annotation. Comparison of yeast and human complexes showed that conservation across species extends from single proteins to their molecular environment. Our analysis provides an outline of the eukaryotic proteome as a network of protein complexes at a level of organization beyond binary interactions. This higher-order map contains fundamental biological information and offers the context for a more reasoned and informed approach to drug discovery.

Cells, Cultured↗

Chromosome-level assembly and annotation of the yellow-shelled fish (Barbodes Wynaadensis).

Barbodes wynaadensis, a unique cyprinid species native to Yunnan Province in China, stands out as an allotetraploid (AABB) fish with a complex evolutionary history. Leveraging a multi-platform sequencing strategy combining MGI short-read, PacBio long-read, and Hi-C scaffolding technologies, we assembled the first chromosome-level genome for B. wynaadensis. The final assembled genome spans 1.76 Gb in length with a contig N50 of 33.53 Mb, demonstrating high assembly continuity. Hi-C scaffolding enabled the reconstruction of 50 pseudochromosomes, representing 99.94% of the total genome assembly. Genome annotation identified 46,121 protein-coding genes, with a functional annotation rate of 99.76%. Repetitive elements constituted 48.26% of the genomic sequences, including lineage-specific expansions of DNA transposons (29.26%) and LTRs (6.36%). This high-quality assembly resolves challenges in polyploid genome reconstruction and provides a critical resource for investigating Cyprinidae evolution, particularly subgenome divergence and adaptation. The dataset also enables practical applications, such as molecular marker development for population monitoring, supporting conservation efforts for this threatened endemic species amid habitat degradation in the Nujiang River basin.

Animals↗