Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Expanding the human proteome with microproteins and peptideins.

A major scientific drive is to characterize the protein-coding genome, which is a primary basis for studying human health. But the fundamental question remains of what has been missed in previous analyses. Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states1-3, with major implications for biomedical science. However, a key gap in knowledge has been which ncORFs produce small microproteins or alternative protein molecules that contribute to the human proteome. Here we report the collaborative efforts of the TransCODE Consortium4 to produce a consensus landscape of protein-level evidence for ncORFs. We show that about 25% of a set of 7,264 ncORFs gives rise to detectable peptides in a large-scale analysis of 95,520 proteomics experiments. We develop an annotation framework for ncORF-encoded microproteins as human proteins and codify the new conceptual model of 'peptideins' as microproteins that have indeterminate potential as functional proteins. To probe the biological implications of peptideins, we create an evolutionary analysis approach, termed ORF relative branch length (ORBL), and determine that evolutionary constraint is common and associates with observation of ncORF-derived peptides. We then characterize a pan-essential cellular phenotype for one peptidein from the OLMALINC long non-coding RNA. Overall, we generate public research tools supported by GENCODE and PeptideAtlas and advance biomedical discovery for understudied components of the human proteome.

Humans↗

AnaGram: protein function assignment.

SUMMARY: AnaGram is a web service for protein function assignment based on identity detection of small significant fragments (protomotifs) that can act as modular pieces in peptide construction. The system is able to assign function by finding correlations between protomotifs and functional annotations contained in SWISS-PROT and Medline databases. In addition, function ontologies are used for hierarchical organization of the predicted functions. Extensive tests have been carried out to evaluate the accuracy and performance of the system. AVAILABILITY: http://jaguar.genetica.uma.es/anagram.htm

Algorithms↗

The etiological agent of Lyme disease, Borrelia burgdorferi, appears to contain only a few small RNA molecules.

Small regulatory RNAs (sRNAs) have recently been shown to be the main controllers of several regulatory pathways. The function of sRNAs depends in many cases on the RNA-binding protein Hfq, especially for sRNAs with an antisense function. In this study, the genome of Borrelia burgdorferi was subjected to different searches for sRNAs, including direct homology and comparative genomics searches and ortholog- and annotation-based search strategies. Two new sRNAs were found, one of which showed complementarity to the rpoS region, which it possibly controls by an antisense mechanism. The role of the other sRNA is unknown, although observed complementarities against particular mRNA sequences suggest an antisense mechanism. We suggest that the low level of sRNAs observed in B. burgdorferi is at least partly due to the presumed lack of both functional Hfq protein and RNase E activity.

Animals↗

Protein structure prediction for the male-specific region of the human Y chromosome.

The complete sequence of the male-specific region of the human Y chromosome (MSY) has been determined recently; however, detailed characterization for many of its encoded proteins still remains to be done. We applied state-of-the-art protein structure prediction methods to all 27 distinct MSY-encoded proteins to provide better understanding of their biological functions and their mechanisms of action at the molecular level. The results of such large-scale structure-functional annotation provide a comprehensive view of the MSY proteome, shedding light on MSY-related processes. We found that, in total, at least 60 domains are encoded by 27 distinct MSY genes, of which 42 (70%) were reliably mapped to currently known structures. The most challenging predictions include the unexpected but confident 3D structure assignments for three domains identified here encoded by the USP9Y, UTY, and BPY2 genes. The domains with unknown 3D structures that are not predictable with currently available theoretical methods are established as primary targets for crystallographic or NMR studies. The data presented here set up the basis for additional scientific discoveries in human biology of the Y chromosome, which plays a fundamental role in sex determination.

Amino Acid Sequence↗

A novel sequence similarity searching and visualization method based on overlappingly translated nucleic acids: the blastNP.

Sequence data are stored in nucleic acid and protein databases. Searching the nucleic acid databases is very specific but rather insensitive method. Searching protein databases is sensitive but not very specific procedure. It was expected that the combination of these methods might provide an optimal approach. Therefore an alternative method to TblastX has been developed, known as blastNP. Nucleic acids in database and query sequences were translated into overlapping protein-like sequences (overlappingly translated sequences or OTSs) before searching with blastP. Thus, each nucleic acid sequence is represented by a single "protein like" sequence (instead of three hypothetical proteins in different reading frames). The blastNP method is defined as a blastP that is performed on an overlappingly translated nucleic acid database using a similarly converted nucleic acid query. The specificity and sensitivity of blastNP and TblastX is very similar, however blastNP is more sensitive to detect short sequence similarities (less than 50 residues). BlastNP combines the advantages of nucleotide and protein blasts and bypasses many difficulties: (1). it is more sensitive to weak sequence similarities than blastN, (2). codon redundancy is eliminated, (3). the sensitivity to single nucleotide polymorphism, mutation and sequencing errors are reduced, (4). it is insensitive to frame shifts. This novel method was proved to find significant sequence similarities which remained hidden for other methods and is a promising tool for further understanding (and annotating) the function of many old and new sequences.

Amino Acid Sequence↗

NAVIP: Unraveling the influence of neighboring small sequence variants on functional impact prediction.

Once a suitable reference sequence has been generated, intra-species variation is often assessed by re-sequencing. Variant calling processes can reveal all differences between strains, accessions, genotypes, or individuals. These variants can be enriched with predictions about their functional implications based on available structural annotations, i.e., gene models. Although these functional impact predictions on a per-variant basis are often accurate, some challenging cases require the simultaneous incorporation of multiple adjacent variants into this prediction process. Examples include neighboring variants which modify each other's functional impact. The Neighborhood-Aware Variant Impact Predictor (NAVIP) considers all variants within a given protein coding sequence when predicting the effect. As a proof of concept, variants between the Arabidopsis thaliana accessions Columbia-0 and Niederzenz-1 were annotated. NAVIP is freely available on GitHub (https://github.com/bpucker/NAVIP) and accessible through a web server (https://pbb-tools.de).

Arabidopsis↗

Uncovering hub genes and key pathways responsive to drought stress in rice via meta-analysis of transcriptomic data.

Drought stress presents a formidable threat to global rice cultivation, triggering complex molecular responses that impact plant growth and productivity. To decipher the underlying gene expression dynamics, we performed a comprehensive meta-analysis of transcriptomic datasets derived from drought-tolerant rice genotypes. Via microarray data from three independent studies, we identified a set of consistently expressed differentially expressed genes (DEGs) under drought conditions. Integration of functional annotation tools, including GO and KEGG pathway enrichment, revealed key biological processes and signaling cascades involved in stress mitigation, such as ABA signaling, protein folding, and photosynthesis suppression. Protein-protein interaction (PPI) network construction, followed by hub gene identification via maximal clique centrality (MCC), highlighted pivotal regulators including LEA proteins, dehydrins, HSP70, and several transcription factors. Machine learning approaches further prioritize potential biomarkers, with Random Forest models achieving high classification accuracy and pinpointing key predictive genes. Chromosomal localization analysis provided spatial insights into the distribution of these hub genes, whose expression patterns were further compared against qRT-PCR data from previously published studies. This integrative approach identifies candidate genomic markers and mechanistic insights that may support future breeding strategies for drought-tolerant rice, pending experimental validation.

Cytoscape↗

Identification of 9 novel transcripts and two RGSL genes within the hereditary prostate cancer region (HPC1) at 1q25.

We applied a systematic bioinformatics approach, followed by careful manual inspection and experimental validation to identify additional expressed sequences located at the Hereditary Prostate Cancer Region (HPC1) between D1S2818 and D1S1642 on chromosome 1q25. All transcripts already described for the 1q25 region were identified and we were able to define 11 additional expressed sequences within this region (three full-length cDNA clone sequences and eight ESTs), increasing the total number of gene count in this region by 38%. Five out of the 11 expressed sequences identified were shown to be expressed in prostate tissue and thus represent novel disease gene candidates for the HPC1 region. Here, we report a detailed characterization of these five novel disease gene candidates, their expression pattern in various tissues, their genomic organization and functional annotation. Two candidates (RGSL1 and RGSL2) correspond to novel members of the RGS family, which is involved in the regulation of G-protein signaling. RGSL1 and RGLS2 expression was detected by real-time polymerase chain reaction in normal prostate tissue, but could not be detected in prostate tumor cell lines, suggesting they might have a role in prostate cancer.

Chromosome Mapping↗

Platelet-derived microvesicles induce differential gene expression in monocytic cells: a DNA microarray study.

Platelet-derived microvesicles (PMV) that are shed from the plasma membrane of activated platelets, expose various platelet-type antigens on their surface and are able to adhere to other blood cells and endothelial cells. There are several clinical conditions with markedly increased numbers of PMV, e.g. acute coronary syndrome, thrombotic microangiopathy and sepsis. To prove whether PMV may contribute to an inflammatory response we used DNA microarray technology to study the effect of PMV on gene expression in the prototypic monocytic cell line MonoMac 6 (MM6). PMV were generated by activating human platelets in plasma with collagen and subsequent removal of platelets and plasma by repeated centrifugation. MM6 were incubated for 2 h with PMV in a ratio corresponding to 75 platelets/cell, or saline as control. After RNA isolation, reverse transcription and fluorescence labelling, cDNA was hybridized on a medium density microarray comprising 5308 probes addressing 4868 transcripts of 4730 human genes relevant to inflammation, immune response and related processes. The formation of PMV-MM6 conjugates was associated with significant variations in gene expression, i.e. 93 genes were found to be differentially expressed (P < 0.001; q < 0.087). Among them, 47 genes with annotated transcripts and proteins were identified. Using Ingenuity Pathway Analysis, 37 of the differentially expressed genes were identified as parts of networks associated with functional pathways including cell-to-cell signalling, cellular growth and proliferation, regulation of gene expression and lipid metabolism. For sphingosine kinase-1 the increased expression could be confirmed exemplarily not only by RT-PCR but also on the enzyme activity level. The data indicate that PMV signal differential expression of inflammation-relevant genes in monocytic cells and may represent a novel link between hemostasis and inflammation.

Blood Platelets↗

Predicting protein function from protein/protein interaction data: a probabilistic approach.

MOTIVATION: The development of experimental methods for genome scale analysis of molecular interaction networks has made possible new approaches to inferring protein function. This paper describes a method of assigning functions based on a probabilistic analysis of graph neighborhoods in a protein-protein interaction network. The method exploits the fact that graph neighbors are more likely to share functions than nodes which are not neighbors. A binomial model of local neighbor function labeling probability is combined with a Markov random field propagation algorithm to assign function probabilities for proteins in the network. RESULTS: We applied the method to a protein-protein interaction dataset for the yeast Saccharomyces cerevisiae using the Gene Ontology (GO) terms as function labels. The method reconstructed known GO term assignments with high precision, and produced putative GO assignments to 320 proteins that currently lack GO annotation, which represents about 10% of the unlabeled proteins in S. cerevisiae.

Algorithms↗

MAGPIE/EGRET annotation of the 2.9-Mb Drosophila melanogaster Adh region.

Our challenge in annotating the 2.91-Mb Adh region of the Drosophila melanogaster genome was to identify genetic and genomic features automatically, completely, and precisely within a 6-week period. To do so, we augmented the MAGPIE microbial genome annotation system to handle eukaryotic genomic sequence data. The new configuration required the integration of eukaryotic gene-finding tools and DNA repeat tools into the automatic data collection module. It also required us to define in MAGPIE new strategies to combine data about eukaryotic exon predictions with functional data to refine the exon predictions. At the heart of the resulting new eukaryotic genome annotation system is a reverse comparison of public protein and complementary DNA sequences against the input genome to identify missing exons and to refine exon boundaries. The software modules that add eukaryotic genome annotation capability to MAGPIE are available as EGRET (Eukaryotic Genome Rapid Evaluation Tool).

Alcohol Dehydrogenase↗

A relevant in vitro eukaryotic live-cell system for the evaluation of plasmodial protein localization.

Understanding the functional genomics and proteomics of plasmodia underpins the development of new approaches to antimalarial chemotherapy. Although genome databanks (e.g. PlasmoDB) and biocomputing tools (e.g. PlasMit, PlasmoAP, PATS) are useful in providing a global albeit predictive view of the myriad of about 5000 genes, only 40% are annotated, with few cases of endorsed subcellular localizations of the corresponding proteins in animal models. Progress in plasmodial protein trafficking has been hampered by the lack of a simple yet reliable method for studying subcellular localization of plasmodial proteins. In this study, we have used a combination of fluorescent markers, organelle-specific probes, phase contrast microscopy, and confocal microscopy to locate a selection of signal peptides from 10 plasmodial proteins in CHO-K1 cells. These eukaryotic cells serve as an in vitro living system for studying the cellular destinations of four mitochondrial-targeted TCA cycle proteins (citrate synthase, CS; isocitrate dehydrogenase, ICDH; branched chain alpha-keto-acid dehydrogenase E1alpha subunit, BCKDH; succinate dehydrogenase flavoprotein-subunit, SDH), two nuclear-targeted proteins (histone deacetylase, HDAC; RNA polymerase, RPOL), two apicoplast-targeted proteins (pyruvate kinase 2, PK2; glutamate dehydrogenase, GDH), and two cytoplasmic resident proteins (malate dehydrogenase, MDH; glycerol kinase, GK). The respective localizations of these malarial proteins have complied with the selected molecular targets, viz. mitochondrial, nuclear and cytoplasmic. Interestingly, MDH that is widely known to be resident in eukaryotic mitochondria was found to be cytoplasmic, probably due to the absence of molecular target sequences. Since the localization of plasmodial proteins is central to the authentication of their pathophysiological roles, this experimental system will serve as a useful a priori approach.

3-Methyl-2-Oxobutanoate Dehydrogenase (Lipoamide)↗

Microbiome Datahub: an open-access platform integrating environmental metadata, taxonomy, and functional annotation for comprehensive metagenome-assembled genome datasets.

BACKGROUND: Metagenome-assembled genomes (MAGs) provide crucial insights into the genomic diversity of uncultured microbes. However, MAG datasets deposited in public repositories such as INSDC are often difficult to reuse due to heterogeneous quality, inconsistent taxonomic and functional annotations, and insufficiently curated environmental metadata. While secondary MAG databases such as MGnify, IMG/M, and SPIRE provide standardized resources, they reconstruct MAGs de novo from public metagenomic reads and therefore do not represent the original MAGs reported in publications. RESULTS: To address this gap, we developed Microbiome Datahub, an open-access platform that systematically aggregates and re-annotates original MAGs from INSDC. We collected 214,427 MAGs, predicted genes by DFAST, performed quality assessment with CheckM, standardized taxonomic assignments with GTDB-Tk, inferred 27 phenotypic traits using Bac2Feature, assigned proteins to MBGD ortholog clusters and KEGG Orthology IDs using PZLAST, and annotated environmental metadata with the Metagenome and Microbes Environmental Ontology. Across these MAGs, the average completeness was 80.5% and contamination 1.8%; notably, the most frequent values were&#x2009;>95% completeness and&#x2009;<1% contamination, indicating that the majority of MAGs are of high quality. Comparative analyses showed that Microbiome Datahub provides phylogenetically and environmentally diverse MAGs: while the majority originated from vertebrate gut environments, a substantial number were also recovered from other habitats such as groundwater, including nearly 10,000 MAGs from the Patescibacteria. Inference of 27 phenotypic traits, including optimum growth temperature, further revealed ecological differentiation across phyla. Protein clustering revealed 56 million identity 40% clusters, with the majority unique compared with MGnify and GlobDB, and&#x2009;~19% of proteins unassigned to MBGD ortholog clusters, underscoring their novelty. CONCLUSIONS: Microbiome Datahub integrates MAG genome sequences, gene and protein predictions, quality metrics, environmental and taxonomic annotations, ortholog cluster assignments, and phenotype predictions, all accessible via a web interface, API, and bulk downloads. By combining original MAGs with curated metadata and functional annotations, Microbiome Datahub constitutes a comprehensive and reusable resource that will accelerate microbiome and microbial genomics research. Video Abstract.

Metagenome↗

Molecular evolution of the AMP-forming Acetyl-CoA synthetase.

Acetyl-CoA-Synthetase (ACS) is involved in the production of acetate, a major metabolite in numerous organisms. There are two forms of this enzyme: ADP-forming ACS and ATP-forming ACS. We focus mainly on the AMP-forming ACS gene, which is relatively well conserved in eubacteria, archeaebacteria, and eukaryotes. BLAST searches in databases showed 30 protein sequences significantly related to the ACS. Most of these sequences were identified as ACS but three of them, belonging to the mammalian species, were annotated as another gene named: the SA gene, which is involved in the essential hypertension. The ACS and SA genes probably derived from a duplication of an ancestral gene but have acquired different functions. Six conserved regions of the ACS protein were defined across the three domains of life. While the precise function of the conserved regions remains unknown, they are probably involved in the enzymatic activity. Among eukaryotes, we found a high variability with respect to the number and the position of introns. However, some positions are conserved between fungi and a nematode. A maximum likelihood tree based upon the conserved regions showed that all sequences except the one from B. subtilis, belong to two basic groups: one the SA-like group including sequences from Archaeoglobus fulgidus and Streptomyces coelicolor, and second, the ACS group. The later can be further divided in two parts: a prokaryotic one including eubacteria and an archaebacterium, and a eukaryotic group within which two proteobacterial sequences branch including ACS from the alpha-proteobacterium Rhodobacter capsulatus. Within the eukaryotic group, bootstrap support is very low, but overall the data are consistent with the view that eukaryotes acquired their ACS gene from the ancestors of mitochondria. The localization of this enzyme in eukaryotic mitochondria is the additional evidence in favor of this interpretation.

Acetate-CoA Ligase↗

Evaluation of protein fold comparison servers.

When a new protein structure has been determined, comparison with the database of known structures enables classification of its fold as new or belonging to a known class of proteins. This in turn may provide clues about the function of the protein. A large number of fold comparison programs have been developed, but they have never been subjected to a comprehensive and critical comparative analysis. Here we describe an evaluation of 11 publicly available, Web-based servers for automatic fold comparison. Both their functionality (e.g., user interface, presentation, and annotation of results) and their performance (i.e., how well established structural similarities are recognized) were assessed. The servers were subjected to a battery of performance tests covering a broad spectrum of folds as well as special cases, such as multidomain proteins, Calpha-only models, new folds, and NMR-based models. The CATH structural classification system was used as a reference. These tests revealed the strong and weak sides of each server. On the whole, CE, DALI, MATRAS, and VAST showed the best performance, but none of the servers achieved a 100% success rate. Where no structurally similar proteins are found by any individual server, it is recommended to try one or two other servers before any conclusions concerning the novelty of a fold are put on paper.

Computational Biology↗

Cytoscape: a software environment for integrated models of biomolecular interaction networks.

Cytoscape is an open source software project for integrating biomolecular interaction networks with high-throughput expression data and other molecular states into a unified conceptual framework. Although applicable to any system of molecular components and interactions, Cytoscape is most powerful when used in conjunction with large databases of protein-protein, protein-DNA, and genetic interactions that are increasingly available for humans and model organisms. Cytoscape's software Core provides basic functionality to layout and query the network; to visually integrate the network with expression profiles, phenotypes, and other molecular states; and to link the network to databases of functional annotations. The Core is extensible through a straightforward plug-in architecture, allowing rapid development of additional computational analyses and features. Several case studies of Cytoscape plug-ins are surveyed, including a search for interaction pathways correlating with changes in gene expression, a study of protein complexes involved in cellular recovery to DNA damage, inference of a combined physical/functional interaction network for Halobacterium, and an interface to detailed stochastic/kinetic gene regulatory models.

Algorithms↗

Enzyme classification by ligand binding.

The problem of assigning a biochemical function to newly discovered proteins has been traditionally approached by expert enzymological analysis, sequence analysis, and structural modeling. In recent years, the appearance of databases containing protein-ligand interaction data for large numbers of protein classes and chemical compounds have provided new ways of investigating proteins for which the biochemical function is not completely understood. In this work, we introduce a method that utilizes ligand-binding data for functional classification of enzymes. The method makes use of the existing Enzyme Commission (EC) classification scheme and the data on interactions of small molecules with enzymes from the BRENDA database. A set of ligands that binds to an enzyme with unknown biochemical function serves as a query to search a protein-ligand interaction database for enzyme classes that are known to interact with a similar set of ligands. These classes provide hypotheses of the query enzyme's function and complement other computational annotations that take advantage of sequence and structural information. Similarity between sets of ligands is computed using point set similarity measures based upon similarity between individual compounds. We present the statistics of classification of the enzymes in the database by a cross-validation procedure and illustrate the application of the method on several examples.

5'-Nucleotidase↗

High-throughput expression, purification, and characterization of recombinant Caenorhabditis elegans proteins.

Modern proteomics approaches include techniques to examine the expression, localization, modifications, and complex formation of proteins in cells. In order to address issues of protein function in vitro using classical biochemical and biophysical approaches, high-throughput methods of cloning the appropriate reading frames, and expressing and purifying proteins efficiently are an important goal of modern proteomics approaches. This process becomes more difficult as functional proteomics efforts focus on the proteins from higher organisms, since issues of correctly identifying intron-exon boundaries and efficiently expressing and solubilizing the (often) multi-domain proteins from higher eukaryotes are challenging. Recently, 12,000 open-reading-frame (ORF) sequences from Caenorhabditis elegans have become available for functional proteomics studies [Nat. Gen. 34 (2003) 35]. We have implemented a high-throughput screening procedure to express, purify, and analyze by mass spectrometry hexa-histidine-tagged C. elegans ORFs in Escherichia coli using metal affinity ZipTips. We find that over 65% of the expressed proteins are of the correct mass as analyzed by matrix-assisted laser desorption MS. Many of the remaining proteins indicated to be "incorrect" can be explained by high-throughput cloning or genome database annotation errors. This provides a general understanding of the expected error rates in such high-throughput cloning projects. The ZipTip purified proteins can be further analyzed under both native and denaturing conditions for functional proteomics efforts.

Animals↗