Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Integrated proteomic network analysis reveals PTPRC as a central hub protein orchestrating co-expression modules and metabolic dysregulation in renal carcinoma: PTPRC protein molecular action.

The occurrence of renal carcinoma is closely related to a variety of molecular mechanisms and metabolic disorders. PTPRC (protein tyrosine phosphatase receptor C), as an important regulatory protein, was studied to reveal the role of PTPRC in renal carcinoma through comprehensive proteomic network analysis, especially its core position in the coordination of co-expression modules and metabolic disorders. This study was the first to download and process multiple publicly available renal cancer transcriptome data to conduct differential gene expression analysis across datasets. Functional enrichment and disease ontology analysis were performed on the transcriptome of renal cancer, and weighted gene co-expression network (WGCNA) was constructed. The results showed that comprehensive principal component analysis revealed significant differences in the transcriptome of renal cancer, and functional annotation revealed specific pathways associated with renal cancer. WGCNA analysis identified tumor-associated co-expression modules, while multi-omics analysis further identified core regulatory networks including PTPRC. As a central hub protein, PTPRC plays an important coordinating role in the co-expression module and metabolic dysregulation of renal carcinoma. This discovery provides a new perspective for understanding the molecular mechanism of kidney cancer.

Humans↗

Metagenomic Analysis of Gut Microbiome of Persistent Pulmonary Hypertension of the Newborn.

Persistent pulmonary hypertension of the newborn (PPHN) is one of the most common diseases in the neonatal intensive care unit which severely affects neonatal survival. Gut microbes play an increasingly important role in human health, but there are rarely reported how gut microbiota contribute to PPHN. In our study, the metagenomic sequencing of feces from 12 PPHN's neonates and 8 controls were performed to expose the relation between neonatal gut microbes and PPHN disease. Firstly, we found that the abundance of Actinobacteria, Proteobacteria, Bacteroidetes were significantly increased in PPHN compared with controls, but the Firmicutes components was reduced. And some pathogenic strains (like Vibrio metschnikovii) were significantly enriched in the PPHN compared with controls. Secondly, functional annotation of genes found that PPHN up-regulated transmembrane transport, but down-regulated ribosome and ATP binding. Lastly, microbial metabolic pathway enrichment analysis indicated that some metabolic pathway in PPHN were conflicting and contradictory, showed that an abnormally increased metabolism, disturbed protein synthesis and genomic instability in the PPHN neonate. Our results contribute to understanding the changes in the species and function of gut microbiota in PPHN, thus providing a theoretical basis for the explanation and treatment of PPHN.

Gastrointestinal Microbiome↗

A scale of functional divergence for yeast duplicated genes revealed from analysis of the protein-protein interaction network.

BACKGROUND: Studying the evolution of the function of duplicated genes usually implies an estimation of the extent of functional conservation/divergence between duplicates from comparison of actual sequences. This only reveals the possible molecular function of genes without taking into account their cellular function(s). We took into consideration this latter dimension of gene function to approach the functional evolution of duplicated genes by analyzing the protein-protein interaction network in which their products are involved. For this, we derived a functional classification of the proteins using PRODISTIN, a bioinformatics method allowing comparison of protein function. Our work focused on the duplicated yeast genes, remnants of an ancient whole-genome duplication. RESULTS: Starting from 4,143 interactions, we analyzed 41 duplicated protein pairs with the PRODISTIN method. We showed that duplicated pairs behaved differently in the classification with respect to their interactors. The different observed behaviors allowed us to propose a functional scale of conservation/divergence for the duplicated genes, based on interaction data. By comparing our results to the functional information carried by GO annotations and sequence comparisons, we showed that the interaction network analysis reveals functional subtleties, which are not discernible by other means. Finally, we interpreted our results in terms of evolutionary scenarios. CONCLUSIONS: Our analysis might provide a new way to analyse the functional evolution of duplicated genes and constitutes the first attempt of protein function evolutionary comparisons based on protein-protein interactions.

Computational Biology↗

Short tandem repeats are associated with diverse mRNAs encoding membrane-targeted proteins.

Within the genomes of multicellular organisms, short tandem repeating sequences (STRs) are ubiquitous, yet usage patterns remain obscure. The repeats (AC)n and (GU)n appear frequently in the untranslated regions (UTRs) of messenger RNAs (mRNAs). To investigate STR usage patterns, we used three approaches: (1) comparisons of individual mRNA database sequences including annotations and linked references, (2) statistical analysis of complete, UTR databases and (3) study of a large gene family, the aquaporins. Among 500 (AC)n- or (GU)n-containing mRNAs, 58 (12%) had known functions. Of these, 50 (86%) encoded proteins whose activities involved membranes or lipids, including integral membrane proteins, peripheral membrane proteins, ion channels, lipid enzymes, receptors and secreted proteins. A control sequence (AU)n also occurred in mRNAs, but only 5% encoded membrane-related functions. Investigation of all reported 3' UTR sequences, demonstrated that the STR (AC)n was 9 times more common in mRNAs encoding membrane functions than in the total UTR database (P < 0.001). Similarly, (GU)n was 8 times more common in membrane-function mRNAs than in the total database (P < 0.001). These observations suggest that (AC)n and (GU)n may be UTR signals for some mRNAs encoding membrane-targeted proteins.

3' Untranslated Regions↗

Structures of phosphate and trivanadate complexes of Bacillus stearothermophilus phosphatase PhoE: structural and functional analysis in the cofactor-dependent phosphoglycerate mutase superfamily.

Bacillus stearothermophilus phosphatase PhoE is a member of the cofactor-dependent phosphoglycerate mutase superfamily possessing broad specificity phosphatase activity. Its previous structural determination in complex with glycerol revealed probable bases for its efficient hydrolysis of both large, hydrophobic, and smaller, hydrophilic substrates. Here we report two further structures of PhoE complexes, to higher resolution of diffraction, which yield a better and thorough understanding of its catalytic mechanism. The environment of the phosphate ion in the catalytic site of the first complex strongly suggests an acid-base catalytic function for Glu83. It also reveals how the C-terminal tail ordering is linked to enzyme activation on phosphate binding by a different mechanism to that seen in Escherichia coli phosphoglycerate mutase. The second complex structure with an unusual doubly covalently bound trivanadate shows how covalent modification of the phosphorylable His10 is accompanied by small structural changes, presumably to catalytic advantage. When compared with structures of related proteins in the cofactor-dependent phosphoglycerate mutase superfamily, an additional phosphate ligand, Gln22, is observed in PhoE. Functional constraints lead to the corresponding residue being conserved as Gly in fructose-2,6-bisphosphatases and Thr/Ser/Cys in phosphoglycerate mutases. A number of sequence annotation errors in databases are highlighted by this analysis. B. stearothermophilus PhoE is evolutionarily related to a group of enzymes primarily present in Gram-positive bacilli. Even within this group substrate specificity is clearly variable highlighting the difficulties of computational functional annotation in the cofactor-dependent phosphoglycerate mutase superfamily.

Amino Acid Sequence↗

An ontology for pharmaceutical ligands and its application for in silico screening and library design.

Annotation efforts in biosciences have focused in past years mainly on the annotation of genomic sequences. Only very limited effort has been put into annotation schemes for pharmaceutical ligands. Here we propose annotation schemes for the ligands of four major target classes, enzymes, G protein-coupled receptors (GPCRs), nuclear receptors (NRs), and ligand-gated ion channels (LGICs), and outline their usage for in silico screening and combinatorial library design. The proposed schemes cover ligand functionality and hierarchical levels of target classification. The classification schemes are based on those established by the EC, GPCRDB, NuclearDB, and LGICDB. The ligands of the MDL Drug Data Report (MDDR) database serve as a reference data set of known pharmacologically active compounds. All ligands were annotated according to the schemes when attribution was possible based on the activity classification provided by the reference database. The purpose of the ligand-target classification schemes is to allow annotation-based searching of the ligand database. In addition, the biological sequence information of the target is directly linkable to the ligand, hereby allowing sequence similarity-based identification of ligands of next homologous receptors. Ligands of specified levels can easily be retrieved to serve as comprehensive reference sets for cheminformatics-based similarity searches and for design of target class focused compound libraries. Retrospective in silico screening experiments within the MDDR01.1 database, searching for structures binding to dopamine D2, all dopamine receptors and all amine-binding class A GPCRs using known dopamine D2 binding compounds as a reference set, have shown that such reference sets are in particular useful for the identification of ligands binding to receptors closely related to the reference system. The potential for ligand identification drops with increasing phylogenetic distance. The analysis of the focus of a tertiary amine based combinatorial library compared to known amine binding class A GPCRs, peptide binding class A GPCRs, and LGIC ligands constitutes a second application scenario which illustrates how the focus of a combinatorial library can be treated quantitatively. The provided annotation schemes, which bridge chem- and bioinformatics by linking ligands to sequences, are expected to be of key utility for further systematic chemogenomics exploration of previously well explored target families.

Combinatorial Chemistry Techniques↗

Actin binding LIM protein 3 (abLIM3).

LIM domain proteins were demonstrated to play key roles in various biological processes such as embryonic development, cell lineage determination, and cancer differentiation. Actin binding LIM protein 1 (abLIM1) was reported to be localized in a genomic region often deleted in human cancers and suggested to be involved in axon guidance. Recently, existence of a second family member was reported, actin binding LIM protein 2. By means of computational biology and comparative genomics, we now characterized an additional, third member of the actin binding LIM protein subgroup, actin binding LIM protein 3 (abLIM3). The human mRNA sequence was previously annotated as differentially regulated in hepatoblastoma compared to normal livers. Conservation of key structural features of abLIM1 and abLIM2, four LIM domains and a VHD domain, suggested comparable biological function of abLIM3 as a linker between actin cytoskeleton and cell signaling pathways. AbLIM3 was found to be conserved in vertebrates, as orthologous sequences were characterized for mouse, fish, and frog. In addition, we report the existence of abLIM2 orthologs in fish and frog, suggesting a similar degree of evolutionary conservation. The intracellular localization of the abLIM3 protein was predicted to be nuclear by means of Reinhardt's neural network and the k-nearest neighbor algorithm. The corresponding abLIM3 gene was localized to chromosome 5q32 and spanned 119 kb, organized in 24 exons. An RT-PCR based expression profile available from the human unidentified gene-encoded (HUGE) database demonstrated highest expression for abLIM3 in heart, lung, liver, and brain/cerebellum accompanied by lower expression in multiple other tissues. Furthermore, abLIM3 was expressed in fetal liver, CNS, and spinal cord.

Amino Acid Sequence↗

On the spatial disposition of the fifth transmembrane helix and the structural integrity of the transmembrane binding site in the opioid and ORL1 G protein-coupled receptor family.

Evidence from statistical cluster analyses of a multiple sequence alignment of G protein-coupled receptor seven-helix folds supports the existence of structurally conserved transmembrane (TM) ligand binding sites in the opioid/opioid receptor-like (ORL1) and amine receptor families. Based on the expectation that functionally conserved regions in homologous proteins will display locally higher levels of sequence identity compared with global sequence similarities that pertain to the overall fold, this approach may have wider applications in functional genomics to annotate sequence data. Binding sites in models of the kappa-opioid receptor seven-helix bundle built from the rhodopsin templates of Baldwin et al. (1997) [J. Mol. Biol., 272, 144-164] and Herzyk and Hubbard (1998) [J. Mol. Biol., 281, 742-751] are compared. The Herzyk and Hubbard template is found to be in better accord with experimental studies of amine, opioid and rhodopsin receptors owing to the reduced physical separation of the extracellular parts of TM helices V and VI and differences in the rotational orientation of the N-terminal of helix V that reveal side chain accessibilities in the Baldwin et al. structure to be out of phase with relative alkylation rates of engineered cysteine residues in the TM binding site of the alpha(2A)-adrenergic receptor. TM helix V in the Baldwin et al. template has been remodelled with a different proline kink to satisfy experimental constraints. A recent proposal that rotation of helix V is associated with receptor activation is critically discussed.

Binding Sites↗

Comparative serial analysis of gene expression of transcript profiles of tomato roots infected with cyst nematode.

We analyzed global transcripts for tomato roots infected with the cyst nematode Globodera rostochiensis using serial analysis of gene expression (SAGE). SAGE libraries were made from nematode-infected roots and uninfected roots at 14 days after inoculation, and the clones including SAGE tags were sequenced. Genes were identified by matching the SAGE tags to tomato expressed sequence tags and cDNA databases. We then compiled a list of numerous genes according to the mRNA levels that were altered after cyst nematode infection. Our SAGE results showed significant changes in expression of many unreported genes involved in nematode infection. Of these, for discussion we selected five SAGE tags of RSI-1, BURP domain-containing protein, hexose transporter, P-rich protein, and PHAP2A that were activated by cyst nematode infection. Over 20% of the tags that were upregulated in the infected root have unknown functions (non-annotated), suggesting that we can obtain information on previously unreported and uncharacterized genes by SAGE. We can also obtain information on previously reported genes involved in nematode infection (e.g., multicystatin, peroxidase, catalase, pectin esterase, and S-adenosylmethionine transferase). To evaluate the validity of our SAGE results, seven genes were further analyzed by semiquantitative reverse transcriptase-polymerase chain reaction and Northern blot hybridization; the results agreed well with the SAGE data.

Animals↗

Modeling the evolution of protein domain architectures using maximum parsimony.

Domains are basic evolutionary units of proteins and most proteins have more than one domain. Advances in domain modeling and collection are making it possible to annotate a large fraction of known protein sequences by a linear ordering of their domains, yielding their architecture. Protein domain architectures link evolutionarily related proteins and underscore their shared functions. Here, we attempt to better understand this association by identifying the evolutionary pathways by which extant architectures may have evolved. We propose a model of evolution in which architectures arise through rearrangements of inferred precursor architectures and acquisition of new domains. These pathways are ranked using a parsimony principle, whereby scenarios requiring the fewest number of independent recombination events, namely fission and fusion operations, are assumed to be more likely. Using a data set of domain architectures present in 159 proteomes that represent all three major branches of the tree of life allows us to estimate the history of over 85% of all architectures in the sequence database. We find that the distribution of rearrangement classes is robust with respect to alternative parsimony rules for inferring the presence of precursor architectures in ancestral species. Analyzing the most parsimonious pathways, we find 87% of architectures to gain complexity over time through simple changes, among which fusion events account for 5.6 times as many architectures as fission. Our results may be used to compute domain architecture similarities, for example, based on the number of historical recombination events separating them. Domain architecture "neighbors" identified in this way may lead to new insights about the evolution of protein function.

Cluster Analysis↗

Complete genome sequencing of Anaplasma marginale reveals that the surface is skewed to two superfamilies of outer membrane proteins.

The rickettsia Anaplasma marginale is the most prevalent tick-borne livestock pathogen worldwide and is a severe constraint to animal health. A. marginale establishes lifelong persistence in infected ruminants and these animals serve as a reservoir for ticks to acquire and transmit the pathogen. Within the mammalian host, A. marginale generates antigenic variants by changing a surface coat composed of numerous proteins. By sequencing and annotating the complete 1,197,687-bp genome of the St. Maries strain of A. marginale, we show that this surface coat is dominated by two families containing immunodominant proteins: the msp2 superfamily and the msp1 superfamily. Of the 949 annotated coding sequences, just 62 are predicted to be outer membrane proteins, and of these, 49 belong to one of these two superfamilies. The genome contains unusual functional pseudogenes that belong to the msp2 superfamily and play an integral role in surface coat antigenic variation, and are thus distinctly different from pseudogenes described as byproducts of reductive evolution in other Rickettsiales.

Anaplasma marginale↗

Functional bioinformatics for Arabidopsis thaliana.

MOTIVATION: The genome of Arabidopsis thaliana, which has the best understood plant genome, still has approximately one-third of its genes with no functional annotation at all from either MIPS or TAIR. We have applied our Data Mining Prediction (DMP) method to the problem of predicting the functional classes of these protein sequences. This method is based on using a hybrid machine-learning/data-mining method to identify patterns in the bioinformatic data about sequences that are predictive of function. We use data about sequence, predicted secondary structure, predicted structural domain, InterPro patterns, sequence similarity profile and expressions data. RESULTS: We predicted the functional class of a high percentage of the Arabidopsis genes with currently unknown function. These predictions are interpretable and have good test accuracies. We describe in detail seven of the rules produced.

Algorithms↗

Recent additions and improvements to the Onto-Tools.

The Onto-Tools suite is composed of an annotation database and six seamlessly integrated, web-accessible data mining tools: Onto-Express, Onto-Compare, Onto-Design, Onto-Translate, Onto-Miner and Pathway-Express. The Onto-Tools database has been expanded to include various types of data from 12 new databases. Our database now integrates different types of genomic data from 19 sequence, gene, protein and annotation databases. Additionally, our database is also expanded to include complete Gene Ontology (GO) annotations. Using the enhanced database and GO annotations, Onto-Express now allows functional profiling for 24 organisms and supports 17 different types of input IDs. Onto-Translate is also enhanced to fully utilize the capabilities of the new Onto-Tools database with an ultimate goal of providing the users with a non-redundant and complete mapping from any type of identification system to any other type. Currently, Onto-Translate allows arbitrary mappings between 29 types of IDs. Pathway-Express is a new tool that helps the users find the most interesting pathways for their input list of genes. Onto-Tools are freely available at http://vortex.cs.wayne.edu/Projects.html.

Animals↗

The MitoDrome database annotates and compares the OXPHOS nuclear genes of Drosophila melanogaster, Drosophila pseudoobscura and Anopheles gambiae.

The oxidative phosphorylation (OXPHOS) is the primary energy-producing process of all aerobic organisms and the only cellular function under the dual control of both the mitochondrial and the nuclear genomes. Functional characterization and evolutionary study of the OXPHOS system is of great importance for the understanding of many as yet unclear aspects of nucleus-mitochondrion genomic co-evolution and co-regulation gene networks. The MitoDrome database is a web-based database which provides genomic annotations about nuclear genes of Drosophila melanogaster encoding for mitochondrial proteins. Recently, MitoDrome has included a new section annotating genomic information about OXPHOS genes in Drosophila pseudoobscura and Anopheles gambiae and their comparative analysis with their Drosophila melanogaster and human counterparts. The introduction of this new comparative annotation section into MitoDrome is expected to be a useful resource for both functional and structural genomics related to the OXPHOS system.

Animals↗

Mass spectrometry of the M. smegmatis proteome: protein expression levels correlate with function, operons, and codon bias.

The fast-growing bacterium Mycobacterium smegmatis is a model mycobacterial system, a nonpathogenic soil bacterium that nonetheless shares many features with the pathogenic Mycobacterium tuberculosis, the causative agent of tuberculosis. The study of M. smegmatis is expected to shed light on mechanisms of mycobacterial growth and complex lipid metabolism, and provides a tractable system for antimycobacterial drug development. Although the M. smegmatis genome sequence is not yet completed, we used multidimensional chromatography and tandem mass spectrometry, in combination with the partially completed genome sequence, to detect and identify a total of 901 distinct proteins from M. smegmatis over the course of 25 growth conditions, providing experimental annotation for many predicted genes with an approximately 5% false-positive identification rate. We observed numerous proteins involved in energy production (9.8% of expressed proteins), protein translation (8.7%), and lipid biosynthesis (5.4%); 33% of the 901 proteins are of unknown function. Protein expression levels were estimated from the number of observations of each protein, allowing measurement of differential expression of complete operons, and the comparison of the stationary and exponential phase proteomes. Expression levels are correlated with proteins' codon biases and mRNA expression levels, as measured by comparison with codon adaptation indices, principle component analysis of codon frequencies, and DNA microarray data. This observation is consistent with notions that either (1) prokaryotic protein expression levels are largely preset by codon choice, or (2) codon choice is optimized for consistency with average expression levels regardless of the mechanism of regulating expression.

Bacterial Proteins↗

TcRRMs and Tcp28 genes are intercalated and differentially expressed in Trypanosoma cruzi life cycle.

The identification and characterization of RNA binding proteins in Trypanosoma cruzi are particularly relevant as they play key roles in the regulatory mechanisms of gene expression. In this work, we have identified coding sequences for the proteins, named TcRRM1 and TcRRM2, in the EST database generated by the T. cruzi genomic initiative. TcRRM1 and TcRRM2 contain two RNA binding domains (RRM) and are very similar to two Trypanosoma brucei RNA binding proteins previously reported, Tbp34 and Tbp37, and to a not yet annotated ORF in Leishmania major genome project. The T. cruzi RRM genes are organized in tandem, alternating with copies of Tcp28, a gene of unknown function. However, TcRRM transcript accumulation is higher in the spheromastigote stage, while Tcp28 transcripts accumulate more in the trypomastigote stage suggesting developmental regulation.

Amino Acid Sequence↗

Go molecular function terms are predictive of subcellular localization.

A protein's function is closely linked to its subcellular localization. Use of Gene Ontology (GO) molecular function terms to extend sequence-based subcellular localization prediction has been previously shown to improve predictive performance. Here, we explore directly the relationship between GO function annotations and localization information, identifying both highly predictive single terms, and terms with large information gain with respect to location. The results identify a number of predictive and informative GO terms with respect to subcellular location, particularly nucleus, extracellular space, membrane, mitochondrion, endoplasmic reticulum and Golgi. There are several clear examples illustrating why the addition of function information provides additional predictive power over sequence alone. Other interesting phenomena can also be seen in the results. Most predictive or informative terms are imperfect, and incorrect prediction may often call out significant biological phenomena. Finally, these results may be useful in the GO annotation process.

Amino Acid Sequence↗

An XRE-type regulator in Streptococcus mutans plays an important role in brpA expression and oxidative stress tolerance response.

This study used a functional genomics approach to explore the role of a xenobiotic response element (XRE)-type regulator (SMU.405c) in Streptococcus mutans physiology, including the expression of biofilm regulatory protein BrpA. Results showed that deletional mutation of xre significantly reduced the ability of the deficient mutant to grow in the presence of methyl viologen, a commonly used oxidative stressor (P < 0.001). When challenged in a hydrogen peroxide killing assay, the survival rate of the &#x2206;xre mutant was >2-log less than the parent strain after 60 min (P < 0.001). Luciferase reporter fusion assays showed that xre deficiency had no significant effect on luciferase expression when it was under the control of the intact brpA promoter, but the reporter activity increased by >6-fold (P < 0.001) when the reporter gene was fused to a brpA promoter derivative with deletion of a putative XRE-binding box. Electrophoretic mobility shift assay (EMSA) showed that recombinant XRE interacted with the brpA promoter, resulting in an electrophoretic shift of the promoter probes. In vitro transcription assay also showed that inclusion of XRE caused transcription to fall off, significantly reducing full-length brpA transcripts. RNA-seq analysis revealed that deficiency of XRE led to altered expression of >102 genes by >2-fold (P < 0.05), including 28 with increased expression, and 74 with decreased expression. Among the down-regulated were genes for DNA repair and oxidative stress tolerance response. These results suggest that XRE (SMU.405c) in S. mutans plays an important role in brpA expression and oxidative stress tolerance response.IMPORTANCEStreptococcus mutans, a keystone pathogen in human dental caries, primarily lives in the highly diverse microbiota on tooth surfaces, where the conditions are often harsh and fluctuate frequently. Locus SMU.405c was annotated to encode a xenobiotic response element (XRE)-like transcriptional regulator, but no information is available concerning the role of this protein in S. mutans pathophysiology. This study used a functional genomics approach along with molecular and transcriptomic analysis to characterize a deletional xre mutant, and the results showed that xre deficiency in S. mutans resulted in weakened oxidative stress tolerance response and alterations in transcription of >102 genes, including those known to play an important role in cell envelope biogenesis and stress tolerance response. Reporter fusion assay, electrophoretic mobility shift assay (EMSA), and in vitro transcription further demonstrated that the XRE-like regulator encoded by SMU.405c is a repressor of brpA expression and plays an important role in oxidative stress tolerance response.

Streptococcus mutans↗