Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Improving the precision of the structure-function relationship by considering phylogenetic context.

Understanding the relationship between protein structure and function is one of the foremost challenges in post-genomic biology. Higher conservation of structure could, in principle, allow researchers to extend current limitations of annotation. However, despite significant research in the area, a precise and quantitative relationship between biochemical function and protein structure has been elusive. Attempts to draw an unambiguous link have often been complicated by pleiotropy, variable transcriptional control, and adaptations to genomic context, all of which adversely affect simple definitions of function. In this paper, I report that integrating genomic information can be used to clarify the link between protein structure and function. First, I present a novel measure of functional proximity between protein structures (F-score). Then, using F-score and other entirely automatic methods measuring structure and phylogenetic similarity, I present a three-dimensional landscape describing their inter-relationship. The result is a "well-shaped" landscape that demonstrates the added value of considering genomic context in inferring function from structural homology. A generalization of methodology presented in this paper can be used to improve the precision of annotation of genes in current and newly sequenced genomes.

Journal Article↗

A cross-species analysis of the rodent uterotrophic program: elucidation of conserved responses and targets of estrogen signaling.

Physiological, morphological, and transcriptional alterations elicited by ethynyl estradiol in the uteri of Sprague-Dawley rats and C57BL/6 mice were assessed using comparable study designs, microarray platforms, and analysis methods to identify conserved estrogen signaling networks. Comparative analysis identified 153 orthologous gene pairs that were positively correlated, suggesting conserved transcriptional targets important in uterine proliferation. Functional annotation for these responses were associated with angiogenesis, water and solute transport, cell cycle control, redox control, DNA replication, protein synthesis and transport, xenobiotic metabolism, cell-cell communication, energetics, and cholesterol and fatty acid regulation. The identification of conserved temporal expression patterns of these orthologs provides experimental support for the transfer of functional annotation from mouse orthologs to 44 previously unannotated rat expressed sequence tags based on their homology and co-expression patterns. The identification of comparable temporal phenotypic responses linked to related gene expression profiles demonstrates the ability of systematic comparative genomic assessments to elucidate important conserved mechanisms in rodent estrogen signaling during uterine proliferation.

Animals↗

COCO-CL: hierarchical clustering of homology relations based on evolutionary correlations.

MOTIVATION: Determining orthology relations among genes across multiple genomes is an important problem in the post-genomic era. Identifying orthologous genes can not only help predict functional annotations for newly sequenced or poorly characterized genomes, but can also help predict new protein-protein interactions. Unfortunately, determining orthology relation through computational methods is not straightforward due to the presence of paralogs. Traditional approaches have relied on pairwise sequence comparisons to construct graphs, which were then partitioned into putative clusters of orthologous groups. These methods do not attempt to preserve the non-transitivity and hierarchic nature of the orthology relation. RESULTS: We propose a new method, COCO-CL, for hierarchical clustering of homology relations and identification of orthologous groups of genes. Unlike previous approaches, which are based on pairwise sequence comparisons, our method explores the correlation of evolutionary histories of individual genes in a more global context. COCO-CL can be used as a semi-independent method to delineate the orthology/paralogy relation for a refined set of homologous proteins obtained using a less-conservative clustering approach, or as a refiner that removes putative out-paralogs from clusters computed using a more inclusive approach. We analyze our clustering results manually, with support from literature and functional annotations. Since our orthology determination procedure does not employ a species tree to infer duplication events, it can be used in situations when the species tree is unknown or uncertain. CONTACT: jothi@mail.nih.gov, przytyck@mail.nih.gov SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online.

Algorithms↗

A new domain family in the superfamily of alkaline phosphatases.

During the course of our large-scale genome analysis a conserved domain, currently detectable only in the genomes of Drosophila melanogaster, Caenorhabditis elegans and Anopheles gambiae, has been identified. The function of this domain is currently unknown and no function annotation is provided for this domain in the publicly available genomic, protein family and sequence databases. The search for the homologues of this domain in the non-redundant sequence database using PSI-BLAST, resulted in identification of distant relationship between this family and the alkaline phosphatase-like superfamily, which includes families of aryl sulfatase, N-acetylgalactosomine-4-sulfatase, alkaline phosphatase and 2,3-bisphosphoglycerate-independent phosphoglycerate mutase (iPGM). The fold recognition procedures showed that this new domain could adopt a similar 3-D fold as for this superfamily. Most of the phosphatases and sulfatases of this superfamily are characterized by functional residues Ser and Cys respectively in the topologically equivalent positions. This functionally important site aligns with Ser/Thr in the members of the new family. Additionally, set of residues responsible for a metal binding site in phosphatases and sulphtases are conserved in the new family. The in-depth analysis suggests that the new family could possess phosphatase activity.

Alkaline Phosphatase↗

Zinc finger proteins and other transcription regulators as response proteins in benzo[a]pyrene exposed cells.

Proteomic analysis, which combines two-dimensional electrophoresis (2-DE) and mass spectrometry (MS), is an important approach to screen proteins responsive to specific stimuli. Benzo[a]pyrene (B[a]P), a prototype of polycyclic hydrocarbons (PAHs), is a potent procarcinogen generated from the combustion of fossil fuel and cigarette smoke. To further probe the molecular mechanism of mutagenesis and carcinogenesis, and to find potential molecular markers involved in cellular responses to B[a]P exposure, we performed proteomic analysis of whole cellular proteins in human amnion epithelial cells after B[a]P-treatment. Image visualization and statistical analysis indicated that more than 40 proteins showed significant changes following B[a]P-treatment (P < 0.05). Among them, 20 proteins existed only in the control groups, while six were only present in B[a]P-treated cells. In addition, the expression of 10 proteins increased whereas 11 decreased after B[a]P-treatment. These proteins were subjected to in-gel tryptic digestion followed by matrix-assisted laser desorption/ionization-time of flight mass spectrometry (MALDI-TOF-MS) analysis. Using peptide mass fingerprinting (PMF) to search the nrNCBI database, we identified 22 proteins. Most of these proteins have unknown functions and have not been previously connected to a response to B[a]P exposure. To further annotate the characteristics of these proteins, GOblet analysis was carried out and results indicated that they were involved in multiple biological processes including regulation of transcription, cell proliferation, cell aging and other processes. However, expression changes were noted in a number of transcription regulators, including eight zinc finger proteins as well as SNF2L1 (SWI/SNF related, matrix associated, actin dependent regulator of chromatin, subfamily a, member 1), which is closely linked to the chromatin remodeling process. These data may provide new clues to further understand the implication of these proteins in cellular responses to carcinogen exposure as well as the molecular mechanisms of B[a]P-induced mutagenesis and carcinogenesis.

Amino Acid Sequence↗

Proteomic analysis on metastasis-associated proteins of human hepatocellular carcinoma tissues.

PURPOSE: A comparative proteomic approach was used to identify and analyze proteins related to metastasis of hepatocellular carcinoma (HCC). METHODS: Proteins extracted from 12 HCC tissue specimens (six with metastases and six without) were separated by two-dimensional gel electrophoresis (2-DE). The protein spots exhibiting statistical alternations between the two groups through computerized image analysis were then identified by mass spectrometry. In addition immunohistochemistry (IHC), Western blotting and RT-PCR were performed to verify the expression of certain candidate proteins. RESULTS: 16 proteins including HSP27, S100A11, CK18 were annotated by mass spectrometry, relevant to chaperone function, cell mobility, cytoskeletal architecture, respectively. Most were previously unconnected with metastasis of HCC. Of these HSP27 was found overexpressed consistently in 2-DE patterns of all metastatic HCC tissues compared with nonmetastatic ones. IHC and Western blotting of HCC tissues confirmed this difference while RT-PCR did not. CONCLUSION: There are various proteins joined together in HCC metastasis. The overexpression of HSP27 may serve as a biomarker for early detection and therapeutic targets unique to the metastatic phenotype of HCC.

Actins↗

Identification and analysis of the mouse basic/Helix-Loop-Helix transcription factor family.

The basic/Helix-Loop-Helix (bHLH) proteins are a family of transcription factors that regulates a variety of biological processes. Based on a previously defined consensus motif, we identified the complete set of bHLH protein family from the mouse proteome databases and carried out a series of bioinformatics analysis. As results, 124 mouse bHLH proteins were identified in this study, and 28 of them were additional bHLH proteins beyond the previous report. These 124 mouse bHLH proteins were classified into groups from A to F by the nomenclature and phylogenetic analysis. Statistic analysis of the Gene Ontology annotation of these proteins showed that the bHLH proteins tend to perform functions related to cell differentiation and development. Gene function enrichment analysis among six groups illuminated that the proteins in certain group tend to have special biology functions, so that the molecular function of the uncharacterized proteins in groups could be inferred.

Amino Acid Sequence↗

Predicting the solvent accessibility of transmembrane residues from protein sequence.

In this study, we propose a novel method to predict the solvent accessible surface areas of transmembrane residues. For both transmembrane alpha-helix and beta-barrel residues, the correlation coefficients between the predicted and observed accessible surface areas are around 0.65. On the basis of predicted accessible surface areas, residues exposed to the lipid environment or buried inside a protein can be identified by using certain cutoff thresholds. We have extensively examined our approach based on different definitions of accessible surface areas and a variety of sets of control parameters. Given that experimentally determining the structures of membrane proteins is very difficult and membrane proteins are actually abundant in nature, our approach is useful for theoretically modeling membrane protein tertiary structures, particularly for modeling the assembly of transmembrane domains. This approach can be used to annotate the membrane proteins in proteomes to provide extra structural and functional information.

Cell Membrane↗

Zinc through the three domains of life.

Zinc is one of the metal ions essential for life, as it is required for the proper functioning of a large number of proteins. Despite its importance, the annotation of zinc-binding proteins in gene banks or protein domain databases still has significant room for improvement. In the present work, we compiled a list of known zinc-binding protein domains and of known zinc-binding sequence motifs (zinc-binding patterns), and then used them jointly to analyze the proteome of 57 different organisms to obtain an overview of zinc usage by archaeal, bacterial, and eukaryotic organisms. Zinc-binding proteins are an abundant fraction of these proteomes, ranging between 4% and 10%. The number of zinc-binding proteins correlates linearly with the total number of proteins encoded by the genome of an organism, but the proportionality constant of Eukaryota (8.8%) is significantly higher than that observed in Bacteria and Archaea (from 5% to 6%). Most of this enrichment is due to the larger portfolio of regulatory proteins in Eukaryota.

Animals↗

Chromosome-level genome assembly of bivalve mollusk, Xishishe Coelomactra antiquata.

Coelomactra antiquata, a significant marine economic shellfish in China, is experiencing a natural population decline due to habitat destruction and overfishing, making the restoration and conservation of its natural resources an urgent priority. This study provides a high - quality chromosome - level genome assembly for C. antiquata, created by PacBio and Hi - C sequencing and resulting in a 19 - chromosome map. The assembly encompasses a genome size of 807.31&#x2009;Mb, with a contig N50 of 17.35&#x2009;Mb and a scaffold N50 of 42.90&#x2009;Mb. A total of 28,070 protein - coding genes were identified, 25,959 of which were functionally annotated. Overall, this study offers a chromosome - level genome for C. antiquata that is highly continuous and complete, providing an indispensable resource for subsequent molecular and genetic studies of this species.

Animals↗

TRILOGY: Discovery of sequence-structure patterns across diverse proteins.

We describe a new computer program, trilogy, for the automated discovery of sequence-structure patterns in proteins. trilogy implements a pattern discovery algorithm that begins with an exhaustive analysis of flexible three-residue patterns; a subset of these patterns are selected as seeds for an extension process in which longer patterns are identified. A key feature of the method is explicit treatment of both the sequence and structure components of these motifs: each trilogy pattern is a pair consisting of a sequence pattern and a structure pattern. Matches to both these component patterns are identified independently, allowing the program to assign a significance score to each sequence-structure pattern that assesses the degree of correlation between the corresponding sequence and structure motifs. trilogy identifies several thousand high-scoring patterns that occur across protein families. These include both previously identified and potentially novel motifs. We expect that these sequence-structure patterns will be useful in predicting protein structure from sequence, annotating newly determined protein structures, and identifying novel motifs of potential functional or structural significance. Further details on 7,768 significant patterns identified by trilogy can be found at http://theory.lcs.mit.edu/trilogy.

Algorithms↗

Duplicated genes evolve slower than singletons despite the initial rate increase.

BACKGROUND: Gene duplication is an important mechanism that can lead to the emergence of new functions during evolution. The impact of duplication on the mode of gene evolution has been the subject of several theoretical and empirical comparative-genomic studies. It has been shown that, shortly after the duplication, genes seem to experience a considerable relaxation of purifying selection. RESULTS: Here we demonstrate two opposite effects of gene duplication on evolutionary rates. Sequence comparisons between paralogs show that, in accord with previous observations, a substantial acceleration in the evolution of paralogs occurs after duplication, presumably due to relaxation of purifying selection. The effect of gene duplication on evolutionary rate was also assessed by sequence comparison between orthologs that have paralogs (duplicates) and those that do not (singletons). It is shown that, in eukaryotes, duplicates, on average, evolve significantly slower than singletons. Eukaryotic ortholog evolutionary rates for duplicates are also negatively correlated with the number of paralogs per gene and the strength of selection between paralogs. A tally of annotated gene functions shows that duplicates tend to be enriched for proteins with known functions, particularly those involved in signaling and related cellular processes; by contrast, singletons include an over-abundance of poorly characterized proteins. CONCLUSIONS: These results suggest that whether or not a gene duplicate is retained by selection depends critically on the pre-existing functional utility of the protein encoded by the ancestral singleton. Duplicates of genes of a higher biological import, which are subject to strong functional constraints on the sequence, are retained relatively more often. Thus, the evolutionary trajectory of duplicated genes appears to be determined by two opposing trends, namely, the post-duplication rate acceleration and the generally slow evolutionary rate owing to the high level of functional constraints.

Animals↗

From fold predictions to function predictions: automation of functional site conservation analysis for functional genome predictions.

A database of functional sites for proteins with known structures, SITE, is constructed and used in conjunction with a simple pattern matching program SiteMatch to evaluate possible function conservation in a recently constructed database of fold predictions for Escherichia coli proteins (Rychlewski L et al., 1999, Protein Sci 8:614-624). In this and other prediction databases, fold predictions are based on algorithms that can recognize weak sequence similarities and putatively assign new proteins into already characterized protein families. It is not clear whether such sequence similarities arise from distant homologies or general similarity of physicochemical features along the sequence. Leaving aside the important question of nature of relations within fold superfamilies, it is possible to assess possible function conservation by looking at the pattern of conservation of crucial functional residues. SITE consists of a multilevel function description based on structure annotations and structure analyses. In particular, active site residues, ligand binding residues, and patterns of hydrophobic residues on the protein surface are used to describe different functional features. SiteMatch, a simple pattern matching program, is designed to check the conservation of residues involved in protein activity in alignments generated by any alignment method. Here, this procedure is used to study conservation of functional features in alignments between protein sequences from the E. coli genome and their optimal structural templates. The optimal templates were identified and alignments taken from the database of genomic structural predictions was described in a previous publication (Rychlewski L et al., 1999, Protein Sci 8:614-624). An automated assessment of function conservation is used to analyze the relation between fold and function similarity for a large number of fold predictions. For instance, it is shown that identifying low significance predictions with a high level of functional residue conservations can be used to extend the prediction sensitivity for fold prediction methods. Over 100 new fold/function predictions in this class were obtained in the E. coli genome. At the same time, about 30% of our previous fold predictions are not confirmed as function predictions, further highlighting the problem of function divergence in fold superfamilies.

Algorithms↗

Genome-scale, biochemical annotation method based on the wheat germ cell-free protein synthesis system.

Since the complete genomic DNA sequencing of various species, attention has turned to the structural properties, and functional characteristics of proteins. Current cell-free protein expression systems from eukaryotes are capable of synthesizing proteins with high speed and accuracy; however, the yields are low due to their instability over time. This report reviews the high-throughput, genome-scale biochemical annotation method based on the cell-free system prepared from wheat embryos. We first briefly reviewed our highly efficient and robust wheat germ cell-free protein synthesis system, and then showed an application of the system for materialization and characterization of genetic information taking a cDNA library of protein kinase from Arabidopsis thaliana as an example. The procedure consists of: (1) fusion of the gene-of-interest to a purification-tag, amplified by the split-primer PCR method; (2) transcription and purification of mRNA; (3) cell-free protein synthesis in the bilayer system using 96-well titer plate; (4) affinity purification and activity measurement. We took 439 cDNAs encoding kinases among 1064 genes annotated so far, and they were translated in parallel into protein. Subsequent assay revealed 207 products having autophosphorylation activity. Furthermore, seven proteins out of 26 calcium-dependent protein kinase genes tested did phosphorylate a synthetic peptide substrate in the presence of calcium ion, demonstrating that the translation products, retained their substrate specificity. The information on biochemical function of gene products accumulated should revolutionize our understanding of biology and fundamentally alter the practice of medicine and influence other industries as well.

Arabidopsis↗

Differential brain transcriptome of beta4 nAChR subunit-deficient mice: is it the effect of the null mutation or the background strain?

Studies using mice with beta4 nicotinic acetylcholine receptor (nAChR) subunit deficiency (beta4-/- mice) helped reveal the roles of this subunit in bradycardiac response to vagal stimulation, nicotine-induced seizure activity and anxiety. To identify genes that might be related to beta4-containing nAChRs activity, we compared the mRNA expression profiles of brains from beta4-/- and wild-type mice using Affymetrix U74Av2 microarray. Seventy-seven genes significantly differentiated between these two experimental groups. Of them, the two most downregulated were spastic paraplegia 21 (human) homolog (Spg21) and 6-pyruvoyl-tetrahydropterin synthase (Pts) genes. Since the targeted mutagenesis of the beta4 nAChR subunit was done by using two mouse strains, 129SvEv and C57BL/6J, it is possible that the genes closely linked to the mutated beta4 gene represent the 129SvEv allele and not the control C57BL/6J-driven allele. We examined this possibility by using public database and quantitative RT-PCR. The expression levels of Spg21 and Pts genes that, like the beta4 gene, are localized on mouse chromosome 9, as well as the expression levels of other genes located on this chromosome, were dependent on the mouse background strain. The 67 differentially expressed genes that are not located on chromosome 9 were further analyzed for overrepresented functional annotations and transcription regulatory elements compared with the entire microarray. Genes encoding for proteins involved in tyrosine phosphatase activity, calcium ion binding, cell growth and/or maintenance, and chromosome organization were overrepresented. Our data enhance the understanding of the molecular interactions involved in the beta4 nAChR subunit function. They also emphasize the need for careful interpretation of expression microarray studies done on genetically manipulated animals.

Animals↗

A high-throughput, near-saturating screen for type III effector genes from Pseudomonas syringae.

Pseudomonas syringae strains deliver variable numbers of type III effector proteins into plant cells during infection. These proteins are required for virulence, because strains incapable of delivering them are nonpathogenic. We implemented a whole-genome, high-throughput screen for identifying P. syringae type III effector genes. The screen relied on FACS and an arabinose-inducible hrpL sigma factor to automate the identification and cloning of HrpL-regulated genes. We determined whether candidate genes encode type III effector proteins by creating and testing full-length protein fusions to a reporter called Delta79AvrRpt2 that, when fused to known type III effector proteins, is translocated and elicits a hypersensitive response in leaves of Arabidopsis thaliana expressing the RPS2 plant disease resistance protein. Delta79AvrRpt2 is thus a marker for type III secretion system-dependent translocation, the most critical criterion for defining type III effector proteins. We describe our screen and the collection of type III effector proteins from two pathovars of P. syringae. This stringent functional criteria defined 29 type III proteins from P. syringae pv. tomato, and 19 from P. syringae pv. phaseolicola race 6. Our data provide full functional annotation of the hrpL-dependent type III effector suites from two sequenced P. syringae pathovars and show that type III effector protein suites are highly variable in this pathogen, presumably reflecting the evolutionary selection imposed by the various host plants.

Arabidopsis↗

Protein kinase resource: an integrated environment for phosphorylation research.

The protein kinase superfamily is an important group of enzymes controlling cellular signaling cascades. The increasing amount of available experimental data provides a foundation for deeper understanding of details of signaling systems and the underlying cellular processes. Here, we describe the Protein Kinase Resource, an integrated online service that provides access to information relevant to cell signaling and enables kinase researchers to visualize and analyze the data directly in an online environment. The data set is synchronized with Uniprot and Protein Data Bank (PDB) databases and is regularly updated and verified. Additional annotation includes interactive display of domain composition, cross-references between orthologs and functional mapping to OMIM records. The Protein Kinase Resource provides an integrated view of the protein kinase superfamily by linking data with their visual representation. Thus, human kinases can be mapped onto the human kinome tree via an interactive display. Sequence and structure data can be easily displayed using applications developed for the PKR and integrated with the website and the underlying database. Advanced search mechanisms, such as multiparameter lookup, sequence pattern, and blast search, enable fast access to the desired information, while statistics tools provide the ability to analyze the relationships among the kinases under study. The integration of data presentation and visualization implemented in the Protein Kinase Resource can be adapted by other online providers of scientific data and should become an effective way to access available experimental information.

Amino Acid Sequence↗

Systematic discovery of pathogen effector functions across human pathogens and pathways.

Pathogens deploy effector proteins to exploit host cell biology, and most effector open reading frames (ORFs) are rapidly evolving and lack functional annotation. We developed the effector ORFeome (eORFeome), a scalable functional genomics platform encompassing 3,835 effector ORFs from diverse viruses, bacteria, and parasites. High-throughput barcoded screens across nuclear factor &#x3ba;B (NF-&#x3ba;B), apoptosis, p53, cGAS-STING, and major histocompatibility complex class I (MHC class I) pathways revealed novel pathway-modulating functions for hundreds of uncharacterized eORFs, unexpected activities of known effectors, and distinct pathway-specific functions encoded by single ORFs. Illustrating the power of this approach, we identified HHV6A U14 as a p53 antagonist, HHV7 U21 as a dual-function STING antagonist and MHC-I antigen display inhibitor, and adenoviral 13.6K/i-leader protein as a de novo-evolved TAP inhibitor that suppresses MHC-I display. These results establish a general framework for systematic effector annotation, uncover new mechanisms of host-pathogen interaction across kingdoms, and highlight pathogen effectors as a versatile toolkit for rewiring and probing human cellular pathways.

Humans↗