Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Biological master games: using biologists' reasoning to guide algorithm development for integrated functional genomics.

We review some powerful new algorithms that build on the intuitive biological interpretation techniques for statistical analysis of functional genomics experiments. Although they were originally designed for transcriptomics, we argue that these algorithms are applicable to any type of -omics study (transcriptomics, proteomics, metabolomics). Rank Products (RP), a strictly non-parametric test statistic to detect differentially regulated elements (genes, proteins, metabolites) in genome-wide screens. RP is particularly powerful for noisy data and low numbers of replicates and makes full use of the availability of a large number of parallel measurements that is typical of modern large-scale experiments. Iterative Group Analysis (iGA), a statistical method that makes the transition from regulated single elements to significant classes of elements, and thus provides an automatic functional annotation of an experiment. Graph-based iGA (GiGA), an extension of iGA that combines experimental data with a broad variety of biological annotations to highlight physiologically relevant regions in a given "evidence graph" (e.g., metabolic networks, signaling pathway diagrams, protein interaction maps). The sequential application of these techniques yields an increasingly abstract interpretation of experimental data that is at the same time quantitative, statistically rigorous, and biologically significant. The results can be used either as helpful tools to guide data visualization and exploration, or as the input for downstream computational applications in a systems biology framework.

Algorithms↗

Elucidation of subfamily segregation and intramolecular coevolution of the olfactomedin-like proteins by comprehensive phylogenetic analysis and gene expression pattern assessment.

The categorization of genes by structural distinctions relevant to biological characteristics is very important for understanding of gene functions and predicting functional implications of uncharacterized genes. It was absolutely necessary to deploy an effective and efficient strategy to deal with the complexity of the large olfactomedin-like (OLF) gene family sharing sequence similarity but playing diversified roles in many important biological processes, as the simple highest-hit homology analysis gave incomprehensive results and led to inappropriate annotation for some uncharacterized OLF members. In light of evolutionary information that may facilitate the classification of the OLF family and proper association of novel OLF genes with characterized homologs, we performed phylogenetic analysis on all 116 OLF proteins currently available, including two novel members cloned by our group. The OLF family segregated into seven subfamilies and members with similar domain compositions or functional properties all fell into relevant subfamilies. Furthermore, our Northern blot analysis and previous studies revealed that the typical human OLF members in each subfamily exhibited tissue-specific expression patterns, which in turn supported the segregation of the OLF subfamilies with functional divergence. Interestingly, the phylogenetic tree topology for the OLF domains alone was almost identical with that of the full-length tree representing the unique phylogenetic feature of full-length OLF proteins and their particular domain compositions. Moreover, each of the major functional domains of OLF proteins kept the same phylogenetic feature in defining similar topology of the tree. It indicates that the OLF domain and the various domains in flanking non-OLF regions have coevolved and are likely to be functionally interdependent. Expanded by a plausible gene duplication and domain couplings scenario, the OLF family comprises seven evolutionarily and functionally distinct subfamilies, in which each member shares similar structural and functional characteristics including the composition of coevolved and interdependent domains. The phylogenetically classified and preliminarily assessed subfamily framework may greatly facilitate the studying on the OLF proteins. Furthermore, it also demonstrated a feasible and reliable strategy to categorize novel genes and predict the functional implications of uncharacterized proteins based on the comprehensive phylogenetic classification of the subfamilies and their relevance to preliminary functional characteristics.

Amino Acid Sequence↗

Functional proteomics mapping of a human signaling pathway.

Access to the human genome facilitates extensive functional proteomics studies. Here, we present an integrated approach combining large-scale protein interaction mapping, exploration of the interaction network, and cellular functional assays performed on newly identified proteins involved in a human signaling pathway. As a proof of principle, we studied the Smad signaling system, which is regulated by members of the transforming growth factor beta (TGFbeta) superfamily. We used two-hybrid screening to map Smad signaling protein-protein interactions and to establish a network of 755 interactions, involving 591 proteins, 179 of which were poorly or not annotated. The exploration of such complex interaction databases is improved by the use of PIMRider, a dedicated navigation tool accessible through the Web. The biological meaning of this network is illustrated by the presence of 18 known Smad-associated proteins. Functional assays performed in mammalian cells including siRNA knock-down experiments identified eight novel proteins involved in Smad signaling, thus validating this integrated functional proteomics approach.

Adaptor Proteins, Signal Transducing↗

A picture of gene sampling/expression in model organisms using ESTs and KOG proteins.

The expressed sequence tag (EST) is an instrument of gene discovery. When available in large numbers, ESTs may be used to estimate gene expression. We analyzed gene expression by EST sampling, using the KOG database, which includes 24,154 proteins from Arabidopsis thaliana (Ath), 17,101 from Caenorhabditis elegans (Cel), 10,517 from Drosophila melanogaster (Dme), and 26,324 from Homo sapiens (Hsa), and 178,538 ESTs for Ath, 215,200 for Cel, 261,404 for Dme, and 1,941,556 for Hsa. BLAST similarity searches were performed to assign KOG annotation to all ESTs. We determined the amount of gene sampling or expression dedicated to each KOG functional category by each model organism. We found that the 25% most-expressed genes are frequently shared among these organisms. The KOG protein classification allowed the EST sampling calculation throughout the glycolysis pathway. We calculated the KOG cluster coverage and inferred that 50 to 80 K ESTs would efficiently cover 80-85% of the KOG database clusters in a transcriptome project. Since KOG is a database biased towards housekeeping genes, this is probably the number of ESTs needed to include the more commonly expressed genes in these organisms. We also examined a still unaddressed question: what is the minimum number of ESTs that should be produced in a transcriptome project?

Animals↗

The complement of enzymatic sets in different species.

We present here a comprehensive analysis of the complement of enzymes in a large variety of species. As enzymes are a relatively conserved group there are several classification systems available that are common to all species and link a protein sequence to an enzymatic function. Enzymes are therefore an ideal functional group to study the relationship between sequence expansion, functional divergence and phenotypic changes. By using information retrieved from the well annotated SWISS-PROT database together with sequence information from a variety of fully sequenced genomes and information from the EC functional scheme we have aimed here to estimate the fraction of enzymes in genomes, to determine the extent of their functional redundancy in different domains of life and to identify functional innovations and lineage specific expansions in the metazoa lineage. We found that prokaryote and eukaryote species differ both in the fraction of enzymes in their genomes and in the pattern of expansion of their enzymatic sets. We observe an increase in functional redundancy accompanying an increase in species complexity. A quantitative assessment was performed in order to determine the degree of functional redundancy in different species. Finally, we report a massive expansion in the number of mammalian enzymes involved in signalling and degradation.

Animals↗

Global molecular and morphological effects of 24-hour chromium(VI) exposure on Shewanella oneidensis MR-1.

The biological impact of 24-h ("chronic") chromium(VI) [Cr(VI) or chromate] exposure on Shewanella oneidensis MR-1 was assessed by analyzing cellular morphology as well as genome-wide differential gene and protein expression profiles. Cells challenged aerobically with an initial chromate concentration of 0.3 mM in complex growth medium were compared to untreated control cells grown in the absence of chromate. At the 24-h time point at which cells were harvested for transcriptome and proteome analyses, no residual Cr(VI) was detected in the culture supernatant, thus suggesting the complete uptake and/or reduction of this metal by cells. In contrast to the untreated control cells, Cr(VI)-exposed cells formed apparently aseptate, nonmotile filaments that tended to aggregate. Transcriptome profiling and mass spectrometry-based proteomic characterization revealed that the principal molecular response to 24-h Cr(VI) exposure was the induction of prophage-related genes and their encoded products as well as a number of functionally undefined hypothetical genes that were located within the integrated phage regions of the MR-1 genome. In addition, genes with annotated functions in DNA metabolism, cell division, biosynthesis and degradation of the murein (peptidoglycan) sacculus, membrane response, and general environmental stress protection were upregulated, while genes encoding chemotaxis, motility, and transport/binding proteins were largely repressed under conditions of 24-h chromate treatment.

Bacterial Proteins↗

Using expressed sequence tag databases to identify ovarian genes of interest.

GenBank contains 4879 expressed sequence tags (EST) derived from four non-normalized human ovarian cDNA libraries. Of these EST, 2646 are contributors to UniGene clusters and have UniGene numbers. The EST map to 1206 distinct UniGenes. A gene expression profile was established for the human ovary by identifying the abundance of each UniGene cluster and its corresponding annotation. The most highly expressed transcripts were for proteins associated with protein synthesis (ribosomal proteins, elongation factors, thymosins, etc.). However, there are also transcripts for genes of unknown function that are ovary-specific. This ovarian gene expression profile provides useful data for the design of DNA microarrays targeted at ovarian function and highlights novel sequences that warrant further investigation.

Databases, Nucleic Acid↗

Multiple alignment of complete sequences (MACS) in the post-genomic era.

Multiple alignment, since its introduction in the early seventies, has become a cornerstone of modern molecular biology. It has traditionally been used to deduce structure / function by homology, to detect conserved motifs and in phylogenetic studies. There has recently been some renewed interest in the development of multiple alignment techniques, with current opinion moving away from a single all-encompassing algorithm to iterative and / or co-operative strategies. The exploitation of multiple alignments in genome annotation projects represents a qualitative leap in the functional analysis process, opening the way to the study of the co-evolution of validated sets of proteins and to reliable phylogenomic analysis. However, the alignment of the highly complex proteins detected by today's advanced database search methods is a daunting task. In addition, with the explosion of the sequence databases and with the establishment of numerous specialized biological databases, multiple alignment programs must evolve if they are to successfully rise to the new challenges of the post-genomic era. The way forward is clearly an integrated system bringing together sequence data, knowledge-based systems and prediction methods with their inherent unreliability. The incorporation of such heterogeneous, often non-consistent, data will require major changes to the fundamental alignment algorithms used to date. Such an integrated multiple alignment system will provide an ideal workbench for the validation, propagation and presentation of this information in a format that is concise, clear and intuitive.

Amino Acid Sequence↗

Integrated LiP-MS and quantitative proteomics reveal coordinated alterations in protein conformation and expression across tumor and peritumoral regions in hepatocellular carcinoma.

Hepatocellular carcinoma (HCC) exhibits substantial molecular heterogeneity, yet protein-level alterations beyond abundance remain insufficiently characterized. Here, we integrated limited proteolysis mass spectrometry (Lip-MS) with 4D label-free quantitative proteomics to investigate conformational accessibility and protein abundance across tumor, peritumoral-near, and peritumoral-far tissues from HCC patients. Differential LiP peptides identified by both DDA and DIA corresponded to 725, 674, and 33 differentially conformed proteins in the Tumor vs. Peritumor-far, Tumor vs. Peritumor-near, and Peritumor-near vs. Peritumor-far comparisons, respectively. Quantitative proteomics identified 405, 365, and 4 differentially expressed proteins in the corresponding comparisons. Integrated analysis identified 488 and 469 conformation-specific altered proteins (CSAPs), which showed altered conformational accessibility without significant abundance changes, and 237 and 205 conformation-expression coupled proteins (CECPs) in the two tumor-involved comparisons. LiP peptide and protein abundance changes were positively correlated, with Spearman coefficients of 0.69-0.72, and more than 99% of CECPs showed concordant directions. Among them, 169 region-conserved CECPs (rcCECPs) were predominantly associated with metabolic and redox-related pathways. Protein-protein interaction analysis identified 30 hub rcCECPs. ACLY, ALDH18A1, GMPS, and DHX9 showed increased representative LiP peptide signals and protein abundance, elevated transcript expression in HCC, and associations with poorer overall survival. Peptide mapping further localized their differential LiP signals to specific sequence regions and annotated domains. Collectively, these findings provide an integrated view of regional conformational accessibility and protein abundance alterations in HCC and identify candidate proteins for further structural and functional investigation.

Humans↗

From the inside out--processing of the Chlamydial autotransporter PmpD and its role in bacterial adhesion and activation of human host cells.

Polymorphic membrane protein (Pmp)21 otherwise known as PmpD is the longest of 21 Pmps expressed by Chlamydophila pneumoniae. Recent bioinformatical analyses annotated PmpD as belonging to a family of exported Gram-negative bacterial proteins designated autotransporters. This prediction, however, was never experimentally supported, nor was the function of PmpD known. Here, using 1D and 2D PAGE we demonstrate that PmpD is processed into two parts, N-terminal (N-pmpD), middle (M-pmpD) and presumably third, C-terminal part (C-pmpD). Based on localization of the external part on the outer membrane as shown by immunofluorescence, immuno-electron microscopy and immunoblotting combined with trypsinization, we demonstrate that N-pmpD translocates to the surface of bacteria where it non-covalently binds other components of the outer membrane. We propose that N-pmpD functions as an adhesin, as antibodies raised against N-pmpD blocked chlamydial infectivity in the epithelial cells. In addition, recombinant N-pmpD activated human monocytes in vitro by upregulating their metabolic activity and by stimulating IL-8 release in a dose-dependent manner. These results demonstrate that N-PmpD is an autotransporter component of chlamydial outer membrane, important for bacterial invasion and host inflammation.

Amino Acid Sequence↗

DIG--a system for gene annotation and functional discovery.

SUMMARY: We describe a database and information discovery system named DIG (Duke Integrated Genomics) designed to facilitate the process of gene annotation and the discovery of functional context. The DIG system collects and organizes gene annotation and functional information, and includes tools that support an understanding of genes in a functional context by providing a framework for integrating and visualizing gene expression, protein interaction and literature-based interaction networks.

Chromosome Mapping↗

Identification of open reading frames unique to a select agent: Ralstonia solanacearum race 3 biovar 2.

An 8x draft genome was obtained and annotated for Ralstonia solanacearum race 3 biovar 2 (R3B2) strain UW551, a United States Department of Agriculture Select Agent isolated from geranium. The draft UW551 genome consisted of 80,169 reads resulting in 582 contigs containing 5,925,491 base pairs, with an average 64.5% GC content. Annotation revealed a predicted 4,454 protein coding open reading frames (ORFs), 43 tRNAs, and 5 rRNAs; 2,793 (or 62%) of the ORFs had a functional assignment. The UW551 genome was compared with the published genome of R. solanacearum race 1 biovar 3 tropical tomato strain GMI1000. The two phylogenetically distinct strains were at least 71% syntenic in gene organization. Most genes encoding known pathogenicity determinants, including predicted type III secreted effectors, appeared to be common to both strains. A total of 402 unique UW551 ORFs were identified, none of which had a best hit or >45% amino acid sequence identity with any R. solanacearum predicted protein; 16 had strong (E < 10(-13)) best hits to ORFs found in other bacterial plant pathogens. Many of the 402 unique genes were clustered, including 5 found in the hrp region and 38 contiguous, potential prophage genes. Conservation of some UW551 unique genes among R3B2 strains was examined by polymerase chain reaction among a group of 58 strains from different races and biovars, resulting in the identification of genes that may be potentially useful for diagnostic detection and identification of R3B2 strains. One 22-kb region that appears to be present in GMI1000 as a result of horizontal gene transfer is absent from UW551 and encodes enzymes that likely are essential for utilization of the three sugar alcohols that distinguish biovars 3 and 4 from biovars 1 and 2.

Arginine↗

Gene functional similarity search tool (GFSST).

BACKGROUND: With the completion of the genome sequences of human, mouse, and other species and the advent of high throughput functional genomic research technologies such as biomicroarray chips, more and more genes and their products have been discovered and their functions have begun to be understood. Increasing amounts of data about genes, gene products and their functions have been stored in databases. To facilitate selection of candidate genes for gene-disease research, genetic association studies, biomarker and drug target selection, and animal models of human diseases, it is essential to have search engines that can retrieve genes by their functions from proteome databases. In recent years, the development of Gene Ontology (GO) has established structured, controlled vocabularies describing gene functions, which makes it possible to develop novel tools to search genes by functional similarity. RESULTS: By using a statistical model to measure the functional similarity of genes based on the Gene Ontology directed acyclic graph, we developed a novel Gene Functional Similarity Search Tool (GFSST) to identify genes with related functions from annotated proteome databases. This search engine lets users design their search targets by gene functions. CONCLUSION: An implementation of GFSST which works on the UniProt (Universal Protein Resource) for the human and mouse proteomes is available at GFSST Web Server. GFSST provides functions not only for similar gene retrieval but also for gene search by one or more GO terms. This represents a powerful new approach for selecting similar genes and gene products from proteome databases according to their functions.

Chromosome Mapping↗

The Pyridoxal-5'-Phosphate-Dependent Enzymes of Mycobacterium tuberculosis.

Enzymes that depend on the cofactor pyridoxal 5'-phosphate (PLP) catalyze a remarkable variety of biochemical reactions in all organisms. In particular, the genome of Mycobacterium tuberculosis, the causative agent of tuberculosis (TB), encodes 45 bona fide PLP-dependent enzymes plus a few related proteins that presumably do not have enzymic function. The large majority of the 45 enzymes have been characterized in terms of catalytic activity and structure. Several of them have been shown to be central to the bacterium's survival and pathogenicity, while some of these enzymes are targets of an extant drug (d-cycloserine). Herein, the annotated catalog of the PLP-dependent enzymes in M. tuberculosis is presented and analyzed with three main goals in mind. The first will be to assess the specific aspects of mycobacterial metabolism that rely most on PLP-dependent enzymes. A second goal will be to signal those enzymes whose function is still uncertain and whose functional characterization may help to further understand the biology of M. tuberculosis. Finally, we will examine the potential and limitations of targeting the PLP-dependent enzymes for the development of new antimycobacterial drugs.

Mycobacterium tuberculosis↗

Building an automated classification of DNA-binding protein domains.

Intensive growth in 3D structure data on DNA-protein complexes as reflected in the Protein Data Bank (PDB) demands new approaches to the annotation and characterization of these data and will lead to a new understanding of critical biological processes involving these data. These data and those from other protein structure classifications will become increasingly important for the modeling of complete proteomes. We propose a fully automated classification of DNA-binding protein domains based on existing 3D-structures from the PDB. The classification, by domain, relies on the Protein Domain Parser (PDP) and the Combinatorial Extension (CE) algorithm for structural alignment. The approach involves the analysis of 3D-interaction patterns in DNA-protein interfaces, assignment of structural domains interacting with DNA, clustering of domains based on structural similarity and DNA-interacting patterns. Comparison with existing resources on describing structural and functional classifications of DNA-binding proteins was used to validate and improve the approach proposed here. In the course of our study we defined a set of criteria and heuristics allowing us to automatically build a biologically meaningful classification and define classes of functionally related protein domains. It was shown that taking into consideration interactions between protein domains and DNA considerably improves the classification accuracy. Our approach provides a high-throughput and up-to-date annotation of DNA-binding protein families which can be found at http://spdc.sdsc.edu.

Artificial Intelligence↗

Integrated proteomic network analysis reveals PTPRC as a central hub protein orchestrating co-expression modules and metabolic dysregulation in renal carcinoma: PTPRC protein molecular action.

The occurrence of renal carcinoma is closely related to a variety of molecular mechanisms and metabolic disorders. PTPRC (protein tyrosine phosphatase receptor C), as an important regulatory protein, was studied to reveal the role of PTPRC in renal carcinoma through comprehensive proteomic network analysis, especially its core position in the coordination of co-expression modules and metabolic disorders. This study was the first to download and process multiple publicly available renal cancer transcriptome data to conduct differential gene expression analysis across datasets. Functional enrichment and disease ontology analysis were performed on the transcriptome of renal cancer, and weighted gene co-expression network (WGCNA) was constructed. The results showed that comprehensive principal component analysis revealed significant differences in the transcriptome of renal cancer, and functional annotation revealed specific pathways associated with renal cancer. WGCNA analysis identified tumor-associated co-expression modules, while multi-omics analysis further identified core regulatory networks including PTPRC. As a central hub protein, PTPRC plays an important coordinating role in the co-expression module and metabolic dysregulation of renal carcinoma. This discovery provides a new perspective for understanding the molecular mechanism of kidney cancer.

Humans↗

Metagenomic Analysis of Gut Microbiome of Persistent Pulmonary Hypertension of the Newborn.

Persistent pulmonary hypertension of the newborn (PPHN) is one of the most common diseases in the neonatal intensive care unit which severely affects neonatal survival. Gut microbes play an increasingly important role in human health, but there are rarely reported how gut microbiota contribute to PPHN. In our study, the metagenomic sequencing of feces from 12 PPHN's neonates and 8 controls were performed to expose the relation between neonatal gut microbes and PPHN disease. Firstly, we found that the abundance of Actinobacteria, Proteobacteria, Bacteroidetes were significantly increased in PPHN compared with controls, but the Firmicutes components was reduced. And some pathogenic strains (like Vibrio metschnikovii) were significantly enriched in the PPHN compared with controls. Secondly, functional annotation of genes found that PPHN up-regulated transmembrane transport, but down-regulated ribosome and ATP binding. Lastly, microbial metabolic pathway enrichment analysis indicated that some metabolic pathway in PPHN were conflicting and contradictory, showed that an abnormally increased metabolism, disturbed protein synthesis and genomic instability in the PPHN neonate. Our results contribute to understanding the changes in the species and function of gut microbiota in PPHN, thus providing a theoretical basis for the explanation and treatment of PPHN.

Gastrointestinal Microbiome↗

A scale of functional divergence for yeast duplicated genes revealed from analysis of the protein-protein interaction network.

BACKGROUND: Studying the evolution of the function of duplicated genes usually implies an estimation of the extent of functional conservation/divergence between duplicates from comparison of actual sequences. This only reveals the possible molecular function of genes without taking into account their cellular function(s). We took into consideration this latter dimension of gene function to approach the functional evolution of duplicated genes by analyzing the protein-protein interaction network in which their products are involved. For this, we derived a functional classification of the proteins using PRODISTIN, a bioinformatics method allowing comparison of protein function. Our work focused on the duplicated yeast genes, remnants of an ancient whole-genome duplication. RESULTS: Starting from 4,143 interactions, we analyzed 41 duplicated protein pairs with the PRODISTIN method. We showed that duplicated pairs behaved differently in the classification with respect to their interactors. The different observed behaviors allowed us to propose a functional scale of conservation/divergence for the duplicated genes, based on interaction data. By comparing our results to the functional information carried by GO annotations and sequence comparisons, we showed that the interaction network analysis reveals functional subtleties, which are not discernible by other means. Finally, we interpreted our results in terms of evolutionary scenarios. CONCLUSIONS: Our analysis might provide a new way to analyse the functional evolution of duplicated genes and constitutes the first attempt of protein function evolutionary comparisons based on protein-protein interactions.

Computational Biology↗