Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Transcriptional coordination of the metabolic network in Arabidopsis.

Patterns of coexpression can reveal networks of functionally related genes and provide deeper understanding of processes requiring multiple gene products. We performed an analysis of coexpression networks for 1,330 genes from the AraCyc database of metabolic pathways in Arabidopsis (Arabidopsis thaliana). We found that genes associated with the same metabolic pathway are, on average, more highly coexpressed than genes from different pathways. Positively coexpressed genes within the same pathway tend to cluster close together in the pathway structure, while negatively correlated genes typically occupy more distant positions. The distribution of coexpression links per gene is highly skewed, with a small but significant number of genes having numerous coexpression partners but most having fewer than 10. Genes with multiple connections (hubs) tend to be single-copy genes, while genes with multiple paralogs are coexpressed with fewer genes, on average, than single-copy genes, suggesting that the network expands through gene duplication, followed by weakening of coexpression links involving duplicate nodes. Using a network-analysis algorithm based on coexpression with multiple pathway members (pathway-level coexpression), we identified and prioritized novel candidate pathway members, regulators, and cross pathway transcriptional control points for over 140 metabolic pathways. To facilitate exploration and analysis of the results, we provide a Web site (http://www.transvar.org/at_coexpress/analysis/web) listing analyzed pathways with links to regression and pathway-level coexpression results. These methods and results will aid in the prioritization of candidates for genetic analysis of metabolism in plants and contribute to the improvement of functional annotation of the Arabidopsis genome.

Arabidopsis↗

Database searching by flexible protein structure alignment.

We have recently developed a flexible protein structure alignment program (FATCAT) that identifies structural similarity, at the same time accounting for flexibility of protein structures. One of the most important applications of a structure alignment method is to aid in functional annotations by identifying similar structures in large structural databases. However, none of the flexible structure alignment methods were applied in this task because of a lack of significance estimation of flexible alignments. In this paper, we developed an estimate of the statistical significance of FATCAT alignment score, allowing us to use it as a database-searching tool. The results reported here show that (1) the distribution of the similarity score of FATCAT alignment between two unrelated protein structures follows the extreme value distribution (EVD), adding one more example to the current collection of EVDs of sequence and structure similarities; (2) introducing flexibility into structure comparison only slightly influences the sensitivity and specificity of identifying similar structures; and (3) the overall performance of FATCAT as a database searching tool is comparable to that of the widely used rigid-body structure comparison programs DALI and CE. Two examples illustrating the advantages of using flexible structure alignments in database searching are also presented. The conformational flexibilities that were detected in the first example may be involved with substrate specificity, and the conformational flexibilities detected in the second example may reflect the evolution of structures by block building.

Databases, Protein↗

Molecular modeling of family GH16 glycoside hydrolases: potential roles for xyloglucan transglucosylases/hydrolases in cell wall modification in the poaceae.

Family GH16 glycoside hydrolases can be assigned to five subgroups according to their substrate specificities, including xyloglucan transglucosylases/hydrolases (XTHs), (1,3)-beta-galactanases, (1,4)-beta-galactanases/kappa-carrageenases, "nonspecific" (1,3/1,3;1,4)-beta-D-glucan endohydrolases, and (1,3;1,4)-beta-D-glucan endohydrolases. A structured family GH16 glycoside hydrolase database has been constructed (http://www.ghdb.uni-stuttgart.de) and provides multiple sequence alignments with functionally annotated amino acid residues and phylogenetic trees. The database has been used for homology modeling of seven glycoside hydrolases from the GH16 family with various substrate specificities, based on structural coordinates for (1,3;1,4)-beta-D-glucan endohydrolases and a kappa-carrageenase. In combination with multiple sequence alignments, the models predict the three-dimensional (3D) dispositions of amino acid residues in the substrate-binding and catalytic sites of XTHs and (1,3/1,3;1,4)-beta-d-glucan endohydrolases; there is no structural information available in the databases for the latter group of enzymes. Models of the XTHs, compared with the recently determined structure of a Populus tremulos x tremuloides XTH, reveal similarities with the active sites of family GH11 (1,4)-beta-D-xylan endohydrolases. From a biological viewpoint, the classification, molecular modeling and a new 3D structure of the P. tremulos x tremuloides XTH establish structural and evolutionary connections between XTHs, (1,3;1,4)-beta-D-glucan endohydrolases and xylan endohydrolases. These findings raise the possibility that XTHs from higher plants could be active not only on cell wall xyloglucans, but also on (1,3;1,4)-beta-D-glucans and arabinoxylans, which are major components of walls in grasses. A role for XTHs in (1,3;1,4)-beta-D-glucan and arabinoxylan modification would be consistent with the apparent overrepresentation of XTH sequences in cereal expressed sequence tags databases.

Amino Acid Sequence↗

Rv0216, a conserved hypothetical protein from Mycobacterium tuberculosis that is essential for bacterial survival during infection, has a double hotdog fold.

The Mycobacterium tuberculosis genome contains about 4000 genes, of which approximately a third code for proteins of unknown function or are classified as conserved hypothetical proteins. We have determined the three-dimensional structure of one of these, the rv0216 gene product, which has been shown to be essential for M. tuberculosis growth in vivo. The structure exhibits the greatest similarity to bacterial and eukaryotic hydratases that catalyse the R-specific hydration of 2-enoyl coenzyme A. However, only part of the catalytic machinery is conserved in Rv0216 and it showed no activity for the substrate crotonyl-CoA. The structure of Rv0216 allows us to assign new functional annotations to a family of seven other M. tuberculosis proteins, a number if which are essential for bacterial survival during infection and growth.

Acyl Coenzyme A↗

A novel representation of protein sequences for prediction of subcellular location using support vector machines.

As the number of complete genomes rapidly increases, accurate methods to automatically predict the subcellular location of proteins are increasingly useful to help their functional annotation. In order to improve the predictive accuracy of the many prediction methods developed to date, a novel representation of protein sequences is proposed. This representation involves local compositions of amino acids and twin amino acids, and local frequencies of distance between successive (basic, hydrophobic, and other) amino acids. For calculating the local features, each sequence is split into three parts: N-terminal, middle, and C-terminal. The N-terminal part is further divided into four regions to consider ambiguity in the length and position of signal sequences. We tested this representation with support vector machines on two data sets extracted from the SWISS-PROT database. Through fivefold cross-validation tests, overall accuracies of more than 87% and 91% were obtained for eukaryotic and prokaryotic proteins, respectively. It is concluded that considering the respective features in the N-terminal, middle, and C-terminal parts is helpful to predict the subcellular location.

Amino Acids, Basic↗

Solution structure of Archaeglobus fulgidis peptidyl-tRNA hydrolase (Pth2) provides evidence for an extensive conserved family of Pth2 enzymes in archea, bacteria, and eukaryotes.

The solution structure of protein AF2095 from the thermophilic archaea Archaeglobus fulgidis, a 123-residue (13.6-kDa) protein, has been determined by NMR methods. The structure of AF2095 is comprised of four alpha-helices and a mixed beta-sheet consisting of four parallel and anti-parallel beta-strands, where the alpha-helices sandwich the beta-sheet. Sequence and structural comparison of AF2095 with proteins from Homo sapiens, Methanocaldococcus jannaschii, and Sulfolobus solfataricus reveals that AF2095 is a peptidyl-tRNA hydrolase (Pth2). This structural comparison also identifies putative catalytic residues and a tRNA interaction region for AF2095. The structure of AF2095 is also similar to the structure of protein TA0108 from archaea Thermoplasma acidophilum, which is deposited in the Protein Data Bank but not functionally annotated. The NMR structure of AF2095 has been further leveraged to obtain good-quality structural models for 55 other proteins. Although earlier studies have proposed that the Pth2 protein family is restricted to archeal and eukaryotic organisms, the similarity of the AF2095 structure to human Pth2, the conservation of key active-site residues, and the good quality of the resulting homology models demonstrate a large family of homologous Pth2 proteins that are conserved in eukaryotic, archaeal, and bacterial organisms, providing novel insights in the evolution of the Pth and Pth2 enzyme families.

Archaea↗

Analysis and prediction of functionally important sites in proteins.

The rapidly increasing volume of sequence and structure information available for proteins poses the daunting task of determining their functional importance. Computational methods can prove to be very useful in understanding and characterizing the biochemical and evolutionary information contained in this wealth of data, particularly at functionally important sites. Therefore, we perform a detailed survey of compositional and evolutionary constraints at the molecular and biological function level for a large set of known functionally important sites extracted from a wide range of protein families. We compare the degree of conservation across different functional categories and provide detailed statistical insight to decipher the varying evolutionary constraints at functionally important sites. The compositional and evolutionary information at functionally important sites has been compiled into a library of functional templates. We developed a module that predicts functionally important columns (FIC) of an alignment based on the detection of a significant "template match score" to a library template. Our template match score measures an alignment column's similarity to a library template and combines a term explicitly representing a column's residue composition with various evolutionary conservation scores (information content and position-specific scoring matrix-derived statistics). Our benchmarking studies show good sensitivity/specificity for the prediction of functional sites and high accuracy in attributing correct molecular function type to the predicted sites. This prediction method is based on information derived from homologous sequences and no structural information is required. Therefore, this method could be extremely useful for large-scale functional annotation.

Binding Sites↗

A novel clan of zinc metallopeptidases with possible intramembrane cleavage properties.

Computer-based database searching and protein multiple sequence alignment has identified a novel clan of zinc metallopeptidases, which, by phylogenetic analysis, has been shown to contain six subfamilies. The family is characterized by four common transmembrane segments and three conserved sequence motifs. The combination of topology analysis and motif identification has detected three potential Zn2+ coordinating residues. Only two of the sequences of this novel zinc metallopeptidase clan possess any functional annotation, one of which is able to cleave its substrate within a cytosol/transmembrane segment junction. A number of observations suggest that the remaining members of this novel clan may also cleave their substrates within transmembrane segments.

Animals↗

Charting protein complexes, signaling pathways, and networks in the immune system.

Systematic deciphering of protein-protein interactions has the potential to generate comprehensive and instructive signaling networks and to fuel new therapeutic and diagnostic strategies. Here, we describe how recent advances in high-throughput proteomic technologies, involving biochemical purification methods and mass spectrometry analysis, can be applied systematically to the characterization of protein complexes and the computation of molecular networks. The networks obtained form the basis for further functional analyses, such as knockdown by RNA interference, ultimately leading to the identification of nodes that represent candidate targets for pharmacological exploitation. No individual experimental approach can accurately elucidate all critical modulatory components and biological aspects of a signaling network. Such functionally annotated protein-protein interaction networks, however, represent an ideal platform for the integration of additional datasets. By providing links between molecules, they also provide links to all previous observations associated with these molecules, be they of genetic, pharmacological, or other origin. As exemplified here by the analysis of the tumor necrosis factor (TNF)-alpha/nuclear factor-kappaB (NF-kappaB) signaling pathway, the approach is applicable to any mammalian cellular signaling pathway in the immune system.

Animals↗

Dictyostelium transcriptional host cell response upon infection with Legionella.

Differential gene expression of Dictyostelium discoideum after infection with Legionella pneumophila was investigated using DNA microarrays. Investigation of a 48 h time course of infection revealed several clusters of co-regulated genes, an enrichment of preferentially up- or downregulated genes in distinct functional categories and also showed that most of the transcriptional changes occurred 24 h after infection. A detailed analysis of the 24 h time point post infection was performed in comparison to three controls, uninfected cells and co-incubation with Legionella hackeliae and L. pneumophilaDeltadotA. One hundred and thirty-one differentially expressed D. discoideum genes were identified as common to all three experiments and are thought to be involved in the pathogenic response. Functional annotation of the differentially regulated genes revealed that apart from triggering a stress response Legionella apparently not only interferes with intracellular vesicle fusion and destination but also profoundly influences and exploits the metabolism of its host. For some of the identified genes, e.g. rtoA involvement in the host response has been demonstrated in a recent study, for others such a role appears plausible. The results provide the basis for a better understanding of the complex host-pathogen interactions and for further studies on the Dictyostelium response to Legionella infection.

Animals↗

An allelic resolution gene atlas for tetraploid potato provides insights into tuberization and stress resilience.

Tubers are modified underground stems that enable asexual, clonal reproduction and serve as a mechanism for overwintering and avoidance of herbivory. Potato (Solanum tuberosum L.) is cultivated for its tubers, which serve as a major crop. Genes responsible for tuber initiation and disease resistance have been characterized in potato including StSP6A, a homolog of flowering time, that functions as a tuberigen, the equivalent of a florigen. To elucidate additional molecular and genetic mechanisms underlying potato biology including tuber initiation, tuber development, and stress responses, we generated a developmental and abiotic/biotic-stress gene expression atlas from 34 tissues and treatments of the tetraploid potato cultivar, Atlantic. Using the haplotype-phased tetraploid Atlantic genome assembly and expression abundances of 129 218 genes, we constructed gene coexpression modules that represent networks associated with distinct developmental stages as well as stress responses. Functional annotations were given to modules and used to identify genes involved in tuberization and stress resilience. Structural variation from a pan-genomic analysis across four cultivated potato genome assemblies as well as domestication and wild introgression data allowed for deeper insights into the modules to identify key genes involved in tuberization and stress responses. This study underscores the importance of transcriptional regulation in tuberization and provides a comprehensive framework for future research on potato development and improvement.

Solanum tuberosum↗

Moving forward with chemical mutagenesis in the mouse.

The study of genetic variation in mice offers a powerful experimental platform for understanding gene function. Complex trait analysis, gene-targeting and gene-trapping technologies, as well as insertional and chemical mutagenesis approaches are becoming increasingly sophisticated and provide a variety of options for cataloguing gene activities and interactions. In this review we discuss fundamental and practical concepts related to chemical mutagenesis and we highlight the growing list of strategies for performing mutagenesis screens in mice. Gene-driven and diverse types of phenotype-driven screens provide several options for the recovery of the invaluable variety of alleles generated by chemical mutagenesis. The unique advantages offered using chemical mutagenesis compare favourably to and complement the spectrum of approaches available for functional annotation of the mammalian genome.

Animals↗

Serine-threonine phosphoregulation by PknB and Stp contributes to quiescence and antibiotic tolerance in Staphylococcus aureus.

Staphylococcus aureus can cause infections that are often chronic and difficult to treat, even when the bacteria are not antibiotic resistant because most antibiotics act only on metabolically active cells. Subpopulations of persister cells are metabolically quiescent, a state associated with delayed growth, reduced protein synthesis, and increased tolerance to antibiotics. Serine-threonine kinases and phosphatases similar to those found in eukaryotes can fine-tune essential bacterial cellular processes, such as metabolism and stress signaling. We found that acid stress-mimicking conditions that S. aureus experiences in host tissues delayed growth, globally altered the serine and threonine phosphoproteome, and increased threonine phosphorylation of the activation loop of the serine-threonine protein kinase B (PknB). The deletion of stp, which encodes the only annotated functional serine-threonine phosphatase in S. aureus, increased the growth delay and phenotypic heterogeneity under different stress challenges, including growth in acidic conditions, the intracellular milieu of human cells, and abscesses in mice. This growth delay was associated with reduced protein translation and intracellular ATP concentrations and increased antibiotic tolerance. Using phosphopeptide enrichment and mass spectrometry-based proteomics, we identified targets of serine-threonine phosphorylation that may regulate bacterial growth and metabolism. Together, our findings highlight the importance of phosphoregulation in mediating bacterial quiescence and antibiotic tolerance and suggest that targeting PknB or Stp might offer a future therapeutic strategy to prevent persister formation during S. aureus infections.

Animals↗

Cellular response of Shewanella oneidensis to strontium stress.

The physiology and transcriptome dynamics of the metal ion-reducing bacterium Shewanella oneidensis strain MR-1 in response to nonradioactive strontium (Sr) exposure were investigated. Studies indicated that MR-1 was able to grow aerobically in complex medium in the presence of 180 mM SrCl2 but showed severe growth inhibition at levels above that concentration. Temporal gene expression profiles were generated from aerobically grown, mid-exponential-phase MR-1 cells shocked with 180 mM SrCl2 and analyzed for significant differences in mRNA abundance with reference to data for nonstressed MR-1 cells. Genes with annotated functions in siderophore biosynthesis and iron transport were among the most highly induced (>100-fold [P < 0.05]) open reading frames in response to acute Sr stress, and a mutant (SO3032::pKNOCK) defective in siderophore production was found to be hypersensitive to SrCl2 exposure, compared to parental and wild-type strains. Transcripts encoding multidrug and heavy metal efflux pumps, proteins involved in osmotic adaptation, sulfate ABC transporters, and assimilative sulfur metabolism enzymes also were differentially expressed following Sr exposure but at levels that were several orders of magnitude lower than those for iron transport genes. Precipitate formation was observed during aerobic growth of MR-1 in broth cultures amended with 50, 100, or 150 mM SrCl2 but not in cultures of the SO3032::pKNOCK mutant or in the abiotic control. Chemical analysis of this precipitate using laser-induced breakdown spectroscopy and static secondary ion mass spectrometry indicated extracellular solid-phase sequestration of Sr, with at least a portion of the heavy metal associated with carbonate phases.

Bacterial Proteins↗

Comparative and functional genomic analysis of prokaryotic nickel and cobalt uptake transporters: evidence for a novel group of ATP-binding cassette transporters.

The transition metals nickel and cobalt, essential components of many enzymes, are taken up by specific transport systems of several different types. We integrated in silico and in vivo methods for the analysis of various protein families containing both nickel and cobalt transport systems in prokaryotes. For functional annotation of genes, we used two comparative genomic approaches: identification of regulatory signals and analysis of the genomic positions of genes encoding candidate nickel/cobalt transporters. The nickel-responsive repressor NikR regulates many nickel uptake systems, though the NikR-binding signal is divergent in various taxonomic groups of bacteria and archaea. B(12) riboswitches regulate most of the candidate cobalt transporters in bacteria. The nickel/cobalt transporter genes are often colocalized with genes for nickel-dependent or coenzyme B(12) biosynthesis enzymes. Nickel/cobalt transporters of different families, including the previously known NiCoT, UreH, and HupE/UreJ families of secondary systems and the NikABCDE ABC-type transporters, showed a mosaic distribution in prokaryotic genomes. In silico analyses identified CbiMNQO and NikMNQO as the most widespread groups of microbial transporters for cobalt and nickel ions. These unusual uptake systems contain an ABC protein (CbiO or NikO) but lack an extracytoplasmic solute-binding protein. Experimental analysis confirmed metal transport activity for three members of this family and demonstrated significant activity for a basic module (CbiMN) of the Salmonella enterica serovar Typhimurium transporter.

ATP-Binding Cassette Transporters↗

Transcriptional profiling reveals a possible role for the timing of the inflammatory response in determining susceptibility to a viral infection.

Using a novel cDNA microarray prepared from sources of actively responding immune system cells, we have investigated the changes in gene expression in the target tissue during the early stages of infection of neonatal chickens with infectious bursal disease virus. Infections of two lines of chickens previously documented as genetically resistant and sensitive to infection were compared in order to ascertain early differences in the response to infection that might provide clues to the mechanism of differential genetic resistance. In addition to major changes that could be explained by previously described changes in infected tissue, some differences in gene expression on infection, and differences between the two chicken lines, were observed that led to a model for resistance in which a more rapid inflammatory response and more-extensive p53-related induction of apoptosis in the target B cells might limit viral replication and consequent pathology. Ironically, the effect in the asymptomatic neonatal infection is that more-severe B-cell depletion is seen in the more genetically resistant chicken. Changes of expression of many chicken genes of unknown function, indicating possible roles in the response to infection, may aid in the functional annotation of these genes.

Animals↗

Small interfering RNA screens reveal enhanced cisplatin cytotoxicity in tumor cells having both BRCA network and TP53 disruptions.

RNA interference technology allows the systematic genetic analysis of the molecular alterations in cancer cells and how these alterations affect response to therapies. Here we used small interfering RNA (siRNA) screens to identify genes that enhance the cytotoxicity (enhancers) of established anticancer chemotherapeutics. Hits identified in drug enhancer screens of cisplatin, gemcitabine, and paclitaxel were largely unique to the drug being tested and could be linked to the drug's mechanism of action. Hits identified by screening of a genome-scale siRNA library for cisplatin enhancers in TP53-deficient HeLa cells were significantly enriched for genes with annotated functions in DNA damage repair as well as poorly characterized genes likely having novel functions in this process. We followed up on a subset of the hits from the cisplatin enhancer screen and validated a number of enhancers whose products interact with BRCA1 and/or BRCA2. TP53(+/-) matched-pair cell lines were used to determine if knockdown of BRCA1, BRCA2, or validated hits that associate with BRCA1 and BRCA2 selectively enhances cisplatin cytotoxicity in TP53-deficient cells. Silencing of BRCA1, BRCA2, or BRCA1/2-associated genes enhanced cisplatin cytotoxicity approximately 4- to 7-fold more in TP53-deficient cells than in matched TP53 wild-type cells. Thus, tumor cells having disruptions in BRCA1/2 network genes and TP53 together are more sensitive to cisplatin than cells with either disruption alone.

Antineoplastic Agents↗

On the inference of parsimonious indel evolutionary scenarios.

Given a multiple alignment of orthologous DNA sequences and a phylogenetic tree for these sequences, we investigate the problem of reconstructing a most parsimonious scenario of insertions and deletions capable of explaining the gaps observed in the alignment. This problem, called the Indel Parsimony Problem, is a crucial component of the problem of ancestral genome reconstruction, and its solution provides valuable information to many genome functional annotation approaches. We first show that the problem is NP-complete. Second, we provide an algorithm, based on the fractional relaxation of an integer linear programming formulation. The algorithm is fast in practice, and the solutions it produces are, in most cases, provably optimal. We describe a divide-and-conquer approach that makes it possible to solve very large instances on a simple desktop machine, while retaining guaranteed optimality. Our algorithms are tested and shown efficient and accurate on a set of 1.8 Mb mammalian orthologous sequences in the CFTR region.

Algorithms↗