Search PubMed⌕ Search

Biomedical subjects

Karsten Suhre

Publications and source records attributed to Karsten Suhre.

13 recordsLinked to original sources

Refining the Genetic Contribution to Type 2 Diabetes Subtypes.

BACKGROUND: Type 2 diabetes (T2D) is a complex and highly heterogeneous disease driven in part by genetic predisposition and can be stratified into clinical subgroups to aid disease management. We recently grouped T2D subjects in the Qatar Biobank (QBB) cohort into Severe Insulin-Deficient Diabetes (SIDD), Severe Insulin-Resistant Diabetes (SIRD), Mild Obesity-Related Diabetes (MOD) and Mild Age-Related Diabetes (MARD) subtypes. Herein, we focused on the genetic makeup of these subtypes. METHODS: We used the QBB cohort (n = 13,808), of whom 2687 were with T2D, and comprehensively assessed polygenic risk scores (PGS) across T2D subtypes, investigated genetic loci associated with each subtype by leveraging the most recent and largest GWAS for T2D, evaluated SNP associations across T2D genetic clusters, and identified protein interaction pathways associated with these distinct T2D subtypes. RESULTS: MOD showed consistently lower PGS compared with other T2D subtypes across all tested scores. SIDD showed more associations with SNPs mapping to residual glycemic cluster compared with other T2D subtypes. The incremental analysis of PGS004838 demonstrated a high ΔAUC of 0.101 for SIDD and a moderate ΔAUC of 0.068 for SIRD, but not for MOD and MARD. Protein interaction analyses identified candidate subtype-associated gene networks linked to pathways related to glucose homeostasis in SIDD, insulin signalling and hepatic metabolism in SIRD, body fat distribution in MOD and vascular-related processes in MARD. CONCLUSION: We found heterogeneous genetic architectures across clinically defined T2D subtypes in a Middle Eastern population. Our findings provide evidence supporting differential polygenic burden, subtype genetic associations and subtype-associated biological pathways across T2D subtypes. These observations support the utility of subtype-based genetic analyses for improving biological understanding of T2D heterogeneity.

Humans↗

3DCoffee: combining protein sequences and structures within multiple sequence alignments.

Most bioinformatics analyses require the assembly of a multiple sequence alignment. It has long been suspected that structural information can help to improve the quality of these alignments, yet the effect of combining sequences and structures has not been evaluated systematically. We developed 3DCoffee, a novel method for combining protein sequences and structures in order to generate high-quality multiple sequence alignments. 3DCoffee is based on TCoffee version 2.00, and uses a mixture of pairwise sequence alignments and pairwise structure comparison methods to generate multiple sequence alignments. We benchmarked 3DCoffee using a subset of HOMSTRAD, the collection of reference structural alignments. We found that combining TCoffee with the threading program Fugue makes it possible to improve the accuracy of our HOMSTRAD dataset by four percentage points when using one structure only per dataset. Using two structures yields an improvement of ten percentage points. The measures carried out on HOM39, a HOMSTRAD subset composed of distantly related sequences, show a linear correlation between multiple sequence alignment accuracy and the ratio of number of provided structure to total number of sequences. Our results suggest that in the case of distantly related sequences, a single structure may not be enough for computing an accurate multiple sequence alignment.

Protein Conformation↗

Phydbac2: improved inference of gene function using interactive phylogenomic profiling and chromosomal location analysis.

Phydbac (phylogenomic display of bacterial genes) implemented a method of phylogenomic profiling using a distance measure based on normalized BLAST scores. This method was able to increase the predictive power of phylogenomic profiling by about 25% when compared to the classical approach based on Hamming distances. Here we present a major extension of Phydbac (named here Phydbac2), that extends both the concept and the functionality of the original web-service. While phylogenomic profiles remain the central focus of Phydbac2, it now integrates chromosomal proximity and gene fusion analyses as two additional non-similarity-based indicators for inferring pairwise gene functional relationships. Moreover, all presently available (January 2004) fully sequenced bacterial genomes and those of three lower eukaryotes are now included in the profiling process, thus increasing the initial number of reference genomes (71 in Phydbac) to 150 in Phydbac2. Using the KEGG metabolic pathway database as a benchmark, we show that the predictive power of Phydbac2 is improved by 27% over the previous version. This gain is accounted for on one hand, by the increased number of reference genomes (11%) and on the other hand, as a result of including chromosomal proximity into the distance measure (16%). The expanded functionality of Phydbac2 now allows the user to query more than 50 different genomes, including at least one member of each major bacterial group, most major pathogens and potential bio-terrorism agents. The search for co-evolving genes based on consensus profiles from multiple organisms, the display of Phydbac2 profiles side by side with COG information, the inclusion of KEGG metabolic pathway maps the production of chromosomal proximity maps, and the possibility of collecting and processing results from different Phydbac queries in a common shopping cart are the main new features of Phydbac2. The Phydbac2 web server is available at http://igs-server.cnrs-mrs.fr/phydbac/.

Artificial Gene Fusion↗

ElNemo: a normal mode web server for protein movement analysis and the generation of templates for molecular replacement.

Normal mode analysis (NMA) is a powerful tool for predicting the possible movements of a given macromolecule. It has been shown recently that half of the known protein movements can be modelled by using at most two low-frequency normal modes. Applications of NMA cover wide areas of structural biology, such as the study of protein conformational changes upon ligand binding, membrane channel opening and closure, potential movements of the ribosome, and viral capsid maturation. Another, newly emerging field of NMA is related to protein structure determination by X-ray crystallography, where normal mode perturbed models are used as templates for diffraction data phasing through molecular replacement (MR). Here we present ElNémo, a web interface to the Elastic Network Model that provides a fast and simple tool to compute, visualize and analyse low-frequency normal modes of large macro-molecules and to generate a large number of different starting models for use in MR. Due to the 'rotation-translation-block' (RTB) approximation implemented in ElNémo, there is virtually no upper limit to the size of the proteins that can be treated. Upon input of a protein structure in Protein Data Bank (PDB) format, ElNémo computes its 100 lowest-frequency modes and produces a comprehensive set of descriptive parameters and visualizations, such as the degree of collectivity of movement, residue mean square displacements, distance fluctuation maps, and the correlation between observed and normal-mode-derived atomic displacement parameters (B-factors). Any number of normal mode perturbed models for MR can be generated for download. If two conformations of the same (or a homologous) protein are available, ElNémo identifies the normal modes that contribute most to the corresponding protein movement. The web server can be freely accessed at http://igs-server.cnrs-mrs.fr/elnemo/index.html.

Crystallography, X-Ray↗

3DCoffee@igs: a web server for combining sequences and structures into a multiple sequence alignment.

This paper presents 3DCoffee@igs, a web-based tool dedicated to the computation of high-quality multiple sequence alignments (MSAs). 3D-Coffee makes it possible to mix protein sequences and structures in order to increase the accuracy of the alignments. Structures can be either provided as PDB identifiers or directly uploaded into the server. Given a set of sequences and structures, pairs of structures are aligned with SAP while sequence-structure pairs are aligned with Fugue. The resulting collection of pairwise alignments is then combined into an MSA with the T-Coffee algorithm. The server and its documentation are available from http://igs-server.cnrs-mrs.fr/Tcoffee/.

Internet↗

CaspR: a web server for automated molecular replacement using homology modelling.

Molecular replacement (MR) is the method of choice for X-ray crystallography structure determination when structural homologues are available in the Protein Data Bank (PDB). Although the success rate of MR decreases sharply when the sequence similarity between template and target proteins drops below 35% identical residues, it has been found that screening for MR solutions with a large number of different homology models may still produce a suitable solution where the original template failed. Here we present the web tool CaspR, implementing such a strategy in an automated manner. On input of experimental diffraction data, of the corresponding target sequence and of one or several potential templates, CaspR executes an optimized molecular replacement procedure using a combination of well-established stand-alone software tools. The protocol of model building and screening begins with the generation of multiple structure-sequence alignments produced with T-COFFEE, followed by homology model building using MODELLER, molecular replacement with AMoRe and model refinement based on CNS. As a result, CaspR provides a progress report in the form of hierarchically organized summary sheets that describe the different stages of the computation with an increasing level of detail. For the 10 highest-scoring potential solutions, pre-refined structures are made available for download in PDB format. Results already obtained with CaspR and reported on the web server suggest that such a strategy significantly increases the fraction of protein structures which may be solved by MR. Moreover, even in situations where standard MR yields a solution, pre-refined homology models produced by CaspR significantly reduce the time-consuming refinement process. We expect this automated procedure to have a significant impact on the throughput of large-scale structural genomics projects. CaspR is freely available at http://igs-server.cnrs-mrs.fr/Caspr/.

Escherichia coli Proteins↗

Set1 is required for meiotic S-phase onset, double-strand break formation and middle gene expression.

The Set1 protein of Saccharomyces cerevisiae is a histone methyltransferase (HMTase) acting on lysine 4 of histone H3. Inactivation of the SET1 gene in a diploid leads to a sporulation defect. We have studied various processes that take place during meiotic differentiation in set1delta diploid cells. The absence of Set1 leads to a delay of meiotic S-phase onset, which reflects a defect in DNA replication initiation. The timely induction of meiotic DNA replication does not require the Set1 HMTase activity, but depends on the SET domain. In addition, set1delta displays a severe impairment of the DNA double-strand break formation, which is not only the consequence of the replication delay. Transcriptional profiling experiments show that the induction of middle meiotic genes, but not of early meiotic genes, is affected by the loss of Set1. In contrast to meiotic replication, the transcriptional induction of the middle meiotic genes appears to depend on the methylation of H3-K4. Our results unveil multiple roles of Set1 in meiotic differentiation and distinguish between HMTase-dependent and -independent Set1 functions.

DNA Damage↗

On the potential of normal-mode analysis for solving difficult molecular-replacement problems.

Molecular replacement (MR) is the method of choice for X-ray crystallographic data phasing when structural data of suitable homologues are available. However, MR may fail even in cases of high sequence homology when conformational changes arising for example from ligand binding or different crystallogenic conditions come into play. In this work, the potential of normal-mode analysis as an extension to MR to allow recovery from such drawbacks is demonstrated. Three examples are presented in which screening for MR solutions with templates perturbed in the direction of one or two normal modes allows a valid MR solution to be found where MR using the original template failed to yield a model that could ultimately be refined. It has been shown recently that half of the known protein movements can be modelled by displacing the studied structure using at most two low-frequency normal modes. This suggests that normal-mode analysis has the potential to break tough MR problems in up to 50% of cases. Moreover, even in cases where an MR solution is available, this method can be used to further improve the starting model prior to refinement, eventually reducing the time spent on manual model construction (in particular for low-resolution data sets).

Carrier Proteins↗

FusionDB: a database for in-depth analysis of prokaryotic gene fusion events.

FusionDB (http://igs-server.cnrs-mrs.fr/FusionDB/) constitutes a resource dedicated to in-depth analysis of bacterial and archaeal gene fusion events. Such events can provide the 'Rosetta stone' in the search for potential protein-protein interactions, as well as metabolic and regulatory networks. However, the false positive rate of this approach may be quite high, prompting a detailed scrutiny of putative gene fusion events. FusionDB readily provides much of the information required for that task. Moreover, FusionDB extends the notion of gene fusion from that of a single gene to that of a family of genes by assembling pairs of genes from different genomes that belong to the same Cluster of Orthogonal Groups (COG). Multiple sequence alignments and phylogenetic tree reconstruction for the N- and C-terminal parts of these 'COG fusion' events are provided to distinguish single and multiple fusion events from cases of gene fission, pseudogenes and other false positives. Finally, gene fusion events with matches to known structures of heterodimers in the Protein Data Bank (PDB) are identified and may be visualized. FusionDB is fully searchable with access to sequence and alignment data at all levels. A number of different scores are provided to easily differentiate 'real' from 'questionable' cases, especially when larger database searches are performed. FusionDB is cross-linked with the 'Phylogenomic Display of Bacterial Genes' (PhydBac) online web server. Together, these servers provide the complete set of information required for in-depth analysis of non-homology-based gene function attribution.

Artificial Gene Fusion↗

Phydbac (phylogenomic display of bacterial genes): An interactive resource for the annotation of bacterial genomes.

Phydbac is a web interactive resource based on phylogenomic profiling, designed to help microbiologists to annotate bacterial proteins. Phylogenomic annotation is based on the assumption that functionally linked protein-coding genes must evolve in a coordinated manner. The detection of subsets of co-evolving genes within a given genome involves the computation of protein sequence conservation profiles across a spectrum of microbial species, followed by the identification of significant pairwise correlations between them. Many ongoing studies are devoted to the problem of computing the most biologically significant phylogenomic profiles and how best identifying clusters of 'functionally interacting' genes. Here we introduce a web tool, Phydbac, allowing the dynamic construction of phylogenomic profiles of protein sequences of interest and their interactive display. In addition, Phydbac can identify Escherichia coli proteins exhibiting the evolution pattern most similar to arbitrary query protein sequences, hence providing functional hints for open reading frames (ORFs) of hypothetical or unknown function. The phylogenomic profiles of all E.coli K-12 protein-coding genes are pre-computed, allowing queries about E.coli genes to be answered instantaneously. The profiles and phylogenomic neighborhoods are computed using an original method shown to perform better than previous ones. An extension of Phydbac, including precomputed profiles for all available bacterial genomes (including major pathogens) will soon be available. Phydbac can be accessed at: http://igs-server.cnrs-mrs.fr/phydbac/.

Bacterial Proteins↗

Genomic correlates of hyperthermostability, an update.

It has been shown (Cambillau, C., and Claverie, J. M. (2000) J. Biol. Chem. 275, 32383-32386) that a large difference between the proportions of charged versus polar (non-charged) amino acids (CvP-bias) was an adequate, if empirical, signature of the proteome of hyperthermophilic organisms (T(growth) >80 degrees C). Since that study, the number of available microbial genomes has more than doubled, raising the possibility that the simple CvP-bias rule might no longer hold. Taking advantage of the new sequence data, we re-analyzed the genomes of 9 fully sequenced thermophiles, 9 hyperthermophiles, and 53 mesothermophile microorganisms to identify the genomic correlates of hyperthermostability on a wider data set. Our new results confirm that the CvP-bias previously identified on a much smaller data set still holds. Moreover, we show that it is an optimal criterion, in the sense that it corresponds to the most discriminating factor between hyperthermophilic and mesothermophilic microorganisms in a principal component analysis. In parallel, we evaluated two other recently proposed correlates of hyperthermostability, the proteome average pI and the dinucleotide statistical index (Kawashima, T., Amano, N., Koike, H., Makino, S., Higuchi, S., Kawashima-Ohya, Y., Watanabe, K., Yamazaki, M., Kanehori, K., Kawamoto, T., Nunoshiba, T., Yamamoto, Y., Aramaki, H., Makino, K., and Suzuki, M. (2000) Proc. Natl. Acad. Sci. 97, 14257-14262). We show that the CvP-bias is the sole criterion that is able to clearly discriminate hyperthermophile from mesothermophile microorganisms on a global genomic basis.

Adaptation, Physiological↗

Structural genomics of highly conserved microbial genes of unknown function in search of new antibacterial targets.

With more than 100 antibacterial drugs at our disposal in the 1980's, the problem of bacterial infection was considered solved. Today, however, most hospital infections are insensitive to several classes of antibacterial drugs, and deadly strains of Staphylococcus aureus resistant to vancomycin--the last resort antibiotic--have recently begin to appear. Other life-threatening microbes, such as Enterococcus faecalis and Mycobacterium tuberculosis are already able to resist every available antibiotic. There is thus an urgent, and continuous need for new, preferably large-spectrum, antibacterial molecules, ideally targeting new biochemical pathways. Here we report on the progress of our structural genomics program aiming at the discovery of new antibacterial gene targets among evolutionary conserved genes of uncharacterized function. A series of bioinformatic and comparative genomics analyses were used to identify a set of 221 candidate genes common to Gram-positive and Gram-negative bacteria. These genes were split between two laboratories. They are now submitted to a systematic 3-D structure determination protocol including cloning, protein expression and purification, crystallization, X-ray diffraction, structure interpretation, and function prediction. We describe here our strategies for the 111 genes processed in our laboratory. Bioinformatics is used at most stages of the production process and out of 111 genes processed--and 17 months into the project--108 have been successfully cloned, 103 have exhibited detectable expression, 84 have led to the production of soluble protein, 46 have been purified, 12 have led to usable crystals, and 7 structures have been determined.

Acid Anhydride Hydrolases↗

Tropheryma whipplei Twist: a human pathogenic Actinobacteria with a reduced genome.

The human pathogen Tropheryma whipplei is the only known reduced genome species (<1 Mb) within the Actinobacteria [high G+C Gram-positive bacteria]. We present the sequence of the 927303-bp circular genome of T. whipplei Twist strain, encoding 808 predicted protein-coding genes. Specific genome features include deficiencies in amino acid metabolisms, the lack of clear thioredoxin and thioredoxin reductase homologs, and a mutation in DNA gyrase predicting a resistance to quinolone antibiotics. Moreover, the alignment of the two available T. whipplei genome sequences (Twist vs. TW08/27) revealed a large chromosomal inversion the extremities of which are located within two paralogous genes. These genes belong to a large cell-surface protein family defined by the presence of a common repeat highly conserved at the nucleotide level. The repeats appear to trigger frequent genome rearrangements in T. whipplei, potentially resulting in the expression of different subsets of cell surface proteins. This might represent a new mechanism for evading host defenses. The T. whipplei genome sequence was also compared to other reduced bacterial genomes to examine the generality of previously detected features. The analysis of the genome sequence of this previously largely unknown human pathogen is now guiding the development of molecular diagnostic tools and more convenient culture conditions.

Actinomycetales↗