Search PubMed⌕ Search

Biomedical subjects

Thomas Lengauer

Publications and source records attributed to Thomas Lengauer.

At least 19 recordsLinked to original sources

Improved scoring of functional groups from gene expression data by decorrelating GO graph structure.

MOTIVATION: The result of a typical microarray experiment is a long list of genes with corresponding expression measurements. This list is only the starting point for a meaningful biological interpretation. Modern methods identify relevant biological processes or functions from gene expression data by scoring the statistical significance of predefined functional gene groups, e.g. based on Gene Ontology (GO). We develop methods that increase the explanatory power of this approach by integrating knowledge about relationships between the GO terms into the calculation of the statistical significance. RESULTS: We present two novel algorithms that improve GO group scoring using the underlying GO graph topology. The algorithms are evaluated on real and simulated gene expression data. We show that both methods eliminate local dependencies between GO terms and point to relevant areas in the GO graph that remain undetected with state-of-the-art algorithms for scoring functional terms. A simulation study demonstrates that the new methods exhibit a higher level of detecting relevant biological terms than competing methods.

Algorithms↗

Computational recognition of potassium channel sequences.

MOTIVATION: Potassium channels are mainly known for their role in regulating and maintaining the membrane potential. Since this is one of the key mechanisms of signal transduction, malfunction of these potassium channels leads to a wide variety of severe diseases. Thus potassium channels are priority targets of research for new drugs, despite the fact that this protein family is highly variable and closely related to other channels, which makes it very difficult to identify new types of potassium channel sequences. RESULTS: Here we present a new method for identifying potassium channel sequences (PSM, Property Signature Method), which-in contrast to the known methods for protein classification-is directly based on physicochemical properties of amino acids rather than on the amino acids themselves. A signature for the pore region including the selectivity filter has been created, representing the most common physicochemical properties of known potassium channels. This string enables genome-wide screening for sequences with similar features despite a very low degree of amino acid similarity within a protein family.

Algorithms↗

CpG island methylation in human lymphocytes is highly correlated with DNA sequence, repeats, and predicted DNA structure.

CpG island methylation plays an important role in epigenetic gene control during mammalian development and is frequently altered in disease situations such as cancer. The majority of CpG islands is normally unmethylated, but a sizeable fraction is prone to become methylated in various cell types and pathological situations. The goal of this study is to show that a computational epigenetics approach can discriminate between CpG islands that are prone to methylation from those that remain unmethylated. We develop a bioinformatics scoring and prediction method on the basis of a set of 1,184 DNA attributes, which refer to sequence, repeats, predicted structure, CpG islands, genes, predicted binding sites, conservation, and single nucleotide polymorphisms. These attributes are scored on 132 CpG islands across the entire human Chromosome 21, whose methylation status was previously established for normal human lymphocytes. Our results show that three groups of DNA attributes, namely certain sequence patterns, specific DNA repeats, and a particular DNA structure, are each highly correlated with CpG island methylation (correlation coefficients of 0.64, 0.66, and 0.49, respectively). We predicted, and subsequently experimentally examined 12 CpG islands from human Chromosome 21 with unknown methylation patterns and found more than 90% of our predictions to be correct. In addition, we applied our prediction method to analyzing Human Epigenome Project methylation data on human Chromosome 6 and again observed high prediction accuracy. In summary, our results suggest that DNA composition of CpG islands (sequence, repeats, and structure) plays a significant role in predisposing CpG islands for DNA methylation. This finding may have a strong impact on our understanding of changes in CpG island methylation in development and disease.

CpG Islands↗

Recco: recombination analysis using cost optimization.

MOTIVATION: Recombination plays an important role in the evolution of many pathogens, such as HIV or malaria. Despite substantial prior work, there is still a pressing need for efficient and effective methods of detecting recombination and analyzing recombinant sequences. RESULTS: We introduce Recco, a novel fast method that, given a multiple sequence alignment, scores the cost of obtaining one of the sequences from the others by mutation and recombination. The algorithm comes with an illustrative visualization tool for locating recombination breakpoints. We analyze the sequence alignment with respect to all choices of the parameter alpha weighting recombination cost against mutation cost. The analysis of the resulting cost curve yields additional information as to which sequence might be recombinant. On random genealogies Recco is comparable in its power of detecting recombination with the algorithm Geneconv (Sawyer, 1989). For specific relevant recombination scenarios Recco significantly outperforms Geneconv.

Algorithms↗

Mutations in the NB-ARC domain of I-2 that impair ATP hydrolysis cause autoactivation.

Resistance (R) proteins in plants confer specificity to the innate immune system. Most R proteins have a centrally located NB-ARC (nucleotide-binding adaptor shared by APAF-1, R proteins, and CED-4) domain. For two tomato (Lycopersicon esculentum) R proteins, I-2 and Mi-1, we have previously shown that this domain acts as an ATPase module that can hydrolyze ATP in vitro. To investigate the role of nucleotide binding and hydrolysis for the function of I-2 in planta, specific mutations were introduced in conserved motifs of the NB-ARC domain. Two mutations resulted in autoactivating proteins that induce a pathogen-independent hypersensitive response upon expression in planta. These mutant forms of I-2 were found to be impaired in ATP hydrolysis, but not in ATP binding, suggesting that the ATP- rather than the ADP-bound state of I-2 is the active form that triggers defense signaling. In addition, upon ADP binding, the protein displayed an increased affinity for ADP suggestive of a change of conformation. Based on these data, we propose that the NB-ARC domain of I-2, and likely of related R proteins, functions as a molecular switch whose state (on/off) depends on the nucleotide bound (ATP/ADP).

Adenosine Triphosphatases↗

NOXclass: prediction of protein-protein interaction types.

BACKGROUND: Structural models determined by X-ray crystallography play a central role in understanding protein-protein interactions at the molecular level. Interpretation of these models requires the distinction between non-specific crystal packing contacts and biologically relevant interactions. This has been investigated previously and classification approaches have been proposed. However, less attention has been devoted to distinguishing different types of biological interactions. These interactions are classified as obligate and non-obligate according to the effect of the complex formation on the stability of the protomers. So far no automatic classification methods for distinguishing obligate, non-obligate and crystal packing interactions have been made available. RESULTS: Six interface properties have been investigated on a dataset of 243 protein interactions. The six properties have been combined using a support vector machine algorithm, resulting in NOXclass, a classifier for distinguishing obligate, non-obligate and crystal packing interactions. We achieve an accuracy of 91.8% for the classification of these three types of interactions using a leave-one-out cross-validation procedure. CONCLUSION: NOXclass allows the interpretation and analysis of protein quaternary structures. In particular, it generates testable hypotheses regarding the nature of protein-protein interactions, when experimental results are not available. We expect this server will benefit the users of protein structural models, as well as protein crystallographers and NMR spectroscopists. A web server based on the method and the datasets used in this study are available at http://noxclass.bioinf.mpi-inf.mpg.de/.

Algorithms↗

Local protein structure prediction using discriminative models.

BACKGROUND: In recent years protein structure prediction methods using local structure information have shown promising improvements. The quality of new fold predictions has risen significantly and in fold recognition incorporation of local structure predictions led to improvements in the accuracy of results. We developed a local structure prediction method to be integrated into either fold recognition or new fold prediction methods. For each local sequence window of a protein sequence the method predicts probability estimates for the sequence to attain particular local structures from a set of predefined local structure candidates. The first step is to define a set of local structure representatives based on clustering recurrent local structures. In the second step a discriminative model is trained to predict the local structure representative given local sequence information. RESULTS: The step of clustering local structures yields an average RMSD quantization error of 1.19 A for 27 structural representatives (for a fragment length of 7 residues). In the prediction step the area under the ROC curve for detection of the 27 classes ranges from 0.68 to 0.88. CONCLUSION: The described method yields probability estimates for local protein structure candidates, giving signals for all kinds of local structure. These local structure predictions can be incorporated either into fold recognition algorithms to improve alignment quality and the overall prediction accuracy or into new fold prediction methods.

Algorithms↗

Diversity and functional plasticity of eukaryotic selenoproteins: identification and characterization of the SelJ family.

Selenoproteins are a diverse group of proteins that contain selenocysteine (Sec), the 21st amino acid. In the genetic code, UGA serves as a termination signal and a Sec codon. This dual role has precluded the automatic annotation of selenoproteins. Recent advances in the computational identification of selenoprotein genes have provided a first glimpse of the size, functions, and phylogenetic diversity of eukaryotic selenoproteomes. Here, we describe the identification of a selenoprotein family named SelJ. In contrast to known selenoproteins, SelJ appears to be restricted to actinopterygian fishes and sea urchin, with Cys homologues only found in cnidarians. SelJ shows significant similarity to the jellyfish J1-crystallins and with them constitutes a distinct subfamily within the large family of ADP-ribosylation enzymes. Consistent with its potential role as a structural crystallin, SelJ has preferential and homogeneous expression in the eye lens in early stages of zebrafish development. A structural role for SelJ would be in contrast to the majority of known selenoenzymes. The unusually highly restricted phylogenetic distribution of SelJ, its specialization, and the comparative analysis of eukaryotic selenoproteomes reveal the diversity and functional plasticity of selenoproteins and point to a mosaic evolution of the use of Sec in proteins.

Adenosine Diphosphate Ribose↗

Multiple-ligand-based virtual screening: methods and applications of the MTree approach.

We present a novel approach for ligand-based virtual screening by combining query molecules into a multiple feature tree model called MTree. All molecules are described by the established feature tree descriptor, which is derived from a topological molecular graph. A new pairwise alignment algorithm leads to a consistent topological molecular alignment based on chemically reasonable matching of corresponding functional groups. These multiple feature tree models find application in ligand-based virtual screening to identify new lead structures for chemical optimization. Retrospective virtual screening with MTree models generated for angiotensin-converting enzyme and the alpha1a receptor on a large candidate database yielded enrichment factors up to 71 for the first 1% of the screened database. MTree models outperformed database searches using single feature trees in terms of hit rates and quality and additionally identified alternative molecular scaffolds not included in any of the query molecules. Furthermore, relevant molecular features, which are known to be important for affinity to the target, are identified by this new methodology.

Adrenergic alpha-1 Receptor Antagonists↗

Clinical significance of in vitro replication-enhancing mutations of the hepatitis C virus (HCV) replicon in patients with chronic HCV infection.

BACKGROUND: Mutations in nonstructural (NS) hepatitis C virus (HCV) proteins enhance replication in HCV-1a/b replicons. The prevalence of such mutations and their clinical significance in vivo are unknown. METHODS: Parts of HCV NS3 and NS4B-NS5B genes that included 31 in vitro replication-enhancing sites were sequenced for 26 patients with chronic HCV genotype 1 infection. RESULTS: Five patients showed specific mutations within NS3 at sites enhancing replication in the replicon. Those mutations were associated with a slower decrease in HCV RNA concentration during interferon (IFN)- alpha -based therapy (P = .007). Neither specific nor other mutations within NS3 and NS4B-NS5B were associated with baseline HCV RNA concentrations. Within NS5A, fewer mutations in the major HCV strain (P = .001) and increased quasi-species complexity (P = .02) and diversity (P = .02) correlated with increasing baseline HCV RNA concentrations. In silico analyses of NS3 protein structures suggested that the majority of observed mutations did not lead to major conformational changes. CONCLUSIONS: Specific mutations leading to enhanced replication in the replicon system were detected in 5 of 26 patients in vivo and were not associated with baseline HCV RNA concentrations but were associated with a slower decrease in HCV RNA concentration during IFN- alpha -based therapy. Quasi-species heterogeneity of NS5A correlated with baseline HCV RNA concentrations.

Adult↗

Computational methods for the design of effective therapies against drug resistant HIV strains.

The development of drug resistance is a major obstacle to successful treatment of HIV infection. The extraordinary replication dynamics of HIV facilitates its escape from selective pressure exerted by the human immune system and by combination drug therapy. We have developed several computational methods whose combined use can support the design of optimal antiretroviral therapies based on viral genomic data.

Database Management Systems↗

Decomposing protein networks into domain-domain interactions.

UNLABELLED: The application of novel experimental techniques has generated large networks of protein-protein interactions. Frequently, important information on the structure and cellular function of protein-protein interactions can be gained from the domains of interacting proteins. We have designed a Cytoscape plugin that decomposes interacting proteins into their respective domains and computes a putative network of corresponding domain-domain interactions. To this end, the network graph of proteins has been extended by additional node and edge types for domain interactions, including different node and edge shapes and coloring schemes used for visualization. An additional plugin provides supplementary web links to Internet resources on domain function and structure. AVAILABILITY: Both Cytoscape plugins can be downloaded from http://www.cytoscape.org

Algorithms↗

BiQ Analyzer: visualization and quality control for DNA methylation data from bisulfite sequencing.

SUMMARY: Manual processing of DNA methylation data from bisulfite sequencing is a tedious and error-prone task. Here we present an interactive software tool that provides start-to-end support for this process. In an easy-to-use manner, the tool helps the user to import the sequence files from the sequencer, to align them, to exclude or correct critical sequences, to document the experiment, to perform basic statistics and to produce publication-quality diagrams. Emphasis is put on quality control: The program automatically assesses data quality and provides warnings and suggestions for dealing with critical sequences. The BiQ Analyzer program is implemented in the Java programming language and runs on any platform for which a recent Java virtual machine is available. AVAILABILITY: The program is available without charge for non-commercial users and can be downloaded from http://biq-analyzer.bioinf.mpi-inf.mpg.de/

DNA↗

Dissection of the inflammatory bowel disease transcriptome using genome-wide cDNA microarrays.

BACKGROUND: The differential pathophysiologic mechanisms that trigger and maintain the two forms of inflammatory bowel disease (IBD), Crohn disease (CD), and ulcerative colitis (UC) are only partially understood. cDNA microarrays can be used to decipher gene regulation events at a genome-wide level and to identify novel unknown genes that might be involved in perpetuating inflammatory disease progression. METHODS AND FINDINGS: High-density cDNA microarrays representing 33,792 UniGene clusters were prepared. Biopsies were taken from the sigmoid colon of normal controls (n = 11), CD patients (n = 10) and UC patients (n = 10). 33P-radiolabeled cDNA from purified poly(A)+ RNA extracted from biopsies (unpooled) was hybridized to the arrays. We identified 500 and 272 transcripts differentially regulated in CD and UC, respectively. Interesting hits were independently verified by real-time PCR in a second sample of 100 individuals, and immunohistochemistry was used for exemplary localization. The main findings point to novel molecules important in abnormal immune regulation and the highly disturbed cell biology of colonic epithelial cells in IBD pathogenesis, e.g., CYLD (cylindromatosis, turban tumor syndrome) and CDH11 (cadherin 11, type 2). By the nature of the array setup, many of the genes identified were to our knowledge previously uncharacterized, and prediction of the putative function of a subsection of these genes indicate that some could be involved in early events in disease pathophysiology. CONCLUSION: A comprehensive set of candidate genes not previously associated with IBD was revealed, which underlines the polygenic and complex nature of the disease. It points out substantial differences in pathophysiology between CD and UC. The multiple unknown genes identified may stimulate new research in the fields of barrier mechanisms and cell signalling in the context of IBD, and ultimately new therapeutic approaches.

Adolescent↗

Ataxin-2 and huntingtin interact with endophilin-A complexes to function in plastin-associated pathways.

Spinocerebellar ataxia type 2 is an inherited neurodegenerative disorder that is caused by an expanded trinucleotide repeat in the SCA2 gene, encoding a polyglutamine stretch in the gene product ataxin-2. Although evidence has been provided that ataxin-2 is involved in RNA metabolism, the physiological function of ataxin-2 remains unclear. Here, we demonstrate that ataxin-2 interacts with two members of the endophilin family, endophilin-A1 and endophilin-A3. To elucidate the physiological implications of these interactions, we exploited yeast as a model system and discovered that expression of ataxin-2 as well as both endophilin proteins is toxic for yeast lacking the SAC6 gene product fimbrin, a protein involved in actin filament organization and endocytotic processes. Intriguingly, expression of huntingtin, another polyglutamine protein interacting with endophilin-A3, was also toxic in Deltasac6 yeast. These effects can be suppressed by simultaneous expression of one of the two human fimbrin orthologs, L- or T-plastin. Moreover, we have discovered that ataxin-2 associates with L- and T-plastin and that overexpression of ataxin-2 leads to accumulation of T-plastin in mammalian cells. Thus, our findings suggest an interplay between ataxin-2, endophilin proteins and huntingtin in plastin-associated cellular pathways.

Adaptor Proteins, Signal Transducing↗

ROCR: visualizing classifier performance in R.

UNLABELLED: ROCR is a package for evaluating and visualizing the performance of scoring classifiers in the statistical language R. It features over 25 performance measures that can be freely combined to create two-dimensional performance curves. Standard methods for investigating trade-offs between specific performance measures are available within a uniform framework, including receiver operating characteristic (ROC) graphs, precision/recall plots, lift charts and cost curves. ROCR integrates tightly with R's powerful graphics capabilities, thus allowing for highly adjustable plots. Being equipped with only three commands and reasonable default values for optional parameters, ROCR combines flexibility with ease of usage. AVAILABILITY: http://rocr.bioinf.mpi-sb.mpg.de. ROCR can be used under the terms of the GNU General Public License. Running within R, it is platform-independent. CONTACT: tobias.sing@mpi-sb.mpg.de.

Computer Graphics↗

Structural and functional analysis of a novel mutation of CYP21B in a heterozygote carrier of 21-hydroxylase deficiency.

Congenital adrenal hyperplasia (CAH) due to 21-hydroxylase deficiency is one of the most common autosomal recessive disorders and occurs in its non-classical form in up to 6% of hirsute women. We report on a young woman with the clinical diagnosis of non-classical CAH and a novel, heterozygous missense mutation CTG-->GTG in exon 8, codon 317, of the steroid 21-hydroxylase CYP21B and complete loss of pseudogenes. Protein sequences of closely related P450 cytochromes and a homology-based 3D model of CYP21B were used for further functional analyses. We found that the mutated residue is part of a large cluster of hydrophobic residues. This cluster has three important features: (1) it is located directly next to the binding pocket, in close vicinity of the heme-cofactor, (2) all amino acids of the cluster are directly connected to two important binding regions, and (3) the packing within the cluster is very dense. Due to the tight packing in the cluster and its direct connection to the binding pocket region, any changes induced by the mutation of residue 317 can be expected to lead to structural shifts within the binding pocket and can explain the clinically observed impairment of 21-hydroxylase activity. In conclusion, the novel mutation L317V of the steroid 21-hydroxylase gene is associated with reduced steroid 21-hydroxylase activity probably due to structural shifts within the binding pocket and a mild phenotype of steroid 21-hydroxylase deficiency. In addition, the results support previous findings in which heterozygous CYP21 mutations are associated with symptoms of hyperandrogenism in susceptible individuals.

Adolescent↗

Confirmation of human protein interaction data by human expression data.

BACKGROUND: With microarray technology the expression of thousands of genes can be measured simultaneously. It is well known that the expression levels of genes of interacting proteins are correlated significantly more strongly in Saccharomyces cerevisiae than those of proteins that are not interacting. The objective of this work is to investigate whether this observation extends to the human genome. RESULTS: We investigated the quantitative relationship between expression levels of genes encoding interacting proteins and genes encoding random protein pairs. Therefore we studied 1369 interacting human protein pairs and human gene expression levels of 155 arrays. We were able to establish a statistically significantly higher correlation between the expression levels of genes whose proteins interact compared to random protein pairs. Additionally we were able to provide evidence that genes encoding proteins belonging to the same GO-class show correlated expression levels. CONCLUSION: This finding is concurrent with the naive hypothesis that the scales of production of interacting proteins are linked because an efficient interaction demands that involved proteins are available to some degree. The goal of further research in this field will be to understand the biological mechanisms behind this observation.

Cluster Analysis↗