Search PubMed⌕ Search

Biomedical subjects

Thomas Lengauer

Publications and source records attributed to Thomas Lengauer.

72 records · Page 4Linked to original sources

Methods for optimizing antiviral combination therapies.

MOTIVATION: Despite some progress with antiretroviral combination therapies, therapeutic success in the management of HIV-infected patients is limited. The evolution of drug-resistant genetic variants in response to therapy plays a key role in treatment failure and finding a new potent drug combination after therapy failure is considered challenging. RESULTS: To estimate the activity of a drug combination against a particular viral strain, we develop a scoring function whose independent variables describe a set of antiviral agents and viral DNA sequences coding for the molecular targets of the respective drugs. The construction of this activity score involves (1) predicting phenotypic drug resistance from genotypes for each drug individually, (2) probabilistic modeling of predicted resistance values and integration into a score for drug combinations, and (3) searching through the mutational neighborhood of the considered strain in order to estimate activity on nearby mutants. For a clinical data set, we determine the optimal search depth and show that the scoring scheme is predictive of therapeutic outcome. Properties of the activity score and applications are discussed.

Algorithms↗

Simple consensus procedures are effective and sufficient in secondary structure prediction.

We have analyzed the performance of majority voting on minimal combination sets of three state-of-the-art secondary structure prediction methods in order to obtain a consensus prediction. Using three large benchmark sets from the EVA server, our results show a significant improvement in the average Q3 prediction accuracy of up to 1.5 percentage points by consensus formation. The application of an additional trivial filtering procedure for predicted secondary structure elements that are too short, does not significantly affect the prediction accuracy. Our analysis also provides valuable insight into the similarity of the results of the prediction methods that we combine as well as the higher confidence in consistently predicted secondary structure.

Computational Biology↗

Tenofovir resistance and resensitization.

Human immunodeficiency viruses in 321 samples from tenofovir-naïve patients were retrospectively evaluated for resistance to this nucleotide analogue. All virus strains with insertions between amino acids 67 and 70 of the reverse transcriptase (n = 6) were highly resistant. Virus strains with the Q151M mutation were divided into susceptible (n = 12) and highly resistant (n = 8) viruses. This difference was due to the absence or presence of the K65R mutation, which was confirmed by site-directed mutagenesis. Viral clones with various combinations of the mutations M41L, K70R, L210W, and T215F or T215Y were analyzed for cross-resistance induced by thymidine analogue mutations (TAMs). The levels of increased resistance induced by single, double, and triple mutations at the indicated positions could be ranked as follows: for mutants with single mutations, mutations at positions 41 > 215 > 70; for mutants with double mutations, mutations at positions 41 and 215 > 70 and 215 = 210 and 215 > 41 and 70; for mutants with triple mutations, mutations at positions 41, 210, and 215 > 41, 70, and 215. Viral clones with M184V or M184I exhibited slightly increased susceptibilities to tenofovir (0.7-fold). Almost all clones with TAM-induced resistance were resensitized when M184V was present (P < 0.001). Among the viruses in the clinical samples, the rate of tenofovir resistance significantly increased with the number of TAMs both in the samples with 184M and in those with 184V (P = 0.005 and P = 0.003, respectively). A resensitizing effect of M184V was confirmed for all samples exhibiting at least one TAM (P = 0.03). However, accumulation of at least two TAMs resulted in more than 2.0-fold reduced susceptibility to tenofovir, irrespective of the presence of M184V. Decision tree building, a classical machine learning technique, was used to generate models for the interpretation of mutations with respect to tenofovir resistance. The application of previously proposed cutoffs for a reduced response to therapy and treatment failure demonstrated the central roles of positions 215 and 65 for 1.5- and 4.0-fold reduced susceptibilities, respectively. Thus, clinically relevant resistance may be conferred by the accumulation of TAMs, and the resensitizing effect of M184V should be considered only minor.

Adenine↗

Protein function from sequence and structure data.

With the large amount of genomics and proteomics data that we are confronted with, computational support for the elucidation of protein function becomes more and more pressing. Many different kinds of biological data harbour signals of protein function, but these signals are often concealed. Computational methods that use protein sequence and structure data can be used for discovering these signals. They provide information that can substantially speed up experimental function elucidation. In this review we concentrate on such methods.

Amino Acid Sequence↗

Diversity and complexity of HIV-1 drug resistance: a bioinformatics approach to predicting phenotype from genotype.

Drug resistance testing has been shown to be beneficial for clinical management of HIV type 1 infected patients. Whereas phenotypic assays directly measure drug resistance, the commonly used genotypic assays provide only indirect evidence of drug resistance, the major challenge being the interpretation of the sequence information. We analyzed the significance of sequence variations in the protease and reverse transcriptase genes for drug resistance and derived models that predict phenotypic resistance from genotypes. For 14 antiretroviral drugs, both genotypic and phenotypic resistance data from 471 clinical isolates were analyzed with a machine learning approach. Information profiles were obtained that quantify the statistical significance of each sequence position for drug resistance. For the different drugs, patterns of varying complexity were observed, including between one and nine sequence positions with substantial information content. Based on these information profiles, decision tree classifiers were generated to identify genotypic patterns characteristic of resistance or susceptibility to the different drugs. We obtained concise and easily interpretable models to predict drug resistance from sequence information. The prediction quality of the models was assessed in leave-one-out experiments in terms of the prediction error. We found prediction errors of 9.6-15.5% for all drugs except for zalcitabine, didanosine, and stavudine, with prediction errors between 25.4% and 32.0%. A prediction service is freely available at http://cartan.gmd.de/geno2pheno.html.

Computational Biology↗

Confidence measures for protein fold recognition.

MOTIVATION: We present an extensive evaluation of different methods and criteria to detect remote homologs of a given protein sequence. We investigate two associated problems: first, to develop a sensitive searching method to identify possible candidates and, second, to assign a confidence to the putative candidates in order to select the best one. For searching methods where the score distributions are known, p-values are used as confidence measure with great success. For the cases where such theoretical backing is absent, we propose empirical approximations to p-values for searching procedures. RESULTS: As a baseline, we review the performances of different methods for detecting remote protein folds (sequence alignment and threading, with and without sequence profiles, global and local). The analysis is performed on a large representative set of protein structures. For fold recognition, we find that methods using sequence profiles generally perform better than methods using plain sequences, and that threading methods perform better than sequence alignment methods. In order to assess the quality of the predictions made, we establish and compare several confidence measures, including raw scores, z-scores, raw score gaps, z-score gaps, and different methods of p-value estimation. We work our way from the theoretically well backed local scores towards more explorative global and threading scores. The methods for assessing the statistical significance of predictions are compared using specificity--sensitivity plots. For local alignment techniques we find that p-value methods work best, albeit computationally cheaper methods such as those based on score gaps achieve similar performance. For global methods where no theory is available methods based on score gaps work best. By using the score gap functions as the measure of confidence we improve the more powerful fold recognition methods for which p-values are unavailable. AVAILABILITY: The benchmark set is available upon request.

Amino Acid Sequence↗

Co-clustering of biological networks and gene expression data.

MOTIVATION: Large scale gene expression data are often analysed by clustering genes based on gene expression data alone, though a priori knowledge in the form of biological networks is available. The use of this additional information promises to improve exploratory analysis considerably. RESULTS: We propose constructing a distance function which combines information from expression data and biological networks. Based on this function, we compute a joint clustering of genes and vertices of the network. This general approach is elaborated for metabolic networks. We define a graph distance function on such networks and combine it with a correlation-based distance function for gene expression measurements. A hierarchical clustering and an associated statistical measure is computed to arrive at a reasonable number of clusters. Our method is validated using expression data of the yeast diauxic shift. The resulting clusters are easily interpretable in terms of the biochemical network and the gene expression data and suggest that our method is able to automatically identify processes that are relevant under the measured conditions.

Algorithms↗

ProML--the protein markup language for specification of protein sequences, structures and families.

We propose a specification language ProML for protein sequences, structures, and families based on the open XML standard. The language allows for portable, system-independent, machine-parsable and human-readable representation of essential features of proteins. The language is of immediate use for several bioinformatics applications: we discuss clustering of proteins into families and the representation of the specific shared features of the respective clusters. Moreover, we use ProML for specification of data used in fold recognition bench-marks exploiting experimentally derived distance constraints.

Programming Languages↗

Improving fold recognition of protein threading by experimental distance constraints.

We present a comprehensive analysis of methods for improving the fold recognition rate of the threading approach to protein structure prediction by the utilization of few additional distance constraints. The distance constraints between protein residues may be obtained by experiments such as mass spectrometry or NMR spectroscopy. We applied a post-filtering step with new scoring functions incorporating measures of constraint satisfaction to ranking lists of 123D threading alignments. The detailed analysis of the results on a small representative benchmark set show that the fold recognition rate can be improved significantly by up to 30% from about 54%-65% to 77%-84%, approaching the maximal attainable performance of 90% estimated by structural superposition alignments. This gain in performance adds about 10% to the recognition rate already achieved in our previous study with cross-link constraints only. Additional recent results on a larger benchmark set involving a confidence function for threading predictions also indicate notable improvements by our combined approach, which should be particularly valuable for rapid structure determination and validation of protein models.

Nuclear Magnetic Resonance, Biomolecular↗

A hypergraph-based method for unification of existing protein structure- and sequence-families.

Classification of proteins is a major challenge in bioinformatics. Here an approach is presented, that unifies different existing classifications of protein structures and sequences. Protein structural domains are represented as nodes in a hypergraph. Shared memberships in sequence families result in hyperedges in the graph. The presented method partitions the hypergraph into clusters of structural domains. Each computed cluster is based on a set of shared sequence family memberships. Thus, the clusters put existing protein sequence families into the context of structural family hierarchies. Conversely, structural domains are related to their sequence family memberships, which can be used to gain further knowledge about the respective structural families.

Databases, Protein↗

Automatic generation of complementary descriptors with molecular graph networks.

We describe a method for the automatic generation of weakly correlated descriptors for molecular data sets. The method can be regarded as a statistical learning procedure that turns the molecular graph, representing the 2D formula of the compound, into an adaptive whole molecule composite descriptor. By translating the molecular graph structure into a dynamical system, the algorithm can compute an output value that is highly sensitive to the molecular topology. This system can be trained by gradient descent techniques, which rely on the efficient calculation of the gradient by back-propagation. We present computational experiments concerning the classification of the Developmental Therapeutics Program AIDS antiviral screen data set on which the performance of the method compares with that of approaches based on substructure comparison.

Algorithms↗

Ensemble methods for classification in cheminformatics.

We describe the application of ensemble methods to binary classification problems on two pharmaceutical compound data sets. Several variants of single and ensembles models of k-nearest neighbors classifiers, support vector machines (SVMs), and single ridge regression models are compared. All methods exhibit robust classification even when more features are given than observations. On two data sets dealing with specific properties of drug-like substances (cytochrome P450 inhibition and "Frequent Hitters", i.e., unspecific protein inhibition), we achieve classification rates above 90%. We are able to reduce the cross-validated misclassification rate for the Frequent Hitters problem by a factor of 2 compared to previous results obtained for the same data set with different modeling techniques.

Journal Article↗

POEM: Parameter Optimization using Ensemble Methods: application to target specific scoring functions.

In computational biology processes such as docking, binding, and folding are often described by simplified, empirical models. These models are fitted to physical properties of the process by adjustable parameters. An appropriate choice of these parameters is crucial for the quality of the models. Locating the best choices for the parameters is often is a difficult task, depending on the complexity of the model. We describe a new method and program, POEM (Parameter Optimization using Ensemble Methods), for this task. In POEM we combine the DOE (Design Of Experiment) procedure with ensembles of different regression methods. We apply the method to the optimization of target specific scoring functions in molecular docking. The method consists of an iterative procedure that uses alternate evaluation and prediction steps. During each cycle of optimization we fit an approximate function to a defined loss function landscape and improve the quality of this fit from cycle to cycle by constantly augmenting our data set. As test applications we fitted the FlexX and Screenscore scoring functions to the kinase and ATPase protein classes. The results are promising: Starting from random parameters we are able to locate parameter sets which show superior performance compared to the original values. The POEM approach converges quickly and the approximated loss function landscapes are smooth, thus making the approach a suitable method for optimizations on rugged landscapes.

Adenosine Triphosphatases↗

A fully computational model for predicting percutaneous drug absorption.

The prediction of transdermal absorption for arbitrary penetrant structures has several important applications in the pharmaceutical industry. We propose a new data-driven, predictive model for skin permeability coefficients k(p) based on an ensemble model using k-nearest-neighbor models and ridge regression. The model was trained and validated with a newly assembled data set containing experimental data and structures for 110 compounds. On the basis of three purely computational descriptors (molecular weight, calculated octanol/water partition coefficient, and solvation free energy), we have developed a model allowing for the reliable, purely computational prediction of skin permeability coefficients. The model is both accurate and robust, as we showed in an extensive validation (correlation coefficient for leave-one-out cross validation: Q = 0.948, mean standard error: 0.2 for log k(p)).

Animals↗

Fully automated flexible docking of ligands into flexible synthetic receptors using forward and inverse docking strategies.

The prediction of the structure of host-guest complexes is one of the most challenging problems in supramolecular chemistry. Usual procedures for docking of ligands into receptors do not take full conformational freedom of the host molecule into account. We describe and apply a new docking approach which performs a conformational sampling of the host and then sequentially docks the ligand into all receptor conformers using the incremental construction technique of the FlexX software platform. The applicability of this approach is validated on a set of host-guest complexes with known crystal structure. Moreover, we demonstrate that due to the interchangeability of the roles of host and guest, the docking process can be inverted. In this inverse docking mode, the receptor molecule is docked around its ligand. For all investigated test cases, the predicted structures are in good agreement with the experiment for both normal (forward) and inverse docking. Since the ligand is often smaller than the receptor and, thus, its conformational space is more restricted, the inverse docking approach leads in most cases to considerable speed-up. By having the choice between two alternative docking directions, the application range of the method is significantly extended. Finally, an important result of this study is the suitability of the simple energy function used here for structure prediction of complexes in organic media.

Algorithms↗

Flexible docking of ligands into synthetic receptors using a two-sided incremental construction algorithm.

We present a new algorithm for the fast and reliable structure prediction of synthetic receptor-ligand complexes. Our method is based on the protein-ligand docking program FlexX and extends our recently introduced docking technique for synthetic receptors, which has been implemented in the program FlexR. To handle the flexibility of the relevant molecules, we apply a novel docking strategy that uses an adaptive two-sided incremental construction algorithm which incorporates the structural flexibility of both the ligand and synthetic receptor. We follow an adaptive strategy, in which one molecule is expanded by attaching its next fragment in all possible torsion angles, whereas the other (partially assembled) molecule serves as a rigid binding partner. Then the roles of the molecules are exchanged. Geometric filters are used to discard partial conformations that cannot realize a targeted interaction pattern derived in a graph-based precomputation phase. The process is repeated until the entire complex is built up. Our algorithm produces promising results on a test data set comprising 10 complexes of synthetic receptors and ligands. The method generated near-native solutions compared to crystal structures in all but one case. It is able to generate solutions within a couple of minutes and has the potential of being used as a virtual screening tool for searching for suitable guest molecules for a given synthetic receptor in large databases of guests and vice versa.

Algorithms↗

Learning multiple evolutionary pathways from cross-sectional data.

We introduce a mixture model of trees to describe evolutionary processes that are characterized by the ordered accumulation of permanent genetic changes. The basic building block of the model is a directed weighted tree that generates a probability distribution on the set of all patterns of genetic events. We present an EM-like algorithm for learning a mixture model of K trees and show how to determine K with a maximum likelihood approach. As a case study, we consider the accumulation of mutations in the HIV-1 reverse transcriptase that are associated with drug resistance. The fitted model is statistically validated as a density estimator, and the stability of the model topology is analyzed. We obtain a generative probabilistic model for the development of drug resistance in HIV that agrees with biological knowledge. Further applications and extensions of the model are discussed.

Algorithms↗