Search PubMed⌕ Search

Biomedical subjects

Morten Nielsen

Publications and source records attributed to Morten Nielsen.

17 recordsLinked to original sources

The validity of predicted T-cell epitopes.

High-performing MHC class I binding predictions have been available for more than a decade; however, their value in terms of actual epitope finding has only now been estimated in a large-scale investigation undertaken by the group of Sette. This work underlines the importance of bioinformatics as a resource-saving tool in the field of epitope discovery. In addition, the data can be used to benchmark the performance of other new or existing CTL epitope-prediction tools.

Algorithms↗

Prediction of residues in discontinuous B-cell epitopes using protein 3D structures.

Discovery of discontinuous B-cell epitopes is a major challenge in vaccine design. Previous epitope prediction methods have mostly been based on protein sequences and are not very effective. Here, we present DiscoTope, a novel method for discontinuous epitope prediction that uses protein three-dimensional structural data. The method is based on amino acid statistics, spatial information, and surface accessibility in a compiled data set of discontinuous epitopes determined by X-ray crystallography of antibody/antigen protein complexes. DiscoTope is the first method to focus explicitly on discontinuous epitopes. We show that the new structure-based method has a better performance for predicting residues of discontinuous epitopes than methods based solely on sequence information, and that it can successfully predict epitope residues that have been identified by different techniques. DiscoTope detects 15.5% of residues located in discontinuous epitopes with a specificity of 95%. At this level of specificity, the conventional Parker hydrophilicity scale for predicting linear B-cell epitopes identifies only 11.0% of residues located in discontinuous epitopes. Predictions by the DiscoTope method can guide experimental epitope mapping in both rational vaccine design and development of diagnostic tools, and may lead to more efficient epitope identification.

Animals↗

Improved method for predicting linear B-cell epitopes.

BACKGROUND: B-cell epitopes are the sites of molecules that are recognized by antibodies of the immune system. Knowledge of B-cell epitopes may be used in the design of vaccines and diagnostics tests. It is therefore of interest to develop improved methods for predicting B-cell epitopes. In this paper, we describe an improved method for predicting linear B-cell epitopes. RESULTS: In order to do this, three data sets of linear B-cell epitope annotated proteins were constructed. A data set was collected from the literature, another data set was extracted from the AntiJen database and a data sets of epitopes in the proteins of HIV was collected from the Los Alamos HIV database. An unbiased validation of the methods was made by testing on data sets on which they were neither trained nor optimized on. We have measured the performance in a non-parametric way by constructing ROC-curves. CONCLUSION: The best single method for predicting linear B-cell epitopes is the hidden Markov model. Combining the hidden Markov model with one of the best propensity scale methods, we obtained the BepiPred method. When tested on the validation data set this method performs significantly better than any of the other methods tested. The server and data sets are publicly available at http://www.cbs.dtu.dk/services/BepiPred.

Journal Article↗

Attempts to predict the long-term decrease in lung function due to radiotherapy of non-small cell lung cancer.

PURPOSE: To obtain a model which can predict long-term decrease in lung function due to radiation damage from dose-volume data for patients with non-small cell lung cancer. PATIENTS AND METHODS: 27 patients were included, all long-term survivors after radical radiation therapy. For each patient a regression analysis was performed on a post-RT succession of measurements of FEV1 in order to estimate the decrease after 2 years and a standard error (SE) on this regression estimate. The modelling was based on dose-volume histograms (DVH) exported from the treatment planning system, and involved fits of threshold models, a mean lung dose model as well as more complex models based on the relative damaged volume (rdV). RESULTS: Decreases after 2 years of up to 28% in FEV1 was measured (median 10%), with significant day-to-day variation in FEV1 for the individual patient. The threshold models predicted the long-term decrease in FEV1 well when the SE was interpreted as the uncertainty of the measured decrease. The best threshold value, marginally, was 30 Gy with an R(2) of 0.46. The mean lung dose model did not perform so well. A complex model based on rdV performed better than any of the other models (R(2)=0.52). CONCLUSION: The long-term decrease in FEV1 could be predicted from a simple dose-volume model when the SE was interpreted as the uncertainty of the measured decrease.

Carcinoma, Non-Small-Cell Lung↗

No evidence for the use of DIR, D-D fusions, chromosome 15 open reading frames or VH replacement in the peripheral repertoire was found on application of an improved algorithm, JointML, to 6329 human immunoglobulin H rearrangements.

Antibody diversity is created by imprecise joining of the variability (V), diversity (D) and joining (J) gene segments of the heavy and light chain loci. Analysis of rearrangements is complicated by somatic hypermutations and uncertainty concerning the sources of gene segments and the precise way in which they recombine. It has been suggested that D genes with irregular recombination signal sequences (DIR) and chromosome 15 open reading frames (OR15) can replace conventional D genes, that two D genes or inverted D genes may be used and that the repertoire can be further diversified by heavy chain V gene (VH) replacement. Safe conclusions require large, well-defined sequence samples and algorithms minimizing stochastic assignment of segments. Two computer programs were developed for analysis of heavy chain joints. JointHMM is a profile hidden Markow model, while JointML is a maximum-likelihood-based method taking the lengths of the joint and the mutational status of the VH gene into account. The programs were applied to a set of 6329 clonally unrelated rearrangements. A conventional D gene was found in 80% of unmutated sequences and 64% of mutated sequences, while D-gene assignment was kept below 5% in artificial (randomly permutated) rearrangements. No evidence for the use of DIR, OR15, multiple D genes or VH replacements was found, while inverted D genes were used in less than 1 per thousand of the sequences. JointML was shown to have a higher predictive performance for D-gene assignment in mutated and unmutated sequences than four other publicly available programs. An online version 1.0 of JointML is available at http://www.cbs.dtu.dk/services/VDJsolver.

Algorithms↗

The role of the proteasome in generating cytotoxic T-cell epitopes: insights obtained from improved predictions of proteasomal cleavage.

Cytotoxic T cells (CTLs) perceive the world through small peptides that are eight to ten amino acids long. These peptides (epitopes) are initially generated by the proteasome, a multi-subunit protease that is responsible for the majority of intra-cellular protein degradation. The proteasome generates the exact C-terminal of CTL epitopes, and the N-terminal with a possible extension. CTL responses may diminish if the epitopes are destroyed by the proteasomes. Therefore, the prediction of the proteasome cleavage sites is important to identify potential immunogenic regions in the proteomes of pathogenic microorganisms (or humans). We have recently shown that NetChop, a neural network-based prediction method, is the best method available at the moment to do such predictions; however, its performance is still lower than desired. Here, we use novel sequence encoding methods and show that the new version of NetChop predicts approximately 10% more of the cleavage sites correctly while lowering the number of false positives with close to 15%. With this more reliable prediction tool, we study two important questions concerning the function of the proteasome. First, we estimate the N-terminal extension of epitopes after proteasomal cleavage and find that the average extension is relatively short. However, more than 30% of the peptides have N-terminal extensions of three amino acids or more, and thus, N-terminal trimming might play an important role in the presentation of a substantial fraction of the epitopes. Second, we show that good TAP ligands have an increased chance of being cleaved by the proteasome, i.e., the specificity of TAP has evolved to fit the specificity of the proteasome. This evolutionary relationship allows for a more efficient antigen presentation.

Animals↗

An integrative approach to CTL epitope prediction: a combined algorithm integrating MHC class I binding, TAP transport efficiency, and proteasomal cleavage predictions.

Reverse immunogenetic approaches attempt to optimize the selection of candidate epitopes, and thus minimize the experimental effort needed to identify new epitopes. When predicting cytotoxic T cell epitopes, the main focus has been on the highly specific MHC class I binding event. Methods have also been developed for predicting the antigen-processing steps preceding MHC class I binding, including proteasomal cleavage and transporter associated with antigen processing (TAP) transport efficiency. Here, we use a dataset obtained from the SYFPEITHI database to show that a method integrating predictions of MHC class I binding affinity, TAP transport efficiency, and C-terminal proteasomal cleavage outperforms any of the individual methods. Using an independent evaluation dataset of HIV epitopes from the Los Alamos database, the validity of the integrated method is confirmed. The performance of the integrated method is found to be significantly higher than that of the two publicly available prediction methods BIMAS and SYFPEITHI. To identify 85% of the epitopes in the HIV dataset, 9% and 10% of all possible nonamers in the HIV proteins must be tested when using the BIMAS and SYFPEITHI methods, respectively, for the selection of candidate epitopes. This number is reduced to 7% when using the integrated method. In practical terms, this means that the experimental effort needed to identify an epitope in a hypothetical protein with 85% probability is reduced by 20-30% when using the integrated method. The method is available at http://www.cbs.dtu.dk/services/NetCTL. Supplementary material is available at http://www.cbs.dtu.dk/suppl/immunology/CTL.php.

ATP-Binding Cassette Transporters↗

Synthesis of linear and tripoidal oligo(phenylene ethynylene)-based building blocks for application in modular DNA-programmed assembly.

Rigid linear and tripoidal organic modules based on the oligo(phenylene ethynylene) backbone having salicylaldehyde-derived termini are synthesized. A highly functionalized 5-iodosalicyl aldehyde was prepared and coupled to each ethynyl group of 1,4-diethynylbenzene or 1,3,5-triethynylbenzene in Sonogashira couplings. The two or three termini of the compounds are functionalized for incorporation in linear and branched oligonucleotide strands. For the linear module (LM), the two termini are equipped with amide spacers, and one of these was functionalized with a DMTr (dimethoxytrityl)-protected hydroxy group and the other with a phosphoramidite. One of the tripoidal modules is prepared with DMTr groups in two of its three termini. A tripoidal module is also synthesized with three different groups on its hydroxy termini: a phosphoramidite, a DMTr group, and an Fmoc group. Extended studies have shown that these rigid linear and tripoidal organic modules can be incorporated into short oligonucleotides. Several of these modules can be applied for DNA-directed assembly and covalent coupling into structures of predetermined connectivity. Such structures have potential application for molecular electronics and nanotechnology.

Alkynes↗

Definition of supertypes for HLA molecules using clustering of specificity matrices.

Major histocompatibility complex (MHC) proteins are encoded by extremely polymorphic genes and play a crucial role in immunity. However, not all genetically different MHC molecules are functionally different. Sette and Sidney (1999) have defined nine HLA class I supertypes and showed that with only nine main functional binding specificities it is possible to cover the binding properties of almost all known HLA class I molecules. Here we present a comprehensive study of the functional relationship between all HLA molecules with known specificities in a uniform and automated way. We have developed a novel method for clustering sequence motifs. We construct hidden Markov models for HLA class I molecules using a Gibbs sampling procedure and use the similarities among these to define clusters of specificities. These clusters are extensions of the previously suggested ones. We suggest splitting some of the alleles in the A1 supertype into a new A26 supertype, and some of the alleles in the B27 supertype into a new B39 supertype. Furthermore the B8 alleles may define their own supertype. We also use the published specificities for a number of HLA-DR types to define clusters with similar specificities. We report that the previously observed specificities of these class II molecules can be clustered into nine classes, which only partly correspond to the serological classification. We show that classification of HLA molecules may be done in a uniform and automated way. The definition of clusters allows for selection of representative HLA molecules that can cover the HLA specificity space better. This makes it possible to target most of the known HLA alleles with known specificities using only a few peptides, and may be used in construction of vaccines. Supplementary material is available at http://www.cbs.dtu.dk/researchgroups/immunology/supertypes.html.

Amino Acid Motifs↗

Improved prediction of MHC class I and class II epitopes using a novel Gibbs sampling approach.

MOTIVATION: Prediction of which peptides will bind a specific major histocompatibility complex (MHC) constitutes an important step in identifying potential T-cell epitopes suitable as vaccine candidates. MHC class II binding peptides have a broad length distribution complicating such predictions. Thus, identifying the correct alignment is a crucial part of identifying the core of an MHC class II binding motif. In this context, we wish to describe a novel Gibbs motif sampler method ideally suited for recognizing such weak sequence motifs. The method is based on the Gibbs sampling method, and it incorporates novel features optimized for the task of recognizing the binding motif of MHC classes I and II. The method locates the binding motif in a set of sequences and characterizes the motif in terms of a weight-matrix. Subsequently, the weight-matrix can be applied to identifying effectively potential MHC binding peptides and to guiding the process of rational vaccine design. RESULTS: We apply the motif sampler method to the complex problem of MHC class II binding. The input to the method is amino acid peptide sequences extracted from the public databases of SYFPEITHI and MHCPEP and known to bind to the MHC class II complex HLA-DR4(B1*0401). Prior identification of information-rich (anchor) positions in the binding motif is shown to improve the predictive performance of the Gibbs sampler. Similarly, a consensus solution obtained from an ensemble average over suboptimal solutions is shown to outperform the use of a single optimal solution. In a large-scale benchmark calculation, the performance is quantified using relative operating characteristics curve (ROC) plots and we make a detailed comparison of the performance with that of both the TEPITOPE method and a weight-matrix derived using the conventional alignment algorithm of ClustalW. The calculation demonstrates that the predictive performance of the Gibbs sampler is higher than that of ClustalW and in most cases also higher than that of the TEPITOPE method.

Algorithms↗

Modular DNA-programmed assembly of linear and branched conjugated nanostructures.

A new strategy for self-assembly and covalent coupling of encoded molecular modules into nanostructures with predetermined connectivity has been developed. The method uses DNA-functionalized oligo(phenylene ethynylene)-derived organic modules for controlling the assembly and covalent coupling of multiple modules. Rigid linear modules (LM) and tripoidal modules (TM) were functionalized with short oligonucleotides at each terminus. They can hybridize and thereby link up modules containing complementary sequences. Each terminus of the oligo(phenylene ethynylene) modules also consists of a salicylaldehyde moiety, which can form metal-salen complexes with other modules. The salicylaldehyde groups of two modules are brought in proximity when their adjoining DNA sequences are complementary, and they selectively form a manganese-salen complex in the presence of ethylenediamine and manganese acetate. The resulting structures consist of a matrix of linear and branched oligo(phenylene ethynylene)s which are linked by conjugated and rigid manganese-salen complexes. These nanostructures are potential conductors for applications in molecular electronics.

Acrylic Resins↗

Reliable prediction of T-cell epitopes using neural networks with novel sequence representations.

In this paper we describe an improved neural network method to predict T-cell class I epitopes. A novel input representation has been developed consisting of a combination of sparse encoding, Blosum encoding, and input derived from hidden Markov models. We demonstrate that the combination of several neural networks derived using different sequence-encoding schemes has a performance superior to neural networks derived using a single sequence-encoding scheme. The new method is shown to have a performance that is substantially higher than that of other methods. By use of mutual information calculations we show that peptides that bind to the HLA A*0204 complex display signal of higher order sequence correlations. Neural networks are ideally suited to integrate such higher order correlations when predicting the binding affinity. It is this feature combined with the use of several neural networks derived from different and novel sequence-encoding schemes and the ability of the neural network to be trained on data consisting of continuous binding affinities that gives the new method an improved performance. The difference in predictive performance between the neural network methods and that of the matrix-driven methods is found to be most significant for peptides that bind strongly to the HLA molecule, confirming that the signal of higher order sequence correlation is most strongly present in high-binding peptides. Finally, we use the method to predict T-cell epitopes for the genome of hepatitis C virus and discuss possible applications of the prediction method to guide the process of rational vaccine design.

Amino Acid Sequence↗

Selecting informative data for developing peptide-MHC binding predictors using a query by committee approach.

Strategies for selecting informative data points for training prediction algorithms are important, particularly when data points are difficult and costly to obtain. A Query by Committee (QBC) training strategy for selecting new data points uses the disagreement between a committee of different algorithms to suggest new data points, which most rationally complement existing data, that is, they are the most informative data points. In order to evaluate this QBC approach on a real-world problem, we compared strategies for selecting new data points. We trained neural network algorithms to obtain methods to predict the binding affinity of peptides binding to the MHC class I molecule, HLA-A2. We show that the QBC strategy leads to a higher performance than a baseline strategy where new data points are selected at random from a pool of available data. Most peptides bind HLA-A2 with a low affinity, and as expected using a strategy of selecting peptides that are predicted to have high binding affinities also lead to more accurate predictors than the base line strategy. The QBC value is shown to correlate with the measured binding affinity. This demonstrates that the different predictors can easily learn if a peptide will fail to bind, but often conflict in predicting if a peptide binds. Using a carefully constructed computational setup, we demonstrate that selecting peptides with a high QBC performs better than low QBC peptides independently from binding affinity. When predictors are trained on a very limited set of data they cannot be expected to disagree in a meaningful way and we find a data limit below which the QBC strategy fails. Finally, it should be noted that data selection strategies similar to those used here might be of use in other settings in which generation of more data is a costly process.

Algorithms↗

From lanosterol to cholesterol: structural evolution and differential effects on lipid bilayers.

Cholesterol is an important molecular component of the plasma membranes of mammalian cells. Its precursor in the sterol biosynthetic pathway, lanosterol, has been argued by Konrad Bloch (Bloch, K. 1965. Science. 150:19-28; 1983. CRC Crit. Rev. Biochem. 14:47-92; 1994. Blonds in Venetian Paintings, the Nine-Banded Armadillo, and Other Essays in Biochemistry. Yale University Press, New Haven, CT.) to also be a precursor in the molecular evolution of cholesterol. We present a comparative study of the effects of cholesterol and lanosterol on molecular conformational order and phase equilibria of lipid-bilayer membranes. By using deuterium NMR spectroscopy on multilamellar lipid-sterol systems in combination with Monte Carlo simulations of microscopic models of lipid-sterol interactions, we demonstrate that the evolution in the molecular chemistry from lanosterol to cholesterol is manifested in the model lipid-sterol membranes by an increase in the ability of the sterols to promote and stabilize a particular membrane phase, the liquid-ordered phase, and to induce collective order in the acyl-chain conformations of lipid molecules. We also discuss the biological relevance of our results, in particular in the context of membrane domains and rafts.

Calorimetry, Differential Scanning↗

DNA-directed coupling of organic modules by multiple parallel reductive aminations and subsequent cleavage of selected DNA sequences.

A new method for DNA-directed assembly of organic modules by multiple parallel reductive aminations is presented. Linear oligonucleotide-functionalized modules (LOMs) consist of a rigid oligo(phenylene ethynylene) backbone with two salicylaldehyde termini, and each terminus is conjugated with an oligonucleotide sequence. The stability of the tetrahydrosalen-linked modules toward elevated temperature, low pH, nucleophiles, and metal chelators is studied and compared to the analogous metal-salen-linked modules. A linear oligonucleotide-functionalized disulfide-linked module (LOSM) containing cleavable linkers between the organic module and the two DNA sequences is coupled by DNA-directed reductive aminations to non-modified LOM modules. This enables selective cleavage of the DNA strands of a central module in a structure consisting of three modules, and the reactions are analyzed by electrophoresis and 32P-labeling of one of the DNA sequences of the central LOSM.

Amination↗