Search PubMed⌕ Search

Biomedical subjects

Rainer Spang

Publications and source records attributed to Rainer Spang.

At least 19 recordsLinked to original sources

Quantifying uncertainty of predictions from cancer progression models.

MOTIVATION: Cancer progresses through the accumulation of genomic events. Cancer progression models such as Mutual Hazard Networks (MHNs) describe this dynamic, enabling prediction of temporal event positions and patient-specific risks of acquiring mutations. However, current MHN analyses rely on single most likely models and do not quantify the uncertainty inherent to parameter estimation. Assessing forecast stability is essential before using them to anticipate treatment-relevant mutations, adapt targeted therapies, or prioritize monitoring of patients at elevated progression risk. RESULTS: We address a key prerequisite for the responsible clinical use of cancer progression models by making MHN-derived predictions uncertainty-aware. We present a Bayesian framework for MHN that uses Markov Chain Monte Carlo to sample from the posterior distributions of model parameters and derived predictions. For practical use we implemented the Random-Walk Metropolis, Metropolis-Adjusted Langevin Algorithm (MALA), and simplified manifold MALA samplers as part of the existing mhn Python package. Only MALA and smMALA were successful in sampling from MHN posteriors, with MALA performing best. While most MHN parameters and predictions showed low posterior variance, a small subset displayed greater variability across the posterior distribution. This differentiation cannot be obtained from a single most likely model, emphasizing the need for uncertainty quantification, especially in clinical contexts. As an illustrative example, posterior sampling identified a subgroup of STK11$-$, KRAS$+$ lung adenocarcinoma patients with a high predicted short-term risk-with low variance across posterior samples-to develop an STK11 mutation. This subgroup exhibited poorer survival under immunotherapy, resembling patterns observed in STK11+ patients. AVAILABILITY AND IMPLEMENTATION: Our implementation is part of version 1.2.0 of the mhn package (https://github.com/spang-lab/LearnMHN). All analyses including the code to produce all figures in this article can be found under https://github.com/huy29433/MCMC-sampling-for-MHN (https://doi.org/10.5281/zenodo.21160219).

Humans↗

Selecting normalization genes for small diagnostic microarrays.

BACKGROUND: Normalization of gene expression microarrays carrying thousands of genes is based on assumptions that do not hold for diagnostic microarrays carrying only few genes. Thus, applying standard microarray normalization strategies to diagnostic microarrays causes new normalization problems. RESULTS: In this paper we point out the differences of normalizing large microarrays and small diagnostic microarrays. We suggest to include additional normalization genes on the small diagnostic microarrays and propose two strategies for selecting them from genomewide microarray studies. The first is a data driven univariate selection of normalization genes. The second is multivariate and based on finding a balanced diagnostic signature. Finally, we compare both methods to standard normalization protocols known from large microarrays. CONCLUSION: Not including additional genes for normalization on small microarrays leads to a loss of diagnostic information. Using house keeping genes from the literature for normalization fails to work for certain datasets. While a data driven selection of additional normalization genes works well, the best results were obtained using a balanced signature.

Algorithms↗

Expression of late cell cycle genes and an increased proliferative capacity characterize very early relapse of childhood acute lymphoblastic leukemia.

PURPOSE: In childhood acute lymphoblastic leukemia (ALL), approximately 25% of patients suffer from relapse. In recurrent disease, despite intensified therapy, overall cure rates of 40% remain unsatisfactory and survival rates are particularly poor in certain subgroups. The probability of long-term survival after relapse is predicted from well-established prognostic factors (i.e., time and site of relapse, immunophenotype, and minimal residual disease). However, the underlying biological determinants of these prognostic factors remain poorly understood. EXPERIMENTAL DESIGN: Aiming at identifying molecular pathways associated with these clinically well-defined prognostic factors, we did gene expression profiling on 60 prospectively collected samples of first relapse patients enrolled on the relapse trial ALL-REZ BFM 2002 of the Berlin-Frankfurt-Münster study group. RESULTS: We show here that patients with very early relapse of ALL are characterized by a distinctive gene expression pattern. We identified a set of 83 genes differentially expressed in very early relapsed ALL compared with late relapsed disease. The vast majority of genes were up-regulated and many were late cell cycle genes with a function in mitosis. In addition, samples from patients with very early relapse showed a significant increase in the percentage of S and G(2)-M phase cells and this correlated well with the expression level of cell cycle genes. CONCLUSIONS: Very early relapse of ALL is characterized by an increased proliferative capacity of leukemic blasts and up-regulated mitotic genes. The latter suggests that novel drugs, targeting late cell cycle proteins, might be beneficial for these patients that typically face a dismal prognosis.

Cell Cycle↗

OrderedList--a bioconductor package for detecting similarity in ordered gene lists.

UNLABELLED: OrderedList is a Bioconductor compliant package for meta-analysis based on ordered gene lists like those resulting from differential gene expression analysis. Our package quantifies the similarity between gene lists. The significance of the similarity score is estimated from random scores computed on perturbed data. OrderedList illustrates list similarity in intuitive plots and determines the score-driving genes for further analysis. AVAILABILITY: http://www.bioconductor.org CONTACT: claudio.lottaz@molgen.mpg.de SUPPLEMENTARY INFORMATION: Please visit our webpage on http://compdiag.molgen.mpg.de/software.

Algorithms↗

A biologic definition of Burkitt's lymphoma from transcriptional and genomic profiling.

BACKGROUND: The distinction between Burkitt's lymphoma and diffuse large-B-cell lymphoma is unclear. We used transcriptional and genomic profiling to define Burkitt's lymphoma more precisely and to distinguish subgroups in other types of mature aggressive B-cell lymphomas. METHODS: We performed gene-expression profiling using Affymetrix U133A GeneChips with RNA from 220 mature aggressive B-cell lymphomas, including a core group of 8 Burkitt's lymphomas that met all World Health Organization (WHO) criteria. A molecular signature for Burkitt's lymphoma was generated, and chromosomal abnormalities were detected with interphase fluorescence in situ hybridization and array-based comparative genomic hybridization. RESULTS: We used the molecular signature for Burkitt's lymphoma to identify 44 cases: 11 had the morphologic features of diffuse large-B-cell lymphomas, 4 were unclassifiable mature aggressive B-cell lymphomas, and 29 had a classic or atypical Burkitt's morphologic appearance. Also, five did not have a detectable IG-myc Burkitt's translocation, whereas the others contained an IG-myc fusion, mostly in simple karyotypes. Of the 176 lymphomas without the molecular signature for Burkitt's lymphoma, 155 were diffuse large-B-cell lymphomas. Of these 155 cases, 21 percent had a chromosomal breakpoint at the myc locus associated with complex chromosomal changes and an unfavorable clinical course. CONCLUSIONS: Our molecular definition of Burkitt's lymphoma clarifies and extends the spectrum of the WHO criteria for Burkitt's lymphoma. In mature aggressive B-cell lymphomas without a gene signature for Burkitt's lymphoma, chromosomal breakpoints at the myc locus were associated with an adverse clinical outcome.

Algorithms↗

Automated in-silico detection of cell populations in flow cytometry readouts and its application to leukemia disease monitoring.

BACKGROUND: Identification of minor cell populations, e.g. leukemic blasts within blood samples, has become increasingly important in therapeutic disease monitoring. Modern flow cytometers enable researchers to reliably measure six and more variables, describing cellular size, granularity and expression of cell-surface and intracellular proteins, for thousands of cells per second. Currently, analysis of cytometry readouts relies on visual inspection and manual gating of one- or two-dimensional projections of the data. This procedure, however, is labor-intensive and misses potential characteristic patterns in higher dimensions. RESULTS: Leukemic samples from patients with acute lymphoblastic leukemia at initial diagnosis and during induction therapy have been investigated by 4-color flow cytometry. We have utilized multivariate classification techniques, Support Vector Machines (SVM), to automate leukemic cell detection in cytometry. Classifiers were built on conventionally diagnosed training data. We assessed the detection accuracy on independent test data and analyzed marker expression of incongruently classified cells. SVM classification can recover manually gated leukemic cells with 99.78% sensitivity and 98.87% specificity. CONCLUSION: Multivariate classification techniques allow for automating cell population detection in cytometry readouts for diagnostic purposes. They potentially reduce time, costs and arbitrariness associated with these procedures. Due to their multivariate classification rules, they also allow for the reliable detection of small cell populations.

Computational Biology↗

Similarities of ordered gene lists.

MOTIVATION: Many applications of microarray technology in clinical cancer studies aim at detecting molecular features for refined diagnosis. In this paper, we follow an opposite rationale: we try to identify common molecular features shared by phenotypically distinct types of cancer using a meta-analysis of several microarray studies. We present a novel algorithm to uncover that two lists of differentially expressed genes are similar, even if these similarities are not apparent to the eye. The method is based on the ordering in the lists. RESULTS: In a meta-analysis of five clinical microarray studies we were able to detect significant similarities in five of the ten possible comparisons of ordered gene lists. We included studies, where not a single gene can be significantly associated to outcome. The detection of significant similarities of gene lists from different microarray studies is a novel and promising approach. It has the potential to improve upon specialized cancer studies by exploring the power of several studies in one single analysis. Our method is complementary to previous methods in that it does not rely on strong effects of differential gene expression in a single study but on consistent ones across multiple studies.

Algorithms↗

Non-transcriptional pathway features reconstructed from secondary effects of RNA interference.

MOTIVATION: Cellular signaling pathways, which are not modulated on a transcriptional level, cannot be directly deduced from expression profiling experiments. The situation changes, when external interventions such as RNA interference or gene knock-outs come into play. Even if the expression of the signaling genes is not changed, secondary effects in downstream genes shed light on the pathway, and allow partial reconstruction of its topology. RESULTS: We introduce an algorithm to infer non-transcriptional pathway features based on differential gene expression in silencing assays. We demonstrate the power of our algorithm in the controlled setting of simulation studies, and explain its practical use in the context of an RNA interference dataset investigating the response to microbial challenge in Drosophila melanogaster.

Algorithms↗

stam--a Bioconductor compliant R package for structured analysis of microarray data.

BACKGROUND: Genome wide microarray studies have the potential to unveil novel disease entities. Clinically homogeneous groups of patients can have diverse gene expression profiles. The definition of novel subclasses based on gene expression is a difficult problem not addressed systematically by currently available software tools. RESULTS: We present a computational tool for semi-supervised molecular disease entity detection. It automatically discovers molecular heterogeneities in phenotypically defined disease entities and suggests alternative molecular sub-entities of clinical phenotypes. This is done using both gene expression data and functional gene annotations. We provide stam, a Bioconductor compliant software package for the statistical programming environment R. We demonstrate that our tool detects gene expression patterns, which are characteristic for only a subset of patients from an established disease entity. We call such expression patterns molecular symptoms. Furthermore, stam finds novel sub-group stratifications of patients according to the absence or presence of molecular symptoms. CONCLUSION: Our software is easy to install and can be applied to a wide range of datasets. It provides the potential to reveal so far indistinguishable patient sub-groups of clinical relevance.

Calibration↗

Early diagnostic marker panel determination for microarray based clinical studies.

We present a novel, cost efficient two-phase design for predictive clinical gene expression studies: early marker panel determination (EMPD). In Phase-1, genome-wide microarrays are used only for a small number of individual patient samples. From this Phase-1 data a panel of marker genes is derived. In Phase-2, the expression values of these marker panel genes are measured for a large group of patients and a predictive classification model is learned from this data. Phase-2 does not require the use of expensive whole genome microarrays, thus making EMPD a cost efficient alternative for current trials. The expected performance loss of EMPD is compared to designs which use genome-wide microarrays for all patients. We also examine the trade-off between the number of patients included in Phase-1 and the number of marker genes required in Phase-2. By analysis of five published datasets we find that in Phase-1 already 16 patients per group are sufficient to determine a suitable marker panel of 10 genes, and that this early decision compromises the final performance only marginally.

Journal Article↗

twilight; a Bioconductor package for estimating the local false discovery rate.

UNLABELLED: twilight is a Bioconductor compatible package for analysing the statistical significance of differentially expressed genes. It is based on the concept of the local false discovery rate (FDR), a generalization of the frequently used global FDR. twilight implements the heuristic search algorithm for estimating the local FDR introduced in our earlier work. In addition to the raw significance measures, it produces diagnostic plots, which provide insight into the extent of differential expression across genes. AVAILABILITY: http://www.bioconductor.org CONTACT: stefanie.scheid@molgen.mpg.de SUPPLEMENTARY INFORMATION: Please visit our software webpage on http://compdiag.molgen.mpg.de/software.

Algorithms↗

Molecular decomposition of complex clinical phenotypes using biologically structured analysis of microarray data.

MOTIVATION: Today, the characterization of clinical phenotypes by gene-expression patterns is widely used in clinical research. If the investigated phenotype is complex from the molecular point of view, new challenges arise and these have not been addressed systematically. For instance, the same clinical phenotype can be caused by various molecular disorders, such that one observes different characteristic expression patterns in different patients. RESULTS: In this paper we describe a novel algorithm called Structured Analysis of Microarrays (StAM), which accounts for molecular heterogeneity of complex clinical phenotypes. Our algorithm goes beyond established methodology in several aspects: in addition to the expression data, it exploits functional annotations from the Gene Ontology database to build biologically focussed classifiers. These are used to uncover potential molecular disease subentities and associate them to biological processes without compromising overall prediction accuracy. AVAILABILITY: Bioconductor compliant R package SUPPLEMENTARY INFORMATION: Complete analyses are available at http://compdiag.molgen.mpg.de/supplements/lottaz05.

Biomarkers, Tumor↗

Detecting common gene expression patterns in multiple cancer outcome entities.

Most oncological microarray studies focus on molecular distinctions in different cancer entities. Recently, researchers started using microarrays for investigating molecular commonalities of multiple cancer types. This poses novel bioinformatics challenges. In this paper we describe a method that detects common molecular mechanisms in different cancer entities. The method extends previously described concepts by introducing Meta-Analysis Pattern Matches. In an analysis of four prognostic cancer studies, involving breast cancer, leukemia, and mesothelioma, we are able to identify 42 genes that show consistent up- or down-regulation in patients with a poor disease outcome. These genes complement the set of previously published candidates for universal prognostic markers in cancer.

Algorithms↗

Finding disease specific alterations in the co-expression of genes.

MOTIVATION: Standard analysis routines for microarray data aim at differentially expressed genes. In this paper, we address the complementary problem of detecting sets of differentially co-expressed genes in two phenotypically distinct sets of expression profiles. RESULTS: We introduce a score for differential co-expression and suggest a computationally efficient algorithm for finding high scoring sets of genes. The use of our novel method is demonstrated in the context of simulations and on real expression data from a clinical study.

Algorithms↗

Differential myocardial gene expression in the development and rescue of murine heart failure.

Numerous murine models of heart failure (HF) have been described, many of which develop progressive deterioration of cardiac function. We have recently demonstrated that several of these can be "rescued" or prevented by transgenic cardiac expression of a peptide inhibitor of the beta-adrenergic receptor kinase (betaARKct). To uncover genomic changes associated with cardiomyopathy and/or its phenotypic rescue by the betaARKct, oligonucleotide microarray analysis of left ventricular (LV) gene expression was performed in a total of 53 samples, including 12 each of Normal, HF, and Rescue. Multiple statistical analyses demonstrated significant differences between all groups and further demonstrated that betaARKct Rescue returned gene expression toward that of Normal. In our statistical analyses, we found that the HF phenotype is blindly predictable based solely on gene expression profile. To investigate the progression of HF, LV gene expression was determined in young mice with mildly diminished cardiac function and in older mice with severely impaired cardiac function. Interestingly, mild and advanced HF mice shared similar gene expression profiles, and importantly, the mild HF mice were predicted as having a HF phenotype when blindly subjected to our predictive model described above. These data not only validate our predictive model but further demonstrate that, in these mice, the HF gene expression profile appears to already be set in the early stages of HF progression. Thus we have identified methodologies that have the potential to be used for predictive genomic profiling of cardiac phenotype, including cardiovascular disease.

Animals↗

Expression profiling of human idiopathic dilated cardiomyopathy.

OBJECTIVE: To investigate the global changes accompanying human dilated cardiomyopathy (DCM) we performed a large-scale expression screen using myocardial biopsies from a group of DCM patients with moderate heart failure. By hierarchical clustering and functional annotation of the deregulated genes we examined extensive changes in the cellular and molecular processes associated to DCM. METHODS: The expression profiles were obtained using a whole genome covering library (UniGene RZPD1) comprising 30336 cDNA clones and amplified RNA from myocardiac biopsies from 10 DCM patients in comparison to tissue samples from four non-failing, healthy donors. RESULTS: By setting stringent selection criteria 364 differentially expressed, sequence-verified non-redundant transcripts were identified with a false discovery rate of <0.001. Numerous genes and ESTs were identified representing previously recognised, as well as novel DCM-associated transcripts. Many of them were found to be upregulated and involved in cardiomyocyte energetics, muscle contraction or signalling. Two hundred and twenty-two deregulated transcripts were functionally annotated and hierarchically clustered providing an insight into the pathophysiology of DCM. Data was validated using the MLP-deficient mouse, in which several differentially expressed transcripts identified in the human DCM biopsies could be confirmed. CONCLUSIONS: We report the first genome-wide expression profile analysis using cardiac biopsies from DCM patients at various stages of the disease. Although there is a diversity of links between the cytoskeleton and the initiation of DCM, we speculate that genes implicated in intracellular signalling and in muscle contraction are associated with early stages of the disease. Altogether this study represents the most comprehensive and inclusive molecular portrait of human cardiomyopathy to date.

Adult↗

A novel approach to remote homology detection: jumping alignments.

We describe a new algorithm for protein classification and the detection of remote homologs. The rationale is to exploit both vertical and horizontal information of a multiple alignment in a well-balanced manner. This is in contrast to established methods such as profiles and profile hidden Markov models which focus on vertical information as they model the columns of the alignment independently and to family pairwise search which focuses on horizontal information as it treats given sequences separately. In our setting, we want to select from a given database of "candidate sequences" those proteins that belong to a given superfamily. In order to do so, each candidate sequence is separately tested against a multiple alignment of the known members of the superfamily by means of a new jumping alignment algorithm. This algorithm is an extension of the Smith-Waterman algorithm and computes a local alignment of a single sequence and a multiple alignment. In contrast to traditional methods, however, this alignment is not based on a summary of the individual columns of the multiple alignment. Rather, the candidate sequence is at each position aligned to one sequence of the multiple alignment, called the "reference sequence." In addition, the reference sequence may change within the alignment, while each such jump is penalized. To evaluate the discriminative quality of the jumping alignment algorithm, we compare it to profiles, profile hidden Markov models, and family pairwise search on a subset of the SCOP database of protein domains. The discriminative quality is assessed by median false positive counts (med-FP-counts). For moderate med-FP-counts, the number of successful searches with our method is considerably higher than with the competing methods.

Algorithms↗

Estimating amino acid substitution models: a comparison of Dayhoff's estimator, the resolvent approach and a maximum likelihood method.

Evolution of proteins is generally modeled as a Markov process acting on each site of the sequence. Replacement frequencies need to be estimated based on sequence alignments. Here we compare three approaches: First, the original method by Dayhoff, Schwartz, and Orcutt (1978) Atlas Protein Seq. Struc. 5:345-352, secondly, the resolvent method (RV) by Müller and Vingron (2000) J. Comput. Biol. 7(6):761-776, and finally a maximum likelihood approach (ML) developed in this paper. We evaluate the methods using a highly divergent and inhomogeneous set of sequence alignments as an input to the estimation procedure. ML is the method of choice for small sets of input data. Although the RV method is computationally much less demanding it performs only slightly worse than ML. Therefore, it is perfectly appropriate for large-scale applications.

Algorithms↗