Search PubMed⌕ Search

Biomedical subjects

Jens Ledet Jensen

Publications and source records attributed to Jens Ledet Jensen.

7 recordsLinked to original sources

Validation of the use of DNA pools and primer extension in association studies of sporadic colorectal cancer for selection of candidate SNPs.

Colorectal cancer (CRC) is a multifactorial disease that involves both lifestyle and genetic factors. To identify single nucleotide polymorphisms (SNPs) associated with sporadic CRC, we used pooled DNA samples representing 230 cases with sporadic CRC and 540 controls. The allele frequency of the SNPs was estimated in the two pools using a genotyping method based on primer extension and capillary electrophoresis (CE). The sensitivity of the method was high, which permitted the detection of an odds ratio (OR) of 1.5. Validation of the method showed that it is robust, linear, sensitive, and reproducible. Of the 224 SNPs investigated, 20 potential candidates associated with CRC were identified, including IL6 -174G>C (g.22062318G>C), XRCC1 c.685 C>T (p.Arg194Trp), PPARGC1A g.92945042C>T (3'UTR 96516), GSTP1 c.342A>C (p.Ile105Val), GSTM1 c.573C>G (p.Lys173Asn), and SULT1A1 g.19934792G>A (p.Arg213His). All were borderline significant, and none were significant at the 5% level. A high number of the SNPs (40%) were not polymorphic in our population. We conclude that instead of looking for single risk factors, investigators should examine individual combinations of potential risk factors to clarify the genetic predisposition to CRC.

Aged↗

Statistical inference in evolutionary models of DNA sequences via the EM algorithm.

We describe statistical inference in continuous time Markov processes of DNA sequences related by a phylogenetic tree. The maximum likelihood estimator can be found by the expectation maximization (EM) algorithm and an expression for the information matrix is also derived. We provide explicit analytical solutions for the EM algorithm and information matrix.

Journal Article↗

Bayesian coestimation of phylogeny and sequence alignment.

BACKGROUND: Two central problems in computational biology are the determination of the alignment and phylogeny of a set of biological sequences. The traditional approach to this problem is to first build a multiple alignment of these sequences, followed by a phylogenetic reconstruction step based on this multiple alignment. However, alignment and phylogenetic inference are fundamentally interdependent, and ignoring this fact leads to biased and overconfident estimations. Whether the main interest be in sequence alignment or phylogeny, a major goal of computational biology is the co-estimation of both. RESULTS: We developed a fully Bayesian Markov chain Monte Carlo method for coestimating phylogeny and sequence alignment, under the Thorne-Kishino-Felsenstein model of substitution and single nucleotide insertion-deletion (indel) events. In our earlier work, we introduced a novel and efficient algorithm, termed the "indel peeling algorithm", which includes indels as phylogenetically informative evolutionary events, and resembles Felsenstein's peeling algorithm for substitutions on a phylogenetic tree. For a fixed alignment, our extension analytically integrates out both substitution and indel events within a proper statistical model, without the need for data augmentation at internal tree nodes, allowing for efficient sampling of tree topologies and edge lengths. To additionally sample multiple alignments, we here introduce an efficient partial Metropolized independence sampler for alignments, and combine these two algorithms into a fully Bayesian co-estimation procedure for the alignment and phylogeny problem. Our approach results in estimates for the posterior distribution of evolutionary rate parameters, for the maximum a-posteriori (MAP) phylogenetic tree, and for the posterior decoding alignment. Estimates for the evolutionary tree and multiple alignment are augmented with confidence estimates for each node height and alignment column. Our results indicate that the patterns in reliability broadly correspond to structural features of the proteins, and thus provides biologically meaningful information which is not existent in the usual point-estimate of the alignment. Our methods can handle input data of moderate size (10-20 protein sequences, each 100-200 bp), which we analyzed overnight on a standard 2 GHz personal computer. CONCLUSION: Joint analysis of multiple sequence alignment, evolutionary trees and additional evolutionary parameters can be now done within a single coherent statistical framework.

Algorithms↗

Applications of hidden Markov models for characterization of homologous DNA sequences with a common gene.

Identifying and characterizing the structure in genome sequences is one of the principal challenges in modern molecular biology, and comparative genomics offers a powerful tool. In this paper, we introduce a hidden Markov model that allows a comparative analysis of multiple sequences related by a phylogenetic tree, and we present an efficient method for estimating the parameters of the model. The model integrates structure prediction methods for one sequence, statistical multiple alignment methods, and phylogenetic information. This unified model is particularly useful for a detailed characterization of DNA sequences with a common gene. We illustrate the model on a variety of homologous sequences.

Agrobacterium tumefaciens↗

Normalization of real-time quantitative reverse transcription-PCR data: a model-based variance estimation approach to identify genes suited for normalization, applied to bladder and colon cancer data sets.

Accurate normalization is an absolute prerequisite for correct measurement of gene expression. For quantitative real-time reverse transcription-PCR (RT-PCR), the most commonly used normalization strategy involves standardization to a single constitutively expressed control gene. However, in recent years, it has become clear that no single gene is constitutively expressed in all cell types and under all experimental conditions, implying that the expression stability of the intended control gene has to be verified before each experiment. We outline a novel, innovative, and robust strategy to identify stably expressed genes among a set of candidate normalization genes. The strategy is rooted in a mathematical model of gene expression that enables estimation not only of the overall variation of the candidate normalization genes but also of the variation between sample subgroups of the sample set. Notably, the strategy provides a direct measure for the estimated expression variation, enabling the user to evaluate the systematic error introduced when using the gene. In a side-by-side comparison with a previously published strategy, our model-based approach performed in a more robust manner and showed less sensitivity toward coregulation of the candidate normalization genes. We used the model-based strategy to identify genes suited to normalize quantitative RT-PCR data from colon cancer and bladder cancer. These genes are UBC, GAPD, and TPT1 for the colon and HSPCB, TEGT, and ATP5B for the bladder. The presented strategy can be applied to evaluate the suitability of any normalization gene candidate in any kind of experimental design and should allow more reliable normalization of RT-PCR data.

Biomarkers, Tumor↗

Recursions for statistical multiple alignment.

Algorithms are presented that allow the calculation of the probability of a set of sequences related by a binary tree that have evolved according to the Thorne-Kishino-Felsenstein model for a fixed set of parameters. The algorithms are based on a Markov chain generating sequences and their alignment at nodes in a tree. Depending on whether the complete realization of this Markov chain is decomposed into the first transition and the rest of the realization or the last transition and the first part of the realization, two kinds of recursions are obtained that are computationally similar but probabilistically different. The running time of the algorithms is O(Pi id=1 Li), where Li is the length of the ith observed sequences and d is the number of sequences. An alternative recursion is also formulated that uses only a Markov chain involving the inner nodes of a tree.

Algorithms↗

Identifying distinct classes of bladder carcinoma using microarrays.

Bladder cancer is a common malignant disease characterized by frequent recurrences. The stage of disease at diagnosis and the presence of surrounding carcinoma in situ are important in determining the disease course of an affected individual. Despite considerable effort, no accepted immunohistological or molecular markers have been identified to define clinically relevant subsets of bladder cancer. Here we report the identification of clinically relevant subclasses of bladder carcinoma using expression microarray analysis of 40 well characterized bladder tumors. Hierarchical cluster analysis identified three major stages, Ta, T1 and T2-4, with the Ta tumors further classified into subgroups. We built a 32-gene molecular classifier using a cross-validation approach that was able to classify benign and muscle-invasive tumors with close correlation to pathological staging in an independent test set of 68 tumors. The classifier provided new predictive information on disease progression in Ta tumors compared with conventional staging (P < 0.005). To delineate non-recurring Ta tumors from frequently recurring Ta tumors, we analyzed expression patterns in 31 tumors by applying a supervised learning classification methodology, which classified 75% of the samples correctly (P < 0.006). Furthermore, gene expression profiles characterizing each stage and subtype identified their biological properties, producing new potential targets for therapy.

Disease Progression↗