Search PubMed⌕ Search

Biomedical subjects

Martin T Wells

Publications and source records attributed to Martin T Wells.

7 recordsLinked to original sources

PROLONG: penalized regression for outcome guided longitudinal omics analysis with network and group constraints.

MOTIVATION: There is a growing interest in longitudinal omics data paired with some longitudinal clinical outcome. Given a large set of continuous omics variables and some continuous clinical outcome, each measured for a few subjects at only a few time points, we seek to identify those variables that co-vary over time with the outcome. To motivate this problem we study a dataset with hundreds of urinary metabolites along with Tuberculosis mycobacterial load as our clinical outcome, with the objective of identifying potential biomarkers for disease progression. For such data clinicians usually apply simple linear mixed effects models which often lack power given the low number of replicates and time points. We propose a penalized regression approach on the first differences of the data that extends the lasso + Laplacian method [Li and Li (Network-constrained regularization and variable selection for analysis of genomic data. Bioinformatics 2008;24:1175-82.)] to a longitudinal group lasso + Laplacian approach. Our method, PROLONG, leverages the first differences of the data to increase power by pairing the consecutive time points. The Laplacian penalty incorporates the dependence structure of the variables, and the group lasso penalty induces sparsity while grouping together all contemporaneous and lag terms for each omic variable in the model. RESULTS: With an automated selection of model hyper-parameters, PROLONG correctly selects target metabolites with high specificity and sensitivity across a wide range of scenarios. PROLONG selects a set of metabolites from the real data that includes interesting targets identified during EDA. AVAILABILITY AND IMPLEMENTATION: An R package implementing described methods called "prolong" is available at https://github.com/stevebroll/prolong. Code snapshot available at 10.5281/zenodo.14804245.

Humans↗

Multiple endoplasmic reticulum-to-nucleus signaling pathways coordinate phospholipid metabolism with gene expression by distinct mechanisms.

In many organisms the coordinated synthesis of membrane lipids is controlled by feedback systems that regulate the transcription of target genes. However, a complete description of the transcriptional changes that accompany the remodeling of membrane phospholipids has not been reported. To identify metabolic signaling networks that coordinate phospholipid metabolism with gene expression, we profiled the sequential and temporal changes in genome-wide expression that accompany alterations in phospholipid metabolism induced by inositol supplementation in yeast. This analysis identified six distinct expression responses, which included phospholipid biosynthetic genes regulated by Opi1p, endoplasmic reticulum (ER) luminal protein folding chaperone and oxidoreductase genes regulated by the unfolded protein response pathway, lipid-remodeling genes regulated by Mga2p, as well as genes involved in ribosome biogenesis, cytosolic stress response, and purine and amino acid metabolism. We also report that the unfolded protein response pathway is rapidly inactivated by inositol supplementation and demonstrate that the response of the unfolded protein response pathway to inositol is separable from the response mediated by Opi1p. These data indicate that altering phospholipid metabolism produces signals that are relayed through numerous distinct ER-to-nucleus signaling pathways and, thereby, produce an integrated transcriptional response. We propose that these signals are generated in the ER by increased flux through the pathway of phosphatidylinositol synthesis.

Amino Acids↗

Multiplicative background correction for spotted microarrays to improve reproducibility.

We propose a simple approach, the multiplicative background correction, to solve a perplexing problem in spotted microarray data analysis: correcting the foreground intensities for the background noise, especially for spots with genes that are weakly expressed or not at all. The conventional approach, the additive background correction, directly subtracts the background intensities from foreground intensities. When the foreground intensities marginally dominate the background intensities, the additive background correction provides unreliable estimates of the differential gene expression levels and usually presents M-A plots with fishtails or fans. Unreliable additive background correction makes it preferable to ignore the background noise, which may increase the number of false positives. Based on the more realistic multiplicative assumption instead of the conventional additive assumption, we propose to logarithmically transform the intensity readings before the background correction, with the logarithmic transformation symmetrizing the skewed intensity readings. This approach not only precludes the fishtails and fans in the M-A plots, but provides highly reproducible background-corrected intensities for both strongly and weakly expressed genes. The superiority of the multiplicative background correction to the additive one as well as the no background correction is justified by publicly available self-hybridization datasets.

Arabidopsis↗

Bayesian normalization and identification for differential gene expression data.

Commonly accepted intensity-dependent normalization in spotted microarray studies takes account of measurement errors in the differential expression ratio but ignores measurement errors in the total intensity, although the definitions imply the same measurement error components are involved in both statistics. Furthermore, identification of differentially expressed genes is usually considered separately following normalization, which is statistically problematic. By incorporating the measurement errors in both total intensities and differential expression ratios, we propose a measurement-error model for intensity-dependent normalization and identification of differentially expressed genes. This model is also flexible enough to incorporate intra-array and inter-array effects. A Bayesian framework is proposed for the analysis of the proposed measurement-error model to avoid the potential risk of using the common two-step procedure. We also propose a Bayesian identification of differentially expressed genes to control the false discovery rate instead of the ad hoc thresholding of the posterior odds ratio. The simulation study and an application to real microarray data demonstrate promising results.

Bayes Theorem↗

Genome-wide analysis reveals inositol, not choline, as the major effector of Ino2p-Ino4p and unfolded protein response target gene expression in yeast.

In the yeast Saccharomyces cerevisiae, the transcription of many genes encoding enzymes of phospholipid biosynthesis are repressed in cells grown in the presence of the phospholipid precursors inositol and choline. A genome-wide approach using cDNA microarray technology was used to profile the changes in the expression of all genes in yeast that respond to the exogenous presence of inositol and choline. We report that the global response to inositol is completely distinct from the effect of choline. Whereas the effect of inositol on gene expression was primarily repressing, the effect of choline on gene expression was activating. Moreover, the combination of inositol and choline increased the number of repressed genes compared with inositol alone and enhanced the repression levels of a subset of genes that responded to inositol. In all, 110 genes were repressed in the presence of inositol and choline. Two distinct sets of genes exhibited differential expression in response to inositol or the combination of inositol and choline in wild-type cells. One set of genes contained the UASINO sequence and were bound by Ino2p and Ino4p. Many of these genes were also negatively regulated by OPI1, suggesting a common regulatory mechanism for Ino2p, Ino4p, and Opi1p. Another nonoverlapping set of genes was coregulated by the unfolded protein response pathway, an ER-localized stress response pathway, but was not dependent on OPI1 and did not show further repression when choline was present together with inositol. These results suggest that inositol is the major effector of target gene expression, whereas choline plays a minor role.

Basic Helix-Loop-Helix Proteins↗

Mapping multiple Quantitative Trait Loci by Bayesian classification.

We developed a classification approach to multiple quantitative trait loci (QTL) mapping built upon a Bayesian framework that incorporates the important prior information that most genotypic markers are not cotransmitted with a QTL or their QTL effects are negligible. The genetic effect of each marker is modeled using a three-component mixture prior with a class for markers having negligible effects and separate classes for markers having positive or negative effects on the trait. The posterior probability of a marker's classification provides a natural statistic for evaluating credibility of identified QTL. This approach performs well, especially with a large number of markers but a relatively small sample size. A heat map to visualize the results is proposed so as to allow investigators to be more or less conservative when identifying QTL. We validated the method using a well-characterized data set for barley heading values from the North American Barley Genome Mapping Project. Application of the method to a new data set revealed sex-specific QTL underlying differences in glucose-6-phosphate dehydrogenase enzyme activity between two Drosophila species. A simulation study demonstrated the power of this approach across levels of trait heritability and when marker data were sparse.

Animals↗

Mathematical model of Listeria monocytogenes cross-contamination in a fish processing plant.

Listeriosis is a foodborne disease caused by the bacterium Listeria monocytogenes. The food industry and government agencies devote considerable resources to reducing contamination of ready-to-eat foods with L. monocytogenes. Because inactivation treatments can effectively eliminate L. monocytogenes present on raw materials, postprocessing cross-contamination from the processing plant environment appears to be responsible for most L. monocytogenes food contamination events. An improved understanding of cross-contamination pathways is critical to preventing L. monocytogenes contamination. Therefore, a plant-specific mathematical model of L. monocytogenes cross-contamination was developed, which described the transmission of L. monocytogenes contamination among food, food contact surfaces, employees' gloves, and the environment. A smoked fish processing plant was used as a model system. The model estimated that 10.7% (5th and 95th percentile, 0.05% and 22.3%, respectively) of food products in a lot are likely to be contaminated with L. monocytogenes. Sensitivity analysis identified the most significant input parameters as the frequency with which employees' gloves contact food and food contact surfaces, and the frequency of changing gloves. Scenario analysis indicated that the greatest reduction of the within-lot prevalence of contaminated food products can be achieved if the raw material entering the plant is free of contamination. Zero contamination of food products in a lot was possible but rare. This model could be used in a risk assessment to quantify the potential public health benefits of in-plant control strategies to reduce cross-contamination.

Animals↗