Search PubMed⌕ Search

Biomedical subjects

Paul Marjoram

Publications and source records attributed to Paul Marjoram.

At least 19 recordsLinked to original sources

Estimating recombination rates from single-nucleotide polymorphisms using summary statistics.

We describe a novel method for jointly estimating crossing-over and gene-conversion rates from population genetic data using summary statistics. The performance of our method was tested on simulated data sets and compared with the composite-likelihood method of R. R. Hudson. For several realistic parameter values, the new method performed similarly to the composite-likelihood approach for estimating crossing-over rates and better when estimating gene-conversion rates. We used our method to analyze a human data set recently genotyped by Perlegen Sciences.

Computer Simulation↗

Cluster analysis for DNA methylation profiles having a detection threshold.

BACKGROUND: DNA methylation, a molecular feature used to investigate tumor heterogeneity, can be measured on many genomic regions using the MethyLight technology. Due to the combination of the underlying biology of DNA methylation and the MethyLight technology, the measurements, while being generated on a continuous scale, have a large number of 0 values. This suggests that conventional clustering methodology may not perform well on this data. RESULTS: We compare performance of existing methodology (such as k-means) with two novel methods that explicitly allow for the preponderance of values at 0. We also consider how the ability to successfully cluster such data depends upon the number of informative genes for which methylation is measured and the correlation structure of the methylation values for those genes. We show that when data is collected for a sufficient number of genes, our models do improve clustering performance compared to methods, such as k-means, that do not explicitly respect the supposed biological realities of the situation. CONCLUSION: The performance of analysis methods depends upon how well the assumptions of those methods reflect the properties of the data being analyzed. Differing technologies will lead to data with differing properties, and should therefore be analyzed differently. Consequently, it is prudent to give thought to what the properties of the data are likely to be, and which analysis method might therefore be likely to best capture those properties.

Artificial Intelligence↗

Inferring population parameters from single-feature polymorphism data.

This article is concerned with a statistical modeling procedure to call single-feature polymorphisms from microarray experiments. We use this new type of polymorphism data to estimate the mutation and recombination parameters in a population. The mutation parameter can be estimated via the number of single-feature polymorphisms called in the sample. For the recombination parameter, a two-feature sampling distribution is derived in a way analogous to that for the two-locus sampling distribution with SNP data. The approximate-likelihood approach using the two-feature sampling distribution is examined and found to work well. A coalescent simulation study is used to investigate the accuracy and robustness of our method. Our approach allows the utilization of single-feature polymorphism data for inference in population genetics.

Arabidopsis↗

Fast "coalescent" simulation.

BACKGROUND: The amount of genome-wide molecular data is increasing rapidly, as is interest in developing methods appropriate for such data. There is a consequent increasing need for methods that are able to efficiently simulate such data. In this paper we implement the sequentially Markovian coalescent algorithm described by McVean and Cardin and present a further modification to that algorithm which slightly improves the closeness of the approximation to the full coalescent model. The algorithm ignores a class of recombination events known to affect the behavior of the genealogy of the sample, but which do not appear to affect the behavior of generated samples to any substantial degree. RESULTS: We show that our software is able to simulate large chromosomal regions, such as those appropriate in a consideration of genome-wide data, in a way that is several orders of magnitude faster than existing coalescent algorithms. CONCLUSION: This algorithm provides a useful resource for those needing to simulate large quantities of data for chromosomal-length regions using an approach that is much more efficient than traditional coalescent models.

Algorithms↗

Association mapping with single-feature polymorphisms.

We develop methods for exploiting "single-feature polymorphism" data, generated by hybridizing genomic DNA to oligonucleotide expression arrays. Our methods enable the use of such data, which can be regarded as very high density, but imperfect, polymorphism data, for genomewide association or linkage disequilibrium mapping. We use a simulation-based power study to conclude that our methods should have good power for organisms like Arabidopsis thaliana, in which linkage disequilibrium is extensive, the reason being that the noisiness of single-feature polymorphism data is more than compensated for by their great number. Finally, we show how power depends on the accuracy with which single-feature polymorphisms are called.

Algorithms↗

Modern computational approaches for analysing molecular genetic variation data.

An explosive growth is occurring in the quantity, quality and complexity of molecular variation data that are being collected. Historically, such data have been analysed by using model-based methods. Models are useful for sharpening intuition, for explanation and for prediction: they add to our understanding of how the data were formed, and they can provide quantitative answers to questions of interest. We outline some of these model-based approaches, including the coalescent, and discuss the applicability of the computational methods that are necessary given the highly complex nature of current and future data sets.

Computational Biology↗

Fine mapping--19th century style.

BACKGROUND: There is great interest in the use of computationally intensive methods for fine mapping of marker data. In this paper we develop methods based upon ideas originally proposed 100 years ago in the context of spatial clustering. METHODS: We use spatial clustering of haplotypes as a low-dimensional surrogate for the unobserved genealogy underlying a set of genotype data. In doing so we hope to avoid the computational complexity inherent in explicitly modelling details of the ancestry of the sample, while at the same time capturing the key correlations induced by that ancestry at a much lower computational cost. RESULTS: We benchmark our methods using the simulated Genetic Analysis Workshop 14 data, using 100 replicates of 4 phenotypes to indicate the power of our method. When a functional mutation relating to a trait is actually present, we find evidence for that mutation in 97 out of 100 replicates, on average. CONCLUSION: Our results show that our method has the ability to accurately infer the location of functional mutations from unphased genotype data.

Congresses as Topic↗

Relative influences of crossing over and gene conversion on the pattern of linkage disequilibrium in Arabidopsis thaliana.

In this article we infer the rates of gene conversion and crossing over in Arabidopsis thaliana from population genetic data. Our data set is a genomewide survey consisting of 1347 fragments of length 600 bp sequenced in 96 accessions. It has several orders of magnitude more markers than any previous nonhuman study. This allows for more accurate inference as well as a detailed comparison between theoretical expectations and observations. Our methodology is specifically set to account for deviations such as recurrent mutations or a skewed frequency spectrum. We found that even if some components of the model clearly do not fit, the pattern of LD conforms to theoretical expectations quite well. The ratio of gene conversion to crossing over is estimated to be around one. We also find evidence for fine-scale variations of the crossing-over rate.

Arabidopsis↗

Genome-wide association mapping in Arabidopsis identifies previously known flowering time and pathogen resistance genes.

There is currently tremendous interest in the possibility of using genome-wide association mapping to identify genes responsible for natural variation, particularly for human disease susceptibility. The model plant Arabidopsis thaliana is in many ways an ideal candidate for such studies, because it is a highly selfing hermaphrodite. As a result, the species largely exists as a collection of naturally occurring inbred lines, or accessions, which can be genotyped once and phenotyped repeatedly. Furthermore, linkage disequilibrium in such a species will be much more extensive than in a comparable outcrossing species. We tested the feasibility of genome-wide association mapping in A. thaliana by searching for associations with flowering time and pathogen resistance in a sample of 95 accessions for which genome-wide polymorphism data were available. In spite of an extremely high rate of false positives due to population structure, we were able to identify known major genes for all phenotypes tested, thus demonstrating the potential of genome-wide association mapping in A. thaliana and other species with similar patterns of variation. The rate of false positives differed strongly between traits, with more clinal traits showing the highest rate. However, the false positive rates were always substantial regardless of the trait, highlighting the necessity of an appropriate genomic control in association studies.

Arabidopsis↗

Light causes phosphorylation of nonactivated visual pigments in intact mouse rod photoreceptor cells.

Phosphorylation of G-protein-coupled receptors (GPCRs) is a required step in signal deactivation. Rhodopsin, a prototypical GPCR, exhibits high gain phosphorylation in vitro whereby a hundred-fold molar excess of phosphates are incorporated into the rhodopsin pool per molecule of activated rhodopsin. The extent by which high gain phosphorylation occurs in the intact mammalian photoreceptor cell, and the molecular mechanism underlying this reaction in vivo, is not known. Trans-phosphorylation is a mechanism proposed for high gain phosphorylation, whereby rhodopsin kinase, upon phosphorylating the activated receptor, continues to phosphorylate nearby nonactivated rhodopsin. We used two different transgenic mouse models to test whether trans-phosphorylation occurs in the intact photoreceptor cell. The first transgenic model expressed a murine cone pigment, S-opsin, together with the endogenous rhodopsin in the rod cell. We showed that selective stimulation of rhodopsin also led to phosphorylation of S-opsin. The second mouse model expressed the constitutively active human opsin mutant K296E. K296E, in the arrestin-/- background, also led to phosphorylation of endogenous mouse rhodopsin in the dark-adapted retina. Both mouse models provide strong support of trans-phosphorylation as an underlying mechanism of high gain phosphorylation, and provide evidence that a substantial fraction of nonactivated visual pigments becomes phosphorylated through this mechanism. Because activated, phosphorylated receptors exhibit decreased catalytic activity, our results suggest that dephosphorylation would be an important step in the full recovery of visual sensitivity during dark adaptation. These results may also have implications for other GPCR signaling pathways.

Animals↗

Statistical tests of the coalescent model based on the haplotype frequency distribution and the number of segregating sites.

Several tests of neutral evolution employ the observed number of segregating sites and properties of the haplotype frequency distribution as summary statistics and use simulations to obtain rejection probabilities. Here we develop a "haplotype configuration test" of neutrality (HCT) based on the full haplotype frequency distribution. To enable exact computation of rejection probabilities for small samples, we derive a recursion under the standard coalescent model for the joint distribution of the haplotype frequencies and the number of segregating sites. For larger samples, we consider simulation-based approaches. The utility of the HCT is demonstrated in simulations of alternative models and in application to data from Drosophila melanogaster.

Chromosome Segregation↗

The molecular signature of normal squamous esophageal epithelium identifies the presence of a field effect and can discriminate between patients with Barrett's esophagus and patients with Barrett's-associated adenocarcinoma.

BACKGROUND AND AIM: Genetic alterations in the normal tissues surrounding various cancers have been described, but a comprehensive analysis of this carcinogenic field effect in Barrett's-associated adenocarcinoma of the esophagus disease has not been reported. The aim of this study was to analyze the gene expression profile of a panel of highly selected genes in the normal squamous esophagus epihelium of patients with Barrett's esophagus, patients with Barrett's-associated adenocarcinoma, and a healthy control group to define the existence of a carcinogenic field effect, and to investigate the clinical importance of such a field effect in the management of Barrett's disease. METHODS: Forty-nine histologic normal squamous esophageal epithelia collected from 19 patients with Barrett's esophagus, 20 patients with Barrett's-associated esophageal adenocarcinoma, and a healthy control group of 10 patients were studied. A quantitative real-time reverse transcription-PCR method (TaqMan) was used to measure the expression of a panel of genes with known associations with gastrointestinal carcinogenesis. RESULTS: A widespread carcinogenic field effect was detected for more than 50% of the genes analyzed including Bax, BFT, CDX2, COX2, DAPK, DNMT1, GSTP1, RARalpha, RARgamma, RXRalpha, RXRbeta, SPARC, TSPAN, and VEGF. Based on the expression signature of the normal appearing squamous esophagus, a linear discriminant analysis was able to distinguish between the three groups of patients with an error rate of 0%. CONCLUSION: This study provides the first comprehensive investigation of a carcinogenic field effect in Barrett's esophagus disease. Based on the gene expression signature of the normal esophagus, patients could be correctly characterized according to their pathologic classification by applying a linear discriminant analysis. Our results provide evidence that a molecular classification might have clinical importance for the diagnosis and treatment of patients with Barrett's esophagus disease.

Adenocarcinoma↗

Estimating the rate of gene conversion on human chromosome 21.

There is a growing recognition that gene conversion can be an important factor in shaping fine-scale patterns of linkage disequilibrium in the human genome. We devised simple multilocus summary statistics for estimating gene-conversion rates from genomewide polymorphism data sets. In addition to being computationally feasible for very large data sets, these summaries were designed to yield robust estimates of gene-conversion rates in the presence of variation in crossing-over rates. Using our summaries, we analyzed 21,840 biallelic single-nucleotide polymorphisms (SNPs) on human chromosome 21. Our results indicate that models including both crossing over and gene conversion fit the overall short-range data (0-5 kb) of chromosome 21 much better than do models including crossing over alone. The estimated ratio of gene-conversion rate to crossing-over rate has a range of 1.6-9.4, depending on the assumed conversion tract length (in the range of 500-50 bp). Removal of the 5,696 SNPs that occur in known mutational hotspots (CpG sites) did not significantly change our conclusions, suggesting that recurrent mutations alone cannot explain our data.

Chromosome Aberrations↗

A multigene expression panel for the molecular diagnosis of Barrett's esophagus and Barrett's adenocarcinoma of the esophagus.

In order to identify genes or combination of genes that have the power to discriminate between premalignant Barrett's esophagus and Barrett's associated adenocarcinoma, we analysed a panel of 23 genes using quantitative real-time RT-PCR (qRT-PCR, Taqman and bioinformatic tools. The genes chosen were either known to be associated with Barrett's carcinogenesis or were filtered from a previous cDNA microarray study on Barrett's adenocarcinoma. A total of 98 tissues, obtained from 19 patients with Barrett's esophagus (BE group) and 20 patients with Barrett's associated esophageal adenocarcinoma (EA group), were studied. Triplicate analysis for the full 23 gene of interest panel, and analysis of an internal control gene, was performed for all samples, for a total of more than 9016 single PCR reactions. We found distinct classes of gene expression patterns in the different types of tissues. The most informative genes clustered in six different classes and had significantly different expression levels in Barrett's esophagus tissues compared to adenocarcinoma tissues. Linear discriminant analysis (LDA) distinguished four genetically different groups. The normal squamous esophagus tissues from patients with BE or EA were not distinguishable from one another, but Barrett's esophagus tissues could be distinguished from adenocarcinoma tissues. Using the most informative genes, obtained from a logistic regression analysis, we were able to completely distinguish between benign Barrett's and Barrett's adenocarcinomas. This study provides the first non-array parallel mRNA quantitation analysis of a panel of genes in the Barrett's esophagus model of multistage carcinogenesis. Our results suggest that mRNA expression quantitation of a panel of genes can discriminate between premalignant and malignant Barrett's disease. Logistic regression and LDAs can be used to further identify, from the complete panel, gene subsets with the power to make these diagnostic distinctions. Expression analysis of a limited number of highly selected genes may have clinical usefulness for the treatment of patients with this disease.

Adenocarcinoma↗

Nicotine-responsive genes in cultured embryonic mouse lung buds: interaction of nicotine and superoxide dismutase.

Nicotine exposure during prenatal development may be a cause of the abnormal lung function seen in infants born to smoking women. Previously we used an organ culture system to demonstrate that nicotine directly affects branching morphogenesis and gene expression in embryonic mouse lung buds. Here we attempt to identify genes potentially involved in the nicotine response and explore the relationship between gene expression changes and stimulation of branching. DNA microarray technology, analyzed by DChip software, and semi-quantitative RT-PCR were applied to RNA samples from embryonic lung buds grown in presence or absence of nicotine. Four genes, BAX, calcyclin, osteopontin and Cu-Zn superoxide dismutase (SOD1), identified by the microarray as showing changes in mRNA level with nicotine treatment were investigated in detail. RT-PCR showed that nicotine exposure resulted in significant decreases in mRNA levels for BAX, calcyclin and osteopontin, but nicotine did not affect the mRNA level of SOD1. Nicotine-induced changes in BAX, calcyclin and osteopontin mRNAs showed a general correlation with stimulation of branching, implying a common mechanism for effects of nicotine on branching and on gene expression. BAX, calcyclin and osteopontin mRNA levels were found to be developmentally regulated, but only the effect of nicotine on BAX mRNA was parallel to the developmental change in vivo, suggesting that nicotine action cannot be explained simply as a stimulation of the embryonic lung's developmental program. Addition of exogenous SOD to the culture medium resulted in increased branching similar to that caused by nicotine, but, unexpectedly, branching was not increased relative to control when nicotine and SOD were co-administered, suggesting interfering mechanisms of action of the two agents. Exogenous SOD was found to alter mRNA levels of BAX, calcyclin and osteopontin in a pattern that differed from that seen in response to nicotine. Gene expression changes seen with co-administration of nicotine and SOD yielded further evidence of interaction between these agents. In conclusion, three putative nicotine-responsive genes were identified whose expression was also influenced by developmental stage and by exogenously added SOD. A common mechanism likely underlies nicotine's effects on both branching and gene expression in this system. Our evidence also suggests that nicotine and SOD stimulate branching by distinct but interacting mechanisms.

Animals↗

A survey of current Bayesian gene mapping methods.

Recently, there has been much interest in the use of Bayesian statistical methods for performing genetic analyses. Many of the computational difficulties previously associated with Bayesian analysis, such as multidimensional integration, can now be easily overcome using modern high-speed computers and Markov chain Monte Carlo (MCMC) methods. Much of this new technology has been used to perform gene mapping, especially through the use of multi-locus linkage disequilibrium techniques. This review attempts to summarise some of the currently available methods and the software available to implement these methods.

Bayes Theorem↗

Haplotype structure and phenotypic associations in the chromosomal regions surrounding two Arabidopsis thaliana flowering time loci.

The feasibility of using linkage disequilbrium (LD) to fine-map loci underlying natural variation in Arabidopsis thaliana was investigated by looking for associations between flowering time and marker polymorphism in the genomic regions containing two candidate genes, FRI and FLC, both of which are known to contribute to natural variation in flowering. A sample of 196 accessions was used, and polymorphism was assessed by sequencing a total of 17 roughly 500-bp fragments. Using a novel Bayesian algorithm based on haplotype similarity, we demonstrate that LD could have been used to fine-map the FRI gene to a roughly 30-kb region and to identify two common loss-of-function alleles. Interestingly, because of genetic heterogeneity, simple single-marker associations would not have been able to map FRI with nearly the same precision. No clear evidence for previously unknown alleles at either locus was found, but the effect of population structure in causing false positives was evident.

Arabidopsis↗

Age-related changes of cardiac gene expression following myocardial ischemia/reperfusion.

Young and old (4 and 25 months of age, respectively) Fisher 344/Brown Norway hybrid female rats were subjected to four 3 min episodes of ischemia separated by 5 min of reperfusion. Corresponding open-chest sham-operated groups received 32 min of no intervention. All rats were allowed to recover, and 24h later hearts were removed and frozen in liquid nitrogen. Global gene profiling in the ischemic and the non-ischemic areas and in the sham-operated hearts as well was carried out by using Affymetrix Gene Chips. Young ischemic hearts demonstrated down-regulation of gene expression associated with early-remodeling including down-regulation of tissue inhibitor of metalloproteinase 1, decorin, collagen, tropoelastin, and fibulin, as well as decreases in hypertrophy-related transcripts. In contrast, old hearts showed a unique injury-related response, which included up-regulation of mRNAs for proteins associated with hypertrophy or apoptosis (including H36-alpha7 integrin, alpha-actin, tubulin, filamin, connective tissue growth factor, calcineurin, serine protease, and apoptosis inducing factor). These injury-related changes in gene expression could in part explain increased gravity of outcomes of ischemia and myocardial infarction in elderly hearts.

Age Factors↗