Search PubMed⌕ Search

Biomedical subjects

Duncan C Thomas

Publications and source records attributed to Duncan C Thomas.

28 records · Page 2Linked to original sources

Segregation and linkage analysis for longitudinal measurements of a quantitative trait.

We present a method for using slopes and intercepts from a linear regression of a quantitative trait as outcomes in segregation and linkage analyses. We apply the method to the analysis of longitudinal systolic blood pressure (SBP) data from the Framingham Heart Study. A first-stage linear model was fit to each subject's SBP measurements to estimate both their slope over time and an intercept, the latter scaled to represent the mean SBP at the average observed age (53.7 years). The subject-specific intercepts and slopes were then analyzed using segregation and linkage analysis. We describe a method for using the standard errors of the first-stage intercepts and slopes as weights in the genetic analyses. For the intercepts, we found significant evidence of a Mendelian gene in segregation analysis and suggestive linkage results (with LOD scores >or= 1.5) for specific markers on chromosomes 1, 3, 5, 9, 10, and 17. For the slopes, however, the data did not support a Mendelian model, and thus no formal linkage analyses were conducted.

Adult Children↗

Genetic Analysis Workshop 13: simulated longitudinal data on families for a system of oligogenic traits.

The Genetic Analysis Workshop 13 simulated data aimed to mimic the major features of the real Framingham Heart Study data that formed Problem 1, but under a known inheritance model and with 100 replicates, so as to allow evaluation of the statistical properties of various methods. The pedigrees used were the 330 real pedigree structures (comprising 4692 individuals) with some minor changes to protect confidentiality. Fifty trait genes and 399 microsatellite markers were simulated by gene dropping on 22 autosomal chromosomes. Assuming random ascertainment of families, a system of eight longitudinal quantitative traits (designed to be similar to those in the real data) was generated with a wide range of heritabilities, including some pleiotropic and interactive effects. Genes could affect either the baseline level or the rate of change of the phenotype. Hypertension diagnosis and treatment were simulated with treatment availability, compliance, and efficacy depending on calendar year. Nongenetic traits of smoking and alcohol were generated as covariates for other traits. Death was simulated as a hazard rate depending upon age, sex, smoking, cholesterol, and systolic blood pressure. After the complete data were simulated, missing data indicators were generated based on logistic models fitted to the real data, involving the subject's history of previous missing values, together with that of their spouses, parents, siblings, and offspring, as well as marital status, only-child indicators, current value at certain simulated traits, and the data collection pattern on the cohort into which each subject was ascertained.

Adult↗

Summary report: Missing data and pedigree and genotyping errors.

Genetic epidemiology is faced with mapping complex traits to genes with relatively small effects whose phenotypes may be modulated by temporal factors. To do this, detailed and accurate data must be available on families, perhaps collected over time. The Framingham Heart Study data supplied to Genetic Analysis Workshop 13 (GAW13), along with its simulated counterpart, contain longitudinal measurements and genomic scan data on 2,885 individuals in 330 families, and offer an opportunity to examine data quality and completeness issues as they affect analytical conclusions. Six GAW13 contributions applied methods to deal with missing data, both phenotypic and genotypic, at a single time point and longitudinally, and with possible errors in pedigree structure and genotypes. The methods included missing phenotypic data imputation by Markov chain Monte Carlo sampling, propensity scoring, regression, and adjusted mean values, as well as the assessment of transmission-disequilibrium tests when missing marker data may be allele-specific. Pedigree structural errors were found by genome-wide allele-sharing probabilities, while Mendelian consistent genotype errors were evaluated through likelihoods of double-recombination events. Each of the methods reviewed here offered insights into how to better take advantage of large, time-dependent, familial data sets. However, no one of them dealt with the longitudinal and familial aspects simultaneously. Overall, more consideration needs to be given to the effects that missing data and data errors have on our ability to map complex traits efficiently and accurately.

Cardiovascular Diseases↗

Modeling and E-M estimation of haplotype-specific relative risks from genotype data for a case-control study of unrelated individuals.

The US National Cancer Institute has recently sponsored the formation of a Cohort Consortium (http://2002.cancer.gov/scpgenes.htm) to facilitate the pooling of data on very large numbers of people, concerning the effects of genes and environment on cancer incidence. One likely goal of these efforts will be generate a large population-based case-control series for which a number of candidate genes will be investigated using SNP haplotype as well as genotype analysis. The goal of this paper is to outline the issues involved in choosing a method of estimating haplotype-specific risk estimates for such data that is technically appropriate and yet attractive to epidemiologists who are already comfortable with odds ratios and logistic regression. Our interest is to develop and evaluate extensions of methods, based on haplotype imputation, that have been recently described (Schaid et al., Am J Hum Genet, 2002, and Zaykin et al., Hum Hered, 2002) as providing score tests of the null hypothesis of no effect of SNP haplotypes upon risk, which may be used for more complex tasks, such as providing confidence intervals, and tests of equivalence of haplotype-specific risks in two or more separate populations. In order to do so we (1) develop a cohort approach towards odds ratio analysis by expanding the E-M algorithm to provide maximum likelihood estimates of haplotype-specific odds ratios as well as genotype frequencies; (2) show how to correct the cohort approach, to give essentially unbiased estimates for population-based or nested case-control studies by incorporating the probability of selection as a case or control into the likelihood, based on a simplified model of case and control selection, and (3) finally, in an example data set (CYP17 and breast cancer, from the Multiethnic Cohort Study) we compare likelihood-based confidence interval estimates from the two methods with each other, and with the use of the single-imputation approach of Zaykin et al. applied under both null and alternative hypotheses. We conclude that so long as haplotypes are well predicted by SNP genotypes (we use the Rh2 criteria of Stram et al. [1]) the differences between the three methods are very small and in particular that the single imputation method may be expected to work extremely well.

Algorithms↗

Bayesian spatial modeling of haplotype associations.

We review methods for relating the risk of disease to a collection of single nucleotide polymorphisms (SNPs) within a small region. Association studies using case-control designs with unrelated individuals could be used either to test for a direct effect of a candidate gene and characterize the responsible variant(s), or to fine map an unknown gene by exploiting the pattern of linkage disequilibrium (LD). We consider a flexible class of logistic penetrance models based on haplotypes and compare them with an alternative formulation based on unphased multilocus genotypes. The likelihood for haplotype-based models requires summation over all possible haplotype assignments consistent with the observed genotype data, and can be fitted using either Expectation-Maximization (E-M) or Markov chain Monte Carlo (MCMC) methods. Subtleties involving ascertainment correction for case-control studies are discussed. There has been great interest in methods for LD mapping based on the coalescent or ancestral recombination graphs as well as methods based on haplotype sharing, both of which we review briefly. Because of their computational complexity, we propose some alternative empirical modeling approaches using techniques borrowed from the Bayesian spatial statistics literature. Here, space is interpreted in terms of a distance metric describing the similarity of any pair of haplotypes to each other, and hence their presumed common ancestry. Specifically, we discuss the conditional autoregressive model and two spatial clustering models: Potts and Voronoi. We conclude with a discussion of the implications of these methods for modeling cryptic relatedness, haplotype blocks, and haplotype tagging SNPs, and suggest a Bayesian framework for the HapMap project.

Algorithms↗

Bayesian modeling of complex metabolic pathways.

Many chronic diseases are the result of a complex sequence of biochemical reactions involving exposures to various environmental agents, metabolized by a number of different genes. Routine epidemiologic analyses of such associations have tended to rely on standard contingency table or logistic regression methods, typically focusing on one variable at a time or pairwise combinations. We consider two statistical alternatives to this approach, one based on Bayesian model averaging, one based on pharmacokinetic modeling of the biochemical pathways. These approaches are illustrated using data from a case-control study of colorectal polyps in relation to tobacco smoking and consumption of well done red meat, both viewed as sources of heterocyclic amines and polycyclic aromatic hydrocarbons. The new analyses are structured in a manner that attempts to take advantage of prior knowledge of the metabolism of these classes of compounds and the various genes that regulate these pathways.

Bayes Theorem↗

Traffic density and the risk of childhood leukemia in a Los Angeles case-control study.

PURPOSE: To investigate the relationship between traffic density and the risk of childhood leukemia. METHODS: The study group consisted of 212 cases and 202 controls from the London et al. (1991) study of childhood leukemia conducted in the Los Angeles area during 1978 to 1984. Using GIS methods, traffic counts on all streets within 1500 feet of each subject's residence of longest duration were determined. From these counts, an integrated distance-weighted traffic density measure was calculated for each subject for use as the analytic variable. Additional information, including magnetic fields and wire-code, was obtained from the original case-control study. Association between traffic density and leukemia, and confounding and effect modification by other variables, were assessed using standard matched case-control analyses. RESULTS: Although the unadjusted traffic density-childhood leukemia rate ratios were slightly elevated, this weak association was explained by confounding by wire code. Wire code remained associated with leukemia after controlling for traffic density. There was little evidence of effect modification between traffic density and magnetic fields, wire code or other variables. CONCLUSIONS: There is no evidence of an association of traffic density with childhood leukemia in the Los Angeles case-control study.

Case-Control Studies↗

A two-stage model for multiple time series data of counts.

We propose a two-stage model for time series data of counts from multiple locations. This method fits first-stage model(s) using the technique of iteratively weighted filtered least squares (IWFLS) to obtain location-specific intercepts and slopes, with possible lagged effects via polynomial distributed lag modeling. These slopes and/or intercepts are then taken to a second-stage mixed-effects meta-regression model in order to stabilize results from various locations. The representation of the models from the stages into a combined mixed-effects model, issues of inference and choices of the parameters in modeling the lag structure are discussed. We illustrate this proposed model via detailed analysis on the effect of air pollution on school absenteeism based on data from the Southern California Children's Health Study.

Journal Article↗