Search PubMed⌕ Search

Biomedical subjects

K Roeder

Publications and source records attributed to K Roeder.

At least 19 recordsLinked to original sources

Cladistic analysis of human apolipoprotein a4 polymorphisms in relation to quantitative plasma lipid risk factors of coronary heart disease.

Genetic variation in several genes involved in lipid metabolism is known to affect population variation in quantitative lipid risk factor profiles for coronary heart disease (CHD). The apolipoprotein A-IV gene (APOA4) is one such candidate gene. We genotyped five polymorphisms in the APOA4 gene (codon 127, codon 130, codon347, codon 360 and 3' VNTR) and investigated their impact on plasma lipid trait levels in three populations comprising 604 U.S. non-Hispanic Whites (NHWs), 408 U.S. Hispanics and 708 Nigerian Blacks. Cladistic analysis was carried out to identify 5-site haplotypes that were associated with significant phenotypic differences in each population. The distribution of APOA4 genotypes was significantly different between ethnic groups. The Africans were monomorphic for two of the five sites (codons 130 and 360), but possess a unique 12 bp insertion that was not observed in NHWs and Hispanics. Due to linkage disequilibrium between the sites, only 6 haplotypes were observed in NHWs and Hispanics, and 4 in Africans. Several gender-and ethnic-specific associations between genotypes and plasma lipid traits were observed when single sites were used. Several haplotypes were identified by cladistic analysis that may carry functional mutations that affect plasma lipid trait levels.

Adult↗

Genome-wide multipoint linkage analyses of multiplex schizophrenia pedigrees from the oceanic nation of Palau.

The oceanic nation of Palau has been geographically and culturally isolated over most of its 2000 year history. As part of a study of the genetic basis of schizophrenia in Palau, we genotyped five large, multigenerational schizophrenia pedigrees using markers every 10 cM (CHLC/Weber screening set 6). The number of affected/unaffected individuals genotyped per family ranged from 11/21 to 5/5. Thus the pedigrees varied in their information for linkage, but each was capable of producing a substantial LOD score. We fitted a simple dominant and recessive model to these data using multipoint linkage analysis implemented by Simwalk2. Predictably, the most informative pedigrees produced the best linkage results. After genotyping additional markers in the region, one pedigree produced a LOD = 3.4 (5q distal) under the dominant model. Seven of nine schizophrenics in the pedigree, mostly 3rd-4th degree relatives, share a 15-cM, 7-marker haplotype. For a different pedigree, another promising signal occurred on distal 3q, LOD = 2.6, for the recessive model. For two other pedigrees, the best LODs were modest, slightly better than 2.0 on 5q and 9p, while the fifth pedigree produced no noteworthy linkage signal. Similar to the results for other populations, our results suggest there are multiple genes conferring liability to schizophrenia even in the small population of Palau (roughly 21,000 individuals) in remote Oceania.

Genome, Human↗

Transmission/disequilibrium test meets measured haplotype analysis: family-based association analysis guided by evolution of haplotypes.

Family data teamed with the transmission/disequilibrium test (TDT), which simultaneously evaluates linkage and association, is a powerful means of detecting disease-liability alleles. To increase the information provided by the test, various researchers have proposed TDT-based methods for haplotype transmission. Haplotypes indeed produce more-definitive transmissions than do the alleles comprising them, and this tends to increase power. However, the larger number of haplotypes, relative to alleles at individual loci, tends to decrease power, because of the additional degrees of freedom required for the test. An optimal strategy would focus the test on particular haplotypes or groups of haplotypes. In this report we develop such an approach by combining the theory of TDT with that of measured haplotype analysis (MHA). MHA uses the evolutionary relationships among haplotypes to produce a limited set of hypothesis tests and to increase the interpretability of these tests. The theory of our approach, called the "evolutionary tree" (ET)-TDT, is developed for two cases: when haplotype transmission is certain and when it is not. Simulations show the ET-TDT can be more powerful than other proposed methods under reasonable conditions. More importantly, our results show that, when multiple polymorphisms are found within the gene, the ET-TDT can be useful for determining which polymorphisms affect liability.

Alleles↗

A Bayesian hierarchical model for allele frequencies.

Genetic epidemiological methodologies, such as linkage analysis, often require accurate estimates of allele frequencies. When studies involve multiple sub-populations with different evolutionary histories, accurate estimates can be difficult to obtain because the number of subjects per sub-population tends to be limited. Given allele counts for a collection of loci and sub-populations, we propose a Bayesian hierarchical model that extends existing empirical Bayesian approaches by allowing for explicit inclusion of prior information about both allele frequencies and inter-population divergence. We describe how such information can be derived from published data and then incorporated into the model via prior distributions for model parameters. By analysis of simulated data, we highlight how the hierarchical model, as implemented in the publicly available program AllDist, combines prior information with the observed data to refine allele frequency estimates.

Alleles↗

Unbiased methods for population-based association studies.

Large, population-based samples and large-scale genotyping are being used to evaluate disease/gene associations. A substantial drawback to such samples is the fact that population substructure can induce spurious associations between genes and disease. We review two methods, called genomic control (GC) and structured association (SA), that obviate many of the concerns about population substructure by using the features of the genomes present in the sample to correct for stratification. The GC approach exploits the fact that population substructure generates "over dispersion" of statistics used to assess association. By testing multiple polymorphisms throughout the genome, only some of which are pertinent to the disease of interest, the degree of overdispersion generated by population substructure can be estimated and taken into account. The SA approach assumes that the sampled population, although heterogeneous, is composed of subpopulations that are themselves homogeneous. By using multiple polymorphisms throughout the genome, this "latent class method" estimates the probability sampled individuals derive from each of these latent subpopulations. GC has the advantage of robustness, simplicity, and wide applicability, even to experimental designs such as DNA pooling. SA is a bit more complicated but has the advantage of greater power in some realistic settings, such as admixed populations or when association varies widely across subpopulations. It, too, is widely applicable. Both also have weaknesses, as elaborated in our review.

Analysis of Variance↗

Genomic control, a new approach to genetic-based association studies.

During the past decade, mutations affecting liability to human disease have been discovered at a phenomenal rate, and that rate is increasing. For the most part, however, those diseases have a relatively simple genetic basis. For diseases with a complex genetic and environmental basis, new approaches are needed to pave the way for more rapid discovery of genes affecting liability. One such approach exploits large, population-based samples and large-scale genotyping to evaluate disease/gene associations. A substantial drawback to such samples is the fact that population heterogeneity can induce spurious associations between genes and disease. We describe a method called genomic control (GC), which obviates many of the concerns about population substructure by using the features of the genomes present in the sample to correct for stratification. Two such approaches are now available. The GC approach exploits the fact that population substructure generate "overdispersion" of statistics used to assess association. By testing multiple polymorphisms throughout the genome, only some of which are pertinent to the disease of interest, the degree of overdispersion generated by population substructure can be estimated and taken into account. The other approach, called Structured Association (SA), assumes that the sampled population, while heterogeneous, is composed of subpopulations that are themselves homogeneous. By using multiple polymorphisms throughout the genome, SA probabilistically assigns sampled individuals to these latent subpopulations. We review in detail the overdispersion GC. In addition to outlining the published ideas on this method, we describe several extensions: quantitative trait studies and case-control studies with haplotypes and multiallelic markers. For each study design our goal is to achieve control similar to that obtained for a family-based study, but with the convenience found in a population-based design.

Alleles↗

Genome-wide distribution of linkage disequilibrium in the population of Palau and its implications for gene flow in Remote Oceania.

Linkage disequilibrium (LD) between alleles on the same human chromosome results from various evolutionary processes and is thus telling about the history of populations. Recently, LD has garnered substantial interest for its value to map and fine-map disease genes. We examine the distribution of LD between short tandem repeat alleles on autosomes and sex chromosomes in the Remote Oceanic population of Palau to evaluate whether the data are consistent with a recent hypothesis about the origins of genetic variation in Palau, specifically that the population experienced extensive male-biased gene flow following initial settlement. Consistent with evolutionary theory based on effective population size, LD between X-linked alleles is stochastically greater than LD between autosomal alleles, however, small but detectable LD occurs for autosomal markers separated by substantial distances. By contrast, while Y-linked alleles experience only one-third the effective population size of X-linked alleles, their mean value for pairwise LD is only slightly larger than X-linked alleles. For a small population known to experience at least two extreme bottlenecks, 56 six-locus Y haplotypes exhibit remarkable diversity (0.96), comparable to Y diversity of Europeans, however, autosomal and X-linked markers display significantly less diversity, as measured by heterozygosity (4.1% less). Palauan Y haplotypes also fall into distinct clusters, again unlike that of Europe. We argue these data are consistent with waves of male-biased gene flow.

Alleles↗

The power of genomic control.

Although association analysis is a useful tool for uncovering the genetic underpinnings of complex traits, its utility is diminished by population substructure, which can produce spurious association between phenotype and genotype within population-based samples. Because family-based designs are robust against substructure, they have risen to the fore of association analysis. Yet, if population substructure could be ignored, this robustness can come at the price of power. Unfortunately it is rarely evident when population substructure can be ignored. Devlin and Roeder recently have proposed a method, termed "genomic control" (GC), which has the robustness of family-based designs even though it uses population-based data. GC uses the genome itself to determine appropriate corrections for population-based association tests. Using the GC method, we contrast the power of two study designs, family trios (i.e., father, mother, and affected progeny) versus case-control. For analysis of trios, we use the TDT test. When population substructure is absent, we find GC is always more powerful than TDT; furthermore, contrary to previous results, we show that as a disease becomes more prevalent the discrepancy in power becomes more extreme. When population substructure is present, however, the results are more complex: TDT is more powerful when population substructure is substantial, and GC is more powerful otherwise. We also explore general issues of power and implementation of GC within the case-control setting and find that, economically, GC is at least comparable to and often less expensive than family-based methods. Therefore, GC methods should prove a useful complement to family-based methods for the genetic analysis of complex traits.

Alleles↗

Haplotype fine mapping by evolutionary trees.

To refine the location of a disease gene within the bounds provided by linkage analysis, many scientists use the pattern of linkage disequilibrium between the disease allele and alleles at nearby markers. We describe a method that seeks to refine location by analysis of "disease" and "normal" haplotypes, thereby using multivariate information about linkage disequilibrium. Under the assumption that the disease mutation occurs in a specific gap between adjacent markers, the method first combines parsimony and likelihood to build an evolutionary tree of disease haplotypes, with each node (haplotype) separated, by a single mutational or recombinational step, from its parent. If required, latent nodes (unobserved haplotypes) are incorporated to complete the tree. Once the tree is built, its likelihood is computed from probabilities of mutation and recombination. When each gap between adjacent markers is evaluated in this fashion and these results are combined with prior information, they yield a posterior probability distribution to guide the search for the disease mutation. We show, by evolutionary simulations, that an implementation of these methods, called "FineMap," yields substantial refinement and excellent coverage for the true location of the disease mutation. Moreover, by analysis of hereditary hemochromatosis haplotypes, we show that FineMap can be robust to genetic heterogeneity.

Algorithms↗

Genomic control for association studies: a semiparametric test to detect excess-haplotype sharing.

Individuals who share a disease mutation from a common ancestor often share alleles at genetic markers adjacent to the mutation, even if the common ancestor is remote. The alleles at these adjacent markers, called the haplotype, can be visualized as a string of realizations of random variables, which may be dependent when individuals are related in some fashion. Ideally, for a sample of individuals all having the same (genetic) disease, this dependence-measured as haplotype-sharing-will be greater in the vicinity of disease genes than in other regions of the genome. In this paper we present a semiparametric test for haplotype-sharing. We begin by developing a model assuming that the ancestral haplotype is known and thus the extent of haplotype-sharing from a common ancestor can be determined unambiguously. The amount of overlap at markers far from the disease is treated as a random variable with an unknown distribution F, which we estimate non-parametrically. Overlap of markers surrounding disease genes are modeled as a mixture pF(x - theta) + (1 - p)F(x), in which p is the fraction of subjects with the disease mutation. Testing for a disease gene then amounts to testing whether p = 0. Next we drop the assumption that the ancestral haplotype is known. To detect excess clustering of haplotypes, we measure the pairwise overlap of a set of haplotypes. As in the simpler scenario, this distribution is modeled as a location-shift mixture. To test the hypothesis we construct a score test with a simple limiting distribution.

Journal Article↗

Flexible parametric measurement error models.

Inferences in measurement error models can be sensitive to modeling assumptions. Specifically, if the model is incorrect, the estimates can be inconsistent. To reduce sensitivity to modeling assumptions and yet still retain the efficiency of parametric inference, we propose using flexible parametric models that can accommodate departures from standard parametric models. We use mixtures of normals for this purpose. We study two cases in detail: a linear errors-in-variables model and a change-point Berkson model.

Biometry↗

Genomic control for association studies.

A dense set of single nucleotide polymorphisms (SNP) covering the genome and an efficient method to assess SNP genotypes are expected to be available in the near future. An outstanding question is how to use these technologies efficiently to identify genes affecting liability to complex disorders. To achieve this goal, we propose a statistical method that has several optimal properties: It can be used with case control data and yet, like family-based designs, controls for population heterogeneity; it is insensitive to the usual violations of model assumptions, such as cases failing to be strictly independent; and, by using Bayesian outlier methods, it circumvents the need for Bonferroni correction for multiple tests, leading to better performance in many settings while still constraining risk for false positives. The performance of our genomic control method is quite good for plausible effects of liability genes, which bodes well for future genetic analyses of complex disorders.

Bayes Theorem↗

The heritability of IQ.

IQ heritability, the portion of a population's IQ variability attributable to the effects of genes, has been investigated for nearly a century, yet it remains controversial. Covariance between relatives may be due not only to genes, but also to shared environments, and most previous models have assumed different degrees of similarity induced by environments specific to twins, to non-twin siblings (henceforth siblings), and to parents and offspring. We now evaluate an alternative model that replaces these three environments by two maternal womb environments, one for twins and another for siblings, along with a common home environment. Meta-analysis of 212 previous studies shows that our 'maternal-effects' model fits the data better than the 'family-environments' model. Maternal effects, often assumed to be negligible, account for 20% of covariance between twins and 5% between siblings, and the effects of genes are correspondingly reduced, with two measures of heritability being less than 50%. The shared maternal environment may explain the striking correlation between the IQs of twins, especially those of adult twins that were reared apart. IQ heritability increases during early childhood, but whether it stabilizes thereafter remains unclear. A recent study of octogenarians, for instance, suggests that IQ heritability either remains constant through adolescence and adulthood, or continues to increase with age. Although the latter hypothesis has recently been endorsed, it gathers only modest statistical support in our analysis when compared to the maternal-effects hypothesis. Our analysis suggests that it will be important to understand the basis for these maternal effects if ways in which IQ might be increased are to be identified.

Adolescent↗

A statistical model for locating regulatory regions in genomic DNA.

In addition to genes, chromosomal DNA contains sequences that serve as signals for turning on and off gene expression. These signals are thought to be distributed as clusters in the regulatory regions of genes. We develop a Bayesian model that views locating regulatory regions in genomic DNA as a change-point problem, with the beginning of regulatory and non-regulatory regions corresponding to the change points. The model is based on a hidden Markov chain. The data consist of nucleotide positions of protein-binding elements in a genomic DNA sequence. These positions are identified using a reference catalogue containing elements that interact with transcription factors implicated in controlling the expression of protein-encoding genes. Among the protein-binding elements in a genomic DNA sequence, the statistical model automatically selects those that tend to predict regulatory regions. We test the model using viral sequences that include known regulatory regions and provide the results obtained for human genomic DNA corresponding to the beta globin locus on chromosome 11.

Adenoviridae↗