Genetics. Harvesting medical information from the human family tree.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to Andrew G Clark.
Explore the source record for details and available documents.
There is a critical need for data-mining methods that can identify SNPs that predict among individual variation in a phenotype of interest and reverse-engineer the biological network of relationships between SNPs, phenotypes, and other factors. This problem is both challenging and important in light of the large number of SNPs in many genes of interest and across the human genome. A potentially fruitful form of exploratory data analysis is the Bayesian or Belief network. A Bayesian or Belief network provides an analytic approach for identifying robust predictors of among-individual variation in a disease endpoints or risk factor levels. We have applied Belief networks to SNP variation in the human APOE gene and plasma apolipoprotein E levels from two samples: 702 African-Americans from Jackson, MS, and 854 non-Hispanic whites from Rochester, MN. Twenty variable sites in the APOE gene were genotyped in both samples. In Jackson, MS, SNPs 4036 and 4075 were identified to influence plasma apoE levels. In Rochester, MN, SNPs 3937 and 4075 were identified to influence plasma apoE levels. All three SNPs had been previously implicated in affecting measures of lipid and lipoprotein metabolism. Like all data-mining methods, Belief networks are meant to complement traditional hypothesis-driven methods of data analysis. These results document the utility of a Belief network approach for mining large scale genotype-phenotype association data.
We have sequenced the genome of a second Drosophila species, Drosophila pseudoobscura, and compared this to the genome sequence of Drosophila melanogaster, a primary model organism. Throughout evolution the vast majority of Drosophila genes have remained on the same chromosome arm, but within each arm gene order has been extensively reshuffled, leading to a minimum of 921 syntenic blocks shared between the species. A repetitive sequence is found in the D. pseudoobscura genome at many junctions between adjacent syntenic blocks. Analysis of this novel repetitive element family suggests that recombination between offset elements may have given rise to many paracentric inversions, thereby contributing to the shuffling of gene order in the D. pseudoobscura lineage. Based on sequence similarity and synteny, 10,516 putative orthologs have been identified as a core gene set conserved over 25-55 million years (Myr) since the pseudoobscura/melanogaster divergence. Genes expressed in the testes had higher amino acid sequence divergence than the genome-wide average, consistent with the rapid evolution of sex-specific proteins. Cis-regulatory sequences are more conserved than random and nearby sequences between the species--but the difference is slight, suggesting that the evolution of cis-regulatory elements is flexible. Overall, a pattern of repeat-mediated chromosomal rearrangement, and high coadaptation of both male genes and cis-regulatory sequences emerges as important themes of genome divergence between these species of Drosophila.
Large-scale SNP genotyping studies rely on an initial assessment of nucleotide variation to identify sites in the DNA sequence that harbor variation among individuals. This "SNP discovery" sample may be quite variable in size and composition, and it has been well established that properties of the SNPs that are found are influenced by the discovery sampling effort. The International HapMap project relied on nearly any piece of information available to identify SNPs-including BAC end sequences, shotgun reads, and differences between public and private sequences-and even made use of chimpanzee data to confirm human sequence differences. In addition, the ascertainment criteria shifted from using only SNPs that had been validated in population samples, to double-hit SNPs, to finally accepting SNPs that were singletons in small discovery samples. In contrast, Perlegen's primary discovery was a resequencing-by-hybridization effort using the 24 people of diverse origin in the Polymorphism Discovery Resource. Here we take these two data sets and contrast two basic summary statistics, heterozygosity and F(ST), as well as the site frequency spectra, for 500-kb windows spanning the genome. The magnitude of disparity between these samples in these measures of variability indicates that population genetic analysis on the raw genotype data is ill advised. Given the knowledge of the discovery samples, we perform an ascertainment correction and show how the post-correction data are more consistent across these studies. However, discrepancies persist, suggesting that the heterogeneity in the SNP discovery process of the HapMap project resulted in a data set resistant to complete ascertainment correction. Ascertainment bias will likely erode the power of tests of association between SNPs and complex disorders, but the effect will likely be small, and perhaps more importantly, it is unlikely that the bias will introduce false-positive inferences.
Explore the source record for details and available documents.
Detecting selective sweeps from genomic SNP data is complicated by the intricate ascertainment schemes used to discover SNPs, and by the confounding influence of the underlying complex demographics and varying mutation and recombination rates. Current methods for detecting selective sweeps have little or no robustness to the demographic assumptions and varying recombination rates, and provide no method for correcting for ascertainment biases. Here, we present several new tests aimed at detecting selective sweeps from genomic SNP data. Using extensive simulations, we show that a new parametric test, based on composite likelihood, has a high power to detect selective sweeps and is surprisingly robust to assumptions regarding recombination rates and demography (i.e., has low Type I error). Our new test also provides estimates of the location of the selective sweep(s) and the magnitude of the selection coefficient. To illustrate the method, we apply our approach to data from the Seattle SNP project and to Chromosome 2 data from the HapMap project. In Chromosome 2, the most extreme signal is found in the lactase gene, which previously has been shown to be undergoing positive selection. Evidence for selective sweeps is also found in many other regions, including genes known to be associated with disease risk such as DPP10 and COL4A3.
Genetic variation in the apolipoprotein A-V gene (APOA5) has been associated with variation in plasma triglyceride (TG) levels in African American and white females and males older than 40 years and/or at increased risk of coronary artery disease. We have examined whether plasma TG levels are associated with 16 APOA5 polymorphisms in young (18-30 years) African American (1,075 females and 783 males) and white (1,041 females and 932 males) individuals of the Coronary Artery Risk Development in Young Adults (CARDIA) Study selected without regard to health. Plasma TG was significantly (P < 0.01) associated with markers 27376 and 28837 (-3A/G) in both white females and males, with 27709 (-1131T/C) and 29085 in white males, with 29009 (S19W) in African American females and white males, and with 30966 in African American females. No statistically significant associations were observed in African American males. These six single-nucleotide polymorphisms individually accounted for 0-0.78% of lnTG variation among white females, 0-2.46% among white males, and 0-0.69% among African American females. The results of our study suggest a small but replicable context-dependent influence of the APOA5 gene region on plasma TG levels in young, healthy individuals.
Nearly one in eight US women will develop breast cancer in their lifetime. Most breast cancer is not associated with a hereditary syndrome, occurs in postmenopausal women, and is estrogen and progesterone receptor-positive. Estrogen exposure is an epidemiologic risk factor for breast cancer and estrogen is a potent mammary mitogen. We studied single nucleotide polymorphisms (SNPs) in estrogen receptors in 615 healthy subjects and 1011 individuals with histologically confirmed breast cancer, all from New York City. We analyzed 13 SNPs in the progesterone receptor gene (PGR), 17 SNPs in estrogen receptor 1 gene (ESR1), and 8 SNPs in the estrogen receptor 2 gene (ESR2). We observed three common haplotypes in ESR1 that were associated with a decreased risk for breast cancer [odds ratio (OR), approximately O.4; 95% confidence interval (CI), 0.2-0.8; P < 0.01]. Another haplotype was associated with an increased risk of breast cancer (OR, 2.1; 95% CI, 1.2-3.8; P < 0.05). A unique risk haplotype was present in approximately 7% of older Ashkenazi Jewish study subjects (OR, 1.7; 95% CI, 1.2-2.4; P < 0.003). We narrowed the ESR1 risk haplotypes to the promoter region and first exon. We define several other haplotypes in Ashkenazi Jews in both ESR1 and ESR2 that may elevate susceptibility to breast cancer. In contrast, we found no association between any PGR variant or haplotype and breast cancer. Genetic epidemiology study replication and functional assays of the haplotypes should permit a better understanding of the role of steroid receptor genetic variants and breast cancer risk.
A binding site for the repressor protein BP1, which contains a tandem (AT)x(T)y repeat, is located approximately 530 bp 5' to the human beta-globin gene (HBB). There is accumulating evidence that BP1 binds to the (AT)9(T)5 allele more strongly than to other alleles, thereby reducing the expression of HBB. In this study, we investigated polymorphisms in the (AT)x(T)y repeat in 57 individuals living in Thailand, including three homozygotes for the hemoglobin E variant (HbE; beta26Glu-->Lys), 22 heterozygotes, and 32 normal homozygotes. We found that (AT)9(T)5 and (AT)7(T)7 alleles were predominant in the studied population and that the HbE variant is in strong linkage disequilibrium with the (AT)9(T)5 allele, which can explain why the betaE chain is inefficiently synthesized compared to the normal betaA chain. Moreover, the mildness of the HbE disease compared to other hemoglobinopathies in Thai may be due, in part, to the presence of the (AT)9(T)5 repeat on the HbE chromosome. In addition, a novel (AC)n polymorphism adjacent to the (AT)x(T)y repeat (i.e., (AC)3(AT)7(T)5) was found through the variation screening in this study.
We report a genome-wide search of Y-linked genes in Drosophila pseudoobscura. All six identifiable orthologs of the D. melanogaster Y-linked genes have autosomal inheritance in D. pseudoobscura. Four orthologs were investigated in detail and proved to be Y-linked in D. guanche and D. bifasciata, which shows that less than 18 million years ago the ancestral Drosophila Y chromosome was translocated to an autosome in the D. pseudoobscura lineage. We found 15 genes and pseudogenes in the current Y of D. pseudoobscura, and none are shared with the D. melanogaster Y. Hence, the Y chromosome in the D. pseudoobscura lineage appears to have arisen de novo and is not homologous to the D. melanogaster Y.
We developed a classification approach to multiple quantitative trait loci (QTL) mapping built upon a Bayesian framework that incorporates the important prior information that most genotypic markers are not cotransmitted with a QTL or their QTL effects are negligible. The genetic effect of each marker is modeled using a three-component mixture prior with a class for markers having negligible effects and separate classes for markers having positive or negative effects on the trait. The posterior probability of a marker's classification provides a natural statistic for evaluating credibility of identified QTL. This approach performs well, especially with a large number of markers but a relatively small sample size. A heat map to visualize the results is proposed so as to allow investigators to be more or less conservative when identifying QTL. We validated the method using a well-characterized data set for barley heading values from the North American Barley Genome Mapping Project. Application of the method to a new data set revealed sex-specific QTL underlying differences in glucose-6-phosphate dehydrogenase enzyme activity between two Drosophila species. A simulation study demonstrated the power of this approach across levels of trait heritability and when marker data were sparse.
Multiple mating by females establishes the opportunity for postcopulatory sexual selection favoring males whose sperm is preferentially employed in fertilizations. Here we use natural variation in a wild population of Drosophila melanogaster to investigate the genetic basis of sperm competitive ability. Approximately 101 chromosome 2 substitution lines were scored for components of sperm competitive ability (P1', P2', fecundity, remating rate, and refractoriness), genotyped at 70 polymorphic markers in 10 male reproductive genes, and measured for transcript abundance of those genes. Permutation tests were applied to quantify the statistical significance of associations between genotype and phenotype. Nine significant associations were identified between polymorphisms in the male reproductive genes and sperm competitive ability and 13 were identified between genotype and transcript abundance, but no significant associations were found between transcript abundance and sperm competitive ability. Pleiotropy was evident in two genes: a polymorphism in Acp33A associated with both P1' and P2' and a polymorphism in CG17331 associated with both elevated P2' and reduced refractoriness. The latter case is consistent with antagonistic pleiotropy and may serve as a mechanism maintaining genetic variation.
Most of the available SNP data have eluded valid population genetic analysis because most population genetical methods do not correctly accommodate the special discovery process used to identify SNPs. Most of the available SNP data have allele frequency distributions that are biased by the ascertainment protocol. We here show how this problem can be corrected by obtaining maximum-likelihood estimates of the true allele frequency distribution. In simple cases, the ML estimate of the true allele frequency distribution can be obtained analytically, but in other cases computational methods based on numerical optimization or the EM algorithm must be used. We illustrate the new correction method by analyzing some previously published SNP data from the SNP Consortium. Appropriate treatment of SNP ascertainment is vital to our ability to make correct inferences from the data of the International HapMap Project.
In Drosophila melanogaster, sperm and accessory gland proteins ("Acps," a major component of seminal fluid) transferred by males during mating trigger many physiological and behavioral changes in females (reviewed in ). Determining the genetic changes triggered in females by male-derived molecules and cells is a crucial first step in understanding female responses to mating and the female's role in postcopulatory processes such as sperm competition, cryptic female choice, and sexually antagonistic coevolution. We used oligonucleotide microarrays to compare gene expression in D. melanogaster females that were either virgin, mated to normal males, mated to males lacking sperm, or mated to males lacking both sperm and Acps. Expression of up to 1783 genes changed as a result of mating, most less than 2-fold. Of these, 549 genes were regulated by the receipt of sperm and 160 as a result of Acps that females received from their mates. The remaining genes whose expression levels changed were modulated by nonsperm/non-Acp aspects of mating. The mating-dependent genes that we have identified contribute to many biological processes including metabolism, immune defense, and protein modification.
Differences in gene expression are central to evolution. Such differences can arise from cis-regulatory changes that affect transcription initiation, transcription rate and/or transcript stability in an allele-specific manner, or from trans-regulatory changes that modify the activity or expression of factors that interact with cis-regulatory sequences. Both cis- and trans-regulatory changes contribute to divergent gene expression, but their respective contributions remain largely unknown. Here we examine the distribution of cis- and trans-regulatory changes underlying expression differences between closely related Drosophila species, D. melanogaster and D. simulans, and show functional cis-regulatory differences by comparing the relative abundance of species-specific transcripts in F1 hybrids. Differences in trans-regulatory activity were inferred by comparing the ratio of allelic expression in hybrids with the ratio of gene expression between species. Of 29 genes with interspecific expression differences, 28 had differences in cis-regulation, and these changes were sufficient to explain expression divergence for about half of the genes. Trans-regulatory differences affected 55% (16 of 29) of genes, and were always accompanied by cis-regulatory changes. These data indicate that interspecific expression differences are not caused by select trans-regulatory changes with widespread effects, but rather by many cis-acting changes spread throughout the genome.
The hemoglobin E variant (HbE; ( beta )26Glu-->Lys) is concentrated in parts of Southeast Asia where malaria is endemic, and HbE carrier status has been shown to confer some protection against Plasmodium falciparum malaria. To examine the effect of natural selection on the pattern of linkage disequilibrium (LD) and to infer the evolutionary history of the HbE variant, we analyzed biallelic markers surrounding the HbE variant in a Thai population. Pairwise LD analysis of HbE and 43 surrounding biallelic markers revealed LD of HbE extending beyond 100 kb, whereas no LD was observed between non-HbE variants and the same markers. The inferred haplotype network suggests a single origin of the HbE variant in the Thai population. Forward-in-time computer simulations under a variety of selection models indicate that the HbE variant arose 1,240-4,440 years ago. These results support the conjecture that the HbE mutation occurred recently, and the allele frequency has increased rapidly. Our study provides another clear demonstration that a high-resolution LD map across the human genome can detect recent variants that have been subjected to positive selection.
While there is considerable appeal to the idea of selecting a few SNPs to represent all, or much, of the DNA sequence variability in a local chromosomal region, it is also important to quantify what detail is lost in adopting such an approach. To address this issue, we compared high- and low-resolution depictions of sequence diversity for the same genomic region, the APOA1/C3/A4/A5 gene cluster on chromosome 11. First, extensive re-sequencing identified all nucleotide and sequence haplotype variation of the linked apolipoprotein genes in 72 individuals from three populations: African-Americans from Jackson, Miss., Europeans from North Karelia, Finland, and European-Americans from Rochester, Minn. We identified 124 SNPs in 17.7 kb and significant differences in variation among genes. APOC3 gene diversity was particularly distinctive at high resolution, showing large allele frequency differences ( F(ST) values >0.250) between Jackson and the other two samples, and divergent population-specific haplotype lineages. Next, we selected haplotype-tagging SNPs (htSNPs) for each gene, at a density of approximately one SNP per kb, using an algorithm suggested by Stram et al. (2003). The 17 htSNPs identified were then used to reconstruct low-resolution haplotypes, from which inferences about the structure of variation were also drawn. This comparison showed that while the htSNPs successfully tagged common haplotype variation, they also left much underlying sequence diversity undetected and failed, in some cases, to co-classify groups of closely related haplotypes. The implications of these findings for other haplotype-based descriptions of human variation are discussed.