Search PubMed⌕ Search

Biomedical subjects

Andrew G Clark

Publications and source records attributed to Andrew G Clark.

At least 19 recordsLinked to original sources

Asymmetric development and function of paired sperm-storage organs in Drosophila melanogaster.

Paired structures often have similar forms and functions, but the processes underlying their formation can differ. They may originate from a common source or from parallel sources, or arise from distinct precursors that follow separate developmental pathways, ultimately converging on comparable structures and roles. When asymmetries emerge and persist through development, members of the pair can specialize in ways that might increase fitness. Here, we report that the Drosophila melanogaster female's pair of spermathecae, which appear similar and have the common role of sperm storage, derive from different developmental compartments defined by expression of lineage-tracing markers corresponding, respectively, to the key patterning genes engrailed and wingless. We further find that the two spermathecae show significant differences in size, secretory activity, and calcium levels and, perhaps as a consequence, sperm retention dynamics. These results open broad avenues for understanding how developmental, physiological, and behavioral asymmetries arise and impact reproductive success.

Animals↗

A 3.9-centimorgan-resolution human single-nucleotide polymorphism linkage map and screening set.

Recent advances in technologies for high-throughout single-nucleotide polymorphism (SNP)-based genotyping have improved efficiency and cost so that it is now becoming reasonable to consider the use of SNPs for genomewide linkage analysis. However, a suitable screening set of SNPs and a corresponding linkage map have yet to be described. The SNP maps described here fill this void and provide a resource for fast genome scanning for disease genes. We have evaluated 6,297 SNPs in a diversity panel composed of European Americans, African Americans, and Asians. The markers were assessed for assay robustness, suitable allele frequencies, and informativeness of multi-SNP clusters. Individuals from 56 Centre d'Etude du Polymorphisme Humain pedigrees, with >770 potentially informative meioses altogether, were genotyped with a subset of 2,988 SNPs, for map construction. Extensive genotyping-error analysis was performed, and the resulting SNP linkage map has an average map resolution of 3.9 cM, with map positions containing either a single SNP or several tightly linked SNPs. The order of markers on this map compares favorably with several other linkage and physical maps. We compared map distances between the SNP linkage map and the interpolated SNP linkage map constructed by the deCode Genetics group. We also evaluated cM/Mb distance ratios in females and males, along each chromosome, showing broadly defined regions of increased and decreased rates of recombination. Evaluations indicate that this SNP screening set is more informative than the Marshfield Clinic's commonly used microsatellite-based screening set.

Alleles↗

Linkage disequilibrium and inference of ancestral recombination in 538 single-nucleotide polymorphism clusters across the human genome.

The prospect of using linkage disequilibrium (LD) for fine-scale mapping in humans has attracted considerable attention, and, during the validation of a set of single-nucleotide polymorphisms (SNPs) for linkage analysis, a set of data for 4,833 SNPs in 538 clusters was produced that provides a rich picture of local attributes of LD across the genome. LD estimates may be biased depending on the means by which SNPs are first identified, and a particular problem of ascertainment bias arises when SNPs identified in small heterogeneous panels are subsequently typed in larger population samples. Understanding and correcting ascertainment bias is essential for a useful quantitative assessment of the landscape of LD across the human genome. Heterogeneity in the population recombination rate, rho=4Nr, along the genome reflects how variable the density of markers will have to be for optimal coverage. We find that ascertainment-corrected rho varies along the genome by more than two orders of magnitude, implying great differences in the recombinational history of different portions of our genome. The distribution of rho is unimodal, and we show that this is compatible with a wide range of mixtures of hotspots in a background of variable recombination rate. Although rho is significantly correlated across the three population samples, some regions of the genome exhibit population-specific spikes or troughs in rho that are too large to be explained by sampling. This result is consistent with differences in the genealogical depth of local genomic regions, a finding that has direct bearing on the design and utility of LD mapping and on the National Institutes of Health HapMap project.

Alleles↗

Natural variation in human membrane transporter genes reveals evolutionary and functional constraints.

Membrane transporters maintain cellular and organismal homeostasis by importing nutrients and exporting toxic compounds. Transporters also play a crucial role in drug response, serving as drug targets and setting drug levels. As part of a pharmacogenetics project, we screened exons and flanking intronic regions for variation in a set of 24 membrane transporter genes (96 kb; 57% coding) in 247 DNA samples from ethnically diverse populations. We identified 680 single nucleotide polymorphisms (SNPs), of which 175 were synonymous and 155 caused amino acid changes, and 29 small insertions and deletions. Amino acid diversity (pi(NS)) in transmembrane domains (TMDs) was significantly lower than in loop domains, suggesting that TMDs have special functional constraints. This difference was especially striking in the ATP-binding cassette superfamily and did not parallel evolutionary conservation: there was little variation in the TMDs, even in evolutionarily unconserved residues. We used allele frequency distribution to evaluate different scoring systems (Grantham, blosum62, SIFT, and evolutionarily conservedevolutionarily unconserved) for their ability to predict which SNPs affect function. Our underlying assumption was that alleles that are functionally deleterious will be selected against and thus under represented at high frequencies and over represented at low frequencies. We found that evolutionary conservation of orthologous sequences, as assessed by evolutionarily conservedevolutionarily unconserved and SIFT, was the best predictor of allele frequency distribution and hence of function. European Americans had an excess of high frequency alleles in comparison to African Americans, consistent with a historic bottleneck. In addition, African Americans exhibited a much higher frequency of population specific medium-frequency alleles than did European Americans.

DNA↗

Molecular population genetics of inducible antibacterial peptide genes in Drosophila melanogaster.

Insects respond to septic infection in part by producing a suite of antimicrobial peptides that may be subject to host-pathogen coevolutionary dynamics. In order to infer population genetic forces acting on Drosophila antibacterial peptide genes, we examine global properties of polymorphism and divergence in the Drosophila melanogaster defensin, drosocin, metchnikowin, attacin C, diptericin A, and cecropin A, B, and C genes. As a functional class, antibacterial peptides exhibit low levels of interspecific amino acid divergence. There are multiple amino acid polymorphisms segregating within D. melanogaster, however, a high proportion of which change the charge or polarity of the variable residue. These polymorphisms are particularly prevalent in processed signal and propeptide domains. We find that models of coevolutionary "arms races" and selectively maintained hypervariability do not adequately describe the population dynamics of mature antibacterial peptides in D. melanogaster, but that a highly significant excess of high-frequency derived polymorphisms coupled with substantial intralocus linkage disequilibrium suggests that positive selection may act on antibacterial peptide genes. Some attributes of the data may be consistent with a simple demographic model of population founding followed by expansion, but departures from the equilibrium null tend to be more pronounced in the peptide genes than at other loci around the genome.

Amino Acid Sequence↗

Tracing the evolutionary history of Drosophila regulatory regions with models that identify transcription factor binding sites.

Much of evolutionary change is mediated at the level of gene expression, yet our understanding of regulatory evolution remains unsatisfying. In light of recent data indicating that transcription factor binding sites undergo substantial turnover between species, we attempt to quantify the process of binding site turnover in regulatory regions of well-studied genes controlling embryonic patterning in Drosophila. We examine polymorphism and divergence data in Drosophila melanogaster and four related species from regulatory regions of five early development genes for which functional binding sites have been identified. This analysis reveals that Drosophila regulatory regions exhibit patterns of variation consistent with functional constraint. We develop a novel approach to binding site prediction which we use to characterize the process of binding site divergence in regulatory regions. This method uses sets of known binding sites to construct a model that predicts transcription factor specificity and bootstrap sampling to derive significance levels. This approach allows appropriate significance levels to be determined even in the face of skewed base composition in the background sequence. Using this approach, we show that, although functional elements exhibit conservation of sequence, there is substantial potential to gain new functional elements within the regulatory regions. Our results show that application of models that predict transcription factor binding sites can yield insights into the process and dynamics of binding site evolution within regulatory regions.

Animals↗

Finding genes underlying risk of complex disease by linkage disequilibrium mapping.

Identification of genes that harbor variation associated with inter-individual differences in risk of complex diseases remains one of the most challenging and important problems in human genetics. For genetic variants that are sufficiently common and have sufficiently large effects, direct tests of association through linkage disequilibrium with anonymous SNPs may prove effective. But the two critical parameters - the frequency of risk-inflating alleles and the magnitudes of their effect on risk - remain largely unknown. In this review we consider the latest information regarding the likely efficacy of the linkage disequilibrium mapping approach.

Chromosome Mapping↗

Y chromosome and other heterochromatic sequences of the Drosophila melanogaster genome: how far can we go?

Whole genome shotgun assemblies have proven remarkably successful in reconstructing the bulk of euchromatic genes, with the only limit appearing to be determined by the sequencing depth. For genes imbedded in heterochromatin, however, the low cloning efficiency of repetitive sequences, combined with the computational challenges, demand that additional clues be used to annotate the sequences. One approach that has proven very successful in identifying protein coding genes in Y-linked heterochromatin of Drosophila melanogaster has been to make a BLASTable database of the small, unmapped contigs and fragments leftover at the end of a shotgun assembly, and to attempt to capture these by blasting with an appropriate query sequence. This approach often yields a staggered alignment of contigs from the unmapped set to the query sequence, as though the disjoint contigs represent small portions of the gene. Further inspection frequently shows that the contigs are broken by very large, heterochromatic introns. Methods of this sort are being expanded to make best use of all available clues to determine which unmapped contigs are associated with genes. These include use of EST libraries, and, in the case of the Y chromosome, testing of male specific genes and reduced shotgun depth of relevant contigs. It appears much more hopeful than anyone would have imagined that whole genome shotgun assemblies can recover the great bulk of even heterochromatic genes.

Animals↗

Robustness of inference of haplotype block structure.

In this report, we examine the validity of the haplotype block concept by comparing block decompositions derived from public data sets by variants of several leading methods of block detection. We first develop a statistical method for assessing the concordance of two block decompositions. We then assess the robustness of inferred haplotype blocks to the specific detection method chosen, to arbitrary choices made in the block-detection algorithms, and to the sample analyzed. Although the block decompositions show levels of concordance that are very unlikely by chance, the absolute magnitude of the concordance may be low enough to limit the utility of the inference. For purposes of SNP selection, it seems likely that methods that do not arbitrarily impose block boundaries among correlated SNPs might perform better than block-based methods.

Algorithms↗

Bayesian sperm competition estimates.

We introduce a Bayesian method for estimating parameters for a model of multiple mating and sperm displacement from genotype counts of brood-structured data. The model is initially targeted for Drosophila melanogaster, but is easily adapted to other organisms. The method is appropriate for use with field studies where the number of mates and the genotypes of the mates cannot be controlled, but where unlinked markers have been collected for a set of females and a sample of their offspring. Advantages over previous approaches include full use of multilocus information and the ability to cope appropriately with missing data and ambiguities about which alleles are maternally vs. paternally inherited. The advantages of including X-linked markers are also demonstrated.

Animals↗

Sequence diversity and haplotype structure in the human ABCB1 (MDR1, multidrug resistance transporter) gene.

OBJECTIVES: There is increasing evidence that polymorphism of the ABCB1 (MDR1) gene contributes to interindividual variability in bioavailability and tissue distribution of P-glycoprotein substrates. The aim of the present study was to (1) identify and describe novel variants in the ABCB1 gene, (2) understand the extent of variation in ABCB1 at the population level, (3) analyze how variation in ABCB1 is structured in haplotypes, and (4) functionally characterize the effect of the most common amino acid change in P-glycoprotein. METHODS AND RESULTS: Forty-eight variant sites, including 30 novel variants and 13 coding for amino acid changes, were identified in a collection of 247 ethnically diverse DNA samples. These variants comprised 64 statistically inferred haplotypes, 33 of which accounted for 92% of chromosomes analyzed. The two most common haplotypes, ABCB1*1 and ABCB1*13, differed at six sites (three intronic, two synonymous, and one non-synonymous) and were present in 36% of all chromosomes. Significant population substructure was detected at both the nucleotide and haplotype level. Linkage disequilibrium was significant across the entire ABCB1 gene, especially between the variant sites found in ABCB1*13, and recombination was inferred. The Ala893Ser change found in the common ABCB1*13 haplotype did not affect P-glycoprotein function. CONCLUSION: This study represents a comprehensive analysis of ABCB1 nucleotide diversity and haplotype structure in different populations and illustrates the importance of haplotype considerations in characterizing the functional consequences of ABCB1 polymorphisms.

Base Sequence↗

Natural selection shaped regional mtDNA variation in humans.

Human mtDNA shows striking regional variation, traditionally attributed to genetic drift. However, it is not easy to account for the fact that only two mtDNA lineages (M and N) left Africa to colonize Eurasia and that lineages A, C, D, and G show a 5-fold enrichment from central Asia to Siberia. As an alternative to drift, natural selection might have enriched for certain mtDNA lineages as people migrated north into colder climates. To test this hypothesis we analyzed 104 complete mtDNA sequences from all global regions and lineages. African mtDNA variation did not significantly deviate from the standard neutral model, but European, Asian, and Siberian plus Native American variations did. Analysis of amino acid substitution mutations (nonsynonymous, Ka) versus neutral mutations (synonymous, Ks) (kaks) for all 13 mtDNA protein-coding genes revealed that the ATP6 gene had the highest amino acid sequence variation of any human mtDNA gene, even though ATP6 is one of the more conserved mtDNA proteins. Comparison of the kaks ratios for each mtDNA gene from the tropical, temperate, and arctic zones revealed that ATP6 was highly variable in the mtDNAs from the arctic zone, cytochrome b was particularly variable in the temperate zone, and cytochrome oxidase I was notably more variable in the tropics. Moreover, multiple amino acid changes found in ATP6, cytochrome b, and cytochrome oxidase I appeared to be functionally significant. From these analyses we conclude that selection may have played a role in shaping human regional mtDNA variation and that one of the selective influences was climate.

Africa↗

The genome sequence of the malaria mosquito Anopheles gambiae.

Anopheles gambiae is the principal vector of malaria, a disease that afflicts more than 500 million people and causes more than 1 million deaths each year. Tenfold shotgun sequence coverage was obtained from the PEST strain of A. gambiae and assembled into scaffolds that span 278 million base pairs. A total of 91% of the genome was organized in 303 scaffolds; the largest scaffold was 23.1 million base pairs. There was substantial genetic variation within this strain, and the apparent existence of two haplotypes of approximately equal frequency ("dual haplotypes") in a substantial fraction of the genome likely reflects the outbred nature of the PEST strain. The sequence produced a conservative inference of more than 400,000 single-nucleotide polymorphisms that showed a markedly bimodal density distribution. Analysis of the genome sequence revealed strong evidence for about 14,000 protein-encoding transcripts. Prominent expansions in specific families of proteins likely involved in cell adhesion and immunity were noted. An expressed sequence tag analysis of genes regulated by blood feeding provided insights into the physiological adaptations of a hematophagous insect.

Animals↗

Contributions of 18 additional DNA sequence variations in the gene encoding apolipoprotein E to explaining variation in quantitative measures of lipid metabolism.

Apolipoprotein E (ApoE) is a major constituent of many lipoprotein particles. Previous genetic studies have focused on six genotypes defined by three alleles, denoted epsilon2, epsilon3, and epsilon4, encoded by two variable exonic sites that segregate in most populations. We have reported studies of the distribution of alleles of 20 biallelic variable sites in the gene encoding the ApoE molecule within and among samples, ascertained without regard to health, from each of three populations: African Americans from Jackson, Miss.; Europeans from North Karelia, Finland; and non-Hispanic European Americans from Rochester, Minn. Here we ask (1) how much variation in blood levels of ApoE (lnApoE), of total cholesterol (TC), of high-density lipoprotein cholesterol (HDL-C), and of triglyceride (lnTG) is statistically explained by variation among APOE genotypes defined by the epsilon2, epsilon3, and epsilon4 alleles; (2) how much additional variation in these traits is explained by genotypes defined by combining the two variable sites that define these three alleles with one or more additional variable sites; and (3) what are the locations and relative allele frequencies of the sites that define multisite genotypes that significantly improve the statistical explanation of variation beyond that provided by the genotypes defined by the epsilon2, epsilon3, and epsilon4 alleles, separately for each of the six gender-population strata. This study establishes that the use of only genotypes defined by the epsilon2, epsilon3, and epsilon4 alleles gives an incomplete picture of the contribution that the variation in the APOE gene makes to the statistical explanation of interindividual variation in blood measurements of lipid metabolism. The addition of variable sites to the genotype definition significantly improved the ability to explain variation in lnApoE and in TC and resulted in the explanation of variation in HDL-C and in lnTG. The combination of additional sites that explained the greatest amount of trait variation was different for different traits and varied among the six gender-population strata. The role that noncoding variable sites play in the explanation of pleiotropic effects on different measures of lipid metabolism reveals that both regulatory and structural functional variation in the APOE gene influences measures of lipid metabolism. This study demonstrates that resequencing of the complete gene in a sample of >/=20 individuals and an evaluation of all combinations of the identified variable sites, separately for each population and interacting environmental context, may be necessary to fully characterize the impact that a gene has on variation in related traits of a metabolic system.

Alleles↗

Sequence polymorphism at the human apolipoprotein AII gene ( APOA2): unexpected deficit of variation in an African-American sample.

A 3.3-kb region, encompassing the APOA2 gene and 2 kb of 5' and 3' flanking DNA, was re-sequenced in a "core" sample of 24 individuals, sampled without regard to the health from each of three populations: African-Americans from Jackson (Miss., USA), Europeans from North Karelia (Finland), and non-Hispanic European-Americans from Rochester, (Minn., USA). Fifteen variable sites were identified (14 SNPs and one multi-allelic microsatellite, all silent), and these sites segregated as 18 sequence haplotypes (or nine, if SNPs only are considered). The haplotype distribution in the core African-American sample was unusual, with a deficit of particular haplotypes compared with those found in the other two samples, and a significantly (P<0.05) low level of nucleotide diversity relative to patterns of polymorphism and divergence at other human loci. Six of the 14 SNPs, whose variation captured the haplotype structure of the core data, were then genotyped by oligonucleotide ligation assay in an additional 2183 individuals from the same three populations (n=843, n=452, and n=888, respectively). All six sites varied in each of the larger "epidemiological" samples, and together, they defined 19 SNP haplotypes, seven with relative frequencies greater than 1% in the total sample; all of these common haplotypes had been identified earlier in the core re-sequencing survey. Here also, the African-American sample showed significantly lower SNP heterozygosity and haplotype diversity than the other two samples. The deficit of polymorphism is consistent with a population-specific non-neutral increase in the relative frequency of several haplotypes in Jackson.

Alleles↗

Linkage disequilibrium and the mapping of complex human traits.

The potential value of haplotypes defined by several single nucleotide polymorphisms has attracted recent interest. With sufficient linkage disequilibrium (LD), haplotypes could be used in association studies to map common alleles that might influence the susceptibility to common diseases, as well as for reconstructing the evolution of the genome. It has been proposed that a globally useful resource need only be based on high frequency variants, identified from a few modest samples. Rapid progress has been made in quantifying the pattern of human LD and haplotypes defined by such common variants within and among populations. However, the quality and utility of the proposed LD-based resource could be seriously compromised if important sampling and analytical factors are overlooked in its design. The LD map should be based on adequately justified criteria defined by sound population genetic principles.

Chromosome Mapping↗