Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

PSIA: A Comprehensive Knowledgebase of Plant Self-incompatibility.

Self-incompatibility (SI) is an important genetic mechanism in angiosperms that prevents inbreeding and promotes outcrossing, with significant implications for crop breeding, including genetic diversity, hybrid seed production, and yield optimization. In eudicots, SI is typically governed by a single S-locus containing tightly linked pistil and pollen S-determinant genes. Despite major advances in SI research, a centralized, comprehensive resource for SI-related genomic data remains lacking. To address this gap, we developed the Plant Self-Incompatibility Atlas (PSIA), a systematically curated knowledgebase providing an extensive compilation of plant SI, including genomic resources for SI species, S gene annotations, molecular mechanisms, phylogenetic relationships, and comparative genomic analyses. The current release of PSIA includes over 500 genome assemblies from 469 SI species. Using known S genes as queries, we manually identified and rigorously curated 3700 S genes. PSIA provides detailed S-locus information from assembled genomes of SI species and offers an interactive platform for browsing, BLAST searches, S gene analysis, and data retrieval. Additionally, PSIA serves as a unique platform for comparative genomic studies of S-loci, facilitating exploration of the dynamic processes underlying the origin, loss, and regain of SI. As a comprehensive and user-friendly resource, PSIA will greatly advance our understanding of angiosperm SI and serve as a valuable tool for crop breeding and hybrid seed production. PSIA is freely available at http://www.plantsi.cn.

Self-Incompatibility in Flowering Plants↗

Genetic alterations in cancer knowledge system: analysis of gene mutations in mouse and human liver and lung tumors.

Mutational incidence and spectra for genes examined in both human and mouse lung and liver tumors were analyzed using the National Institute of Environmental Health Sciences (NIEHS) Genetic Alterations in Cancer (GAC) knowledge system. GAC is a publicly available, web-based system for evaluating data obtained from peer-reviewed studies of genetic changes in tumors associated with exposure to chemical, physical, or biological agents, as well as spontaneous tumors. In mice, mutations in Kras2 and Hras-1 were the most common events reported for lung and liver tumors, respectively, whether chemically induced or spontaneous. There was a significant difference in Kras2 mutation incidence for spontaneous versus induced mouse lung tumors and in Hras-1 mutation incidence and spectrum for spontaneous versus induced mouse liver tumors. The major gene changes reported for human lung and liver tumors were in KRAS2 (lung only) and TP53. The KRAS2 mutation incidence was similar for spontaneous and asbestos-induced human lung tumors, while the TP53 mutation incidence differed significantly. Aflatoxin B1, hepatitis B virus, hepatitis C virus, and vinyl chloride all caused TP53 mutations in human liver tumors, but the mutation spectrum for each agent differed. The incidence of KRAS2 mutations in human compared to mouse lung tumors differed significantly, as did the incidence of Hras and p53 gene mutations in human compared to mouse liver tumors. Differences observed in the mutation spectra for agent-induced compared to spontaneous tumors and similarities in spectra for structurally similar agents support the concept that mutation spectra can serve as a "fingerprint" of exposure based on chemical structure.

Amines↗

Transferability of tag SNPs to capture common genetic variation in DNA repair genes across multiple populations.

Genetic association studies can be made more cost-effective by exploiting linkage disequilibrium patterns between nearby single-nucleotide polymorphisms (SNPs). The International HapMap Project now offers a dense SNP map across the human genome in four population samples. One question is how well tag SNPs chosen from a resource like HapMap can capture common variation in independent disease samples. To address the issue of tag SNP transferability, we genotyped 2,783 SNPs across 61 genes (with a total span of 6 Mb) involved in DNA repair in 466 individuals from multiple populations. We picked tag SNPs in samples with European ancestry from the Centre d'Etude du Polymorphisme Humain, and evaluated coverage of common variation in the other samples. Our comparative analysis shows that common variation in non-African samples can be captured robustly with only marginal loss in terms of the maximum r2. We also evaluated the transferability of specified multi-marker haplotypes as predictors for untyped SNPs, and demonstrate that they provide equivalent coverage compared to single-marker tests (pairwise tags) while requiring fewer SNPs for genotyping. The efficacy of a tagging-based approach in studying genotype-phenotype correlations in complex traits is strongly supported by our empirical results.

Computational Biology↗

Physical mapping of autonomic/sympathetic candidate genetic loci for hypertension in the human genome: a somatic cell radiation hybrid library approach.

Allelic variation at multiple genetic loci may contribute to hypertension. Since autonomic/sympathetic dysfunction may play an early, pathogenic, heritable role in hypertension, we evaluated candidate loci likely to contribute to such dysfunction, including catecholamine biosynthetic enzymes, catecholamine transporters, neuropeptides, and adrenergic receptors. Since chromosomal locations and physical map positions of many of these loci had not yet been identified, we used the GeneBridge4 human/hamster radiation (somatic cell) hybrid library panel (resolution approximately 1 to approximately 1.5 Mb), along with specifically designed oligonucleotide primers and PCR (200-400 bp products) to position these loci in the human genome. Primers were designed from sequences outside the coding regions (3'-flanking or intronic segments) to avoid cross-species (hamster) amplification. Chromosomal positions were assigned in cR (centi-Ray) units ( approximately 270 Kbp/cR(3000) for GeneBridge 4). A total of 13 loci were newly assigned chromosomal positions; of particular interest was a cluster of adrenergic candidate loci on chromosome 5q (including ADRB2, ADRA1A, DRD1, GPRK6, and NPY6R), a region harbouring linkage peaks for blood pressure. Such physical map positions will enable more precise selection of polymorphic microsatellite and single nucleotide polymorphism markers at these loci, to aid in linkage and association studies of autonomic/sympathetic dysfunction in human hypertension.

Autonomic Nervous System↗

Evidence for a microRNA expansion in the bilaterian ancestor.

Understanding how animal complexity has arisen and identifying the key genetic components of this process is a central goal of evolutionary developmental biology. The discovery of microRNAs (miRNAs) as key regulators of development has identified a new set of candidates for this role. microRNAs are small noncoding RNAs that regulate tissue-specific or temporal gene expression through base pairing with target mRNAs. The full extent of the evolutionary distribution of miRNAs is being revealed as more genomes are scrutinized. To explore the evolutionary origins of metazoan miRNAs, we searched the genomes of diverse animals occupying key phylogenetic positions for homologs of experimentally verified human, fly, and worm miRNAs. We identify 30 miRNAs conserved across bilaterians, almost double the previous estimate. We hypothesize that this larger than previously realized core set of miRNAs was already present in the ancestor of all Bilateria and likely had key roles in allowing the evolution of diverse specialist cell types, tissues, and complex morphology. In agreement with this hypothesis, we found only three, conserved miRNA families in the genome of the sea anemone Nematostella vectensis and no convincing family members in the genome of the demosponge Reniera sp. The dramatic expansion of the miRNA repertoire in bilaterians relative to sponges and cnidarians suggests that increased miRNA-mediated gene regulation accompanied the emergence of triploblastic organ-containing body plans.

Algorithms↗

HaploRec: efficient and accurate large-scale reconstruction of haplotypes.

BACKGROUND: Haplotypes extracted from human DNA can be used for gene mapping and other analysis of genetic patterns within and across populations. A fundamental problem is, however, that current practical laboratory methods do not give haplotype information. Estimation of phased haplotypes of unrelated individuals given their unphased genotypes is known as the haplotype reconstruction or phasing problem. RESULTS: We define three novel statistical models and give an efficient algorithm for haplotype reconstruction, jointly called HaploRec. HaploRec is based on exploiting local regularities conserved in haplotypes: it reconstructs haplotypes so that they have maximal local coherence. This approach--not assuming statistical dependence for remotely located markers--has two useful properties: it is well-suited for sparse marker maps, such as those used in gene mapping, and it can actually take advantage of long maps. CONCLUSION: Our experimental results with simulated and real data show that HaploRec is a powerful method for the large scale haplotyping needed in association studies. With sample sizes large enough for gene mapping it appeared to be the best compared to all other tested methods (Phase, fastPhase, PL-EM, Snphap, Gerbil; simulated data), with small samples it was competitive with the best available methods (real data). HaploRec is several orders of magnitude faster than Phase and comparable to the other methods; the running times are roughly linear in the number of subjects and the number of markers. HaploRec is publicly available at http://www.cs.helsinki.fi/group/genetics/haplotyping.html.

Chromosome Mapping↗

Candidate gene analysis of the Price Foundation anorexia nervosa affected relative pair dataset.

The eating disorders are severe psychiatric illnesses with significant morbidity and mortality that exhibit statistically significant familial risk and heritability, providing support for a molecular genetic approach toward defining etiological factors. An emerging candidate gene literature has concentrated on serotinergic and dopaminergic candidates. With the financial support of the Price Foundation, a group of investigators initiated an international multi-center collaboration (Price Foundation Collaborative Group) in 1995 to study the genetics of anorexia and bulimia nervosa by collecting and analyzing phenotypes and genotypes of individuals and their relatives affected with eating disorders. The first sample of families collected by this collaborative group, known as the Price Foundation Anorexia Nervosa Affected Relative Pair (AN-ARP) dataset, was ascertained on an proband affected with Anorexia Nervosa (AN), with relative pairs affected with the eating disorders AN, Bulimia Nervosa or Eating Disorders Not Otherwise Specified [1]. Biognosis U.S., Inc. was founded to identify and characterize candidate susceptibility genes for anorexia and bulimia nervosa phenotypes in the Price Foundation eating disorder datasets. During 2000-2001, Biognosis U.S., Inc. developed and implemented a research program with a focus on the analysis of candidate genes nominated by neurochemical characteristics of eating disorder patients [2], serotonergic and dopaminergic candidate gene polymorphisms [3], neuroendocrine regulation of appetite [4], and by a positional hypothesis from a linkage analysis of the AN-ARP dataset [5]. This report reviews the anorexia nervosa candidate gene literature through 2001, the candidate gene research program implemented at Biognosis U.S., Inc. and selected candidate gene findings in the AN-ARP dataset derived from that research program.

Animals↗

A network of investigator networks in human genome epidemiology.

The task of identifying genetic determinants for complex, multigenetic diseases is hampered by small studies, publication and reporting biases, and lack of common standards worldwide. The authors propose the creation of a network of networks that include groups of investigators collecting data for human genome epidemiology research. Twenty-three networks of investigators addressing specific diseases or research topics and representing several hundreds of teams have already joined this initiative. For each field, the authors are currently creating a core registry of teams already participating in the respective network. A wider international registry will include all other teams also working in the same field. Independent investigators are invited to join the registries and existing networks and to join forces in creating additional ones as needed. The network of networks aims to register these networks, teams, and investigators; be a resource for information about or connections to the many networks; offer methodological support; promote sound design and standardization of analytical practices; generate inclusive overviews of fields at large; facilitate rapid confirmation of findings; and avoid duplication of effort.

Databases, Genetic↗

Spline-fitting with a genetic algorithm: a method for developing classification structure-activity relationships.

Classification methods allow for the development of structure-activity relationship models when the target property is categorical rather than continuous. We describe a classification method which fits descriptor splines to activities, with descriptors selected using a genetic algorithm. This method, which we identify as SFGA, is compared to the well-established techniques of recursive partitioning (RP) and soft independent modeling by class analogy (SIMCA) using five series of compounds: cyclooxygenase-2 (COX-2) inhibitors, benzodiazepine receptor (BZR) ligands, estrogen receptor (ER) ligands, dihydrofolate reductase (DHFR) inhibitors, and monoamine oxidase (MAO) inhibitors. Only 1-D and 2-D descriptors were used. Approximately 40% of compounds in each series were assigned to a test set, "cherry-picked" from the complete set such that they lie outside the training set as much as possible. SFGA produced models that were more predictive for all but the DHFR set, for which SIMCA was most predictive. RP gave the least predictive models for all but the MAO set. A similar trend was observed when using training and test sets to which compounds were randomly assigned and when gradually eliminating compounds from the (designed) training set. The stability of models was examined for the random and reduced sets, where stability means that classification statistics and the selected descriptors are similar for models derived from different sets. Here, SIMCA produced the most stable models, followed by SFGA and RP. We show that a consensus approach that combines all three methods outperforms the single best model for all data sets.

Algorithms↗

The importance of modelling heterogeneity in complex disease: application to NIMH Schizophrenia Genetics Initiative data.

As for other complex diseases, linkage analyses of schizophrenia (SZ) have produced evidence for numerous chromosomal regions, with inconsistent results reported across studies. The presence of locus heterogeneity appears likely and may reduce the power of linkage analyses if homogeneity is assumed. In addition, when multiple heterogeneous datasets are pooled, inter-sample variation in the proportion of linked families (alpha) may diminish the power of the pooled sample to detect susceptibility loci, in spite of the larger sample size obtained. We compare the significance of linkage findings obtained using allele-sharing LOD scores (LOD(exp))-which assume homogeneity-and heterogeneity LOD scores (HLOD) in European American and African American NIMH SZ families. We also pool these two samples and evaluate the relative power of the LOD(exp) and two different heterogeneity statistics. One of these (HLOD-P) estimates the heterogeneity parameter alpha only in aggregate data, while the second (HLOD-S) determines alpha separately for each sample. In separate and combined data, we show consistently improved performance of HLOD scores over LOD(exp). Notably, genome-wide significant evidence for linkage is obtained at chromosome 10p in the European American sample using a recessive HLOD score. When the two samples are combined, linkage at the 10p locus also achieves genome-wide significance under HLOD-S, but not HLOD-P. Using HLOD-S, improved evidence for linkage was also obtained for a previously reported region on chromosome 15q. In linkage analyses of complex disease, power may be maximised by routinely modelling locus heterogeneity within individual datasets, even when multiple datasets are combined to form larger samples.

Chromosomes, Human, Pair 10↗

Aspects of cancer immunotherapy.

Cancer immunotherapy has traditionally undergone a 'revolution' every decade, from the use of Bacille Calmette-Guérin by scarification in the 1970s, to interleukin-2 therapies in the 1980s, and monoclonal antibody treatments in the early 1990s. Usually the early reports on the use of such agents were encouraging, but when more patients were studied in multiple centres, the initial promising results could not be confirmed. Now in a new century, we have more reagents and methods available than ever before - indeed, with such a plethora of reagents it is difficult to envisage them being fully and appropriately tested within the next decade, by which time there will be even more reagents to test. However, there have been three major advances which should lead to substantial progress in cancer immunotherapy: (1) the widespread use of genetic engineering, enabling identification of candidate vaccine proteins and manipulation of their sequences; (2) the production of antigens, antibodies and cytokines in large amounts by recombinant technologies, and (3) an understanding of the mode of presentation of peptides by major histocompatibility complex Class I and Class II molecules and their recognition by T cells. Despite these advances, there are major problems facing cancer immunotherapy, such as the ability of tumours to mutate and evade the immune system and the difficulty of precisely defining the interactions of effector cells in mediating 'rejection' or destruction of a tumour. There are clearly immunological similarities with diseases such as malaria and schistosomiasis, where the invading foreign organisms can use a variety of strategies to resist an elicited immune response. The failure to find a suitable vaccine for these diseases must lead to some pessimism for the development of immunotherapy for an autologous tumour. However, there are promising studies now in progress which should give an indication of the most important directions to follow. This review provides a commentary on aspects of cancer immunotherapy and in particular will deal with: (1) the selection of antigens as vaccine components; (2) the modes of presentation of antigens, particularly by major histocompatibility complex Class I molecules; and (3) new modes of delivery of vaccine immunogens.

Antigen Presentation↗

Glyceraldehyde-3-phosphate dehydrogenase mediates anoxia response and survival in Caenorhabditis elegans.

Oxygen deprivation has a role in the pathology of many human diseases. Thus it is of interest in understanding the genetic and cellular responses to hypoxia or anoxia in oxygen-deprivation-tolerant organisms such as Caenorhabditis elegans. In C. elegans the DAF-2/DAF-16 pathway, an IGF-1/insulin-like signaling pathway, is involved with dauer formation, longevity, and stress resistance. In this report we compared the response of wild-type and daf-2(e1370) animals to anoxia. Unlike wild-type animals, the daf-2(e1370) animals have an enhanced anoxia-survival phenotype in that they survive long-term anoxia and high-temperature anoxia, do not accumulate significant tissue damage in either of these conditions, and are motile after 24 hr of anoxia. RNA interference was used to screen DAF-16-regulated genes that suppress the daf-2(e1370)-enhanced anoxia-survival phenotype. We identified gpd-2 and gpd-3, two nearly identical genes in an operon that encode the glycolytic enzyme glyceraldehyde-3-phosphate dehydrogenase. We found that not only is the daf-2(e1370)-enhanced anoxia phenotype dependent upon gpd-2 and gpd-3, but also the motility of animals exposed to brief periods of anoxia is prematurely arrested in gpd-2/3(RNAi) and daf-2(e1370);gpd-2/3(RNAi) animals. These data suggest that gpd-2 and gpd-3 may serve a protective role in tissue exposed to oxygen deprivation.

Animals↗

Linkage disequilibrium mapping via cladistic analysis of phase-unknown genotypes and inferred haplotypes in the Genetic Analysis Workshop 14 simulated data.

We recently described a method for linkage disequilibrium (LD) mapping, using cladistic analysis of phased single-nucleotide polymorphism (SNP) haplotypes in a logistic regression framework. However, haplotypes are often not available and cannot be deduced with certainty from the unphased genotypes. One possible two-stage approach is to infer the phase of multilocus genotype data and analyze the resulting haplotypes as if known. Here, haplotypes are inferred using the expectation-maximization (EM) algorithm and the best-guess phase assignment for each individual analyzed. However, inferring haplotypes from phase-unknown data is prone to error and this should be taken into account in the subsequent analysis. An alternative approach is to analyze the phase-unknown multilocus genotypes themselves. Here we present a generalization of the method for phase-known haplotype data to the case of unphased SNP genotypes. Our approach is designed for high-density SNP data, so we opted to analyze the simulated dataset. The marker spacing in the initial screen was too large for our method to be effective, so we used the answers provided to request further data in regions around the disease loci and in null regions. Power to detect the disease loci, accuracy in localizing the true site of the locus, and false-positive error rates are reported for the inferred-haplotype and unphased genotype methods. For this data, analyzing inferred haplotypes outperforms analysis of genotypes. As expected, our results suggest that when there is little or no LD between a disease locus and the flanking region, there will be no chance of detecting it unless the disease variant itself is genotyped.

Chromosome Mapping↗

STABIX: summary-statistic-based GWAS indexing and compression.

MOTIVATION: Genome-wide association studies (GWAS) are widely used to investigate the role of genetics in disease traits, but the resulting file sizes from these studies are large, posing barriers to efficient storage, sharing, and querying. This issue is especially important for biobanks like the UK Biobank that publish GWAS for thousands of traits, increasing the volume of data that must be effectively managed. Current compression and query methods reduce file sizes and allow for quick genomic position-based queries but do not provide utility for quickly finding loci based on their summary statistics. For example, finding all SNVs in a particular p-value range would require decompressing and scanning the whole file. We propose a new tool, STABIX, which introduces summary-statistic-based queries and improves upon the standard bgzip compression and Tabix query tool in both compression ratio and decompression speed. RESULTS: When applied to 10 GWAS files from PanUKBB, STABIX created smaller compressed data and indices than Tabix for all files, where bgzip and tbi files were an average of 1.2 times the size of STABIX compressed files and indexes. In the same 10 files, STABIX per gene decompression was, on average 7× faster than Tabix per gene decompression, and achieved faster per gene decompression times for over 99% of nearly 20,000 genes. AVAILABILITY AND IMPLEMENTATION: Software freely available for download at GitHub: https://github.com/kristen-schneider/stabix/.

Genome-Wide Association Study↗

Identification of a genetic determinant responsible for host specificity in Streptococcus thermophilus bacteriophages.

Phage-host interactions remain poorly understood in lactic acid bacteria and essentially in all Gram-positive bacteria. The aim of this study was to identify the phage genetic determinant (anti-receptor) involved in the recognition of Streptococcus thermophilus hosts. The complete genomic sequence of the lytic S. thermophilus phage DT1 was determined previously, and bioinformatic analysis indicated that orf18 might be the anti-receptor gene. The orf18 of six additional S. thermophilus phages was determined (DT2, DT4, MD1, MD2, MD4 and Q5) and compared with the orf18 of DT1. The deduced ORF18 was divided into three domains. The first domain, which contains the N-terminal part of the protein, was conserved in all seven phages. The second domain was detected in only two phages and flanked by a motif called collagen-like repeats. The second domain also contained a variable region (VR1). All seven phages had a third domain that consisted of the C-terminal section of the protein as well as another variable region (VR2). Chimeric DT1 phages were constructed by recombination; a portion of its orf18 was replaced by the corresponding section in orf18 of the phage MD4. All DT1 chimeric phages acquired the host range of phage MD4. Analysis of the orf18 in the chimeric phages revealed that host specificity in phages DT1 and MD4 resulted from VR2. This is the first report on the identification and characterization of a phage gene involved in the host recognition process of Gram-positive bacteria.

Adsorption↗

Transcript level alterations reflect gene dosage effects across multiple tissues in a mouse model of down syndrome.

Human trisomy 21, which results in Down syndrome (DS), is one of the most complicated congenital genetic anomalies compatible with life, yet little is known about the molecular basis of DS. It is generally accepted that chromosome 21 (Chr21) transcripts are overexpressed by about 50% in cells with an extra copy of this chromosome. However, this assumption is difficult to test in humans due to limited access to tissues, and direct support for this idea is available for only a few Chr21 genes or in a limited number of tissues. The Ts65Dn mouse is widely used as a model for studies of DS because it is at dosage imbalance for the orthologs of about half of the 284 Chr21 genes. Ts65Dn mice have several features that directly parallel developmental anomalies of DS. Here we compared the expression of 136 mouse orthologs of Chr21 genes in nine tissues of the trisomic and euploid mice. Nearly all of the 77 genes which are at dosage imbalance in Ts65Dn showed increased transcript levels in the tested tissues, providing direct support for a simple model of increased transcription proportional to the gene copy number. However, several genes escaped this rule, suggesting that they may be controlled by additional tissue-specific regulatory mechanisms revealed in the trisomic situation.

Animals↗

Analysis of variation in expression of autosomal dominant osteopetrosis type 2: searching for modifier genes.

INTRODUCTION: Autosomal Dominant Osteopetrosis type II (ADO2) is a heritable osteosclerotic disorder that results from heterozygous mutations in the ClCN7 gene. Analysis of ADO2 in our pedigrees indicates that the penetrance is 66%, with a highly variable phenotype. METHODS: To identify genes that modify disease status, we performed a 10 cM genome-wide scan using 400 microsatellite markers in 112 subjects from our 8 largest ADO2 families with mutations in the ClCN7 gene. Results were analyzed by parametric linkage analysis using autosomal dominant and recessive models for affects on disease status. Follow-up genotyping with additional microsatellite markers was performed for regions with LOD scores over 1.5. In addition, we compared the frequency of two nonsynonymous SNPs, rs12926089 (V418M) and rs11559208 (K691E), and one promoter SNP rs960467 in the normal ClCN7 allele between a sample of unaffected gene carriers and clinically affected subjects to test the hypothesis that genetic variation in the non-disease allele within the ClCN7 gene might influence disease expression. RESULTS: We found potential evidence of linkage for a modifier gene(s) on 9q21-22 with a LOD score of 1.89, which is not statistically significant, but interesting. We also found that, for SNP V418M on the non-disease allele with the wild-type ClCN7 sequence, 94.92% (56/59) of clinically affected subjects and 78.13% (25/32) of unaffected gene carriers had a valine while 5.08% (3/59) of the affected subjects and 21.88% (7/32) of unaffected gene carriers had a methionine (P < 0.03). Unfortunately, SNP K691E was not informative in our families. For SNP rs960467, on the non-disease allele with the wild-type ClCN7 gene, 87.93% (51/58) of clinically affected subjects and 62.50% (20/32) of unaffected gene carriers had a C allele while 12.07% (7/58) of the clinically affected subjects and 37.50% (12/32) of unaffected gene carriers had a T allele (P < 0.007). As expected, the polymorphisms on the disease allele were not associated with disease status. CONCLUSIONS: Chromosome 9q21-22 may harbor a modifier gene(s) that affect(s) ADO2 disease status and severity. Additionally, we find the associations between the polymorphisms on the non-disease allele and unaffected gene carrier status.

Alleles↗

Divergent haplotypes and human history as revealed in a worldwide survey of X-linked DNA sequence variation.

The population genetic history of a 10.1-kbp noncoding region of the human X chromosome was studied using the males of the HGDP-CEPH Human Genome Diversity Panel (672 individuals from 52 populations). The geographic distribution of patterns of variation was roughly consistent with previous studies, with the major exception that 1 highly divergent haplotype (haplotype X, hX) was observed at low frequency in widely scattered non-African populations and not at all observed in sub-Saharan African populations. Microsatellite (short tandem repeat) variation within the sequenced region was low among copies of hX, even though the estimated time of ancestry of hX and other sequences was 1.44 Myr. The estimated age of the common ancestor of all hX copies was 5,230 years (95% consistency index: 2,000-75,480 years). To further address the presence of hX in Africa, additional samples from Chad and Tanzania were screened. Five additional copies of hX were observed, consistent with a history in which hX was present in Africa prior to the migration of modern humans out of Africa and with eastern Africa being the source of non-African modern human populations. Taken together, these features of hX-that it is much older than other haplotypes and uncommon and patchily distributed throughout Africa, Europe, and Asia-present a cautionary tale for interpretations of human history.

Chromosomes, Human, X↗