Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

Efficient reconstruction of haplotype structure via perfect phylogeny.

Each person's genome contains two copies of each chromosome, one inherited from the father and the other from the mother. A person's genotype specifies the pair of bases at each site, but does not specify which base occurs on which chromosome. The sequence of each chromosome separately is called a haplotype. The determination of the haplotypes within a population is essential for understanding genetic variation and the inheritance of complex diseases. The haplotype mapping project, a successor to the human genome project, seeks to determine the common haplotypes in the human population. Since experimental determination of a person's genotype is less expensive than determining its component haplotypes, algorithms are required for computing haplotypes from genotypes. Two observations aid in this process: first, the human genome contains short blocks within which only a few different haplotypes occur; second, as suggested by Gusfield, it is reasonable to assume that the haplotypes observed within a block have evolved according to a perfect phylogeny, in which at most one mutation event has occurred at any site, and no recombination occurred at the given region. We present a simple and efficient polynomial-time algorithm for inferring haplotypes from the genotypes of a set of individuals assuming a perfect phylogeny. Using a reduction to 2-SAT we extend this algorithm to handle constraints that apply when we have genotypes from both parents and child. We also present a hardness result for the problem of removing the minimum number of individuals from a population to ensure that the genotypes of the remaining individuals are consistent with a perfect phylogeny. Our algorithms have been tested on real data and give biologically meaningful results. Our webserver (http://www.cs.columbia.edu/compbio/hap/) is publicly available for predicting haplotypes from genotype data and partitioning genotype data into blocks.

Adult↗

High-resolution genomic profiles of human lung cancer.

Lung cancer is the leading cause of cancer mortality worldwide, yet there exists a limited view of the genetic lesions driving this disease. In this study, an integrated high-resolution survey of regional amplifications and deletions, coupled with gene-expression profiling of non-small-cell lung cancer subtypes, adenocarcinoma and squamous-cell carcinoma (SCC), identified 93 focal copy-number alterations, of which 21 span <0.5 megabases and contain a median of five genes. Whereas all known lung cancer genes/loci are contained in the dataset, most of these recurrent copy-number alterations are previously uncharacterized and include high-amplitude amplifications and homozygous deletions. Notably, despite their distinct histopathological phenotypes, adenocarcinoma and SCC genomic profiles showed a nearly complete overlap, with only one clear SCC-specific amplicon. Among the few genes residing within this amplicon and showing consistent overexpression in SCC is p63, a known regulator of squamous-cell differentiation. Furthermore, intersection with the published pancreatic cancer comparative genomic hybridization dataset yielded, among others, two focal amplicons on 8p12 and 20q11 common to both cancer types. Integrated DNA-RNA analyses identified WHSC1L1 and TPX2 as two candidates likely targeted for amplification in both pancreatic ductal adenocarcinoma and non-small-cell lung cancer.

Carcinoma, Non-Small-Cell Lung↗

Identification and analysis of axonemal dynein light chain 1 in primary ciliary dyskinesia patients.

Primary ciliary dyskinesia (PCD) is a genetically heterogeneous disorder characterized by chronic infections of the upper and lower airways, randomization of left/right body asymmetry, and reduced fertility. The phenotype results from dysfunction of motile cilia of the respiratory epithelium, at the embryonic node and of sperm flagella. Ultrastructural defects often involve outer dynein arms (ODAs), that are composed of several light (LCs), intermediate, and heavy (HCs) dynein chains. We recently showed that recessive mutations of DNAH5, the human ortholog of the biflagellate Chlamydomonas ODA gamma-HC, cause PCD. In Chlamydomonas, motor protein activity of the gamma-ODA-HC is regulated by binding of the axonemal LC1. We report the identification of the human (DNAL1) and murine (Dnal1) orthologs of the Chlamydomonas LC1-gene. Northern blot and in situ hybridization analyses revealed specific expression in testis, embryonic node, respiratory epithelium, and ependyma, resembling the DNAH5 expression pattern. In silico protein analysis showed complete conservation of the LC1/gamma-HC binding motif in DNAL1. Protein interaction studies demonstrated binding of DNAL1 and DNAH5. Based on these findings, we considered DNAL1 a candidate for PCD and sequenced all exons of DNAL1 in 86 patients. Mutational analysis was negative, excluding a major role of DNAL1 in the pathogenesis of PCD.

Amino Acid Motifs↗

High-density single-nucleotide polymorphism maps of the human genome.

Here we report a large, extensively characterized set of single-nucleotide polymorphisms (SNPs) covering the human genome. We determined the allele frequencies of 55,018 SNPs in African Americans, Asians (Japanese-Chinese), and European Americans as part of The SNP Consortium's Allele Frequency Project. A subset of 8333 SNPs was also characterized in Koreans. Because these SNPs were ascertained in the same way, the data set is particularly useful for modeling. Our results document that much genetic variation is shared among populations. For autosomes, some 44% of these SNPs have a minor allele frequency > or =10% in each population, and the average allele frequency differences between populations with different continental origins are less than 19%. However, the several percentage point allele frequency differences among the closely related Korean, Japanese, and Chinese populations suggest caution in using mixtures of well-established populations for case-control genetic studies of complex traits. We estimate that approximately 7% of these SNPs are private SNPs with minor allele frequencies <1%. A useful set of characterized SNPs with large allele frequency differences between populations (>60%) can be used for admixture studies. High-density maps of high-quality, characterized SNPs produced by this project are freely available.

Alleles↗

BLISS: binding site level identification of shared signal-modules in DNA regulatory sequences.

BACKGROUND: Regulatory modules are segments of the DNA that control particular aspects of gene expression. Their identification is therefore of great importance to the field of molecular genetics. Each module is composed of a distinct set of binding sites for specific transcription factors. Since experimental identification of regulatory modules is an arduous process, accurate computational techniques that supplement this process can be very beneficial. Functional modules are under selective pressure to be evolutionarily conserved. Most current approaches therefore attempt to detect conserved regulatory modules through similarity comparisons at the DNA sequence level. However, some regulatory modules, despite the conservation of their responsible binding sites, are embedded in sequences that have little overall similarity. RESULTS: In this study, we present a novel approach that detects conserved regulatory modules via comparisons at the binding site level. The technique compares the binding site profiles of orthologs and identifies those segments that have similar (not necessarily identical) profiles. The similarity measure is based on the inner product of transformed profiles, which takes into consideration the p values of binding sites as well as the potential shift of binding site positions. We tested this approach on simulated sequence pairs as well as real world examples. In both cases our technique was able to identify regulatory modules which could not to be identified using sequence-similarity based approaches such as rVista 2.0 and Blast. CONCLUSION: The results of our experiments demonstrate that, for sequences with little overall similarity at the DNA sequence level, it is still possible to identify conserved regulatory modules based solely on binding site profiles.

Animals↗

Peptide deformylase inhibitors as potent antimycobacterial agents.

Peptide deformylase (PDF) catalyzes the hydrolytic removal of the N-terminal formyl group from nascent proteins. This is an essential step in bacterial protein synthesis, making PDF an attractive target for antibacterial drug development. Essentiality of the def gene, encoding PDF from Mycobacterium tuberculosis, was demonstrated through genetic knockout experiments with Mycobacterium bovis BCG. PDF from M. tuberculosis strain H37Rv was cloned, expressed, and purified as an N-terminal histidine-tagged recombinant protein in Escherichia coli. A novel class of PDF inhibitors (PDF-I), the N-alkyl urea hydroxamic acids, were synthesized and evaluated for their activities against the M. tuberculosis PDF enzyme as well as their antimycobacterial effects. Several compounds from the new class had 50% inhibitory concentration (IC50) values of <100 nM. Some of the PDF-I displayed antibacterial activity against M. tuberculosis, including MDR strains with MIC90 values of <1 microM. Pharmacokinetic studies of potential leads showed that the compounds were orally bioavailable. Spontaneous resistance towards these inhibitors arose at a frequency of < or =5 x 10(-7) in M. bovis BCG. DNA sequence analysis of several spontaneous PDF-I-resistant mutants revealed that half of the mutants had acquired point mutations in their formyl methyltransferase gene (fmt), which formylated Met-tRNA. The results from this study validate M. tuberculosis PDF as a drug target and suggest that this class of compounds have the potential to be developed as novel antimycobacterial agents.

Administration, Oral↗

High genetic diversity in the chemoreceptor superfamily of Caenorhabditis elegans.

We investigated genetic polymorphism in the Caenorhabditis elegans srh and str chemoreceptor gene families, each of which consists of approximately 300 genes encoding seven-pass G-protein-coupled receptors. Almost one-third of the genes in each family are annotated as pseudogenes because of apparent functional defects in N2, the sequenced wild-type strain of C. elegans. More than half of these "pseudogenes" have only one apparent defect, usually a stop codon or deletion. We sequenced the defective region for 31 such genes in 22 wild isolates of C. elegans. For 10 of the 31 genes, we found an apparently functional allele in one or more wild isolates, suggesting that these are not pseudogenes but instead functional genes with a defective allele in N2. We suggest the term "flatliner" to describe genes whose functional vs. pseudogene status is unclear. Investigations of flatliner gene positions, d(N)/d(S) ratios, and phylogenetic trees indicate that they are not readily distinguished from functional genes in N2. We also report striking heterogeneity in the frequency of other polymorphisms among these genes. Finally, the large majority of polymorphism was found in just two strains from geographically isolated islands, Hawaii and Madeira. This suggests that our sampling of wild diversity in C. elegans is narrow and that identification of additional strains from similarly isolated regions will greatly expand the diversity available for study.

Alleles↗

Investigation of altering single-nucleotide polymorphism density on the power to detect trait loci and frequency of false positive in nonparametric linkage analyses of qualitative traits.

Genome-wide linkage analysis using microsatellite markers has been successful in the identification of numerous Mendelian and complex disease loci. The recent availability of high-density single-nucleotide polymorphism (SNP) maps provides a potentially more powerful option. Using the simulated and Collaborative Study on the Genetics of Alcoholism (COGA) datasets from the Genetics Analysis Workshop 14 (GAW14), we examined how altering the density of SNP marker sets impacted the overall information content, the power to detect trait loci, and the number of false positive results. For the simulated data we used SNP maps with density of 0.3 cM, 1 cM, 2 cM, and 3 cM. For the COGA data we combined the marker sets from Illumina and Affymetrix to create a map with average density of 0.25 cM and then, using a sub-sample of these markers, created maps with density of 0.3 cM, 0.6 cM, 1 cM, 2 cM, and 3 cM. For each marker set, multipoint linkage analysis using MERLIN was performed for both dominant and recessive traits derived from marker loci. Our results showed that information content increased with increased map density. For the homogeneous, completely penetrant traits we created, there was only a modest difference in ability to detect trait loci. Additionally, as map density increased there was only a slight increase in the number of false positive results when there was linkage disequilibrium (LD) between markers. The presence of LD between markers may have led to an increased number of false positive regions but no clear relationship between regions of high LD and locations of false positive linkage signals was observed.

Alcoholism↗

Identification of cytogenetic subgroups and karyotypic pathways of clonal evolution in follicular lymphomas.

Follicular lymphoma (FL) is characterized by the activation of BCL2 through t(14;18)(q32;q21). Additional acquired mutations are necessary to generate a fully malignant clonal proliferation. Many of these secondary genetic alterations are visible in the clonal karyotype; however, the sequence by which they arise and their influence on clinical behavior have not been determined. The ability to address these issues has been hampered by the lack of computational methods to manipulate complex chromosomal data in a sufficiently large cohort of cases. In the present investigation, we analyzed secondary karyotypic alterations in 336 cases of FL with t(14;18) to identify the most common regions of recurrent chromosomal gain or loss. This revealed 29 recurrent changes present in more than 5% of the tumors. Each tumor karyotype was then assessed for the presence or absence of each of these 29 specific changes. By statistical means, we show that the chromosomal changes arise in an apparent temporal order, with distinct early and late changes. We identify, by principal-components analysis, four possible cytogenetic pathways that characterize the early stages of clonal evolution, which converge to a common route at later stages. We show that FLs with t(14;18) may be classified into cytogenetic subgroups determined by the presence or absence of 6q-, +7, or der(18)t(14;18). Correlation with clinical outcomes in a subset of cases with clinical data revealed del(17p) and +12 to be correlated with an adverse clinical outcome. The clinical implications of these pathways of clonal evolution need to be examined on a prospective basis in a large cohort of FLs.

Chromosome Aberrations↗

On the structural differences between markers and genomic AC microsatellites.

AC microsatellites have proved particularly useful as genetic markers. For some purposes, such as in population biology, the inferences drawn depend on the quantitative values of their mutation rates. This, together with intrinsic biological interest, has led to widespread study of microsatellite mutational mechanisms. Now, however, inconsistencies are appearing in the results of marker-based versus non-marker-based studies of mutational mechanisms. The reasons for this have not been investigated, but one possibility, pursued here, is that the differences result from structural differences between markers and genomic microsatellites. Here we report a comparison between the CEPH AC marker microsatellites and the global population of AC microsatellites in the human genome. AC marker microsatellites are longer than the global average. Controlling for length, marker microsatellites contain on average fewer interruptions, and have longer segments, than their genomic counterparts. Related to this, marker microsatellites show a greater tendency to concentrate the majority of their repeats into one segment. These differences plausibly result from scientists selecting markers for their high polymorphism. In addition to the structural differences, there are differences in the base composition of flanking sequences, marker flanking regions being richer in C and G and poorer in A and T. Our results indicate that there are profound differences between marker and genomic microsatellites that almost certainly affect their mutation rates. There is a need for a unified model of mutational mechanisms that accounts for both marker-derived and genomic observations. A suggestion is made as to how this might be done.

Base Composition↗

Microsatellite linkage analysis, single-nucleotide polymorphisms, and haplotype associations with ECB21 in the COGA data.

This study, part of the Genetic Analysis Workshop 14 (GAW14), explored real Collaborative Study on the Genetics of Alcoholism data for linkage and association mapping between genetic polymorphisms (microsatellite and single-nucleotide polymorphisms (SNPs)) and beta (16.5-20 Hz) oscillations of the brain rhythms (ecb21). The ecb21 phenotype underwent the statistical adjustments for the age of participants, and for attaining a normal distribution. A total of 1,000 subjects' available phenotypes were included in linkage analysis with microsatellite markers. Linkage analysis was performed only for chromosome 4 where a quantitative trait locus with 5.01 LOD score had been previously reported. Previous findings related this location with the gamma-aminobutyric acid type A (GABAA) receptor. At the same location, our analysis showed a LOD score of 2.2. This decrease in the LOD score is the result of a drastic reduction (one-third) of the available GAW14 phenotypic data. We performed SNP and haplotype association analyses with the same phenotypic data under the linkage peak region on chromosome 4. Seven Affymetrix and two Illumina SNPs showed significant associations with ecb21 phenotype. A haplotype, a combination of SNPs TSC0044171 and TSC0551006 (the latter almost under the region of GABAA genes), showed a significant association with ecb21 (p = 0.015) and a relatively high frequency in the sample studied. Our results affirmed that the GABA region has potential of harboring genes that contribute quantitatively to the beta oscillation of the brain rhythms. The inclusion of the remaining 614 subjects, which in the GAW14 had missing data for the ecb21, can improve the strength of the associations as they have already shown that they contribute quite important information in the linkage analysis.

Alcoholism↗

Recombination and selection shape the molecular diversity pattern of nitrogen-fixing Sinorhizobium sp. associated to Medicago.

We investigate the genetic structure and molecular selection pattern of a sympatric population of Sinorhizobium meliloti and Sinorhizobium medicae. These bacteria fix nitrogen in association with plants of the genus Medicago. A set of 116 isolates were obtained from a soil sample, from root nodules of three groups of plants representing among-species, within-species and intraline diversity in the Medicago genus. Bacteria were characterized by sequencing at seven loci evenly distributed along the genome of both Sinorhizobium species, covering the chromosome and the two megaplasmids. We first test whether the diversity of host plants influence the bacterial diversity recovered. Using the same data set, we then analyse the selective pattern at each locus. There was no relationship between the diversity of Medicago plants that were used for sampling and the diversity of their symbionts. However, we found evidence of selection within each of the two main symbiotic regions, located on the two different megaplasmids. Purifying selection or a selective sweep was found to occur in the nod genomic region, which includes genes involved in nodulation specificity, whereas balancing selection was detected in the exo region, close to genes involved in exopolysaccharide production. Such pattern likely reflects the interaction between host plants and bacterial symbionts, with a possible conflict of interest between plants and cheater bacterial genotypes. Recombination appears to occur preferentially within and among loci located on megaplasmids, rather than within the chromosome. Thus, recombination may play an important role in resolving this conflict by allowing different selection patterns at different loci.

Databases, Genetic↗

TEL deletion analysis supports a novel view of relapse in childhood acute lymphoblastic leukemia.

PURPOSE: TEL (ETV6)-AML1 (RUNX1) chimeric gene fusions are frequent genetic abnormalities in childhood acute lymphoblastic leukemia (ALL). They often arise prenatally as early events or initiating events and are complemented by secondary postnatal genetic events of which deletion of the non-rearranged, second TEL allele is the most common. This consistent sequence of molecular pathogenesis facilitates an analysis of the clonal origins of relapse in this leukemia, which has some unusual clinical features. EXPERIMENTAL DESIGN: We compared the boundaries, by microsatellite mapping, of TEL deletions at relapse versus diagnosis in 15 informative patients. Moreover, we compared the relatedness of diagnostic and relapse clones using immunoglobulin and T-cell receptor genes rearrangements and clonotypic TEL-AML1 genomic fusion. RESULTS: Five patients retained the apparent same size TEL deletion, seven had larger deletions, and three had smaller deletions at relapse. In all of the cases evaluated, the clonal relatedness of diagnostic and relapse cells was confirmed by the retention of clonotypic TEL-AML1 genomic sequence and/or at least one identical immunoreceptor gene rearrangement. CONCLUSIONS: These data provide further evidence that TEL deletions are secondary to TEL-AML1 fusions in ALL. They are compatible with the novel idea that in at least some cases of childhood ALL, remission occurs with persistence of a preleukemic "fetal" clone, and subsequent relapse reflects the emergence of a new subclone from this reservoir after an independent "second hit," i.e., independent TEL deletion. To our knowledge, the study is the most extensive and comprehensive analysis of the relationship between diagnostic and relapse clones in childhood ALL presented thus far.

Artificial Gene Fusion↗

Low number of mitochondrial pseudogenes in the chicken (Gallus gallus) nuclear genome: implications for molecular inference of population history and phylogenetics.

BACKGROUND: Mitochondrial DNA has been detected in the nuclear genome of eukaryotes as pseudogenes, or Numts. Human and plant genomes harbor a large number of Numts, some of which have high similarity to mitochondrial fragments and thus may have been inadvertently included in population genetic and phylogenetic studies using mitochondrial DNA. Birds have smaller genomes relative to mammals, and the genome-wide frequency and distribution of Numts is still unknown. The release of a preliminary version of the chicken (Gallus gallus) genome by the Genome Sequencing Center at Washington University, St. Louis provided an opportunity to search this first avian genome for the frequency and characteristics of Numts relative to those in human and plants. RESULTS: We detected at least 13 Numts in the chicken nuclear genome. Identities between Numts and mitochondrial sequences varied from 58.6 to 88.8%. Fragments ranged from 131 to 1,733 nucleotides, collectively representing only 0.00078% of the nuclear genome. Because fewer Numts were detected in the chicken nuclear genome, they do not represent all regions of the mitochondrial genome and are not widespread in all chromosomes. Nuclear integrations in chicken seem to occur by a DNA intermediate and in regions of low gene density, especially in macrochromosomes. CONCLUSION: The number of Numts in chicken is low compared to those in human and plant genomes, and is within the range found for most sequenced eukaryotic genomes. For chicken, PCR amplifications of fragments of about 1.5 kilobases are highly likely to represent true mitochondrial amplification. Sequencing of these fragments should expose the presence of unusual features typical of pseudogenes, unless the nuclear integration is very recent and has not yet been mutated. Metabolic selection for compact genomes with reduced repetitive DNA and gene-poor regions where Numts occur may explain their low incidence in birds.

Animals↗

A corpus of GA4GH phenopackets: Case-level phenotyping for genomic diagnostics and discovery.

The Global Alliance for Genomics and Health (GA4GH) Phenopacket Schema was released in 2022 and approved by ISO as a standard for sharing clinical and genomic information about an individual, including phenotypic descriptions, numerical measurements, genetic information, diagnoses, and treatments. A phenopacket can be used as an input file for software that supports phenotype-driven genomic diagnostics and for algorithms that facilitate patient classification and stratification for identifying new diseases and treatments. There has been a great need for a collection of phenopackets to test software pipelines and algorithms. Here, we present Phenopacket Store. Phenopacket Store v.0.1.19 includes 6,668 phenopackets representing 475 Mendelian and chromosomal diseases associated with 423 genes and 3,834 unique pathogenic alleles curated from 959 different publications. This represents the first large-scale collection of case-level, standardized phenotypic information derived from case reports in the literature with detailed descriptions of the clinical data and will be useful for many purposes, including the development and testing of software for prioritizing genes and diseases in diagnostic genomics, machine learning analysis of clinical phenotype data, patient stratification, and genotype-phenotype correlations. This corpus also provides best-practice examples for curating literature-derived data using the GA4GH Phenopacket Schema.

Humans↗

Identifying susceptibility genes by using joint tests of association and linkage and accounting for epistasis.

Simulated Genetic Analysis Workshop 14 data were analyzed by jointly testing linkage and association and by accounting for epistasis using a candidate gene approach. Our group was unblinded to the "answers." The 48 single-nucleotide polymorphisms (SNPs) within the six disease loci were analyzed in addition to five SNPs from each of two non-disease-related loci. Affected sib-parent data was extracted from the first 10 replicates for populations Aipotu, Kaarangar, and Danacaa, and analyzed separately for each replicate. We developed a likelihood for testing association and/or linkage using data from affected sib pairs and their parents. Identical-by-descent (IBD) allele sharing between sibs was explicitly modeled using a conditional logistic regression approach and incorporating a covariate that represents expected IBD allele sharing given the genotypes of the sibs and their parents. Interactions were accounted for by performing likelihood ratio tests in stages determined by the highest order interaction term in the model. In the first stage, main effects were tested independently, and in subsequent stages, multilocus effects were tested conditional on significant marginal effects. A reduction in the number of tests performed was achieved by prescreening gene combinations with a goodness-of-fit chi square statistic that depended on mating-type frequencies. SNP-specific joint effects of linkage and association were identified for loci D1, D2, D3, and D4 in multiple replicates. The strongest effect was for SNP B03T3056, which had a median p-value of 1.98 x 10(-34). No two- or three-locus effects were found in more than one replicate.

Databases, Genetic↗

Functional study of a novel single deletion in the TITF1/NKX2.1 homeobox gene that produces congenital hypothyroidism and benign chorea but not pulmonary distress.

CONTEXT: We studied two sisters with congenital hypothyroidism and choreoathetosis but not respiratory distress. OBJECTIVE: The aim of this study was to establish the genetic defect that causes this phenotype and study the molecular mechanisms of the pathology by means of functional analysis. DESIGN: Sequencing of DNA, expression vectors generation, EMSAs, transfections experiments as well as bioinformatics analysis were performed. RESULTS: We found a new single deletion (825delC) in one allele of the TITF1/NKX2.1 gene. The mutation located in the C-terminal domain generates a nonsense thyroid transcription factor 1 (TTF1) protein, with 22 amino less and rich in positive charges. This protein shows diminished binding to DNA, does not interfere with wild-type (wt) TTF1 binding, and fails to activate reporter genes harboring the thyroglobulin (Tg), thyroperoxidase (TPO), or surfactant protein B (SP-B) promoters. In addition, the mutant (mut) protein has a dominant-negative effect on the transcriptional activity of wt TTF1 in a promoter-specific manner, inhibiting the transcription of Tg and TPO but not of SP-B. Using a Gal4 reporter system, we demonstrate that the mut protein is not transcriptionally active and does not likely compete with the wild type for coactivators. Interestingly, the mut protein impairs the wt capacity to synergize with paired box 8 (PAX8). This cooperation is necessary for Tg and TPO transcription but dispensable for SP-B expression. CONCLUSION: These results are concordant with the phenotype of the two sisters studied and demonstrate a differential role for TTF1 in the different tissues in which it is expressed.

Amino Acid Sequence↗

Exact tests of Hardy-Weinberg equilibrium and homogeneity of disequilibrium across strata.

Detecting departures from Hardy-Weinberg equilibrium (HWE) of marker-genotype frequencies is a crucial first step in almost all human genetic analyses. When a sample is stratified by multiple ethnic groups, it is important to allow the marker-allele frequencies to differ over the strata. In this situation, it is common to test for HWE by using an exact test within each stratum and then using the minimum P value as a global test. This approach does not account for multiple testing, and, because it does not combine information over strata, it does not have optimal power. Several approximate methods to combine information over strata have been proposed, but most of them sum over strata a measure of departure from HWE; if the departures are in different directions, then summing can diminish the overall evidence of departure from HWE. An exact stratified test is more appealing because it uses the probability of genotype configurations across the strata as evidence for global departures from HWE. We developed an exact stratified test for HWE for diallelic markers, such as single-nucleotide polymorphisms (SNPs), and an exact test for homogeneity of Hardy-Weinberg disequilibrium. By applying our methods to data from Perlegen and HapMap--a combined total of more than five million SNP genotypes, with three to four strata and strata sizes ranging from 23 to 60 subjects--we illustrate that the exact stratified test provides more-robust and more-powerful results than those obtained by either the minimum of exact test P values over strata or approximate stratified tests that sum measures of departure from HWE. Hence, our new methods should be useful for samples composed of multiple ethnic groups.

Computer Simulation↗