AUTOTET: a program for analysis of autotetraploid genotypic data.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
With the advent of high-density DNA marker data sets for the mouse and other model systems, 100 or more genotype are routinely generated from large groups of mice. Issues of the accuracy and reliability of the genotyping are extremely important but often not addressed until genetic analysis is conducted. Simple tests that rely on the robust predictions arising from Mendelian genetics can be made quickly in the molecular laboratory as the data are generated, and require only a spreadsheet program. In this report, genotype data from 392 mice tested at 96 marker sites were analyzed for errors that are typical when handling large volumes of data generated in a repetitive process. The testing consisted of: (1) repeating the genotyping of approximately 1% of the samples; (2) examining the deviation from the expected segregation ratio ( 1:2:1 ) on a marker-by-marker basis; and (3) testing the correlation of the genotype at one marker with that at neighboring genetic markers on a chromosome. These three steps allowed analysis at the level of the microtiter plate, where errors are most likely to occur. A set of 96 dinucleotide repeat markers that are polymorphic between the C57BL/6J and DBA/2J mouse strains and can be multiplexed is reported for use in other genotyping projects.
The problem of inferring kinship structure among a sample of individuals using genetic markers is considered with the objective of developing hypothesis tests for genetic relatedness with nearly optimal properties. The class of tests considered are those that are constrained to be permutation invariant, which in this context defines tests whose properties do not depend on the labeling of the individuals. This is appropriate when all individuals are to be treated identically from a statistical point of view. The approach taken is to derive tests that are probably most powerful for a permutation invariant alternative hypothesis that is, in some sense, close to a null hypothesis of mutual independence. This is analagous to the locally most powerful test commonly used in parametric inference. Although the resulting test statistic is a U-statistic, normal approximation theory is found to be inapplicable because of high skewness. As an alternative it is found that a conditional procedure based on the most powerful test statistic can calculate accurate significance levels without much loss in power. Examples are given in which this type of test proves to be more powerful than a number of alternatives considered in the literature, including Queller and Goodknight's (1989) estimate of genetic relatedness, the average number of shared alleles (Blouin, 1996), and the number of feasible sibling triples (Almudevar and Field, 1999).
The identification of genes contributing to variation in complex phenotypes requires genetic data of high fidelity. Thus, the identification of pedigree and genotyping errors is a crucial prerequisite to the analysis of data from a genome scan for disease genes. The problem has been given little attention in most gene hunting papers; the focus has often been on eliminating mendelian inconsistencies in order that the analysis may proceed, rather than on achieving the best possible data. Though a number of computer programs are available to assist in the identification of genotyping and pedigree errors, the process is still not completely automated. While the Collaborative Study on the Genetics of Alcoholism (COGA) data set for GAW11 is completely compatible with Mendel's rules, there are still some errors present. We inspected the COGA data for the presence of additional errors, and identified five possible pedigree errors.
Explore the source record for details and available documents.
A computer algorithm for numerical evaluation of the statistical power of an exact test of Hardy-Weinberg genotypic proportions (HWP), developed here, indicates that the power is dependent on the number of segregating alleles as well as allele frequencies. While low levels of departure from the null hypothesis are difficult to detect from single-locus data, should such deviation be due to population substructuring, multiple loci, at each of which the number of segregating alleles is large (as seen with hypervariable loci), may easily detect even low levels of departure from HWP. Undetected small levels of departure may still provide conservative estimates of genotype frequencies from allele frequency data, following the current practice in forensic genetics.
Haplotype analysis is important for mapping traits. Recently, methods for estimating haplotype frequencies from genotypes of unrelated individuals based on the expectation-maximization (EM) algorithm have been developed. Our program estimates haplotype frequencies in the population and determines the posterior probability distribution of diplotype configuration (diplotype distribution) for each subject based on the estimated haplotype frequencies. Samples from three ethnic groups for the smoothelin gene (SMTN) and those from three Japanese groups for serum amyloid A genes (SAA@) were analyzed. The estimated diplotype distribution for each individual was concentrated, in most cases, in a single diplotype configuration. The diplotype configuration thus determined was the same as that determined in in vitro experiments, with one exception. Thus, the diplotype configurations determined using the estimated haplotype frequencies from unrelated individuals are reliable. Using this method, the risk of a subject developing a phenotype may be estimated from the diplotype distribution when the phenotype is associated with diplotype configurations.
SNP markers are becoming central for studying genetic determinants of complex diseases. Large SNP data collected in such studies call for the development of specialized analysis tools. We present methods for selecting sets of SNPs that can be associated to sample properties in case/control studies. We also describe how scoring and selection can be statistically tested. This is done at the single locus as well as at the set level.
We generalize an approach suggested by Hill (Heredity, 33, 229-239, 1974) for testing for significant association among alleles at two loci when only genotype and not haplotype frequencies are available. The principle is to use the Expectation-Maximization (EM) algorithm to resolve double heterozygotes into haplotypes and then apply a likelihood ratio test in order to determine whether the resolutions of haplotypes are significantly nonrandom, which is equivalent to testing whether there is statistically significant linkage disequilibrium between loci. The EM algorithm in this case relies on the assumption that genotype frequencies at each locus are in Hardy-Weinberg proportions. This method can accommodate X-linked loci and samples from haplodiploid species. We use three methods for testing significance of the likelihood ratio: the empirical distribution in a large number of randomized data sets, the X2 approximation for the distribution of likelihood ratios, and the Z2 test. The performance of each method is evaluated by applying it to simulated data sets and comparing the tail probability with the tail probability from Fisher's exact test applied to the actual haplotype data. For realistic sample sizes (50-150 individuals) all three methods perform well with two or three alleles per locus, but only the empirical distribution is adequate when there are five to eight alleles per locus, as is typical of hypervariable loci such as microsatellites. The method is applied to a data set of 32 microsatellite loci in a Finnish population and the results confirm the theoretical predictions. We conclude that with highly polymorphic loci, the EM algorithm does lead to a useful test for linkage disequilibrium, but that it is necessary to find the empirical distribution of likelihood ratios in order to perform a test of significance correctly.
The ability to confidently identify or exclude a population as the source of an individual has numerous powerful applications in molecular ecology. Several alternative assignment methods have recently been developed and are yet to be fully evaluated with empirical data. In this study we tested the efficacy of different assignment methods by using a translocated rock-wallaby (Petrogale lateralis) population, of known provenance. Specimens from the translocated population (n = 43), its known source population (n = 30) and four other nearby populations (n = 19-32) were genotyped for 11 polymorphic microsatellite loci. The results identified Bayesian clustering, frequency and Bayesian methods as the most consistent and accurate, correctly assigning 93-100% of individuals up to a significance threshold of P = 0.01. Performance was variable among the distance-based methods, with the Cavalli-Sforza and Edwards chord distance performing best, whereas Goldstein et al.'s (deltamu)2 consistently performed poorly. Using Bayesian clustering, frequency and Bayesian methods we then attempted to determine the source of rock-wallabies which have recently recolonized an outcrop (Gardners) 8 km from the nearest rock-wallaby population. Results indicate that the population at Gardners originated via a recent dispersal event from the eastern end of Mt. Caroline. This is only the second published record of dispersal by rock-wallabies between habitat patches and is the longest movement recorded to date. Molecular techniques and methods of analysis are now available to allow detailed studies of dispersal in rock-wallabies and should also be possible for many other taxa.
Abacavir is frequently used in antiretroviral combination therapies as a potent nucleoside reverse transcriptase inhibitor (NRTI). Four mutations are selected for by abacavir in vitro and in vivo: K65R, L74V, Y115F, and M184V. Abacavir resistance has also been observed in NRTI multidrug-resistant samples. Furthermore, abacavir resistance has been described in the context of zidovudine resistance. To evaluate the genetic basis of abacavir resistance, the viral genotype and phenotypic resistance were analyzed for 307 patient samples. Low- and high-level resistances were defined as 2.5- to 5.5-fold- and >5.5-fold-reduced susceptibility, respectively. If all samples with abacavir-selected and NRTI multidrug resistance-associated mutations were scored as resistant, 27.6% of the samples were misclassified, mainly due to samples falsely scored as susceptible. Therefore, the relative frequencies of other mutations were evaluated. Mutations at codons 44 and 118 were rarely detected in abacavir-susceptible samples but were overrepresented in resistant samples. Site-directed mutagenesis of E44D, V118I, and M184V resulted in low-level resistance for the double mutant 44/184 and the triple mutant. Low-level abacavir resistance was also detected for a viral clone carrying zidovudine mutations only. Additional insertion of M184V into the zidovudine background doubled the resistance, whereas 44/118 did not lead to a further increase. Incorporating combinations of zidovudine mutations and M184V into the scoring system markedly reduced the number of misclassified samples, whereas 44/118 did not improve the prediction. In conclusion, the combination of M184V with zidovudine mutations gives rise to high-level abacavir resistance, which may be clinically relevant. Thus, options for useful sequential combinations of NRTI are limited.
Many African countries currently use a sulfadoxine-pyrimethamine combination (SP) or amodiaquine (AQ) to treat uncomplicated Plasmodium falciparum malaria. Both drugs represent the last inexpensive alternatives to chloroquine. However, resistant P. falciparum populations are largely reported in Africa, and it is compulsory to know the present situation of resistance. The in vivo World Health Organization standard 28-day test was used to assess the efficacy of AQ and SP to treat uncomplicated falciparum malaria in Gabonese children under 10 years of age. To document treatment failures, molecular genotyping to distinguish therapeutic failures from reinfections and drug dosages were undertaken. A total of 118 and 114 children were given AQ or SP, respectively, and were monitored. SP was more effective than AQ, with 14.0 and 34.7% of therapeutic failures, respectively. Three days after initiation of treatment, the mean level of monodesethylamodiaquine (MdAQ) in plasma was 149 ng/ml in children treated with amodiaquine. In those treated with SP, mean levels of sulfadoxine and pyrimethamine in plasma were 100 microg/ml and 212 ng/ml, respectively. Levels of the three drugs were higher in patients successfully treated with AQ (MdAQ plasma levels) or SP (sulfadoxine and pyrimethamine plasma levels). Blood concentration higher than breakpoints of 135 ng/ml for MdAQ, 100 micro g/ml for sulfadoxine, and 175 ng/ml for pyrimethamine were associated with treatment success (odds ratio: 4.5, 9.8, and 11.8, respectively; all P values were <0.009). Genotyping of merozoite surface proteins 1 and 2 demonstrated a mean of 4.0 genotypes per person before treatment. At reappearance of parasitemia, both recrudescent parasites (represented by common bands in both samples) and newly inoculated parasites (represented by bands that were absent before treatment) were present in the blood of most (51.1%) children. Only 3 (6.4%) therapeutic failures were the result not of treatment inefficacy but of new infection. In areas where levels of drug resistance and complexity of infections are high, drug dosage and parasite genotyping may be of limited interest in improving the precision of drug efficacy measurement. Their use should be weighted according to logistical constraints.
We describe an approach for picking haplotype-tagging single nucleotide polymorphisms (htSNPs) that is presently being taken in two large nested case-control studies within a multiethnic cohort (MEC), which are engaged in a search for associations between risk of prostate and breast cancer and common genetic variations in candidate genes. Based on a preliminary sample of 70 control subjects chosen at random from each of the 5 ethnic groups in the MEC we estimate haplotype frequencies using a variant of the Excoffier-Slatkin E-M algorithm after genotyping a high density of SNPs selected every 3-5 kb in and surrounding a candidate gene. In order to evaluate the performance of a candidate set of htSNPS (which will be genotyped in the much larger case-control sample) we treat the haplotype frequencies estimate above as known, and carry out a formal calculation of the uncertainty of the number of copies of common haplotypes carried by an individual, summarizing this calculation as a coefficient of determination, R2h. A candidate set of htSNPS of a given size is chosen so as to maximize the minimum value of R2h over the common haplotypes, h.
Explore the source record for details and available documents.
Crucial to understanding the process of natural selection is characterizing phenotypic selection. Measures of phenotypic selection can be biased by environmental variation among individuals that causes a spurious correlation between a trait and fitness. One solution is analyzing genotypic data, rather than phenotypic data. Genotypic data, however, are difficult to gather, can be gathered from few species, and typically have low statistical power. Environmental correlations may act through traits other than through fitness itself. A path analytic framework, which includes measures of such traits, may reduce environmental bias in estimates of selection coefficients. We tested the efficacy of path analysis to reduce bias by re-analyzing three experiments where both phenotypic and genotypic data were available. All three consisted of plant species (Impatiens capensis, Arabidopsis thaliana, and Raphanus sativus) grown in experimental plots or the greenhouse. We found that selection coefficients estimated by path analysis using phenotypic data were highly correlated with those based on genotypic data with little systematic bias in estimating the strength of selection. Although not a panacea, using path analysis can substantially reduce environmental biases in estimates of selection coefficients. Such confidence in phenotypic selection estimates is critical for progress in the study of natural selection.
Pyrenean brown bears Ursus arctos are threatened with extinction. Management efforts to preserve this population require a comprehensive knowledge of the number and sex of the remaining individuals and their respective home ranges. This goal has been achieved using a combination of noninvasive genetic sampling of hair and faeces collected in the field and corresponding track size data. Genotypic data were collected at 24 microsatellite loci using a rigorous multiple-tubes approach to avoid genotyping errors associated with low quantities of DNA. Based on field and genetic data, the Pyrenean population was shown to be composed at least of one yearling, three adult males, and one adult female. These data indicate that extinction of the Pyrenean brown bear population is imminent without population augmentation. To preserve the remaining Pyrenean gene pool and increase genetic diversity, we suggest that managers consider population augmentation using only females. This study demonstrates that comprehensive knowledge of endangered small populations of mammals can be obtained using noninvasive genetic sampling.
This work has two purposes: (i) empirically selecting levels of significance that maximize the fraction of markers close to a gene (hit rate) when performing linkage analyses of simulated data and (ii) evaluating the utility of a previously reported scan statistic on the same data. Genotype data were simulated from a trait model of seven susceptibility genes. For purpose (i), five statistics were evaluated on all marker loci in fifty replicates; two-point lod and heterogeneity lod scores maximized over dominance (mlod, mhlod), a multi-allelic TDT test, an affected sib-pair test (ASP), and a model-free test on all sib-pairs (ALL_SIBS). Within each replicate the fraction of markers (hit rate) significant at specified levels of significance and also (a) within fifty markers of, or (b) on the same chromosome as a major gene was calculated. For purpose (ii), scan statistics of length 15 were calculated for each chromosome and their empirical significance levels estimated on the basis of 500 replicates generated under no linkage. The scan statistic was applied to the mhlod scores from one replicate (Replicate 5). Empirical p-values for the scan statistic were determined by computing mhlod scores on 500 replicates of simulated null data. For purpose (i), significance levels between 0.001 and 0.01 had the greatest hit rate for all five methods and both criteria. For criterion (a) at the 0.001 level of significance, both mlod and mhlod displayed the highest hit rates, approximately 0.4 for each. For criterion (b), all methods but ALL_SIBS and ASP had hit rates ranging between 0.4 and 0.5. For purpose (ii), the scan statistic proved equally or more powerful than the single-locus statistic for two of the seven susceptibility genes while the remaining five genes were not detected.
Linkage (genotypic) data from the 5q31-33 candidate region for asthma were contributed to Genetic Analysis Workshop 12 by members of the International Consortium on Asthma Genetics (COAG). Data came from five independent studies sampled from five countries. Genotypic data for a total of 26 markers were available, although the number of markers typed in each data set varied. Phenotypic and genotypic data was available from a total of 569 families and 3,175 subjects. The phenotypic data available varied among the studies; however information regarding physician-diagnosed asthma and total serum IgE levels was available in all five studies. This paper describes the ascertainment, data collection methods, phenotypic data, and genotypic data available for the single linkage region analyses undertaken for Genetic Analysis Workshop 12.