Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genotype Data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Methods for analysis and visualization of SNP genotype data for complex diseases.

SNP markers are becoming central for studying genetic determinants of complex diseases. Large SNP data collected in such studies call for the development of specialized analysis tools. We present methods for selecting sets of SNPs that can be associated to sample properties in case/control studies. We also describe how scoring and selection can be statistically tested. This is done at the single locus as well as at the set level.

Bayes Theorem↗

Testing for linkage disequilibrium in genotypic data using the Expectation-Maximization algorithm.

We generalize an approach suggested by Hill (Heredity, 33, 229-239, 1974) for testing for significant association among alleles at two loci when only genotype and not haplotype frequencies are available. The principle is to use the Expectation-Maximization (EM) algorithm to resolve double heterozygotes into haplotypes and then apply a likelihood ratio test in order to determine whether the resolutions of haplotypes are significantly nonrandom, which is equivalent to testing whether there is statistically significant linkage disequilibrium between loci. The EM algorithm in this case relies on the assumption that genotype frequencies at each locus are in Hardy-Weinberg proportions. This method can accommodate X-linked loci and samples from haplodiploid species. We use three methods for testing significance of the likelihood ratio: the empirical distribution in a large number of randomized data sets, the X2 approximation for the distribution of likelihood ratios, and the Z2 test. The performance of each method is evaluated by applying it to simulated data sets and comparing the tail probability with the tail probability from Fisher's exact test applied to the actual haplotype data. For realistic sample sizes (50-150 individuals) all three methods perform well with two or three alleles per locus, but only the empirical distribution is adequate when there are five to eight alleles per locus, as is typical of hypervariable loci such as microsatellites. The method is applied to a data set of 32 microsatellite loci in a Finnish population and the results confirm the theoretical predictions. We conclude that with highly polymorphic loci, the EM algorithm does lead to a useful test for linkage disequilibrium, but that it is necessary to find the empirical distribution of likelihood ratios in order to perform a test of significance correctly.

Algorithms↗

Microsatellite autosomal genotyping data in four indigenous populations from El Salvador.

Fifteen microsatellite loci (D3S1358, TH01, D21S11, D18S51, PENTA E, D5S818, D13S317, D7S820, D16S539, CSF1PO, PENTA D, vWA, D8S1179, TPOX, and FGA) have been genotyped in four indigenous populations from El Salvador (Central America), namely, Conchagua, Izalco, Panchimalco, and San Alejo. Here we have obtained values for several indices of forensic interest for these population samples. Population differentiation test showed no significant statistical differences between these four populations, and an AMOVA test indicates that most of the genetic variation (approximately 100%) occurs within individuals. Population pairwise genetic comparisons with other population samples seem to indicate the existence of a major Native American component in the populations from El Salvador.

DNA Fingerprinting↗

Modeling and E-M estimation of haplotype-specific relative risks from genotype data for a case-control study of unrelated individuals.

The US National Cancer Institute has recently sponsored the formation of a Cohort Consortium (http://2002.cancer.gov/scpgenes.htm) to facilitate the pooling of data on very large numbers of people, concerning the effects of genes and environment on cancer incidence. One likely goal of these efforts will be generate a large population-based case-control series for which a number of candidate genes will be investigated using SNP haplotype as well as genotype analysis. The goal of this paper is to outline the issues involved in choosing a method of estimating haplotype-specific risk estimates for such data that is technically appropriate and yet attractive to epidemiologists who are already comfortable with odds ratios and logistic regression. Our interest is to develop and evaluate extensions of methods, based on haplotype imputation, that have been recently described (Schaid et al., Am J Hum Genet, 2002, and Zaykin et al., Hum Hered, 2002) as providing score tests of the null hypothesis of no effect of SNP haplotypes upon risk, which may be used for more complex tasks, such as providing confidence intervals, and tests of equivalence of haplotype-specific risks in two or more separate populations. In order to do so we (1) develop a cohort approach towards odds ratio analysis by expanding the E-M algorithm to provide maximum likelihood estimates of haplotype-specific odds ratios as well as genotype frequencies; (2) show how to correct the cohort approach, to give essentially unbiased estimates for population-based or nested case-control studies by incorporating the probability of selection as a case or control into the likelihood, based on a simplified model of case and control selection, and (3) finally, in an example data set (CYP17 and breast cancer, from the Multiethnic Cohort Study) we compare likelihood-based confidence interval estimates from the two methods with each other, and with the use of the single-imputation approach of Zaykin et al. applied under both null and alternative hypotheses. We conclude that so long as haplotypes are well predicted by SNP genotypes (we use the Rh2 criteria of Stram et al. [1]) the differences between the three methods are very small and in particular that the single imputation method may be expected to work extremely well.

Algorithms↗

Selecting tagging SNPs for association studies using power calculations from genotype data.

Recent studies have indicated that linkage disequilibrium (LD) between single nucleotide polymorphism (SNP) markers can be used to derive a reduced set of tagging SNPs (tSNPs) for genetic association studies. Previous strategies for identifying tSNPs have focused on LD measures or haplotype diversity, but the statistical power to detect disease-associated variants using tSNPs in genetic studies has not been fully characterized. We propose a new approach of selecting tSNPs based on determining the set of SNPs with the highest power to detect association. Two-locus genotype frequencies are used in the power calculations. To show utility, we applied this power method to a large number of SNPs that had been genotyped in Caucasian samples. We demonstrate that a significant reduction in genotyping efforts can be achieved although the reduction depends on genotypic relative risk, inheritance mode and the prevalence of disease in the human population. The tSNP sets identified by our method are remarkably robust to changes in the disease model when small relative risk and additive mode of inheritance are employed. We have also evaluated the ability of the method to detect unidentified SNPs. Our findings have important implications in applying tSNPs from different data sources in association studies.

Algorithms↗

Haplotype frequency estimation error analysis in the presence of missing genotype data.

BACKGROUND: Increasingly researchers are turning to the use of haplotype analysis as a tool in population studies, the investigation of linkage disequilibrium, and candidate gene analysis. When the phase of the data is unknown, computational methods, in particular those employing the Expectation-Maximisation (EM) algorithm, are frequently used for estimating the phase and frequency of the underlying haplotypes. These methods have proved very successful, predicting the phase-known frequencies from data for which the phase is unknown with a high degree of accuracy. Recently there has been much speculation as to the effect of unknown, or missing allelic data - a common phenomenon even with modern automated DNA analysis techniques - on the performance of EM-based methods. To this end an EM-based program, modified to accommodate missing data, has been developed, incorporating non-parametric bootstrapping for the calculation of accurate confidence intervals. RESULTS: Here we present the results of the analyses of various data sets in which randomly selected known alleles have been relabelled as missing. Remarkably, we find that the absence of up to 30% of the data in both biallelic and multiallelic data sets with moderate to strong levels of linkage disequilibrium can be tolerated. Additionally, the frequencies of haplotypes which predominate in the complete data analysis remain essentially the same after the addition of the random noise caused by missing data. CONCLUSIONS: These findings have important implications for the area of data gathering. It may be concluded that small levels of drop out in the data do not affect the overall accuracy of haplotype analysis perceptibly, and that, given recent findings on the effect of inaccurate data, ambiguous data points are best treated as unknown.

Alleles↗

Tag SNP selection in genotype data for maximizing SNP prediction accuracy.

MOTIVATION: The search for genetic regions associated with complex diseases, such as cancer or Alzheimer's disease, is an important challenge that may lead to better diagnosis and treatment. The existence of millions of DNA variations, primarily single nucleotide polymorphisms (SNPs), may allow the fine dissection of such associations. However, studies seeking disease association are limited by the cost of genotyping SNPs. Therefore, it is essential to find a small subset of informative SNPs (tag SNPs) that may be used as good representatives of the rest of the SNPs. RESULTS: We define a new natural measure for evaluating the prediction accuracy of a set of tag SNPs, and use it to develop a new method for tag SNPs selection. Our method is based on a novel algorithm that predicts the values of the rest of the SNPs given the tag SNPs. In contrast to most previous methods, our prediction algorithm uses the genotype information and not the haplotype information of the tag SNPs. Our method is very efficient, and it does not rely on having a block partition of the genomic region. We compared our method with two state-of-the-art tag SNP selection algorithms on 58 different genotype datasets from four different sources. Our method consistently found tag SNPs with considerably better prediction ability than the other methods. AVAILABILITY: The software is available from the authors on request.

Algorithms↗

Source population of dispersing rock-wallabies (Petrogale lateralis) identified by assignment tests on multilocus genotypic data.

The ability to confidently identify or exclude a population as the source of an individual has numerous powerful applications in molecular ecology. Several alternative assignment methods have recently been developed and are yet to be fully evaluated with empirical data. In this study we tested the efficacy of different assignment methods by using a translocated rock-wallaby (Petrogale lateralis) population, of known provenance. Specimens from the translocated population (n = 43), its known source population (n = 30) and four other nearby populations (n = 19-32) were genotyped for 11 polymorphic microsatellite loci. The results identified Bayesian clustering, frequency and Bayesian methods as the most consistent and accurate, correctly assigning 93-100% of individuals up to a significance threshold of P = 0.01. Performance was variable among the distance-based methods, with the Cavalli-Sforza and Edwards chord distance performing best, whereas Goldstein et al.'s (deltamu)2 consistently performed poorly. Using Bayesian clustering, frequency and Bayesian methods we then attempted to determine the source of rock-wallabies which have recently recolonized an outcrop (Gardners) 8 km from the nearest rock-wallaby population. Results indicate that the population at Gardners originated via a recent dispersal event from the eastern end of Mt. Caroline. This is only the second published record of dispersal by rock-wallabies between habitat patches and is the longest movement recorded to date. Molecular techniques and methods of analysis are now available to allow detailed studies of dispersal in rock-wallabies and should also be possible for many other taxa.

Animals↗

Prediction of abacavir resistance from genotypic data: impact of zidovudine and lamivudine resistance in vitro and in vivo.

Abacavir is frequently used in antiretroviral combination therapies as a potent nucleoside reverse transcriptase inhibitor (NRTI). Four mutations are selected for by abacavir in vitro and in vivo: K65R, L74V, Y115F, and M184V. Abacavir resistance has also been observed in NRTI multidrug-resistant samples. Furthermore, abacavir resistance has been described in the context of zidovudine resistance. To evaluate the genetic basis of abacavir resistance, the viral genotype and phenotypic resistance were analyzed for 307 patient samples. Low- and high-level resistances were defined as 2.5- to 5.5-fold- and >5.5-fold-reduced susceptibility, respectively. If all samples with abacavir-selected and NRTI multidrug resistance-associated mutations were scored as resistant, 27.6% of the samples were misclassified, mainly due to samples falsely scored as susceptible. Therefore, the relative frequencies of other mutations were evaluated. Mutations at codons 44 and 118 were rarely detected in abacavir-susceptible samples but were overrepresented in resistant samples. Site-directed mutagenesis of E44D, V118I, and M184V resulted in low-level resistance for the double mutant 44/184 and the triple mutant. Low-level abacavir resistance was also detected for a viral clone carrying zidovudine mutations only. Additional insertion of M184V into the zidovudine background doubled the resistance, whereas 44/118 did not lead to a further increase. Incorporating combinations of zidovudine mutations and M184V into the scoring system markedly reduced the number of misclassified samples, whereas 44/118 did not improve the prediction. In conclusion, the combination of M184V with zidovudine mutations gives rise to high-level abacavir resistance, which may be clinically relevant. Thus, options for useful sequential combinations of NRTI are limited.

Algorithms↗

Combination of drug level measurement and parasite genotyping data for improved assessment of amodiaquine and sulfadoxine-pyrimethamine efficacies in treating Plasmodium falciparum malaria in Gabonese children.

Many African countries currently use a sulfadoxine-pyrimethamine combination (SP) or amodiaquine (AQ) to treat uncomplicated Plasmodium falciparum malaria. Both drugs represent the last inexpensive alternatives to chloroquine. However, resistant P. falciparum populations are largely reported in Africa, and it is compulsory to know the present situation of resistance. The in vivo World Health Organization standard 28-day test was used to assess the efficacy of AQ and SP to treat uncomplicated falciparum malaria in Gabonese children under 10 years of age. To document treatment failures, molecular genotyping to distinguish therapeutic failures from reinfections and drug dosages were undertaken. A total of 118 and 114 children were given AQ or SP, respectively, and were monitored. SP was more effective than AQ, with 14.0 and 34.7% of therapeutic failures, respectively. Three days after initiation of treatment, the mean level of monodesethylamodiaquine (MdAQ) in plasma was 149 ng/ml in children treated with amodiaquine. In those treated with SP, mean levels of sulfadoxine and pyrimethamine in plasma were 100 microg/ml and 212 ng/ml, respectively. Levels of the three drugs were higher in patients successfully treated with AQ (MdAQ plasma levels) or SP (sulfadoxine and pyrimethamine plasma levels). Blood concentration higher than breakpoints of 135 ng/ml for MdAQ, 100 micro g/ml for sulfadoxine, and 175 ng/ml for pyrimethamine were associated with treatment success (odds ratio: 4.5, 9.8, and 11.8, respectively; all P values were <0.009). Genotyping of merozoite surface proteins 1 and 2 demonstrated a mean of 4.0 genotypes per person before treatment. At reappearance of parasitemia, both recrudescent parasites (represented by common bands in both samples) and newly inoculated parasites (represented by bands that were absent before treatment) were present in the blood of most (51.1%) children. Only 3 (6.4%) therapeutic failures were the result not of treatment inefficacy but of new infection. In areas where levels of drug resistance and complexity of infections are high, drug dosage and parasite genotyping may be of limited interest in improving the precision of drug efficacy measurement. Their use should be weighted according to logistical constraints.

Amodiaquine↗

Choosing haplotype-tagging SNPS based on unphased genotype data using a preliminary sample of unrelated subjects with an example from the Multiethnic Cohort Study.

We describe an approach for picking haplotype-tagging single nucleotide polymorphisms (htSNPs) that is presently being taken in two large nested case-control studies within a multiethnic cohort (MEC), which are engaged in a search for associations between risk of prostate and breast cancer and common genetic variations in candidate genes. Based on a preliminary sample of 70 control subjects chosen at random from each of the 5 ethnic groups in the MEC we estimate haplotype frequencies using a variant of the Excoffier-Slatkin E-M algorithm after genotyping a high density of SNPs selected every 3-5 kb in and surrounding a candidate gene. In order to evaluate the performance of a candidate set of htSNPS (which will be genotyped in the much larger case-control sample) we treat the haplotype frequencies estimate above as known, and carry out a formal calculation of the uncertainty of the number of copies of common haplotypes carried by an individual, summarizing this calculation as a coefficient of determination, R2h. A candidate set of htSNPS of a given size is chosen so as to maximize the minimum value of R2h over the common haplotypes, h.

Aromatase↗

Use of unphased multilocus genotype data in indirect association studies.

It is usually assumed that detection of a disease susceptability gene via marker polymorphisms in linkage disequilibrium with it is facilitated by consideration of marker haplotypes. However, capture of the marker haplotype information requires resolution of gametic phase, and this must usually be inferred statistically. Recently, we questioned the value of the marker haplotype information, and suggested that certain analyses of multivariate marker data, not based on haplotypes explicitly and not requiring resolution of gametic phase, are often more powerful than analyses based on haplotypes. Here, we review this work and assess more carefully the situations in which our conclusions might apply. We also relate these analyses to alternative approaches to haplotype analysis, namely those based on haplotype similarity and those inspired by cladistics.

Genetic Markers↗

Estimating haplotype effects on dichotomous outcome for unphased genotype data using a weighted penalized log-likelihood approach.

OBJECTIVE: To develop a method to estimate haplotype effects on dichotomous outcomes when phase is unknown, that can also estimate reliable effects of rare haplotypes. METHODS: In short, the method uses a logistic regression approach, with weights attached to all possible haplotype combinations of an individual. An EM-algorithm was used: in the E-step the weights are estimated, and the M-step consists of maximizing the joint log-likelihood. When rare haplotypes were present, a penalty function was introduced. We compared four different penalties. To investigate statistical properties of our method, we performed a simulation study for different scenarios. The evaluation criteria are the mean bias of the parameter estimates, the root of the mean squared error, the coverage probability, power, Type I error rate and the false discovery rate. RESULTS: For the unpenalized approach, mean bias was small, coverage probabilities were approximately 95%, power ranged from 15.2 to 44.7% depending on haplotype frequency, and Type I error rate was around 5%. All penalty functions reduced the standard errors of the rare haplotypes, but introduced bias. This trade-off decreased power. CONCLUSION: The unpenalized weighted log-likelihood approach performs well. A penalty function can help to estimate an effect for rare haplotypes.

Algorithms↗

Reducing environmental bias when measuring natural selection.

Crucial to understanding the process of natural selection is characterizing phenotypic selection. Measures of phenotypic selection can be biased by environmental variation among individuals that causes a spurious correlation between a trait and fitness. One solution is analyzing genotypic data, rather than phenotypic data. Genotypic data, however, are difficult to gather, can be gathered from few species, and typically have low statistical power. Environmental correlations may act through traits other than through fitness itself. A path analytic framework, which includes measures of such traits, may reduce environmental bias in estimates of selection coefficients. We tested the efficacy of path analysis to reduce bias by re-analyzing three experiments where both phenotypic and genotypic data were available. All three consisted of plant species (Impatiens capensis, Arabidopsis thaliana, and Raphanus sativus) grown in experimental plots or the greenhouse. We found that selection coefficients estimated by path analysis using phenotypic data were highly correlated with those based on genotypic data with little systematic bias in estimating the strength of selection. Although not a panacea, using path analysis can substantially reduce environmental biases in estimates of selection coefficients. Such confidence in phenotypic selection estimates is critical for progress in the study of natural selection.

Arabidopsis↗

Genetic diversity and differentiation of guanaco populations from Argentina inferred from microsatellite data.

Genotype data from 14 microsatellite markers were used to assess the genetic diversity and differentiation of four guanaco populations from Argentine Patagonia. These animals were recently captured in the wild and maintained in semi-captivity for fibre production. Considerable genetic diversity in these populations was suggested by the finding of a total of 162 alleles, an average mean number of alleles per locus ranging from 6.50 to 8.19, and H(e) values ranging from 0.66 to 0.74. Assessment of population differentiation showed moderate but significant values of F(ST)=0.071 (P=0.000) and R(ST)=0.083 (P=0.000). An amova test showed that the genetic variation among populations was 5.6% while within populations it was 94.4%. A number of 6.6 migrants per generation may support these results. Unambiguous individual assignment to original populations was obtained for the Pilcaniyeu, Las Heras and La Esperanza populations. The erroneous assignment of 18.75% Rio Mayo individuals to the Las Heras population can be explained by the low genetic differentiation found between these two populations. Thirty-nine of 56 loci per population combinations were in Hardy--Weinberg disequilibrium because of guanaco heterozygote deficiency, which may be explained by population subdivision. The high level of genetic diversity of the guanacos analysed here indicates that the Patagonian guanaco constitutes an important genetic resource for conservation or economic utilization programmes.

Analysis of Variance↗

Noninvasive genetic tracking of the endangered Pyrenean brown bear population.

Pyrenean brown bears Ursus arctos are threatened with extinction. Management efforts to preserve this population require a comprehensive knowledge of the number and sex of the remaining individuals and their respective home ranges. This goal has been achieved using a combination of noninvasive genetic sampling of hair and faeces collected in the field and corresponding track size data. Genotypic data were collected at 24 microsatellite loci using a rigorous multiple-tubes approach to avoid genotyping errors associated with low quantities of DNA. Based on field and genetic data, the Pyrenean population was shown to be composed at least of one yearling, three adult males, and one adult female. These data indicate that extinction of the Pyrenean brown bear population is imminent without population augmentation. To preserve the remaining Pyrenean gene pool and increase genetic diversity, we suggest that managers consider population augmentation using only females. This study demonstrates that comprehensive knowledge of endangered small populations of mammals can be obtained using noninvasive genetic sampling.

Animals↗

Two approaches for consolidating results from genome scans of complex traits: selection methods and scan statistics.

This work has two purposes: (i) empirically selecting levels of significance that maximize the fraction of markers close to a gene (hit rate) when performing linkage analyses of simulated data and (ii) evaluating the utility of a previously reported scan statistic on the same data. Genotype data were simulated from a trait model of seven susceptibility genes. For purpose (i), five statistics were evaluated on all marker loci in fifty replicates; two-point lod and heterogeneity lod scores maximized over dominance (mlod, mhlod), a multi-allelic TDT test, an affected sib-pair test (ASP), and a model-free test on all sib-pairs (ALL_SIBS). Within each replicate the fraction of markers (hit rate) significant at specified levels of significance and also (a) within fifty markers of, or (b) on the same chromosome as a major gene was calculated. For purpose (ii), scan statistics of length 15 were calculated for each chromosome and their empirical significance levels estimated on the basis of 500 replicates generated under no linkage. The scan statistic was applied to the mhlod scores from one replicate (Replicate 5). Empirical p-values for the scan statistic were determined by computing mhlod scores on 500 replicates of simulated null data. For purpose (i), significance levels between 0.001 and 0.01 had the greatest hit rate for all five methods and both criteria. For criterion (a) at the 0.001 level of significance, both mlod and mhlod displayed the highest hit rates, approximately 0.4 for each. For criterion (b), all methods but ALL_SIBS and ASP had hit rates ranging between 0.4 and 0.5. For purpose (ii), the scan statistic proved equally or more powerful than the single-locus statistic for two of the seven susceptibility genes while the remaining five genes were not detected.

Chromosome Mapping↗