Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genotype Data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Detecting and genotyping Escherichia coli O157:H7 using multiplexed PCR and nucleic acid microarrays.

Rapid detection and characterization of food borne pathogens such as Escherichia coli O157:H7 is crucial for epidemiological investigations and food safety surveillance. As an alternative to conventional technologies, we examined the sensitivity and specificity of nucleic acid microarrays for detecting and genotyping E. coli O157:H7. The array was composed of oligonucleotide probes (25-30 mer) complementary to four virulence loci (intimin, Shiga-like toxins I and II, and hemolysin A). Target DNA was amplified from whole cells or from purified DNA via single or multiplexed polymerase chain reaction (PCR), and PCR products were hybridized to the array without further modification or purification. The array was 32-fold more sensitive than gel electrophoresis and capable of detecting amplification products from < 1 cell equivalent of genomic DNA (1 fg). Immunomagnetic capture, PCR and a microarray were subsequently used to detect 55 CFU ml(-1) (E. coli O157:H7) from chicken rinsate without the aid of pre-enrichment. Four isolates of E. coli O157:H7 and one isolate of O91:H2, for which genotypic data were available, were unambiguously genotyped with this array. Glass-based microarrays are relatively simple to construct and provide a rapid and sensitive means to detect multiplexed PCR products; the system is amenable to automation.

Animals↗

Comparison of microsatellites and amplified fragment length polymorphism markers for parentage analysis.

This study compares the properties of dominant markers, such as amplified fragment length polymorphisms (AFLPs), with those of codominant multiallelic markers, such as microsatellites, in reconstructing parentage. These two types of markers were used to search for both parents of an individual without prior knowledge of their relationships, by calculating likelihood ratios based on genotypic data, including mistyping. Experimental data on 89 oak trees genotyped for six microsatellite markers and 159 polymorphic AFLP loci were used as a starting point for simulations and tests. Both sets of markers produced high exclusion probabilities, and among dominant markers those with dominant allele frequencies in the range 0.1-0.4 were more informative. Such codominant and dominant markers can be used to construct powerful statistical tests to decide whether a genotyped individual (or two individuals) can be considered as the true parent (or parent pair). Gene flow from outside the study stand (GFO), inferred from parentage analysis with microsatellites, overestimated the true GFO, whereas with AFLPs it was underestimated. As expected, dominant markers are less efficient than codominant markers for achieving this, but can still be used with good confidence, especially when loci are deliberately selected according to their allele frequencies.

DNA, Plant↗

Tuberculosis in the Inuit community of Quebec, Canada.

In low-incidence countries targeting tuberculosis (TB) elimination, TB remains a problem of a few high-risk groups. In Canada, Aboriginals, and particularly the Arctic Inuit communities, have witnessed dramatic decreases in TB during the 1960s to 1970s, but rates remain at least 10 to 20 times higher than the national average. We are describing the results of an integrated traditional and molecular epidemiology study of all culture-positive Mycobacterium tuberculosis cases in the Arctic Inuit communities of Quebec from 1990 until 2000. The demographic characteristics of the 46 TB cases included in the study were most notable for a bimodal age distribution (48% under 25 years). Genotyping analysis using multiple modalities (IS6110 restriction fragment length polymorphism, spoligotype, mycobacterial interspersed repetitive units-variable number tandem repeats) showed that 76% (35/46) of TB cases were clustered (six clusters, median size four cases) and estimated that at least 62.5% of TB cases were due to ongoing transmission. By integrating the epidemiologic and genotyping data, we observed that the genotyping clustering results were concordant with recognized epidemiologic links but most notably identified previously unrecognized intervillage transmission. This study demonstrates significant ongoing transmission in a geographically isolated, low-density population. In a resource-rich country such as Canada, these communities illustrate some of the persistent challenges of TB control and elimination.

Adolescent↗

A transmission disequilibrium test for general pedigrees that is robust to the presence of random genotyping errors and any number of untyped parents.

Two issues regarding the robustness of the original transmission disequilibrium test (TDT) developed by Spielman et al are: (i) missing parental genotype data and (ii) the presence of undetected genotype errors. While extensions of the TDT that are robust to items (i) and (ii) have been developed, there is to date no single TDT statistic that is robust to both for general pedigrees. We present here a likelihood method, the TDT(ae), which is robust to these issues in general pedigrees. The TDT(ae) assumes a more general disease model than the traditional TDT, which assumes a multiplicative inheritance model for genotypic relative risk. Our model is based on Weinberg's work. To assess robustness, we perform simulations. Also, we apply our method to two data sets from actual diseases: psoriasis and sitosterolemia. Maximization under alternative and null hypotheses is performed using Powell's method. Results of our simulations indicate that our method maintains correct type I error rates at the 1, 5, and 10% levels of significance. Furthermore, a Kolmorogov-Smirnoff Goodness of Fit test suggests that the data are drawn from a central chi2 with 2 df, the correct asymptotic null distribution. The psoriasis results suggest two loci as being significantly linked to the disease, even in the presence of genotyping errors and missing data, and the sitosterolemia results show a P-value of 1.5 x 10(-9) for the marker locus nearest to the sitosterolemia disease genes. We have developed software to perform TDT(ae) calculations, which may be accessed from our ftp site.

Computer Simulation↗

Identification and analysis of error types in high-throughput genotyping.

Although it is clear that errors in genotyping data can lead to severe errors in linkage analysis, there is as yet no consensus strategy for identification of genotyping errors. Strategies include comparison of duplicate samples, independent calling of alleles, and Mendelian-inheritance-error checking. This study aimed to develop a better understanding of error types associated with microsatellite genotyping, as a first step toward development of a rational error-detection strategy. Two microsatellite marker sets (a commercial genomewide set and a custom-designed fine-resolution mapping set) were used to generate 118,420 and 22,500 initial genotypes and 10,088 and 8,328 duplicates, respectively. Mendelian-inheritance errors were identified by PedManager software, and concordance was determined for the duplicate samples. Concordance checking identifies only human errors, whereas Mendelian-inheritance-error checking is capable of detection of additional errors, such as mutations and null alleles. Neither strategy is able to detect all errors. Inheritance checking of the commercial marker data identified that the results contained 0.13% human errors and 0.12% other errors (0.25% total error), whereas concordance checking found 0.16% human errors. Similarly, Mendelian-inheritance-error checking of the custom-set data identified 1.37% errors, compared with 2.38% human errors identified by concordance checking. A greater variety of error types were detected by Mendelian-inheritance-error checking than by duplication of samples or by independent reanalysis of gels. These data suggest that Mendelian-inheritance-error checking is a worthwhile strategy for both types of genotyping data, whereas fine-mapping studies benefit more from concordance checking than do studies using commercial marker data. Maximization of error identification increases the likelihood of linkage when complex diseases are analyzed.

Alleles↗

Detection of genotyping errors and pseudo-SNPs via deviations from Hardy-Weinberg equilibrium.

Genotype error can greatly reduce the power of a genetic study. For family data, genotype error can be assessed by examining marker data for non-Mendelian inconsistencies, closely linked markers for double recombination events, and consistency of duplicate genotypes. For case-control data, duplicate samples are genotyped, and controls are tested for deviations from Hardy-Weinberg equilibrium (HWE). Duplicate samples can provide accurate estimates of genotyping error rates, unless systematic genotyping errors have occurred. Although genotyping errors can cause deviations from HWE, these deviations are usually small, and the power to detect them is low except for high rates of genotyping error and/or large sample sizes. An additional problem is that even when deviations from HWE are detected for marker loci, without additional experimentation it is not possible to unequivocally implicate genotyping error as the cause. The power and sample sizes necessary to detect deviations from HWE for single-nucleotide polymorphism (SNP) data are examined for a variety of genotyping error and pseudo-SNP models. For the majority of genotyping models examined, the power is poor to detect deviations from HWE. For example, for 1,000 controls, if an allele with a frequency of 0.1 fails to amplify for 28% of the heterozygous genotypes producing a sample error rate of 0.05, the power is 0.51 to detect a deviation from HWE at an alpha level of 0.05. On the other hand, the detection of deviations from HWE for pseudo-SNPs (paralogous and ectopic sequence variants) for the majority of models examined produces a power of >0.8 for sample sizes as small as 50 individuals.

Gene Frequency↗

Accounting for decay of linkage disequilibrium in haplotype inference and missing-data imputation.

Although many algorithms exist for estimating haplotypes from genotype data, none of them take full account of both the decay of linkage disequilibrium (LD) with distance and the order and spacing of genotyped markers. Here, we describe an algorithm that does take these factors into account, using a flexible model for the decay of LD with distance that can handle both "blocklike" and "nonblocklike" patterns of LD. We compare the accuracy of this approach with a range of other available algorithms in three ways: for reconstruction of randomly paired, molecularly determined male X chromosome haplotypes; for reconstruction of haplotypes obtained from trios in an autosomal region; and for estimation of missing genotypes in 50 autosomal genes that have been completely resequenced in 24 African Americans and 23 individuals of European descent. For the autosomal data sets, our new approach clearly outperforms the best available methods, whereas its accuracy in inferring the X chromosome haplotypes is only slightly superior. For estimation of missing genotypes, our method performed slightly better when the two subsamples were combined than when they were analyzed separately, which illustrates its robustness to population stratification. Our method is implemented in the software package PHASE (v2.1.1), available from the Stephens Lab Web site.

Algorithms↗

Evaluation of the ovine callipyge locus: III. genotypic effects on meat quality traits.

A resource flock of 362 F2 lambs provided phenotypic and genotypic data to estimate effects of callipyge (CLPG) genotypes (NN, NC, CN, and CC) on meat quality traits. The mutant allele is represented as C, the normal allele(s) as N, and the paternal allele of a genotype is given first. Lambs of each genotype born in 1994 and 1995 were serially slaughtered in six groups at 3-wk intervals starting at 23 wk of age. Warner-Bratzler shear force and subjective evaluation of marbling were collected during both years from longissimus. Calpastatin activity was measured on longissimus from the 1994 group, and ELISA quantification of calpastatin protein was obtained from the 1995 group. Significant additive and paternal polar overdominance effects on meat quality traits were detected. This is in contrast to previous research that detected only polar overdominance effects on slaughter and carcass traits in this population. The magnitude of genotypic effects on shear force differed significantly between years; however, additive (P < .01), paternal polar overdominance (P < .001), and maternal dominance (P < .01) effects adjusted for variation in carcass weight were detected within each year. Shear force data adjusted to the mean slaughter age or carcass weight indicated that the means and variances of CN and CC genotypes were greater than values of NC and NN. Shear force values were greatest for CN and were intermediate for CC. The difference in shear force (adjusted for variation in slaughter age) between homozygous genotypes (additive effect) was supported by calpastatin activity data with 2-df F-tests of 3.66 (P < .05) and 11.84 (P < .001) at d 0 and 7 postmortem, respectively. Corresponding values for the paternal polar overdominance effects on calpastatin activity were 53.80 (P < .001) and 87.43 (P < .001). Calpastatin ELISA data (d 0, adjusted for slaughter age) exhibited a paternal polar overdominance effect exclusively with a 2-df F-test of 57.63 (P < .001). Additive and paternal polar overdominance effects on marbling adjusted for slaughter age had F-tests of 6.41 (P < .01) and 93.29 (P < .001), respectively. Consequences of increased longissimus shear force must be addressed if the advantages of CN lambs for dressing percentage and carcass composition are to be realized. Further research is needed to establish whether selection targeted at changing the background genome can mitigate the negative effects of the C allele on meat tenderness.

Animals↗

A new algorithm for haplotype-based association analysis: the Stochastic-EM algorithm.

It is now widely accepted that haplotypic information can be of great interest for investigating the role of a candidate gene in the etiology of complex diseases. In the absence of family data, haplotypes cannot be deduced from genotypes, except for individuals who are homozygous at all loci or heterozygous at only one site. Statistical methodologies are therefore required for inferring haplotypes from genotypic data and testing their association with a phenotype of interest. Two maximum likelihood algorithms are often used in the context of haplotype-based association studies, the Newton-Raphson (NR) and the Expectation-Maximisation (EM) algorithms. In order to circumvent the limitations of both algorithms, including convergence to local minima and saddle points, we here described how a stochastic version of the EM algorithm, referred to as SEM, could be used for testing haplotype-phenotype association. Statistical properties of the SEM algorithm were investigated through a simulation study for a large range of practical situations, including small/large samples and rare/frequent haplotypes, and results were compared to those obtained by use of the standard NR algorithm. Our simulation study indicated that the SEM algorithm provides results similar to those of the NR algorithm, making the SEM algorithm of great interest for haplotype-based association analysis, especially when the number of polymorphisms is quite large.

Algorithms↗

Global landscape of recent inferred Darwinian selection for Homo sapiens.

By using the 1.6 million single-nucleotide polymorphism (SNP) genotype data set from Perlegen Sciences [Hinds, D. A., Stuve, L. L., Nilsen, G. B., Halperin, E., Eskin, E., Ballinger, D. G., Frazer, K. A. & Cox, D. R. (2005) Science 307, 1072-1079], a probabilistic search for the landscape exhibited by positive Darwinian selection was conducted. By sorting each high-frequency allele by homozygosity, we search for the expected decay of adjacent SNP linkage disequilibrium (LD) at recently selected alleles, eliminating the need for inferring haplotype. We designate this approach the LD decay (LDD) test. By these criteria, 1.6% of Perlegen SNPs were found to exhibit the genetic architecture of selection. These results were confirmed on an independently generated data set of 1.0 million SNP genotypes (International Human Haplotype Map Phase I freeze). Simulation studies indicate that the LDD test, at the megabase scale used, effectively distinguishes selection from other causes of extensive LD, such as inversions, population bottlenecks, and admixture. The approximately 1,800 genes identified by the LDD test were clustered according to Gene Ontology (GO) categories. Based on overrepresentation analysis, several predominant biological themes are common in these selected alleles, including host-pathogen interactions, reproduction, DNA metabolism/cell cycle, protein metabolism, and neuronal function.

Alleles↗

Mapping a new genetic locus for X linked retinitis pigmentosa to Xq28.

We have defined a new genetic locus for an X linked form of retinitis pigmentosa (RP) on chromosome Xq28. We examined 15 members of a family in which RP appeared to be transmitted in an X linked manner. Ocular examinations were performed, and fundus photographs and electroretinograms were obtained for selected patients. Blood samples were obtained from all patients and an additional seven family members who were not given examinations. Visual acuity in four affected individuals ranged from 20/40 to 20/80+. Patients described the onset of night blindness and colour vision defects in the second decade of life, with the earliest at 13 years of age. Examined affected individuals had constricted visual fields and retinal findings compatible with RP. Based on full field electroretinography, cone function was more severely reduced than rod function. Female carriers had no ocular signs or symptoms and slightly reduced cone electroretinographic responses. Affected and non-affected family members were genotyped for 20 polymorphic markers on the X-chromosome spaced at 10 cM intervals. Genotyping data were analysed using GeneMapper software. Genotyping and linkage analyses identified significant linkage to markers DXS8061, DXS1073, and DXS1108 with two point LOD scores of 2.06, 2.17, and 2.20, respectively. Haplotype analysis revealed segregation of the disease phenotype with markers at Xq28.

Adolescent↗

The effect of missing data on linkage disequilibrium mapping and haplotype association analysis in the GAW14 simulated datasets.

We used our newly developed linkage disequilibrium (LD) plotting software, JLIN, to plot linkage disequilibrium between pairs of single-nucleotide polymorphisms (SNPs) for three chromosomes of the Genetic Analysis Workshop 14 Aipotu simulated population to assess the effect of missing data on LD calculations. Our haplotype analysis program, SIMHAP, was used to assess the effect of missing data on haplotype-phenotype association. Genotype data was removed at random, at levels of 1%, 5%, and 10%, and the LD calculations and haplotype association results for these levels of missingness were compared to those for the complete dataset. It was concluded that ignoring individuals with missing data substantially affects the number of regions of LD detected which, in turn, could affect tagging SNPs chosen to generate haplotypes.

Chromosome Mapping↗

Identifying Mycobacterium tuberculosis complex strain families using spoligotypes.

We present a novel approach for analysis of Mycobacterium tuberculosis complex (MTC) strain genotyping data. Our work presents a first step in an ongoing project dedicated to the development of decision support tools for tuberculosis (TB) epidemiologists exploiting both genotyping and epidemiological data. We focus on spacer oligonucleotide typing (spoligotyping), a genotyping method based on analysis of a direct repeat (DR) locus. We use mixture models to identify strain families of MTC based on their spoligotyping patterns. Our algorithm, SPOTCLUST, incorporates biological information on spoligotype evolution, without attempting to derive the full phylogeny of MTC. We applied our algorithm to 535 different spoligotype patterns identified among 7166 MTC strains isolated between 1996 and 2004 from New York State TB patients. Two models were employed and validated: a 36-component model based on global spoligotype database SpolDB3, and a randomly initialized model (RIM) containing 48 components. Our analysis both confirmed previously expert-defined families of MTC strains and suggested certain new families. SPOTCLUST, which is available online, can be further improved by incorporating data obtained using additional strain genetic markers and epidemiological information. We demonstrate on New York City (NYC) patient data how the resulting models can potentially form the basis of TB control tools using genotyping.

Adolescent↗

A test of homogeneity of Hardy-Weinberg disequilibrium across strata.

For genotype data being sampled from several strata with different allele frequencies, it is necessary to verify the assumption of homogeneity of Hardy-Weinberg disequilibrium across strata before testing Hardy-Weinberg law across strata. In practice, disequilibrium can be measured via fixation coefficients (ie, ratios of genotypic frequencies) or disequilibrium coefficients (ie, differences of genotypic frequencies). Test for homogeneity of Hardy-Weinberg disequilibrium using data from several populations has been derived according to fixation coefficients. In this article, using the likelihood score theory extended to nuisance parameters, we derive a homogeneity score test for comparing disequilibrium coefficients across several independent strata. Simulation results demonstrate that the homogeneity score test performs satisfactorily in the sense that its empirical size seldom exceeds the pre-chosen nominal level by more than 10% even for small sample sizes. Corresponding power and sample size formulae are provided as well. We illustrate our test with a real glyoxalase genotype data set.

Gene Frequency↗

Effects of differential genotyping error rate on the type I error probability of case-control studies.

OBJECTIVES: It is well known that genotyping error adversely affects the power of genetic case-control association studies but there is little research on its effects on type I error, and none that has addressed possible differences in genotype error rates between cases and controls. METHODS: We used simulations to examine the influence of genotyping error on the type I error probability given by case-control studies. The effect of genotyping error on the magnitude of type I error was explored for a single marker of varying minor allele frequency (MAF), and for haplotypic tests based on two markers with varying MAF and linkage disequilibrium (LD) measure r(2). RESULTS: We show that even with low genotyping error rates (<0.01), systematic differences in the error rate between samples can result in type I error rates substantially above 0.05. The effect was maximal for markers with small MAF, markers in strong LD, and where a common allele is more frequently misclassified as a rare allele than vice versa. The problem was also exacerbated by the use of large samples. CONCLUSIONS: Our results show that small differential genotyping error rates between cases and controls pose significant problems for association analyses. Differential genotyping error rates are particularly likely to arise where genotype data are combined from multiple sites, or where case genotypes are examined against archived reference population cohort genotypes that are being generated in several countries. Although these strategies may be necessary to obtain adequately powered samples, our data show the importance of stringent quality control. Furthermore, associations based on rare haplotypes should be treated with caution.

Alleles↗

Inference on recombination and block structure using unphased data.

In this study compatibility with a tree for unphased genotype data is discussed. If the data are compatible with a tree, the data are consistent with an assumption of no recombination in its evolutionary history. Further, it is said that there is a solution to the perfect phylogeny problem; i.e., for each individual a pair of haplotypes can be defined and the set of all haplotypes can be explained without invoking recombination. A new algorithm to decide whether or not a sample is compatible with a tree is derived. The new algorithm relies on an equivalence relation between sites that mutually determine the phase of each other. (The previous algorithm was based on advanced graph theoretical tools.) The equivalence relation is used to derive the number of solutions to the perfect phylogeny problem. Further, a series of statistics, R ( j ) ( M ), j >or= 2, are defined. These can be used to detect recombination events in the sample's history and to divide the sample into regions that are compatible with a tree. The new statistics are applied to real data from human genes. The results from this application are discussed with reference to recent suggestions that recombination in the human genome is highly heterogeneous.

Biometry↗

Generation and exploration of a dense genetic map in a region of a QTL affecting corpora lutea in a Meishan x Yorkshire cross.

Previously genomic scans revealed quantitative trait loci (QTL) on porcine Chromosome 8 (SSC8) as significantly affecting the number of corpora lutea (CL) in swine. In one study, statistical evidence for the putative QTL was found in the chromosomal region defined by the microsatellites (MS) SW205, SW444, SW206, and SW29. A Yeast Artificial Chromosome library was screened by using the corresponding primers for clones containing these MS by PCR. From five positive YAC clones, 10 additional MS were isolated and mapped to SSC8 with the INRA-University of Minnesota porcine Radiation Hybrid (IMpRH) panel. The genetic map position of the QTL has been refined by addition of these 10 markers. The QTL evaluation included pedigrees of F2-intercross Meishan x Yorkshire design, with phenotypic data of 108 F2 female offspring and genotypic data for 29 MS markers on SSC8. The analysis was performed by using the least squares regression method. The calculated QTL effect for CL obtained by the multilocus least squares method showed a maximum test statistic (F value = 13.98) at position 99 cM between three MS derived from YACs containing SW205 and SW1843 spanning an interval of 7.1 cM. The point-wise (nominal) P-value was 5.21 x 10-6 corresponding to a genome-wide P-value of 0.009. The additive QTL effect explained 17.4% of the phenotypic variance.

Animals↗

An examination of the genotyping error detection function of SIMWALK2.

This investigation was undertaken to assess the sensitivity and specificity of the genotyping error detection function of the computer program SIMWALK2. We chose to examine chromosome 22, which had 7 microsatellite markers, from a single simulated replicate (330 pedigrees with a pattern of missing genotype data similar to the Framingham families). We created genotype errors at five overall frequencies (0.0, 0.025, 0.050, 0.075, and 0.100) and applied SIMWALK2 to each of these five data sets, respectively assuming that the total error rate (specified in the program), was at each of these same five levels. In this data set, up to an assumed error rate of 10%, only 50% of the Mendelian-consistent mistypings were found under any level of true errors. And since as many as 70% of the errors detected were false-positives, blanking suspect genotypes (at any error probability) will result in a reduction of statistical power due to the concomitant blanking of correctly typed alleles. This work supports the conclusion that allowing for genotyping errors within likelihood calculations during statistical analysis may be preferable to choosing an arbitrary cut-off.

Adult Children↗