Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genotype Data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Evaluation of genotype data in clinical risk assessment: methods and application to BRCA1, BRCA2, and N-acetyl transferase-2 genotypes in breast cancer.

Associations of numerous susceptibility genes with disease risk have been reported. However, objective methods have not been developed to evaluate the conditions under which translation of knowledge about susceptibility genotypes may be clinically informative. We describe and apply a statistical approach to evaluate when genotype information may be clinically informative in disease risk assessment. We estimate an interval of cumulative cancer incidences where it may be appropriate to use these genes in disease risk assessment. We also estimate the magnitude of a log odds ratio (H) that measures genotype-disease association. We illustrate this method with three breast cancer susceptibility genotypes: population screening data evaluating the 185delAG mutation at BRCA1 and the 6174delT mutation at BRCA2 in a Ashkenazi Jewish population, and case control data for the slow acetylation genotype at the N-acetyl transferase 2 (NAT2) gene in combination with smoking. Knowledge of the 185delAG mutation in BRCA1 (HdelAG = 3.42; 95% CI: 3.04, 3.79) or the 6174delT mutation in BRCA2 (HdelT = 1.98; 95% CI: 1.16, 2.30) can be clinically informative in distinguishing individuals who are and are not at breast cancer risk in populations with cumulative breast cancer incidences of > or = 4% and > or = 13%, respectively. NAT2 genotypes alone are much less clinically informative in predicting breast cancer risk (HNAT2 = 0.10). However, knowledge of both heavy smoking 20 years ago and NAT2 genotype is a more clinically informative predictor of postmenopausal breast cancer risk with HNAT2 = 2.19, when the cumulative breast cancer incidence in the target population is at least 31%. These results indicate that knowledge of the 185delG mutation-status may be clinically informative even in populations with low cumulative breast cancer incidences, whereas the 6174delT mutation and NAT2 genotypes may only be clinically informative in a population with higher cumulative breast cancer incidence. The proposed approach can be used to objectively evaluate the conditions under which susceptibility genotypes may be applied for risk assessment or genetic screening.

Adult↗

Lithuanian haemophilia A and B registry comprising phenotypic and genotypic data.

Haemophilia represents the most common hereditary severe bleeding disorder in humans. About 100 families with this condition live in Lithuania, one of the Baltic states with a population of 3.7 million. Haemophilia care and genetic counselling are still rendered difficult owing to limited availability of clotting factor concentrate and molecular genetic diagnosis. In the present study, a haemophilia registry, comprising phenotypic and genotypic data of the majority of Lithuanian haemophilia A and B patients, was established. The phenotype includes the degree of severity, factor VIII:C, factor VIII:Ag, factor IX:C, von Willebrand factor and antigen (VWF:RiCoF, vWF:Ag) and inhibitor status. Genotyping of the factor VIII and IX genes was performed using mutation screening methods and direct sequencing. In 61 out of 63 patients with haemophilia A (96.8%) and all eight patients with haemophilia B (100%), the causative mutations could be detected. Nineteen of the factor VIII gene defects and two of the factor IX gene mutations are reported for the first time. Identified mutations allowed direct carrier diagnosis in 83 female relatives revealing 44 carriers, 38 non-carriers and one somatic mosaicism. The information provided by this registry will be helpful for monitoring the treatment of Lithuanian haemophilia patients and also for reliable genetic counselling of the affected families in the future.

Factor IX↗

Haplotype association analysis of human disease traits using genotype data of unrelated individuals.

Haplotype inference has become an important part of human genetic data analysis due to its functional and statistical advantages over the single-locus approach in linkage disequilibrium mapping. Different statistical methods have been proposed for detecting haplotype - disease associations using unphased multi-locus genotype data, ranging from the early approach by the simple gene-counting method to the recent work using the generalized linear model. However, these methods are either confined to case - control design or unable to yield unbiased point and interval estimates of haplotype effects. Based on the popular logistic regression model, we present a new approach for haplotype association analysis of human disease traits. Using haplotype-based parameterization, our model infers the effects of specific haplotypes (point estimation) and constructs confidence interval for the risks of haplotypes (interval estimation). Based on the estimated parameters, the model calculates haplotype frequency conditional on the trait value for both discrete and continuous traits. Moreover, our model provides an overall significance level for the association between the disease trait and a group or all of the haplotypes. Featured by the direct maximization in haplotype estimation, our method also facilitates a computer simulation approach for correcting the significance level of individual haplotype to adjust for multiple testing. We show, by applying the model to an empirical data set, that our method based on the well-known logistic regression model is a useful tool for haplotype association analysis of human disease traits.

Computer Simulation↗

Haplotype reconstruction from genotype data using Imperfect Phylogeny.

UNLABELLED: Critical to the understanding of the genetic basis for complex diseases is the modeling of human variation. Most of this variation can be characterized by single nucleotide polymorphisms (SNPs) which are mutations at a single nucleotide position. To characterize the genetic variation between different people, we must determine an individual's haplotype or which nucleotide base occurs at each position of these common SNPs for each chromosome. In this paper, we present results for a highly accurate method for haplotype resolution from genotype data. Our method leverages a new insight into the underlying structure of haplotypes that shows that SNPs are organized in highly correlated 'blocks'. In a few recent studies, considerable parts of the human genome were partitioned into blocks, such that the majority of the sequenced genotypes have one of about four common haplotypes in each block. Our method partitions the SNPs into blocks, and for each block, we predict the common haplotypes and each individual's haplotype. We evaluate our method over biological data. Our method predicts the common haplotypes perfectly and has a very low error rate (<2% over the data) when taking into account the predictions for the uncommon haplotypes. Our method is extremely efficient compared with previous methods such as PHASE and HAPLOTYPER. Its efficiency allows us to find the block partition of the haplotypes, to cope with missing data and to work with large datasets. AVAILABILITY: The algorithm is available via a Web server at http://www.calit2.net/compbio/hap/

Algorithms↗

Accuracy of haplotype frequency estimation for biallelic loci, via the expectation-maximization algorithm for unphased diploid genotype data.

Haplotype analyses have become increasingly common in genetic studies of human disease because of their ability to identify unique chromosomal segments likely to harbor disease-predisposing genes. The study of haplotypes is also used to investigate many population processes, such as migration and immigration rates, linkage-disequilibrium strength, and the relatedness of populations. Unfortunately, many haplotype-analysis methods require phase information that can be difficult to obtain from samples of nonhaploid species. There are, however, strategies for estimating haplotype frequencies from unphased diploid genotype data collected on a sample of individuals that make use of the expectation-maximization (EM) algorithm to overcome the missing phase information. The accuracy of such strategies, compared with other phase-determination methods, must be assessed before their use can be advocated. In this study, we consider and explore sources of error between EM-derived haplotype frequency estimates and their population parameters, noting that much of this error is due to sampling error, which is inherent in all studies, even when phase can be determined. In light of this, we focus on the additional error between haplotype frequencies within a sample data set and EM-derived haplotype frequency estimates incurred by the estimation procedure. We assess the accuracy of haplotype frequency estimation as a function of a number of factors, including sample size, number of loci studied, allele frequencies, and locus-specific allelic departures from Hardy-Weinberg and linkage equilibrium. We point out the relative impacts of sampling error and estimation error, calling attention to the pronounced accuracy of EM estimates once sampling error has been accounted for. We also suggest that many factors that may influence accuracy can be assessed empirically within a data set-a fact that can be used to create "diagnostics" that a user can turn to for assessing potential inaccuracies in estimation.

Algorithms↗

A method for distinguishing consanguinity and population substructure using multilocus genotype data.

We use the patterns of homozygosity at multiple loci to distinguish between excess homozygosity caused by consanguineous mating and that due to undetected population subdivision (the Wahlund effect). Clarification of the underlying causes of excess homozygosity is of practical importance in explaining the occurrence of recessive genetic disorders and in forensic match probability calculations. We calculated a likelihood surface for two parameters: C, the proportion of the population practicing consanguinity, and theta, the genetic correlation due population subdivision. To illustrate the method, we applied it to multilocus genotypic data of two U.K. Asian populations, one practicing a high frequency of cousin marriage, and another in which caste endogamy was suspected. The method was able to successfully distinguish the different patterns of relatedness. The method also returned accurate estimates of C and theta using simulated data sets. We show how our method can be extended to allow for degrees of inbreeding closer than cousin unions, including selfing. With closer inbreeding, the relatedness of recent ancestors beyond the parents becomes an issue.

Algorithms↗

Dissecting trait heterogeneity: a comparison of three clustering methods applied to genotypic data.

BACKGROUND: Trait heterogeneity, which exists when a trait has been defined with insufficient specificity such that it is actually two or more distinct traits, has been implicated as a confounding factor in traditional statistical genetics of complex human disease. In the absence of detailed phenotypic data collected consistently in combination with genetic data, unsupervised computational methodologies offer the potential for discovering underlying trait heterogeneity. The performance of three such methods--Bayesian Classification, Hypergraph-Based Clustering, and Fuzzy k-Modes Clustering--appropriate for categorical data were compared. Also tested was the ability of these methods to detect trait heterogeneity in the presence of locus heterogeneity and/or gene-gene interaction, which are two other complicating factors in discovering genetic models of complex human disease. To determine the efficacy of applying the Bayesian Classification method to real data, the reliability of its internal clustering metrics at finding good clusterings was evaluated using permutation testing. RESULTS: Bayesian Classification outperformed the other two methods, with the exception that the Fuzzy k-Modes Clustering performed best on the most complex genetic model. Bayesian Classification achieved excellent recovery for 75% of the datasets simulated under the simplest genetic model, while it achieved moderate recovery for 56% of datasets with a sample size of 500 or more (across all simulated models) and for 86% of datasets with 10 or fewer nonfunctional loci (across all simulated models). Neither Hypergraph Clustering nor Fuzzy k-Modes Clustering achieved good or excellent cluster recovery for a majority of datasets even under a restricted set of conditions. When using the average log of class strength as the internal clustering metric, the false positive rate was controlled very well, at three percent or less for all three significance levels (0.01, 0.05, 0.10), and the false negative rate was acceptably low (18 percent) for the least stringent significance level of 0.10. CONCLUSION: Bayesian Classification shows promise as an unsupervised computational method for dissecting trait heterogeneity in genotypic data. Its control of false positive and false negative rates lends confidence to the validity of its results. Further investigation of how different parameter settings may improve the performance of Bayesian Classification, especially under more complex genetic models, is ongoing.

Algorithms↗

Simple tests to detect errors in high-throughput genotype data in the molecular laboratory.

With the advent of high-density DNA marker data sets for the mouse and other model systems, 100 or more genotype are routinely generated from large groups of mice. Issues of the accuracy and reliability of the genotyping are extremely important but often not addressed until genetic analysis is conducted. Simple tests that rely on the robust predictions arising from Mendelian genetics can be made quickly in the molecular laboratory as the data are generated, and require only a spreadsheet program. In this report, genotype data from 392 mice tested at 96 marker sites were analyzed for errors that are typical when handling large volumes of data generated in a repetitive process. The testing consisted of: (1) repeating the genotyping of approximately 1% of the samples; (2) examining the deviation from the expected segregation ratio ( 1:2:1 ) on a marker-by-marker basis; and (3) testing the correlation of the genotype at one marker with that at neighboring genetic markers on a chromosome. These three steps allowed analysis at the level of the microtiter plate, where errors are most likely to occur. A set of 96 dinucleotide repeat markers that are polymorphic between the C57BL/6J and DBA/2J mouse strains and can be multiplexed is reported for use in other genotyping projects.

Algorithms↗

Automatic scoring and quality assessment using accuracy bounds for FP-TDI SNP genotyping data.

BACKGROUND: Human diversity, namely single nucleotide polymorphisms (SNPs), is becoming a focus of biomedical research. Despite the binary nature of SNP determination, the majority of genotyping assay data need a critical evaluation for genotype calling. We applied statistical models to improve the automated analysis of 2-dimensional SNP data. METHODS: We derived several quantities in the framework of Gaussian mixture models that provide figures of merit to objectively measure the data quality. The accuracy of individual observations is scored as the probability of belonging to a certain genotype cluster, while the assay quality is measured by the overlap between the genotype clusters. RESULTS: The approach was extensively tested with a dataset of 438 nonredundant SNP assays comprising >150,000 datapoints. The performance of our automatic scoring method was compared with manual assignments. The agreement for the overall assay quality is remarkably good, and individual observations were scored differently by man and machine in 2.6% of cases, when applying stringent probability threshold values. CONCLUSION: Our definition of bounds for the accuracy for complete assays in terms of misclassification probabilities goes beyond other proposed analysis methods. We expect the scoring method to minimise human intervention and provide a more objective error estimate in genotype calling.

Algorithms↗

Using approximate Bayesian computation to estimate tuberculosis transmission parameters from genotype data.

Tuberculosis can be studied at the population level by genotyping strains of Mycobacterium tuberculosis isolated from patients. We use an approximate Bayesian computational method in combination with a stochastic model of tuberculosis transmission and mutation of a molecular marker to estimate the net transmission rate, the doubling time, and the reproductive value of the pathogen. This method is applied to a published data set from San Francisco of tuberculosis genotypes based on the marker IS6110. The mutation rate of this marker has previously been studied, and we use those estimates to form a prior distribution of mutation rates in the inference procedure. The posterior point estimates of the key parameters of interest for these data are as follows: net transmission rate, 0.69/year [95% credibility interval (C.I.) 0.38, 1.08]; doubling time, 1.08 years (95% C.I. 0.64, 1.82); and reproductive value 3.4 (95% C.I. 1.4, 79.7). These figures suggest a rapidly spreading epidemic, consistent with observations of the resurgence of tuberculosis in the United States in the 1980s and 1990s.

Algorithms↗

Computing the minimum recombinant haplotype configuration from incomplete genotype data on a pedigree by integer linear programming.

We study the problem of reconstructing haplotype configurations from genotypes on pedigree data with missing alleles under the Mendelian law of inheritance and the minimum-recombination principle, which is important for the construction of haplotype maps and genetic linkage/association analyses. Our previous results show that the problem of finding a minimum-recombinant haplotype configuration (MRHC) is in general NP-hard. This paper presents an effective integer linear programming (ILP) formulation of the MRHC problem with missing data and a branch-and-bound strategy that utilizes a partial order relationship and some other special relationships among variables to decide the branching order. Nontrivial lower and upper bounds on the optimal number of recombinants are introduced at each branching node to effectively prune the search tree. When multiple solutions exist, a best haplotype configuration is selected based on a maximum likelihood approach. The paper also shows for the first time how to incorporate marker interval distance into a rule-based haplotyping algorithm. Our results on simulated data show that the algorithm could recover haplotypes with 50 loci from a pedigree of size 29 in seconds on a Pentium IV computer. Its accuracy is more than 99.8% for data with no missing alleles and 98.3% for data with 20% missing alleles in terms of correctly recovered phase information at each marker locus. A comparison with a statistical approach SimWalk2 on simulated data shows that the ILP algorithm runs much faster than SimWalk2 and reports better or comparable haplotypes on average than the first and second runs of SimWalk2. As an application of the algorithm to real data, we present some test results on reconstructing haplotypes from a genome-scale SNP dataset consisting of 12 pedigrees that have 0.8% to 14.5% missing alleles.

Algorithms↗

Most powerful permutation invariant tests for relatedness hypotheses using genotypic data.

The problem of inferring kinship structure among a sample of individuals using genetic markers is considered with the objective of developing hypothesis tests for genetic relatedness with nearly optimal properties. The class of tests considered are those that are constrained to be permutation invariant, which in this context defines tests whose properties do not depend on the labeling of the individuals. This is appropriate when all individuals are to be treated identically from a statistical point of view. The approach taken is to derive tests that are probably most powerful for a permutation invariant alternative hypothesis that is, in some sense, close to a null hypothesis of mutual independence. This is analagous to the locally most powerful test commonly used in parametric inference. Although the resulting test statistic is a U-statistic, normal approximation theory is found to be inapplicable because of high skewness. As an alternative it is found that a conditional procedure based on the most powerful test statistic can calculate accurate significance levels without much loss in power. Examples are given in which this type of test proves to be more powerful than a number of alternatives considered in the literature, including Queller and Goodknight's (1989) estimate of genetic relatedness, the average number of shared alleles (Blouin, 1996), and the number of feasible sibling triples (Almudevar and Field, 1999).

Alleles↗

Cleaning genotype data.

The identification of genes contributing to variation in complex phenotypes requires genetic data of high fidelity. Thus, the identification of pedigree and genotyping errors is a crucial prerequisite to the analysis of data from a genome scan for disease genes. The problem has been given little attention in most gene hunting papers; the focus has often been on eliminating mendelian inconsistencies in order that the analysis may proceed, rather than on achieving the best possible data. Though a number of computer programs are available to assist in the identification of genotyping and pedigree errors, the process is still not completely automated. While the Collaborative Study on the Genetics of Alcoholism (COGA) data set for GAW11 is completely compatible with Mendel's rules, there are still some errors present. We inspected the COGA data for the presence of additional errors, and identified five possible pedigree errors.

Alcoholism↗

Streamlined analysis of pooled genotype data in SNP-based association studies.

Several groups have developed methods for estimating allele frequencies in DNA pools as a fast and cheap way for detecting allelic association between genetic markers and disease. To obtain accurate estimates of allele frequencies, a correction factor k for the degree to which measurement of allele-specific products is biased is generally applied. Factor k is usually obtained as the ratio of the two allele-specific signals in samples from heterozygous individuals, a step that can significantly impair throughput and increase cost. We have systematically investigated the properties of k through the use of empirical and simulated data. We show that for the dye terminator primer extension genotyping method we have applied, the correction factor k is substantially influenced by the dye terminators incorporated, but also by the terminal 3' base of the extension primer. We also show that the variation in k is large enough to result in unacceptable error rates if association studies are conducted without regard to k. We show that the impact of ignoring k can be neutralized by applying a correction factor k(max) that can be easily derived, but this at the potential cost of an increase in type I error. Finally, based upon observed distributions for k we derive a method allowing the estimation of the probability pooled data reflects significant differences in the allele frequencies between the subjects comprising the pools. By controlling the error rates in the absence of knowledge of the appropriate SNP-specific correction factors, each approach enhances the performance of DNA pooling, while considerably streamlining the method by reducing time and cost.

Alleles↗

Statistical power of an exact test of Hardy-Weinberg proportions of genotypic data at a multiallelic locus.

A computer algorithm for numerical evaluation of the statistical power of an exact test of Hardy-Weinberg genotypic proportions (HWP), developed here, indicates that the power is dependent on the number of segregating alleles as well as allele frequencies. While low levels of departure from the null hypothesis are difficult to detect from single-locus data, should such deviation be due to population substructuring, multiple loci, at each of which the number of segregating alleles is large (as seen with hypervariable loci), may easily detect even low levels of departure from HWP. Undetected small levels of departure may still provide conservative estimates of genotype frequencies from allele frequency data, following the current practice in forensic genetics.

Algorithms↗

Determination of probability distribution of diplotype configuration (diplotype distribution) for each subject from genotypic data using the EM algorithm.

Haplotype analysis is important for mapping traits. Recently, methods for estimating haplotype frequencies from genotypes of unrelated individuals based on the expectation-maximization (EM) algorithm have been developed. Our program estimates haplotype frequencies in the population and determines the posterior probability distribution of diplotype configuration (diplotype distribution) for each subject based on the estimated haplotype frequencies. Samples from three ethnic groups for the smoothelin gene (SMTN) and those from three Japanese groups for serum amyloid A genes (SAA@) were analyzed. The estimated diplotype distribution for each individual was concentrated, in most cases, in a single diplotype configuration. The diplotype configuration thus determined was the same as that determined in in vitro experiments, with one exception. Thus, the diplotype configurations determined using the estimated haplotype frequencies from unrelated individuals are reliable. Using this method, the risk of a subject developing a phenotype may be estimated from the diplotype distribution when the phenotype is associated with diplotype configurations.

Algorithms↗