Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genotype Data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Single region linkage analyses of asthma: description of data sets.

Linkage (genotypic) data from the 5q31-33 candidate region for asthma were contributed to Genetic Analysis Workshop 12 by members of the International Consortium on Asthma Genetics (COAG). Data came from five independent studies sampled from five countries. Genotypic data for a total of 26 markers were available, although the number of markers typed in each data set varied. Phenotypic and genotypic data was available from a total of 569 families and 3,175 subjects. The phenotypic data available varied among the studies; however information regarding physician-diagnosed asthma and total serum IgE levels was available in all five studies. This paper describes the ascertainment, data collection methods, phenotypic data, and genotypic data available for the single linkage region analyses undertaken for Genetic Analysis Workshop 12.

Adolescent↗

Comparison of methods for analysis of selective genotyping survival data.

Survival traits and selective genotyping datasets are typically not normally distributed, thus common models used to identify QTL may not be statistically appropriate for their analysis. The objective of the present study was to compare models for identification of QTL associated with survival traits, in particular when combined with selective genotyping. Data were simulated to model the survival distribution of a population of chickens challenged with Marek disease virus. Cox proportional hazards (CPH), linear regression (LR), and Weibull models were compared for their appropriateness to analyze the data, ability to identify associations of marker alleles with survival, and estimation of effects when all individuals were genotyped (full genotyping) and when selective genotyping was used. Little difference in power was found between the CPH and the LR model for low censoring cases for both full and selective genotyping. The simulated data were not transformed to follow a Weibull distribution and, as a result, the Weibull model generally resulted in less power than the other two models and overestimated effects. Effect estimates from LR and CPH were unbiased when all individuals were genotyped, but overestimated when selective genotyping was used. Thus, LR is preferred for analyzing survival data when the amount of censoring is low because of ease of implementation and interpretation. Including phenotypic data of non-genotyped individuals in selective genotyping analysis increased power, but resulted in LR having an inflated false positive rate, and therefore the CPH model is preferred for this scenario, although transformation of the data may also make the Weibull model appropriate for this case. The results from the research presented herein are directly applicable to interval mapping analyses.

Genotype↗

Oxford genome screen for asthma-associated traits.

A genome screen for linkage of quantitative traits underlying asthma has been carried out previously by our group on 80 families sub-selected for discordant phenotypes from a general population sample. The families contained a total of 203 offspring forming 172 sib-pairs. Genotypic data for at a total of 296 markers were available. This paper describes the ascertainment, phenotypic data, and genotypic data made available for Genetic Analysis Workshop 12.

Asthma↗

Inferences about linkage disequilibrium.

Existing theory for inferences about linkage disequilibrium is restricted to a measure defined on gametic frequencies. Unless gametic frequencies are directly observable, they are inferred from genotypic frequencies under the assumption of random union of gametes. Primary emphasis in this paper is given to genotypic data, and disequilibrium coefficients are defined for all subsets of two or more of the four genes, two at each of two loci, carried by an individual. Linkage disequilibrium coefficients are defined for genes within and between gametes, and methods of estimating and testing these coefficients are given for gametic data. For genotypic data, when coupling and repulsion double heterozygotes cannot be distinguished. Burrows' composite measure of linkage disequilibrium is discussed. In particular, the estimate for this measure and hypothesis tests based on it are compared to the usual maximum likelihood estimate of gametic linkage disequilibrium, and corresponding likelihood ratio or contingency chi-square tests. General use of the composite measure, whether or not random union of gametes is an appropriate assumption, is recommended. Attention is given to small samples, where the non-normality of gene frequencies will have greatest effect on methods of inference based on normal theory. Even tools such as Fisher's z-transformation for the correlation of gene frequencies are found to perform quite satisfactorily.

Gene Frequency↗

Whole genome association mapping by incompatibilities and local perfect phylogenies.

BACKGROUND: With current technology, vast amounts of data can be cheaply and efficiently produced in association studies, and to prevent data analysis to become the bottleneck of studies, fast and efficient analysis methods that scale to such data set sizes must be developed. RESULTS: We present a fast method for accurate localisation of disease causing variants in high density case-control association mapping experiments with large numbers of cases and controls. The method searches for significant clustering of case chromosomes in the "perfect" phylogenetic tree defined by the largest region around each marker that is compatible with a single phylogenetic tree. This perfect phylogenetic tree is treated as a decision tree for determining disease status, and scored by its accuracy as a decision tree. The rationale for this is that the perfect phylogeny near a disease affecting mutation should provide more information about the affected/unaffected classification than random trees. If regions of compatibility contain few markers, due to e.g. large marker spacing, the algorithm can allow the inclusion of incompatibility markers in order to enlarge the regions prior to estimating their phylogeny. Haplotype data and phased genotype data can be analysed. The power and efficiency of the method is investigated on 1) simulated genotype data under different models of disease determination 2) artificial data sets created from the HapMap ressource, and 3) data sets used for testing of other methods in order to compare with these. Our method has the same accuracy as single marker association (SMA) in the simplest case of a single disease causing mutation and a constant recombination rate. However, when it comes to more complex scenarios of mutation heterogeneity and more complex haplotype structure such as found in the HapMap data our method outperforms SMA as well as other fast, data mining approaches such as HapMiner and Haplotype Pattern Mining (HPM) despite being significantly faster. For unphased genotype data, an initial step of estimating the phase only slightly decreases the power of the method. The method was also found to accurately localise the known susceptibility variants in an empirical data set--the DeltaF508 mutation for cystic fibrosis--where the susceptibility variant is already known--and to find significant signals for association between the CYP2D6 gene and poor drug metabolism, although for this dataset the highest association score is about 60 kb from the CYP2D6 gene. CONCLUSION: Our method has been implemented in the Blossoc (BLOck aSSOCiation) software. Using Blossoc, genome wide chip-based surveys of 3 million SNPs in 1000 cases and 1000 controls can be analysed in less than two CPU hours.

Chromosome Mapping↗

Genetic variation in putative regulatory loci controlling gene expression in breast cancer.

Candidate single-nucleotide polymorphisms (SNPs) were analyzed for associations to an unselected whole genome pool of tumor mRNA transcripts in 50 unrelated patients with breast cancer. SNPs were selected from 203 candidate genes of the reactive oxygen species pathway. We describe a general statistical framework for the simultaneous analysis of gene expression data and SNP genotype data measured for the same cohort, which revealed significant associations between subsets of SNPs and transcripts, shedding light on the underlying biology. We identified SNPs in EGF, IL1A, MAPK8, XPC, SOD2, and ALOX12 that are associated with the expression patterns of a significant number of transcripts, indicating the presence of regulatory SNPs in these genes. SNPs were found to act in trans in a total of 115 genes. SNPs in 43 of these 115 genes were found to act both in cis and in trans. Finally, subsets of SNPs that share significantly many common associations with a set of transcripts (biclusters) were identified. The subsets of transcripts that are significantly associated with the same set of SNPs or to a single SNP were shown to be functionally coherent in Gene Ontology and pathway analyses and coexpressed in other independent data sets, suggesting that many of the observed associations are within the same functional pathways. To our knowledge, this article is the first study to correlate SNP genotype data in the germ line with somatic gene expression data in breast tumors. It provides the statistical framework for further genotype expression correlation studies in cancer data sets.

Breast Neoplasms↗

Recombination within the HLA-D region. Correlation of molecular genotyping with functional data.

Molecular genotyping of the HLA-D/DR region in a family correlated with serologic and cellular typing data. It was further possible to predict a subtle difference in SB region-related functions from such molecular studies. A family that included an individual who inherited an HLA haplotype with a paternal recombination between HLA-B and the HLA-D/DR region was identified by classic HLA typing techniques. Segregation of HLA-D/DR region genes in this family was studied by Southern blot analysis using cDNA probes for DR alpha, DR beta, DC alpha, DC beta, and SB beta. Restriction enzyme fragment polymorphisms observed for every gene tested were in concordance with assigned HLA haplotypes (including the individual known to have inherited a paternal recombinant haplotype) with one exception: two HLA identical siblings were observed to have different SB beta restriction fragment patterns. Further testing revealed that one individual inherited a maternal HLA haplotype recombinant between the HLA-D/DR region and SB beta. Although both maternal SB alleles typed as SB4, allelic differences could be detected cellularly by primed lymphocytes and by the differential expression of a class II cell surface antigen using monoclonal antibody. Therefore, predicted and nonpredicted recombinant haplotypes were detected in a family by molecular genotyping.

DNA Restriction Enzymes↗

Quantifying the amount of missing information in genetic association studies.

Many genetic analyses are done with incomplete information; for example, unknown phase in haplotype-based association studies. Measures of the amount of available information can be used for efficient planning of studies and/or analyses. In particular, the linkage disequilibrium (LD) between two sets of markers can be interpreted as the amount of information one set of markers contains for testing allele frequency differences in the second set, and measuring LD can be viewed as quantifying information in a missing data problem. We introduce a framework for measuring the association between two sets of variables; for example, genotype data for two distinct groups of markers, or haplotype and genotype data for a given set of polymorphisms. The goal is to quantify how much information is in one data set, e.g. genotype data for a set of SNPs, for estimating parameters that are functions of frequencies in the second data set, e.g. haplotype frequencies, relative to the ideal case of actually observing the complete data, e.g. haplotypes. In the case of genotype data on two mutually exclusive sets of markers, the measure determines the amount of multi-locus LD, and is equal to the classical measure r(2), if the sets consist each of one bi-allelic marker. In general, the measures are interpreted as the asymptotic ratio of sample sizes necessary to achieve the same power in case-control testing. The focus of this paper is on case-control allele/haplotype tests, but the framework can be extended easily to other settings like regressing quantitative traits on allele/haplotype counts, or tests on genotypes or diplotypes. We highlight applications of the approach, including tools for navigating the HapMap database [The International HapMap Consortium, 2003], and genotyping strategies for positional cloning studies.

Alleles↗

An optimal algorithm for perfect phylogeny haplotyping.

Inferring haplotype data from genotype data is a crucial step in linking SNPs to human diseases. Given n genotypes over m SNP sites, the haplotype inference (HI) problem deals with finding a set of haplotypes so that each given genotype can be formed by a combining a pair of haplotypes from the set. The perfect phylogeny haplotyping (PPH) problem is one of the many computational approaches to the HI problem. Though it was conjectured that the complexity of the PPH problem was O(nm), the complexity of all the solutions presented until recently was O(nm (2)). In this paper, we make complete use of the column-ordering that was presented earlier and show that there must be some interdependencies among the pairwise relationships between SNP sites in order for the given genotypes to allow a perfect phylogeny. Based on these interdependencies, we introduce the FlexTree (flexible tree) data structure that represents all the pairwise relationships in O(m) space. The FlexTree data structure provides a compact representation of all the perfect phylogenies for the given set of genotypes. We also introduce an ordering of the genotypes that allows the genotypes to be added to the FlexTree sequentially. The column ordering, the FlexTree data structure, and the row ordering we introduce make the O(nm) OPPH algorithm possible. We present some results on simulated data which demonstrate that the OPPH algorithm performs quiet impressively when compared to the previous algorithms. The OPPH algorithm is one of the first O(nm) algorithms presented for the PPH problem.

Algorithms↗

A haplotype-based test of association using data from cohort and nested case-control epidemiologic studies.

Haplotype-based risk models can lead to powerful methods for detecting the association of a disease with a genomic region of interest. In population-based studies of unrelated individuals, however, the haplotype status of some subjects may not be discernible without ambiguity from available locus-specific genotype data. A score test for detecting haplotype-based association using genotype data has been developed in the context of generalized linear models for analysis of data from cross-sectional and retrospective studies. In this article, we develop a test for association using genotype data from cohort and nested case-control studies where subjects are prospectively followed until disease incidence or censoring (end of follow-up) occurs. Assuming a proportional hazard model for the haplotype effects, we derive an induced hazard function of the disease given the genotype data, and hence propose a test statistic based on the associated partial likelihood. The proposed test procedure can account for differential follow-up of subjects, can adjust for possibly time-dependent environmental co-factors and can make efficient use of valuable age-at-onset information that is available on cases. We provide an algorithm for computing the test statistic using readily available statistical software. Utilizing simulated data in the context of two genomic regions GPX1 and GPX3, we evaluate the validity of the proposed test for small sample sizes and study its power in the presence and absence of missing genotype data.

Algorithms↗

Estimating the correlation of pairwise relatedness along chromosomes.

The 'spatial' pattern of the correlation of pairwise relatedness among loci within a chromosome is an important aspect for an insight into genomic evolution in natural populations. In this article, a statistical genetic method is presented for estimating the correlation of pairwise relatedness among linked loci. The probabilities of identity-in-state (IIS) are related to the probabilities of identity-by-descent (IBS) for the two- and three-loci cases. By decomposing the joint probabilities of two- or three-loci IBD, the probability of pairwise relatedness at a single locus and its correlation among linked loci can be simultaneously estimated. To provide effective statistical methods for estimation, weighted least square (LS) and maximum likelihood (ML) methods are evaluated through extensive Monte Carlo simulations. Results show that the ML method gives a better performance than the weighted LS method with haploid genotypic data. However, there are no significant differences between the two methods when two- or three-loci diploid genotypic data are employed. Compared with the optimal size for haploid genotypic data, a smaller optimal sample size is predicted with diploid genotypic data.

Animals↗

The Cys allele of the DRD2 Ser311Cys polymorphism has a dominant effect on risk for schizophrenia: evidence from fixed- and random-effects meta-analyses.

Previously we derived independent estimates of the effect of the dopamine D2 receptor (DRD2) Ser311Cys polymorphism on risk for schizophrenia using fixed- and random-effects meta-analyses. Both analyses identified a significant association between the Cys allele and schizophrenia, but neither included all available data. Furthermore, genotype data were not evaluated in either analysis, thus precluding any determination of the mode of inheritance. The present study was conducted to resolve discrepancies between the existing meta-analyses, and provide more comprehensive and accurate estimates of the nature and magnitude of the influence of the Ser311Cys polymorphism on risk for schizophrenia. All discrepancies between the two sets of previously meta-analyzed studies were identified and resolved to the mutual satisfaction of the authors, and the final dataset was analyzed independently by fixed- and random-effects meta-analyses. A total of 27 samples, comprising 3,707 schizophrenia patients and 5,363 control subjects, were included in the analyses of allelic association, while smaller numbers of studies and subjects were included in each of the genotypic association analyses. A significant effect of the Cys allele was observed under both fixed-effects (odds ratio [OR] = 1.4; P = 0.002) and random-effects (OR = 1.4; P = 0.007) models. Cys/Ser heterozygotes were at elevated risk for schizophrenia when compared to Ser/Ser homozygotes (fixed- and random-effects OR = 1.4, p(s) or= 0.948). There was no evidence of heterogeneity, excessive influence of any single study, or publication bias in any of the analyses, suggesting that the effect of this DRD2 polymorphism on schizophrenia risk is reliable and uniform across populations, and our estimates of its magnitude are robust and accurate.

Alleles↗

HapBlock: haplotype block partitioning and tag SNP selection software using a set of dynamic programming algorithms.

UNLABELLED: Recent studies have revealed that linkage disequilibrium (LD) patterns vary across the human genome with some regions of high LD interspersed with regions of low LD. Such LD patterns make it possible to select a set of single nucleotide polymorphism (SNPs; tag SNPs) for genome-wide association studies. We have developed a suite of computer programs to analyze the block-like LD patterns and to select the corresponding tag SNPs. Compared to other programs for haplotype block partitioning and tag SNP selection, our program has several notable features. First, the dynamic programming algorithms implemented are guaranteed to find the block partition with minimum number of tag SNPs for the given criteria of blocks and tag SNPs. Second, both haplotype data and genotype data from unrelated individuals and/or from general pedigrees can be analyzed. Third, several existing measures/criteria for haplotype block partitioning and tag SNP selection have been implemented in the program. Finally, the programs provide flexibility to include specific SNPs (e.g. non-synonymous SNPs) as tag SNPs. AVAILABILITY: The HapBlock program and its supplemental documents can be downloaded from the website http://www.cmb.usc.edu/~msms/HapBlock.

Algorithms↗

Optimizing genetic ancestry adjustment in DNA methylation studies: a comparative analysis of approaches.

BACKGROUND: Genetic ancestry is an important factor to account for in DNA methylation studies because genetic variation influences DNA methylation patterns. One approach uses principal components (PCs) calculated from CpG sites that overlap with common SNPs to adjust for ancestry when genotyping data is not available. However, this method does not remove technical and biological variations, such as sex and age, prior to calculating the PCs. The first PC is therefore often associated with factors other than ancestry. METHODS: We developed and adapted the adapted EpiAnceR+ approach, which includes (1) residualizing the CpG data overlapping with common SNPs for control probe PCs, sex, age, and cell type proportions to remove the effects of technical and biological factors, and (2) integrating the residualized data with genotype calls from the SNP probes (commonly referred to as rs probes) present on the arrays, before calculating PCs and evaluated the clustering ability and relationship to genetic ancestry. RESULTS: The PCs generated by EpiAnceR+ led to improved clustering for repeated samples from the same individual and stronger associations with genetic ancestry groups predicted from genotype information compared to the original approach. EpiAnceR+ also outperformed the use of DNA methylation PCs or surrogate variables for ancestry adjustment. CONCLUSIONS: We show that the EpiAnceR+ approach improves the adjustment for genetic ancestry in DNA methylation studies. EpiAnceR+ can be integrated into existing R pipelines for commercial methylation arrays, such as 450 K, EPIC v1, and EPIC v2. The code is available on GitHub ( https://github.com/KiraHoeffler/EpiAnceR ).

DNA Methylation↗

Somatic allele loss in genetic linkage analysis of cancer.

The ability to detect or reject genetic linkage in studies of human cancer is often diminished because multiple affected relatives in a pedigree are unavailable for analysis. The observation of somatic allele loss in tumors can provide knowledge about gametic phase. Therefore, consideration of tumor genotype data could be used to obtain knowledge about gametic phase ordinarily gained from a larger sample of individuals in cancer families. The objective of the present study is to describe a method for improving the power to detect or reject genetic linkage by using knowledge about somatic genetic changes in tumor tissue. A modification to the lod score method of linkage analysis is proposed in which knowledge of gametic phase in the linkage likelihood is inferred from observations of loss of constitutional heterozygosity (LoH) in tumor tissue. This methodology was evaluated using a double backcross nuclear family with a pair of offspring. The expected lod score improved substantially when tumor genotype data were included in the analysis. For example, when the haplotype remaining in tumor tissue was identical to the inherited haplotype in constitutional tissue 99% of the time, linkage analyses without tumor genotype data would require a 2-5 times larger sample of offspring pairs to conclude linkage with an expected lod score value of 3 or greater, compared to analyses incorporating tumor genotype data. These results suggest that consideration of tumor genotype data using the proposed method can substantially improve the power of linkage analyses in cancer families.

Alleles↗

Upscaling Genotyping by Amplicon Sequencing With GBAS-GUI.

Genotyping by amplicon sequencing (GBAS) is a relatively low-cost approach for generating genotypic data compared with established genomic methods, making it highly scalable and particularly suitable for large-scale genetic monitoring projects. However, most existing analytical pipelines are either marker-specific, insufficiently scalable, or lacking efficient data management systems for the long-term integration of genotypic information, limiting the full potential of GBAS. Here, we address this gap by introducing GBAS-GUI (https://github.com/sonnenbe-dot/GBAS-GUI), a pipeline capable of generating GBAS-based genotypic data for a wide variety of loci at scale. GBAS-GUI integrates a graphical user interface with multiple checkpoints to improve accessibility and robustness. It implements multiprocessing architecture and a relational database that links genotypic data with associated sample metadata to enhance scalability and data management. The pipeline further enables marker screening through automated calculation of polymorphism information content (PIC) and implements a strategy to recover homologous genotypic information from paralogous loci with non-overlapping amplicon length ranges. Using multiple empirical datasets, we demonstrate substantial improvements in processing speed, database management and handling artefacts related to co-amplification of unspecific regions and duplicates of the same genomic region. We further show that incorporating the full sequence information captured by an amplicon increases marker information content beyond what is achievable with length-based genotyping alone and expands the analytical versatility of GBAS. Overall, GBAS-GUI provides a robust, scalable and versatile framework that unlocks the potential of GBAS for large-scale population genetic and phylogeographic studies.

Genotyping Techniques↗

Merging microsatellite data.

Genotype calling procedures vary from laboratory to laboratory for many microsatellite markers. Even within the same laboratory, application of different experimental protocols often leads to ambiguities. The impact of these ambiguities ranges from irksome to devastating. Resolving the ambiguities can increase effective sample size and preserve evidence in favor of disease-marker associations. Because different data sets may contain different numbers of alleles, merging is unfortunately not a simple process of matching alleles one to one. Merging data sets manually is difficult, time-consuming, and error-prone due to differences in genotyping hardware, binning methods, molecular weight standards, and curve fitting algorithms. Merging is particularly difficult if few or no samples occur in common, or if samples are drawn from ethnic groups with widely varying allele frequencies. It is dangerous to align alleles simply by adding a constant number of base pairs to the alleles of one of the data sets. To address these issues, we have developed a Bayesian model and a Markov chain Monte Carlo (MCMC) algorithm for sampling the posterior distribution under the model. Our computer program, MicroMerge, implements the algorithm and almost always accurately and efficiently finds the most likely correct alignment. Common allele frequencies across laboratories in the same ethnic group are the single most important cue in the model. MicroMerge computes the allelic alignments with the greatest posterior probabilities under several merging options. It also reports when data sets cannot be confidently merged. These features are emphasized in our analysis of simulated and real data.

Algorithms↗

Allowing for missing data at highly polymorphic genes when testing for maternal, offspring and maternal-fetal genotype incompatibility effects.

Genes can be associated with disease through an individual's inherited genotype, the maternal genotype or the interaction between these two. When the gene is highly polymorphic, it is more difficult to identify the gene's functional role than for less polymorphic loci, because different alleles at the locus may be associated with the disease through separate and joint effects from maternal and offspring genotypes. Family-based studies are used to test genetic associations because of their robustness to population stratification. However, parental genotype data are often missing, and omitting incompletely genotyped families is inefficient. Methods have been proposed to accommodate incomplete families in family-based association studies. They are not easily generalized to allow simultaneous examination of offspring allelic, maternal allelic and maternal-fetal genotype (MFG) incompatibility effects. Since many MFG incompatibility effects occur through matching between maternal and offspring's genotypes, we present an identity-by-state (IBS) framework to incorporate incomplete families in the MFG test when modeling genetic effects produced by a polymorphic gene. Using simulations, we examine the MFG test's performance with incomplete parental genotype data and an IBS framework. The MFG test using the IBS framework is immune to population stratification and efficiently uses information from incomplete families.

Algorithms↗