Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genotype Data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Integration of genomic data in Electronic Health Records--opportunities and dilemmas.

OBJECTIVES: In this paper we give an overview about the challenge the postgenomic era poses on biomedical informaticists. The occurrence of new (genomic) data types necessitates new data models, new viewing metaphors and methods to deal with the disclosure of genomic data. We discuss integration issues when inferring phenotype and genotype data. Another challenge is to find the right phenotype to genotype data in order to get appropriate case numbers for sound clinical genotype-phenotype inference studies. METHODS: Genomic data could be integrated in an Electronic Health Record (EHR) in several ways. We describe patient-centered and pointer-based integration strategies and the corresponding data types and data models. The inference mechanisms for the interpretation of row data contain different agents. We describe vertical, horizontal and temporal agents. RESULTS: We have to deal with several new data types, not being standardized for EHR integration. Genomic data tends to be more structured than phenotype data. Beyond the development of new data models, vertical, horizontal and temporal agents have to be developed in order to link genotype and phenotype. As the genomic EHR will contain very sensitive data, confidentiality and privacy concerns have to be addressed. CONCLUSIONS: Given the necessity to capture both environment and genomic state of a patient and their interaction, clinical information systems have to be redesigned. While genotyping seems to be automatable easily, this is not the case for clinical information. More integration work on terminologies and ontologies has to be done.

Computational Biology↗

A combined linkage-physical map of the human genome.

We have constructed de novo a high-resolution genetic map that includes the largest set, to our knowledge, of polymorphic markers (N=14,759) for which genotype data are publicly available; that combines genotype data from both the Centre d'Etude du Polymorphisme Humain (CEPH) and deCODE pedigrees; that incorporates single-nucleotide polymorphisms; and that also incorporates sequence-based positional information. The position of all markers on our map is corroborated by both genomic sequence and recombination-based data. This specific combination of features maximizes marker inclusion, coverage, and resolution, making this map uniquely suitable as a comprehensive resource for determining genetic map information (order and distances) for any large set of polymorphic markers.

Chromosome Mapping↗

High-throughput association testing on DNA pools to identify genetic variants that confer susceptibility to acute myeloid leukemia.

We have evaluated the use of allele-specific PCR (AS PCR) on DNA pools as a tool for screening inherited genetic variants that may be associated with risk of adult acute myeloid leukemia (AML). Two DNA pools were constructed, one of 444 AML cases, and another of 823 matched controls. The pools were validated using individual genotyping data for GSTP1 and LTalpha variants. Allele frequencies for variants in GSTP1 and LTalpha were estimated using quantitative AS PCR, and when compared to individual genotyping data, a high degree of concordance was seen. AS primer pairs were designed for nine candidate genetic variants in DNA repair and cell cycle/apoptotic regulatory genes, including Cyclin D1 [codon 870 splice site variant (A>G)]; BRCA1, P871L; ERCC2, K751Q; FAS -1377 (G>A); hMLH1 -93 (G>A) and V219I; p21, S31R; and the XRCC1 R194W and R399Q variants. For six of these assays, there was at least 95% concordance between AS PCR genotyping and an alternative approach carried out on individual samples. Furthermore, these six AS PCR assays all accurately estimated allele frequencies in the pools that had been calculated using individual genotyping data. A significant disease association was seen with AML for the -1377 variant in FAS (odds ratio 1.76, 95% confidence interval 1.26-2.44). These data suggest that quantitative AS PCR can be used as an efficient screening technique for disease associations of genetic variants in DNA pools made from case-control studies.

Base Sequence↗

Statistical approaches to paternity analysis in natural populations and applications to the North Atlantic humpback whale.

We present a new method for paternity analysis in natural populations that is based on genotypic data that can take the sampling fraction of putative parents into account. The method allows paternity assignment to be performed in a decision theoretic framework. Simulations are performed to evaluate the utility and robustness of the method and to assess how many loci are necessary for reliable paternity inference. In addition we present a method for testing hypotheses regarding relative reproductive success of different ecologically or behaviorally defined groups as well as a new method for estimating the current population size of males from genotypic data. This method is an extension of the fractional paternity method to the case where only a proportion of all putative fathers have been sampled. It can also be applied to provide abundance estimates of the number of breeding males from genetic data. Throughout, the methods were applied to genotypic data collected from North Atlantic humpback whales (Megaptera novaeangliae) to test if the males that appear dominant during the mating season have a higher reproductive success than the subdominant males.

Animals↗

Assessment and management of single nucleotide polymorphism genotype errors in genetic association analysis.

Single nucleotide polymorphisms (SNP) may be used in case-control designs to test for association between a marker (the SNP) and a disease. However, such designs usually assume that the genotype data are reported without error. We propose a method, the reduced penetrance model method (RPM) that allows for errors in a case-control design, as compared to the full penetrance model method (FPM), that assumes data are errorless. Pearson's chi 2 applied to a 2 x 2 contingency table is the test statistic considered. Additionally, we provide a likelihood method to estimate error rates using SNP genotype data in CEPH pedigrees. We test our method (RPM) against the standard method (FPM) using simulated data. All SNP loci are assumed to have two alleles, coded 1 and 2. We consider three pairs of error rates, two different sample sizes, and two sets of allele frequencies for the SNP locus. SNP genotype data in two populations are simulated under a null hypothesis (allele frequencies equal in both populations) and under an alternative hypothesis (allele frequencies differ between two populations). The total number of simulations is 24; 12 simulations under the null hypothesis, and 12 simulations under the alternative. The significance level threshold is 5%. For the null case, 9/12 (75%) of the simulations show no increase in type I error under RPM, while 3/12 (25%) show a slight increase (rejecting the null for at most 7% of the replicates). There is no increase in the type I error rate for FPM method, which can also be shown analytically. For the alternative case (power), there is a consistent increase in power for the RPM method as compared to FPM method, and average increase of 0.02 for the simulations considered. When sample sizes are large there is virtually no difference in power between RPM and FPM methods. Also, the RPM method provides consistently more accurate allele frequency estimates for the various populations. Our likelihood method to estimate error rates with CEPH pedigrees provides good estimates on average. The largest difference between a true error rate and our average estimated error rate is 0.006. However, there is a fair amount of variability in the estimates, suggesting the need for multiple experiments or larger numbers of CEPH pedigrees. Researchers may use the methods presented in this paper to (1) estimate error rates for their automated genotyping process, and (2) allow for such errors in association analyses, thereby increasing power to detect differences between allele frequencies in case and control populations when errors are present.

Alleles↗

Methods of quantifying and visualising outbreaks of tuberculosis using genotypic information.

Genotypic data from pathogenic isolates are often used to measure the extent of infectious disease transmission. These methods include phylogenetic reconstruction and the evaluation of clustering indices. The first aim of this paper is to critique current methods used to analyse genotypic data from molecular epidemiological studies of tuberculosis. In particular, by not accounting for the mutation rate of markers, errors arise in making inferences about outbreaks based on genotypic information. The second aim is to suggest a new way to represent genotypic data visually, involving graphs and trees. We also discuss some interpretations and modifications of existing indices. Although our focus is tuberculosis, the methods we discuss are generally applicable to any directly transmissible clonal pathogen.

Cluster Analysis↗

Ideal discrimination of discrete clinical endpoints using multilocus genotypes.

Multifactor Dimensionality Reduction (MDR) is a method for the classification and prediction of discrete clinical endpoints using attributes constructed from multilocus genotype data. Empirical studies with both real and simulated data suggest that MDR has good power for detecting gene-gene interactions in the absence of independent main effects. The purpose of this study is to develop an objective, theory-driven approach to evaluate the strengths and limitations of MDR. To accomplish this goal, we borrow concepts from ideal observer analysis used in visual perception to evaluate the theoretical limits of classifying and predicting discrete clinical endpoints using multilocus genotype data. We conclude that MDR ideally discriminates between low risk and high risk subjects using attributes constructed from multilocus genotype data. We also how that the classification approach used once a multilocus attribute is constructed is similar to that of a naive Bayes classifier. This study provides a theoretical foundation for the continued development, evaluation, and application of the MDR as a data mining tool in the domain of statistical genetics and genetic epidemiology.

Animals↗

Bayesian inference of recent migration rates using multilocus genotypes.

A new Bayesian method that uses individual multilocus genotypes to estimate rates of recent immigration (over the last several generations) among populations is presented. The method also estimates the posterior probability distributions of individual immigrant ancestries, population allele frequencies, population inbreeding coefficients, and other parameters of potential interest. The method is implemented in a computer program that relies on Markov chain Monte Carlo techniques to carry out the estimation of posterior probabilities. The program can be used with allozyme, microsatellite, RFLP, SNP, and other kinds of genotype data. We relax several assumptions of early methods for detecting recent immigrants, using genotype data; most significantly, we allow genotype frequencies to deviate from Hardy-Weinberg equilibrium proportions within populations. The program is demonstrated by applying it to two recently published microsatellite data sets for populations of the plant species Centaurea corymbosa and the gray wolf species Canis lupus. A computer simulation study suggests that the program can provide highly accurate estimates of migration rates and individual migrant ancestries, given sufficient genetic differentiation among populations and sufficient numbers of marker loci.

Animal Migration↗

Fluorescence-based fragment size analysis.

The successful application of capillary electrophoresis technology to the genotyping of various types of polymorphisms has been well documented. The flexibility and automation of the Applied Biosystems 3100 Genetic Analyzer make it an excellent capillary electrophoresis platform for the generation of high quality genotype data. These data are readily applied to pharmacogenomic investigations of various types. Included in this chapter is a protocol for the generation of genotype data using minimal template deoxyribonucleic acid and maximizing the automation of both data analysis and genotype assignment through the use of the Applied Biosystems GeneMapper 3.0 software package.

DNA↗

A consensus linkage map of the chicken genome.

A consensus linkage map has been developed in the chicken that combines all of the genotyping data from the three available chicken mapping populations. Genotyping data were contributed by the laboratories that have been using the East Lansing and Compton reference populations and from the Animal Breeding and Genetics Group of the Wageningen University using the Wageningen/Euribrid population. The resulting linkage map of the chicken genome contains 1889 loci. A framework map is presented that contains 480 loci ordered on 50 linkage groups. Framework loci are defined as loci whose order relative to one another is supported by odds greater then 3. The possible positions of the remaining 1409 loci are indicated relative to these framework loci. The total map spans 3800 cM, which is considerably larger than previous estimates for the chicken genome. Furthermore, although the physical size of the chicken genome is threefold smaller then that of mammals, its genetic map is comparable in size to that of most mammals. The map contains 350 markers within expressed sequences, 235 of which represent identified genes or sequences that have significant sequence identity to known genes. This improves the contribution of the chicken linkage map to comparative gene mapping considerably and clearly shows the conservation of large syntenic regions between the human and chicken genomes. The compact physical size of the chicken genome, combined with the large size of its genetic map and the observed degree of conserved synteny, makes the chicken a valuable model organism in the genomics as well as the postgenomics era. The linkage maps, the two-point lod scores, and additional information about the loci are available at web sites in Wageningen (http://www.zod.wau.nl/vf/ research/chicken/frame_chicken.html) and East Lansing (http://poultry.mph.msu.edu/).

Animals↗

Haplotype-based association analysis in cohort and nested case-control studies.

Genetic epidemiologic studies often collect genotype data at multiple loci within a genomic region of interest from a sample of unrelated individuals. One popular method for analyzing such data is to assess whether haplotypes, i.e., the arrangements of alleles along individual chromosomes, are associated with the disease phenotype or not. For many study subjects, however, the exact haplotype configuration on the pair of homologous chromosomes cannot be derived with certainty from the available locus-specific genotype data (phase ambiguity). In this article, we consider estimating haplotype-specific association parameters in the Cox proportional hazards model, using genotype, environmental exposure, and the disease endpoint data collected from cohort or nested case-control studies. We study alternative Expectation-Maximization algorithms for estimating haplotype frequencies from cohort and nested case-control studies. Based on a hazard function of the disease derived from the observed genotype data, we then propose a semiparametric method for joint estimation of relative-risk parameters and the cumulative baseline hazard function. The method is greatly simplified under a rare disease assumption, for which an asymptotic variance estimator is also proposed. The performance of the proposed estimators is assessed via simulation studies. An application of the proposed method is presented, using data from the Alpha-Tocopherol, Beta-Carotene Cancer Prevention Study.

Algorithms↗

Clinically Relevant Pharmacogenomic Variant Frequencies in Kazakh, Russian, and Uzbek Population Groups Residing in Kazakhstan.

Central Asian populations remain underrepresented in pharmacogenomic research, limiting the availability of population-specific data for genotype-informed prescribing and precision medicine. This study analyzed clinically relevant pharmacogenomic variant frequencies in Kazakh, Russian, and Uzbek population groups residing in Kazakhstan using genome-wide genotype data from 1301 individuals: Kazakh (n = 1111), Russian (n = 156), and Uzbek (n = 34). ClinPGx, a PharmGKB-based clinical annotation framework that prioritizes variant-drug associations according to levels of evidence, was used to select variants with evidence levels 1A, 1B, and 2A. In total, 112 directly genotyped variants were retained for population-specific allele and genotype frequency analysis. All 112 variants were queried against the gnomAD v4.1 genome and exome reference datasets. Of these, matching allele-frequency data for the predefined reported allele were available in at least one of the two gnomAD datasets for 103 variants, whereas for 9 variants the VEP-based query did not return a matching gnomAD frequency for that allele. Frequencies were reported for the same predefined reported allele across all groups, and differences between the study groups were assessed using 95% confidence intervals, Fisher's exact tests, and false discovery rate correction. Genotype counts and the proportions of individuals carrying at least one copy of the reported allele were also summarized for all selected variants. Several pharmacogenomic variants showed population-specific frequency patterns, including NUDT15 rs116855232, SLCO1B1 rs4149056, VKORC1 rs9934438, and UGT1A1 rs10929302. Comparison with gnomAD showed that the observed frequencies were variant-specific and could not be consistently approximated by a single broad genetic ancestry group. Reference-based population structure analysis provided additional ancestry context and supported separate reporting by population group. The study did not evaluate clinical outcomes or make individual prescribing recommendations, and the small Uzbek sample size limits the precision of frequency estimates for this group, particularly for rare variants. Overall, this study provides a clinically prioritized pharmacogenomic frequency resource for underrepresented population groups in Kazakhstan and supports broader Central Asian representation in pharmacogenomic implementation research.

Central Asia↗

Effectiveness of selective genotyping for detection of quantitative trait loci: an analysis of grain and malt quality traits in three barley populations.

Marker genotype data and grain and malt quality phenotype data from three barley (Hordeum vulgare L.) mapping populations were used to investigate the feasibility of selective genotyping for detection of quantitative trait loci (QTLs). With selective genotyping, only individuals with high and low phenotypic values for the trait of interest are genotyped. Here, genotyping of 10 to 70% of each population (i.e., 5 to 35% in each tail of the phenotypic distribution) was considered. Genomic positions detected by selective genotyping were compared to QTL position estimates from interval mapping analysis using marker genotype data from the entire population. Selective genotyping reliably detected almost all of the mapped QTLs, often with only 10% of the population genotyped. Selective genotyping also detected spurious QTLs in regions of the genome where no significant QTL had been mapped. Even with additional genotyping to verify putative QTLs, the total genotyping effort for detection of QTLs for a single trait by selective genotyping was usually less than 30% of that required for conventional interval mapping. Simultaneous investigation of two or more traits by selective genotyping would require additional genotyping effort, but could still be worthwhile.

Genotype↗

Using genetic markers to directly estimate gene flow and reproductive success parameters in plants on the basis of naturally regenerated seedlings.

Estimating seed and pollen gene flow in plants on the basis of samples of naturally regenerated seedlings can provide much needed information about "realized gene flow," but seems to be one of the greatest challenges in plant population biology. Traditional parentage methods, because of their inability to discriminate between male and female parentage of seedlings, unless supported by uniparentally inherited markers, are not capable of precisely describing seed and pollen aspects of gene flow realized in seedlings. Here, we describe a maximum-likelihood method for modeling female and male parentage in a local plant population on the basis of genotypic data from naturally established seedlings and when the location and genotypes of all potential parents within the population are known. The method models female and male reproductive success of individuals as a function of factors likely to influence reproductive success (e.g., distance of seed dispersal, distance between mates, and relative fecundity--i.e., female and male selection gradients). The method is designed to account for levels of seed and pollen gene flow into the local population from unsampled adults; therefore, it is well suited to isolated, but also wide-spread natural populations, where extensive seed and pollen dispersal complicates traditional parentage analyses. Computer simulations were performed to evaluate the utility and robustness of the model and estimation procedure and to assess how the exclusion power of genetic markers (isozymes or microsatellites) affects the accuracy of the parameter estimation. In addition, the method was applied to genotypic data collected in Scots pine (isozymes) and oak (microsatellites) populations to obtain preliminary estimates of long-distance seed and pollen gene flow and the patterns of local seed and pollen dispersal in these species.

Computer Simulation↗

Integrating clinical and laboratory data in genetic studies of complex phenotypes: a network-based data management system.

The identification of genes underlying a complex phenotype can be a massive undertaking, and may require a much larger sample size than thought previously. The integration of such large volumes of clinical and laboratory data has become a major challenge. In this paper we describe a network-based data management system designed to address this challenge. Our system offers several advantages. Since the system uses commercial software, it obviates the acquisition, installation, and debugging of privately-available software, and is fully compatible with Windows and other commercial software. The system uses relational database architecture, which offers exceptional flexibility, facilitates complex data queries, and expedites extensive data quality control. The system is particularly designed to integrate clinical and laboratory data efficiently, producing summary reports, pedigrees, and exported files containing both phenotype and genotype data in a virtually unlimited range of formats. We describe a comprehensive system that manages clinical, DNA, cell line, and genotype data, but since the system is modular, researchers can set up only those elements which they need immediately, expanding later as needed.

Clinical Laboratory Information Systems↗

iHAP--integrated haplotype analysis pipeline for characterizing the haplotype structure of genes.

BACKGROUND: The advent of genotype data from large-scale efforts that catalog the genetic variants of different populations have given rise to new avenues for multifactorial disease association studies. Recent work shows that genotype data from the International HapMap Project have a high degree of transferability to the wider population. This implies that the design of genotyping studies on local populations may be facilitated through inferences drawn from information contained in HapMap populations. RESULTS: To facilitate analysis of HapMap data for characterizing the haplotype structure of genes or any chromosomal regions, we have developed an integrated web-based resource, iHAP. In addition to incorporating genotype and haplotype data from the International HapMap Project and gene information from the UCSC Genome Browser Database, iHAP also provides capabilities for inferring haplotype blocks and selecting tag SNPs that are representative of haplotype patterns. These include block partitioning algorithms, block definitions, tag SNP definitions, as well as SNPs to be "force included" as tags. Based on the parameters defined at the input stage, iHAP performs on-the-fly analysis and displays the result graphically as a webpage. To facilitate analysis, intermediate and final result files can be downloaded. CONCLUSION: The iHAP resource, available at http://ihap.bii.a-star.edu.sg, provides a convenient yet flexible approach for the user community to analyze HapMap data and identify candidate targets for genotyping studies.

Algorithms↗

Selecting additional tag SNPs for tolerating missing data in genotyping.

BACKGROUND: Recent studies have shown that the patterns of linkage disequilibrium observed in human populations have a block-like structure, and a small subset of SNPs (called tag SNPs) is sufficient to distinguish each pair of haplotype patterns in the block. In reality, some tag SNPs may be missing, and we may fail to distinguish two distinct haplotypes due to the ambiguity caused by missing data. RESULTS: We show there exists a subset of SNPs (referred to as robust tag SNPs) which can still distinguish all distinct haplotypes even when some SNPs are missing. The problem of finding minimum robust tag SNPs is shown to be NP-hard. To find robust tag SNPs efficiently, we propose two greedy algorithms and one linear programming relaxation algorithm. The experimental results indicate that (1) the solutions found by these algorithms are quite close to the optimal solution; (2) the genotyping cost saved by using tag SNPs can be as high as 80%; and (3) genotyping additional tag SNPs for tolerating missing data is still cost-effective. CONCLUSION: Genotyping robust tag SNPs is more practical than just genotyping the minimum tag SNPs if we can not avoid the occurrence of missing data. Our theoretical analysis and experimental results show that the performance of our algorithms is not only efficient but the solution found is also close to the optimal solution.

Algorithms↗

Precise definition of anonymization in genetic polymorphism studies.

Anonymization is an essential tool to protect privacy of participants in epidemiological studies. This paper classifies types of anonymization in genetic polymorphism studies, providing precise definitions. They are: 1) unlinkable anonymization at enrollment without a participant list; 2) unlinkable anonymization before genotyping with a participant list; 3) linkable anonymization; 4) unlinkable anonymization for outsiders; and 5) linkable anonymization for outsiders. The classification in view of accessibility to a table including genotype data with directly identifiable data such as names is important; if such tables exist, staff may obtain genotype information about participants. The first three modes are defined here as anonymization unaccessible to genotype data with directly identifiable information for research staff. Anonymization with a key code held by participants is possible with any of the above anonymization modes, by which participants can access to their own genotypes through telephone or internet. A guideline issued on March 29, 2001 with collaboration of three Ministries in Japan defines "anonymization in a linkable fashion" and "anonymization in an unlinkable fashion", "for the purpose of preventing the personal information from being divulged externally in violation of law, the present guidelines or a research protocol", but the contents are not clear in practice. The proposed definitions will be useful when we describe and discuss the preferable mode of anonymization for a given polymorphism study.

Anonymous Testing↗