Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genotype Data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Laboratory Information Management Software for genotyping workflows: applications in high throughput crop genotyping.

BACKGROUND: With the advances in DNA sequencer-based technologies, it has become possible to automate several steps of the genotyping process leading to increased throughput. To efficiently handle the large amounts of genotypic data generated and help with quality control, there is a strong need for a software system that can help with the tracking of samples and capture and management of data at different steps of the process. Such systems, while serving to manage the workflow precisely, also encourage good laboratory practice by standardizing protocols, recording and annotating data from every step of the workflow. RESULTS: A laboratory information management system (LIMS) has been designed and implemented at the International Crops Research Institute for the Semi-Arid Tropics (ICRISAT) that meets the requirements of a moderately high throughput molecular genotyping facility. The application is designed as modules and is simple to learn and use. The application leads the user through each step of the process from starting an experiment to the storing of output data from the genotype detection step with auto-binning of alleles; thus ensuring that every DNA sample is handled in an identical manner and all the necessary data are captured. The application keeps track of DNA samples and generated data. Data entry into the system is through the use of forms for file uploads. The LIMS provides functions to trace back to the electrophoresis gel files or sample source for any genotypic data and for repeating experiments. The LIMS is being presently used for the capture of high throughput SSR (simple-sequence repeat) genotyping data from the legume (chickpea, groundnut and pigeonpea) and cereal (sorghum and millets) crops of importance in the semi-arid tropics. CONCLUSION: A laboratory information management system is available that has been found useful in the management of microsatellite genotype data in a moderately high throughput genotyping laboratory. The application with source code is freely available for academic users and can be downloaded from http://www.icrisat.org/gt-bt/lims/lims.asp.

Algorithms↗

Understanding recurrence in Mycobacterium avium complex pulmonary disease: genotypic strategies to support clinical decision-making.

Pulmonary disease caused by Mycobacterium avium complex (MAC-PD) is a chronic, recurrent disease, and its high recurrence rate after treatment makes clinical management difficult. Distinguishing whether recurrence is due to persistence of existing strains or reinfection with new strains is essential for establishing treatment strategies, preventing overuse of antimicrobials, and establishing infection control measures. According to reports, 54%-74% of MAC-PD recurrence is due to reinfection, which may be mainly related to environmental reservoirs such as household water supply. In this review, we present various clinical scenarios in which MAC-PD recurrence may occur and examine genotyping techniques as a strategy to distinguish and respond to them. From traditional methods such as IS1245-based restriction fragment length polymorphism, pulsed-field gel electrophoresis, and hsp65 and rpoB gene sequencing to high-resolution analysis techniques such as multilocus sequence testing and whole-genome sequencing, the latest molecular typing methods are comprehensively summarized. Integrating these genotype data into clinical settings, standardizing single-nucleotide polymorphism-based interpretation thresholds, and promoting the establishment of a global MAC strain database will make a substantial contribution to more accurately distinguishing the recurrence mechanisms of MAC-PD and establishing personalized treatment strategies.IMPORTANCEThe global burden of nontuberculous mycobacterial pulmonary disease (PD) is increasing, with Mycobacterium avium (MAC)-PD being the most prevalent and clinically challenging form. Its low treatment success rates, high frequency of recurrence, and persistent environmental exposure complicate both diagnosis and management. A critical clinical issue is determining whether recurrence represents true relapse, due to persistence of the original strain, or reinfection with a new strain, as this guides treatment and prevents overtreatment. Genotypic strategies capable of resolving strain-level differences can improve diagnostic accuracy, prevent misclassification, and ultimately support more informed treatment decisions. Therefore, integrating genotyping data into clinical workflows, standardizing single-nucleotide polymorphism thresholds, and establishing a global MAC strain database will not only support personalized treatment but also enhance the broader public health response to this disease.

Humans↗

Intra- and interpopulation genotype reconstruction from tagging SNPs.

The optimal method to be used for tSNP selection, the applicability of a reference LD map to unassayed populations, and the scalability of these methods to genome-wide analysis, all remain subjects of debate. We propose novel, scalable matrix algorithms that address these issues and we evaluate them on genotypic data from 38 populations and four genomic regions (248 SNPs typed for approximately 2000 individuals). We also evaluate these algorithms on a second data set consisting of genotypes available from the HapMap database (1336 SNPs for four populations) over the same genomic regions. Furthermore, we test these methods in the setting of a real association study using a publicly available family data set. The algorithms we use for tSNP selection and unassayed SNP reconstruction do not require haplotype inference and they are, in principle, scalable even to genome-wide analysis. Moreover, they are greedy variants of recently developed matrix algorithms with provable performance guarantees. Using a small set of carefully selected tSNPs, we achieve very good reconstruction accuracy of "untyped" genotypes for most of the populations studied. Additionally, we demonstrate in a quantitative manner that the chosen tSNPs exhibit substantial transferability, both within and across different geographic regions. Finally, we show that reconstruction can be applied to retrieve significant SNP associations with disease, with important genotyping savings.

Algorithms↗

Apolipoprotein J polymorphisms and serum HDL cholesterol levels in African blacks.

Apolipoprotein J (apoJ, protein; APOJ, gene) is found in serum associated with high-density lipoprotein (HDL) subfractions, which also contain apolipoprotein A-I (apoA1) and cholesteryl ester transfer protein. ApoJ has been shown to be involved in a variety of physiological functions, including lipid transport. In earlier studies we reported the existence of a common genetic polymorphism (APOJ*1 and APOJ*2 alleles) using isoelectric focusing (IEF) and immunoblotting. In this study we determined the molecular basis of this polymorphism and together with another polymorphism at codon 328 (G-->A) evaluated its relationship with serum HDL cholesterol and apoA1 levels in 767 African blacks stratified by staff level: junior (less affluent, n = 450) and senior (more affluent, n = 317). The molecular analysis of the cathodally shifted APOJ*2 allele on IEF gels revealed an amino acid substitution of asparagine by histidine resulting from a missense mutation (A-->C) at codon 317 in exon 7. The frequency of the APOJ*2 (C) allele of codon 317 in the total sample was 0.267, whereas that of the less common allele A of codon 328 was 0.04. Despite their close proximity, no linkage disequilibrium was observed between the 2 polymorphisms. The impact of the codon 317 polymorphic variation was significant on serum HDL cholesterol (p = 0.003) and HDL3 cholesterol (p = 0.001) in junior staff. The adjusted mean values of these traits were higher in the codon 317 APOJ*2/*2 genotype than in the *1/*1 and *1/*2 genotypes. Overall, the APOJ codon 317 polymorphism explained 10.2% and 8.3% of the phenotypic variation in HDL cholesterol and HDL3 cholesterol, respectively, in junior staff. The codon 328 polymorphism showed a significant effect on HDL2 cholesterol (p = 0.039) and apoA1 (p = 0.007) only in junior women and accounted for 2.5% and 4.2% of the phenotypic variation in HDL2 cholesterol and apoA1, respectively. We also analyzed the combined effects of these genotypes at the 2 polymorphic sites. Significant effects on HDL cholesterol (p = 0.004) and HDL3 cholesterol (p = 0.008) in junior men and on HDL2 cholesterol (p = 0.003) in junior women were observed in the combined genotype data. The 2-locus genotypes explained 6.0% and 5.3% of the residual phenotypic variation of HDL cholesterol and HDL3 cholesterol in junior men and 10.4% of HDL2 cholesterol in junior women. These data indicate that the effect of the APOJ polymorphism on HDL cholesterol levels is modulated by socioeconomic status, as measured by staff level. Given the association of HDL and its subfractions with cardiovascular disease, these polymorphisms may lead to a better understanding of interracial differences in the risk of cardiovascular disease.

Adult↗

Molecular epidemiology of Legionella pneumophila environmental isolates representing nine different serogroups determined by automated ribotyping and pulsed-field gel electrophoresis.

The purposes of the study were (i) to describe the abundance and epidemiology of Legionellaceae in the man-made environment in a northern Italian area, (ii) to assess the concordance between pulsed-field gel electrophoresis (PFGE) and automated ribotyping (AR) techniques for genotyping L. pneumophila and (iii) to investigate the correlation between serogrouping and genotyping data. Water was sampled from reservoirs in 12 buildings across an area of 80-km radius. Despite the water temperature always being maintained above 55 degrees C, all of the buildings sampled were contaminated with Legionellaceae on at least one occasion and 63 L. pneumophila isolates representing nine different serogroups were collected. The two DNA methods revealed a high degree of genetic heterogeneity, even though identical L. pneumophila clones were recovered at different sites. The AR technique provided a fairly reliable approximation of PFGE results (73% concordance), however there was poor correlation between serogrouping and genotyping data as identical DNA fingerprints were shared by isolates of different serogroups.

Electrophoresis, Gel, Pulsed-Field↗

Applying data mining techniques to the mapping of complex disease genes.

The simulated sequence data for the Genetic Analysis Workshop 12 were analyzed using data mining techniques provided by SAS ENTERPRISE MINER Release 4.0 in addition to traditional statistical tests for linkage and association of genetic markers with disease status. We examined two ways of combining these approaches to make use of the covariate data along with the genotypic data. The result of incorporating data mining techniques with more classical methods is an improvement in the analysis, both by correctly classifying the affection status of more individuals and by locating more single nucleotide polymorphisms related to the disease, relative to analyses that use classical methods alone.

Chromosome Mapping↗

A combined analysis of D22S278 marker alleles in affected sib-pairs: support for a susceptibility locus for schizophrenia at chromosome 22q12. Schizophrenia Collaborative Linkage Group (Chromosome 22).

Several groups have reported weak evidence for linkage between schizophrenia and genetic markers located on chromosome 22q using the lod score method of analysis. However these findings involved different genetic markers and methods of analysis, and so were not directly comparable. To resolve this issue we have performed a combined analysis of genotypic data from the marker D22S278 in multiply affected schizophrenic families derived from 11 independent research groups worldwide. This marker was chosen because it showed maximum evidence for linkage in three independent datasets (Vallada et al., Am J Med Genet 60:139-146, 1995; Polymeropoulos et al., Neuropsychiatr Genet 54:93-99, 1994; Lasseter et al., Am J Med Genet, 60:172-173, 1995. Using the affected sib-pair method as implemented by the program ESPA, the combined dataset showed 252 alleles shared compared with 188 alleles not share (chi-square 9.31, 1df, P = 0.001) where parental genotype data was completely known. When sib-pairs for whom parental data was assigned according to probability were included the number of alleles shared was 514.1 compared with 437.8 not shared (chi-square 6.12, 1df, P = 0.006). Similar results were obtained when a likelihood ratio method for sib-pair analysis was used. These results indicate that may be a susceptibility locus for schizophrenia at 22q12.

Alleles↗

High density linkage disequilibrium mapping using models of haplotype block variation.

MOTIVATION: The presence of millions of single nucleotide polymorphisms (SNPs) in the human genome has spurred interest in genetic mapping methods based on linkage disequilibrium. The recently discovered haplotype block structure of human variation promises to improve the effectiveness of these methods. A key difficulty for mapping techniques is the cost involved in separately identifying the haplotypes on each of an individual's chromosomes. RESULTS: We present a new approach for performing linkage disequilibrium mapping using high density haplotype or genotype data. Our method is based on a statistical model of haplotype block variation, which takes account of recombination hotspots, bottlenecks, genetic drift and mutation. We test our technique on two empirically determined high density datasets, attempting to recover the location of an SNP which was hidden and converted into phenotype information. We compare the results against a mapping method based on individual SNPs as well as a competing haplotype-based approach. We show that our strategy significantly outperforms these other approaches when used as a guide for resequencing and that it can also deal with both unphased genotype data and low penetrance diseases. AVAILABILITY: HaploBlock executables for Linux, Mac OS X and Sun OS, as well as user documentation, are available online at http://bioinfo.cs.technion.ac.il/haploblock/

Artificial Intelligence↗

Correlation between genotype, metabolic data, and clinical presentation in carnitine palmitoyltransferase 2 (CPT2) deficiency.

Carnitine palmitoyltransferase 2 (CPT2) deficiency, the most common inherited disease of the mitochondrial long-chain fatty acid (LCFA) oxidation, may result in distinct clinical phenotypes, namely a mild adult muscular form and a severe hepatocardiomuscular disease with an onset in the neonatal period or in infancy. In order to understand the mechanisms underlying the difference in severity between these phenotypes, we analyzed a cohort of 20 CPT2-deficient patients being affected either with the infantile (seven patients) or the adult onset form of the disease (13 patients). Using a combination of direct sequencing and denaturing gradient gel electrophoresis, 13 CPT2 mutations were identified, including five novel ones, namely: 371G>A (R124Q), 437A>C (N146T), 481C>T (R161W), 983A>G (D328G), and 1823G>C (D608H). After updating the spectrum of CPT2 mutations (n=39) and genotypes (n=38) as well as their consequences on CPT2 activity and LCFA oxidation, it appears that both the type and location of CPT2 mutations and one or several additional genetic factors to be identified would modulate the LCFA flux and therefore the severity of the disease.

Adult↗

Mapping trait loci by use of inferred ancestral recombination graphs.

Large-scale association studies are being undertaken with the hope of uncovering the genetic determinants of complex disease. We describe a computationally efficient method for inferring genealogies from population genotype data and show how these genealogies can be used to fine map disease loci and interpret association signals. These genealogies take the form of the ancestral recombination graph (ARG). The ARG defines a genealogical tree for each locus, and, as one moves along the chromosome, the topologies of consecutive trees shift according to the impact of historical recombination events. There are two stages to our analysis. First, we infer plausible ARGs, using a heuristic algorithm, which can handle unphased and missing data and is fast enough to be applied to large-scale studies. Second, we test the genealogical tree at each locus for a clustering of the disease cases beneath a branch, suggesting that a causative mutation occurred on that branch. Since the true ARG is unknown, we average this analysis over an ensemble of inferred ARGs. We have characterized the performance of our method across a wide range of simulated disease models. Compared with simpler tests, our method gives increased accuracy in positioning untyped causative loci and can also be used to estimate the frequencies of untyped causative alleles. We have applied our method to Ueda et al.'s association study of CTLA4 and Graves disease, showing how it can be used to dissect the association signal, giving potentially interesting results of allelic heterogeneity and interaction. Similar approaches analyzing an ensemble of ARGs inferred using our method may be applicable to many other problems of inference from population genotype data.

Algorithms↗

Reporting, appraising, and integrating data on genotype prevalence and gene-disease associations.

The recent completion of the first draft of the human genome sequence and advances in technologies for genomic analysis are generating tremendous opportunities for epidemiologic studies to evaluate the role of genetic variants in human disease. Many methodological issues apply to the investigation of variation in the frequency of allelic variants of human genes, of the possibility that these influence disease risk, and of assessment of the magnitude of the associated risk. Based on a Human Genome Epidemiology workshop, a checklist for reporting and appraising studies of genotype prevalence and studies of gene-disease associations was developed. This focuses on selection of study subjects, analytic validity of genotyping, population stratification, and statistical issues. Use of the checklist should facilitate the integration of evidence from these studies. The relation between the checklist and grading schemes that have been proposed for the evaluation of observational studies is discussed. Although the limitations of grading schemes are recognized, a robust approach is proposed. Other issues in the synthesis of evidence that are particularly relevant to studies of genotype prevalence and gene-disease association are discussed, notably identification of studies, publication bias, criteria for causal inference, and the appropriateness of quantitative synthesis.

Case-Control Studies↗

[Genetic predisposition factors for multiple sclerosis (based on the data from genotyping patients of the Russian ethnic group)].

Patients' genotyping of Russian ethnic group with multiple sclerosis (MS) was performed for the first time in two loci of the main complex of histocompatibility: in DR HLA class II (gene DRB1) and in the locus of tumor necrosis factors (TNF). There was no difference in the incidence of alleles groups which correlated with some specificity of DR. Meanwhile TNF-a1 and TNF-a9 alleles were encountered significantly more frequently in patients and TNF-a7 in control group. When all the patients and controls examined were divided into groups in dependence on combination of DR and TNF it was found that relations observed between MS and TNF-a7 and TNF-a9 alleles were displayed much more in individuals which carried alleles of gene DRB1, corresponding to DR15 specificity. These are alleles which are known as the main risk factor of MS in Caucasians. The patients with "protective" TNF-a7 allele were characterized by more favorable course of the disease. Thus highly significant genetic markers were revealed for the first time in region of TNF genes which were associated with increased or decreased risk of MS development at least in Russian ethnic group. There was also possibility of their interaction with group of DR15 alleles of DRB1 gene. One of the markers revealed (TNF-a7) occurred to be bound both with decreased risk of MS and with favorable clinical course, which was observed for genetic markers of MS for the first time. Manifestation of one or another property of TNF-a7 marker depends on the presence of alleles of DRB1 gene which corresponds to DR15 specificity.

Adult↗

Frequency of CYP2C9 genotypes among Omani patients receiving warfarin and its correlation with warfarin dose.

OBJECTIVES: This study was conducted to determine the frequency of CYP2C9 alleles in Omani patients receiving warfarin and to correlate genotyping data with warfarin dosage. The Omani population has Asian and African ethnicities. METHODS: CYP2C9 genotypes were determined by the polymerase chain reaction restriction fragment length polymorphism method. Non-parametric Kruskal-Wallis test was used to compare groups of continuous data for significance differences. RESULTS: Genotyping data showed that 12.7 and 5.8% of the samples were heterozygous for the CYP2C9*2 and CYP2C9*3 alleles, respectively. The CYP2C9*2 allele frequency was 0.074 in our population. It was 0.029 for CYP2C9*3. CONCLUSION: This is the first report on the presence of CYP2C9*2 allele homozygocity in any Asian or African population.

Adult↗

Software for analysis and manipulation of genetic linkage data.

We present eight computer programs written in the C programming language that are designed to analyze genotypic data and to support existing software used to construct genetic linkage maps. Although each program has a unique purpose, they all share the common goals of affording a greater understanding of genetic linkage data and of automating tasks to make computers more effective tools for map building. The PIC/HET and FAMINFO programs automate calculation of relevant quantities such as heterozygosity, PIC, allele frequencies, and informativeness of markers and pedigrees. PREINPUT simplifies data submissions to the Centre d'Etude du Polymorphisme Humain (CEPH) data base by creating a file with genotype assignments that CEPH's INPUT program would otherwise require to be input manually. INHERIT is a program written specifically for mapping the X chromosome: by assigning a dummy allele to males, in the nonpseudoautosomal region, it eliminates falsely perceived noninheritances in the data set. The remaining four programs complement the previously published genetic linkage mapping software CRI-MAP and LINKAGE. TWOTABLE produces a more readable format for the output of CRI-MAP two-point calculations; UNMERGE is the converse to CRI-MAP's merge option; and GENLINK and LINKGEN automatically convert between the genotypic data file formats required by these packages. All eight applications read input from the same types of data files that are used by CRI-MAP and LINKAGE. Their use has simplified the management of data, has increased knowledge of the content of information in pedigrees, and has reduced the amount of time needed to construct genetic linkage maps of chromosomes.

Alleles↗

Haplotype association analysis of discrete and continuous traits using mixture of regression models.

We present a regression-based method of haplotype association analysis for quantitative and dichotomous traits in samples consisting of unrelated individuals. The method takes account of uncertain phase by initially estimating haplotype frequencies and obtaining the posterior probabilities of all possible haplotype combinations in each individual, then using these as weights in a finite mixture of regression models. Using this method, different combinations of marker loci can be modeled, to find a parsimonious set of marker loci that are most predictive and therefore most likely to be closely associated with the a quantitative trait locus. The method has the additional advantage of being able to use individuals with some missing genotype data, by considering all possible genotypes at the missing markers. We have implemented this method using the SNPHAP and Mx programs and illustrated its use on published data on idiopathic generalized epilepsy.

Chromosome Mapping↗

Haplotype analysis in the Collaborative Study on the Genetics of Alcoholism data: double recombinants.

The presence of close double recombinants in genotyping data may help identify genotyping errors. Alternatively, putative double recombinants may be associated with genetic mechanisms that may be related to disease. Phase-known apparent double recombination events were identified in the Collaborative Study on the Genetics of Alcoholism data, and compared to the sex-specific genetic map at each region. A number of double recombinants occurred within a short genetic distance. Also, in some families multiple double recombinants were observed flanking the same genetic marker, both suggesting possible genotyping error. An excess of paternal double recombinants was identified, which is consistent with reports of sex-specific differential meiotic interference.

Alcoholism↗

Contribution of different HFE genotypes to iron overload disease: a pooled analysis.

PURPOSE: To determine the contribution of the C282Y and H63D mutations in the HFE gene to clinical expression of hereditary hemochromatosis. METHODS: Pooled analysis of 14 case-control studies reporting HFE genotype data, to evaluate the association of different HFE genotypes with iron overload. In addition, we used data from the pooled analysis and published data to estimate the penetrance of the C282Y/C282Y genotype. RESULTS: Homozygosity for the C282Y mutation carried the largest risk for iron overload (OR = 4383, 95% CI 1374 to >10,000) and accounted for the majority of hemochromatosis cases (attributable fraction (AF) = 0.73). Risks for other genotypes were much smaller: OR = 32 for genotype C282Y/H63D (95% CI 18.5 to 55.4, AF = 0.06); OR = 5.7 for H63D/H63D (95% CI 3.2 to 10.1, AF = 0.01); OR = 4.1 for C282Y heterozygosity (95% CI 2.9 to 5.8, with heterogeneity in study results, making this association uncertain); and OR = 1.6 for H63D heterozygosity (95% CI 1 to 2.6, AF = 0.03). Estimates of penetrance for the C282Y/C282Y genotype were highly sensitive to estimates of the prevalence of iron overload disease. At a prevalence of 2.5 per 1000 or less, penetrance of the C282Y/C282Y genotype is unlikely to exceed 50%. Penetrance of other HFE genotypes is much lower. CONCLUSIONS: C282Y homozygosity confers the highest risk for iron overload but the H63D mutation is also associated with increased risk. Our data indicate a gradient of risk associated with different HFE genotypes and thus suggest the presence of other modifiers, either genetic or environmental, that contribute to the clinical expression of hemochromatosis.

Amino Acid Substitution↗