Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genotype Data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Constructing near-perfect phylogenies with multiple homoplasy events.

MOTIVATION: We explore the problem of constructing near-perfect phylogenies on bi-allelic haplotypes, where the deviation from perfect phylogeny is entirely due to homoplasy events. We present polynomial-time algorithms for restricted versions of the problem. We show that these algorithms can be extended to genotype data, in which case the problem is called the near-perfect phylogeny haplotyping (NPPH) problem. We present a near-optimal algorithm for the H1-NPPH problem, which is to determine if a given set of genotypes admit a phylogeny with a single homoplasy event. The time-complexity of our algorithm for the H1-NPPH problem is O(m2(n + m)), where n is the number of genotypes and m is the number of SNP sites. This is a significant improvement over the earlier O(n4) algorithm. We also introduce generalized versions of the problem. The H(1, q)-NPPH problem is to determine if a given set of genotypes admit a phylogeny with q homoplasy events, so that all the homoplasy events occur in a single site. We present an O(m(q+1)(n + m)) algorithm for the H(1,q)-NPPH problem. RESULTS: We present results on simulated data, which demonstrate that the accuracy of our algorithm for the H1-NPPH problem is comparable to that of the existing methods, while being orders of magnitude faster. AVAILABILITY: The implementation of our algorithm for the H1-NPPH problem is available upon request.

Algorithms↗

Kin selection may influence fostering behaviour in Antarctic fur seals (Arctocephalus gazella).

Fostering confers obvious advantages to the offspring but is seemingly costly to the caregiver. Such behaviour is particularly paradoxical in seals where the energetic investment in milk is very high and has led to the suggestion that this behaviour may have evolved through either kin selection or reciprocity. We used a combination of genetic and behavioural data to investigate whether kin selection plays a role in the fostering behaviour observed in a well-studied population of Antarctic fur seals (Arctocephalus gazella) from Bird Island, South Georgia. Genotypic data from eight highly polymorphic microsatellite markers were used to estimate relatedness among mother-pup pairs, foster mother-pup pairs and the total population. Mean relatedness was found to be significantly higher for foster mother-pup pairs than that observed for the total population, suggesting that kin selection could have a role in the maintenance of fostering behaviour in this species.

Animals↗

A note on phasing long genomic regions using local haplotype predictions.

The common approaches for haplotype inference from genotype data are targeted toward phasing short genomic regions. Longer regions are often tackled in a heuristic manner, due to the high computational cost. Here, we describe a novel approach for phasing genotypes over long regions, which is based on combining information from local predictions on short, overlapping regions. The phasing is done in a way, which maximizes a natural maximum likelihood criterion. Among other things, this criterion takes into account the physical length between neighboring single nucleotide polymorphisms. The approach is very efficient and is applied to several large scale datasets and is shown to be successful in two recent benchmarking studies (Zaitlen et al., in press; Marchini et al., in preparation). Our method is publicly available via a webserver at http://research.calit2.net/hap/.

Algorithms↗

Two common polymorphisms in the APO A-IV coding gene: their evolution and linkage disequilibrium.

Human apolipoprotein A-IV (APO A-IV) exhibits a common protein polymorphism detectable by isoelectric focusing (IEF) due to a single base substitution at codon 360 which replaces the frequently occurring glutamine residue (allele 1) with histidine (allele 2). Recently, sequence analysis of the APO A-IV coding region has revealed another common nucleotide substitution at codon 347 which converts the commonly present threonine residue (allele A) into serine (allele T). In order to investigate the extent of genetic variation at codon 347, we screened DNA samples from 192 unrelated individuals using a polymerase chain reaction based assay. The frequencies of the two alleles, A-IV*A and A-IV*T, were 0.81 and 0.19, respectively, with average heterozygosity 0.31. Genetic screening of the corresponding 192 plasma samples by IEF gave frequencies of 0.922 and 0.078 for the A-IV*1 and A-IV*2 alleles, respectively, at codon 360 with average heterozygosity 0.14. Genotype data at the two polymorphic sites were used to assign unequivocal haplotypes to all the 384 chromosomes. Of the expected four haplotypes (A1, T1, A2, and T2) only three were observed and their frequencies were 0.732 for A1, 0.190 for T1 and 0.078 for A2, with average heterozygosity 0.42. Although our data indicate significant linkage disequilibrium between the two sites (chi 21 = 7.65, P < 0.006, standardized disequilibrium constant phi = -0.14) the degree of nonrandom association varied between alleles at the two sites. Based upon allele frequency data and variable linkage disequilibrium between alleles, we propose that the A2 and T1 haplotypes may have evolved from the parental A1 haplotype by two independent mutations.

Apolipoproteins A↗

Method for using complete and incomplete trios to identify genes related to a quantitative trait.

A number of tests for linkage and association with qualitative traits have been developed, with the most well-known being the transmission/disequilibrium test (TDT). For quantitative traits, varying extensions of the TDT have been suggested. The quantitative trait approach we propose is based on extending the log-linear model for case-parent trio data (Weinberg et al. [1998] Am. J. Hum. Genet. 62:969-978). Like the log-linear approach for qualitative traits, our proposed polytomous logistic approach for quantitative traits allows for population admixture by conditioning on parental genotypes. Compared to other methods, simulations demonstrate good power and robustness of the proposed test under various scenarios of the genotype effect, distribution of the quantitative trait, and population stratification. In addition, missing parental genotype data can be accommodated through an expectation-maximization (EM) algorithm approach. The EM approach allows recovery of most of the lost power due to incomplete trios.

Computer Simulation↗

Polymorphisms and haplotypes in the human immunoglobulin kappa locus.

By comparing the restriction patterns of the DNA from 23 unrelated individuals 16 polymorphisms were defined which allowed us to differentiate between the duplicated copies Op, Ap, Lp and Od, Ad, Ld of the kappa locus (p for the C kappa proximal, d for the distal copy). Some of these duplication-differentiating polymorphisms or DDP revealed also allelic differences between individuals; they are therefore restriction fragment length polymorphism (RFLP) markers at the same time. Three RFLP in the single copy B-J kappa-C kappa region were included into the study. Three basic haplotypes were derived from the combined genotype data, haplotypes N, G and 11. The latter haplotype in which the whole distal copy of the kappa locus is missing was found three times among the 46 haploid genomes studied. The genotypes of the family members of an individual who is homozygous for haplotype 11 are consistent with Mendelian inheritance. Haplotypes N and G are distinguished from each other by eight RFLP markers. Six additional haplotypes, which were found in one or several individuals each, can be derived from the basic haplotypes N and G by hypothetical recombination and/or mutation events.

Chromosome Mapping↗

Analysis of 16S rRNA gene sequences of Vibrio costicola strains: description of Salinivibrio costicola gen. nov., comb. nov.

The phylogenetic positions of six Vibrio costicola strains were determined by direct sequencing and analysis of their PCR-amplified 16S ribosomal DNAs. A comparative analysis of the sequence data revealed that the moderate halophile V. costicola forms a monophyletic branch that is distinct from other Vibrio species and from moderately halophilic species of other genera. These results complement phenotypic and genotypic data determined previously. The molecular evidence, together with several phenotypic differences, distinguishes V. costicola from species of the genus Vibrio and other species belonging to the gamma subclass of the Proteobacteria and indicates that V. costicola should be placed in a new and separate genus. The name Salinivibrio costicola gen. nov., comb. nov. is proposed for this bacterium. The guanine-plus-cytosine content of the DNA is 49.4 to 50.5 mol%. The type strain of S. costicola is strain NCIMB 701 (= ATCC 33508).

Base Sequence↗

HLA-DQB1, -DQA1, -DRB1 linkage disequilibrium and haplotype diversity in a Mestizo population from Guadalajara, Mexico.

HLA-DQB1, -DQA1, and -DRB1 genes were typed by polymerase chain reaction with sequence-specific primer (PCR-SSP) in 159 healthy volunteers from 32 families living in Guadalajara, Mexico. Three-locus genotype data from all family members were used to infer haplotypes in 54 unrelated individuals of the sample, from which estimate of segregating haplotype frequencies and linkage disequilibrium (LD) between loci were computed. Genotype distributions were concordant with Hardy-Weinberg expectations (HWE) for all three loci, and allele distributions were similar to the ones observed in other Latin-American populations. Of the 56 distinct three-site (DQB1-DQA1-DRB1) haplotypes observed in the sample, the five most common (i.e., with frequencies of five counts or more) were: *0302-*0301-*04, *0201-*0201-*07, *0301-*0501-*14, *0402-*0401-*08, and *0501-*0101-*01. These common three-locus haplotypes also contributed to the majority of the significant two-locus linkage disequilibria of these three sites.

HLA-DQ Antigens↗

Pedigree and genotype errors in the Framingham Heart Study.

The pedigree and genotype data from the Framingham Heart Study were examined for errors. Errors in 21 of 329 pedigrees were detected with the program PREST, and of these the errors in 16 pedigrees were resolved. Genotyping errors were then detected with SIMWALK2. Five Mendelian errors were found following the pedigree corrections. Double-recombinant errors were more common, with 142 being detected at mistyping probabilities of 0.25 or greater.

Adult Children↗

Interleukin-10 gene promoter polymorphism in English and Polish healthy controls. Polymerase chain reaction haplotyping using 3' mismatches in forward and reverse primers.

A polymerase chain reaction with sequence-specific primers (PCR-SSP) system using primers with mismatches at the 3' ends was developed to determine polymorphisms in IL-10 promoter region. Three previously described biallelic polymorphisms in IL-10 were linked in a 12 reaction PCR-SSP system and the method used to provide genotype data on 233 UK and 166 Polish controls. There are eight possible polymorphic combinations in IL-10 promoter gene but only three were observed in both control groups. Population frequencies of IL-10 genotypes show, in contrast to HLA, that UK and Polish frequencies are remarkably similar.

Alleles↗

Multicentre evaluation of the VITEK 2 Advanced Expert System for interpretive reading of antimicrobial resistance tests.

Interpretive reading analyses the complete resistance profiles of bacteria to multiple antibiotics and infers the resistance mechanisms present; it aids therapeutic choice and enhances surveillance data. We evaluated the Advanced Expert System (AES), which interprets MICs generated by the VITEK 2. Ten European laboratories tested 42 reference strains and 76-106 of their own strains, representing important resistance genotypes. Interpretive reading by the VITEK 2 AES achieved full agreement with genotype data for 88-89% of strains, with the correct mechanism identified as one of two possibilities for a further 5-6%. Mechanisms inferred with 90% agreement with reference data included methicillin resistance in staphylococci, glycopeptide resistance in enterococci, quinolone resistance in staphylococci and Enterobacteriaceae, AAC(6')-APH(2")-mediated aminoglycoside resistance in Gram-positive cocci, erm-mediated macrolide resistance in pneumococci, extended-spectrum beta-lactamases (ESBLs) in Enterobacteriaceae and Pseudomonas aeruginosa, and acquired penicillinases in Enterobacteriaceae. VanA, VanB and VanC phenotypes of enterococci were distinguished reliably, and ESBL production was accurately inferred in AmpC-inducible species as well as Escherichia coli and Klebsiella spp. Mechanisms identified, but only as possibilities among several, included IRT-type beta-lactamases and individual aminoglycoside-modifying enzymes in Enterobacteriaceae. Most disagreements with reference data concerned pneumococci found to have high-level penicillin resistance by the VITEK 2 AES but previously determined, phenotypically, to have intermediate resistance. When ESBL production was inferred in E. coli and klebsiellae, the VITEK 2 AES edited susceptible results for cephalosporins (except cefoxitin) to resistant; when an acquired penicillinase was inferred in Enterobacteriaceae, piperacillin results were edited to resistant; and when staphylococci were found methicillin resistant, resistance was reported for all beta-lactams. Further editing may be desirable (e.g. of cephalosporin results for salmonellas inferred to have ESBLs).

Data Interpretation, Statistical↗

Approaches to identifying genetic predictors of clinical outcome in rheumatoid arthritis.

Predicting which patients with rheumatoid arthritis (RA), at presentation, are likely to suffer a severe disease course based on genotype data would be a major clinical advance. It would ensure that patients at highest risk of a severe outcome could be targeted with early aggressive therapies. With a better understanding of interactions between genotype and drug response it would be possible to prescribe treatments most likely to be efficacious and safe for specific patient subgroups. While a clear genetic component has been demonstrated in RA severity, the identification of genetic factors poses a challenge to researchers in the field. Initiatives such as the SNP Consortium and advances in genotyping technology have facilitated the investigation of genetic factors in both disease susceptibility and severity. However, several other factors, such as the availability of suitable longitudinal cohorts, definition of outcome measures, study design, selection of genetic markers, and statistical power, will all contribute to the likely success of genetic studies. Several strategies that have been applied in the pursuit of genetic predictors of clinical outcome in RA. While some encouraging results have been generated, it has so far been difficult to quantify the predictive value of genetic markers and extrapolate the results from genetic studies to clinic patients. Establishing high quality prospective inception cohorts, a more systemic approach to defining suitable outcome measures, and understanding the effects of treatment, will be critical to the eventual identification of good predictive genetic markers.

Animals↗

A microsatellite-based multipoint index map of human chromosome 22.

Utilizing the CEPH (Centre d'Etude du Polymorphism Humain) reference panel and genotyping data for 24 simple tandem repeat polymorphism (STRP) markers, we have constructed a 15-locus multipoint genetic framework map of human chromosome 22. The markers form a continuous linkage group of 51 cM in males and 81 cM in females. Likely genetic locations are provided for 9 additional STRP sequences. The map was constructed employing the CRIMAP computational methodology to build the multipoint map via a stepwise algorithm. The quality of the framework map was evaluated using a battery of statistical diagnostics that suggest a typing error frequency of 0.1% for markers within the map.

Base Sequence↗

Two sex-chromosome-linked microsatellite loci show geographic variance among North American Ostrinia nubilalis.

PCR-based O. nubilalis population and pedigree analysis indicated female specificity of a (GAAAAT)n microsatellite, and male specificity of a CAYCARCGTCACTAA repeat unit marker. These loci were respectively named Ostrinia nubilalis W-chromosome 1 (ONW1) and O. nubilalis Z-chromosome 1 (ONZ1). Intact repeats of three, four, or five GAAAAT units are present among ONW1 alleles, and biallelic variation exists at the ONZ1 locus. Screening of 493 male at ONZ1 and 448 heterogametic females at ONZ1 and ONW1 loci from eleven North American sample sites was used to construct genotypic data. Analysis of molecular variance (AMOVA) and F-statistics indicated no female haplotype or male ONZ1 allele frequency differentiation between voltinism ecotypes. Four subpopulations from northern latitudes, Minnesota and South Dakota, showed the absence of a single female haplotype, a significant deviation of ONZ1 data from Hardy-Weinberg expectation, and low-level geographic divergence from other subpopulations. Low ONZ1 and ONW1 allele diversity could be attributed either to large repeat unit sizes, low repeat number, reduced effective population (Ne) size of sex chromosomes, or the result of recent O. nubilalis introduction and population expansion, but likely could not be due to inbreeding.

Animals↗

Variability of the PreS1/PreS2/S regions of hepatitis B virus in Hungary.

Infection with the hepatitis B virus can occur perinatally, parenterally, or sexually, and it can cause acute or chronic liver diseases. Phylogenetic analysis of the virus has led to its classification into eight genotypes (A-H), which show a characteristic worldwide distribution. The aim of this study was to reveal the HBV genotypes present in Hungary and to investigate a nosocomial and an intrafamilial outbreak. The collected samples were tested by nested PCR, and a 650-nucleotide-long segment of the preS1/preS2/S region was sequenced. As no previous genotype data were available from Hungary, sera of 24 HBsAg-positive patients were collected from different regions of the country. They also served as control samples for the molecular epidemiologic study. Nineteen of them carried genotype D of hepatitis B virus, and five of them carried genotype A. Twenty-nine patients from a haemato-oncology unit were affected in a nosocomial outbreak. The patients had haematological and/or oncological diseases, most of them were immunosuppressed. In twenty-eight cases, based on phylogenetic analysis of the viruses, there was presumably a common source of infection, and an epidemiological investigation showed that the infections seemed to be hospital-acquired. In the intrafamilial outbreak, two asymptomatic carrier children infected their foster mother. The three sequences were totally identical.

Amino Acid Sequence↗

Genetic association between Ubiquitin Carboxy-terminal Hydrolase-L1 gene S18Y polymorphism and sporadic Alzheimer's disease in a Chinese Han population.

Increasing evidence indicates that the dysfunction of ubiquitin-proteasome system (UPS) is associated with Alzheimer's disease (AD). In the ubiquitin-proteasome pathway, Ubiquitin Carboxy-terminal Hydrolase-L1 (UCH-L1) plays an important role for the cellular clearance of abnormal proteins. Since a substitution of serine by tyrosine at codon 18, exon 3 (S18Y polymorphism) of the UCH-L1 gene exhibits a protective effect against the development of degenerative disease such as sporadic Parkinson's disease (PD) in several different ethnic groups, we hypothesized that UCH-L1 gene S18Y polymorphism may have that same effect on the pathologic process of AD. We examined UCH-L1 S18Y polymorphism genotypes of 116 sporadic AD patients and 123 healthy subjects in Chinese Han population using PCR-restriction fragment length polymorphism (RFLP) analysis. The allele and genotype data as well as data after stratification by age of onset failed to demonstrate any association between AD and S18Y polymorphism. However, after stratification by gender, female AD patients showed significantly less frequencies of Y allele and YY genotype in S18Y polymorphism than female controls (P = 0.003 and P = 0.015 respectively). We conclude that Y allele and YY genotype of S18Y in the UCH-L1 gene may have a protective effect against sporadic AD in female subjects, probably due to altering the function of UCH-L1 and the interactions among different risk factors.

Adult↗

Sparse polygenic risk score inference with the spike-and-slab LASSO.

MOTIVATION: Large-scale biobanks, with rich phenotypic and genomic data across hundreds of thousands of samples, provide ample opportunities to elucidate the genetics of complex traits and diseases. Consequently, there is growing demand for robust and scalable methods for disease risk prediction from genotype data. Inference in this setting is challenging due to the high-dimensionality of genomic data, especially when coupled with smaller sample sizes. Popular Polygenic Risk Score (PRS) inference methods address this challenge by adopting sparse Bayesian priors or penalized regression techniques, such as the Least Absolute Shrinkage and Selection Operator (LASSO). However, the former class of methods are not as scalable and do not produce exact sparsity, while the latter tends to over-shrink large coefficients. RESULTS: In this study, we present SSLPRS, a novel PRS method based on the Spike-and-Slab LASSO (SSL) prior, which offers a theoretical bridge between the two frameworks. We extend previous work to derive a coordinate-ascent inference algorithm that operates on GWAS summary statistics, which is orders-of-magnitude more efficient than corresponding individual-level-based implementations. To illustrate the statistical properties of the proposed model, we conducted experiments involving nine simulation configurations and nine quantitative phenotypes from the UK Biobank. Our results demonstrate that SSLPRS is competitive with state-of-the-art methods in terms of prediction accuracy and exhibits superior variable selection performance, especially in sparse genetic architectures. In simulations, this translates to upwards of 50% improvement in positive predictive value. In analysis of real phenotypes, we show that selected variants are highly enriched for meaningful genomic annotations and have better replication rates in larger meta-analyses. AVAILABILITY AND IMPLEMENTATION: SSLPRS is available in the open-source package https://github.com/li-lab-mcgill/penprs.

Multifactorial Inheritance↗

Polymorphisms of folate metabolic genes and susceptibility to bladder cancer: a case-control study.

Epidemiological studies have shown an association between low folate intake and an increased cancer risk. Major genes involved in folate metabolism include methylene-tetrahydrofolate reductase (MTHFR) and methionine synthase (MS). We investigated joint effects of polymorphisms of the MTHFR (677 C-->T, 1298A-->C) and MS genes (2756 A-->G), dietary folate intake and cigarette smoking on the risk of bladder cancer in a case-control study. The study population consisted of 457 bladder cancer patients and 457 healthy controls, matched to the cases in terms of age, gender and ethnicity. Genotype data were analyzed in a subset of 410 Caucasian cases and 410 controls. Compared with individuals carrying the MTHFR 677 wild-type (CC) and reporting a high folate intake, those carrying the variant genotype (CT or TT) and reporting a low folate intake were at a significantly 3.51-fold increased risk of bladder cancer (95% CI: 1.59-6.52). In contrast, individuals carrying a variant genotype and reporting a high folate intake were at only a 1.39-fold increased risk (95% CI: 0.71-2.70), and those carrying the wild-type and reporting a low folate intake were at only 1.56-fold increased risk (95% CI: 0.82-2.97). The interaction between genetic polymorphisms and folate intake was significant on the multiplicative scale (P = 0.01). When analyzed in the context of smoking status, compared with never smokers with the MTHFR 677 wild-type, the risk increased to 6.56-fold (95% CI: 3.28-13.12) in current smokers carrying the variant genotype. Analyses of the MTHFR 1298, MS 2756 genes revealed similar results. In addition, age at cancer onset in former smokers increased as the proportion of the heteromorphic haplotype in the individual increased (P = 0.005). Our results strongly suggest that polymorphisms of the MTHFR and MS genes act together with low folate intake and smoking to increase bladder cancer risk. These results have important implications for cancer prevention in susceptible populations.

5-Methyltetrahydrofolate-Homocysteine S-Methyltran↗