Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genotype Data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Location on the human genetic linkage map of 26 genes involved in blood coagulation.

Several human genetic linkage maps have been constructed as part of the Human Genome Project. These maps show the positional order of closely linked, highly informative AC-repeat polymorphisms on each human chromosome, and are extremely useful in genetic linkage analysis of inheritable diseases. For a candidate gene approach the current linkage maps are less useful, since they consist mainly of anonymous markers rather than of specific genes. This situation also applies for inheritable disorders of blood coagulation. Numerous genes are involved in the blood coagulation cascade and its regulation, and can be considered as candidate genes for unexplained haemophilia and thrombophilia. We have selected 29 candidate genes that seem to be the ones most likely to be involved in thrombophilia. For 19 genes genotype data were already present in the CEPH database (version 7.0). We typed 7 additional genes in the CEPH reference families, i.e. the factor V, factor XII, protein C, protein S, prothrombin, thrombomodulin, and heparin cofactor II gene. The genotype data were used to integrate these 26 genes in the current genetic linkage map, and to identify closely linked AC-repeat polymorphisms. This information will benefit the investigation of inheritable disorders of blood coagulation, especially thrombophilia.

Base Sequence↗

Diagnosis and genotyping of Helicobacter pylori by polymerase chain reaction of bacterial DNA from gastric juice.

BACKGROUND: Efficient and accurate detection of Helicobacter pylori infection as well as identification of virulence-associated alleles are important for the treatment of gastroduodenal diseases caused by this gastric pathogen. The present study was performed to test the efficiency of gastric juice polymerase chain reaction (PCR) method for the rapid detection of H. pylori infection and to determine the bacterial genotypes without the need for culture, which is often not feasible especially in developing countries. METHODS: DNA was extracted from gastric juice samples collected from 45 subjects and was used to amplify urease B gene (ureB) for H. pylori. Results obtained from this method were further confirmed by rapid urease test (RUT), histology and culture. Genotypes of the infected strains predicted from gastric juice PCR were compared to the genotype data obtained from the isolated strains. RESULTS: Among 45 cases, 32 were positive by RUT, 37 by histological examination, 25 by gastric juice PCR method, while culture yielded positive results for 19 samples. Except for one case, all the 19 culture-positive strains gave the same genotype with the gastric juice PCR result. It was found that the gastric juice PCR is more efficient for detection of multiple-strain infection as compared to genotype data obtained from strains isolated as pooled culture. CONCLUSIONS: This moderately sensitive technique could be employed with good efficiency, particularly in cases where it is difficult to obtain biopsy. Moreover, with this method bacterial genotype could be obtained.

Bangladesh↗

Fine mapping--19th century style.

BACKGROUND: There is great interest in the use of computationally intensive methods for fine mapping of marker data. In this paper we develop methods based upon ideas originally proposed 100 years ago in the context of spatial clustering. METHODS: We use spatial clustering of haplotypes as a low-dimensional surrogate for the unobserved genealogy underlying a set of genotype data. In doing so we hope to avoid the computational complexity inherent in explicitly modelling details of the ancestry of the sample, while at the same time capturing the key correlations induced by that ancestry at a much lower computational cost. RESULTS: We benchmark our methods using the simulated Genetic Analysis Workshop 14 data, using 100 replicates of 4 phenotypes to indicate the power of our method. When a functional mutation relating to a trait is actually present, we find evidence for that mutation in 97 out of 100 replicates, on average. CONCLUSION: Our results show that our method has the ability to accurately infer the location of functional mutations from unphased genotype data.

Congresses as Topic↗

Rhizobium gallicum sp. nov. and Rhizobium giardinii sp. nov., from Phaseolus vulgaris nodules.

Thirty-one strains of two new genomic species (genomic species 1 and 2) of rhizobia isolated from root nodules of Phaseolus vulgaris and originating from various locations in France were compared with reference strains of rhizobia by performing a numerical analysis of 64 phenotypic features. Each genomic species formed a distinct phenon and was separated from the other rhizobial species. A comparison of the complete 16S rRNA gene sequences of a representative of genomic species 1 (strain R602spT) and a representative of genomic species 2 (strain H152T) with the sequences of other rhizobia and related bacteria revealed that each genomic species formed a lineage independent of the lineages formed by the previously recognized species of rhizobia. Genomic species 1 clustered with the species that include the bean-nodulating rhizobia, Rhizobium leguminosarum, Rhizobium etli, and Rhizobium tropici, and branched with unclassified rhizobial strain OK50, which was isolated from root nodules of Pterocarpus klemmei in Japan. Genomic species 2 was distantly related to all other Rhizobium species and related taxa, and the most closely related organisms were Rhizobium galegae and several Agrobacterium species. On the basis of the results of phenotypic and phylogenetic analyses and genotypic data previously published and reviewed in this paper, two new species of the genus Rhizobium, Rhizobium gallicum and Rhizobium giardinii, are proposed for genomic species 1 and 2, respectively. Each species could be divided in two subgroups on the basis of symbiotic characteristics, as shown by phenotypic (host range and nitrogen fixation effectiveness) and genotypic data. For each species, one subgroup had the same symbiotic characteristics as R. leguminosarum biovar phaseoli and R. etli biovar phaseoli. The other subgroup had a species-specific symbiotic phenotype and genotype. Therefore, we propose that each species should be subdivided into two biovars, as follows: R. gallicum biovar gallicum and R. gallicum biovar phaseoli; and R. giardinii biovar giardinii and R. giardinii biovar phaseoli.

Colony Count, Microbial↗

Cladistic analysis: its applications in association studies of complex diseases.

INTRODUCTION: With the increase in genotype data generated by high throughput typing technologies, there is currently a lack of complexity-oriented analytical methods that can maximise the information obtained from these raw data for the study of complex diseases. We introduce the cladistic analysis that is traditionally applied in evolution studies and taxonomy, to specify relevant comparisons of traits associated with each haplotype/genotype in a population sample. METHODS: Haplotypes are determined from the genotype data and linked to each other by their evolutionary relationships to form a cladogram. This is then used to specify relevant statistical comparisons. The central assumption is that any functionally important genetic variation causing a phenotypic effect at any point in the course of evolution will be embedded in the framework of haplotypes represented by the cladogram. APPLICATIONS: There are various applications of cladistic analysis in the study of complex diseases. Basically, it helps in the identification of haplotypes that are associated with a disease state or a significantly altered level of quantitative trait. However, its limitations are that only polymorphic sites on the same DNA strand can be analysed and that recombination events must be relatively rarer than mutational events. CONCLUSIONS: In the absence of methods that can recognise the complexity of the genotype organization, and given its ability to exploit evolutionary information for optimising the analytical strategy, cladistic analysis would be a method of choice for studying multi-loci effects on a quantitative trait or disease outcome.

Genetic Variation↗

A simulated annealing algorithm for maximum likelihood pedigree reconstruction.

The calculation of maximum likelihood pedigrees for related organisms using genotypic data is considered. The problem is formulated so that the domain of optimization is a permutation space. This is a feature shared by the travelling salesman problem, for which simulated annealing is known to be effective. Using this technique it is found that pedigrees can be reconstructed with minimal error using genotypic data of a quality currently realizable. In complex pedigrees accurate reconstruction can be done with no a priori age or sex information. For smaller numbers of individuals a method of efficiently enumerating all admissible pedigrees of nonzero likelihood is given.

Algorithms↗

Genome-wide cis-expression Quantitative Trait Loci (eQTL) and transcriptomic signals reveal distinct molecular regulation across correlated feed efficiency traits.

INTRODUCTION: Feed efficiency (FE) is a complex trait which determines livestock production profitability, yet the molecular mechanisms behind it remain unclear. This study investigated the blood transcriptomic profile of lambs, alongside genotype data with the aim to uncover the genetic basis of FE traits such as absolute dry matter intake (DMIabsolute), DMI adjusted for body size (DMIadjusted), average daily live weight gain (ADG), and residual feed intake (RFI). MATERIALS AND METHODS: Bulk RNA-Seq and genotype data were analysed using three complementary approaches: differential gene expression (DGE) analysis, weighted gene co-expression network analysis (WGCNA), and cis-expression Quantitative Trait Loci (cis-eQTL) mapping. These methods were used independently to identify genes and regulatory networks associated with FE traits and to investigate evidence supporting multi-trait candidate gene selection. RESULTS: DGE analysis revealed 2, 24, 85 and 4 differentially expressed genes for DMIabsolute, DMIadjusted, ADG, and RFI (Padjusted < 0.05), functionally enriched in sensory perception, ATP-dependent chromatin remodeling, Notch signaling and immune response pathways. 9 gene modules significantly associated with the FE traits (P &#x2264; 0.05) with correlations ranging from r = -0.56 to 0.49, were identified using WGCNA. Single nucleotide polymorphism (SNP)-level cis-eQTL analysis identified 93 eSNPs associated with 74 genes (false discovery rate (FDR) < 0.05), while permutation-derived gene level analysis identified 280 eGenes (FDR < 0.2, empirical P < 0.03). Across the three analyses, applying thresholds of DGE (Padjusted < 0.05), WGCNA (correlation, P &#x2264; 0.05), and cis-eQTL gene-level significance (empirical P < 0.05), multiple overlapping genes were identified including DNMT3A, KANSL1, NCOR1 for DMIadjusted, ACOX2, FANCF, CIMIP2B, LOC101115106, ARMH2, LOC132657496 for ADG, and LOC114114576 for RFI representing regulators of variations in FE. DISCUSSION: The integration of DGE, WGCNA, and cis-eQTL analyses identified key genes and regulatory mechanisms associated with variation in FE traits. These results highlight that integrated multi-trait candidate gene identification approaches can reveal key genes that lower feed intake while maintaining animal growth, supporting breeding strategies aimed at improving efficiency and long-term economic sustainability in sheep.

average daily gain (ADG)↗

Quantitative-trait-locus analysis of body-mass index and of stature, by combined analysis of genome scans of five Finnish study groups.

In recent years, many genomewide screens have been performed, to identify novel loci predisposing to various complex diseases. Often, only a portion of the collected clinical data from the study subjects is used in the actual analysis of the trait, and much of the phenotypic data is ignored. With proper consent, these data could subsequently be used in studies of common quantitative traits influencing human biology, and such a reanalysis method would be further justified by the nonbiased ascertainment of study individuals. To make our point, we report here a quantitative-trait-locus (QTL) analysis of body-mass index (BMI) and stature (i.e., height), with genotypic data from genome scans of five Finnish study groups. The combined study group was composed of 614 individuals from 247 families. Five study groups were originally ascertained in genetic studies on hypertension, obesity, osteoarthritis, migraine, and familial combined hyperlipidemia. Most of the families are from the Finnish Twin Cohort, which represents a population-wide sample. In each of the five genome scans, approximately 350 evenly spaced markers were genotyped on 22 autosomes. In analyzing the genotype data by a variance-component method, we found, on chromosome 7pter (maximum multipoint LOD score of 2.91), evidence for QTLs affecting stature, and a second locus, with suggestive evidence for linkage to stature, was detected on chromosome 9q (maximum multipoint LOD score of 2.61). Encouragingly, the locus on chromosome 7 is supported by the data reported by Hirschhorn et al. (in this issue), who used a similar method. We found no evidence for QTLs affecting BMI.

Body Height↗

Estimation of sib-pair IBD sharing and multipoint polymorphism information content by linear regression.

A simple method of estimating IBD sharing (pi) of sibling pairs from multipoint genotype data based on linear regression is presented. The new method is an extension of that developed by Fulker, Cherny, and Cardon (1995) and involves a measure of sib-pair specific information, W, defined as the difference between the unconditional variance of pi and the conditional variance of pi given the marker genotype data. When markers are not fully informative, the W method provides estimates closer to those obtained from the exact hidden Markov method (HMM) than the original Fulker method. Using W, we also derive a generalisation for the polymorphism information content (PIC) for multiple markers. This multipoint polymorphism information content (MPIC) can be evaluated at any location along a marker map and is proportional to the noncentrality parameter for linkage. We illustrate MPIC by assessing the relative information content of single nucleotide polymorphism (SNP) and microsatellite markers.

Alleles↗

Treatment switch guided by HIV-1 genotyping in Brazil.

We assessed the performance of HIV-1 genotyping tests in rescue therapy. Patients were divided into two groups: group 1 (genotyped), included those switching to new antiretroviral drugs based on HIV-1 genotyping data, and group 2 (standard of care -SOC), comprised those in rescue therapy who had not used this test. This was an open and non-randomized study, with 74 patients, followed up for a mean period of 12 months, from February 2002 to May 2003. The groups differed in the duration of antiretroviral use, experience with diverse drug classes (non-nucleoside reverse transcriptase inhibitors and protease inhibitors) and viral load <2.6 log10 copies/mL at any time during treatment. In 23 patients (group 1), the switch in antiretroviral (ARV) regimen was based on genotyping data; this test was not used for 51 patients (group 2). Two CD4 + lymphocyte counts and viral load counts were made for each patient during the study. Data from the pharmacy where patients received antiretroviral agents, medical charts, and direct interviews with patients to assess compliance to treatment, were analyzed. In the genotyped group, the average drop in viral load was 2.8 log10, compared with a 1.5 log10 difference in group 2; the difference was significant in the first assessment performed six months after switching (p=0.001). Considering the patients with viral load < 2.6 log10 (400 copies/mL) after switching, the patients in group 1 had a better performance in the first assessment (73.9% versus 31.1% in groups 1 and 2, respectively); this difference was significant (p=0.001). In multivariate analysis, the variables associated with a greater drop in viral load in the first assessment were the patients whose switching was based on genotyping (group 1), those with a past history of viral load < 2.6 log10 and correct use of antiretroviral agents. In conclusion, the genotyping test and adherence were found to be independent factors for success in the management of patients who failed treatment.

Adult↗

An integrated genetic and physical map of the bovine X chromosome.

Genotypic data for 56 microsatellites (ms) generated from maternal full sib families nested within paternal half sib pedigrees were used to construct a linkage map of the bovine X Chromosome (Chr) (BTX) that spans 150 cM (ave. interval 2.7 cM). The linkage map contains 36 previously unlinked ms; seven generated from a BTXp library. Genotypic data from these 36 ms was merged into an existing linkage map to more than double the number of informative BTX markers. A male specific linkage map of the pseudoautosomal region was also constructed from five ms at the distal end of BTXq. Four informative probes physically assigned by fluorescence in situ hybridization defined the extent of coverage, confirmed the position of the pseudoautosomal region on the q-arm, and identified a 4.1-cM marker interval containing the centromere of BTX.

Animals↗

Multinomial logistic regression approach to haplotype association analysis in population-based case-control studies.

BACKGROUND: The genetic association analysis using haplotypes as basic genetic units is anticipated to be a powerful strategy towards the discovery of genes predisposing human complex diseases. In particular, the increasing availability of high-resolution genetic markers such as the single-nucleotide polymorphisms (SNPs) has made haplotype-based association analysis an attractive alternative to single marker analysis. RESULTS: We consider haplotype association analysis under the population-based case-control study design. A multinomial logistic model is proposed for haplotype analysis with unphased genotype data, which can be decomposed into a prospective logistic model for disease risk as well as a model for the haplotype-pair distribution in the control population. Environmental factors can be readily incorporated and hence the haplotype-environment interaction can be assessed in the proposed model. The maximum likelihood estimation with unphased genotype data can be conveniently implemented in the proposed model by applying the EM algorithm to a prospective multinomial logistic regression model and ignoring the case-control design. We apply the proposed method to the hypertriglyceridemia study and identifies 3 haplotypes in the apolipoprotein A5 gene that are associated with increased risk for hypertriglyceridemia. A haplotype-age interaction effect is also identified. Simulation studies show that the proposed estimator has satisfactory finite-sample performances. CONCLUSION: Our results suggest that the proposed method can serve as a useful alternative to existing methods and a reliable tool for the case-control haplotype-based association analysis.

Algorithms↗

Optimal step length EM algorithm (OSLEM) for the estimation of haplotype frequency and its application in lipoprotein lipase genotyping.

BACKGROUND: Haplotype based linkage disequilibrium (LD) mapping has become a powerful and cost-effective method for performing genetic association studies, particularly in the search for genetic markers in linkage disequilibrium with complex disease loci. Various methods (e.g. Monte-Carlo (Gibbs sampling); EM (expectation maximization); and Clark's method) have been used to estimate haplotype frequencies from routine genotyping data. RESULTS: These algorithms can be very slow for large number of SNPs. In order to speed them up, we have developed a new algorithm using numerical analysis technology, a so-called optimal step length EM (OSLEM) that accelerates the calculation. By optimizing approximately the step length of the EM algorithm, OSLEM can run at about twice the speed of EM. This algorithm has been used for lipoprotein lipase (LPL) genotyping analysis. CONCLUSIONS: This new optimal step length EM (OSLEM) algorithm can accelerate the calculation for haplotype frequency estimation for genotyping data without pedigree information. An OSLEM on-line server is available, as well as a free downloadable version.

Algorithms↗

Genome-wide linkage disequilibrium and haplotype maps.

There is currently a broad effort to produce genome-wide high-density linkage disequilibrium (LD) maps with single nucleotide polymorphisms. The hope is that the resulting maps can be exploited to find genes that affect the onset and severity of at least some common human diseases. These maps may also be useful for identifying genes that affect drug response or the likelihood of drug toxicities. The goal of this review is to provide a broad overview of some of the key concerns motivating the design of a major international project called the International Haplotype Map Project. The process of map production requires the identification of very large numbers of polymorphic sites, implementation of facile, highly accurate and inexpensive genotyping production pipelines, and provision for public access to the genotype data. Great progress has been made recently in genotyping methods and these advances are allowing very large-scale data collection. A major goal of these efforts is to enable the selection of subsets of markers that capture useful genetic information in short genomic intervals, while optimally reducing the number of markers that must be genotyped. Standard measures of LD provide a starting point but may not fully capture the complexity of the information inherent in the data. Extremely dense genotype data in several broadly representative populations (European, Chinese, Japanese, and Yoruba) should yield important insights into the genetic structure of most genes. Further study is required to determine how broadly applicable the data will be to other population groups. Significant challenges lie ahead in determining the best methods for the selection of markers in disease/phenotype studies, large-scale genotyping, and analysis of the resulting genetic data.

Animals↗

Microsatellite null alleles and estimation of population differentiation.

Microsatellite null alleles are commonly encountered in population genetics studies, yet little is known about their impact on the estimation of population differentiation. Computer simulations based on the coalescent were used to investigate the evolutionary dynamics of null alleles, their impact on F(ST) and genetic distances, and the efficiency of estimators of null allele frequency. Further, we explored how the existing method for correcting genotype data for null alleles performed in estimating F(ST) and genetic distances, and we compared this method with a new method proposed here (for F(ST) only). Null alleles were likely to be encountered in populations with a large effective size, with an unusually high mutation rate in the flanking regions, and that have diverged from the population from which the cloned allele state was drawn and the primers designed. When populations were significantly differentiated, F(ST) and genetic distances were overestimated in the presence of null alleles. Frequency of null alleles was estimated precisely with the algorithm presented in Dempster et al. (1977). The conventional method for correcting genotype data for null alleles did not provide an accurate estimate of F(ST) and genetic distances. However, the use of the genetic distance of Cavalli-Sforza and Edwards (1967) corrected by the conventional method gave better estimates than those obtained without correction. F(ST) estimation from corrected genotype frequencies performed well when restricted to visible allele sizes. Both the proposed method and the traditional correction method have been implemented in a program that is available free of charge at http://www.montpellier.inra.fr/URLB/. We used 2 published microsatellite data sets based on original and redesigned pairs of primers to empirically confirm our simulation results.

Alleles↗

Continuous covariates in genetic association studies of case-parent triads: gene and gene-environment interaction effects, population stratification, and power analysis.

We propose a multinomial logistic regression method which permits estimation and likelihood ratio tests for allele effects, their interactions with continuous covariates, and assessment of the degree of population stratification in genetic association studies of case-parent triads. Our approach overcomes the constraint imposed by the categorical nature of explanatory variables in the log-linear model. We also demonstrate that the multinomial logistic method can yield efficient inference in the presence of missing parental genotype data via the use of the Expectation-Maximization (EM) algorithm. We performed simulations to compare the multinomial logistic model with the case-pseudosibling conditional logistic model approach, both of which permit the incorporation of continuous covariates. Simulation results indicate that the multinomial logistic model and the conditional logistic model lead to similar estimates in large samples. A simulation-based method of sample size estimation is also used to show that the two models are approximately equivalent in sample size requirements. When parental genotype data are missing, either completely at random or dependent on covariates, the use of the EM algorithm gives multinomial logistic model greater power. Since the multinomial logistic model offers the possibility of assessing the degree of population stratification in the sample and can also provide efficient inference in the presence of missing parental genotypes, the proposed model has an important application in epidemiological family-based association studies.

Journal Article↗

Construction of the model for the Genetic Analysis Workshop 14 simulated data: genotype-phenotype relationships, gene interaction, linkage, association, disequilibrium, and ascertainment effects for a complex phenotype.

The Genetic Analysis Workshop 14 simulated dataset was designed 1) To test the ability to find genes related to a complex disease (such as alcoholism). Such a disease may be given a variety of definitions by different investigators, have associated endophenotypes that are common in the general population, and is likely to be not one disease but a heterogeneous collection of clinically similar, but genetically distinct, entities. 2) To observe the effect on genetic analysis and gene discovery of a complex set of gene x gene interactions. 3) To allow comparison of microsatellite vs. large-scale single-nucleotide polymorphism (SNP) data. 4) To allow testing of association to identify the disease gene and the effect of moderate marker x marker linkage disequilibrium. 5) To observe the effect of different ascertainment/disease definition schemes on the analysis. Data was distributed in two forms. Data distributed to participants contained about 1,000 SNPs and 400 microsatellite markers. Internet-obtainable data consisted of a finer 10,000 SNP map, which also contained data on controls. While disease characteristics and parameters were constant, four "studies" used varying ascertainment schemes based on differing beliefs about disease characteristics. One of the studies contained multiplex two- and three-generation pedigrees with at least four affected members. The simulated disease was a psychiatric condition with many associated behaviors (endophenotypes), almost all of which were genetic in origin. The underlying disease model contained four major genes and two modifier genes. The four major genes interacted with each other to produce three different phenotypes, which were themselves heterogeneous. The population parameters were calibrated so that the major genes could be discovered by linkage analysis in most datasets. The association evidence was more difficult to calibrate but was designed to find statistically significant association in 50% of datasets. We also simulated some marker x marker linkage disequilibrium around some of the genes and also in areas without disease genes. We tried two different methods to simulate the linkage disequilibrium.

Computer Simulation↗

Emergence of drug resistance mutations in a group of HIV-infected children taking nelfinavir-containing regimens.

HIV-1-infected children are often treated with therapy regimens including protease inhibitors (PIs). We monitored the virologic response in a small group of pediatric patients undergoing therapy with regimens including the PI nelfinavir and determined whether new drug resistance mutations were present immediately after virologic failure. Seventeen reverse transcriptase inhibitor (RTI)-experienced children starting nelfinavir-containing therapy regimens were studied. After virologic failure, HIV-1 protease (PR) and RT sequences were examined for drug resistance mutations. Viral load levels decreased to <400 HIV RNA copies/ml in six patients and remained at <400 HIV RNA copies/ml in four patients. Three patients did not respond virologically; all three had mutations specific for one or more of their regimen drugs either before or soon after nelfinavir initiation. The virologic response was transient in eight patients whose viral loads did not decrease to <400 HIV RNA copies/ml. Genotypic data from seven of the eight patients revealed mutations specific for one or more of their regimen drugs after virologic rebound. PI resistance mutations occurred in eight patients: D30N in six, and L90M in three. In three patients, the only new mutation after failure was the RT mutation M184V. Despite virologic failure, sustained increases in CD4+ lymphocyte counts were noted in eight patients. We conclude that in this small group of pediatric patients, virologic failure occurred in all patients whose viral loads did not become undetectable after the switch to a nelfinavir-containing regimen. After failure, new drug resistance mutations were found in either PR or RT. Studies of larger cohorts are warranted to determine whether HIV-1 genotypic data can help in the formulation of effective salvage therapies in children.

CD4 Lymphocyte Count↗