Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical genetics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Some statistical and regulatory issues in the evaluation of genetic and genomic tests.

The genomics revolution is reverberating throughout the worlds of pharmaceutical drugs, genetic testing and statistical science. This revolution, which uses single nucleotide polymorphisms (SNPs) and gene expression technology, including cDNA and oligonucleotide microarrays, for a range of tests from home-brews to high-complexity lab kits, can allow the selection or exclusion of patients for therapy (responders or poor metabolizers). The wide variety of US regulatory mechanisms for these tests is discussed. Clinical studies to evaluate the performance of such tests need to follow statistical principles for sound diagnostic test design. Statistical methodology to evaluate such studies can be wide ranging, including receiver operating characteristic (ROC) methodology, logistic regression, discriminant analysis, multiple comparison procedures resampling, Bayesian hierarchical modeling, recursive partitioning, as well as exploratory techniques such as data mining. Recent examples of approved genetic tests are discussed.

Animals↗

The power and statistical behaviour of allele-sharing statistics when applied to models with two disease loci.

We have evaluated the power for detecting a common trait determined by two loci, using seven statistics, of which five are implemented in the computer program SimWalk2, and two are implemented in GENEHUNTER. Unlike most previous reports which involve evaluations of the power of allele-sharing statistics for a single disease locus, we have used a simulated data set of general pedigrees in which a two-locus disease is segregating and evaluated several nonparametric linkage statistics implemented in the two programs. We found that the power for detecting linkage using the S(all) statistic in GENEHUNTER (GH, version 2.1), implemented as statistic E in SimWalk2 (version 2.82), is different in the two. The P values associated with statistic E output by SimWalk2 are consistently more conservative than those from GENEHUNTER except when the underlying model includes heterogeneity at a level of 50% where the P values output are very comparable. On the other hand, when the thresholds are determined empirically under the null hypothesis, S(all) in GENEHUNTER and statistic E have similar power.

Alleles↗

Determining the sample size for co-dominant molecular marker-assisted linkage detection for a monogenic qualitative trait by controlling the type-I and type-II errors in a segregating F2 population.

Tests for linkage are usually performed using the lod score method. A critical question in linkage analyses is the choice of sample size. The appropriate sample size depends on the desired type-I error and power of the test. This paper investigates the exact type-I error and power of the lod score method in a segregating F(2) population with co-dominant markers and a qualitative monogenic dominant-recessive trait. For illustration, a disease-resistance trait is considered, where the susceptible allele is recessive. A procedure is suggested for finding the appropriate sample size. It is shown that recessive plants have about twice the information content of dominant plants, so the former should be preferred for linkage detection. In some cases the exact alpha-values for a given nominal alpha may be rather small due to the discrete nature of the sampling distribution in small samples. We show that a gain in power is possible by using exact methods.

Crosses, Genetic↗

The population genetics of sporophytic self-incompatibility in Senecio squalidus L. (Asteraceae) I: S allele diversity in a natural population.

Twenty-six individuals of the sporophytic self-incompatible (SSI) weed, Senecio squalidus were crossed in a full diallel to determine the number and frequency of S alleles in an Oxford population. Incompatibility phenotypes were determined by fruit-set results and the mating patterns observed fitted a SSI model that allowed us to identify six S alleles. Standard population S allele number estimators were modified to deal with S allele data from a species with SSI. These modified estimators predicted a total number of approximately six S alleles for the entire Oxford population of S. squalidus. This estimate of S allele number is low compared to other estimates of S allele diversity in species with SSI. Low S allele diversity in S. squalidus is expected to have arisen as a consequence of a disturbed population history since its introduction and subsequent colonisation of the British Isles. Other features of the SSI system in S. squalidus were also investigated: (a) the strength of self-incompatibility response; (b) the nature of S allele dominance interactions; and (c) the relative frequencies of S phenotypes. These are discussed in view of the low S allele diversity estimates and the known population history of S. squalidus.

Crosses, Genetic↗

Precision and high-resolution mapping of quantitative trait loci by use of recurrent selection, backcross or intercross schemes.

Dissecting quantitative genetic variation into genes at the molecular level has been recognized as the greatest challenge facing geneticists in the twenty-first century. Tremendous efforts in the last two decades were invested to map a wide spectrum of quantitative genetic variation in nearly all important organisms onto their genome regions that may contain genes underlying the variation, but the candidate regions predicted so far are too coarse for accurate gene targeting. In this article, the recurrent selection and backcross (RSB) schemes were investigated theoretically and by simulation for their potential in mapping quantitative trait loci (QTL). In the RSB schemes, selection plays the role of maintaining the recipient genome in the vicinity of the QTL, which, at the same time, are rapidly narrowed down over multiple generations of backcrossing. With a high-density linkage map of DNA polymorphisms, the RSB approach has the potential of dissecting the complex genetic architecture of quantitative traits and enabling the underlying QTL to be mapped with the precision and resolution needed for their map-based cloning to be attempted. The factors affecting efficiency of the mapping method were investigated, suggesting guidelines under which experimental designs of the RSB schemes can be optimized. Comparison was made between the RSB schemes and the two popular QTL mapping methods, interval mapping and composite interval mapping, and showed that the scenario of genomic distribution of QTL that was unlocked by the RSB-based mapping method is qualitatively distinguished from those unlocked by the interval mapping-based methods.

Animals↗

Quantitative trait loci (QTL) detection in multicross inbred designs: recovering QTL identical-by-descent status information from marker data.

Mapping quantitative trait loci in plants is usually conducted using a population derived from a cross between two inbred lines. The power of such QTL detection and the parameter estimates depend largely on the choice of the two parental lines. Thus, the QTL detected in such populations represent only a small part of the genetic architecture of the trait. In addition, the effects of only two alleles are characterized, which is of limited interest to the breeder, while common pedigree breeding material remains unexploited for QTL mapping. In this study, we extend QTL mapping methodology to a generalized framework, based on a two-step IBD variance component approach, applicable to any type of breeding population obtained from inbred parents. We then investigate with simulated data mimicking conventional breeding programs the influence of different estimates of the IBD values on the power of QTL detection. The proposed method would provide an alternative to the development of specifically designed recombinant populations, by utilizing the genetic variation actually managed by plant breeders. The use of these detected QTL in assisting breeding would thus be facilitated.

Computer Simulation↗

Estimating effective population size or mutation rate with microsatellites.

Microsatellites are short tandem repeats that are widely dispersed among eukaryotic genomes. Many of them are highly polymorphic; they have been used widely in genetic studies. Statistical properties of all measures of genetic variation at microsatellites critically depend upon the composite parameter theta = 4Nmicro, where N is the effective population size and micro is mutation rate per locus per generation. Since mutation leads to expansion or contraction of a repeat number in a stepwise fashion, the stepwise mutation model has been widely used to study the dynamics of these loci. We developed an estimator of theta, theta; (F), on the basis of sample homozygosity under the single-step stepwise mutation model. The estimator is unbiased and is much more efficient than the variance-based estimator under the single-step stepwise mutation model. It also has smaller bias and mean square error (MSE) than the variance-based estimator when the mutation follows the multistep generalized stepwise mutation model. Compared with the maximum-likelihood estimator theta; (L) by, theta; (F) has less bias and smaller MSE in general. theta; (L) has a slight advantage when theta is small, but in such a situation the bias in theta; (L) may be more of a concern.

Alleles↗

Genetic network models and statistical properties of gene expression data in knock-out experiments.

It is shown here how gene knock-out experiments can be simulated in Random Boolean Networks (RBN), which are well-known simplified models of genetic networks. The results of the simulations are presented and compared with those of actual experiments in S. cerevisiae. RBN with two incoming links per node have been considered, and the Boolean functions have been chosen at random among the set of so-called canalizing functions. Genes are knocked-out (i.e. silenced) one at a time, and the variations in the expression levels of the other genes, with respect to the unperturbed case, are considered. Two important variables are defined: (i) avalanches, which measure the size of the perturbation generated by knocking out a single gene, and (ii) susceptibilities, which measure how often the expression of a given gene is modified in these experiments. A remarkable observation is that the distributions of avalanches and susceptibilities are very robust, i.e. they are very similar in different random networks; this should be contrasted with the distribution of other variables that show a high variance in RBN. Moreover, the distribution of avalanches and susceptibilities of the RBN models are close to those observed in actual experiments performed with S. cerevisiae, where the changes in gene expression levels have been recorded with DNA microarrays. These findings suggest that these distributions might be "generic" properties, common to a wide range of genetic models and real genetic networks. The importance of such generic properties is discussed.

Computational Biology↗

[Multiple nonlinear statistical method of population genetic structure based on the allelic polymorphism data].

The distribution and structure of the allelic polymorphism data are analyzed and it is pointed out that the distribution of allelic polymorphism data reveals the characteristic of closed data (also named as compositional data or data of constant sum). It is interpreted that the correlation structure of the allelic polymorphism data contains null correlations introduced by "closure" and the statistical distribution of the data is not normal because of its constant row sum, which resulted in great difficulties in analyzing the data with traditional multiple linear statistical methods such as principal component analysis, factor analysis, cluster analysis and canonical correlation analysis. Based on the theory of compositional data analysis proposed by Aitchison in 1982, a multiple nonlinear statistical method originating from the "logratios" approach to the statistical analysis of compositional data is put forward in this paper. As an example, the "logratios" method was used to analyze the genetic structure of TH01 polymorphic loci in Chinese population and the results were compared with those of multiple linear methods such as component principal. It is concluded that the "logratios" multiple nonlinear principle component analysis is a better method with the virtue of sensitivity and specificity for analyzing the genetic structure of population from the data of allelic polymorphism.

Alleles↗

GFINDer: genetic disease and phenotype location statistical analysis and mining of dynamically annotated gene lists.

Phenotype analysis is commonly recognized to be of great importance for gaining insight into genetic interaction underlying inherited diseases. However, few computational contributions have been proposed for this purpose, mainly owing to lack of controlled clinical information easily accessible and structured for computational genome-wise analyses. We developed and made available through GFINDer web server an original approach for the analysis of genetic disorder related genes by exploiting the information on genetic diseases and their clinical phenotypes present in textual form within the Online Mendelian Inheritance in Man (OMIM) database. Because several synonyms for the same name and different names for overlapping concepts are often used in OMIM, we first normalized phenotype location descriptions reducing them to a list of unique controlled terms representing phenotype location categories. Then, we hierarchically structured them and the correspondent genetic diseases according to their topology and granularity of description, respectively. Thus, in GFINDer we could implement specific Genetic Disorders modules for the analysis of these structured data. Such modules allow to automatically annotate user-classified gene lists with updated disease and clinical information, classify them according to the genetic syndrome and the phenotypic location categories, and statistically identify the most relevant categories in each gene class. GFINDer is available for non-profit use at http://www.bioinformatics.polimi.it/GFINDer/.

Data Interpretation, Statistical↗

Statistical validity for testing associations between genetic markers and quantitative traits in family data.

In genetic analysis it is often of interest to analyze associations between traits of unknown genetic etiology and genetic markers from pedigree data. Statistical methods that assume independence of pedigree members cannot be used because they disregard the statistical dependencies of members in a pedigree. For quantitative traits, a regression model proposed by George and Elston [Genet Epidemiol 4:193-201, 1987] uses an asymptotic likelihood ratio test and incorporates a correlation structure that allows for statistical dependence among the pedigree members. The statistical validity of this test is assessed for finite samples by measuring the discrepancy between the empirical and theoretical chi-square distributions. The variance of the mean of the dependent variable is determined to be related to this discrepancy and can be used to determine whether a pedigree structure is large enough for making valid statistical inferences on the basis of the asymptotic test. A multi-generational pedigree of 200 or so individuals should in many cases be sufficient for valid results when using the asymptotic likelihood ratio test for the association between markers and continuous traits.

Family↗

Identifying quantitative trait locus by genetic background interactions in association studies.

Association studies are designed to identify main effects of alleles across a potentially wide range of genetic backgrounds. To control for spurious associations, effects of the genetic background itself are often incorporated into the linear model, either in the form of subpopulation effects in the case of structure or in the form of genetic relationship matrices in the case of complex pedigrees. In this context epistatic interactions between loci can be captured as an interaction effect between the associated locus and the genetic background. In this study I developed genetic and statistical models to tie the locus by genetic background interaction idea back to more standard concepts of epistasis when genetic background is modeled using an additive relationship matrix. I also simulated epistatic interactions in four-generation randomly mating pedigrees and evaluated the ability of the statistical models to identify when a biallelic associated locus was epistatic to other loci. Under additive-by-additive epistasis, when interaction effects of the associated locus were quite large (explaining 20% of the phenotypic variance), epistasis was detected in 79% of pedigrees containing 320 individuals. The epistatic model also predicted the genotypic value of progeny better than a standard additive model in 78% of simulations. When interaction effects were smaller (although still fairly large, explaining 5% of the phenotypic variance), epistasis was detected in only 9% of pedigrees containing 320 individuals and the epistatic and additive models were equally effective at predicting the genotypic values of progeny. Epistasis was detected with the same power whether the overall epistatic effect was the result of a single pairwise interaction or the sum of nine pairwise interactions, each generating one ninth of the epistatic variance. The power to detect epistasis was highest (94%) at low QTL minor allele frequency, fell to a minimum (60%) at minor allele frequency of about 0.2, and then plateaued at about 80% as alleles reached intermediate frequencies. The power to detect epistasis declined when the linkage disequilibrium between the DNA marker and the functional polymorphism was not complete.

Animals↗

[Dynamics of population genetics parameters and their statistical during assessment selection for quantitative characters. I. An additive model. One character].

The selection for a single additively inherited quantitative character is studied using computer models of 3 types: 1) all the individuals had the same viability, and the paratypic deviation does not depend on their genotype; 2) differential viability of genotypes is taken into account with respect to a number of heterozygous loci; 3) differential paratypic deviation is estimated, it is introduced like viability in the model 2. Two types of trancation selection, stabilizing and directed, under the selection coefficient of approximately 0.5 are studied. There are studied dynamics of genotypic (omega2gamma) and phenotypic (omega2phi) variances, the heritability index and it estimates (for the correlation progeny-parent--rho; for the regression progeny-parent--b; for half-sibses--v), and the non-equilibrium of a population for the models. The number of generation was 10; the population number in every generation was 200; the number of loci in the main experiment was 10.h2 values were calculated in the progeny before selection (p1, b1, v1) and after selection (p2, b2, v2). It is shown that stabilizing selection results in the formation of balanced gene complexes; the rate of decreasing omega2phi depends on the genetic length of a chromosome region in which genes, determining the character, are located, and not on the number of genes. The distribution of a character under directed selection depends on the type of the model. p2, b2, and v2 values are the worst. The best is the b1 value. It is concluded that the problem of predicting the selection effect using statistical estimates of heretability is connected with the problem of investigation of population heterogeneity and integrating their genetical structure.

Computers↗

Molecular variability of sunflower downy mildew, Plasmopara halstedii, from different continents.

Downy mildew of sunflower (Helianthus annuus L.), caused by the pathogen Plasmopara halstedii, is a potentially devastating disease. Seventy-seven isolates of P. halstedii collected in twelve countries from four continents were investigated for RAPD polymorphism with 21 primers. The study led to a binary matrix, which was subjected to various complementary analyses. This is the first report on the international genetic diversity of the pathogen. Similarity indices ranged from 89% to 100%. Neither a consensus unweighted pair group method with arithmetic means (UPGMA) tree constructed after bootstrap resampling of markers nor a principal component analysis based on distance matrix revealed very consistent clusterings of the isolates, and groups did not fit race or geographical origins. Phylogenies were probably obscured by limited diversity. Analysis of molecular variance (AMOVA) and Nei's genetic diversity statistics gave similar conclusions. Most of the genetic diversity was attributable to individual differences. The most differentiated races also had the lowest within-diversity indices, which suggest that they appeared recently with strong bottleneck effects. Our analyses suggest that this pathogen is probably homothallic or has an asexual mode of reproduction and that gene flow among countries can occur through commercial exchanges. Knowledge of the downy mildew populations' structure at the international level will help to devise strategies for controlling this potentially devastating disease.

DNA↗