Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical genetics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Transition event statistics in genetics and disordered kinetics. Theoretical approaches for extracting rate distributions from experimental data.

We study the analogies between the theory of rate processes in disordered systems and the overdispersed molecular clocks in evolutionary biology. A biological "molecular clock" expresses the statistics of the number of amino acid or nucleotide substitutions during evolution. Random variations of the evolution rates lead to statistical (overdispersed) molecular clocks which are described by random point processes with random substitution rates. We find that the models for overdispersed molecular clocks are equivalent to those of the random-rate or random channel models used in disordered kinetics. The number of transport (reaction) events in disordered kinetics plays the same role as the number of substitution events in molecular biology. We study the connections between the (observed) statistics of the transition events and the statistics of random rate coefficients and random channels; a unified approach is developed which is valid both in molecular biology and in disordered kinetics. We develop methods for extracting statistical information about the variations of rate coefficients from experimental or observed data regarding the fluctuations of the numbers of substitution, reaction, or transport events. For systems with static disorder, the observed statistics of the number of reaction events, expressed in terms of probabilities at a given time or by the cumulants of the number of transition events at a given time, contains the information necessary for evaluating the cumulants or the probability density of the rate coefficients or the density of states for random channel kinetics. For dynamic disorder this is not possible; further information about multitime probability distributions of the reaction events is needed.

Amino Acid Substitution↗

Logical implications of applying the principles of population genetics to the interpretation of DNA profiling evidence.

There have been several efforts to codify the approach to interpreting DNA evidence [National Research Council, The Evaluation of Forensic DNA Evidence, National Academy Press, Washington, DC, 1996; I.W. Evett, B.S. Weir, Interpreting DNA Evidence: Statistical Genetics for Forensic Scientists Sinauer, Sunderland, MA, 1998]. Despite these efforts there are still aspects of ad hoc decision making in modern DNA interpretation. This article discusses some of the remaining areas of concern in this respect. Because of the immense discriminating power of DNA evidence it is unlikely that these concerns would contribute to a miscarriage of justice. They are more likely to lead to lengthy and wasteful debate in court, and to potential appeals. We advocate a previously developed approach to DNA evidence [Sci. Justice 39 (4) (1999) 257; B.S. Weir, in: D.J. Balding, C. Cannings, M. Bishop (Eds.), Handbook of Statistical Genetics, Wiley Series in Probability and Statistics, Wiley, New York, ISBN: 0-471-86094-8, 2001; J. R. Stat. Soc. A 158 (1) (1995) 21] that would give a more solid logical foundation and hopefully lead to sounder and less debatable testimony.

DNA Fingerprinting↗

Identification and simulation of new non-random statistical properties common to different populations of eukaryotic non-coding genes.

The autocorrelation function analysing the occurrence probability of the i-motif YRY(N)iYRY in genes allows the identification of mainly two periodicities modulo 2, 3 and the preferential occurrence of the motif YRY(N)6YRY (R = purine = adenine or guanine, Y = pyrimidine = cytosine or thymine, N = R or Y). These non-random genetic statistical properties can be simulated by an independent mixing of the three oligonucleotides YRYRYR, YRYYRY and YRY(N)6 (Arquès & Michel, 1990b). The problem investigated in this study is whether new properties can be identified in genes with other autocorrelation functions and also simulated with an oligonucleotide mixing model. The two autocorrelation functions analysing the occurrence probability of the i-motifs RRR(N)iRRR and YYY(N)iYYY simultaneously identify three new non-random genetic statistical properties: a short linear decrease, local maxima for i identical to 3[6] (i = 3, 9, etc) and a large exponential decrease. Furthermore, these properties are common to three different populations of eukaryotic non-coding genes: 5' regions, introns and 3' regions (see section 2). These three non-random properties can also be simulated by an independent mixing of the four oligonucleotides R8, Y8, RRRYRYRRR, YYYRYRYYY and large alternating R/Y series. The short linear decrease is a result of R8 and Y8, the local maxima for i identical to 3[6], of RRRYRYRRR and YYYRYRYYY, and the large exponential decrease, of large alternating R/Y series (section 3). The biological meaning of these results and their relation to the previous oligonucleotide mixing model are presented in the Discussion.

Animals↗

Quantitative trait locus mapping using human pedigrees.

In the past decade phenomenal progress has been made in molecular and statistical genetic methods for localizing quantitative trait loci. Because of these advances, we can anticipate a long period of active genetic research in which the genes influencing human quantitative variability will be mapped and their effects accurately evaluated. Here, we review the current state of the science in statistical genetic methods for quantitative trait linkage analysis. In particular, we detail a variance component-based framework for localizing quantitative trait loci and for accurately estimating their relative effect sizes. Attention is paid to the optimal design of human family studies for localizing genes of small to moderate effect. In addition, methods and strategies are described for dealing with the most important complications of quantitative variation, including the assessment of genotype x environment interaction and epistasis.

Bias↗

HLA-SD antigens and schizophrenia: statistical and genetical considerations.

The HLA-SD phenotype distributions of hebephrenic and paranoid schizophrenics, and of the two groups combined, in an Italian population and in a combined group from the Swedish population have been analyzed statistically. There is a significantly decreased frequency of HLA-A10 in all of these. Theae are some preliminary indications of an increased frequency (a positive association) for some of the other antigens of the HLA-SD series, but there is insufficient data at present for evaluating the significance of these findings. Differences between hebephrenic and paranoid schizophrenics have been detected.

Diagnosis, Differential↗

Quantitative trait nucleotide analysis using Bayesian model selection.

Although much attention has been given to statistical genetic methods for the initial localization and fine mapping of quantitative trait loci (QTLs), little methodological work has been done to date on the problem of statistically identifying the most likely functional polymorphisms using sequence data. In this paper we provide a general statistical genetic framework, called Bayesian quantitative trait nucleotide (BQTN) analysis, for assessing the likely functional status of genetic variants. The approach requires the initial enumeration of all genetic variants in a set of resequenced individuals. These polymorphisms are then typed in a large number of individuals (potentially in families), and marker variation is related to quantitative phenotypic variation using Bayesian model selection and averaging. For each sequence variant a posterior probability of effect is obtained and can be used to prioritize additional molecular functional experiments. An example of this quantitative nucleotide analysis is provided using the GAW12 simulated data. The results show that the BQTN method may be useful for choosing the most likely functional variants within a gene (or set of genes). We also include instructions on how to use our computer program, SOLAR, for association analysis and BQTN analysis.

Bayes Theorem↗

Genetic and statistical properties of residual feed intake.

Residual feed intake is defined as the difference between actual feed intake and that predicted on the basis of requirements for production and maintenance of body weight. Formulas were developed to obtain genetic parameters of residual feed intake from knowledge of the genetic and phenotypic parameters of the component traits. Genetic parameters of residual feed intake were determined for a range of heritabilities (h2 = .1, .3, or .5) for component traits of feed intake and production, and genetic (rg = .1, .5, or .9) and environmental (re = .1, .5, or .9) correlations between them. Resulting heritability of residual feed intake ranged from .03 to .84 and the genetic correlation between residual feed intake and production ranged from -.90 to .87. Heritability of residual feed intake depends considerably on the environmental correlation between feed intake and production. Residual feed intake based on phenotypic regression of feed intake on production usually contains a genetic component due to production. Residual feed intake based on genotypic regression of feed intake on production is genetically independent of production and its use is equivalent to use of a selection index restricted to hold production constant. Multiple-trait selection on residual feed intake, based on either phenotypic or genetic regressions, and production is equivalent to multiple-trait selection on feed intake and production. Residual energy intake in dairy cattle was examined as an example. Heritability of residual energy intake based on genotypic regression was close to zero and indicated that measurement of feed intake provides little additional genetic information over and above that provided by milk production and body weight. The principles outlined in this study have broader application than just to residual feed intake and apply to any trait that is defined as a linear function of other traits.

Animals↗

Tension versus ecological zones in a two-locus system.

Previous theories show that tension and ecological zones are indistinguishable in terms of gene frequency clines. Here I analytically show that these two types of zones can be distinguished in terms of genetic statistics other than gene frequency. A two-locus cline model is examined with the assumptions of random mating, weak selection, no drift, no mutation, and multiplicative viabilities. The genetic statistics for distinguishing the two types of zones are the deviations of one- or two-locus genotypic frequencies from Hardy-Weinberg equilibrium (HWE) or from random association of gametes (RAG), and the deviations of additive and dominance variances from the values at HWE. These deviations have a discontinuous distribution in space and different extents of interruptions in the ecological zone with a sharp boundary, but exhibit a continuous distribution in the tension zone. Linkage disequilibrium enhances the difference between the deviations from HWE and from RAG for any two-locus genotypic frequency.

Ecosystem↗

Theory and practice in quantitative genetics.

With the rapid advances in molecular biology, the near completion of the human genome, the development of appropriate statistical genetic methods and the availability of the necessary computing power, the identification of quantitative trait loci has now become a realistic prospect for quantitative geneticists. We briefly describe the theoretical biometrical foundations underlying quantitative genetics. These theoretical underpinnings are translated into mathematical equations that allow the assessment of the contribution of observed (using DNA samples) and unobserved (using known genetic relationships) genetic variation to population variance in quantitative traits. Several statistical models for quantitative genetic analyses are described, such as models for the classical twin design, multivariate and longitudinal genetic analyses, extended twin analyses, and linkage and association analyses. For each, we show how the theoretical biometrical model can be translated into algebraic equations that may be used to generate scripts for statistical genetic software packages, such as Mx, Lisrel, SOLAR, or MERLIN. For using the former program a web-library (available from http://www.psy.vu.nl/mxbib) has been developed of freely available scripts that can be used to conduct all genetic analyses described in this paper.

Genetic Linkage↗

[Allelic variations of DPB1, DQA1, DQB1 and DRB1 and rheumatoid arthritis: further genetic and statistical considerations].

We used polymerase chain reaction amplification and hybridization with specific oligonucleotides to analyze the distribution of DPB1, DQA1, DQB1, and DRB1 allelic variants in 48 patients with rheumatoid arthritis (RA) and compared our results with those from 109 randomly chosen, healthy control subjects. Our work confirms a previously reported increase in DR4 specificity in RA: in particular, we found a statistically significant positive association of the DRB1*0401 and DRB1*0404 alleles with RA. When we compared the DR4 groups, however, none of the DRB1*04 alleles were increased in the RA group. Molecular analysis of the other DRB1 polymorphic variants disclosed the trend of a positive association of DRB1*0101 (DR1) in DR4 negative patients vs DR4 negative healthy control subjects, and an increase in DRw6 (DRB1*13,*14) in the DR4 and/or DR1 negative patient group. Moreover, analysis of the association between RA and a heptapeptide motif (positions 67-74) in the third hypervariable region confirmed that this epitope confers enhanced risk for the development of RA with respect to the allele DRB1*0404 (etiologic fraction = 0.53 vs 0.12). We also observed a statistically significant increase in DQA1*0301 and DQB1*0302 accompanied by a significant decrease in DQA1*0202, DQA1*0501 and DQB1*0201 in RA patients. Analysis of DPB1 alleles disclosed no significant differences between RA patients and healthy control subjects.

Adult↗

A genetic and statistical study of the respiratory distress syndrome.

The hospital records of 197 infants with the respiratory distress syndrome (RDS) were reviewed and the families of 111 of them subsequently contacted to obtain a family history. After correcting for biasis of ascertainment, the incidence of RDS among the full sibs was found to be between 12 and 19% depending on whether the individuals diagnosed as "possible RDS" were counted as affected. Among the low birth weight (LBW, less than or equal to 2.5 kg) and/or preterm (less than or equal to 37 weeks gestation) infants in the sibships, the incidence of RDS was 32-50%. Considering only sibs born after the probands yielded the empiric recurrence risk of 17--27% for all younger sibs and 39--67% for LBW/preterm younger sibs. The risk for maternal half-sibs was of about the same magnitude as that for full sibs, while the risk for paternal half-sibs was minimal. Among the LBW/preterm first cousins of probands, only the infants of maternal aunts showed an RDS incidence clearly higher than that in the general population. We think these data suggest a genetically determined maternal factor predisposing the infants of certain mothers to RDS. Other significant findings include: 1) an excess of males among the probands but a normal sex ratio among the sibs of the probands; 2) a decrease in mean birth weight and mean length of gestation for not only the probands but also their sibs; 3) a decrease in the mean parental ages at the birth of the probands; 4) a relative dearth of first-born and an excess of second-born infants among the probands; 5) an increased incidence of stillbirths in the sibships; 6) an increased number of probands born by cesarean section; and 7) a twin concordance of 75%.

Age Factors↗

LRTae: improving statistical power for genetic association with case/control data when phenotype and/or genotype misclassification errors are present.

BACKGROUND: In the field of statistical genetics, phenotype and genotype misclassification errors can substantially reduce power to detect association with genetic case/control studies. Misclassification also can bias population frequency parameters such as genotype, haplotype, or multi-locus genotype frequencies. These problems are of particular concern in case/control designs because, short of repeated sampling, there is no way to detect misclassification errors. We developed a double-sampling procedure for case/control genetic association using a likelihood ratio test framework. Different approaches have been proposed to deal with misclassification errors. We have chosen the likelihood framework because of the ease with which misclassification probabilities may be incorporated into in the statistical framework and hypothesis testing. The statistic is called the Likelihood Ratio Test allowing for errors (LRTae) and is freely available via software download. RESULTS: We applied our procedure to 10,000 replicates of simulated case/control data in which we introduced phenotype misclassification errors. The phenotype considered is Ankylosing Spondylitis (AS). The LRTae method power was always greater than LRTstd power for the significance levels considered (5%, 1%, 0.1%, 0.01%). Power gains for the LRTae method over the LRTstd method increased as the significance level became more stringent. Multi-locus genotype frequency estimates using LRTae method were more accurate than estimates using LRTstd method. CONCLUSION: The LRTae method can be applied to single-locus genotypes, multi-locus genotypes, or multi-locus haplotypes in a case/control framework and can be more powerful to detect association in case/control studies when both genotype and/or phenotype errors are present. Furthermore, the LRTae method provides asymptotically unbiased estimates of case and control genotype frequencies, as well as estimates of phenotype and/or genotype misclassification rates.

Case-Control Studies↗