Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical genetics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Statistical methods in psychiatric genetics.

Most psychiatric disorders are determined by the complex interplay between genetic and environmental factors. Aetiological research into these complex disorders raises many different questions which require a variety of statistical methods. These include survival analysis for the estimation of morbid risk, structural equation models for the partitioning of phenotypic variances and covariances into genetic and other components, complex segregation analysis to detect loci of major effect, and linkage and association analysis for the localisation and identification of susceptibility genes. Future developments in psychiatric genetics will involve the integration of genetic and epidemiological statistics in order to study the interplay between genetic and environmental factors in the complex pathways which lead to mental disorders.

Chromosome Mapping↗

The triangle test statistic (TTS): a test of genetic homogeneity using departure from the triangle constraints in IBD distribution among affected sib-pairs.

The proportions of affected sibs sharing 2, 1 or 0 identical by descent parental marker alleles have been shown to conform to the 'triangle constraints' (Suarez, 1978; Holmans, 1993). It has also been shown (Dudoit & Speed, 1999) that the constraints are verified provided certain assumptions hold. In this study we explore a realistic situation in which the constraints fail due to the presence of a factor in which the sibs differ, a factor on which penetrance depends. This factor may be a characteristic of the trait (severe vs. mild form), or the presence/absence of an associated trait or an environmental factor. We show that under such situations, using the triangle constraints may lead to important loss of power to detect linkage by the MLS test. We propose here an alternative approach in order to detect both linkage and heterogeneity.

Alleles↗

Stability distribution in the phage lambda-DNA double helix: a correlation between physical and genetic structure.

Statistical analyses on the positional correlation of physical-stability and base-sequence distribution maps with genetic map are made for the whole DNA (48502 bases) of lambda-phage. The susceptibility to a double-helix unfolding perturbation and the fraction of the transient opening of a particular region of the double helix are adopted to define this physical stability. The principal features obtained are: A) The DNA double strand of protein coding regions is found to have homostabilizing propensity around a defined stability which is characteristic to each individual gene. B) The stability of the double helix in non-protein coding region fluctuates, on average over the whole region, more than that in protein coding region. C) Boundary regions of protein coding and non-protein coding regions are regions of high stability-fluctuation. Stability especially fluctuates at the protein-coding-region side of the boundary. Contrary to the quiet feature of the interior part of protein coding region rather noisy part exists at its edge. D) One frequently opening region coincides with the attaching site for the site specific recombination between phage and bacterial DNA. There are two possible ways to explain the noisy feature in the stability distribution in non-protein coding regions: 1) The region has been used as the locus of recombination as evolution took place. Thus DNAs which were homostabilized around a different value characteristic to each individual DNA, have been joined there many times, so that the noise has accumulated as a remnant of evolutional history; and/or 2) the base-composition homogenizing or double-helix homostabilizing mechanism does not work in unneeded region such as non-protein coding region or introns. Since corresponding characteristics have been found in our previous analyses on other viral and globin-gene DNAs, the rules mentioned above may be comprehensively extended to other DNAs.

Bacteriophage lambda↗

[Bayesian statistics-based method for genetic linkage analysis].

Bayesian School as one of the important statistical schools is different from the Classical Statistics, and the Bayesian methods have been widely used in many fields of modern sciences. In the present paper, we discussed the application of Bayesian method in linkage analysis, including the Bayesian estimation of recombination fraction, linkage testing based on the Bayes Factor and the Bayesian approach for genetic linkage map construction via Markov chain Monte Carlo algorithm. Simulation study and real data analysis were performed using SAS/IML software, and the validity and practicability of Bayesian method in genetic linkage analysis were thus verified.

Bayes Theorem↗

External noise and feedback regulation: steady-state statistics of auto-regulatory genetic network.

The steady-state statistics of a single gene auto-regulatory genetic network with the additive external Gaussian white noises is investigated. The main result shows that the negative feedback will result in that the mRNA noise has a positive contribution to the protein noise, but the positive feedback will result in that the mRNA noise has a negative contribution to the protein noise. If there is no feed back, then the contribution of mRNA noise to protein noise is always positive. On the other hand, the analysis and numerical simulations of linear and nonlinear feedback show that it is possible that the negative feedback increases, but the positive feedback decreases, the protein noise.

Animals↗

A statistical model for the genetic origin of allometric scaling laws in biology.

Many biological processes, from cellular metabolism to population dynamics, are characterized by particular allometric scaling (power-law) relationships between size and rate. Although such allometric relationships may be under genetic determination, their precise genetic mechanisms have not been clearly understood due to a lack of a statistical analytical method. In this paper, we present a basic statistical framework for mapping quantitative genes (or quantitative trait loci, QTL) responsible for universal quarter-power scaling laws of organic structure and function with the entire body size. Our model framework allows the testing of whether a single QTL affects the allometric relationship of two traits or whether more than one linked QTL is segregating. Like traditional multi-trait mapping, this new model can increase the power to detect the underlying QTL and the precision of its localization on the genome. Beyond the traditional method, this model is integrated with pervasive scaling laws to take advantage of the mechanistic relationships of biological structures and processes. Simulation studies indicate that the estimation precision of the QTL position and effect can be improved when the scaling relationship of the two traits is considered. The application of our model in a real example from forest trees leads to successful detection of a QTL governing the allometric relationship of third-year stem height with third-year stem biomass. The model proposed here has implications for genetic, evolutionary, biomedicinal and breeding research.

Animals↗

Statistics of selectively neutral genetic variation.

Random models of evolution are instrumental in extracting rates of microscopic evolutionary mechanisms from empirical observations on genetic variation in genome sequences. In this context it is necessary to know the statistical properties of empirical observables (such as the local homozygosity, for instance). Previous work relies on numerical results or assumes Gaussian approximations for the corresponding distributions. In this paper we give an analytical derivation of the statistical properties of the local homozygosity and other empirical observables assuming selective neutrality. We find that such distributions can be very non-Gaussian.

Alleles↗

How to foster citizens' statistical reasoning: implications for genetic counseling.

OBJECTIVES: Our aim is to provide an overview of key research findings from cognitive psychology regarding effective ways of communicating statistical information, and to point out the implications of these findings for genetic testing. METHOD: We review the literature on the presentation of statistical information in diagnostic test results, discuss various representations that invite misunderstandings, and propose alternative representations that foster understanding. RESULTS: Single-event probabilities, conditional probabilities and relative risks are easily misunderstood. Specifying the class of events to which a probability refers and using natural frequency statements improve understanding. CONCLUSIONS: Cognitive psychology has identified simple and effective tools for improving statistical reasoning. They can help to improve the public's understanding of diagnostic test results.

Data Interpretation, Statistical↗

Testing for association in the presence of population stratification: a simulation study comparing the S-TDT, STRAT and the GC.

A novel approach for association testing in the presence of population stratification has been introduced by Pritchard et al. (2000a) and Pritchard et al. (2000b). The structured association approach is a two-tiered procedure that first estimates the population structure and then tests the null hypothesis H0: 'no association within subpopulations' in the second step. A power comparison of the stratified test for association (STRAT) (Pritchard et al., 2000b) and the Transmission-Disequilibrium-Test (TDT) (Spielman and Ewens, 1993a) in a simulation framework showed superiority of STRAT if allele frequencies or associations between allele and disease differ strongly in subpopulations. In more homogeneous situations, the TDT had greater power than STRAT. However, the TDT, based on family trios,that uses population controls, needs 50% more genotyping compared to STRAT. The Sib-Transmission-Disequilibrium-Test (S-TDT) needs the same amount of genotyping since it relays in its minimal configuration on pairs of siblings. This raises the question how the S-TDT (Spielman and Ewens, 1998a) performs compared to the population based methods STRAT and Genomic Controls (GC). In this paper, we present a simulation study accounting for two different models of population stratification in different settings of allele frequencies and under different risk models. The results showed that under a discrete as well as under an admixed population model, STRAT strongly outperformed the S-TDT and the GC when different alleles were associated in different subpopulations. In contrast, the S-TDT had greater power than STRAT when the same allele was associated in both subpopulations. Here, the GC was sometimes even more powerful than the S-TDT, depending on the population model and the allele frequency differences. A general recommendation for the use of one of the tests can therefore not be given.

Algorithms↗

A genetic factor model for the statistical analysis of multilocus DNA fingerprints.

A novel concept is described for the statistical analysis of multilocus DNA fingerprints. Utilizing this method, it is shown by simulation that the application of multilocus DNA fingerprints to paternity testing is robust against deviations from idealistic assumptions made about underlying models and parameters. Partial homozygosity, allelism and linkage at the DNA loci involved, as well as variations in estimates of band-sharing probabilities were studied for effects on the resulting paternity probabilities. None of the above-mentioned phenomena appear to change these values to an extent relevant for decision making in paternity cases.

Alleles↗

Individual-specific risk factors for anorexia nervosa: a pilot study using a discordant sister-pair design.

BACKGROUND: The aim of this pilot study was to examine which unique factors (genetic and environmental) increase the risk for developing anorexia nervosa by using a case-control design of discordant sister pairs. METHODS: Forty-five sister-pairs, one of whom had anorexia nervosa and the other did not, were recruited. Both sisters completed the Oxford Risk Factor Interview for Eating Disorders and measures for eating disorder traits, and sib-pair differences. Blood or cheek cell samples were taken for genetic analysis. Statistical power of the genetic analysis of discordant same-sex siblings was calculated using a specially written program, DISCORD. RESULTS: The sisters with anorexia nervosa differed from their healthy sisters in terms of personal vulnerability traits and exposure to high parental expectations and sexual abuse. Factors within the dieting risk domain did not differ. However, there was evidence of poor feeding in childhood. No difference in the distribution of genotypes or alleles of the DRD4, COMT, the 5HT2A and 5HT2C receptor genes was detected. These results are preliminary because our calculations indicate that there is insufficient power to detect the expected effect on risk with this sample size. CONCLUSIONS: A combination of intrinsic and extrinsic factors increases the risk of developing anorexia nervosa. It would, therefore, be informative to undertake a larger study to examine in more detail the unique genetic and environmental factors that are associated with various forms of eating disorders.

Adolescent↗

The family based association test method: strategies for studying general genotype--phenotype associations.

With possibly incomplete nuclear families, the family based association test (FBAT) method allows one to evaluate any test statistic that can be expressed as the sum of products (covariance) between an arbitrary function of an offspring's genotype with an arbitrary function of the offspring's phenotype. We derive expressions needed to calculate the mean and variance of these test statistics under the null hypothesis of no linkage. To give some guidance on using the FBAT method, we present three simple data analysis strategies for different phenotypes: dichotomous (affection status), quantitative and censored (eg, the age of onset). We illustrate the approach by applying it to candidate gene data of the NIMH Alzheimer Disease Initiative. We show that the RC-TDT is equivalent to a special case of the FBAT method. This result allows us to generalise the RC-TDT to dominant, recessive and multi-allelic marker codings. Simulations compare the resulting FBAT tests to the RC-TDT and the S-TDT. The FBAT software is freely available.

Alzheimer Disease↗

Gap statistics for whole genome shotgun DNA sequencing projects.

MOTIVATION: Investigators utilize gap estimates for DNA sequencing projects. Standard theories assume sequences are independently and identically distributed, leading to appreciable under-prediction of gaps. RESULTS: Using a statistical scaling factor and data from 20 representative whole genome shotgun projects, we construct regression equations that relate coverage to a normalized gap measure. Prokaryotic genomes do not correlate to sequence coverage, while eukaryotes show strong correlation if the chaff is ignored. Gaps decrease at an exponential rate of only about one-third of that predicted via theory alone. Case studies suggest that departure from theory can largely be attributed to assembly difficulties for repeat-rich genomes, but bias and coverage anomalies are also important when repeats are sparse. Such factors cannot be readily characterized a priori, suggesting upper limits on the accuracy of gap prediction. We also find that diminishing coverage probability discussed in other studies is a theoretical artifact that does not arise for the typical project.

Animals↗

Statistical models for estimating the genetic basis of repeated measures and other function-valued traits.

The genetic analysis of characters that are best considered as functions of some independent and continuous variable, such as age, can be a complicated matter, and a simple and efficient procedure is desirable. Three methods are common in the literature: random regression, orthogonal polynomial approximation, and character process models. The goals of this article are (i) to clarify the relationships between these methods; (ii) to develop a general extension of the character process model that relaxes correlation stationarity, its most stringent assumption; and (iii) to compare and contrast the techniques and evaluate their performance across a range of actual and simulated data. We find that the character process model, as described in 1999 by Pletcher and Geyer, is the most successful method of analysis for the range of data examined in this study. It provides a reasonable description of a wide range of different covariance structures, and it results in the best models for actual data. Our analysis suggests genetic variance for Drosophila mortality declines with age, while genetic variance is constant at all ages for reproductive output. For growth in beef cattle, however, genetic variance increases linearly from birth, and genetic correlations are high across all observed ages.

Aging↗

Detecting statistically significant common insertion sites in retroviral insertional mutagenesis screens.

Retroviral insertional mutagenesis screens, which identify genes involved in tumor development in mice, have yielded a substantial number of retroviral integration sites, and this number is expected to grow substantially due to the introduction of high-throughput screening techniques. The data of various retroviral insertional mutagenesis screens are compiled in the publicly available Retroviral Tagged Cancer Gene Database (RTCGD). Integrally analyzing these screens for the presence of common insertion sites (CISs, i.e., regions in the genome that have been hit by viral insertions in multiple independent tumors significantly more than expected by chance) requires an approach that corrects for the increased probability of finding false CISs as the amount of available data increases. Moreover, significance estimates of CISs should be established taking into account both the noise, arising from the random nature of the insertion process, as well as the bias, stemming from preferential insertion sites present in the genome and the data retrieval methodology. We introduce a framework, the kernel convolution (KC) framework, to find CISs in a noisy and biased environment using a predefined significance level while controlling the family-wise error (FWE) (the probability of detecting false CISs). Where previous methods use one, two, or three predetermined fixed scales, our method is capable of operating at any biologically relevant scale. This creates the possibility to analyze the CISs in a scale space by varying the width of the CISs, providing new insights in the behavior of CISs across multiple scales. Our method also features the possibility of including models for background bias. Using simulated data, we evaluate the KC framework using three kernel functions, the Gaussian, triangular, and rectangular kernel function. We applied the Gaussian KC to the data from the combined set of screens in the RTCGD and found that 53% of the CISs do not reach the significance threshold in this combined setting. Still, with the FWE under control, application of our method resulted in the discovery of eight novel CISs, which each have a probability less than 5% of being false detections.

Chromosome Mapping↗

Relationship estimation in affected sib pair analysis of late-onset diseases.

In linkage studies, errors in pedigree structure will often be uncovered through Mendelian inconsistencies. In affected sib pair analysis of diseases with late onset, however, such mistakes will usually go undetected since parental genotypes are commonly not known. Cases of nonpaternity, unrecorded adoption or accidental sample swap in the laboratory will then not be noticed. Typically, such relationship errors lead to a decrease in power for linkage. In this paper, a method is presented which allows verification of the relationship between stated sibs using their marker genotypes. The method is likelihood-based and incorporates a Bayesian approach to compute posterior relationship probabilities. It is shown that sibs, half-sibs and unrelated individuals can be distinguished from each other quite reliably using numbers of markers that should be available in most sib pair studies. It is demonstrated that elimination of false sib pairs increases the power to detect linkage in affected sib pair studies. The gain in power may be large if relationship errors occur quite frequently; the gain will be only moderate if relationship errors are very infrequent. Software for relationship estimation is provided.

Bayes Theorem↗

A unified approach to study hypervariable polymorphisms: statistical considerations of determining relatedness and population distances.

Relatedness between individuals as well as evolutionary relationships between populations can be studied by comparing genotypic similarities between individuals. When hypervariable loci are used to describe genotypes, it is shown that both of these problems can be approached with a unified theory based on allele sharing between individuals. The distributions of the number of shared alleles between individuals indicate their kin relationships. Extending this, we obtain statistics for genetic distances between populations based on average number of alleles shared between individuals within and between two different populations. Traditional statistical inferential procedure can be used to establish specific kinship relationships between individuals. We derive estimates of the number of hypervariable loci needed for a specified reliability of such an inference. Evolutionary dynamics of genetic distance statistics based on allele sharing is also studied. It shows that such measures of genetic distances remain linear with the time of divergence for a period comparable to that of the gene frequency-based measures of genetic distances. Statistical properties of measures based on allele sharing establish that for using such summary statistics it is not necessary to know the full characteristics of all loci used. It is enough to know the degree of heterozygosity per locus and the number of loci. Therefore, in principle, this approach can also be used for DNA fingerprinting data in the studies of relatedness between individuals as well as between populations. The possible compromising features of multilocus DNA fingerprinting data are also discussed.

Alleles↗

Score statistic to test for genetic correlation for proband-family design.

In genetic epidemiological studies informative families are often oversampled to increase the power of a study. For a proband-family design, where relatives of probands are sampled, we derive the score statistic to test for clustering of binary and quantitative traits within families due to genetic factors. The derived score statistic is robust to ascertainment scheme. We considered correlation due to unspecified genetic effects and/or due to sharing alleles identical by descent (IBD) at observed marker locations in a candidate region. A simulation study was carried out to study the distribution of the statistic under the null hypothesis in small data-sets. To illustrate the score statistic, data from 33 families with type 2 diabetes mellitus (DM2) were analyzed. In addition to the binary outcome DM2 we also analyzed the quantitative outcome, body mass index (BMI). For both traits familial aggregation was highly significant. For DM2, also including IBD sharing at marker D3S3681 as a cause of correlation gave an even more significant result, which suggests the presence of a trait gene linked to this marker. We conclude that for the proband-family design the score statistic is a powerful and robust tool for detecting clustering of outcomes.

Alleles↗