Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

A Guide for Exploring Pleiotropic Associations in Genome-Wide Association Studies Using Summary Statistics.

Genome-wide association studies (GWAS) have shown that pleiotropy, whereby a single genetic variant or gene influences multiple traits, is common in complex human diseases. Detecting cross-phenotype associations from GWAS summary statistics remains challenging because of small effect sizes, extensive multiple testing, heterogeneous effects, and possible differences in effect direction across traits. Methods that jointly analyze multiple traits can improve the ability to detect pleiotropic signals while retaining the practical advantages of summary statistic-based analyses. Although a range of statistical approaches has been developed for this purpose, practical guidance on their application, assumptions, and interpretation remains limited. This tutorial reviews several widely used methods for pleiotropy detection from GWAS summary statistics, including ASSET, PLACO, GPA, CPBayes, and GCPBayes, and demonstrates their application using breast and thyroid cancer datasets. We also highlight the importance of accounting for effect heterogeneity, correlation, and biological group structure at the gene and pathway levels in the detection and interpretation of pleiotropic association signals.

Genome-Wide Association Study↗

Multilocus statistics to uncover epistasis and heterogeneity in complex diseases: revisiting a set of multiple sclerosis data.

New statistics are developed to gather the contribution of many alleles at different loci to common diseases. Both inferential and descriptive statistics are included in order to uncover epistatic effects as well as heterogeneity. The problem of multiple testing is circumvented by considering a global null hypothesis. Global testing is supplemented by descriptive methods that make use of measures like odds ratio or the P-value of individually tested allele combinations. Visualization helps to reflect complex data sets. The methods described here have been scrutinized by statistical simulations, and we show that power gains can be substantial as compared to single locus statistics. Typing data of multiple sclerosis patients and controls are investigated, representing an example of larger scale information in screening candidate genes for their impact on complex diseases. New insights emerge from this data set demonstrating genetic heterogeneity and evidence for epistasis.

Adult↗

Statistical practices: the seven deadly sins.

This paper discusses selected problems in applied statistical analysis: (a) over-reliance on null hypothesis statistical testing, (b) failing to perform a power analysis prior to conducting the study, (c) using asymptotic statistical approximations with small samples, (d) ignoring missing data, (e) failing to deal with the multiplicity problem when performing multiple statistical comparisons, (f) using stepwise procedures to select variables in regression analysis, and (g) failing to perform or report model diagnostics. Suggestions and guidelines to address these issues in manuscripts concerning child neuropsychological research are offered.

Child↗

Bootstrap-corrected ADF test statistics in covariance structure analysis.

The asymptotically distribution-free (ADF) test statistic for covariance structure analysis (CSA) has been reported to perform very poorly in simulation studies, i.e. it leads to inaccurate decisions regarding the adequacy of models of psychological processes. It is shown in the present study that the poor performance of the ADF test statistic is due to inadequate estimation of the weight matrix (W = gamma -1), which is a critical quantity in the ADF theory. Bootstrap procedures based on Hall's bias reduction perspective are proposed to correct the ADF test statistic. It is shown that the bootstrap correction of additive bias on the ADF test statistic yields the desired tail behaviour as the sample size reaches 500 for a 15-variable-3-factor confirmatory factor-analytic model, even if the distribution of the observed variables is not multivariate normal and the latent factors are dependent. These results help to revive the ADF theory in CSA.

Factor Analysis, Statistical↗

Using nuclear morphometry to discriminate the tumorigenic potential of cells: a comparison of statistical methods.

Despite interest in the use of nuclear morphometry for cancer diagnosis and prognosis as well as to monitor changes in cancer risk, no generally accepted statistical method has emerged for the analysis of these data. To evaluate different statistical approaches, Feulgen-stained nuclei from a human lung epithelial cell line, BEAS-2B, and a human lung adenocarcinoma (non-small cell) cancer cell line, NCI-H522, were subjected to morphometric analysis using a CAS-200 imaging system. The morphometric characteristics of these two cell lines differed significantly. Therefore, we proceeded to address the question of which statistical approach was most effective in classifying individual cells into the cell lines from which they were derived. The statistical techniques evaluated ranged from simple, traditional, parametric approaches to newer machine learning techniques. The multivariate techniques were compared based on a systematic cross-validation approach using 10 fixed partitions of the data to compute the misclassification rate for each method. For comparisons across cell lines at the level of each morphometric feature, we found little to distinguish nonparametric from parametric approaches. Among the linear models applied, logistic regression had the highest percentage of correct classifications; among the nonlinear and nonparametric methods applied, the Classification and Regression Trees model provided the highest percentage of correct classifications. Classification and Regression Trees has appealing characteristics: there are no assumptions about the distribution of the variables to be used, there is no need to specify which interactions to test, and there is no difficulty in handling complex, high-dimensional data sets containing mixed data types.

Adenocarcinoma↗

Influence of protein structure databases on the predictive power of statistical pair potentials.

A long standing goal in protein structure studies is the development of reliable energy functions that can be used both to verify protein models derived from experimental constraints as well as for theoretical protein folding and inverse folding computer experiments. In that respect, knowledge-based statistical pair potentials have attracted considerable interests recently mainly because they include the essential features of protein structures as well as solvent effects at a low computing cost. However, the basis on which statistical potentials are derived have been questioned. In this paper, we investigate statistical pair potentials derived from protein three-dimensional structures, addressing in particular questions related to the form of these potentials, as well as to the content of the database from which they are derived. We have shown that statistical pair potentials depend on the size of the proteins included in the database, and that this dependence can be reduced by considering only pairs of residue close in space (i.e., with a cutoff of 8 A). We have shown also that statistical potentials carry a memory of the quality of the database in terms of the amount and diversity of secondary structure it contains. We find, for example, that potentials derived from a database containing alpha-proteins will only perform best on alpha-proteins in fold recognition computer experiments. We believe that this is an overall weakness of these potentials, which must be kept in mind when constructing a database.

Chemical Phenomena↗

Traversing the conceptual divide between biological and statistical epistasis: systems biology and a more modern synthesis.

Epistasis plays an important role in the genetic architecture of common human diseases and can be viewed from two perspectives, biological and statistical, each derived from and leading to different assumptions and research strategies. Biological epistasis is the result of physical interactions among biomolecules within gene regulatory networks and biochemical pathways in an individual such that the effect of a gene on a phenotype is dependent on one or more other genes. In contrast, statistical epistasis is defined as deviation from additivity in a mathematical model summarizing the relationship between multilocus genotypes and phenotypic variation in a population. The goal of this essay is to review definitions and examples of biological and statistical epistasis and to explore the relationship between the two. Specifically, we present and discuss the following two questions in the context of human health and disease. First, when does statistical evidence of epistasis in human populations imply underlying biomolecular interactions in the etiology of disease? Second, when do biomolecular interactions produce patterns of statistical epistasis in human populations? Answers to these two reciprocal questions will provide an important framework for using genetic information to improve our ability to diagnose, prevent and treat common human diseases. We propose that systems biology will provide the necessary information for addressing these questions and that model systems such as bacteria, yeast and digital organisms will be a useful place to start.

Animals↗

A scan statistic for identifying chromosomal patterns of SNP association.

We have developed a single nucleotide polymorphism (SNP) association scan statistic that takes into account the complex distribution of the human genome variation in the identification of chromosomal regions with significant SNP associations. This scan statistic has wide applicability for genetic analysis, whether to identify important chromosomal regions associated with common diseases based on whole-genome SNP association studies or to identify disease susceptibility genes based on dense SNP positional candidate studies. To illustrate this method, we analyzed patterns of SNP associations on chromosome 19 in a large cohort study. Among 2,944 SNPs, we found seven regions that contained clusters of significantly associated SNPs. The average width of these regions was 35 kb with a range of 10-72 kb. We compared the scan statistic results to Fisher's product method using a sliding window approach, and detected 22 regions with significant clusters of SNP associations. The average width of these regions was 131 kb with a range of 10.1-615 kb. Given that the distances between SNPs are not taken into consideration in the sliding window approach, it is likely that a large fraction of these regions represents false positives. However, all seven regions detected by the scan statistic were also detected by the sliding window approach. The linkage disequilibrium (LD) patterns within the seven regions were highly variable indicating that the clusters of SNP associations were not due to LD alone. The scan statistic developed here can be used to make gene-based or region-based SNP inferences about disease association.

Chromosome Mapping↗

Statistics-based approach for aneurysm volume measurements.

PURPOSE: To evaluate the ability of high-resolution MRA to monitor changes in intracranial aneurysm volume, and devise a highly reliable technique for obtaining these measurements. MATERIALS AND METHODS: To obtain a baseline estimate of the repeatability of MRA scans and validate the statistics-based technique for aneurysm volume measurement, multiple scans were obtained on individual subjects over a period of up to 1 year. These 3D MRA data sets were coregistered and then analyzed using the volumetric analysis of segmented data and the proposed statistical method. RESULTS: It was shown that high-resolution MRA provides highly repeatable data sets. Both methods used for the aneurysm volume measurements showed consistent results. However, the proposed statistical method had lower error and was much less sensitive to the choice of segmentation parameter than the volumetric analysis of segmented data. A change of 1 mm in the average radius of the aneurysm was detectable with the statistics-based technique. CONCLUSIONS: This study demonstrates that the statistical method of aneurysm volume measurement in high-resolution MRA allows reliable and accurate assessments of aneurysm volume changes.

Cerebrovascular Circulation↗

The statistical performance of an MCF-7 cell culture assay evaluated using generalized linear mixed models and a score test.

Biological assays often utilize experimental designs where observations are replicated at multiple levels, and where each level represents a separate component of the assay's overall variance. Statistical analysis of such data usually ignores these design effects, whereas more sophisticated methods would improve the statistical power of assays. This report evaluates the statistical performance of an in vitro MCF-7 cell proliferation assay (E-SCREEN) by identifying the optimal generalized linear mixed model (GLMM) that accurately represents the assay's experimental design and variance components. Our statistical assessment found that 17beta-oestradiol cell culture assay data were best modelled with a GLMM configured with a reciprocal link function, a gamma error distribution, and three sources of design variation: plate-to-plate; well-to-well, and the interaction between plate-to-plate variation and dose. The gamma-distributed random error of the assay was estimated to have a coefficient of variation (COV) = 3.2 per cent, and a variance component score test described by X. Lin found that each of the three variance components were statistically significant. The optimal GLMM also confirmed the estrogenicity of five weakly oestrogenic polychlorinated biphenyls (PCBs 17, 49, 66, 74, and 128). Based on information criteria, the optimal gamma GLMM consistently out-performed equivalent naive normal and log-normal linear models, both with and without random effects terms. Because the gamma GLMM was by far the best model on conceptual and empirical grounds, and requires only trivially more effort to use, we encourage its use and suggest that naive models be avoided when possible.

Biological Assay↗

Evaluation of an adjusted chi-square statistic as applied to observational studies involving clustered binary data.

A simple adjustment to the Pearson chi-square test has been proposed for comparing proportions estimated from clustered binary observations. However, the assumptions needed to assure the validity of this test have not yet been thoroughly addressed. These assumptions will hold for experimental comparisons, but could be violated for some observational comparisons. In this paper we investigate the conditions under which the adjusted chi-square statistic is valid and examine its performance when these assumptions are violated. We also introduce some alternative test statistics that do not require these assumptions. The test statistics considered are then compared through simulation and an example presented based on real data. The simulation study shows that the adjusted chi-square statistic generally produces empirical type I errors close to nominal under the assumption of a common intracluster correlation coefficient. Even if the intracluster correlations are different, the adjusted chi-square statistic performs well when the groups have equal numbers of clusters.

Chi-Square Distribution↗

Thresholding of statistical maps in functional neuroimaging using the false discovery rate.

Finding objective and effective thresholds for voxelwise statistics derived from neuroimaging data has been a long-standing problem. With at least one test performed for every voxel in an image, some correction of the thresholds is needed to control the error rates, but standard procedures for multiple hypothesis testing (e.g., Bonferroni) tend to not be sensitive enough to be useful in this context. This paper introduces to the neuroscience literature statistical procedures for controlling the false discovery rate (FDR). Recent theoretical work in statistics suggests that FDR-controlling procedures will be effective for the analysis of neuroimaging data. These procedures operate simultaneously on all voxelwise test statistics to determine which tests should be considered statistically significant. The innovation of the procedures is that they control the expected proportion of the rejected hypotheses that are falsely rejected. We demonstrate this approach using both simulations and functional magnetic resonance imaging data from two simple experiments.

Adult↗

A unified approach to study hypervariable polymorphisms: statistical considerations of determining relatedness and population distances.

Relatedness between individuals as well as evolutionary relationships between populations can be studied by comparing genotypic similarities between individuals. When hypervariable loci are used to describe genotypes, it is shown that both of these problems can be approached with a unified theory based on allele sharing between individuals. The distributions of the number of shared alleles between individuals indicate their kin relationships. Extending this, we obtain statistics for genetic distances between populations based on average number of alleles shared between individuals within and between two different populations. Traditional statistical inferential procedure can be used to establish specific kinship relationships between individuals. We derive estimates of the number of hypervariable loci needed for a specified reliability of such an inference. Evolutionary dynamics of genetic distance statistics based on allele sharing is also studied. It shows that such measures of genetic distances remain linear with the time of divergence for a period comparable to that of the gene frequency-based measures of genetic distances. Statistical properties of measures based on allele sharing establish that for using such summary statistics it is not necessary to know the full characteristics of all loci used. It is enough to know the degree of heterozygosity per locus and the number of loci. Therefore, in principle, this approach can also be used for DNA fingerprinting data in the studies of relatedness between individuals as well as between populations. The possible compromising features of multilocus DNA fingerprinting data are also discussed.

Alleles↗

Visual discrimination of textures with identical third-order statistics.

We found a new class of two-dimensional random textures with identical third-order statistics that can be effortlessly discriminated. Discrimination is based on local "granularity" differences between these iso-trigon texture pairs. This is the more surprising since it is commonly assumed that texture granularity (grain) is determined by the power spectrum which, in turn, can be obtained from the second-order statistics. Because textures with identical third-order statistics must have identical second-order statistics (i.e., identical power spectra), visible texture granularity is not controlled by power spectra, and not even by third-order statistics.

Discrimination, Psychological↗

Correspondence between statistically derived behavior problem syndromes and child psychiatric diagnoses in a community sample.

The correspondence between Diagnostic and Statistical Manual (3rd ed.) (DSM-III) diagnoses and statistically derived syndromes was examined within a community sample of children and adolescents in Puerto Rico. Specifically, the extent to which behavior dimensions, derived from the Child Behavior Checklist and the Youth Self-Report, corresponded to psychiatric diagnoses, derived from parent and child versions of the Diagnostic Interview Schedule for Children, was examined. The alternative approaches for assessing psychopathology in children and adolescents were compared against external validators. The results indicated a meaningful convergence between DSM-III diagnoses and statistical syndromes; however, a one-to-one correspondence did not emerge. Little evidence was found for "diagnostic thresholds." There was no evidence of the superiority of either the statistically derived syndromes or the DSM-III diagnoses. The incorporation of a measure of impairment improved the validity of both approaches. Adding parental reports to the self-reports of adolescents yielded little gain in the validity of either the statistical or diagnostic approach. The implications for the definition and assessment of child and adolescent disorders are discussed.

Adolescent↗

Perceptual and statistical analysis of cardiac phase and amplitude images.

A perceptual experiment was conducted using cardiac phase and amplitude images. Estimates of statistical parameters were derived from the images and the diagnostic potential of human and statistical decisions compared. Five methods were used to generate the images from 75 gated cardiac studies, 39 of which were classified as pathological. The images were presented to 12 observers experienced in nuclear medicine. The observers rated the images using a five-category scale based on their confidence of an abnormality presenting. Circular and linear statistics were used to analyse phase and amplitude image data, respectively. Estimates of mean, standard deviation (SD), skewness, kurtosis and the first term of the spatial correlation function were evaluated in the region of the left ventricle. A receiver operating characteristic analysis was performed on both sets of data and the human and statistical decisions compared. For phase images, circular SD was shown to discriminate better between normal and abnormal than experienced observers, but no single statistic discriminated as well as the human observer for amplitude images.

Analysis of Variance↗

Alignment statistic for identifying related protein sequences.

Closely related proteins show an obvious kinship by having numerous matching amino acids in their aligned sequences. Kinship between anciently separated proteins requires a statistical evaluation to rule out fortuitous similarities. A simple statistic is developed which assumes equal probability for all codon pairs, and a table of critical values for amino acid sequence alignments of lengthnments of length 200 or less is presented. Applying this statistic to V and C regions of immunoglobulin chains, aligned on the basis of shared features of three-dimensional structure, provides evidence that the V and C sequences descended from a common ancestor. Similarly the distant evolutionary relationship of dehydrogenases, flavdoxin, and subtilisin, suggested by structural alignments, is verified. On the other hand, the statistic does not verify a common evolutionary origin for the heme binding pocket in globins and cytochrome bs. Empirical evidence from the distribution of MMD values of amino acid pairs in comparisons of misaligned polypeptide chains and from Monte Carlo trials of sequences aligned with arbitrary gaps supports the validity of the statistic.

Amino Acid Sequence↗

Statistics for colon and rectal surgeons.

To read the literature critically, it is important to understand the fundamental principles of statistical analysis. In a review of 190 articles of interest to colon and rectal surgeons, it was found that only 29 percent of articles contained no statistics at all. Twenty-four percent contained descriptive statistics (mean, median, standard, deviation) and 47 percent contained the use of formal statistical tests. In this article, several basic statistical concepts are reviewed. The three most frequently used tests in the colon and rectal literature, the t test, chi-square, and nonparametric tests, are described. An example of each test is given to illustrate how the test is used and how the results are interpreted.

Colorectal Surgery↗