Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Statistics in medical journals.

The general standard of statistics in medical journals is poor. This paper considers the reasons for this with illustrations of the types of error that are common. The consequences of incorrect statistics in published papers are discussed; these involve scientific and ethical issues. Suggestions are made about ways in which the standard of statistics may be improved. Particular emphasis is given to the necessity for medical journals to have proper statistical refereeing of submitted papers.

Ethics↗

Identifying important results from multiple statistical tests.

When many statistical tests are performed simultaneously, the overall chance of a type I error (incorrect rejection of a true null hypothesis) can substantially exceed the nominal error rate used in each individual test. Numerous techniques exist to adjust results of individual tests to control this problem. In general, these techniques apply a more stringent criterion of statistical significance (a smaller P-value) to each individual test than normally needed to maintain the experimentwise type I error. With an analysis that seeks to identify results for further research, however, such a conservative technique may not be appropriate. We present a new approach that uses a mixture of several distributions to model the set of P-values or of test statistics. One component models the results consistent with a failure to reject the null hypothesis, while the other distribution(s) in the mixture represent results inconsistent with the null hypothesis. These latter results may not achieve statistical significance based on a conventional P-value. We illustrate the use of the method on national mortality data and on several data sets analysed previously.

Cause of Death↗

A comparison of statistical methods for combining event rates from clinical trials.

We compare two statistical methods for combining event rates from several studies. Both methods treat each study as a separate stratum. The Peto-modified Mantel-Haenszel (Peto) method estimates a combined odds ratio assuming homogeneity across strata and provides a test for heterogeneity. The DerSimonian and Laird modified Cochran method (D&L) produces a weighted average of rate differences, where the weights allow for among-study variability. We analyse 22 meta-analyses from ten reports by both methods. The pooled estimates are divided by their standard errors to produce a Z-statistic. A t-test comparing Z-statistics from all 22 studies suggests that the D&L method tends to be more conservative [d(Peto - D&L) = 0.29, t = 2.53, p = 0.02]. For a subset of 14 non-heterogeneous studies, the difference is smaller and non-significant (d = 0.09, t = 0.72, p = 0.49). The results from the methods correlate well (r = 0.66 for all 22 studies, r = 0.95 for 14 non-heterogeneous studies). Thus, the presence of heterogeneity influences our conclusion. We discuss the statistical and scientific implications of these findings.

Clinical Trials as Topic↗

Applications of global statistics in analysing quality of life data.

Quality of life (QOL) instruments usually consist of a number of components, each of which deals specifically with a particular functionally related dysfunction. In a clinical trial whose primary aim is the evaluation of the treatment by means of QOL instruments, analysis of each of the components usually consists of either univariate analysis of variance (ANOVA) or some non-parametric methods. This multiple testing approach can produce an increase in false positive findings. One attempt to correct for this is the Bonferroni adjustment. Another approach is to apply global statistics (parametric or non-parametric) for the null hypothesis of no treatment difference versus the alternative hypothesis that one treatment is uniformly better than the other for QOL instruments as a whole. Data from a randomized double-blind trial of 111 congestive heart failure patients, which involved four QOL instruments, were analysed with univariate ANOVA, Bonferroni adjustment, parametric and non-parametric global statistics. The global statistics complemented the univariate methods and made the presentation of QOL data very effective. I recommend the general use of global statistics in analysis of QOL data.

Attitude to Health↗

Who should teach medical statistics, when, how and where should it be taught?

As far as we can tell, there is no single answer to any of these questions. This paper describes some of the approaches adopted in the teaching of medical statistics in the U.K. medical schools. It is suggested that collaboration between non-statistically qualified teachers and medical statisticians is beneficial, with an emphasis in the application of statistical principles to interesting and 'relevant' medical topics. A block of 'laboratory based' teaching in the early years may be followed by occasional, clinically focused sessions later in the undergraduate course. Research oriented courses, made available when postgraduates have a real and pressing need for information, are thought likely to be most valuable and rewarding for students and statisticians. It is thought that the use of information technology to improve the communication of concepts and 'facts' during lectures, and for ad hoc enquiries by students, is likely to make the most of the limited teaching resources. The future is thought to be in the greater use of small group or individualized teaching which confirms or tests the knowledge gained from students use of I.T. supported activities. Unless lecturers collaborate in evaluative studies which compare different teaching methods, it will never be possible to provide valid generalizable advice to teachers of medical statistics.

Education, Medical↗

Statistical perspectives on confidentiality and data access in public health.

Confidentiality and disclosure limitation are topics that are inherently statistical but, until recently, they have received limited attention from statistical methodologists. That situation has changed considerably in the present decade. In this paper, we provide an introduction and overview of some statistical disclosure limitation issues that are of special relevance to public health studies and surveys, and the linkages to current research on bounds for multi-dimensional contingency table entries and 'simulated' categorical data. We also describe how these research methods relate to a new data access query system being developed for use by NCHS and other statistical agencies.

Confidentiality↗

ReMeDy: A Flexible Statistical Framework for Region-Based Detection of DNA Methylation Dysregulation.

Region-based epigenome-wide association studies have demonstrated improved statistical power and biological interpretability compared with probe-wise analyses of DNA methylation data. However, most existing region-based methods characterize methylation dysregulation primarily through changes in mean methylation levels associated with a phenotype of interest. Substantial evidence indicates that phenotype-associated methylation alterations may also manifest through changes in methylation variability or through joint shifts in mean and variability. Despite this, no existing statistical framework jointly models mean-variance methylation changes in a region-based manner. We propose ReMeDy, a flexible statistical framework that uses a hierarchical likelihood approach within a generalized linear model setting to identify differentially methylated regions, variably methylated regions, and regions exhibiting joint differential and variable methylation at a genome-wide scale. Unlike existing models, ReMeDy operates directly on biologically defined co-methylated regions, allowing it to naturally capture spatial correlation inherent in DNA methylation array data, while avoiding reliance on heuristic, user-defined tuning parameters such as smoothing spans and kernel bandwidths that can substantially influence results and introduce subjectivity. Through extensive simulation studies and comprehensive benchmarking against popular models, we demonstrate that ReMeDy maintains false discovery and Type-I error rates at nominal levels while achieving consistently higher statistical power across a wide range of realistic scenarios. Application to population-level DNA methylation data further shows that ReMeDy identifies biologically meaningful regions and pathways implicated in complex human diseases that are not captured by conventional mean-based analyses alone. ReMeDy is implemented as an open-source R package and is freely available at https://github.com/SChatLab/ReMeDy.

DNA Methylation↗

Vital statistics linked birth/infant death and hospital discharge record linkage for epidemiological studies.

A methodology for linking vital statistics linked birth/death data and hospital discharge data is described. The resulting data set combines information on a neonate's sociodemographic characteristics, prenatal care, and mortality aspects and connects it to detailed health outcome and resource utilization data, thus establishing an extensive database for epidemiological studies. In the absence of a universal identifier common to both databases, our linkage strategy relied on using a virtual identifier based on variables common to both data sets. In the case of multiple incidences of the same virtual identifier we used secondary health status information to optimize the likelihood of linking low birth weight or premature infants in one database to infants of similar health status in the other while randomizing cases in which no secondary information was present. Applying our method to the 1992 California birth cohort, we could link 563,114 out of 571,189 eligible births (98.59%). Of these links, 91.2% were established on the basis of unique virtual identifiers. The link was internally consistent and no bias was evident when comparing variable distributions for all single live births in the vital statistics linked birth/death file and linked births in the linked vital statistics linked birth/death and hospital discharge file. Multiple imputation techniques showed that the prediction error incurred by randomization was negligible. Even though computationally intensive, our method for linking the vital statistics linked birth/death file and the hospital discharge file appeared to be effective. However, it is important to be aware of the limitations of the resulting data set, in particular the fact that it cannot be used for tracking individual cases. The method provides a database suitable for a variety of perinatal epidemiological analyses, such as descriptive studies of disease distribution in neonates, studies of the geographic distribution of disease, and studies of the relationship between risk and outcome.

Algorithms↗

Statistical sulcal shape comparisons: application to the detection of genetic encoding of the central sulcus shape.

Principal Component Analysis allows a quantitative description of shape variability with a restricted number of parameters (or modes) which can be used to quantify the difference between two shapes through the computation of a modal distance. A statistical test can then be applied to this set of measurements in order to detect a statistically significant difference between two groups. We have applied this methodology to highlight evidence of genetic encoding of the shape of neuroanatomical structures. To investigate genetic constraint, we studied if shapes were more similar within 10 pairs of monozygotic twins than within interpairs and compared the results with those obtained from 10 pairs of dizygotic twins. The statistical analysis was performed using a Mantel permutation test. We show, using simulations, that this statistical test applied on modal distances can detect a possible genetic encoding. When applied to real data, this study highlighted genetic constraints on the shape of the central sulcus. We found from 10 pairs of monozygotic twins that the intrapair modal distance of the central sulcus was significantly smaller than the interpair modal distance, for both the left central sulcus (Z = -2.66; P < 0.005) and the right central sulcus (Z = -2.26; P < 0.05). Genetic constraints on the definition of the central sulcus shape were confirmed by applying the same experiment to 10 pairs of normal young individuals (Z = -1.39; Z = -0.63, i.e., values not significant at the P < 0.05 level) and 10 pairs of dizygotic twins (Z = 0.47; Z = 0.03, i.e., values not significant at the P < 0.05 level).

Adult↗

Selection of an adaptive test statistic for use with multiple comparison analyses of neuroimaging data.

Statistical analysis of neuroimages is commonly approached with intergroup comparisons made by repeated application of univariate or multivariate tests performed on the set of the regions of interest sampled in the acquired images. The use of such large numbers of tests requires application of techniques for correction for multiple comparisons. Standard multiple comparison adjustments (such as the Bonferroni) may be overly conservative when data are correlated and/or not normally distributed. Resampling-based step-down procedures that successfully account for unknown correlation structures in the data have recently been introduced. We combined resampling step-down procedures with the Minimum Variance Adaptive method, which allows selection of an optimal test statistic from a predefined class of statistics for the data under analysis. As shown in simulation studies and analysis of autoradiographic data, the combined technique exhibits a significant increase in statistical power, even for small sample sizes (n = 8, 9, 10).

Algorithms↗

Statistical parametric mapping of hypoxic tissue identified by [(18)F]fluoromisonidazole and positron emission tomography following acute ischemic stroke.

Positron emission tomography (PET) and the ligand [(18)F]fluoromisonidazole ((18)F-FMISO) have been used to image hypoxic tissue in the brain following acute stroke. Existing region of interest (ROI)-based methods of analysis are time consuming and operator-dependent. We describe and validate a method of statistical parametric mapping to identify regions of increased (18)F-FMISO uptake. The (18)F-FMISO PET images were transformed into a standardized coordinate space and intensity normalized. Then t statistic maps were created using a pooled estimate of variance. Statistical inference was based on the theory of Gaussian Random Fields. We examined the homogeneity of variance in normal subjects and the influence of normalization by mean whole brain activity versus mean activity in the contralateral hemisphere. Validity of the distributional assumptions inherent in parametric analysis was tested by comparison with a non-parametric method. The results of parametric analysis were also compared with those obtained with the existing ROI-based method. Variance in uptake at each voxel in normal subjects was homogeneous and not affected by mean voxel activity or distance from the centre of the image. The method of normalization influenced results significantly. Normalization by whole brain mean activity resulted in a smaller volume of tissue being classified as hypoxic compared to normalisation by mean activity in the contralateral hemisphere. The ROI-based method was subject to interobserver variability with a coefficient of variability of 16%. The volumes of hypoxic tissue identified by parametric and nonparametric methods were highly correlated (r = 0.99). These findings suggest that using a pooled variance and contralateral hemisphere normalisation, statistical parametric mapping can be used to objectively identify regions of increased (18)F-FMISO uptake following acute stroke in individual subjects.

Acute Disease↗

Development of a friendly, self-teaching, interactive statistical package for analysis of clinical research data. The BRIGHT STAT-PACK.

We have developed a new statistical analysis package for use by the clinical investigator, the clinician, and the laboratory researcher. This package attempts to implement the following philosophy: The programs should be essentially self-teaching, friendly, and forgiving; the programs should educate the user regarding the underlying theory, assumptions, and interpretation of the statistical methods involved; the programs should automatically test relevant assumptions and warn the user when these assumptions appear to have been violated; the programs should make recommendations about the availability of alternative statistical methods and automatically perform such analyses when indicated; the programs should interpret the results; and the programs should mimic, insofar as possible, the logic used in a routine, elementary statistical consultation. Several programs have been developed, extensively tested, and used.

Computer-Assisted Instruction↗

The multidisciplinary feeding profile: a statistically based protocol for assessment of dependent feeders.

The development of the multidisciplinary feeding profile entailed a level of statistical analyses not commonly utilized in test development. This paper describes the statistical analyses and offers an explanation of why specific statistical tests were chosen. It also serves to identify where clinical knowledge and experience overrode specific statistical tests.

Persons with Disabilities↗

A statistical approach for classifying change in cognitive function in individuals following pharmacologic challenge: an example with alprazolam.

INTRODUCTION: The effects of any drug treatment on cognitive function are typically studied in groups of subjects. Observations made about the behavior of the drug, in the study sample, are then generalized to the population from which the sample was drawn. However, the magnitude and pharmacodynamic qualities of the response to many central nervous system-active drugs are known to vary in the population. Therefore, it is useful to consider statistical models for the detection of cognitive change in response to a drug treatment in individual subjects. MATERIALS AND METHODS: In this report, we first outline the statistical assumptions and requirements for the reliable estimation of clinically relevant individual change in cognition. We then used the sedative benzodiazepine, alprazolam, as a pharmacologic challenge in healthy volunteer subjects to test our statistical model, using a parallel groups placebo-controlled study design. After treatment, the nature and severity of alprazolam-induced cognitive change was determined for each individual. RESULTS: Our proposed method and analysis showed an excellent sensitivity and specificity for alprazolam-related cognitive deterioration in individuals. DISCUSSION AND CONCLUSIONS: These findings, although preliminary, suggest that statistically reliable decisions about the effects of sedative drugs on cognition can be made for individuals.

Adult↗

Statistical downscaling of general-circulation-model- simulated average monthly air temperature to the beginning of flowering of the dandelion (Taraxacum officinale) in Slovenia.

Phenological observations are a valuable source of information for investigating the relationship between climate variation and plant development. Potential climate change in the future will shift the occurrence of phenological phases. Information about future climate conditions is needed in order to estimate this shift. General circulation models (GCM) provide the best information about future climate change. They are able to simulate reliably the most important mean features on a large scale, but they fail on a regional scale because of their low spatial resolution. A common approach to bridging the scale gap is statistical downscaling, which was used to relate the beginning of flowering of Taraxacum officinale in Slovenia with the monthly mean near-surface air temperature for January, February and March in Central Europe. Statistical models were developed and tested with NCAR/NCEP Reanalysis predictor data and EARS predict and data for the period 1960-1999. Prior to developing statistical models, empirical orthogonal function (EOF) analysis was employed on the predictor data. Multiple linear regression was used to relate the beginning of flowering with expansion coefficients of the first three EOF for the Janauary, Febrauary and March air temperatures, and a strong correlation was found between them. Developed statistical models were employed on the results of two GCM (HadCM3 and ECHAM4/OPYC3) to estimate the potential shifts in the beginning of flowering for the periods 1990-2019 and 2020-2049 in comparison with the period 1960-1989. The HadCM3 model predicts, on average, 4 days earlier occurrence and ECHAM4/OPYC3 5 days earlier occurrence of flowering in the period 1990-2019. The analogous results for the period 2020-2049 are a 10- and 11-day earlier occurrence.

Asteraceae↗

The distribution of fetal death in control mice and its implications on statistical tests for dominant lethal effects.

In dominant lethal testing fetal death is generally assumed to follow either a Poisson or binomial distribution. However, both of these models were found to be inappropriate when three large sets of mouse control data and other data sets from the literature were examined. The validity of statistical test procedures based on these inappropriate models was then studied in detail. It was found that chi-square tests (which assume an underlying binomial distribution) may seriously exaggerate the level of significance and hence should not be used. In contrast, the inappropriateness of the underlying Poisson or binomial model appeared to have little effect on the validity of pairwise comparisons by analysis of variance procedures. Unlike chi-square, these procedures regard the pregnant female rather than the individual implant as the experimental unit. However, a statistical analysis of dominant lethal data generally involves more than a series of pairwise comparisons, and it is unclear how an invalid underlying model may affect statistical test procedures in this more complex situation. Moreover, it is difficult to justify the use of statistical models that are demonstrably invalid when a reasonable alternative exists. Thus, until a satisfactory parametric model can be found and appropriate test procedures derived, we prefer to analyze dominant lethal data by non-parametric (distribution-free) methods. Proportion of dead implants per female appears to be a more meaningful measure of fetal death than number of dead implants per female for several seasons which include (1) analyses based on proportions take the total number of implants per female into account and (2) analyses based on proportions make more reasonable assumptions concerning pre-implantation losses and are more powerful when such losses occur. Despite our concern with the appropriateness of the underlying model, in practice we have found few instances in which non-parametric and analysis of variance procedures have led to markedly different conclusions.

Animals↗

Discrimination of changes in the second-order statistics of natural and synthetic images.

It has been suggested that the second-order statistics of different natural images are all remarkably similar and that neurones and channels in the visual system may exploit this similarity. We have measured the ability of human observers to discriminate changes in these statistics using different natural and synthetic stimulus images and have found that the dependence of their discrimination thresholds upon the reference second-order statistics is similar in form, for both kinds of stimuli. However, there is some variety in the magnitudes of the thresholds for the natural stimulus images; in fact, the second-order statistics of different natural images are more diverse than previously suggested. The discrimination task can be modelled as the discrimination of changes in local contrast within restricted spatial frequency bands and is similar to the discrimination of blur.

Contrast Sensitivity↗

The autopsy and vital statistics.

Vital statistics in the United States are collected through a decentralized, cooperative system of various levels of government administrated by the National Center for Health Statistics. Although registration of all deaths is virtually complete and demographic items are accurate, the reliability of cause of death data is hampered by the current state of medical knowledge, the incompleteness of information available at the time of death, the way in which physicians complete death certificates, and the system of classification of underlying cause. The need for quality assurance in national cause of death statistics can be met in large part by connecting the autopsy to the mainstream of vital statistics. Through case by case individual linkage of death certificates and autopsies in designated demographic and/or geographic areas, a representative, continuously collected, population-based system of aggregated autopsy data would be created. Demographic and clinical selection bias should be checked and adjusted through traditional methods of epidemiologic standardization. Such a use of autopsy information could further pathology's goals of understanding disease and improving the public health.

Autopsy↗