Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “STATISTICS”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Simple solution to a common statistical problem: interpreting multiple tests.

BACKGROUND: The misinterpretation of the results of multiple statistical tests is an error commonly made in scientific literature. When testing several outcome variables simultaneously, many researchers declare a statistically significant result for each test having a P value of <0.05, for example. This approach ignores the fact that, based on a probability result called the Bonferroni inequality, the risk of incorrectly declaring as significant > or =1 test result increases with the number of tests conducted. The implication of this practice is that many scientific results are presented as statistically significant when the underlying data do not adequately support such a claim (sometimes referred to as false-positive results). Although the sequentially rejective Bonferroni test is well known among statisticians, it is not used routinely in scientific literature. OBJECTIVE: The intent of this article was to increase the awareness and understanding of the sequentially rejective Bonferroni test, thereby expanding its use. METHODS: This article describes the statistical problem and demonstrates how the use of the sequentially rejective Bonferroni test ensures that incorrect declarations of statistical significance for > or =1 test result are bounded by 0.05, for example. CONCLUSION: The sequentially rejective Bonferroni test is an easily applied, versatile statistical tool that enables researchers to make simultaneous inferences from their data without risking an unacceptably high overall type I error rate.

Confidence Intervals↗

Statistical mechanical treatment of protein conformation. 5. A multistate model for specific-sequence copolymers of amino acids.

One-dimensional short-range interaction models for specific-sequence copolymers of amino acids have been developed in this series of papers. In the present paper, a multistate model (involving right-handed helical (hR), extended (epsilon), chain-reversal (R and S), left-handed helical (hL), right-handed bridge-region (zota R), left-handed bridge-region (zota L), and coil (or other) (c) states) is developed for the prediction of protein backbone conformation. This model involves ten parameters (WhR, UPSILONHR, V epsilon, VR, VS, WhL, VhL, U zota R, U zota L, and Uc) and requires a 10X10 statistical weight matrix. Assuming that the left-handed helical sequence cannot occur in proteins, this 10X10 matrix can be reduced to a 9X9 matrix with nine parameters (WhR, VhR, V epsilon, VR, VS, VhL, U zota R, U zota L, and Uc). A nearest neighbor approximation of this multistate model is also formulated; with the omission of left-handed helical sequences, and the inclusion of the left-handed bridge region in the c state, this approximate model requires a 7X7 matrix with statistical weights WhR, VhR, VS, VhL, U zota R, and Uc, expressed as values relative to the statistical weight of the epsilon state. The statistical weights for the multistate model are evaluated from the atomic coordinates of the X-ray structures of 26 native proteins. These statistical weights and the multistate model are applied in the prediction of the backbone conformations of proteins. The conformational probabilities of finding a residue in hR, epsilon, R, S, hL, zota R, or c states, defined as relative values with respect to their average values over the whole molecule, are calculated for bovine pancreatic trypsin inhibitor and clostridial flavodoxin, in order to select the most probable conformation for each residue of these proteins. The predicted results are compared to experimental observations and are discussed together with the reliability of the statistical weights. In the Appendix, the property of asymmetric nucleation of helical sequences is introduced into the (nearest neighbor) multistate model.

Amino Acid Sequence↗

The statistical significance of protein identification results as a function of the number of protein sequences searched.

The potential for obtaining a true mass spectrometric protein identification result depends on the choice of algorithm as well as on experimental factors that influence the information content in the mass spectrometric data. Current methods can never prove definitively that a result is true, but an appropriate choice of algorithm can provide a measure of the statistical risk that a result is false, i.e., the statistical significance. We recently demonstrated an algorithm, Probity, which assigns the statistical significance to each result. For any choice of algorithm, the difficulty of obtaining statistically significant results depends on the number of protein sequences in the sequence collection searched. By simulations of random protein identifications and using the Probity algorithm, we here demonstrate explicitly how the statistical significance depends on the number of sequences searched. We also provide an example on how the practitioner's choice of taxonomic constraints influences the statistical significance.

Algorithms↗

Statistical power when testing for genetic differentiation.

A variety of statistical procedures are commonly employed when testing for genetic differentiation. In a typical situation two or more samples of individuals have been genotyped at several gene loci by molecular or biochemical means, and in a first step a statistical test for allele frequency homogeneity is performed at each locus separately, using, e.g. the contingency chi-square test, Fisher's exact test, or some modification thereof. In a second step the results from the separate tests are combined for evaluation of the joint null hypothesis that there is no allele frequency difference at any locus, corresponding to the important case where the samples would be regarded as drawn from the same statistical and, hence, biological population. Presently, there are two conceptually different strategies in use for testing the joint null hypothesis of no difference at any locus. One approach is based on the summation of chi-square statistics over loci. Another method is employed by investigators applying the Bonferroni technique (adjusting the P-value required for rejection to account for the elevated alpha errors when performing multiple tests simultaneously) to test if the heterogeneity observed at any particular locus can be regarded significant when considered separately. Under this approach the joint null hypothesis is rejected if one or more of the component single locus tests is considered significant under the Bonferroni criterion. We used computer simulations to evaluate the statistical power and realized alpha errors of these strategies when evaluating the joint hypothesis after scoring multiple loci. We find that the 'extended' Bonferroni approach generally is associated with low statistical power and should not be applied in the current setting. Further, and contrary to what might be expected, we find that 'exact' tests typically behave poorly when combined in existing procedures for joint hypothesis testing. Thus, while exact tests are generally to be preferred over approximate ones when testing each particular locus, approximate tests such as the traditional chi-square seem preferable when addressing the joint hypothesis.

Animals↗

[Statistical aspects in planning psychotropic drug trials (author's transl)].

It has become virtually unthinkable to objectively judge trials of comparisons without statistical inference. For studies involving psychopharmaceuticals the requirements have especially grown, since variations of efficacy are difficult to differentiate and objectivate. The basic principles of statistical planning must be applied to studies with psychopharmaceuticals. On the basis of these general principles the relevant points of the clinical and statistical models and their relationship to one another are discussed. Special aspects of the study objectives, the clinical model, the medical trial plan, the statistical model and design, the execution of the trials, statistical evaluation and its interpretation are presented. The setting up of hypotheses, balancing and maintenance of factors of disturbances, sample sizes, randomizing techniques and the confrontation of clinical relevance with statistical significance are some of the important points discussed.

Clinical Trials as Topic↗

Comparison of nonparametric statistics for detection of linkage in nuclear families: single-marker evaluation.

We have evaluated 23 different statistics, from a total of 10 popular software packages for model-free linkage analysis of nuclear-family data, by applying them to single-marker data simulated under several two-locus disease models. The statistics that we examined fall into two broad categories: (1) those that test directly for increased identity-by-state or identity-by-descent sharing (by use of the programs APM, Genetic Analysis System [GAS] SIBSTATE and SIBDES, SAGE SIBPAL, ERPA, SimIBD, and Genehunter NPL) and (2) those that are based on likelihood-ratio tests and that report LOD scores (by use of the programs Splink, SIBPAIR, Mapmaker/Sibs, ASPEX, and GAS SIBMLS). For each of eight two-locus disease models, we analyzed six data sets; the first three data sets consisted of two-child families with both sibs affected and zero, one, or both parents typed, whereas the other three data sets consisted of four-child families with at least two affected sibs and zero, one, or both parents typed. We report false-positive rates, overall rank by power, and the power for each statistic. We give rough recommendations regarding which programs provide the most powerful tests for linkage, as well as the programs to be avoided under certain conditions. For the likelihood-ratio-based statistics, we examined the effects of various treatments of sibships with multiple affected individuals. Finally, we explored the use of some simple two-of-three composite statistics and found that such tests are of only marginal benefit over the most powerful single statistic.

Adult↗

The effect of statistical uncertainty on inverse treatment planning based on Monte Carlo dose calculation.

The effect of the statistical uncertainty, or noise, in inverse treatment planning for intensity modulated radiotherapy (IMRT) based on Monte Carlo dose calculation was studied. Sets of Monte Carlo beamlets were calculated to give uncertainties at Dmax ranging from 0.2% to 4% for a lung tumour plan. The weights of these beamlets were optimized using a previously described procedure based on a simulated annealing optimization algorithm. Several different objective functions were used. It was determined that the use of Monte Carlo dose calculation in inverse treatment planning introduces two errors in the calculated plan. In addition to the statistical error due to the statistical uncertainty of the Monte Carlo calculation, a noise convergence error also appears. For the statistical error it was determined that apparently successfully optimized plans with a noisy dose calculation (3% 1sigma at Dmax), which satisfied the required uniformity of the dose within the tumour, showed as much as 7% underdose when recalculated with a noise-free dose calculation. The statistical error is larger towards the tumour and is only weakly dependent on the choice of objective function. The noise convergence error appears because the optimum weights are determined using a noisy calculation, which is different from the optimum weights determined for a noise-free calculation. Unlike the statistical error, the noise convergence error is generally larger outside the tumour, is case dependent and strongly depends on the required objectives.

Algorithms↗

A statistical method for identifying differential gene-gene co-expression patterns.

MOTIVATION: To understand cancer etiology, it is important to explore molecular changes in cellular processes from normal state to cancerous state. Because genes interact with each other during cellular processes, carcinogenesis related genes may form differential co-expression patterns with other genes in different cell states. In this study, we develop a statistical method for identifying differential gene-gene co-expression patterns in different cell states. RESULTS: For efficient pattern recognition, we extend the traditional F-statistic and obtain an Expected Conditional F-statistic (ECF-statistic), which incorporates statistical information of location and correlation. We also propose a statistical method for data transformation. Our approach is applied to a microarray gene expression dataset for prostate cancer study. For a gene of interest, our method can select other genes that have differential gene-gene co-expression patterns with this gene in different cell states. The 10 most frequently selected genes, include hepsin, GSTP1 and AMACR, which have recently been proposed to be associated with prostate carcinogenesis. However, genes GSTP1 and AMACR cannot be identified by studying differential gene expression alone. By using tumor suppressor genes TP53, PTEN and RB1, we identify seven genes that also include hepsin, GSTP1 and AMACR. We show that genes associated with cancer may have differential gene-gene expression patterns with many other genes in different cell states. By discovering such patterns, we may be able to identify carcinogenesis related genes.

Algorithms↗

Empowering research: statistical power in general practice research.

BACKGROUND: Statistical power is a measure of the extent to which a study is capable of discerning differences or associations which exist within the population under investigation, and is of critical importance whenever a hypothesis is tested by statistics. Conventionally, studies should reach a power level of 0.8, such that four times out of five a false null hypothesis will be rejected by a study. Statistical power may most easily be increased by increasing sample size. OBJECTIVE: We aimed to assess the level of statistical power of general practice research. METHODS: A total of 1422 statistical tests in 85 quantitative original papers in the British Journal of General Practice were analysed for statistical power. RESULTS: The median power of tests analysed was 0.71, representing a slightly greater than two-thirds likelihood of rejecting false null hypotheses. Of 85 studies, 37 (44%) attained power of 0.8 or more. Ten studies had power of more than 0.99 suggesting 'over-powering'. Twenty-one of the papers surveyed (25%) had a likelihood of gaining significant results poorer than that obtained by tossing a coin when a null hypothesis is false. CONCLUSION: While achieving higher power than studies in similar surveys of other disciplines, the power of general practice research falls short of the 0.8 convention. Adequate power is essential so that effects which exist are not missed. Recommendations are made concerning power calculations prior to the start of research and reporting of results in journal articles.

Bias↗

Validity of cerebrovascular disease mortality statistics in Bulgaria.

BACKGROUND: Cerebrovascular disease (CVD) is the leading cause of death in Bulgaria but the increasing mortality could be explained by inaccuracy of the statistical data and so investigation of the validity of CVD mortality statistics is of primary importance. METHODS: The investigation comprised three phases. An adequate questionnaire, requiring a reliable decision on the presence/absence of CVD in deceased patients was developed. During the first phase the questionnaire was validated on the basis of 325 inpatients aged 20 years and over. In the second phase the applicability of the questionnaire was proved and verified in patients who died outside hospital. This was performed by using 'twin'-copies of each questionnaire, completed in the first phase. In the third phase the applicability of the questionnaire for evaluation of CVD mortality statistics was checked using a sample of 119 death certificates. Statistical analysis of the information from each of the three phases was intended to evaluate the sensitivity and specificity of the questionnaire and to assess the validity of mortality statistics. RESULTS: High sensitivity of the questionnaire was established and it remained at the same value during each of the three phases of the study. Specificity was considered lower when the questionnaire was applied for those who died outside hospital and when used by non-neurologists. An underestimation of CVD by 8.9% was obtained in the first phase and it amounted to 37.81% in the third phase. A high proportion of incomplete or unsystematically completed death certificates was found. This represents a potential source of inaccuracy in mortality statistics. CONCLUSION: The questionnaire developed for presence/absence of CVD in decreased patients proved to be a reliable instrument for certifying CVD mortality.

Aged↗

A statistical sampling algorithm for RNA secondary structure prediction.

An RNA molecule, particularly a long-chain mRNA, may exist as a population of structures. Further more, multiple structures have been demonstrated to play important functional roles. Thus, a representation of the ensemble of probable structures is of interest. We present a statistical algorithm to sample rigorously and exactly from the Boltzmann ensemble of secondary structures. The forward step of the algorithm computes the equilibrium partition functions of RNA secondary structures with recent thermodynamic parameters. Using conditional probabilities computed with the partition functions in a recursive sampling process, the backward step of the algorithm quickly generates a statistically representative sample of structures. With cubic run time for the forward step, quadratic run time in the worst case for the sampling step, and quadratic storage, the algorithm is efficient for broad applicability. We demonstrate that, by classifying sampled structures, the algorithm enables a statistical delineation and representation of the Boltzmann ensemble. Applications of the algorithm show that alternative biological structures are revealed through sampling. Statistical sampling provides a means to estimate the probability of any structural motif, with or without constraints. For example, the algorithm enables probability profiling of single-stranded regions in RNA secondary structure. Probability profiling for specific loop types is also illustrated. By overlaying probability profiles, a mutual accessibility plot can be displayed for predicting RNA:RNA interactions. Boltzmann probability-weighted density of states and free energy distributions of sampled structures can be readily computed. We show that a sample of moderate size from the ensemble of an enormous number of possible structures is sufficient to guarantee statistical reproducibility in the estimates of typical sampling statistics. Our applications suggest that the sampling algorithm may be well suited to prediction of mRNA structure and target accessibility. The algorithm is applicable to the rational design of small interfering RNAs (siRNAs), antisense oligonucleotides, and trans-cleaving ribozymes in gene knock-down studies.

Algorithms↗

Assessment of surveillance and vital statistics data for monitoring abortion mortality, United States, 1972-1975.

To assess the usefulness of vital statistics and surveillance for monitoring abortion mortality, the authors compared data from two systems of classification: 1) deaths classified according to the underlying cause by the National Center for Health Statistics (NCHS) under the International Classification of Disease, Adapted (ICDA) code numbers 640-645 (abortion) for 1972-1975; and 2) abortion-related deaths reported to the Center for Disease Control (CDC) through its epidemiologic surveillance of abortion mortality for the same years. Vital statistics classifications dealing with the underlying cause of death are based on criteria defined by ICDA guidelines applied to all available information listed on death certificates, and exclude some deaths classified as abortion-related by CDC. Surveillance classifications are based on broader criteria developed by CDC for expanded data gathered by individual case investigation. Results showed that the surveillance techniques had identified more deaths as abortion-related and had resolved more cases into the specific abortion categories of legal, illegal, and spontaneous than vital statistics tabulations based on death certificates. The authors estimate that the surveillance system alone reported 88% of all abortion-related deaths, the vital statistics system 52%, and the two systems combined a total of 94%. Inadequate physician documentation on the death certificate was the primary reason vital statistics data contained a smaller number of reported abortion deaths than surveillance data.

Abortion, Illegal↗

Use of the scan statistic to detect time-space clustering.

A test for time-space clustering is proposed based on the scan statistic, the maximum number of events in a 365-day period in each of several geographic units. The data under consideration should consist of the exact date and geographic unit for each event, and data should be available for several years for which the risk of disease can be assumed constant. The statistic is the ratio of the excess number of events summed over all the geographic regions, to the square root of the sum of the variances. This statistic is similar in construction to the Ederer-Myers-Mantel statistic (Biometrics 1964;20:626-38), but does not require that attention be limited to calendar years (January 1-December 31). Unlike other tests for time-space clustering, the scan statistic allows one to calculate measures of attributable risk and effect size. Data concerning adolescent suicide are used to illustrate the procedure. The tables and asymptotic formulas given for the mean and variance of the proposed statistic should be useful in the evaluation of both clustering in time and in time-space.

Adolescent↗

Statistical issues in long-term followup studies.

The thesis of this article is that since followup studies of patients with schizophrenia have provided a rich body of informative knowledge, emphasis should now be on hypothesis testing in future studies. The consumer of research needs to have a general appreciation of statistical thinking, design, and methods to make an informed synthesis of results presented. This article presents a brief discussion of relevant statistics and statistical issues, as free of jargon as possible. Given the multivariable nature of psychiatric research, the natural focus is on multivariate statistics. A strong background in statistics is not required to understand the methods described and how and why they might be applied. Reference material for more detailed discussions is provided. It is a truism to say that the highest quality scientific work is dependent on an intimate collaboration with the expert in statistics and the expert in clinical methods, and this merits repeating only because it is so often ignored.

Follow-Up Studies↗

Statistical issues in the analysis of low-dose endocrine disruptor data.

UNLABELLED: The National Institute of Environmental Health Sciences (NIEHS) and the U.S. Environmental Protection Agency (U.S. EPA) recently cosponsored the Endocrine Disruptors Low-Dose Peer REVIEW: The purpose of this meeting was to examine data supporting the presence or absence of low-dose effects of endocrine disruptors in specific studies and then to evaluate the likelihood and significance of these and/or other potential low-dose effects for humans. All invited speakers agreed to provide their raw data in advance of the meeting to a Statistics Subpanel, which was asked to reevaluate the authors' experimental design, data analysis, and interpretation of experimental results. The purpose of this statistical reevaluation was to provide an independent assessment of the experimental design and data analysis used in each of the studies and to identify key statistical issues relevant to the evaluation and interpretation of the data. This paper presents a summary of the Statistics Subpanel's evaluation. Specific examples are presented to illustrate problems that arose in the experimental design and data analysis of certain studies. The statistical principles and issues that are discussed in this paper are not unique to endocrine disruptor studies and should provide important guidelines regarding appropriate experimental design and statistical analysis for other types of laboratory investigations.

Analysis of Variance↗

Statistical approach to the phase problem.

The minimal function and its minimal principle employed in the traditional Shake-and-Bake algorithm rely on the probabilistic estimates of the cosines of the structure invariants. In this paper, a novel statistical approach to the phase problem, which utilizes statistical properties of the structure invariants, is proposed. The statistical maximal function and its maximal principle are formulated, and the corresponding statistical Shake-and-Bake algorithm and its associated statistical parameter-shift procedure are proposed and tested. The test results show that the statistical approach to the phase problem is a simple, reliable, less computationally intensive and more efficient procedure for phase determination in X-ray crystallography.

Crystallography, X-Ray↗

Comparison of maximum statistics for hypothesis testing when a nuisance parameter is present only under the alternative.

In many practical problems, a hypothesis testing involves a nuisance parameter which appears only under the alternative hypothesis. Davies (1977, Biometrika 64, 247-254) proposed the maximum of the score statistics over the whole range of the nuisance parameter as a test statistic for this type of hypothesis testing. Freidlin, Podgor, and Gastwirth (1999, Biometrics 55, 883-886) studied two other simpler maximum test statistics, the maximum of the score statistics at two extreme points of the nuisance parameter, and the maximum of the score statistics at three points of the nuisance parameter including the two extreme points. In this article, we compare the powers of these three maximum-type statistics in the context of three genetic problems.

Biometry↗

Patients and medical statistics. Interest, confidence, and ability.

BACKGROUND: People are increasingly presented with medical statistics. There are no existing measures to assess their level of interest or confidence in using medical statistics. OBJECTIVE: To develop 2 new measures, the STAT-interest and STAT-confidence scales, and assess their reliability and validity. DESIGN: Survey with retest after approximately 2 weeks. SUBJECTS: Two hundred and twenty-four people were recruited from advertisements in local newspapers, an outpatient clinic waiting area, and a hospital open house. MEASURES: We developed and revised 5 items on interest in medical statistics and 3 on confidence understanding statistics. RESULTS: Study participants were mostly college graduates (52%); 25% had a high school education or less. The mean age was 53 (range 20 to 84) years. Most paid attention to medical statistics (6% paid no attention). The mean (SD) STAT-interest score was 68 (17) and ranged from 15 to 100. Confidence in using statistics was also high: the mean (SD) STAT-confidence score was 65 (19) and ranged from 11 to 100. STAT-interest and STAT-confidence scores were moderately correlated (r=.36, P<.001). Both scales demonstrated good test-retest repeatability (r=.60, .62, respectively), internal consistency reliability (Cronbach's alpha=0.70 and 0.78), and usability (individual item nonresponse ranged from 0% to 1.3%). Scale scores correlated only weakly with scores on a medical data interpretation test (r=.15 and .26, respectively). CONCLUSION: The STAT-interest and STAT-confidence scales are usable and reliable. Interest and confidence were only weakly related to the ability to actually use data.

Adult↗