Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “STATISTICS”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Weighing the results of differing 'low dose' studies of the mouse prostate by Nagel, Cagen, and Ashby: quantification of experimental power and statistical results.

Differing experimental findings with respect to "low dose" responses in the mouse prostate after in utero exposure have generated considerable controversy. An analysis of such controversies requires a broad strength and weight of the evidence approach. For example, a National Toxicology Program review panel acquired the raw data from nearly 50 studies and then statistically reanalyzed these data in a common and comparable approach. However, the statistical power of the various studies was not calculated and the quantitative p values were not reported in this reanalysis. Such calculations and values address vital strength- and weight-of-the-evidence questions: (1) how sensitive were the various studies to detect changes in prostate weight, particularly the negative replicate studies and (2) what were the p values; were negative studies robust or only marginal in their inability to find an effect? We first examined the statistical power of the studies to detect a positive effect on prostate weight. Preliminary calculations indicated that the two subsequent replicating studies were indeed more sensitive to changes in prostate weight in comparison to the original study, having reasonable power to detect an effect at only 50% of the response reported in the original study. Additional calculations were performed using the raw data available from one negative replicating study and the methods recommended by the statistics subpanel of the original review. This analysis used Dunnett's multiple comparison procedure for groups with p<0.05 to infer statistical significance, employed an analysis-of-covariance model with body weight as a covariate, and addressed litter as a nested random effect. The quantitated p values for this replicated study, comparing the two Bisphenol A treatment groups (2 and 20 microg/kg/day) to the control, were 0.821 and 0.972, respectively. This indicates this study was indeed robust in finding no treatment-related effect. Thus, the weight and strength of the evidence, based on sensitivity and quantitative p value, was that it is highly unlikely for this negative replicating study to have missed a true effect. In the future, we recommend a similar use of statistical power analysis for those designing experimental studies and for those conducting weight-of-the-evidence reviews, and we also recommend the clear quantitation and reporting of p values to support the review's interpretation and conclusions.

Algorithms↗

Transition event statistics in genetics and disordered kinetics. Theoretical approaches for extracting rate distributions from experimental data.

We study the analogies between the theory of rate processes in disordered systems and the overdispersed molecular clocks in evolutionary biology. A biological "molecular clock" expresses the statistics of the number of amino acid or nucleotide substitutions during evolution. Random variations of the evolution rates lead to statistical (overdispersed) molecular clocks which are described by random point processes with random substitution rates. We find that the models for overdispersed molecular clocks are equivalent to those of the random-rate or random channel models used in disordered kinetics. The number of transport (reaction) events in disordered kinetics plays the same role as the number of substitution events in molecular biology. We study the connections between the (observed) statistics of the transition events and the statistics of random rate coefficients and random channels; a unified approach is developed which is valid both in molecular biology and in disordered kinetics. We develop methods for extracting statistical information about the variations of rate coefficients from experimental or observed data regarding the fluctuations of the numbers of substitution, reaction, or transport events. For systems with static disorder, the observed statistics of the number of reaction events, expressed in terms of probabilities at a given time or by the cumulants of the number of transition events at a given time, contains the information necessary for evaluating the cumulants or the probability density of the rate coefficients or the density of states for random channel kinetics. For dynamic disorder this is not possible; further information about multitime probability distributions of the reaction events is needed.

Amino Acid Substitution↗

Diagnosing item score patterns on a test using item response theory-based person-fit statistics.

Person-fit statistics have been proposed to investigate the fit of an item score pattern to an item response theory (IRT) model. The author investigated how these statistics can be used to detect different types of misfit. Intelligence test data were analyzed using person-fit statistics in the context of the G. Rasch (1960) model and R. J. Mokken's (1971, 1997) IRT models. The effect of the choice of an IRT model to detect misfitting item score patterns and the usefulness of person-fit statisticsfor diagnosis of misfit are discussed. Results showed that different types of person-fit statistics can be used to detect different kinds of person misfit. Parametric person-fit statistics had more power than nonparametric person-fit statistics.

Adult↗

Statistics in physiology and pharmacology: a slow and erratic learning curve.

1. Learning how to apply statistical analyses to the results of experimental or clinical studies may take a lifetime of trial (and sometimes error), as it has done in the author's case. There is no evidence that biomedical investigators of the present generation are on a steeper learning curve. Gross misunderstandings of the purpose and functions of statistical analysis are apparent in applications to research grant-giving bodies and ethics committees, in manuscripts submitted to journals and sometimes in published papers. 2. Although estimation of minimal group (sample) size for a given power is an essential step in planning clinical studies, it seems to be used rarely in laboratory experimental work. This is despite exhortations to restrict the number of animals used to a minimum. 3. Most investigators use hypothesis testing to analyse their results, but their understanding of the meaning of the resultant P-values is slight. 4. A flaw found almost universally in biomedical manuscripts is to make multiple inferences from the results of a single study. The goal of statistical analysis is to maintain the familywise type I error rate (risk of false-positive inference) at a predetermined level (usually 5%). But, when multiple inferences are made from the same experiment, the risk of false-positive error is inflated. There are two solutions to this problem: (i) use a multiple comparison procedure to control the familywise type I error rate; and (ii) test a single, global hypothesis. 5. Biomedical investigators have been quick to acquire computer statistics software and to use it to analyse their experiments. However, they have been slow to recognize the limitations of this software. These include: (i) inadequate documentation of routines, so that neither the user nor the reader of published papers can be sure how the tests have been executed; (ii) flawed algorithms for the execution of statistical procedures; and (iii) failure to recognize that the best software for their purposes is that which takes them just beyond their statistical horizons. 6. The obvious solution to these difficulties is to recruit a biomedical statistician into every research group, at a relatively trivial cost. However, properly qualified biostatisticians are in desperately short supply in Australia. It follows that research groups, national grant-giving agencies and academic institutions must make provision for the proper training and subsequent employment of biostatisticians.

Data Interpretation, Statistical↗

A geometric approach to tree shape statistics.

This article presents a new way to quantify the descriptive ability of tree shape statistics. Where before, tree shape statistics were chosen by their ability to distinguish between macroevolutionary models, the resolution presented in this paper quantifies the ability of a statistic to differentiate between similar and different trees. This is termed the geometric approach to differentiate it from the model-based approach previously explored. A distinct advantage of this perspective is that it allows evaluation of multiple tree shape statistics describing different aspects of tree shape. After developing the methodology, it is applied here to make specific recommendations for a suite of three statistics that may prove useful in applications. The article ends with an application of the statistics to clarify the impact of taxa omission on tree shape.

Classification↗

Statistical methods of translating microarray data into clinically relevant diagnostic information in colorectal cancer.

MOTIVATION: It is a common practice in cancer microarray experiments that a normal tissue is collected from the same individual from whom the tumor tissue was taken. The indirect design is usually adopted for the experiment that uses a common reference RNA hybridized both to normal and tumor tissues. However, it is often the case that the test material is not large enough for the experimenter to extract enough RNA to conduct the microarray experiment. Hence, collecting n cases does not necessarily end up with a matched pair sample of size n. Instead we usually have a matched pair sample of size n1, and two independent samples of sizes n2 and n3, respectively, for 'reference versus normal tissue only' and 'reference versus tumor tissue only' hybridizations (n=n1 + n2 + n3). Standard statistical methods need to be modified and new statistical procedures are developed for analyzing this mixed dataset. RESULTS: We propose a new test statistic, t3, as a means of combining all the information in the mixed dataset for detecting differentially expressed (DE) genes between normal and tumor tissues. We employed the extended receiver operating characteristic approach to the mixed dataset. We devised a measure of disagreement between a RT-PCR experiment and a microarray experiment. Hotelling's T2 statistic is employed to detect a set of DE genes and its prediction rate is compared with the prediction rate of a univariate procedure. We observe that Hotelling's T2 statistic detects DE genes more efficiently than a univariate procedure and that further research is warranted on the formal test procedure using Hotelling's T2 statistic. CONTACT: bskim@yonsei.ac.kr.

Adult↗

Identifying differentially expressed genes from microarray experiments via statistic synthesis.

MOTIVATION: A common objective of microarray experiments is the detection of differential gene expression between samples obtained under different conditions. The task of identifying differentially expressed genes consists of two aspects: ranking and selection. Numerous statistics have been proposed to rank genes in order of evidence for differential expression. However, no one statistic is universally optimal and there is seldom any basis or guidance that can direct toward a particular statistic of choice. RESULTS: Our new approach, which addresses both ranking and selection of differentially expressed genes, integrates differing statistics via a distance synthesis scheme. Using a set of (Affymetrix) spike-in datasets, in which differentially expressed genes are known, we demonstrate that our method compares favorably with the best individual statistics, while achieving robustness properties lacked by the individual statistics. We further evaluate performance on one other microarray study.

Algorithms↗

The contribution of statistical parametric mapping in the assessment of precuneal and medial temporal lobe perfusion by 99mTc-HMPAO SPECT in mild Alzheimer's and Lewy body dementia.

AIM: To assess the role of 99mTc-hexamethylpropyleneamine oxime single-photon emission computed tomography (99mTc-HMPAO SPECT) imaging of the precuneus and medial temporal lobe in the individual patient with mild Alzheimer's disease and dementia with Lewy bodies (DLB) using statistical parametric mapping and visual image interpretation. METHODS: Thirty-four patients with mild late-onset Alzheimer's disease, 20 patients with early-onset Alzheimer's disease, 15 patients with DLB and 31 healthy controls were studied. All patients fulfilled appropriate clinical criteria; the DLB patients also had evidence of dopaminergic presynaptic terminal loss on 123I-N-omega-fluoropropyl-2beta-carbomethoxy-3beta-(4-iodophenyl)-tropane imaging. 99mTc-HMPAO SPECT brain scans were acquired on a multidetector gamma camera and images were assessed separately by visual interpretation and with SPM99. RESULTS: Statistical parametric maps were significantly more accurate than visual image interpretation in all disease categories. In patients with mild late-onset Alzheimer's disease, statistical parametric mapping demonstrated significant hypoperfusion to the precuneus in 59% and to the medial temporal lobe in 53%. Seventy-six per cent of these patients had a defect in either location. No controls had precuneal or medial temporal lobe hypoperfusion (specificity, 100%). Statistical parametric mapping also demonstrated 73% of patients with DLB to have precuneal abnormalities, but only 6% had medial temporal lobe involvement. CONCLUSION: These findings illustrate the capability of statistical parametric mapping to demonstrate reliable abnormalities in the majority, but not all, patients with either mild Alzheimer's disease or DLB. Precuneal hypoperfusion is not specific to Alzheimer's disease and is equally likely to be found in DLB. In this study, medial temporal hypoperfusion was significantly more common in Alzheimer's disease than in DLB. Statistical parametric maps appear to be considerably more reliable than simple visual interpretation of 99mTc-HMPAO images for these regions.

Adult↗

Statistical power of MRI monitored trials in multiple sclerosis: new data and comparison with previous results.

OBJECTIVES: To evaluate the durations of the follow up and the reference population sizes needed to achieve optimal and stable statistical powers for two period cross over and parallel group design clinical trials in multiple sclerosis, when using the numbers of new enhancing lesions and the numbers of active scans as end point variables. METHODS: The statistical power was calculated by means of computer simulations performed using MRI data obtained from 65 untreated relapsing-remitting or secondary progressive patients who were scanned monthly for 9 months. The statistical power was calculated for follow up durations of 2, 3, 6, and 9 months and for sample sizes of 40-100 patients for parallel group and of 20-80 patients for two period cross over design studies. The stability of the estimated powers was evaluated by applying the same procedure on random subsets of the original data. RESULTS: When using the number of new enhancing lesions as the end point, the statistical power increased for all the simulated treatment effects with the duration of the follow up until 3 months for the parallel group design and until 6 months for the two period cross over design. Using the number of active scans as the end point, the statistical power steadily increased until 6 months for the parallel group design and until 9 months for the two period cross over design. The power estimates in the present sample and the comparisons of these results with those obtained by previous studies with smaller patient cohorts suggest that statistical power is significantly overestimated when the size of the reference data set decreases for parallel group design studies or the duration of the follow up decreases for two period cross over studies. CONCLUSIONS: These results should be used to determine the duration of the follow up and the sample size needed when planning MRI monitored clinical trials in multiple sclerosis.

Adolescent↗

Parametric vs. non-parametric statistics of low resolution electromagnetic tomography (LORETA).

This study compared the relative statistical sensitivity of non-parametric and parametric statistics of 3-dimensional current sources as estimated by the EEG inverse solution Low Resolution Electromagnetic Tomography (LORETA). One would expect approximately 5% false positives (classification of a normal as abnormal) at the P < .025 level of probability (two tailed test) and approximately 1% false positives at the P < .005 level. EEG digital samples (2 second intervals sampled 128 Hz, 1 to 2 minutes eyes closed) from 43 normal adult subjects were imported into the Key Institute's LORETA program. We then used the Key Institute's cross-spectrum and the Key Institute's LORETA output files (*.lor) as the 2,394 gray matter pixel representation of 3-dimensional currents at different frequencies. The mean and standard deviation *.lor files were computed for each of the 2,394 gray matter pixels for each of the 43 subjects. Tests of Gaussianity and different transforms were computed in order to best approximate a normal distribution for each frequency and gray matter pixel. The relative sensitivity of parametric vs. non-parametric statistics were compared using a "leave-one-out" cross validation method in which individual normal subjects were withdrawn and then statistically classified as being either normal or abnormal based on the remaining subjects. Log10 transforms approximated Gaussian distribution in the range of 95% to 99% accuracy. Parametric Z score tests at P < .05 cross-validation demonstrated an average misclassification rate of approximately 4.25%, and range over the 2,394 gray matter pixels was 27.66% to 0.11%. At P < .01 parametric Z score cross-validation false positives were 0.26% and ranged from 6.65% to 0% false positives. The non-parametric Key Institute's t-max statistic at P < .05 had an average misclassification error rate of 7.64% and ranged from 43.37% to 0.04% false positives. The nonparametric t-max at P < .01 had an average misclassification rate of 6.67% and ranged from 41.34% to 0% false positives of the 2,394 gray matter pixels for any cross-validated normal subject. In conclusion, adequate approximation to Gaussian distribution and high cross-validation can be achieved by the Key Institute's LORETA programs by using a log10 transform and parametric statistics, and parametric normative comparisons had lower false positive rates than the non-parametric tests.

Adolescent↗

The statistics of identifying differentially expressed genes in Expresso and TM4: a comparison.

BACKGROUND: Analysis of DNA microarray data takes as input spot intensity measurements from scanner software and returns differential expression of genes between two conditions, together with a statistical significance assessment. This process typically consists of two steps: data normalization and identification of differentially expressed genes through statistical analysis. The Expresso microarray experiment management system implements these steps with a two-stage, log-linear ANOVA mixed model technique, tailored to individual experimental designs. The complement of tools in TM4, on the other hand, is based on a number of preset design choices that limit its flexibility. In the TM4 microarray analysis suite, normalization, filter, and analysis methods form an analysis pipeline. TM4 computes integrated intensity values (IIV) from the average intensities and spot pixel counts returned by the scanner software as input to its normalization steps. By contrast, Expresso can use either IIV data or median intensity values (MIV). Here, we compare Expresso and TM4 analysis of two experiments and assess the results against qRT-PCR data. RESULTS: The Expresso analysis using MIV data consistently identifies more genes as differentially expressed, when compared to Expresso analysis with IIV data. The typical TM4 normalization and filtering pipeline corrects systematic intensity-specific bias on a per microarray basis. Subsequent statistical analysis with Expresso or a TM4 t-test can effectively identify differentially expressed genes. The best agreement with qRT-PCR data is obtained through the use of Expresso analysis and MIV data. CONCLUSION: The results of this research are of practical value to biologists who analyze microarray data sets. The TM4 normalization and filtering pipeline corrects microarray-specific systematic bias and complements the normalization stage in Expresso analysis. The results of Expresso using MIV data have the best agreement with qRT-PCR results. In one experiment, MIV is a better choice than IIV as input to data normalization and statistical analysis methods, as it yields as greater number of statistically significant differentially expressed genes; TM4 does not support the choice of MIV input data. Overall, the more flexible and extensive statistical models of Expresso achieve more accurate analytical results, when judged by the yardstick of qRT-PCR data, in the context of an experimental design of modest complexity.

Algorithms↗

Statistics in medical research.

The role of statistics in medical research starts at the planning stage of a clinical trial or laboratory experiment to establish the design and size of an experiment that will ensure a good prospect of detecting effects of clinical or scientific interest. Statistics is again used during the analysis of data (sample data) to make inferences valid in a wider population. In simple situations computation of simple quantities such as P-values, confidence intervals, standard deviations, standard errors or application of some standard parametric or nonparametric tests may suffice. Despite their wide use even these simple notions are sometimes misunderstood or misinterpreted by research workers in other disciplines who have only a limited knowledge of statistics. More sophisticated research projects often need advanced statistical methods including the formulation and testing of mathematical models to make relevant inferences from observed data. Such advanced methods should only be applied with a clear understanding both of their purposes and the implication of any conclusions based upon their use. Close collaboration between statisticians, whether professionals in that field or medical research workers with a sound statistical background, and other members of a research team is needed to ensure a seamless integration of the statistical elements into the reporting and discussion of research outcomes. Some suggestions are made as to how that collaboration is best achieved.

Biomedical Research↗

Statistical inference by confidence intervals: issues of interpretation and utilization.

This article examines the role of the confidence interval (CI) in statistical inference and its advantages over conventional hypothesis testing, particularly when data are applied in the context of clinical practice. A CI provides a range of population values with which a sample statistic is consistent at a given level of confidence (usually 95%). Conventional hypothesis testing serves to either reject or retain a null hypothesis. A CI, while also functioning as a hypothesis test, provides additional information on the variability of an observed sample statistic (ie, its precision) and on its probable relationship to the value of this statistic in the population from which the sample was drawn (ie, its accuracy). Thus, the CI focuses attention on the magnitude and the probability of a treatment or other effect. It thereby assists in determining the clinical usefulness and importance of, as well as the statistical significance of, findings. The CI is appropriate for both parametric and nonparametric analyses and for both individual studies and aggregated data in meta-analyses. It is recommended that, when inferential statistical analysis is performed, CIs should accompany point estimates and conventional hypothesis tests wherever possible.

Bias↗

The use of sampling for vital registration and vital statistics.

In this paper, the author does not so much try to give a blueprint for the application of sampling methods to vital registration and vital statistics as to show the opportunities for their use and the advantages to be derived from them. In the less developed areas of the world, modern sampling methods make it possible to obtain very accurate national statistics in the early stages of the establishment of a vital registration and vital statistics system and will lead to its more orderly and efficient development. In areas where more or less complete registration exists, the use of sampling may result in a reduction of costs and an improvement in the quality and currency of the data obtained.The sample vital statistics system proposed by the author should comprise complete primary registration units or combinations of them, representative of the entire universe for which statistics are wanted. The selection of these, however, must be made at random; but, in order to avoid bias, the units should be taken with probabilities proportionate to their size.After discussing the ways of carrying out his proposal and the relation of a sample vital statistics system to health programmes, the author considers the use of the sample system as a supplement to a complete system and the advantages of sampling for quality control, checking the completeness of registration, preparing advance tabulations, and conducting supplemental surveys and research.

Data Collection↗

Securing wide appreciation of health statistics.

All the authors are agreed on the need for a certain publicizing of health statistics, but do Amaral Pyrrait points out that the medical profession prefers to convince itself rather than to be convinced. While there is great utility in articles and reviews in the professional press (especially for paramedical personnel) Aubenque, de Groot, and Kohn show how appreciation can effectively be secured by making statistics more easily understandable to the non-expert by, for instance, including readable commentaries in official publications, simplifying charts and tables, and preparing simple manuals on statistical methods. Aubenque and Kohn also stress the importance of linking health statistics to other economic and social information. Benjamin suggests that the principles of market research could to advantage be applied to health statistics to determine the precise needs of the "consumers". At the same time, Aubenque points out that the value of the ultimate results must be clear to those who provide the data; for this, Kohn suggests that the enumerators must know exactly what is wanted and why.There is general agreement that some explanation of statistical methods and their uses should be given in the curricula of medical schools and that lectures and postgraduate courses should be arranged for practising physicians.

Humans↗

MEDICAL LIBRARY STATISTICS.

Four compilations of medical library statistics have been published to date, namely those by Louise Darling in 1956, by the Medical Library Association in its 1959 Directory, by Harold Bloomquist in 1962, and by the author in 1964. In addition to these sources and to the annual statistics compiled by the Library Services Branch of the U. S. Office of Education, surveys of pharmacy, hospital, and medical society libraries have been completed recently. Standards for medical and special libraries are being considered by the Medical Library Association through its Guidelines Survey and by the Special Libraries Association through its Statistics Coordinating Project and its Standards Survey. To coordinate the collection of medical library statistics and to make the information readily available to the profession, it is suggested that the Medical Library Association support the collection and publication of statistics of representative medical libraries until such time as the Library Services Branch is able to implement fully its program of library statistics.

Data Collection↗

Assessment of statistical methods used in library-based approaches to microbial source tracking.

Several commonly used statistical methods for fingerprint identification in microbial source tracking (MST) were examined to assess the effectiveness of pattern-matching algorithms to correctly identify sources. Although numerous statistical methods have been employed for source identification, no widespread consensus exists as to which is most appropriate. A large-scale comparison of several MST methods, using identical fecal sources, presented a unique opportunity to assess the utility of several popular statistical methods. These included discriminant analysis, nearest neighbour analysis, maximum similarity and average similarity, along with several measures of distance or similarity. Threshold criteria for excluding uncertain or poorly matched isolates from final analysis were also examined for their ability to reduce false positives and increase prediction success. Six independent libraries used in the study were constructed from indicator bacteria isolated from fecal materials of humans, seagulls, cows and dogs. Three of these libraries were constructed using the rep-PCR technique and three relied on antibiotic resistance analysis (ARA). Five of the libraries were constructed using Escherichia coli and one using Enterococcus spp. (ARA). Overall, the outcome of this study suggests a high degree of variability across statistical methods. Despite large differences in correct classification rates among the statistical methods, no single statistical approach emerged as superior. Thresholds failed to consistently increase rates of correct classification and improvement was often associated with substantial effective sample size reduction. Recommendations are provided to aid in selecting appropriate analyses for these types of data.

Animals↗

Application of statistical methods in the papers published in the East African Medical Journal (EAMJ) since 1923.

The subject Statistics was hardly known in the research world at the time the EAMJ was being started in 1923. It was at this time, scientists working in their field of specialization were busy developing statistical methods with the aim of solving problems affecting them in their research work. Using a random sample of some of the papers published by the EAMJ, this paper evaluates the usage of statistical methods since 1923. Usage was very low before 1965, but started picking-up with a rising trend since then but still not impressive. Scientists should strive to apply statistical methods correctly in their research work. EAMJ should lay more emphasis on statistical refereeing as a policy to raise the quality of published papers in the Journal. If statisticians were not available, which often may be the case, scientists should be encouraged to show their work to other research colleagues working in similar areas, before sending the paper to the EAMJ. By adopting this approach, it is possible that the quality of papers will go up and the usage of statistical methods will increase.

Bias↗