Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “STATISTICS”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Optimal allele-sharing statistics for genetic mapping using affected relatives.

The choice of allele-sharing statistics can have a great impact on the power of robust affected relative methods. Similarly, when allele-sharing statistics from several pedigrees are combined, the weight applied to each pedigree's statistic can affect power. Here we describe the direct connection between the affected relative methods and traditional parametric linkage analysis, and we use this connection to give explicit formulae for the optimal sharing statistics and weights, applicable to all pedigree types. One surprising consequence is that under any single gene model, the value of the optimal allele-sharing statistic does not depend on whether observed sharing is between more closely or more distantly related affected relatives. This result also holds for any multigene model with loci unlinked, additivity between loci, and all loci having small effect. For specific classes of two-allele models, we give the most powerful statistics and optimal weights for arbitrary pedigrees. When the effect size is small, these also extend to multigene models with additivity between loci. We propose a useful new statistic, S(rob dom), which performs well for dominant and additive models with varying phenocopy rates and varying predisposing allele frequency. We find that the statistic S(_#alleles), performs well for recessive models with varying phenocopy rates and varying redisposing allele frequency. We also find that for models with large deviation from null sharing, the correspondence between allele-sharing statistics and the models for which they are optimal may also depend on which method is used to test for linkage.

Alleles↗

The organization of statistical activities at medical research institutions.

In this paper, we discuss the organizational structure for the conduct of statistical activities at medical Colleges and Universities. Here, we have particularly focused on two significant statistical activities, i.e., statistical consultation and research. The consulting service for medical researchers consists of practical statistical analysis, instruction on computer manipulation and software, response to reviewer's comments and statistical design of prospective studies. From our various experiences, we describe the actual implementation of statistical consultations for medical researchers. It has played an important role in supporting medical research. In addition, we also outline some research, with respect to new ways of applying statistics in medical science. This paper concludes that it may be practical for the existing computer center or department of medical informatics, in charge of the computing service, to conduct statistical activities until formal organizations are established at the academic institution. To realize the conduct of statistical activities by Department of Medical Informatics, it needs a team of biostatisticians, data analysts and computer personnel.

Computer Systems↗

Statistics for nonparametric linkage analysis of X-linked traits in general pedigrees.

We have compared the power of several allele-sharing statistics for "nonparametric" linkage analysis of X-linked traits in nuclear families and extended pedigrees. Our rationale was that, although several of these statistics have been implemented in popular software packages, there has been no formal evaluation of their relative power. Here, we evaluate the relative performance of five test statistics, including two new test statistics. We considered sibships of sizes two through four, four different extended pedigrees, 15 different genetic models (12 single-locus models and 3 two-locus models), and varying recombination fractions between the marker and the trait locus. We analytically estimated the sample sizes required for 80% power at a significance level of.001 and also used simulation methods to estimate power for a sample size of 10 families. We tried to identify statistics whose power was robust over a wide variety of models, with the idea that such statistics would be particularly useful for detection of X-linked loci associated with complex traits. We found that a commonly used statistic, S(all), generally performed well under various conditions and had close to the optimal sample sizes in most cases but that there were certain cases in which it performed quite poorly. Our two new statistics did not perform any better than those already in the literature. We also note that, under dominant and additive models, regardless of the statistic used, pedigrees with all-female siblings have very little power to detect X-linked loci.

Alleles↗

Application of the bootstrap procedure provides an alternative to standard statistical procedures in the estimation of the vitamin B-6 requirement.

The bootstrap procedure is a versatile statistical tool for the estimation of standard errors and confidence intervals. It is useful when standard statistical methods are not available or are poorly behaved, e.g., for nonlinear functions or when assumptions of a statistical model have been violated. Inverse regression estimation is an example of a statistical tool with a wide application in human nutrition. In a recent study, inverse regression was used to estimate the vitamin B-6 requirement of young women. In the present statistical application, both standard statistical methods and the bootstrap technique were used to estimate the mean vitamin B-6 requirement, standard errors and 95% confidence intervals for the mean. The bootstrap procedure produced standard error estimates and confidence intervals that were similar to those calculated by using standard statistical estimators. In a Monte Carlo simulation exploring the behavior of the inverse regression estimators, bootstrap standard errors were found to be nearly unbiased, even when the basic assumptions of the regression model were violated. On the other hand, the standard asymptotic estimator was found to behave well when the assumptions of the regression model were met, but behaved poorly when the assumptions were violated. In human metabolic studies, which are often restricted to small sample sizes, or when statistical methods are not available or are poorly behaved, bootstrap estimates for calculating standard errors and confidence intervals may be preferred. Investigators in human nutrition may find that the bootstrap procedure is superior to standard statistical procedures in cases similar to the examples presented in this paper.

Adult↗

Scaled test statistics and robust standard errors for non-normal data in covariance structure analysis: a Monte Carlo study.

Research studying robustness of maximum likelihood (ML) statistics in covariance structure analysis has concluded that test statistics and standard errors are biased under severe non-normality. An estimation procedure known as asymptotic distribution free (ADF), making no distributional assumption, has been suggested to avoid these biases. Corrections to the normal theory statistics to yield more adequate performance have also been proposed. This study compares the performance of a scaled test statistic and robust standard errors for two models under several non-normal conditions and also compares these with the results from ML and ADF methods. Both ML and ADF test statistics performed rather well in one model and considerably worse in the other. In general, the scaled test statistic seemed to behave better than the ML test statistic and the ADF statistic performed the worst. The robust and ADF standard errors yielded more appropriate estimates of sampling variability than the ML standard errors, which were usually downward biased, in both models under most of the non-normal conditions. ML test statistics and standard errors were found to be quite robust to the violation of the normality assumption when data had either symmetric and platykurtic distributions, or non-symmetric and zero kurtotic distributions.

Analysis of Variance↗

Incongruence between test statistics and P values in medical papers.

BACKGROUND: Given an observed test statistic and its degrees of freedom, one may compute the observed P value with most statistical packages. It is unknown to what extent test statistics and P values are congruent in published medical papers. METHODS: We checked the congruence of statistical results reported in all the papers of volumes 409-412 of Nature (2001) and a random sample of 63 results from volumes 322-323 of BMJ (2001). We also tested whether the frequencies of the last digit of a sample of 610 test statistics deviated from a uniform distribution (i.e., equally probable digits). RESULTS: 11.6% (21 of 181) and 11.1% (7 of 63) of the statistical results published in Nature and BMJ respectively during 2001 were incongruent, probably mostly due to rounding, transcription, or type-setting errors. At least one such error appeared in 38% and 25% of the papers of Nature and BMJ, respectively. In 12% of the cases, the significance level might change one or more orders of magnitude. The frequencies of the last digit of statistics deviated from the uniform distribution and suggested digit preference in rounding and reporting. CONCLUSIONS: This incongruence of test statistics and P values is another example that statistical practice is generally poor, even in the most renowned scientific journals, and that quality of papers should be more controlled and valued.

Confidence Intervals↗

Multivariable risk prediction can greatly enhance the statistical power of clinical trial subgroup analysis.

BACKGROUND: When subgroup analyses of a positive clinical trial are unrevealing, such findings are commonly used to argue that the treatment's benefits apply to the entire study population; however, such analyses are often limited by poor statistical power. Multivariable risk-stratified analysis has been proposed as an important advance in investigating heterogeneity in treatment benefits, yet no one has conducted a systematic statistical examination of circumstances influencing the relative merits of this approach vs. conventional subgroup analysis. METHODS: Using simulated clinical trials in which the probability of outcomes in individual patients was stochastically determined by the presence of risk factors and the effects of treatment, we examined the relative merits of a conventional vs. a "risk-stratified" subgroup analysis under a variety of circumstances in which there is a small amount of uniformly distributed treatment-related harm. The statistical power to detect treatment-effect heterogeneity was calculated for risk-stratified and conventional subgroup analysis while varying: 1) the number, prevalence and odds ratios of individual risk factors for risk in the absence of treatment, 2) the predictiveness of the multivariable risk model (including the accuracy of its weights), 3) the degree of treatment-related harm, and 5) the average untreated risk of the study population. RESULTS: Conventional subgroup analysis (in which single patient attributes are evaluated "one-at-a-time") had at best moderate statistical power (30% to 45%) to detect variation in a treatment's net relative risk reduction resulting from treatment-related harm, even under optimal circumstances (overall statistical power of the study was good and treatment-effect heterogeneity was evaluated across a major risk factor [OR = 3]). In some instances a multi-variable risk-stratified approach also had low to moderate statistical power (especially when the multivariable risk prediction tool had low discrimination). However, a multivariable risk-stratified approach can have excellent statistical power to detect heterogeneity in net treatment benefit under a wide variety of circumstances, instances under which conventional subgroup analysis has poor statistical power. CONCLUSION: These results suggest that under many likely scenarios, a multivariable risk-stratified approach will have substantially greater statistical power than conventional subgroup analysis for detecting heterogeneity in treatment benefits and safety related to previously unidentified treatment-related harm. Subgroup analyses must always be well-justified and interpreted with care, and conventional subgroup analyses can be useful under some circumstances; however, clinical trial reporting should include a multivariable risk-stratified analysis when an adequate externally-developed risk prediction tool is available.

Clinical Trials as Topic↗

[Use and presentation of statistics methods in original articles published in MEDICINA CLINICA in 1993].

BACKGROUND: In the last few years a marked increase has been observed in the use of statistical techniques in biomedical publications. Some characteristics of the statistical quality of the first 100 articles consecutively published in 1993 in the section of originals and surveys of the journal MEDICINA CLINICA are presented in this study. METHODS: In each original one reviewer identified errors and/or criticisms in statistical methodology. An adaptation of the protocol of statistical revision made by the journal The Lancet since november 1990 was used in addition to its own classification. Likewise, an error and/or criticism involving minimum statistical quality necessary was considered as major and that which only had a decrease in optimum statistical level was considered as minor. Statistical analysis consisted in descriptive tables of the number of originals in which each of the classified errors and/or criticisms were observed. RESULTS: Sixty-seven percent of the originals (CI 95%: 64;70) presented major errors and/or criticisms in design, analysis and inference. The most frequent were found in power and sample size (design), need for better analysis (analysis) and lack of confidence intervals (inference). Only 15% (CI 95%: 13;17) had major errors in statistical presentation and 82% (CI 95%: 80;84) minor errors among which the so-called orphaned p was of note. CONCLUSIONS: The need for accompanying the p values by the confidence intervals or taking the calculation of the required sample size into account are of note. Furthermore, the diffusion of explicit recommendations concerning the carrying out and presentation of statistical analysis is necessary.

Periodicals as Topic↗

[Evaluation of the use of statistical techniques in original articles published in the Medicina Clínica during 3 decades (1962-1992)].

BACKGROUND: The incorrect use of statistical techniques in medical articles may seriously compromise the validity of conclusions. This finding otherwise is relatively common. METHODS: A total of 84 original articles published in Medicina Clínica between 1962 and 1992 were reviewed with the aim of assessing the use and appropriateness of statistical techniques. The use of statistics, the quality of the analyses performed, and the inaccuracy of the statistical techniques used were evaluated. We also classified the statistical techniques most commonly used throughout the study period. RESULTS: There was a marked increase in the use of statistical analyses, from 8.3% in 1962 to 83.3% in 1992. It should be noted that a substantial part of this increase has been due to the use of inferential tests, which accounted up to 70% in the sample of articles published in 1992. This finding, however, was associated with an increase in the number of incorrect analyses. The most common statistical errors included assumption of normal distribution of data (with no mention of the test used to confirm this fact), mistake between standard deviation and standard error of the mean, inadequate inferences on the basis of the sample size, inappropriate use of the Student's t test, chi-square test, nonparametric tests or multivariate analyses as well as misunderstanding of linear regression and correlation. CONCLUSIONS: High standards in scientific research have been accompanied by a significant increase in the number of clinical studies with statistical analysis of data. However, this apparently favorable situation has been associated with an increase in the number of inaccurate analyses. It has been found that sophisticated statistical tests are rarely used in articles published in Medicina Clínica.

Chi-Square Distribution↗

The development of national vital statistics in Canada: Part 1--From 1605 to 1945.

This article describes the key events in the development of the national vital statistics system in Canada. Particular emphasis is placed on the role played by Statistics Canada, known as the Dominion Bureau of Statistics from 1918 to 1971. There were many obstacles to uniform national compilations, including differences in provincial legislation, the incomplete registration of vital events, a lack of uniform standards in classification and methods of presentation, the omission of important data, the use of fiscal instead of calendar years, and periodic breaks in the annual publications prepared by the provinces and territories. To overcome these obstacles, collaboration among federal and provincial/territorial governments was necessary. This two-part article chronicles the evolution of this collaboration, which led to the production of national vital statistics in Canada. Part 1 covers the years 1605 to 1945, from the time explorers, the Catholic Church and census takers first recorded details about the European population in New France, to the establishment of a system of national vital statistics. It ends by noting the important role national vital statistics played in launching Family Allowances in 1945. Part 2, scheduled to appear in a future issue of Health Reports, will cover the years 1945 to the present. It will focus on the creation of the National Vital Statistics Index, the Vital Statistics Council, computerization, record linkage, and occupational and environmental health statistics.

Canada↗

Advice on statistical analysis for Circulation Research.

Since the late 1970s when many journals published articles warning about the misuse of statistical methods in the analysis of data, researchers have become more careful about statistical analysis, but errors including low statistical power and inadequate analysis of repeated-measurement studies are still prevalent. In this review, several statistical methods are introduced that are not always familiar to basic and clinical cardiologists but may be useful for revealing the correct answer from the data. The aim of this review is not only to draw the attention of investigators to these tests but also to stress the conditions in which they are applicable. These methods are now generally available in statistical program packages. Researchers need not know how to calculate the statistics from the data but are required to select the correct method from the menu and interpret the statistical results accurately. With the choice of appropriate statistical programs, the issue is no longer how to do the test but when to do it.

Analysis of Variance↗

On the statistical assessment of classifiers using DNA microarray data.

BACKGROUND: In this paper we present a method for the statistical assessment of cancer predictors which make use of gene expression profiles. The methodology is applied to a new data set of microarray gene expression data collected in Casa Sollievo della Sofferenza Hospital, Foggia--Italy. The data set is made up of normal (22) and tumor (25) specimens extracted from 25 patients affected by colon cancer. We propose to give answers to some questions which are relevant for the automatic diagnosis of cancer such as: Is the size of the available data set sufficient to build accurate classifiers? What is the statistical significance of the associated error rates? In what ways can accuracy be considered dependant on the adopted classification scheme? How many genes are correlated with the pathology and how many are sufficient for an accurate colon cancer classification? The method we propose answers these questions whilst avoiding the potential pitfalls hidden in the analysis and interpretation of microarray data. RESULTS: We estimate the generalization error, evaluated through the Leave-K-Out Cross Validation error, for three different classification schemes by varying the number of training examples and the number of the genes used. The statistical significance of the error rate is measured by using a permutation test. We provide a statistical analysis in terms of the frequencies of the genes involved in the classification. Using the whole set of genes, we found that the Weighted Voting Algorithm (WVA) classifier learns the distinction between normal and tumor specimens with 25 training examples, providing e = 21% (p = 0.045) as an error rate. This remains constant even when the number of examples increases. Moreover, Regularized Least Squares (RLS) and Support Vector Machines (SVM) classifiers can learn with only 15 training examples, with an error rate of e = 19% (p = 0.035) and e = 18% (p = 0.037) respectively. Moreover, the error rate decreases as the training set size increases, reaching its best performances with 35 training examples. In this case, RLS and SVM have error rates of e = 14% (p = 0.027) and e = 11% (p = 0.019). Concerning the number of genes, we found about 6000 genes (p < 0.05) correlated with the pathology, resulting from the signal-to-noise statistic. Moreover the performances of RLS and SVM classifiers do not change when 74% of genes is used. They progressively reduce up to e = 16% (p < 0.05) when only 2 genes are employed. The biological relevance of a set of genes determined by our statistical analysis and the major roles they play in colorectal tumorigenesis is discussed. CONCLUSIONS: The method proposed provides statistically significant answers to precise questions relevant for the diagnosis and prognosis of cancer. We found that, with as few as 15 examples, it is possible to train statistically significant classifiers for colon cancer diagnosis. As for the definition of the number of genes sufficient for a reliable classification of colon cancer, our results suggest that it depends on the accuracy required.

Aged↗

Issues in the selection of a summary statistic for meta-analysis of clinical trials with binary outcomes.

Meta-analysis of binary data involves the computation of a weighted average of summary statistics calculated for each trial. The selection of the appropriate summary statistic is a subject of debate due to conflicts in the relative importance of mathematical properties and the ability to intuitively interpret results. This paper explores the process of identifying a summary statistic most likely to be consistent across trials when there is variation in control group event rates. Four summary statistics are considered: odds ratios (OR); risk differences (RD) and risk ratios of beneficial (RR(B)); and harmful outcomes (RR(H)). Each summary statistic corresponds to a different pattern of predicted absolute benefit of treatment with variation in baseline risk, the greatest difference in patterns of prediction being between RR(B) and RR(H). Selection of a summary statistic solely based on identification of the best-fitting model by comparing tests of heterogeneity is problematic, principally due to low numbers of trials. It is proposed that choice of a summary statistic should be guided by both empirical evidence and clinically informed debate as to which model is likely to be closest to the expected pattern of treatment benefit across baseline risks. Empirical investigations comparing the four summary statistics on a sample of 551 systematic reviews provide evidence that the RR and OR models are on average more consistent than RD, there being no difference on average between RR and OR. From a second sample of 114 meta-analyses evidence indicates that for interventions aimed at preventing an undesirable event, greatest absolute benefits are observed in trials with the highest baseline event rates, corresponding to the model of constant RR(H). The appropriate selection for a particular meta-analysis may depend on understanding reasons for variation in control group event rates; in some situations uncertainty about the choice of summary statistic will remain.

Anti-Inflammatory Agents, Non-Steroidal↗

Evaluation of statistical packages for suitability for use by clinical investigators in medicine.

With the increased availability of personal computers and statistical software packages, it is inevitable that there will be increasing attempts by clinical investigators to perform data management and statistical analysis. Reviews of statistical packages are abundant in computer and statistical journals. However the majority of them were not written for clinical investigators in medicine. This paper presents an analytic approach to evaluate the suitability of statistical packages for use by clinical investigators for data-management and preliminary statistical-analysis purposes. The evaluation scheme addresses five areas of concern: availability of data-management features; availability of basic statistical-analysis features; ease of use; documentation; and quality of programs. Among six statistical packages reviewed by this process, CRISP is recommended as the most suitable package for clinical investigators to use for data-management and preliminary statistical-analysis purposes.

Evaluation Studies as Topic↗

Craniofacial reconstruction using a combined statistical model of face shape and soft tissue depths: methodology and validation.

Forensic facial reconstruction aims at estimating the facial outlook associated with an unidentified skull specimen. Estimation is generally based on tabulated average values of soft tissue thicknesses measured at a sparse set of landmarks on the skull. Traditional 'plastic' methods apply modeling clay or plasticine on a cast of the skull, approximating the estimated tissue depths at the landmarks and interpolating in between. Current computerized techniques mimic this landmark interpolation procedure using a single static facial surface template. However, the resulting reconstruction is biased by the specific choice of the template and no face-specific regularization is used during the interpolation process. We reduce the template bias by using a flexible statistical model of a dense set of facial surface points, combined with an associated sparse set of skull-based landmarks. This statistical model is constructed from a facial database of (N = 118) individuals and limits the reconstructions to statistically plausible outlooks. The actual reconstruction is obtained by fitting the skull-based landmarks of the template model to the corresponding landmarks indicated on a digital copy of the skull to be reconstructed. The fitting process changes the face-specific statistical model parameters in a regularized way and interpolates the remaining landmark fit error using a minimal bending thin-plate spline (TPS)-based deformation. Furthermore, estimated properties of the skull specimen (BMI, age and gender, e.g.) can be incorporated as conditions on the reconstruction by removing property-related shape variation from the statistical model description before the fitting process. The proposed statistical method is validated, both in terms of accuracy and identification success rate, based on leave-one-out cross-validation tests applied on the facial database. Accuracy results are obtained by statistically analyzing the local 3D facial surface differences of the reconstructions and their corresponding ground truth. Identification success rate is obtained by comparing, based on correlation, Euclidean distance matrix (EDM) signatures of the reconstructed and the original 3D facial surfaces in the database. A subjective identification success rate is quantified based on face-pool tests. Finally a qualitative comparison is made between facial reconstructions of a real-case skull, based on two typical static face models and our statistical model, showing the shortcomings of current face models and the improved performance of the statistical model.

Adolescent↗

[Effect of statistical review on manuscript quality in Medicina Clínica (Barcelona): a randomized study].

BACKGROUND AND OBJECTIVE: The statistical review of biomedical articles should result in an improved quality. The objective of this study was to compare the effects of clinical review and joint clinical and statistical review on manuscript quality, in articles submitted to Medicina Clínica (Barcelona), a Spanish weekly journal of internal medicine. METHOD: Original papers arriving between May 2000 and February 2001 were randomized either to a clinical review group or a clinical and statistical review group. Two evaluators, blinded to the paper's group, assessed the quality improvement in both groups, from submission to publication using a modified version ot the Goodman et al. scale. The protocol required that final versions arrived before the end of May 2001. RESULTS: Final sample size was 43 manuscripts, evaluated before and after peer review. On the intention to treat analysis, the estimated effect of statistical review was 1.35 (95% CI: -0.45 to 3.16) positive, but not statistically significant. The analysis of the reviewers' comments revealed some protocol deviations. Taking into account the spontaneous inclusion of statistical experts in the clinical group, the estimated effect was statistically significant, with a confidence interval of 0.3 to 3.7. CONCLUSION: The inclusion of a statistical expert in the peer review process improves manuscript quality, although in the intention to treat analysis the improvement was not statistically significant.

Humans↗

Statistics usage in the American Journal of Obstetrics and Gynecology: has anything changed?

OBJECTIVE: Our purpose was to compare statistical listing and usage between articles published in the American Journal of Obstetrics and Gynecology in 1994 with those published in 1999. STUDY DESIGN: All papers included in the obstetrics, fetus-placenta-newborn, and gynecology sections and the transactions of societies sections of the January through June 1999 issues of the American Journal of Obstetrics and Gynecology (volume 180, numbers 1 to 6) were reviewed for statistical usage. Each paper was given a rating for the cataloging of applied statistics and a rating for the appropriateness of statistical usage, when possible. These results were compared with the data collected on a similar review of articles published in 1994. RESULTS: Of the 238 available articles, 195 contained statistics and were reviewed. In comparison to the articles published in 1994, there were significantly more articles that completely cataloged applied statistics (74.3% vs 47.4%) (P <.0001), and there was a significant improvement in appropriateness of statistical usage (56.4% vs 30.3%) (P <.0001). CONCLUSION: Changes in the Instructions to Authors regarding the description of applied statistics and probable changes in the behavior of researchers and Editors have led to an improvement in the quality of statistics in papers published in the American Journal of Obstetrics and Gynecology.

Gynecology↗

Phylogenetically enhanced statistical tools for RNA structure prediction.

MOTIVATION: Methods that predict the structure of molecules by looking for statistical correlation have been quite effective. Unfortunately, these methods often disregard phylogenetic information in the sequences they analyze. Here, we present a number of statistics for RNA molecular-structure prediction. Besides common pair-wise comparisons, we consider a few reasonable statistics for base-triple predictions, and present an elaborate analysis of these methods. All these statistics incorporate phylogenetic relationships of the sequences in the analysis to varying degrees, and the different nature of these tests gives a wide choice of statistical tools for RNA structure prediction. RESULTS: Starting from statistics that incorporate phylogenetic information only as independent sequence evolution models for each position of a multiple alignment, and extending this idea to a joint evolution model of two positions, we enhance the usual purely statistical methods (e.g. methods based on the Mutual Information statistic) with the use of phylogenetic information available in the sequences. In particular, we present a joint model based on the HKY evolution model, and consequently a X(2) test of independence for two positions. A significant part of this work is devoted to some mathematical analysis of these methods. We tested these statistics on regions of 16S and 23S rRNA, and tRNA.

Base Sequence↗