Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “STATISTICS”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Evaluating national cause-of-death statistics: principles and application to the case of China.

Mortality statistics systems provide basic information on the levels and causes of mortality in populations. Only a third of the world's countries have complete civil registration systems that yield adequate cause-specific mortality data for health policy-making and monitoring. This paper describes the development of a set of criteria for evaluating the quality of national mortality statistics and applies them to China as an example. The criteria cover a range of structural, statistical and technical aspects of national mortality data. Little is known about cause-of-death data in China, which is home to roughly one-fifth of the world's population. These criteria were used to evaluate the utility of data from two mortality statistics systems in use in China, namely the Ministry of Health-Vital Registration (MOH-VR) system and the Disease Surveillance Point (DSP) system. We concluded that mortality registration was incomplete in both. No statistics were available for geographical subdivisions of the country to inform resource allocation or for the monitoring of health programmes. Compilation and publication of statistics is irregular in the case of the DSP, and they are not made publicly available at all by the MOH-VR. More research is required to measure the content validity of cause-of-death attribution in the two systems, especially due to the use of verbal autopsy methods in rural areas. This framework of criteria-based evaluation is recommended for the evaluation of national mortality data in developing countries to determine their utility and to guide efforts to improve their value for guiding policy.

Adolescent↗

A brief history of numbers and statistics with cytometric applications.

A brief history of numbers and statistics traces the development of numbers from prehistory to completion of our current system of numeration with the introduction of the decimal fraction by Viete, Stevin, Burgi, and Galileo at the turn of the 16th century. This was followed by the development of what we now know as probability theory by Pascal, Fermat, and Huygens in the mid-17th century which arose in connection with questions in gambling with dice and can be regarded as the origin of statistics. The three main probability distributions on which statistics depend were introduced and/or formalized between the mid-17th and early 19th centuries: the binomial distribution by Pascal; the normal distribution by de Moivre, Gauss, and Laplace, and the Poisson distribution by Poisson. The formal discipline of statistics commenced with the works of Pearson, Yule, and Gosset at the turn of the 19th century when the first statistical tests were introduced. Elementary descriptions of the statistical tests most likely to be used in conjunction with cytometric data are given and it is shown how these can be applied to the analysis of difficult immunofluorescence distributions when there is overlap between the labeled and unlabeled cell populations.

Cell Count↗

Comparison of multiple point and statistical motor unit number estimation.

This study compares two common techniques for motor unit number estimation, multiple point stimulation and statistical method, to determine which is more reproducible. Surface recorded motor unit action potentials (SMUPs) of the left hypothenar muscle group were measured on 20 controls and 10 ALS patients. For multiple point, 10 different threshold SMUPs were recorded. For statistical method, mean SMUP amplitude was measured at several stimulus levels, typically spanning >40% of CMAP amplitude range. Both techniques were performed twice, results averaged, electrodes changed, and all recording repeated. For controls, mean of two motor unit number estimation (MUNE) (+/- standard deviation) was 60 (+/-5) for statistical method, and 108 (+/-38) for multiple point. For ALS patients, these values were 21 (+/-16) for statistical method and 55 (+/-39) for multiple point. Test-retest correlation coefficients and coefficients of variation for mean of two MUNE were 0.98 and 7% for statistical method, and 0.90 and 12% for multiple point, respectively. Statistical method was more reproducible and faster than multiple point, supporting its utility in monitoring rates of MUNE change.

Adult↗

Fine-needle aspiration biopsy of Hurthle cell lesions of the thyroid gland: A cytomorphologic study of 139 cases with statistical analysis.

BACKGROUND: Lesions of the thyroid gland composed of Hurthle cells encompass pathologic entities ranging from hyperplastic nodules with Hurthle cell metaplasia to Hurthle cell carcinomas. The cytologic distinction between these entities can be diagnostically challenging. Many cytologic features of Hurthle cell lesions that distinguish neoplastic Hurthle cell lesions requiring surgery from those that are benign and nonneoplastic have been described, but with variable usefulness. This is due, in part, to the small numbers of cases examined in previous studies and the limited application of statistical analysis. A morphologic study was made of 139 Hurthle cell lesions of the thyroid gland and statistical analysis applied to identify a set of cytomorphologic features that distinguish benign Hurthle cell lesions (BHCL) from Hurthle cell neoplasms (HCN). METHODS: Fine-needle aspiration biopsies (FNABs) of thyroid nodules with a predominant Hurthle cell component and corresponding histologic followup were included in the study. Cases were divided into BHCL and HCN groups on the basis of the histologic diagnosis. All cases were reviewed to assess the following 14 cytologic features: overall cellularity, cytoarchitecture, percentage of Hurthle cells, percentage of single cells, percentage of follicular cells observed as naked Hurthle cell nuclei, background colloid, chronic inflammation, cystic change, transgressing blood vessels (TBV), intracytoplasmic lumina, presence of multinucleated Hurthle cells, nuclear to cytoplasmic ratio, nuclear pleomorphism/atypia, and nucleolar prominence. The results were evaluated by using univariate and stepwise logistic regression (SLR) analysis; statistical significance was achieved at P-values < 0.05. RESULTS: One hundred thirty-nine FNAB specimens, corresponding to 56 HCN and 83 BHCL, fulfilled the study criteria. Six of the 14 cytologic features evaluated were shown by univariate analysis to be statistically significant in predicting HCN: nonmacrofollicular architecture (P < 0.001), absence of background colloid (P < 0.001), absence of chronic inflammation (P < 0.001), presence of TBV (P < 0.001), > 90% Hurthle cells (P < 0.001), and >10% single Hurthle cells (P = 0.014). The first four of these features were also shown to be statistically significant in the SLR analysis (P = 0.005, 0.010, 0.016, and 0.045, respectively), and when all four of these features were present HCN was correctly identified 86% of the time. CONCLUSIONS: In the current study of 139 FNAB specimens of thyroid Hurthle cell nodules, 14 cytologic features were examined and 6 were found to be statistically significant in identifying HCN. The following four features, when found in combination, were found to be highly predictive of HCN: nonmacrofollicular architecture, absence of colloid, absence of inflammation, and presence of TBV.

Adenoma↗

Statistical potentials for fold assessment.

A protein structure model generally needs to be evaluated to assess whether or not it has the correct fold. To improve fold assessment, four types of a residue-level statistical potential were optimized, including distance-dependent, contact, Phi/Psi dihedral angle, and accessible surface statistical potentials. Approximately 10,000 test models with the correct and incorrect folds were built by automated comparative modeling of protein sequences of known structure. The criterion used to discriminate between the correct and incorrect models was the Z-score of the model energy. The performance of a Z-score was determined as a function of many variables in the derivation and use of the corresponding statistical potential. The performance was measured by the fractions of the correctly and incorrectly assessed test models. The most discriminating combination of any one of the four tested potentials is the sum of the normalized distance-dependent and accessible surface potentials. The distance-dependent potential that is optimal for assessing models of all sizes uses both C(alpha) and C(beta) atoms as interaction centers, distinguishes between all 20 standard residue types, has the distance range of 30 A, and is derived and used by taking into account the sequence separation of the interacting atom pairs. The terms for the sequentially local interactions are significantly less informative than those for the sequentially nonlocal interactions. The accessible surface potential that is optimal for assessing models of all sizes uses C(beta) atoms as interaction centers and distinguishes between all 20 standard residue types. The performance of the tested statistical potentials is not likely to improve significantly with an increase in the number of known protein structures used in their derivation. The parameters of fold assessment whose optimal values vary significantly with model size include the size of the known protein structures used to derive the potential and the distance range of the accessible surface potential. Fold assessment by statistical potentials is most difficult for the very small models. This difficulty presents a challenge to fold assessment in large-scale comparative modeling, which produces many small and incomplete models. The results described in this study provide a basis for an optimal use of statistical potentials in fold assessment.

Algorithms↗

A spatial scan statistic for ordinal data.

Spatial scan statistics are widely used for count data to detect geographical disease clusters of high or low incidence, mortality or prevalence and to evaluate their statistical significance. Some data are ordinal or continuous in nature, however, so that it is necessary to dichotomize the data to use a traditional scan statistic for count data. There is then a loss of information and the choice of cut-off point is often arbitrary. In this paper, we propose a spatial scan statistic for ordinal data, which allows us to analyse such data incorporating the ordinal structure without making any further assumptions. The test statistic is based on a likelihood ratio test and evaluated using Monte Carlo hypothesis testing. The proposed method is illustrated using prostate cancer grade and stage data from the Maryland Cancer Registry. The statistical power, sensitivity and positive predicted value of the test are examined through a simulation study.

Cluster Analysis↗

Publications from multicentre clinical trials: statistical techniques and accessibility to the reader.

Articles from multicentre randomized clinical trials were analysed by methods adapted from Emerson and Colditz and Juzych et al. to compare the frequency with which different statistical methods are used by clinical trials investigators with the frequency used by other researchers, and to determine how much statistical knowledge is required to interpret the statistical treatment of data from clinical trials. We observed differences between the frequency of usage of statistical methods and the accessibility of the clinical trials publications and those of all medical research articles published by specific journals. Clinical trials publications are less accessible than others in medical journals to the reader who knows only descriptive statistics, t-tests, contingency tables, power calculations, and life table methods. Many more statistical methods must be known by a reader to understand fully publications regarding treatment group comparisons for the primary outcomes of interest from clinical trials.

Data Interpretation, Statistical↗

Teaching statistics to non-specialists.

Teaching statistics to non-specialists is a challenge for which most statisticians are unprepared by their own training within an academic mathematics department. Most statistics courses for medical undergraduates still focus on research statistics, whereas it would be more appropriate to concentrate on statistics relevant to clinical decision-making about an individual patient. Teaching statistics to Master of Public Health students presents further challenges because of the wide variety of their backgrounds and the greater demands from mature postgraduates. Whatever the audience, however, the same principles apply: medical statistics should be taught as non-mathematically as possible, only introducing formulae when absolutely necessary and explaining their components; plenty of practical applications should be given; there should be ample opportunity for practice to gain hands-on experience using both calculators and computers (preferably with MINITAB); and tutorials should be streamed according to perceived mathematical ability, with remedial mathematics teaching available to those who need it.

Curriculum↗

Statistics of sexual size dimorphism.

In comparative studies of sexual size dimorphism (SSD), the methods used to quantify dimorphism are controversial. SSD is commonly expressed as a ratio between species mean values of males and females, such as M/F or (M-F)/([M+F]/2), but a number of investigators have suggested that ratios should not be used, mainly because their distributions usually violate the assumptions of parametric statistical tests, or because they lead to spurious relationships that invalidate the interpretation and statistical significance of regressions and correlations. As an alternative to ratios, the comparative study of SSD can be conducted by a combination of regression with sex-specific data and residuals from this regression. Twenty-five data sets were selected from the literature and used to duplicate a variety of statistical procedures commonly employed in studies of SSD. All analyses were repeated with five different ratios and with methods that avoid the calculation of any ratios. These data and a review of the statistical properties of ratios and residuals indicate that: (1) most of the ratios used in the SSD literature are unnecessary, and several commonly used ratios are statistically inferior to others. Only two ratios are needed, one on a logarithmic scale and one on a linear scale; (2) there is no problem with spurious correlation or non-normality when ratios are used in several types of statistical procedures commonly employed in studies of SSD; (3) residuals cannot replace ratios for the evaluation of many questions regarding the pattern of SSD among species; and (4) residuals usually are used incorrectly, leading to misspecified regression equations. Most of the questions for which residuals are used should be addressed by multiple regression. These results apply to studies using comparative methods with or without adjustments for phylogenetic effects.

Animals↗

Statistical analysis of drug interactions in anesthesia.

Anesthesiologists often use more than one drug in a patient to achieve a target response, such as a desired blood pressure. Isobolographic analysis is a standard mathematical method to test whether a drug-drug interaction exists and, if so, whether the interaction is synergistic or antagonistic. Several experimental protocols are suitable to collect data for interpretation by isobolographic analysis. Traditionally, each subject would receive a single dose of one or more drugs and the presence or absence of the target response in the subject at a specific time would be recorded. An alternative procedure is to infuse one or more drugs into each subject until the target response occurs. This procedure is clinically relevant to studies of anesthesia drugs, which are often titrated to achieve a target response. This latter method of testing for drug interaction requires fewer subjects than the traditional approach because more information is obtained from each subject. We present a statistical test for drug-drug interaction that uses the total doses of drugs given to each subject. This statistical test can be used as part of a complete analysis of drug-drug interaction. We consider (i) selection of appropriate sample sizes and doses; (ii) randomizing subjects to groups; (iii) effects of unequal group variances on the accuracy of the statistical test; (iv) appropriate methods to present the drug-drug interaction data graphically; and (v) tests to ensure that statistical assumptions are satisfied. Using the statistical methods described here, drug-drug interaction can be quantified and evaluated for statistical significance.

Anesthetics↗

The participant effect: mortality in a community-based study compared to vital statistics.

The 20-year mortality experience of the community-based Evans County Heart Study population is compared to local, regional and national vital statistics. Deficit mortality occurred in the study population at younger ages while at older ages mortality was similar to or greater than vital statistics. This was particularly true for white and nonwhite males, whose mortality patterns were statistically significantly different from Evans Co. vital statistics (P less than 0.005). Nonwhite/white mortality ratios in the study were close to those observed in local vital statistics, particularly for males. Sex mortality ratios in the study population were lower than in vital statistics due to a stronger participant effect (lower mortality) in males. Evans Co. was an area of particularly high mortality for whites in the period 1960-1980 compared to other parts of Georgia and the U.S. Results of this study are similar to other reports of participant effects in epidemiologic follow-up studies; implications for bias in estimates of population levels of disease and of disease/exposure relationships are discussed.

Adult↗

The effects of signal conditioning on the statistical analyses of gait EMG.

Ensemble averaged EMG profiles generated for leg muscles during gait have been used to clinically assess disease or injury. Several of the methods that have been reported for conditioning gait EMG signals were compared using data collected from clinically normal subjects walking on a treadmill. Specifically investigated were the effects of filtering and the quantity of data averaged upon several statistical tests that measure the variability of, or differences between, EMG profiles. Our results suggest that the variance ratio (VR) provides a reasonable test of data variability because of its modest sensitivity to both the degree of filtering and the amount of data averaged. They also suggest that of the comparison statistics: Pearson's r, the Kolmogorov-Smirnov T test and the ANOVA F ratio, the T test was the most reliable in detecting differences between given profiles for all test conditions. However, recognition of this ability of the T test must be tempered by the knowledge that while obvious EMG signal differences did exist, observable functional differences in gait did not. The relationship between statistically similar/dissimilar EMG patterns and clinically functional/dysfunctional gait patterns needs to be established. In addition, since all of the test statistics studied were affected to some degree by filtering and averaging, care should be used when comparing statistical results from separate studies unless it is known that the studies were conducted under similar conditions, including data processing. To that end, we recommend that at least 20 strides be used in the averaging process since the statistics we tested have reached or are asymptotically approaching their final values by this point.

Analysis of Variance↗

The future of statistics and statisticians in European regulatory affairs.

Roles of statisticians in drug regulatory agencies include assessment of the adequacy of statistical aspects of the design of the programme of efficacy and safety studies, design of individual studies, data collection and quality assurance in each study, analysis of individual studies, interpretation of individual studies, and interpretation of results of the programme. These assessments should be made in collaboration with medical, pharmacological, and toxicological colleagues, and liaison with company statisticians before, during, and after submission of the application if the procedure is to be efficient. Recommendations for improving statistical quality of drug license applications and provision for their statistical review in Europe include: 1. Employment of more experienced, professional statistical assessors in both national and European Community (EC) regulatory authorities, with a wider group of supporting experts 2. Obligatory statistical coauthorship or validation of expert reports before submission of applications, and 3. Unified statistical guidelines for preparation of applications.

Drug and Narcotic Control↗

Predicting the behaviour of proteins in hydrophobic interaction chromatography. 2. Using a statistical description of their surface amino acid distribution.

This paper focuses on the prediction of the dimensionless retention time (DRT) of proteins in hydrophobic interaction chromatography (HIC) by means of mathematical models based on the statistical description of the amino acid surface distribution. Previous models characterises the protein surface as a whole. However, most of the time it is not the whole protein but some of its specific regions that interact with the environment. It seems much more natural to use local measurements of the characteristics of the surface. Therefore, the statistical characterisation of the distribution of an amino acid property on the protein surface was carried out from the systematic calculation of the local average of this property in a neighbourhood placed sequentially on each of the amino acids on the protein surface. This process allowed us to characterise the distribution of this property quantitatively using three main statistics: average, standard deviation and maximum. In particular, if the property considered is a hydrophobicity scale, these statistics allowed us to characterise the average hydrophobicity and the hydrophobic content of the most hydrophobic cluster or hotspot, as well as the heterogeneity of the hydrophobicity distribution on the protein surface. We tested the performance of the DRT predictive models based on these statistics on a set of 15 proteins. We obtained better predictive results with respect to the models previously reported. The best predictive model was a linear model based on the maximum. This statistic was calculated using an index of the mobilities of amino acids in chromatography. The predictive performance of this model (measured as the Jack Knife MSE) was 26.9% better than those obtained by the best model which does not consider the amino acid distribution and 19.5% better than the model based on the hydrophobic imbalance (HI). In addition, the best performance was obtained by a linear multivariable model based on the HI and the maximum. The difference between the experimental data and the prediction carried out by this model was smaller than those observed previously. In fact, this model obtained better predictive capacities than a previous linear multivariable model decreasing the Jack Knife MSE in 8.7%. In addition, this model allowed us to diminish the number of variables required, increasing, in this way, the degrees of freedom of the model.

Amino Acids↗

On the identification of differentially expressed genes: improving the generalized F-statistics for Affymetrix microarray gene expression data.

It has been shown that the generalized F-statistics can give satisfactory performances in identifying differentially expressed genes with microarray data. However, for some complex diseases, it is still possible to identify a high proportion of false positives because of the modest differential expressions of disease related genes and the systematic noises of microarrays. The main purpose of this study is to develop statistical methods for Affymetrix microarray gene expression data so that the impact on false positives from non-expressed genes can be reduced. I proposed two novel generalized F-statistics for identifying differentially expressed genes and a novel approach for estimating adjusting factors. The proposed statistical methods systematically combine filtering of non-expressed genes and identification of differentially expressed genes. For comparison, the discussed statistical methods were applied to an experimental data set for a type 2 diabetes study. In both two- and three-sample analyses, the proposed statistics showed improvement on the control of false positives.

Computer Simulation↗

Visual working memory for image statistics.

To define the role of statistical features of images in visual working memory, we compared the ability of subjects (N=6) to identify changes in arrays of black and white checks when these changes altered some aspect of their statistical structure, versus when these changes did not. Alteration of luminance statistics or local higher-order statistics improved performance, but alteration of the degree of bilateral symmetry did not. The dependence of performance on the degree of statistical change indicated that statistical information was represented in a graded, rather than categorical, fashion.

Adult↗

Beyond fourth-order texture discrimination: generation of extreme-order and statistically-balanced textures.

Julesz introduced the concept of statistically defined textures and their perceptual discrimination. Julesz showed that discrimination was possible with statistics equated to third-order, specifying fourth-order textures. Klein and Tyler offered a variety of paradigms suggesting that fourth order might be the limit on human texture processing. To go beyond this limit, new texture paradigms are now introduced to avoid contamination by luminance extrema, to control local and long-range texture properties, and to provide textures without global statistical structure. Local luminance contamination is avoided by novel orientation plaids, in which higher-order rules govern the orientation of local elements rather than their coloring. These textures allow evaluation of texture discrimination up to thirty-second order by cortical pattern elements. Long-range processing is studied by random strip rotation and by interlacing of independent textures. Each substantially degrades the visibility of the fourth-order textures, revealing that the fourth-order information is conveyed largely by local rather than long-range perturbations from random statistics. Finally, textures equated at all orders can be defined in terms of their global statistics, but may nevertheless readily be discriminated in human vision. The discrimination on the basis of local perturbations implies that human vision assesses textures through a local sampling window, and is largely insensitive to longer-range statistical properties.

Discrimination, Psychological↗

Statistical analysis of synaptic transmission: model discrimination and confidence limits.

Procedures for discriminating between competing statistical models of synaptic transmission, and for providing confidence limits on the parameters of these models, have been developed. These procedures were tested against simulated data and were used to analyze the fluctuations in synaptic currents evoked in hippocampal neurones. All models were fitted to data using the Expectation-Maximization algorithm and a maximum likelihood criterion. Competing models were evaluated using the log-likelihood ratio (Wilks statistic). When the competing models were not nested, Monte Carlo sampling of the model used as the null hypothesis (H0) provided density functions against which H0 and the alternate model (H1) were tested. The statistic for the log-likelihood ratio was determined from the fit of H0 and H1 to these probability densities. This statistic was used to determine the significance level at which H0 could be rejected for the original data. When the competing models were nested, log-likelihood ratios and the chi 2 statistic were used to determine the confidence level for rejection. Once the model that provided the best statistical fit to the data was identified, many estimates for the model parameters were calculated by resampling the original data. Bootstrap techniques were then used to obtain the confidence limits of these parameters.

Algorithms↗