Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “STATISTICS”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Role of attention and perceptual grouping in visual statistical learning.

Statistical learning has been widely proposed as a mechanism by which observers learn to decompose complex sensory scenes. To determine how robust statistical learning is, we investigated the impact of attention and perceptual grouping on statistical learning of visual shapes. Observers were presented with stimuli containing two shapes that were either connected by a bar or unconnected. When observers were required to attend to both locations at which shapes were presented, the degree of statistical learning was unaffected by whether the shapes were connected or not. However, when observers were required to attend to just one of the shapes' locations, statistical learning was observed only when the shapes were connected. These results demonstrate that visual statistical learning is not just a passive process. It can be modulated by both attention and connectedness, and in natural scenes these factors may constrain the role of stimulus statistics in learning.

Attention↗

Application of statistical methods in quantitative microscopy.

The purpose of the paper is twofold: first, to describe the aspects of quantitative microscopy where statistical ideas are being applied today; and second, to describe some ways in which more complete use of statistics could improve quantitative microscopy. The typical estimation problem in quantitative microscopy is described with emphasis on the modelling of variability in estimation. A Poisson field model for imbedded particulates is used to illustrate the value of theoretical treatments of variability. Sampling, estimation and multivariate analysis are cited as areas of statistics presently used by quantitative microscopists. The potential role of statistics in quantitative microscopy is discussed. Examples from particle sizing and characterization of alveolar lung structure illustrate the value of statistical ideas. Automatic image analysers, while making the collection of large amounts of data feasible, have complicated statistical estimation. An investigation of the size distribution of random chord lengths through alveolar chambers illustrates how statistical methodology must be modified for valid inference.

Animals↗

Comparative phylogeographic summary statistics for testing simultaneous vicariance.

Testing for simultaneous vicariance across comparative phylogeographic data sets is a notoriously difficult problem hindered by mutational variance, the coalescent variance, and variability across pairs of sister taxa in parameters that affect genetic divergence. We simulate vicariance to characterize the behaviour of several commonly used summary statistics across a range of divergence times, and to characterize this behaviour in comparative phylogeographic datasets having multiple taxon-pairs. We found Tajima's D to be relatively uncorrelated with other summary statistics across divergence times, and using simple hypothesis testing of simultaneous vicariance given variable population sizes, we counter-intuitively found that the variance across taxon pairs in Nei and Li's net nucleotide divergence (pi(net)), a common measure of population divergence, is often inferior to using the variance in Tajima's D across taxon pairs as a test statistic to distinguish ancient simultaneous vicariance from variable vicariance histories. The opposite and more intuitive pattern is found for testing more recent simultaneous vicariance, and overall we found that depending on the timing of vicariance, one of these two test statistics can achieve high statistical power for rejecting simultaneous vicariance, given a reasonable number of intron loci (> 5 loci, 400 bp) and a range of conditions. These results suggest that components of these two composite summary statistics should be used in future simulation-based methods which can simultaneously use a pool of summary statistics to test comparative the phylogeographic hypotheses we consider here.

Classification↗

Assessment of statistical procedures used in papers in the Australian Veterinary Journal.

One hundred and thirty-three papers (80 Original Articles and 53 Short Contributions) of 279 papers in 23 consecutive issues of the Australian Veterinary Journal were examined for their statistical content. Only 38 (29%) would have been acceptable to a statistical referee without revision, revision would have been indicated in 88 (66%), and the remaining 7 (5%) had major flaws. Weaknesses in design were found in 40 (30%), chiefly in respect to randomisation and to the size of the experiment. Deficiencies in analysis in 60 (45%) were in methods, application and calculation, and in the failure to use appropriate methods for multiple comparisons and repeated measures. Problems were detected in presentation in 44 (33%) of papers, with insufficient information about the data or its statistical analysis and presentation of statistics (appropriate missing or inappropriate shown) the main problems. Conclusions were considered to be inconsistent with the analysis in 35 (26%) of papers, due mainly to their interpretation of the results of significance testing. It is suggested that statistical refereeing, the publication of statistical guidelines for authors and statistical advice to Animal Experimentation Ethics Committees could all play a part in achieving improvement.

Analysis of Variance↗

Statistics and mathematics anxiety in social science students: some interesting parallels.

This study illuminates some interesting parallels between statistics anxiety and mathematics anxiety in social science students. Parallel to what is confirmed for mathematics anxiety, two factors were observed to underly statistics anxiety scores, namely, statistics test anxiety and content anxiety. The study revealed modest though significant correlations between student attributes and the two confirmed dimensions of statistics anxiety. Furthermore, parallel to the inverse correlation reported for mathematics anxiety and maths course performance, statistics anxiety correlated negatively with high school matriculation scores in maths as well as self perceptions of maths abilities. These data lend support to the hypothesis that aversive prior experiences with mathematics, prior poor achievement in maths, and a low sense of maths self-efficacy are meaningful antecedent correlates of statistics anxiety and thus lend some credence to the "deficit" interpretation of statistics anxiety.

Achievement↗

Normal theory based test statistics in structural equation modelling.

Even though data sets in psychology are seldom normal, the statistics used to evaluate covariance structure models are typically based on the assumption of multivariate normality. Consequently, many conclusions based on normal theory methods are suspect. In this paper, we develop test statistics that can be correctly applied to the normal theory maximum likelihood estimator. We propose three new asymptotically distribution-free (ADF) test statistics that technically must yield improved behaviour in samples of realistic size, and use Monte Carlo methods to study their actual finite sample behaviour. Results indicate that there exists an ADF test statistic that also performs quite well in finite sample situations. Our analysis shows that various forms of ADF test statistics are sensitive to model degrees of freedom rather than to model complexity. A new index is proposed for evaluating whether a rescaled statistic will be robust. Recommendations are given regarding the application of each test statistic.

Humans↗

Fundamental concepts in statistics: elucidation and illustration.

Fundamental concepts in statistics form the cornerstone of scientific inquiry. If we fail to understand fully these fundamental concepts, then the scientific conclusions we reach are more likely to be wrong. This is more than supposition: for 60 years, statisticians have warned that the scientific literature harbors misunderstandings about basic statistical concepts. Original articles published in 1996 by the American Physiological Society's journals fared no better in their handling of basic statistical concepts. In this review, we summarize the two main scientific uses of statistics: hypothesis testing and estimation. Most scientists use statistics solely for hypothesis testing; often, however, estimation is more useful. We also illustrate the concepts of variability and uncertainty, and we demonstrate the essential distinction between statistical significance and scientific importance. An understanding of concepts such as variability, uncertainty, and significance is necessary, but it is not sufficient; we show also that the numerical results of statistical analyses have limitations.

Humans↗

Statistical Viewer: a tool to upload and integrate linkage and association data as plots displayed within the Ensembl genome browser.

BACKGROUND: To facilitate efficient selection and the prioritization of candidate complex disease susceptibility genes for association analysis, increasingly comprehensive annotation tools are essential to integrate, visualize and analyze vast quantities of disparate data generated by genomic screens, public human genome sequence annotation and ancillary biological databases. We have developed a plug-in package for Ensembl called "Statistical Viewer" that facilitates the analysis of genomic features and annotation in the regions of interest defined by linkage analysis. RESULTS: Statistical Viewer is an add-on package to the open-source Ensembl Genome Browser and Annotation System that displays disease study-specific linkage and/or association data as 2 dimensional plots in new panels in the context of Ensembl's Contig View and Cyto View pages. An enhanced upload server facilitates the upload of statistical data, as well as additional feature annotation to be displayed in DAS tracts, in the form of Excel Files. The Statistical View panel, drawn directly under the ideogram, illustrates lod score values for markers from a study of interest that are plotted against their position in base pairs. A module called "Get Map" easily converts the genetic locations of markers to genomic coordinates. The graph is placed under the corresponding ideogram features a synchronized vertical sliding selection box that is seamlessly integrated into Ensembl's Contig- and Cyto- View pages to choose the region to be displayed in Ensembl's "Overview" and "Detailed View" panels. To resolve Association and Fine mapping data plots, a "Detailed Statistic View" plot corresponding to the "Detailed View" may be displayed underneath. CONCLUSION: Features mapping to regions of linkage are accentuated when Statistic View is used in conjunction with the Distributed Annotation System (DAS) to display supplemental laboratory information such as differentially expressed disease genes in private data tracks. Statistic View is a novel and powerful visual feature that enhances Ensembl's utility as valuable resource for integrative genomic-based approaches to the identification of candidate disease susceptibility genes. At present there are no other tools that provide for the visualization of 2-dimensional plots of quantitative data scores against genomic coordinates in the context of a primary public genome annotation browser.

Chromosome Mapping↗

Statistics on continuous IBD data: exact distribution evaluation for a pair of full(half)-sibs and a pair of a (great-) grandchild with a (great-) grandparent.

BACKGROUND: Pairs of related individuals are widely used in linkage analysis. Most of the tests for linkage analysis are based on statistics associated with identity by descent (IBD) data. The current biotechnology provides data on very densely packed loci, and therefore, it may provide almost continuous IBD data for pairs of closely related individuals. Therefore, the distribution theory for statistics on continuous IBD data is of interest. In particular, distributional results which allow the evaluation of p-values for relevant tests are of importance. RESULTS: A technology is provided for numerical evaluation, with any given accuracy, of the cumulative probabilities of some statistics on continuous genome data for pairs of closely related individuals. In the case of a pair of full-sibs, the following statistics are considered: (i) the proportion of genome with 2 (at least 1) haplotypes shared identical-by-descent (IBD) on a chromosomal segment, (ii) the number of distinct pieces (subsegments) of a chromosomal segment, on each of which exactly 2 (at least 1) haplotypes are shared IBD. The natural counterparts of these statistics for the other relationships are also considered. Relevant Maple codes are provided for a rapid evaluation of the cumulative probabilities of such statistics. The genomic continuum model, with Haldane's model for the crossover process, is assumed. CONCLUSIONS: A technology, together with relevant software codes for its automated implementation, are provided for exact evaluation of the distributions of relevant statistics associated with continuous genome data on closely related individuals.

Chromosome Mapping↗

Improving the statistical detection of regulated genes from microarray data using intensity-based variance estimation.

BACKGROUND: Gene microarray technology provides the ability to study the regulation of thousands of genes simultaneously, but its potential is limited without an estimate of the statistical significance of the observed changes in gene expression. Due to the large number of genes being tested and the comparatively small number of array replicates (e.g., N = 3), standard statistical methods such as the Student's t-test fail to produce reliable results. Two other statistical approaches commonly used to improve significance estimates are a penalized t-test and a Z-test using intensity-dependent variance estimates. RESULTS: The performance of these approaches is compared using a dataset of 23 replicates, and a new implementation of the Z-test is introduced that pools together variance estimates of genes with similar minimum intensity. Significance estimates based on 3 replicate arrays are calculated using each statistical technique, and their accuracy is evaluated by comparing them to a reliable estimate based on the remaining 20 replicates. The reproducibility of each test statistic is evaluated by applying it to multiple, independent sets of 3 replicate arrays. Two implementations of a Z-test using intensity-dependent variance produce more reproducible results than two implementations of a penalized t-test. Furthermore, the minimum intensity-based Z-statistic demonstrates higher accuracy and higher or equal precision than all other statistical techniques tested. CONCLUSION: An intensity-based variance estimation technique provides one simple, effective approach that can improve p-value estimates for differentially regulated genes derived from replicated microarray datasets. Implementations of the Z-test algorithms are available at http://vessels.bwh.harvard.edu/software/papers/bmcg2004.

DNA, Complementary↗

Teaching statistics to medical students using problem-based learning: the Australian experience.

BACKGROUND: Problem-based learning (PBL) is gaining popularity as a teaching method in UK medical schools, but statistics and research methods are not being included in this teaching. There are great disadvantages in omitting statistics and research methods from the main teaching. PBL is well established in Australian medical schools. The Australian experience in teaching statistics and research methods in curricula based on problem-based learning may provide guidance for other countries, such as the UK, where this method is being introduced. METHODS: All Australian medical schools using PBL were visited, with two exceptions. Teachers of statistics and medical education specialists were interviewed. For schools which were not visited, information was obtained by email. RESULTS: No Australian medical school taught statistics and research methods in a totally integrated way, as part of general PBL teaching. In some schools, statistical material was integrated but taught separately, using different tutors. In one school, PBL was used only for 'public health' related subjects. In some, a parallel course using more traditional techniques was given alongside the PBL teaching of other material. This model was less successful than the others. CONCLUSIONS: There are several difficulties in implementing an integrated approach. However, not integrating is detrimental to statistics and research methods teaching, which is of particular concern in the age of evidence-based medicine. Some possible ways forward are suggested.

Australia↗

An assessment of recently published gene expression data analyses: reporting experimental design and statistical factors.

BACKGROUND: The analysis of large-scale gene expression data is a fundamental approach to functional genomics and the identification of potential drug targets. Results derived from such studies cannot be trusted unless they are adequately designed and reported. The purpose of this study is to assess current practices on the reporting of experimental design and statistical analyses in gene expression-based studies. METHODS: We reviewed hundreds of MEDLINE-indexed papers involving gene expression data analysis, which were published between 2003 and 2005. These papers were examined on the basis of their reporting of several factors, such as sample size, statistical power and software availability. RESULTS: Among the examined papers, we concentrated on 293 papers consisting of applications and new methodologies. These papers did not report approaches to sample size and statistical power estimation. Explicit statements on data transformation and descriptions of the normalisation techniques applied prior to data analyses (e.g. classification) were not reported in 57 (37.5%) and 104 (68.4%) of the methodology papers respectively. With regard to papers presenting biomedical-relevant applications, 41(29.1 %) of these papers did not report on data normalisation and 83 (58.9%) did not describe the normalisation technique applied. Clustering-based analysis, the t-test and ANOVA represent the most widely applied techniques in microarray data analysis. But remarkably, only 5 (3.5%) of the application papers included statements or references to assumption about variance homogeneity for the application of the t-test and ANOVA. There is still a need to promote the reporting of software packages applied or their availability. CONCLUSION: Recently-published gene expression data analysis studies may lack key information required for properly assessing their design quality and potential impact. There is a need for more rigorous reporting of important experimental factors such as statistical power and sample size, as well as the correct description and justification of statistical methods applied. This paper highlights the importance of defining a minimum set of information required for reporting on statistical design and analysis of expression data. By improving practices of statistical analysis reporting, the scientific community can facilitate quality assurance and peer-review processes, as well as the reproducibility of results.

Analysis of Variance↗

[Statistical considerations for preparation of a study protocol in pharmacological studies].

In order to develop a new drug with scientific rationale, it is important to design a study properly. We suggested appropriate statistical considerations in the course of planning the study that are helpful for correct evaluation of the results of a pharmacological study, the information of which will be reflected in the planning of subsequent pharmacological and clinical studies. The statistical considerations for designing a pharmacological study are as follows: 1. clarification about the purpose of the study, 2. clear statement about the endpoint of the study, 3. framing of hypotheses, 4. selection of statistical analysis method, 5. statistical considerations about study design, 6. statistical considerations of planning about statistical analysis and describing the statistical analysis. We expect researchers will be able to obtain a more reliable conclusion by preparing a study protocol taking our suggestions into consideration.

Pharmacology↗

A comparison of three statistics for detecting differences in digitized dental radiographs: a simulation study.

OBJECTIVES: Because of methodology-induced structural differences in dental radiographs, determination of change has always depended upon expert interpretation. However, new methods should be able to considerably reduce structured error in digitized subtracted images. Once true change in density is obscured only by random variation in pixel density, statistical methods may be brought to bear on the problem of detecting change. The most appropriate statistic is not obvious, however, since density change can be quantified with respect to both magnitude and dimensional extent. Whereas mean density loss is often intuitively defined as the average density of those pixels losing density (to preclude gaining pixels from offsetting losing pixels), the extent of change may be defined in a variety of ways. In this study, extent was defined as either (a) the total number of pixels losing density, or (b) the size of the largest cluster of losing pixels. The object was to evaluate the comparative statistical power of three possible statistics (based on mean density, number of losing pixels, and size of largest losing cluster) for detecting change. METHODS: In a series of simulations of comparative clinical trials, density was reduced in the centre of 1600-pixel square regions of interest by either one or 10 grey-scale units, and t-tests, based on the three statistics, were then compared for their ability to detect differences. RESULTS: Each of the three statistics was shown to exhibit superior relative power under particular conditions of loss magnitude, loss distribution, and pixel threshold for change. CONCLUSION: Selection of the appropriate statistic for identifying change between radiographs will require further information about the anticipated distribution of density changes for the different disease processes under investigation.

Absorptiometry, Photon↗

Effect of varying the case mix on the standardized mortality ratio and W statistic: A simulation study.

OBJECTIVE: To evaluate the validity of using the standardized mortality ratio (SMR) and the W statistic as risk-adjusted measures of hospital mortality to judge ICU performance. DESIGN: APACHE (acute physiology and chronic health evaluation) II data were collected prospectively from the surgical ICU (SICU) at a single institution using all adult admissions (n = 6806) over an 8-year period (excluding cardiac surgical patients, burn patients, and patients under 16 years of age). Using a computer simulation technique, virtual ICUs (VICUs) with mortality rates between 5% and 16% were constructed. After first dividing the original data set into deciles of risk, each VICU was constructed by randomly resampling between 10 and 680 patients from each decile. The SMR, W statistic, and Z statistic were calculated for 10,000 different case mixes. SETTING: The SICU at a 450-bed teaching hospital. PATIENTS: A group of 6,806 adult patient admissions, excluding cardiac surgical patients and burn patients. MEASUREMENTS AND RESULTS: VICUs were created from a data set of actual patients treated at one institution in order to test the hypothesis that the SMR and W statistic would remain invariant when applied to subsets of patients from a single institution. Instead, the SMR and W statistic were found to be very sensitive to changes in case mix. The SMR and W statistic were linear functions of the simulated ICU mortality rate. CONCLUSION: This simulation demonstrates that the SMR and the W statistic based on APACHE II cannot be used to compare outcomes of ICUs. We have proposed a revision of the SMR that eliminates the effect of case mix and allows for more accurate comparisons of ICU performance.

APACHE↗

[Diversity and frequency of scientific research design and statistical methods in the "Arquivos Brasileiros de Oftalmologia": a systematic review of the "Arquivos Brasileiros de Oftalmologia"--1993-2002].

PURPOSE: To verify the frequency of study design, applied statistical analysis and approval by institutional review offices (Ethics Committee) of articles published in the "Arquivos Brasileiros de Oftalmologia" during a 10-year interval, with later comparative and critical analysis by some of the main international journals in the field of Ophthalmology. METHODS: Systematic review without metanalysis was performed. Scientific papers published in the "Arquivos Brasileiros de Oftalmologia" between January 1993 and December 2002 were reviewed by two independent reviewers and classified according to the applied study design, statistical analysis and approval by the institutional review offices. To categorize those variables, a descriptive statistical analysis was used. RESULTS: After applying inclusion and exclusion criteria, 584 articles for evaluation of statistical analysis and, 725 articles for evaluation of study design were reviewed. Contingency table (23.10%) was the most frequently applied statistical method, followed by non-parametric tests (18.19%), Student's t test (12.65%), central tendency measures (10.60%) and analysis of variance (9.81%). Of 584 reviewed articles, 291 (49.82%) presented no statistical analysis. Observational case series (26.48%) was the most frequently used type of study design, followed by interventional case series (18.48%), observational case description (13.37%), non-random clinical study (8.96%) and experimental study (8.55%). CONCLUSION: We found a higher frequency of observational clinical studies, lack of statistical analysis in almost half of the published papers. Increase in studies with approval by institutional review Ethics Committee was noted since it became mandatory in 1996.

Bibliometrics↗

BLMT: statistical sequence analysis using N-grams.

UNLABELLED: Statistical analysis of amino acid and nucleotide sequences, especially sequence alignment, is one of the most commonly performed tasks in modern molecular biology. However, for many tasks in bioinformatics, the requirement for the features in an alignment to be consecutive is restrictive and "n-grams" (aka k-tuples) have been used as features instead. N-grams are usually short nucleotide or amino acid sequences of length n, but the unit for a gram may be chosen arbitrarily. The n-gram concept is borrowed from language technologies where n-grams of words form the fundamental units in statistical language models. Despite the demonstrated utility of n-gram statistics for the biology domain, there is currently no publicly accessible generic tool for the efficient calculation of such statistics. Most sequence analysis tools will disregard matches because of the lack of statistical significance in finding short sequences. This article presents the integrated Biological Language Modeling Toolkit (BLMT) that allows efficient calculation of n-gram statistics for arbitrary sequence datasets. AVAILABILITY: BLMT can be downloaded from http://www.cs.cmu.edu/~blmt/source and installed for standalone use on any Unix platform or Unix shell emulation such as Cygwin on the Windows platform. Specific tools and usage details are described in a "readme" file. The n-gram computations carried out by the BLMT are part of a broader set of tools borrowed from language technologies and modified for statistical analysis of biological sequences; these are available at http://flan.blm.cs.cmu.edu/.

Algorithms↗

Using statistics in employment discrimination cases.

Statistics play an important role in employment discrimination cases, and this role will expand as the legal profession becomes increasingly aware of the utility of the increasingly sophisticated statistical methods available. Historically, plaintiffs in employment discrimination cases have used statistics to establish prima facie cases concerning inequities in areas such as wage rates, personnel selection, promotion, layoff, and termination decisions. Defendants have also employed statistics to demonstrate the fairness of employment practices and policies--for example, by providing statistical evidence of the validity of a test used in personnel selection. This article provides an overview of the role of statistics and the major statistical techniques employed in discrimination cases.

Black or African American↗