Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “STATISTICS”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

Statistical methods in the Fourier domain to enhance and classify images.

A mathematical model, for which rigorous methods of statistical inference are available, is described and techniques for image enhancement and linear discriminant analysis of groups are developed. Since the gray values of neighboring pixels in tomographically produced medical images are spatially correlated, the calculations are carried out in the Fourier domain to insure statistical independence of the variables. Furthermore, to increase the power of statistical tests the known spatial covariance was used to specify constraints in the spectral domain. These methods were compared to statistical procedures carried out in the spatial domain. Positron emission tomography (PET) images of alcoholics with organic brain disorders were compared by these techniques to age-matched normal volunteers. Although these techniques are employed to analyze group characteristics of functional images, they provide a comprehensive set of mathematical and statistical procedures in the spectral domain that can also be applied to images of other modalities, such as computed tomography (CT) or magnetic resonance imaging (MRI).

Aged↗

Applications of computer-intensive statistical methods to environmental research.

Conventional statistical approaches rely heavily on the properties of the central limit theorem to bridge the gap between the characteristics of a sample and some theoretical sampling distribution. Problems associated with nonrandom sampling, unknown population distributions, heterogeneous variances, small sample sizes, and missing data jeopardize the assumptions of such approaches and cast skepticism on conclusions. Conventional nonparametric alternatives offer freedom from distribution assumptions, but design limitations and loss of power can be serious drawbacks. With the data-processing capacity of today's computers, a new dimension of distribution-free statistical methods has evolved that addresses many of the limitations of conventional parametric and nonparametric methods. Computer-intensive statistical methods involve reshuffling, resampling, or simulating a data set thousands of times to empirically define a sampling distribution for a chosen test statistic. The only assumption necessary for valid results is the random assignment of experimental units to the test groups or treatments. Application to a real data set illustrates the advantages of these methods, including freedom from distribution assumptions without loss of power, complete choice over test statistics, easy adaptation to design complexities and missing data, and considerable intuitive appeal. The illustrations also reveal that computer-intensive methods can be more time consuming than conventional methods and the amount of computer code required to orchestrate reshuffling, resampling, or simulation procedures can be appreciable.

Analysis of Variance↗

Statistical analysis of behavioral toxicology data and studies.

One of the areas of toxicology in which a wide variation occurs in the statistical procedures used to analyze experimental data is behavioral toxicology. Due to either limitations in statistical training in the toxicologists or to a lack of understanding of the underlying biological mechanisms on the part of the statisticians, data is frequently analyzed by methodologies which either do not have optimal characteristics of sensitivity and power or for which the underlying assumptions as to the nature of the data are not valid. To establish a firm basis for an identification of the optimal and most appropriate forms of statistical analysis of behavioral toxicology data (and to design efficient and sensitive studies), the four general types of data (observational scores, response rates, error rates, and times-to-endpoints) and one special class of data (teratology and reproduction) are examined in detail. The present practices as to statistical analysis are then reviewed and suggestions as to optimal methods (based on experience, with data sets presented as examples) are developed and presented. The underlying key to this entire process is to establish the biological and statistical nature of the data being generated and to design and analyze experiments accordingly.

Amphetamine↗

A theory of preattentive texture discrimination based on first-order statistics of textons.

The many indistinguishable texture pairs having identical second-, but different third- and higher-order statistics, led to the conjecture that globally the preattentive texture discrimination system cannot process statistical parameters of third- or higher-order. Thus in cases when iso-second-order textures yield discrimination this must be based on local conspicuous features called textons (Julesz, 1980). Here it is shown that globally even second-order statistical parameters, such as autocorrelation, cannot be processed by the textural system, and texture discrimination is solely the result of first-order statistics (density) of textons. It is also shown that the perceivable distance of statistical constraints (coherence distance) in densely packed stochastic textures is very short, four dots or less. As of now, only three texton classes were found: color, elongated blobs (line segments) of given width, orientation, and length, and the terminators (end-points) of these elongated blobs. The strength of these textons is demonstrated by several examples.

Form Perception↗

Empirical evaluation of two-sample statistical tests for differences of stepping phase during insect walking.

In order to determine which statistical tests can validly be applied to data that describe a temporal relationship between two or more repetitive movements by an animal, we evaluated empirically seven two-sample tests that seemed potentially useful: Student's t test, the Watson Williams test for means, the variance-ratio F test, the Watson Williams test for the concentration parameter k, the Wallraff test, the Mann Whitney test and the Watson U2 test. Evaluations were carried out on the timing (phases) of bursts of muscular activity in one leg relative to those in another during free walking in cockroaches. Each statistical test was evaluated by dividing randomly a single parent set of data into two subsets, each subset containing about half the original data set. This division was repeated 400 times, thus generating 400 different pairs of subsets. Each statistical test was used separately on the pairs of subsets to test the null hypothesis that the two samples of each pair came from the same population; this procedure generated 400 statistics for each test, one for each pair of subsets. An estimate of the reliability of each statistical test was obtained by comparing the number of times the test actually indicated a significant difference between subsets to the number of times it might be expected to do so out (20 out of 400 when tested at the 5% level of significance). This procedure was repeated on ten different sets of data. The outcome of the evaluation suggested that, from an empirical point of view, Student's t, the Mann Whitney, the Wallraff and the Watson U2 tests may be useful in assessing differences among the data we analyzed. The variance-ratio F test and the Watson Williams test for the concentration parameter k were clearly not usable. The Watson Williams test for means might be useful in some circumstances. Performing an arcsine transformation of the data did not significantly alter these results. Possible causes of the inapplicability of some of these tests to phase data are discussed.

Animals↗

The ethics of randomised controlled trials: a matter of statistical belief?

This paper outlines the approaches of two apparently competing schools of statistics. The criticisms made by supporters of Bayesian statistics about conventional Frequentist statistics are explained, and the Bayesian claim that their method enables research into new treatments without the need for clinical trials is examined in detail. Several further important issues are considered, including: the use of historical controls and data routinely collected on patients; balance in randomised trials; the possibility of giving information to patients; patient choice and patient autonomy; and how widely the results of clinical trials can be used. It is concluded that good statistical techniques in the design and analysis of medical studies are essential, but the statistical school used in developing such techniques is relatively unimportant.

Bayes Theorem↗

Estimating and testing autocorrelation with small samples: a comparison of the C-statistic to a modified estimator.

Huitema and McKean (Psychological Bulletin, 110, 291-304, 1991) recently showed, in a Monte-Carlo study, that five conventional estimators of first-order autocorrelation perform poorly for small (< 50) sample sizes. They suggested a modified estimator and a test for autocorrelation. We examine an estimator not considered by Huitema and McKean: the C-statistic (Young, Annals of Mathematical Statistics, 12, 293-300, 1941). A Monte-Carlo study of the small sample properties of the C-statistic shows that it performs as well or better than the modified estimator suggested by Huitema and McKean (1991). The C-statistic is also shown to be closely related to the d-statistic of the widely used Durbin-Watson test.

Biometry↗

Uses of statistical editing of real-time ambulatory electrocardiographic recordings for quantitative ventricular ectopic beat counts.

Real-time ambulatory monitoring analyzes each heart beat, counts events, and stores ECG samples for later visual verification. Typically, a physician examines these to determine whether the computer algorithm accurately identified arrhythmias. Physician editing is performed using best clinical judgement. We developed a simple statistical editing procedure for adjusting false positive and false negative computer detections. In 20 subjects having 24-hr monitoring we compared statistically edited ventricular premature beat (VPB) counts and pair/run counts with the unedited monitor counts and with physician assessment using a visual counted gold standard. The agreement of the statistically edited count with the visual standard was 65% for total VPB, 85% for VPB pair/runs, and 90% for a risk score based on ventricular ectopy. Corresponding agreements for unedited monitor count were 15, 25, and 30%, respectively. Physician assessment was not sufficiently precise to allow quantitative count estimates. This study indicates a statistical editing procedure substantially increases the level of agreement between the visual standard and the monitor count of VPB frequency and complexity. Statistically edited data are suitable for quantitative counts of VPB and other arrhythmic events in research and in medical diagnosis and treatment. This editing procedure can be a useful adjunct to any ambulatory monitoring system.

Arrhythmias, Cardiac↗

Signal statistics in objective auditory evoked potential (AEP) detection by the phase spectral method.

This paper reports on statistical aspects relevant to the use of the phase spectrum of post-stimulus EEG, in objective detection of the auditory evoked potential. The sampling statistics of two statistical estimators are discussed: the mean phase vector magnitude, and the standard deviation of an ensemble of post-stimulus EEG phases. These two estimators are circular statistics, and subject to strong sample size bias. Their confidence intervals have been derived empirically for sample sizes routinely used in clinical audiometry. A trial example illustrates the use of the objective phase statistics developed here; it is noted that the method may also be more efficient than the visual scoring of averaged responses.

Adult↗

Description and reporting of statistical methods.

A review of the last three volumes of AJIC (1990, 1991, and 1992, numbers 1 through 3) revealed a modest use of statistics in infection control research. Discussion of the statistics most commonly used results in the following recommendations: (1) Researchers who use the t test should report the t test statistic, exact significance probabilities, and a discussion of assumptions and requirements underlying the analysis. (2) A 95% confidence interval for estimated quantities is preferred to p values. (3) The potential impact of nonresponse bias should be included in reports of survey data. (4) An adequate description of statistical methods should be included in all reports of survey data. (5) An adequate description of statistical methods should be included in all research articles.

Confidence Intervals↗

Exact statistical tests for any carcinogenic effect in animal bioassays.

Standard statistical treatment of data from carcinogenicity bioassays generally involves separate analyses of data from many tumor responses in each sex in two species. There are two serious difficulties with this approach: the excessive probability of one or more false positive findings due to the large number of individual tests applied and the lack of mutual support among the separate tests (e.g., results that are close to significant from several organs should be allowed to reinforce each other, but such mutual support does not formally occur in statistical tests currently employed). In this paper we propose a class of tests that deals with both of these difficulties. The test statistics proposed are functions of p values from multiple conventional tests. The significance levels are computed by a Monte Carlo randomization procedure that treats individual animals (rather than tumor-specific response scores) as units of variation, so that the assumption of independence of tumors at different sites is not required. A single overall test statistic is derived from results from all individual tumor sites; thus there is proper control for the false positive rate. Mutual support from results from different tumor sites can be obtained by using a test statistic such as the product of the K smallest p values from conventional tests. Suggestions are made regarding specific tests that could be applied routinely to carcinogenesis bioassay data. The usefulness of the proposed tests is demonstrated by applying them to data from a National Toxicology Program bioassay of decabromodiphenyl oxide.

Animals↗

Bayesian ranking of sites for engineering safety improvements: decision parameter, treatability concept, statistical criterion, and spatial dependence.

In recent years, there has been a renewed interest in applying statistical ranking criteria to identify sites on a road network, which potentially present high traffic crash risks or are over-represented in certain type of crashes, for further engineering evaluation and safety improvement. This requires that good estimates of ranks of crash risks be obtained at individual intersections or road segments, or some analysis zones. The nature of this site ranking problem in roadway safety is related to two well-established statistical problems known as the small area (or domain) estimation problem and the disease mapping problem. The former arises in the context of providing estimates using sample survey data for a small geographical area or a small socio-demographic group in a large area, while the latter stems from estimating rare disease incidences for typically small geographical areas. The statistical problem is such that direct estimates of certain parameters associated with a site (or a group of sites) with adequate precision cannot be produced, due to a small available sample size, the rareness of the event of interest, and/or a small exposed population or sub-population in question. Model based approaches have offered several advantages to these estimation problems, including increased precision by "borrowing strengths" across the various sites based on available auxiliary variables, including their relative locations in space. Within the model based approach, generalized linear mixed models (GLMM) have played key roles in addressing these problems for many years. The objective of the study, on which this paper is based, was to explore some of the issues raised in recent roadway safety studies regarding ranking methodologies in light of the recent statistical development in space-time GLMM. First, general ranking approaches are reviewed, which include naïve or raw crash-risk ranking, scan based ranking, and model based ranking. Through simulations, the limitation of using the naïve approach in ranking is illustrated. Second, following the model based approach, the choice of decision parameters and consideration of treatability are discussed. Third, several statistical ranking criteria that have been used in biomedical, health, and other scientific studies are presented from a Bayesian perspective. Their applications in roadway safety are then demonstrated using two data sets: one for individual urban intersections and one for rural two-lane roads at the county level. As part of the demonstration, it is shown how multivariate spatial GLMM can be used to model traffic crashes of several injury severity types simultaneously and how the model can be used within a Bayesian framework to rank sites by crash cost per vehicle-mile traveled (instead of by crash frequency rate). Finally, the significant impact of spatial effects on the overall model goodness-of-fit and site ranking performances are discussed for the two data sets examined. The paper is concluded with a discussion on possible directions in which the study can be extended.

Accidents, Traffic↗

Misuse of statistical tests in Archives of Clinical Neuropsychology publications.

This article reviews the (mis)use of statistical tests in neuropsychology research studies published in the Archives of Clinical Neuropsychology in the years 1990-1992 and 1996-2000, and 2001-2004, prior to, commensurate with the internet-based and paper-based release, and following the release of the American Psychological Association's Task Force on Statistical Inference. The authors focused on four statistical errors: inappropriate use of null hypothesis tests, inappropriate use of P-values, neglect of effect size, and inflation of Type I error rates. Despite the recommendations of the Task Force on Statistical Inference published in 1999, the present study recorded instances of these statistical errors both pre- and post-APA's report, with only the reporting of effect size increasing after the release of the report. Neuropsychologists involved in empirical research should be better aware of the limitations and boundaries of hypothesis testing as well as the theoretical aspects of research methodology.

Bibliometrics↗

The use of United States vital statistics in perinatal and obstetric research.

Vital statistics data have been used to track maternal and child health in the United States since the early 1900s. The breadth of information collected on birth and death certificates coupled with advances in computer processing have made possible critical perinatal and obstetric research. These enhancements also facilitate potentially problematic uses of the same data. This commentary explores characteristics of the United States Vital Statistics System and presents some thoughts with regard to the appropriate use of these data. The advantages of vital statistics include representativeness and the ability to examine subpopulations. Limitations include possible underreporting of medical conditions and procedures, lack of ability to ascertain clinical intent, and the well-known issues with gestational age reporting. Analyses based on vital statistics are important in informing future clinical research projects. However, respecting the limitations of vital statistics data enhance their appropriate role in obstetric and perinatal research.

Epidemiologic Research Design↗

Ability of static and statistical mechanics posturographic measures to distinguish between age and fall risk.

Traditional posturographic analysis and four statistical mechanics techniques were applied to center-of-pressure (COP) trajectories of young, older "low-fall-risk" and older "high-fall-risk" individuals. Low-fall-risk older adults were active 3 days per week in a cardiac rehabilitation program, while high-fall-risk older adults were diagnosed with perilymph fistula. Subjects diagnosed with perilymph fistula must have experienced two of the following vestibular findings: constant disequilibrium, positional vertigo and/or a positive fistula test. Non-parametric statistical tests were used to determine whether the posturographic measures could detect differences between the young and older "low-fall-risk" groups (age comparison) and between the older "low-" and "high-risk" groups (risk of falling comparison). The statistical mechanics techniques were more sensitive than the traditional measures: detecting significant differences between the young and older "low-risk" groups, while none of the traditional measures were significantly different. In addition, interpretation of the statistical mechanics techniques may offer more insight into the nature of the process controlling the COP trajectories. However, the methods offered slightly different explanations. For instance, the Hurst rescaled range analysis suggests that the movement of the COP is governed solely by anti-persistent behavior, whereas the stabilogram diffusion analysis suggests a short-term persistence balanced by a long-term anti-persistence. These discrepancies highlight the need for a model that incorporates the biological systems responsible for maintaining balance and experimental methods to directly quantify their status and roles. Until such a model exists, however, the statistical mechanics techniques appear to have some advantages over traditional posturographic measures for studying balance control.

Accidental Falls↗

Statistical learning of tone sequences by human infants and adults.

Previous research suggests that language learners can detect and use the statistical properties of syllable sequences to discover words in continuous speech (e.g. Aslin, R.N., Saffran, J.R., Newport, E.L., 1998. Computation of conditional probability statistics by 8-month-old infants. Psychological Science 9, 321-324; Saffran, J.R., Aslin, R.N., Newport, E.L., 1996. Statistical learning by 8-month-old infants. Science 274, 1926-1928; Saffran, J., R., Newport, E.L., Aslin, R.N., (1996). Word segmentation: the role of distributional cues. Journal of Memory and Language 35, 606-621; Saffran, J.R., Newport, E.L., Aslin, R.N., Tunick, R.A., Barrueco, S., 1997. Incidental language learning: Listening (and learning) out of the corner of your ear. Psychological Science 8, 101-195). In the present research, we asked whether this statistical learning ability is uniquely tied to linguistic materials. Subjects were exposed to continuous non-linguistic auditory sequences whose elements were organized into 'tone words'. As in our previous studies, statistical information was the only word boundary cue available to learners. Both adults and 8-month-old infants succeeded at segmenting the tone stream, with performance indistinguishable from that obtained with syllable streams. These results suggest that a learning mechanism previously shown to be involved in word segmentation can also be used to segment sequences of non-linguistic stimuli.

Adult↗

A new statistic for the analysis of association between trait and polymorphic marker loci.

Inference for detecting the existence of an association between a diallelic marker and a trait locus is based on the chi-squared statistic with one degree of freedom. For polymorphic markers with m alleles (2), three approaches are mainly used in practice. First, one may use Pearson's chi-squared statistic with m-1 degrees of freedom (d.f.) but this leads to a loss in test power. Second, one can select an allele to be the most associated and then collapse the other allele categories into a single class. This reduces in a biased way, the locus to a diallelic system. Third, one may use the Terwilliger [J.D. Terwilliger, Am. J. Hum. Genet. 56 (1995) 777] likelihood ratio statistic which has a non-standard unknown limiting probability distribution. In this paper, we propose a new statistic, L(D), based on the second testing approach. We derive the asymptotic probability distribution of L(D) in an easy way. Simulation studies show that L(D) is more powerful than Pearson's chi-squared statistic with m-1 d.f.

Alleles↗

Extrapolating traditional DNA microarray statistics to tiling and protein microarray technologies.

A credit to microarray technology is its broad application. Two experiments--the tiling microarray experiment and the protein microarray experiment--are exemplars of the versatility of the microarrays. With the technology's expanding list of uses, the corresponding bioinformatics must evolve in step. There currently exists a rich literature developing statistical techniques for analyzing traditional gene-centric DNA microarrays, so the first challenge in analyzing the advanced technologies is to identify which of the existing statistical protocols are relevant and where and when revised methods are needed. A second challenge is making these often very technical ideas accessible to the broader microarray community. The aim of this chapter is to present some of the most widely used statistical techniques for normalizing and scoring traditional microarray data and indicate their potential utility for analyzing the newer protein and tiling microarray experiments. In so doing, we will assume little or no prior training in statistics of the reader. Areas covered include background correction, intensity normalization, spatial normalization, and the testing of statistical significance.

Animals↗