Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

CAN'T MISS--conquer any number task by making important statistics simple. Part 2. Probability, populations, samples, and normal distributions.

Healthcare quality improvement professionals need to understand and use inferential statistics to interpret sample data from their organizations. In quality improvement and healthcare research studies all the data from a population often are not available, so investigators take samples and make inferences about the population by using inferential statistics. This three-part series will give readers an understanding of the concepts of inferential statistics as well as the specific tools for calculating confidence intervals for samples of data. This article, Part 2, describes probability, populations, and samples. The uses of descriptive and inferential statistics are outlined. The article also discusses the properties and probability of normal distributions, including the standard normal distribution.

Data Interpretation, Statistical↗

Change of statistical parameters of transmitter release during various kinetic tests in unparalysed voltage-clamped rat diaphragm.

1. The statistical nature of transmitter release was studied in unparalysed cut rat diaphragm using a voltage clamp technique at room temperature (23 degrees C). 2. While the binomial distribution described observed amplitude histograms well, the Poisson distribution was clearly inadequate. Values of m ranged from 40 to 45, while values for n varied from 45 to 53 and p was between 0.82 and 0.90. 3. Estimation of the probability of release from the transient decay of e.p.c.s in short tetanic trains (11 pulses, 150 Hz) gave values of p from 0.031 to 0.054, which are more than one order of magnitude lower than statistical estimates. 4. As a result of the short tetanic stimulation (5 pulses, 20, 50 and 100 Hz) there is an initial transient facilitation which afterwards becomes masked by depression. Statistical analysis suggests that the changes in the average numbers of quanta released (m) could be attributed to the change in the immediately available store (n). 5. During long tetanic stimulation (4000 pulses, 10-100 Hz) statistical analysis suggests that the decrease in the average number of quanta released (m) could be attributed almost entirely to the decrease in the immediately available store (n). The probability of release (p) decreased only slightly. The extent of the post-tetanic potentiation indicates that it cannot be explained on the grounds of increased probability of release (p) only. There should be an increase as well in the immediately available store (n). 6. It is suggested that while depression is most likely caused by the depletion of the immediately available store due to insufficient replenishment, the facilitation is probably caused by the increase in the capacity of the immediately available store to contain transmitter.

Animals↗

On Wiener filtering and the physics behind statistical modeling.

The closed-form solution of the so-called statistical multivariate calibration model is given in terms of the pure component spectral signal, the spectral noise, and the signal and noise of the reference method. The "statistical" calibration model is shown to be as much grounded on the physics of the pure component spectra as any of the "physical" models. There are no fundamental differences between the two approaches since both are merely different attempts to realize the same basic idea, viz., the spectrometric Wiener filter. The concept of the application-specific signal-to-noise ratio (SNR) is introduced, which is a combination of the two SNRs from the reference and the spectral data. Both are defined and the central importance of the latter for the assessment and development of spectroscopic instruments and methods is explained. Other statistics like the correlation coefficient, prediction error, slope deficiency, etc., are functions of the SNR. Spurious correlations and other practically important issues are discussed in quantitative terms. Most important, it is shown how to use a priori information about the pure component spectra and the spectral noise in an optimal way, thereby making the distinction between statistical and physical calibrations obsolete and combining the best of both worlds. Companies and research groups can use this article to realize significant savings in cost and time for development efforts.

Algorithms↗

A statistical methodology for mammographic density detection.

A statistical methodology is presented based on a chi-square probability analysis that allows the automated discrimination of radiolucent tissue (fat) from radiographic densities (fibroglandular tissue) in digitized mammograms. The method is based on earlier work developed at this facility that shows mammograms may be considered as evolving from a linear filtering operation where a random input field is passed through a 1/f filtering process. The filtering process is reversible which allows the solution of the input field with knowledge obtained from the raw image (the output). The input field solution is analogous to a prewhitening technique or deconvolution. This field contains all the information of the raw image in a much simplified format that can be approximated and analyzed with parametric methods. In the work presented here evidence indicates that there are two random events occurring in the input field with differing variances: (1) one relating to fat tissue with the smaller variance, and (2) the second relating to all other tissue with the larger variance. A statistical comparison of the variances is made by scanning the image with a small search window. A relaxation method allows for making a reliable estimate of the smaller variance which is considered as the global reference. If a local variance deviates significantly from the reference variance, based on chi-square analysis, it is labeled as nonfat; otherwise it is labeled as fat. This statistical test procedure results in a region by region continuous labeling of fat and nonfat tissue across the image. In the work presented here, the emphasis is on the methodology development with supporting preliminary results that are very encouraging. It is widely accepted that mammographic density is a breast cancer risk factor. An important application of this work is to incorporate density-based risk analysis into the ongoing statistical-based detection work developed at this facility. Additional applications include risk analysis dependent on either percentages or total amounts of fat or dense tissue. This work may be considered as the initial step in introducing many of the known breast cancer risk factors into the actual image data analysis.

Breast↗

A statistical study of cochlear nerve discharge patterns in response to complex speech stimuli.

Cochlear nerve discharge patterns in response to the synthesized consonant-vowel stimulus /da/ were collected from a population of 223 auditory-nerve fibers from a single cat. For each nerve fiber, discharges were measured from multiple, independent stimulus presentations, with the means and variances of the post-stimulus time histograms and Fourier transforms of response generated from the ensemble of stimulus presentations. The statistics were not consistent with those predicted via an inhomogeneous Poisson counting process model. Specifically, the synchronized components as measured by the Fourier transforms of post-stimulus time histogram responses have variances that are as much as a factor of 3 times lower than the predicted by the Poisson model. To account for the non-Poisson nature of the statistics, the Markov process model of Siebert/Gaumond was adopted. Using the maximum-likelihood and minimum description length algorithms, introduced by Miller [J. Acoust. Soc. Am. 77, 1452-1464 (1985)] and Mark and Miller [J. Acoust. Soc. Am. 91, 989-1002 (1992)], estimates of the stimulus and recovery functions were computed for each nerve fiber. Then, Markov point processes were simulated with the stimulus and recovery functions generated from these nerve fibers. The statistics of the simulated Markov processes are shown to have almost identical first- and second-order statistics as those measured for the population of auditory-nerve fibers, and demonstrates the effectiveness of the Markov point process model in accounting for the correlation effects associated with the discharge history-dependent refractory properties of auditory nerve response.

Acoustic Stimulation↗

Statistical problems in ESP research.

In search of repeatable ESP experiments, modern investigators are using more complex targets, richer and freer responses, feedback, and more naturalistic conditions. This makes tractable statistical models less applicable. Moreover, controls often are so loose that no valid statistical analysis is possible. Some common problems are multiple end points, subject cheating, and unconscious sensory cueing. Unfortunately, such problems are hard to recognize from published records of the experiments in which they occur; rather, these problems are often uncovered by reports of independent skilled observers who were present during the experiment. This suggests that magicians and psychologists be regularly used as observers. New statistical ideas have been developed for some of the new experiments. For example, many modern ESP studies provide subjects with feedback--partial information about previous guesses--to reward the subjects for correct guesses in hope of inducing ESP learning. Some feedback experiments can be analyzed with the use of skill-scoring, a statistical procedure that depends on the information available and the way the guessing subject uses this information.

Feedback↗

Statistical methods in microbiology.

Statistical methodology is viewed by the average laboratory scientist, or physician, sometimes with fear and trepidation, occasionally with loathing, and seldom with fondness. Statistics may never be loved by the medical community, but it does not have to be hated by them. It is true that statistical science is sometimes highly mathematical, always philosophical, and occasionally obtuse, but for the majority of medical studies it can be made palatable. The goal of this article has been to outline a finite set of methods of analysis that investigators should choose based on the nature of the variable being studied and the design of the experiment. The reader is encouraged to seek the advice of a professional statistician when there is any doubt about the appropriate method of analysis. A statistician can also help the investigator with problems that have nothing to do with statistical tests, such as quality control, choice of response variable and comparison groups, randomization, and blinding of assessment of response variables.

Humans↗

Kappa statistics as indicators of quality assurance in histopathology and cytopathology.

Kappa statistics are widely-used to assess performance in quality assurance schemes. Low values, however, are difficult to interpret, especially when confidence intervals have not been calculated. A model of a dichotomous decision in pathology (benignancy or malignancy in fine needle aspirates of the breast) was used to calculate kappa statistics (with confidence limits) for increasing false positive rates. It was found that the level at which the upper 95% confidence interval for the kappa statistic fell below 1 was an insensitive method of detecting unsatisfactory performance as at that level the false positive rate was unacceptably high (> 1%) for all populations of specimens less than 800 in number. Either large populations of samples are required in quality assurance schemes which use kappa statistics (which may well be impractical) or other methods of assessing performance, possibly with weighted outcomes, are required.

Biopsy, Needle↗

Statistical issues in analysis of diagnostic imaging experiments with multiple observations per patient.

Many diagnostic imaging experiments are characterized by the presence of several observations for each patient studied. Evaluation of metastases with different imaging modalities in patients with cancer or examination of multiple artery segments in patients with heart abnormalities are some examples of such studies. Data obtained from multiple observations per patient are cluster correlated and should not be analyzed by using standard statistical methods because of correlations within a subject. In this article, positron emission tomographic studies are used as a framework to review statistical methods for the analysis of clustered data. Some simple statistical methods that account for correlation within a subject and that can be applied to conventional and well-known statistical methods, such as the chi(2) and t tests, are introduced. One of these methods is illustrated by using a brief analysis of data from a positron emission tomographic study, which demonstrates how resulting conclusions may be incorrect if appropriate techniques are not applied. Alternative methods that can handle multiple observations and dependency within a subject for diagnostic imaging studies are discussed.

Bone Neoplasms↗

A long closed state of the synaptosomal bursting potassium channel confers a statistical memory.

1. The statistical properties of the bursting potassium channel from fused Torpedo synaptosomes were studied by using patch-clamp recording and time series analysis. 2. Voltage steps produce channel openings; the number of channels opening fluctuates from trial to trial. The maximal current observed in each trial is strongly dependent on the previous history of the membrane patch. Trials with no activity are frequently clumped together and so are trials with intense activity. 3. Autocorrelation analysis reveals a strong interdependence of successive responses, which is voltage dependent. 4. We propose that the strong statistical interdependence of responses to successive depolarizing pulses (the statistical "memory") is a manifestation of a long lived closed state. We speculate that this statistical memory may be of significance in frequency modulation of transmitter release.

Animals↗

A review of caveats in statistical nuclear image analysis.

A large body of the published literature in nuclear image analysis do not evaluate their findings on an independent data set. Hence, if several features are evaluated on a limited data set over-optimistic results are easily achieved. In order to find features that separate different outcome classes of interest, statistical evaluation of the nuclear features must be performed. Furthermore, to classify an unknown sample using image analysis, a classification rule must be designed and evaluated. Unfortunately, statistical evaluation methods used in the literature of nuclear image analysis are often inappropriate. The present article discusses some of the difficulties in statistical evaluation of nuclear image analysis, and a study of cervical cancer is presented in order to illustrate the problems. In conclusion, some of the most severe errors in nuclear image analysis occur in analysis of a large feature set, including few patients, without confirming the results on an independent data set. To select features, Bonferroni correction for multiple test is recommended, together with a standard feature set selection method. Furthermore, we consider that the minimum requirement of performing statistical evaluation in nuclear image analysis is confirmation of the results on an independent data set. We suggest that a consensus of how to perform evaluation of diagnostic and prognostic features is necessary, in order to develop reliable tools for clinical use, based on nuclear image analysis.

Data Interpretation, Statistical↗

Primary unit for statistical analysis in morphometry: patient or cell?

In a series of 16 oxyphilic follicular neoplasms of the thyroid (8 adenomas and 8 carcinomas), three different approaches for the analysis of morphometric data were evaluated. It was shown that the statistical design of morphometric studies is by nature nested due to subsampling of cells within each patient. Therefore, the most appropriate analysis would be to account for this hierarchical structure. However, related statistical methods are not at present well established, especially as far as classification rules are concerned. Therefore, the nested design is converted into the simple factorial one by considering only one kind of statistical unit - either patients or cells. The results of the study presented indicate that ignoring the patient as unit of analysis leads to a substantial error in statistical output, regardless of the particular procedure applied. Moreover, the size of the error can be neither diminished nor controlled. Choosing patients as primary units assures accurate results and also has an advantage of gaining some additional information by calculating several distributional estimates in each patient. However, this approach often requires a reduction of dimensions and, furthermore, is not encouraged in certain fields of quantitative cytology. Advantages and disadvantages of all approaches have been summarized and practical recommendations for their use have been worked out.

Adenocarcinoma, Follicular↗

Tissue counter analysis of histologic sections of melanoma: influence of mask size and shape, feature selection, statistical methods and tissue preparation.

BACKGROUND: Tissue counter analysis is an image analysis tool designed for the detection of structures in complex images at the macroscopic or microscopic scale. As a basic principle, small square or circular measuring masks are randomly placed across the image and image analysis parameters are obtained for each mask. Based on learning sets, statistical classification procedures are generated which facilitate an automated classification of new data sets. OBJECTIVE: To evaluate the influence of the size and shape of the measuring masks as well as the importance of feature selection, statistical procedures and technical preparation of slides on the performance of tissue counter analysis in microscopic images. As main quality measure of the final classification procedure, the percentage of elements that were correctly classified was used. STUDY DESIGN: HE-stained slides of 25 primary cutaneous melanomas were evaluated by tissue counter analysis for the recognition of melanoma elements (section area occupied by tumour cells) in contrast to other tissue elements and background elements. Circular and square measuring masks, various subsets of image analysis features and classification and regression trees compared with linear discriminant analysis as statistical alternatives were used. The percentage of elements that were correctly classified by the various classification procedures was assessed. In order to evaluate the applicability to slides obtained from different laboratories, the best procedure was automatically applied in a test set of another 50 cases of primary melanoma derived from the same laboratory as the learning set and two test sets of 20 cases each derived from two different laboratories, and the measurements of melanoma area in these cases were compared with conventional assessment of vertical tumour thickness. RESULTS: Square measuring masks were slightly superior to circular masks, and larger masks (64 or 128 pixels in diameter) were superior to smaller masks (8 to 32 pixels in diameter). As far as the subsets of image analysis features were concerned, colour features were superior to densitometric and Haralick texture features. Statistical moments of the grey level distribution were of least significance. CART (classification and regression tree) analysis turned out to be superior to linear discriminant analysis. In the best setting, 95% of melanoma tissue elements were correctly recognized. Automated measurement of melanoma area in the independent test sets yielded a correlation of r=0.846 with vertical tumour thickness (p<0.001), similar to the relationship reported for manual measurements. The test sets obtained from different laboratories yielded comparable results. CONCLUSIONS: Large, square measuring masks, colour features and CART analysis provide a useful setting for the automated measurement of melanoma tissue in tissue counter analysis, which can also be used for slides derived from different laboratories.

Adult↗

[Use of DNA chips (microarrays) in medicine: technical foundations and basic procedures for statistical analysis of results].

DNA microarray technology allows the assessment of genetic analyses on thousands of genes simultaneously. The statistical analyses of these experiments are challenging since a high number of multiple hypotheses are tested and classical statistical methods need to adapt to this situation. Furthermore, the great variability observed in the experiments and their high cost of them needs a careful design. In this review we will explain what is a cDNA microarray, how it works and its potential uses. Later we will deal with statistical issues of design and analysis, from the image processing and data quality control, to the statistical test of hypothesis to detect interesting genes. Finally we will comment on multivariate methods to detect patterns in gene expression.

DNA↗

Statistical properties of Teng and Risch's sibship type tests for detecting an association between disease and a candidate allele.

Risch and Teng [Genome Res 1998;8:1273-1288] and Teng and Risch [Genome Res 1999;9:234-241] proposed a class of transmission/disequilibrium test-like statistical tests based on the difference between the estimated allele frequencies in the affected and control populations. They evaluated the power of a variety of family-based and nonfamily-based designs for detecting an association between a candidate allele and disease. Because they were concerned with diseases with low penetrances, their power calculations assumed that unaffected individuals can be treated as a random sample from the population. They predicted that this assumption rendered their sample size calculations slightly conservative. We generalize their partial ascertainment conditioning by including the status of the unaffected sibs in the calculations of the distribution and power of the statistic used to compare the allele frequency in affected offspring to the estimated frequency in the parents, based on sibships with genotyped affected and unaffected sibs. Sample size formulas for our full ascertainment methods are presented. The sample sizes for our procedure are compared to those of Teng and Risch. The numerical results and simulations indicate that the simplifying assumption used in Teng and Risch can produce both conservative and anticonservative results. The magnitude of the difference between the sample sizes needed by their partial ascertainment approximation and the full ascertainment is small in the circumstances they focused on but can be appreciable in others, especially when the baseline penetrances are moderate. Two other statistics, using different estimators for the variance of the basic statistic comparing the allele frequencies in the affected and unaffected sibs are introduced. One of them incorporates an estimate of the null variance obtained from an auxiliary sample and appears to noticeably decrease the sample sizes required to achieve a prespecified power.

Case-Control Studies↗

Influence of number of surfaces and observers on statistical power in a multiobserver ROC radiographic caries detection study.

The aim of the study was to evaluate the influence of the number of surfaces (N(SURF)) and the number of observers (N(OBS)) on the statistical power of a study comparing the diagnostic accuracies of radiographic systems used for approximal caries lesion detection. A data set consisting of 338 surfaces examined by 10 independent observers using four radiographic systems was available. The presence of a caries lesion was assessed from a 5-point confidence scale. The true lesion diagnosis was established by histological validation. ROC curve areas (A(z)s) were used to express the diagnostic accuracy of the observers with the radiographic systems. Assuming that the A(z)s were tested by a two-way analysis of variance, we performed a simulation study in order to evaluate how the power of this statistical analysis depended on N(SURF) and N(OBS). As a measure of the statistical power we used the standard error of the difference between the expected A(z)s of two systems. The simulations were made with N(SURF) in the range from 25 to 338 and N(OBS) from 2 to 10. The simulations showed that the power increased as a function of the total number of evaluations per system (N(SURF) x N(OBS)), but how this number was attained in relation to the number of surfaces and observers had only marginal influence on the power. Thus, from a statistical point of view it may be concluded, provided that data are analyzed by a two-way analysis of variance, that study designs for comparing the accuracy of several systems can be composed freely in relation to the number of surfaces and observers as long as the total number of evaluations per system are identical.

Analysis of Variance↗

New statistical software for intralaboratory and interlaboratory quality control in clinical cytology. Validation in a simulation study on clinical samples.

OBJECTIVE: To design a statistical software package to provide automated calculations of normal and weighted and 3 indices. STUDY DESIGN: Prompted by the lack of commonly available software to compute weighted kappa and the nonproportionate workload needed to calculate our 3 variability indices manually, the new statistical software package was designed. To demonstrate the performance of the new CONQUISTADOR software, a simulation study (both intralaboratory and interlaboratory) was designed using 5,000 clinical samples randomly selected from a data file of > or = 200,000 conventional Pap smears and programmed to become "analyzed" by 12 cytologists in 5 imaginary laboratories. RESULTS: A representative set of both complete and partial outputs provided by the software, in Excel format (Microsoft, Redmond, Washington, U.S.A.) are shown to illustrate the different functions of the program. In the interlaboratory mode, the software calculates accuracy indicators (sensitivity, specificity, positive and negative predictive value, and their 95% CI), which are not common features of regular statistical packages; kappa and weighted kappa; and their 95% CI (comparison of single laboratories to all laboratories and pairwise comparisons between single laboratories). The 3 diagnostic variability indices can be computed separately for all samples or for only the positive samples. In the intralaboratory mode, the software calculates the same indices for individual cytologists. CONCLUSION: The CONQUISTADOR statistical package has properties that are useful in monitoring cytologic laboratory quality in both intralaboratory and interlaboratory settings. The software will be distributed by the National Institute of Health, Rome, for the delivery costs only.

Clinical Laboratory Techniques↗

Statistical characterization of real-world illumination.

Although studies of vision and graphics often assume simple illumination models, real-world illumination is highly complex, with reflected light incident on a surface from almost every direction. One can capture the illumination from every direction at one point photographically using a spherical illumination map. This work illustrates, through analysis of photographically acquired, high dynamic range illumination maps, that real-world illumination possesses a high degree of statistical regularity. The marginal and joint wavelet coefficient distributions and harmonic spectra of illumination maps resemble those documented in the natural image statistics literature. However, illumination maps differ from typical photographs in that illumination maps are statistically nonstationary and may contain localized light sources that dominate their power spectra. Our work provides a foundation for statistical models of real-world illumination, thereby facilitating the understanding of human material perception, the design of robust computer vision systems, and the rendering of realistic computer graphics imagery.

Contrast Sensitivity↗