Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Genome comparison using Gene Ontology (GO) with statistical testing.

BACKGROUND: Automated comparison of complete sets of genes encoded in two genomes can provide insight on the genetic basis of differences in biological traits between species. Gene ontology (GO) is used as a common vocabulary to annotate genes for comparison. Current approaches calculate the fold of unweighted or weighted differences between two species at the high-level GO functional categories. However, to ensure the reliability of the differences detected, it is important to evaluate their statistical significance. It is also useful to search for differences at all levels of GO. RESULTS: We propose a statistical approach to find reliable differences between the complete sets of genes encoded in two genomes at all levels of GO. The genes are first assigned GO terms from BLAST searches against genes with known GO assignments, and for each GO term the abundance of genes in the two genomes is compared using a chi-squared test followed by false discovery rate (FDR) correction. We applied this method to find statistically significant differences between two cyanobacteria, Synechocystis sp. PCC6803 and Anabaena sp. PCC7120. We then studied how the set of identified differences vary when different BLAST cutoffs are used. We also studied how the results vary when only subsets of the genes were used in the comparison of human vs. mouse and that of Saccharomyces cerevisiae vs. Schizosaccharomyces pombe. CONCLUSION: There is a surprising lack of statistical approaches for comparing complete genomes at all levels of GO. With the rapid increase of the number of sequenced genomes, we hope that the approach we proposed and tested can make valuable contribution to comparative genomics.

Anabaena↗

Statistical inference of chromosomal homology based on gene colinearity and applications to Arabidopsis and rice.

BACKGROUND: The identification of chromosomal homology will shed light on such mysteries of genome evolution as DNA duplication, rearrangement and loss. Several approaches have been developed to detect chromosomal homology based on gene synteny or colinearity. However, the previously reported implementations lack statistical inferences which are essential to reveal actual homologies. RESULTS: In this study, we present a statistical approach to detect homologous chromosomal segments based on gene colinearity. We implement this approach in a software package ColinearScan to detect putative colinear regions using a dynamic programming algorithm. Statistical models are proposed to estimate proper parameter values and evaluate the significance of putative homologous regions. Statistical inference, high computational efficiency and flexibility of input data type are three key features of our approach. CONCLUSION: We apply ColinearScan to the Arabidopsis and rice genomes to detect duplicated regions within each species and homologous fragments between these two species. We find many more homologous chromosomal segments in the rice genome than previously reported. We also find many small colinear segments between rice and Arabidopsis genomes.

Algorithms↗

The case for well-conducted experiments to validate statistical protocols for 2D gels: different pre-processing = different lists of significant proteins.

BACKGROUND: The proteomics literature has seen a proliferation of publications that seek to apply the rapidly improving technology of 2D gels to study various biological systems. However, there is a dearth of systematic studies that have investigated appropriate statistical approaches to analyse the data from these experiments. RESULTS: Comparison of the effects of statistical pre-processing on the results of two sample t-tests suggests that the results of 2D gel experiments and by extension the conclusions derived from these experiments are not independent of the statistical protocol used. CONCLUSIONS: This study suggests that there is a need for well-conducted validation studies to establish optimal statistical techniques to be used on such data sets.

Animals↗

Effect sizes and statistical testing in the determination of clinical significance in behavioral medicine research.

BACKGROUND: The interpretation of clinical significance continues to be an obstacle for researchers in behavioral medicine. PURPOSE: To review selected behavioral medicine research to critically examine the perception among investigators that behavioral effects on health are small based on common metrics of clinical significance. METHODS: Using quantitative findings from recent behavioral medicine research in medical and psychiatric journals, we explored results in terms of several statistical metrics to assess potential clinical significance: r coefficients, risk ratios, risk difference measures, and attributable risk. RESULTS: Translated into r coefficients, even established health predictors such as smoking, obesity, and fitness had only modest effects (rs =.03-.22), and the range of effect sizes were comparable with those based on psychological predictors including depression and stress-reactivity (rs =.06-.22). In contrast, effects for both classes of predictors were suggestive of clinical significance based on public health statistics. CONCLUSIONS: Our choice of statistics for defining "small" and "large" effect sizes affects the perceived importance of behavioral health findings. In the assessment of health outcomes with low incidence rates, effects expressed as correlations using even the most robust predictors will often appear small. In these instances, we challenge researchers to move beyond conventional data analysis approaches and to expand their clinical interpretation efforts by employing additional statistical methods favored in medicine and public health.

Behavioral Medicine↗

Some statistical implications of dose uncertainty in radiation dose-response analyses.

Statistical dose-response analyses in radiation epidemiology can produce misleading results if they fail to account for radiation dose uncertainties. While dosimetries may differ substantially depending on the ways in which the subjects were exposed, the statistical problems typically involve a predominantly linear dose-response curve, multiple sources of uncertainty, and uncertainty magnitudes that are best characterized as proportional rather than additive. We discuss some basic statistical issues in this setting, including the bias and shape distortion induced by classical and Berkson uncertainties, the effect of uncertain dose-prediction model parameters on estimated dose-response curves, and some notes on statistical methods for dose-response estimation in the presence of radiation dose uncertainties.

Artifacts↗

Statistical analytical methods for comparing the incidence of tumors to the historical control data.

Statistical analysis for comparing the incidence of tumors in treated groups to historical control data is not generally performed. In the present study, a number of data exhibiting the lowest, low, moderate and highest incidence of tumors from long-term rodent bioassay studies has been compared with the historical control data. In studies exhibiting the lowest incidence (less than a few percent) of tumors, the Kastenbaum and Bowman test was found to be relevant since it takes into account the sample size of both the historical control data base and each treated group in the study. In studies where a wider range of tumor incidence was exhibited, a statistical method which employs a rejection limits based on the range of incidence in the historical data is recommended. When malignant tumors are evident in treated groups, no matter how low the incidence, they should be analyzed statistically and compared with the incidence in historical control data as well as those in the concurrent control group. Statistical analytical comparisons of study results to historical control data may contribute to more meaningful evaluations in carcinogenicity studies by eliminating possible false positive or false negative results.

Animals↗

Attentional spread in the statistical processing of visual displays.

We tested the hypothesis that distributing attention over an array of similar items makes its statistical properties automatically available. We found that extracting the mean size of sets of circles was easier to combine with tasks requiring distributed or global attention than with tasks requiring focused attention. One explanation may be that extracting the statistical descriptors requires parallel access to all the information in the array. Consistent with this claim, we found an advantage for simultaneous over successive presentation when the total time available was matched. However, the advantage was small; parallel access facilitates statistical processing without being essential. Evidence that statistical processing is automatic when attention is distributed over a display came from the finding that there was no decrement in accuracy relative to single-task performance when mean judgments were made concurrently with another task that required distributed or global attention.

Attention↗

[Weights in Horvitz-Thompson statistic for complex samples].

The problem of using ordinary statistical methods basically for simple random samples in analyzing data from complex surveys for non-simple random samples was discussed. Horvitz-Thompson statistic and the poststratification weights were introduced. Illustrative examples were given to show the validity and superiority of Horvitz-Thompson statistic in comparison with statistic used for simple random samples.

Bias↗

Needs in vital and health statistics in underdeveloped countries.

The author discusses the types of health and vital statistics that would be of the greatest practical value to countries with only slightly developed public-health and vital registration systems and the ways by which these statistics may be obtained.While a regular census is necessary for proper mortality and natality statistics, population estimates may be successfully used until a census can be taken. Natality statistics should include live-births, stillbirths, legitimacy, and age of mother. For morbidity measurement, four sources of information or types of inquiry can be used before complete registration systems are available: sickness surveys by home visits of families; records of notifiable communicable diseases; medical records of sickness in schools; and records of health welfare centres and health visitors, when these exist. The use of infant mortality figures in underdeveloped countries is subject to considerable error, and great effort will be needed to get every living child in the birth register. A useful local index of health is the recording in selected areas of deaths during the first three years of life. Death-rates at higher ages can only be assessed where death registration is fairly complete.It is suggested that reliable information on population changes, child mortality and sickness, and the incidence of disease can be obtained by a continuous study programme in carefully selected model survey and registration districts. Apart from the immediate results from such a programme, it would prepare the ground for the subsequent establishment of a full vital registration system.

Censuses↗

Multiplicity-adjusted sample size requirements: a strategy to maintain statistical power with Bonferroni adjustments.

BACKGROUND: A researcher must carefully balance the risk of 2 undesirable outcomes when designing a clinical trial: false-positive results (type I error) and false-negative results (type II error). In planning the study, careful attention is routinely paid to statistical power (i.e., the complement of type II error) and corresponding sample size requirements. However, Bonferroni-type alpha adjustments to protect against type I error for multiple tests are often resisted. Here, a simple strategy is described that adjusts alpha for multiple primary efficacy measures, yet maintains statistical power for each test. METHOD: To illustrate the approach, multiplicity-adjusted sample size requirements were estimated for effects of various magnitude with statistical power analyses for 2-tailed comparisons of 2 groups using chi2 tests and t tests. These analyses estimated the required sample size for hypothetical clinical trial protocols in which the prespecified number of primary efficacy measures ranged from 1 to 5. Corresponding Bonferroni-adjusted alpha levels were used for these calculations. RESULTS: Relative to that required for 1 test, the sample size increased by about 20% for 2 dependent variables and 30% for 3 dependent variables. CONCLUSION: The strategy described adjusts alpha for multiple primary efficacy measures and, in turn, modifies the sample size to maintain statistical power. Although the strategy is not novel, it is typically overlooked in psychopharmacology trials. The number of primary efficacy measures must be prespecified and carefully limited when a clinical trial protocol is prepared. If multiple tests are designated in the protocol, the alpha-level adjustment should be anticipated and incorporated in sample size calculations.

Clinical Protocols↗

Reengineering vital registration and statistics systems for the United States.

For more than a hundred years, the United States has operated a decentralized vital statistics system as an essential component of public health. Statistics based on births and deaths registered in the United States are a primary source of data used to track health status, to plan, implement, and evaluate health and social services, and to set health policy. The national vital statistics system provides nearly complete, continuous, and comparable federal, state, and local data. The system, however, is based on outmoded vital registration practices and structures, which raises concerns about data quality, timeliness, and the lack of real-time linkage capabilities. While many organizations are working together to address these issues and have made notable achievements, questions remain to be answered. Efforts to rejuvenate the nation's vital statistics system will need to expand dramatically to provide public health with a timely, high-quality, and flexible system to monitor vital health outcomes at the local, state, and national levels.

Birth Certificates↗

[Segmenting lung fields in serial chest radiographs using both population and patient-specific shape statistics].

This paper presents a new deformable model using both population-based and patient-specific shape statistics to segment lung fields from serial chest radiographs. First, a modified scale-invariant feature transform (SIFT) local descriptor is used to characterize the image features in the vicinity of each pixel, so that the deformable model deforms in a way that seeks for the region with similar SIFT local descriptors; second, the deformable model is constrained by both population-based and patient-specific shape statistics. At first, population-based shape statistics plays an leading role when the number of serial images is small, and gradually, patient-specific shape statistics plays a more and more important role after a sufficient number of segmentation results on the same patient have been obtained. The proposed deformable model can adapt to the shape variability of different patients, and obtain more robust and accurate segmentation results.

Algorithms↗

Usage of statistics in the surgical literature and the 'orphan P' phenomenon.

The statistics used in 240 surgical publications were reviewed. Basic parametric statistics were used in 60% of the publications; 21% of publications failed to document a measure of central tendency, 11% of publications contained an undefined '+/-' notation, and 10% of publications did not state the type of evaluative statistic that was used to calculate a P value, that is, an 'orphan P'. These results indicate the need for wider education about the use of descriptive and basic parametric statistics. It is impossible to evaluate the surgical literature critically without these skills.

Australia↗

Rating scales in psychopharmacology. Statistical aspects.

The scaling problem of measurement of clinical data in psychiatry should be differed from, although connected with, the statistical problem of statistical significance testing in psychopharmacological trials. The scaling levels (ordinal versus interval) of a rating scale as well as the statistical tests (non-parametric versus parametric) are different in Phase II and Phase III trials. Sufficient statistics of improvement curves and of recovery rates are discussed.

Clinical Trials as Topic↗

[Assessment of the use of statistical methods in health research].

BACKGROUND: Computer use has lead to a great development of statistical methods. AIM: To assess the use of statistical methods in Chilean medical literature. METHODS: Two hundred sixty four papers appeared in Revista Médica de Chile and Revista Chilena de Pediatría between 1983 and 1993 were reviewed. RESULTS: Student's "t", Fisher's, and chi 2 test are the most frequently used statistical methods in 67% of papers. Correlation coefficients are used in 10% of papers. Multivariate methods are seldom used. CONCLUSIONS: Statistical analysis of papers published in Chilean medical journals is restricted to very few methods.

Chile↗

Statistical significance and clinical relevance: the importance of power in clinical trials in dermatology.

When evaluating the validity of a study, the reader must consider both the clinical and statistical significance of the findings. A study that claims clinical relevance may lack sufficient statistical significance to make a meaningful statement. Conversely, a study that shows a statistically significant difference in 2 treatment options may lack practicality. The concept of power of a clinical trial refers to the probability of detecting a difference between study groups when a true difference exists. We will discuss statistical power by examining studies too small to identify important differences, studies so large as to identify differences that are not clinically significant, difficult-to-design studies without very large patient populations, and those studies with both adequate power and clinically relevant findings. Dermatologists should not focus on small P values alone to decide whether a treatment is clinically useful; it is essential to consider the magnitude of treatment differences and the power of the study.

Clinical Trials as Topic↗

The reporting of statistical techniques in otolaryngology journals.

The domain of this study is the reporting of statistical analyses in the otolaryngology literature during 1983 and 1984. Fewer than ten basic statistical procedures accounted for more than 90% of the statistical techniques reported. Implications for authors, journals, and educators are discussed. We offer suggestions for imparting statistical skills that may be helpful in curriculum design, residency training, and continuing medical education planning.

Humans↗

Disease cluster statistics for imprecise space-time locations.

Health professionals are investigating an increasing number of possible disease clusters, and statistical tests play an important role in cluster description and analysis. Existing cluster statistics assume precise data, when in reality health events are often imprecise (for example, place of residence is known only to the census district or zip code) and uncertain (for example, 'I first became ill sometime in 1985'). This incompatibility--precise methods used to analyse imprecise data--is largely ignored, resulting in test statistics of unknown accuracy. Most cluster statistics can be written as the cross-product of two matrices where one matrix reflects nearest-neighbour, distance or adjacency relationships and the second matrix is health related (for example, case-control identities). This paper explores a general approach to clustering, which incorporates uncertainty regarding space-time locations into these nearest neighbour, distance or adjacency relationships. Because the approach is general it can be used with almost all existing cluster tests, and, because it accounts for imprecise location data, it is suited to the 'real-world' nature of disease cluster investigations.

Bias↗