Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Statistical properties of population differentiation estimators under stepwise mutation in a finite island model.

Microsatellite loci mutate at an extremely high rate and are generally thought to evolve through a stepwise mutation model. Several differentiation statistics taking into account the particular mutation scheme of the microsatellite have been proposed. The most commonly used is R(ST) which is independent of the mutation rate under a generalized stepwise mutation model. F(ST) and R(ST) are commonly reported in the literature, but often differ widely. Here we compare their statistical performances using individual-based simulations of a finite island model. The simulations were run under different levels of gene flow, mutation rates, population number and sizes. In addition to the per locus statistical properties, we compare two ways of combining R(ST) over loci. Our simulations show that even under a strict stepwise mutation model, no statistic is best overall. All estimators suffer to different extents from large bias and variance. While R(ST) better reflects population differentiation in populations characterized by very low gene-exchange, F(ST) gives better estimates in cases of high levels of gene flow. The number of loci sampled (12, 24, or 96) has only a minor effect on the relative performance of the estimators under study. For all estimators there is a striking effect of the number of samples, with the differentiation estimates showing very odd distributions for two samples.

Computer Simulation↗

The burgeoning field of statistical phylogeography.

In the newly emerging field of statistical phylogeography, consideration of the stochastic nature of genetic processes and explicit reference to theoretical expectations under various models has dramatically transformed how historical processes are studied. Rather than being restricted to ad hoc explanations for observed patterns of genetic variation, assessments about the underlying evolutionary processes are now based on statistical tests of various hypotheses, as well as estimates of the parameters specified by the models. A wide range of demographical and biogeographical processes can be accommodated by these new analytical approaches, providing biologically more realistic models. Because of these advances, statistical phylogeography can provide unprecedented insights about a species' history, including decisive information about the factors that shape patterns of genetic variation, species distributions, and speciation. However, to improve our understanding of such processes, a critical examination and appreciation of the inherent difficulties of historical inference and challenges specific to testing phylogeographical hypotheses are essential. As the field of statistical phylogeography continues to take shape many difficulties have been resolved. Nonetheless, careful attention to the complexities of testing historical hypotheses and further theoretical developments are essential to improving the accuracy of our conclusions about a species' history.

Demography↗

Statistical analysis of denaturing gel electrophoresis (DGE) fingerprinting patterns.

Technical developments in molecular biology have found extensive applications in the field of microbial ecology. Among these techniques, fingerprinting methods such as denaturing gel electrophoresis (DGE, including the three options: DGGE, TGGE and TTGE) has been applied to environmental samples over this last decade. Microbial ecologists took advantage of this technique, originally developed for the detection of single mutations, for the analysis of whole bacterial communities. However, until recently, the results of these high quality fingerprinting patterns were restricted to a visual interpretation, neglecting the analytical potential of the method in terms of statistical significance and ecological interpretation. A brief recall is presented here about the principles and limitations of DGE fingerprinting analysis, with an emphasis on the need of standardization of the whole analytical process. The main content focuses on statistical strategies for analysing the gel patterns, from single band examination to the analysis of whole fingerprinting profiles. Applying statistical method make the DGE fingerprinting technique a promising tool. Numerous samples can be analysed simultaneously, permitting the monitoring of microbial communities or simply bacterial groups for which occurrence and relative frequency are affected by any environmental parameter. As previously applied in the fields of plant and animal ecology, the use of statistics provides a significant advantage for the non-ambiguous interpretation of the spatial and temporal functioning of microbial communities.

Bacteria↗

Statistical inference in facial plastic surgery: perspectives and alternatives.

Facial plastic surgeons often must make decisions with imperfect information. Statistical inference is fundamentally the practice of using data to draw conclusions about uncertain phenomena. It is important, therefore, that facial plastic surgeons engaged both in clinical practice and in research have an understanding of statistical concepts to conduct research with results that are meaningful, to assess the validity of published research, and to adopt the most effective techniques and treatments. The purpose of this article is to provide an overview of classical statistical methods that are encountered frequently in facial plastic surgery research, discuss issues of interpretation of results, and introduce an alternative paradigm for conducting statistical inference.

Algorithms↗

Statistical process control geometric Q-chart for nosocomial infection surveillance.

Several authors have proposed the use of statistical process control charting methods for the surveillance of endemic rates of nosocomial infections. The principal goal of such a charting program is to recognize any increase of the endemic rate to an epidemic rate as soon as possible after the change occurs. However, many of the statistical process control charting methods that have been proposed are based on classical charting principles that are effective largely for processes for which sufficient historical data are available. These methods require that a fairly large data set, taken while the infection rate was stable at a low endemic value, must be available to begin the charting process. These data are used both to confirm the appropriateness of the probability distribution and to make a control chart for the infection process based on the distribution. However, such data sets are often not available. The purpose of this article is to inform and demonstrate to readers that recent research in statistics has developed modern statistical process control methods that can be used effectively with or without such prior data. These methods make possible much more effective nosocomial infection surveillance programs that will give timely warnings of the onsets of epidemics or evidence of the effectiveness of infection control initiatives. These warnings will permit earlier correction initiatives and thus avoid much liability.

Clostridioides difficile↗

Statistical modeling of large microarray data sets to identify stimulus-response profiles.

A statistical modeling approach is proposed for use in searching large microarray data sets for genes that have a transcriptional response to a stimulus. The approach is unrestricted with respect to the timing, magnitude or duration of the response, or the overall abundance of the transcript. The statistical model makes an accommodation for systematic heterogeneity in expression levels. Corresponding data analyses provide gene-specific information, and the approach provides a means for evaluating the statistical significance of such information. To illustrate this strategy we have derived a model to depict the profile expected for a periodically transcribed gene and used it to look for budding yeast transcripts that adhere to this profile. Using objective criteria, this method identifies 81% of the known periodic transcripts and 1,088 genes, which show significant periodicity in at least one of the three data sets analyzed. However, only one-quarter of these genes show significant oscillations in at least two data sets and can be classified as periodic with high confidence. The method provides estimates of the mean activation and deactivation times, induced and basal expression levels, and statistical measures of the precision of these estimates for each periodic transcript.

CDC28 Protein Kinase, S cerevisiae↗

Statistical analysis of sparse infection data and its implications for retroviral treatment trials in primates.

Reports on retroviral primate trials rarely publish any statistical analysis. Present statistical methodology lacks appropriate tests for these trials and effectively discourages quantitative assessment. This paper describes the theory behind VACMAN, a user-friendly computer program that calculates statistics for in vitro and in vivo infectivity data. VACMAN's analysis applies to many retroviral trials using i.v. challenges and is valid whenever the viral dose-response curve has a particular shape. Statistics from actual i.v. retroviral trials illustrate some unappreciated principles of effective animal use: dilutions other than 1:10 can improve titration accuracy; infecting titration animals at the lowest doses possible can lower challenge doses; and finally, challenging test animals in small trials with more virus than controls safeguards against false successes, "reuses" animals, and strengthens experimental conclusions. The theory presented also explains the important concept of viral saturation, a phenomenon that may cause in vitro and in vivo titrations to agree for some retroviral strains and disagree for others.

Animals↗

Measurement of genetic structure within populations using Moran's spatial autocorrelation statistics.

Spatial structure of genetic variation within populations, an important interacting influence on evolutionary and ecological processes, can be analyzed in detail by using spatial autocorrelation statistics. This paper characterizes the statistical properties of spatial autocorrelation statistics in this context and develops estimators of gene dispersal based on data on standing patterns of genetic variation. Large numbers of Monte Carlo simulations and a wide variety of sampling strategies are utilized. The results show that spatial autocorrelation statistics are highly predictable and informative. Thus, strong hypothesis tests for neutral theory can be formulated. Most strikingly, robust estimators of gene dispersal can be obtained with practical sample sizes. Details about optimal sampling strategies are also described.

Biological Evolution↗

A unified statistical framework for sequence comparison and structure comparison.

We present an approach for assessing the significance of sequence and structure comparisons by using nearly identical statistical formalisms for both sequence and structure. Doing so involves an all-vs.-all comparison of protein domains [taken here from the Structural Classification of Proteins (scop) database] and then fitting a simple distribution function to the observed scores. By using this distribution, we can attach a statistical significance to each comparison score in the form of a P value, the probability that a better score would occur by chance. As expected, we find that the scores for sequence matching follow an extreme-value distribution. The agreement, moreover, between the P values that we derive from this distribution and those reported by standard programs (e.g., BLAST and FASTA validates our approach. Structure comparison scores also follow an extreme-value distribution when the statistics are expressed in terms of a structural alignment score (essentially the sum of reciprocated distances between aligned atoms minus gap penalties). We find that the traditional metric of structural similarity, the rms deviation in atom positions after fitting aligned atoms, follows a different distribution of scores and does not perform as well as the structural alignment score. Comparison of the sequence and structure statistics for pairs of proteins known to be related distantly shows that structural comparison is able to detect approximately twice as many distant relationships as sequence comparison at the same error rate. The comparison also indicates that there are very few pairs with significant similarity in terms of sequence but not structure whereas many pairs have significant similarity in terms of structure but not sequence.

Animals↗

A statistical human resources costing and accounting model for analysing the economic effects of an intervention at a workplace.

The study had two primary aims. The first aim was to combine a human resources costing and accounting approach (HRCA) with a quantitative statistical approach in order to get an integrated model. The second aim was to apply this integrated model in a quasi-experimental study in order to investigate whether preventive intervention affected sickness absence costs at the company level. The intervention studied contained occupational organizational measures, competence development, physical and psychosocial working environmental measures and individual and rehabilitation measures on both an individual and a group basis. The study is a quasi-experimental design with a non-randomized control group. Both groups involved cleaning jobs at predominantly female workplaces. The study plan involved carrying out before and after studies on both groups. The study included only those who were at the same workplace during the whole of the study period. In the HRCA model used here, the cost of sickness absence is the net difference between the costs, in the form of the value of the loss of production and the administrative cost, and the benefits in the form of lower labour costs. According to the HRCA model, the intervention used counteracted a rise in sickness absence costs at the company level, giving an average net effect of 266.5 Euros per person (full-time working) during an 8-month period. Using an analogue statistical analysis on the whole of the material, the contribution of the intervention counteracted a rise in sickness absence costs at the company level giving an average net effect of 283.2 Euros. Using a statistical method it was possible to study the regression coefficients in sub-groups and calculate the p-values for these coefficients; in the younger group the intervention gave a calculated net contribution of 605.6 Euros with a p-value of 0.073, while the intervention net contribution in the older group had a very high p-value. Using the statistical model it was also possible to study contributions of other variables and interactions. This study established that the HRCA model and the integrated model produced approximately the same monetary outcomes. The integrated model, however, allowed a deeper understanding of the various possible relationships and quantified the results with confidence intervals.

Accounting↗

Making bootstrap statistical inferences: a tutorial.

Bootstrapping is a computer-intensive statistical technique in which extensive computational procedures are heavily dependent on modern high-speed digital computers. The payoff for such intensive computations is freedom from two major limiting factors that have dominated classical statistical theory since its beginning: the assumption that the data conform to a bell-shaped curve, and the need to focus on statistical measures whose theoretical properties can be analysed mathematically. The name "bootstrap" was derived from an old saying about pulling oneself up by one's bootstraps. In this case, bootstrapping means redrawing samples randomly from the original sample with replacement. The key idea, computations, advantages, limitations, and application potential of bootstrapping in the field of physical education and exercise science are introduced and illustrated using a set of national physical fitness testing data. Finally, an example of a bootstrapping application is provided. Through a step-by-step approach, the development and implementation of the bootstrap statistical inference are illustrated.

Data Interpretation, Statistical↗

Epidemiology and statistical methods in prediction of patient outcome.

Substantial gaps exist in the data of the assessment of risk and prognosis that limit our understanding of the complex mechanisms that contribute to the greatest cancer epidemic, prostate cancer, of our time. This report was prepared by an international multidisciplinary committee of the World Health Organization to address contemporary issues of epidemiology and statistical methods in prostate cancer, including a summary of current risk assessment methods and prognostic factors. Emphasis was placed on the relative merits of each of the statistical methods available. We concluded that: 1. An international committee should be created to guide the assessment and validation of molecular biomarkers. The goal is to achieve more precise identification of those who would benefit from treatment. 2. Prostate cancer is a predictable disease despite its biologic heterogeneity. However, the accuracy of predicting it must be improved. We expect that more precise statistical methods will supplant the current staging system. The simplicity and intuitive ease of using the current staging system must be balanced against the serious compromise in accuracy for the individual patient. 3. The most useful new statistical approaches will integrate molecular biomarkers with existing prognostic factors to predict conditional life expectancy (i.e. the expected remaining years of a patient's life) and take into account all-cause mortality.

Adult↗

Simultaneous use of weighted logrank and standardized Kaplan-Meier statistics.

Rank-based test procedures in censored survival data differ considerably in their sensitivity to various alternatives to the hypothesis of equality of underlying distributions. Procedures based on the simultaneous use of multiple statistics provide an appealing approach to obtaining more global sensitivity while maintaining the sensitivity to alternatives of interest. In this article, the joint distribution of weighted logrank statistics (to assess overall differences in time-to-event distributions) and a standardized difference in Kaplan-Meier estimates (to assess differences at a specific prespecified time post randomization) is obtained and this result is used to formulate statistical test procedures based on the simultaneous use of these two types of statistics. Simulations are used to assess small sample properties of this approach and its usefulness is illustrated in important recent oncology clinical trials.

Clinical Trials, Phase III as Topic↗

Statistical issues of quality control in organised breast cancer screening.

BACKGROUND: European guidelines for breast-cancer screening recommend an integrated approach of mammography screening with subsequent assessment and biopsy, if required, in one screening unit under permanent quality control, for which target values are released. Although the calculation of the respective rates (e.g. for participation, assessment, biopsy, or cancer detection) appears trivial, the statistical assessment of their compatibility with the target values is less obvious. This is especially true if subjects with a positive diagnostic result leave the screening-assessment chain prematurely, and information about further diagnostic results outside the organised screening is lacking. METHOD: Statistical models for the basic situation, in which complete information about the screening and assessment outcome is available, as well as for when information is incomplete, are presented. The statistical methods for obtaining the confidence limits, statistical tests and sample sizes needed to obtain a desired power of tests for the process parameters of interest are also given. RESULTS: The sample-size calculations indicate that large numbers of enrolled subjects are required to obtain reasonably narrow confidence limits, and that incomplete information about the outcome of diagnostic procedures among screening positives considerably worsens the feasibility of quality control. CONCLUSIONS: Although the methodology is specified for breast-cancer screening, it should be adaptable easily to other screening issues.

Breast Neoplasms↗

The distribution of Student's t-statistic for small samples from lognormal exposure distributions.

To assess compliance with industrial hygiene exposure criteria (e.g., TLVs), it may be necessary to perform statistical tests of hypotheses based on relatively small samples. For pollutants with long biological half-lives, the parameter most relevant for determining the risk faced by workers is the long-term arithmetic average concentration of the pollutant. In industrial environments it is common for pollutant concentrations to be approximately lognormal. Unfortunately, when based on small samples from lognormal distributions, the ordinary Student's t-statistic has some undesirable characteristics which are not recognized widely by practicing industrial hygienists. The difficulties in using the ordinary Student's t-statistic to evaluate the average exposure have been demonstrated. The properties of alternative test statistics have been explored. Some general observations on the implications of these findings have been made.

Air Pollutants, Occupational↗

Progress report on the guidance for industry for statistical aspects of the design, analysis, and interpretation of chronic rodent carcinogenicity studies of pharmaceuticals.

The U.S. Food and Drug Administration (FDA) is in the process of preparing a draft Guidance for Industry document on the statistical aspects of carcinogenicity studies of pharmaceuticals for public comment. The purpose of the document is to provide statistical guidance for the design of carcinogenicity experiments, methods of statistical analysis of study data, interpretation of study results, presentation of data and results in reports, and submission of electronic study data. This article covers the genesis of the guidance document and some statistical methods in study design, data analysis, and interpretation of results included in the draft FDA guidance document.

Animals↗

Potential use of the scan statistic for quality control in blood product manufacturing.

There are minimal standards for the processing of whole blood components, and to apply those standards requires a system of quality assurances. Excessive indications of failures in compliance trigger inspections and other remedial actions, but the demarcation of what is excessive is a critical issue. Issues of low volume in some production facilities, low expected frequency of nonconformance, multiple nonindependent statistical tests, and controlling both the false-positive and false-negative rates all complicate quality assurance procedures. The scan statistic is a statistic that computes the number of events in a moving window throughout the period of risk. Monitoring plans based on scan statistics are developed to estimate the probability that a process that is under control.

Algorithms↗

Statistical correction for non-parallelism in a urinary enzyme immunoassay.

Our aim was to develop a statistical method to correct for non-parallelism in an estrone-3-glucuronide (E1G) enzyme immunoassay (EIA). Non-parallelism of serially diluted urine specimens with a calibration curve was demonstrated in an EIA for E1G. A linear mixed-effects analysis of 40 urine specimens was used to model the relationship of E1G concentration with urine volume and derive a statistical correction. The model was validated on an independent sample and applied to 30 menstrual cycles from American women. Specificity, detection limit, parallelism, recovery, correlation with serum estradiol, and imprecision of the assay were determined. Intra-and inter-assay CVs were less than 14% for high- and low-urine controls. Urinary E1G across the menstrual cycle was highly correlated with serum estradiol (r= 0.94). Non-parallelism produced decreasing E1G concentration with increase in urine volume (slope = -0.210, p < 0.0001). At 50% inhibition, the assay had 100% cross-reactivity with E1G and 83% with 17beta-estradiol 3-glucuronide. The dose-response curve of the latter did not parallel that of E1G and is a possible cause of the non-parallelism. The statistical correction adjusting E1G concentration to a standardized urine volume produced parallelism in 24 independent specimens (slope = -0.043+/-0.010), and improved the average CV of E1G concentration across dilutions from 19.5%+/-5.6% before correction to 10.3%+/-5.3% after correction. A statistical method based on linear mixed effects modeling is an expedient approach for correction of non-parallelism, particularly for hormone data that will be analyzed in aggregate.

Adult↗