Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “STATISTICS”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Uses and limitations of statistical accounting for random error correlations, in the validation of dietary questionnaire assessments.

OBJECTIVE: To examine statistical models that account for correlation between random errors of different dietary assessment methods, in dietary validation studies. SETTING: In nutritional epidemiology, sub-studies on the accuracy of the dietary questionnaire measurements are used to correct for biases in relative risk estimates induced by dietary assessment errors. Generally, such validation studies are based on the comparison of questionnaire measurements (Q) with food consumption records or 24-hour diet recalls (R). In recent years, the statistical analysis of such studies has been formalized more in terms of statistical models. This made the need of crucial model assumptions more explicit. One key assumption is that random errors must be uncorrelated between measurements Q and R, as well as between replicate measurements R1 and R2 within the same individual. These assumptions may not hold in practice, however. Therefore, more complex statistical models have been proposed to validate measurements Q by simultaneous comparisons with measurements R plus a biomarker M, accounting for correlations between the random errors of Q and R. CONCLUSIONS: The more complex models accounting for random error correlations may work only for validation studies that include markers of diet based on physiological knowledge about the quantitative recovery, e.g. in urine, of specific elements such as nitrogen or potassium, or stable isotopes administered to the study subjects (e.g. the doubly labelled water method for assessment of energy expenditure). This type of marker, however, eliminates the problem of correlation of random errors between Q and R by simply taking the place of R, thus rendering complex statistical models unnecessary.

Diet↗

Assessment of the beryllium lymphocyte proliferation test using statistical process control.

Despite more than 20 years of surveillance and epidemiologic studies using the beryllium blood lymphocyte proliferation test (BeBLPT) as a measure of beryllium sensitization (BeS) and as an aid for diagnosing subclinical chronic beryllium disease (CBD), improvements in specific understanding of the inhalation toxicology of CBD have been limited. Although epidemiologic data suggest that BeS and CBD risks vary by process/work activity, it has proven difficult to reach specific conclusions regarding the dose-response relationship between workplace beryllium exposure and BeS or subclinical CBD. One possible reason for this uncertainty could be misclassification of BeS resulting from variation in BeBLPT testing performance. The reliability of the BeBLPT, a biological assay that measures beryllium sensitization, is unknown. To assess the performance of four laboratories that conducted this test, we used data from a medical surveillance program that offered testing for beryllium sensitization with the BeBLPT. The study population was workers exposed to beryllium at various facilities over a 10-year period (1992-2001). Workers with abnormal results were offered diagnostic workups for CBD. Our analyses used a standard statistical technique, statistical process control (SPC), to evaluate test reliability. The study design involved a repeated measures analysis of BeBLPT results generated from the company-wide, longitudinal testing. Analytical methods included use of (1) statistical process control charts that examined temporal patterns of variation for the stimulation index, a measure of cell reactivity to beryllium; (2) correlation analysis that compared prior perceptions of BeBLPT instability to the statistical measures of test variation; and (3) assessment of the variation in the proportion of missing test results and how time periods with more missing data influenced SPC findings. During the period of this study, all laboratories displayed variation in test results that were beyond what would be expected due to chance alone. Patterns of test results suggested that variations were systematic. We conclude that laboratories performing the BeBLPT or other similar biological assays of immunological response could benefit from a statistical approach such as SPC to improve quality management.

Air Pollutants, Occupational↗

The misuse of statistics: concepts, tools, and a research agenda.

This paper presents concerns regarding misuse of statistics in scientific work, especially in biomedical research. The paper discusses what is meant by "misuse." It appears that misuse arises from various sources: degrees of competence in statistical theory and methods, honest error in the application of methods, egregious negligence, and deliberate deception (misconduct.) The incidence of error is partly due to a perceived need to meet artificial statistical criteria for acceptance of research reports for publication by journals. There has been no systematic research into the prevalence of misuse or its breakdown by type. Nonetheless, there are ways to encourage, or even to enforce, good statistical practice. These can be greatly supported by use of available statistical ethics documents. This article suggests lines of further research that could define the problem more explicitly and that might lead to additional corrective measures.

Biomedical Research↗

Interactive analysis of Belgian vital statistics on the Internet.

The purpose of the Centre for Operational Research in Public Health (CORPH) is to optimize the accessibility to health information, thus making it possible to measure and follow up the health status of the Belgian population. The Standardized Procedures for Mortality Analysis (SPMA) software was developed in order to facilitate the use of vital statistics for health policy-makers and scientific researchers. Nowadays, SPMA is available on the Internet, because accessibility to health information is crucial. SPMA serves via a system of menus as the interface between databases (population, birth, and mortality) on one hand and statistical procedures on the other hand. Users can choose the parameters such as year, cause of death, geographical level, and statistical indicator, and so dynamic reports are produced 'on demand'. These procedures are available for the following modules: overall mortality, specific cause mortality, and perinatal statistics. Analysis can be carried out for one specific year or for a period over time. Pre-defined procedures accessible through menus make SPMA user-friendly, as it can be used without any preliminary knowledge of the statistical package. Tables, charts, or maps display the results. Users need only an Internet browser to access the application.

Adolescent↗

A test statistic to detect errors in sib-pair relationships.

Several authors have proposed algorithms to detect Mendelian errors in human genetic linkage data. Most currently available methods use likelihood-based methods on multiplex family data to identify typing or pedigree errors. These algorithms cannot be applied in many sib-pair collections, because of lack of parental-genotype information. Nonetheless, misspecifying the relationships between individuals has serious consequences for sib-pair linkage studies: false relationships bias the statistics designed to identify linkage with disease phenotypes. To test the hypothesis that two individuals are sibs, we propose a test statistic based on the summation, over a large number of genetic markers, of the number of alleles shared identical by state by a pair of individuals, for each marker. The test statistic has an approximately normal distribution under the null hypothesis, and extreme negative values correspond to nonsib pairs. Power and significance studies show that the test statistic calculated by use of 50 unlinked markers has 96% power to detect half-sibs and has 100% power to detect unrelated individuals as not full-sib pairs, with a 5% false-positive rate. Furthermore, extreme positive values of the test statistic identify sibs as MZ twins.

Algorithms↗

Statistics of assay validation in high throughput cell imaging of nuclear factor kappaB nuclear translocation.

This report describes statistical validation methods implemented on assay data for inhibition of subcellular redistribution of nuclear factor kappaB (NF kappaB) in HeLa cells. We quantified cellular inhibition of cytoplasmic-nuclear translocation of NF kappaB in response to a range of concentrations of interleukin-1 (IL-1) receptor antagonist in the presence of IL-1alpha using eight replicate rows in each four 96-well plates scanned five times on each of 2 days. Translocation was measured as the fractional localized intensity of the nucleus (FLIN), an implementation of our more general fractional localized intensity of the compartments (FLIC), which analyzes whole compartments in the context of the entire cell. The NF kappaB antagonist assay (inhibition of IL-1- induced NF kappaB translocation) data were collected on a Q3DM (San Diego, CA) EIDAQtrade mark 100 high throughput microscopy system. [In 2003, Q3DM was purchased by Beckman Coulter Inc. (Fullerton, CA), which released the IC 100 successor to the EIDAQ 100.] The generalized FLIC method is described along with two-point (minimum-maximum) and multiple point titration statistical methods. As a ratio of compartment intensities that tend to change proportionally, FLIN was resistant to photobleaching errors. Two-point minimum-maximum statistical analyses yielded the following: a Z' of 0.174 with the data as n = 320 independent well samples; Z' by row data in a range of 0.393-0.933, with a mean of 0.766; by-plate Z' data of 0.310, 0.443, 0.545, and 0.794; and by-plate means of columns Z' data of 0.879, 0.927, 0.945, and 0.963. The mean 50% inhibitory concentration (IC50) for IL-1 receptor antagonist over all experiments was 213 ng/ml. The combined IC50 coefficients of variation (CVs) were 0.74%, 0.85%, 2.09%, and 2.52% for the four plates. Repeatability IC50 CVs were as follows: day to day 3.0%, row to row 8.0%, plate to plate 2.8%, and day to day 0.6%. The number of cells required for statistically resolvable differences in dose concentrations, plotted in a family of FLIN sigma/deltamicro (SD/range) curves and tabulated, demonstrated cell-by-cell assay precision with our combined sigma/deltamicro = 0.32 that required approximately 10-fold fewer cells than in a previously reported NF kappaB assay with sigma/deltamicro = 1.52. To better understand the relationship between cell-by-cell measurements and IC50 precision, 500 Monte Carlo simulations with varying cell-measurement SDs were used to explore three-, five-, seven-, and 11-point model titrations. The reductions in deltaIC50 90% confidence intervals from 11- to three-point titrations were 10-fold with the previously reported sigma/deltamicro = 1.52 and twofold with our sigma/deltamicro = 0.32. With these normalized parameters, this report provides a common statistical foundation, independent of the assay details, for evaluating the performance of imaging data on any instrument.

Active Transport, Cell Nucleus↗

Making sense of score statistics for sequence alignments.

The search for similarity between two biological sequences lies at the core of many applications in bioinformatics. This paper aims to highlight a few of the principles that should be kept in mind when evaluating the statistical significance of alignments between sequences. The extreme value distribution is first introduced, which in most cases describes the distribution of alignment scores between a query and a database. The effects of the similarity matrix and gap penalty values on the score distribution are then examined, and it is shown that the alignment statistics can undergo an abrupt phase transition. A few types of random sequence databases used in the estimation of statistical significance are presented, and the statistics employed by the BLAST, FASTA and PRSS programs are compared. Finally the different strategies used to assess the statistical significance of the matches produced by profiles and hidden Markov models are presented.

Amino Acid Sequence↗

A comparative review of statistical methods for discovering differentially expressed genes in replicated microarray experiments.

MOTIVATION: A common task in analyzing microarray data is to determine which genes are differentially expressed across two kinds of tissue samples or samples obtained under two experimental conditions. Recently several statistical methods have been proposed to accomplish this goal when there are replicated samples under each condition. However, it may not be clear how these methods compare with each other. Our main goal here is to compare three methods, the t-test, a regression modeling approach (Thomas et al., Genome Res., 11, 1227-1236, 2001) and a mixture model approach (Pan et al., http://www.biostat.umn.edu/cgi-bin/rrs?print+2001,2001a,b) with particular attention to their different modeling assumptions. RESULTS: It is pointed out that all the three methods are based on using the two-sample t-statistic or its minor variation, but they differ in how to associate a statistical significance level to the corresponding statistic, leading to possibly large difference in the resulting significance levels and the numbers of genes detected. In particular, we give an explicit formula for the test statistic used in the regression approach. Using the leukemia data of Golub et al. (Science, 285, 531-537, 1999), we illustrate these points. We also briefly compare the results with those of several other methods, including the empirical Bayesian method of Efron et al. (J. Am. Stat. Assoc., to appear, 2001) and the Significance Analysis of Microarray (SAM) method of Tusher et al. (PROC: Natl Acad. Sci. USA, 98, 5116-5121, 2001).

Acute Disease↗

Primer design and marker clustering for multiplex SNP-IT primer extension genotyping assay using statistical modeling.

MOTIVATION: The optimization of the primer design is critical for the development of high-throughput SNP genotyping methods. Recently developed statistical models of the SNP-IT primer extension genotyping reaction allow further improvement of primer quality for the assay. RESULTS: Here we describe how the statistical models can be used to improve primer design for the assay. We also show how to optimize clustering of the SNP markers into multiplex panels using statistical model for multiplex SNP-IT. The primer set failure probability calculated by a model is used as a minimization function for both primer selection and primers clustering. Three clustering algorithms for the multiplex genotyping SNP-IT assay are described and their relative performance is evaluated. We also describe the approaches to improve the speed of primer design and clustering calculations when using the statistical models. Our clustering decreases the average failure probability of the marker set by 7-25%. The experimental marker failure rate in the multiplex reaction was reduced dramatically and success rate can be achieved as high as 96%. AVAILABILITY: The primer design using statistical models is freely available from www.autoprimer.com.

Algorithms↗

Differential gene expression detection using penalized linear regression models: the improved SAM statistics.

UNLABELLED: Differential gene expression detection using microarrays has received lots of research interests recently. Many methods have been proposed, including variants of F-statistics, non-parametric approaches and empirical Bayesian methods etc. The SAM statistics has been shown to have good performance in empirical studies. SAM is more like an ad hoc shrinkage method. The idea is that for small sample microarray data, it is often useful to pool information across genes to improve efficiency. Under Bayesian framework Smyth formally derived the test statistics with shrinkage using the hierarchical models. In this paper we cast differential gene expression detection in the familiar framework of linear regression model. Commonly used test statistics correspond to using least squares to estimate the regression parameters. Based on the vast literature of research on linear models, we can naturally consider other alternatives. Here we explore the penalized linear regression. We propose the penalized t-/F-statistics for two-class microarray data based on [Formula: see text] penalty. We will show that the penalized test statistics intuitively makes sense and through applications we illustrate its good performance. AVAILABILITY: Supplementary information including program codes, more detailed analysis results and R functions for the proposed methods can be found at http://www.biostat.umn.edu/~baolin/research CONTACT: baolin@biostat.umn.edu SUPPLEMENTARY INFORMATION: http://www.biostat.umn.edu/~baolin/research.

Cell Line, Tumor↗

Statistical analysis of microarray data: a Bayesian approach.

The potential of microarray data is enormous. It allows us to monitor the expression of thousands of genes simultaneously. A common task with microarray is to determine which genes are differentially expressed between two samples obtained under two different conditions. Recently, several statistical methods have been proposed to perform such a task when there are replicate samples under each condition. Two major problems arise with microarray data. The first one is that the number of replicates is very small (usually 2-10), leading to noisy point estimates. As a consequence, traditional statistics that are based on the means and standard deviations, e.g. t-statistic, are not suitable. The second problem is that the number of genes is usually very large (approximately 10,000), and one is faced with an extreme multiple testing problem. Most multiple testing adjustments are relatively conservative, especially when the number of replicates is small. In this paper we present an empirical Bayes analysis that handles both problems very well. Using different parametrizations, we develop four statistics that can be used to test hypotheses about the means and/or variances of the gene expression levels in both one- and two-sample problems. The methods are illustrated using experimental data with prior knowledge. In addition, we present the result of a simulation comparing our methods to well-known statistics and multiple testing adjustments.

Algorithms↗

An empirical evaluation of genetic distance statistics using microsatellite data from bear (Ursidae) populations.

A large microsatellite data set from three species of bear (Ursidae) was used to empirically test the performance of six genetic distance measures in resolving relationships at a variety of scales ranging from adjacent areas in a continuous distribution to species that diverged several million years ago. At the finest scale, while some distance measures performed extremely well, statistics developed specifically to accommodate the mutational processes of microsatellites performed relatively poorly, presumably because of the relatively higher variance of these statistics. At the other extreme, no statistic was able to resolve the close sister relationship of polar bears and brown bears from more distantly related pairs of species. This failure is most likely due to constraints on allele distributions at microsatellite loci. At intermediate scales, both within continuous distributions and in comparisons to insular populations of late Pleistocene origin, it was not possible to define the point where linearity was lost for each of the statistics, except that it is clearly lost after relatively short periods of independent evolution. All of the statistics were affected by the amount of genetic diversity within the populations being compared, significantly complicating the interpretation of genetic distance data.

Alleles↗

Biological models and statistical interactions: an example from multistage carcinogenesis.

From the assessment of statistical interaction between risk factors it is tempting to infer the nature of the biologic interaction between the factors. However, the use of statistical analyses of epidemiologic data to infer biologic processes can be misleading. as an example, we consider the multistage model of carcinogenesis. Under this biologic model, it is shown, by means of simple hypothetical examples, that even if carcinogenic factors act independently, some pairs may fit an additive statistical model, some a multiplicative statistical model, and some neither. The elucidation of biological interactions by means of statistical models requires the imaginative and prudent use of inductive and deductive reasoning; it cannot be done mechanically.

Cell Transformation, Neoplastic↗

Unveiling non-small cell lung cancer treatment effect heterogeneity: a comparative analysis of statistical methods.

BACKGROUND: For patients with advanced non-small cell lung cancer lacking targetable genomic alterations, the impact of clinicogenomic characteristics on the effectiveness of combining chemotherapy with immunotherapy is unclear. METHODS: We evaluated 4 statistical methods for detecting heterogeneous treatment effects related to clinical factors, including programmed death-ligand 1 expression, tumor mutation burden, and stage at diagnosis, using the American Association for Cancer Research Project Genomics Evidence Neoplasia Exchange BioPharma Collaborative dataset supplemented with institutional data collected under the same data curation model. A 2-sided P value of no more than .05 was used to denote statistical significance for all analyses. RESULTS: The mixture model revealed 2 latent subgroups: in one subgroup, there was no meaningful treatment effect, with average progression-free survival (PFS) only 5% longer with immunotherapy alone (95% confidence interval [CI] = -19% to 35%); in the second subgroup, immunotherapy alone was associated with a 35% decrease in average PFS (95% CI = -59% to 2%), corresponding to a ratio in treatment effects of 1.62 (95% CI = 1.02 to 2.57). There was a marginal association between lower tumor mutation burden levels and membership in the subgroup with improved PFS following receipt of chemoimmunotherapy. The causal survival forest highlighted the importance of tumor mutation burden (variable importance ranking: 1) and programmed death-ligand 1 (variable importance ranking: 3) when assessing heterogeneity. In contrast, the accelerated failure time and Cox proportional hazards models did not detect any statistically significant heterogeneous treatment effects. In simulations, the mixture model identified heterogeneous treatment effects more frequently than other methods, especially with weak covariate relationships, demonstrating its utility for informing personalized treatment approaches. CONCLUSIONS: The application of novel statistical methods to large scale clinico-genomic databases offers an opportunity to more accurately identify heterogeneous treatment effects in some settings as compared to traditional statistical methods. Applying such methods to the AACR Project GENIE BPC non-small cell lung cancer data indicated a potential association between decreasing tumor mutation burden and improved outcomes with chemoimmunotherapy as compared to immunotherapy alone.

Humans↗

Improved statistical methods reveal direct interactions between 16S and 23S rRNA.

Recent biochemical studies have indicated a number of regions in both the 16S and 23S rRNA that are exposed on the ribosomal subunit surface. In order to predict potential interactions between these regions we applied novel phylogenetically-based statistical methods to detect correlated nucleotide changes occurring between the rRNA molecules. With these methods we discovered a number of highly significant correlated changes between different sets of nucleotides in the two ribosomal subunits. The predictions with the highest correlation values belong to regions of the rRNA subunits that are in close proximity according to recent crystal structures of the entire ribosome. We also applied a new statistical method of detecting base triple interactions within these same rRNA subunit regions. This base triple statistic predicted a number of new base triples not detected by pair-wise interaction statistics within the rRNA molecules. Our results suggest that these statistical methods may enhance the ability to detect novel structural elements both within and between RNA molecules.

Animals↗

Statistical genetics concepts and approaches in schizophrenia and related neuropsychiatric research.

Statistical genetics is a research field that focuses on mathematical models and statistical inference methodologies that relate genetic variations (ie, naturally occurring human DNA sequence variations or "polymorphisms") to particular traits or diseases (phenotypes) usually from data collected on large samples of families or individuals. The ultimate goal of such analysis is the identification of genes and genetic variations that influence disease susceptibility. Although of extreme interest and importance, the fact that many genes and environmental factors contribute to neuropsychiatric diseases of public health importance (eg, schizophrenia, bipolar disorder, and depression) complicates relevant studies and suggests that very sophisticated mathematical and statistical modeling may be required. In addition, large-scale contemporary human DNA sequencing and related projects, such as the Human Genome Project and the International HapMap Project, as well as the development of high-throughput DNA sequencing and genotyping technologies have provided statistical geneticists with a great deal of very relevant and appropriate information and resources. Unfortunately, the use of these resources and their interpretation are not straightforward when applied to complex, multifactorial diseases such as schizophrenia. In this brief and largely nonmathematical review of the field of statistical genetics, we describe many of the main concepts, definitions, and issues that motivate contemporary research. We also provide a discussion of the most pressing contemporary problems that demand further research if progress is to be made in the identification of genes and genetic variations that predispose to complex neuropsychiatric diseases.

Chromosome Mapping↗

Methodological and statistical problems in sleep apnea research: the literature on uvulopalatopharyngoplasty.

A comprehensive review of the literature on the surgical treatment of sleep apnea found 37 appropriate papers (total n = 992) on uvulopalatopharyngoplasty (UPPP). Methodological and statistical problems in these papers included the following: 1) There were no randomized studies and few (n = 4) with control groups. 2) Median sample size was only 21.5; thus statistical power was low and clinically important associations were routinely classified as "not statistically significant". 3) Only one paper presented the confidence bounds that might distinguish between statistical and clinical significance. 4) Because of short follow-up time and infrequent repeat follow-ups, little is known about whether UPPP results deteriorate with time. 5) In at least 15 papers, bias caused by retrospective designs and nonrandom loss to follow-up raised questions about the generalizability of results. 6) Few papers associated polysomnographic data with patient-based quality of life measures. 7) Missing data and missing and inconsistent definitions were common. 8) Baseline measures were often biased because the same assessment was inappropriately but routinely used for both screening and baseline. We conclude that because of these and other problems, there is much that is needlessly unknown about UPPP. It is the responsibility of the research and professional communities to define training, editorial and review procedures that will raise the methodological and statistical quality of published research.

Humans↗

A statistical analysis of weekday operating room anesthesia group staffing costs at nine independently managed surgical suites.

UNLABELLED: At many surgical suites, surgeons and patients schedule elective cases on whatever future workday they choose, resulting in there being no limit on the number of cases performed each day. Staff are then scheduled in the manner that satisfies the marketing guarantee to the surgeons, satisfies labor contracts, and minimizes staffing costs. We assessed weekday nurse anesthesia group staffing at nine such suites to determine whether statistical methods can identify staffing solutions whereby all the cases are covered but for which staffing costs are less than those obtained using the staffing plans implemented by anesthesia groups' managers. Two years of operating room information system case duration and staffing data were analyzed. First- and second-shift staffing was assessed using previously published algorithms. The statistical methods identified staffing solutions with significantly decreased labor costs than those currently being used at eight of the nine surgical suites. The statistical methods relied more on overtime than second-shift staffing. The incremental decrease in staffing costs achievable by using overlapping 8-, 10-, and 13-h shifts was negligible. Overall, we found that statistical methods can identify, for some surgical suites, staffing solutions whereby all the cases are covered but for which costs are significantly less and productivity significantly more than those obtained using the plans developed by the managers based on their experience and the data. IMPLICATIONS: Statistical methods can identify, for some surgical suites, anesthesia staffing solutions whereby all the cases are covered but for which labor costs are significantly less than those obtained using the staffing plans developed by the managers based on data and their experience.

Anesthesia↗