Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “STATISTICS”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

New class of level statistics in correlated disordered chains.

We study the properties of the level statistics of 1D disordered systems with long-range spatial correlations. We find a threshold value in the degree of correlations below which in the limit of large system size the level statistics follows a Poisson distribution (as expected for 1D uncorrelated-disordered systems), and above which the level statistics is described by a new class of distribution functions. At the threshold, we find that with increasing system size, the standard deviation of the function describing the level statistics converges to the standard deviation of the Poissonian distribution as a power law. Above the threshold we find that the level statistics is characterized by different functional forms for different degrees of correlations.

Journal Article↗

Accurate and efficient loop selections by the DFIRE-based all-atom statistical potential.

The conformations of loops are determined by the water-mediated interactions between amino acid residues. Energy functions that describe the interactions can be derived either from physical principles (physical-based energy function) or statistical analysis of known protein structures (knowledge-based statistical potentials). It is commonly believed that statistical potentials are appropriate for coarse-grained representation of proteins but are not as accurate as physical-based potentials when atomic resolution is required. Several recent applications of physical-based energy functions to loop selections appear to support this view. In this article, we apply a recently developed DFIRE-based statistical potential to three different loop decoy sets (RAPPER, Jacobson, and Forrest-Woolf sets). Together with a rotamer library for side-chain optimization, the performance of DFIRE-based potential in the RAPPER decoy set (385 loop targets) is comparable to that of AMBER/GBSA for short loops (two to eight residues). The DFIRE is more accurate for longer loops (9 to 12 residues). Similar trend is observed when comparing DFIRE with another physical-based OPLS/SGB-NP energy function in the large Jacobson decoy set (788 loop targets). In the Forrest-Woolf decoy set for the loops of membrane proteins, the DFIRE potential performs substantially better than the combination of the CHARMM force field with several solvation models. The results suggest that a single-term DFIRE-statistical energy function can provide an accurate loop prediction at a fraction of computing cost required for more complicate physical-based energy functions. A Web server for academic users is established for loop selection at the softwares/services section of the Web site http://theory.med.buffalo.edu/.

Membrane Proteins↗

Statistical-acoustics models of energy decay in systems of coupled rooms and their relation to geometrical acoustics.

An improved statistical-acoustics model of high-frequency sound fields in coupled rooms is developed by incorporating into prior models geometrical-acoustics corrections for both energy decay within subrooms and energy transfer between subrooms. The conditions under which statistical-acoustics models of coupled rooms are valid approximations to geometrical acoustics are examined by comparison of computational geometrical-acoustics predictions of decay curves in two- and three-room systems with those of both improved and prior statistical-acoustics models. The accuracy of the decay model used within subrooms is found to have a primary influence on the accuracy of predictions in coupled systems. Likewise, nondiffuse transfer of energy is shown to significantly affect decay of energy in systems of coupled rooms. The decrease in energy density of the reverberant field with distance from the source, which is predicted by geometrical acoustics, is found to result in spatial dependence of decay-curve shape for certain coupling geometries. Geometrical effects are shown to contribute to the failure of statistical-acoustics models in the case of strong coupling between subrooms; thus, previously proposed statistical-acoustics criteria cannot predict the point at which the models break down with consistent accuracy.

Acoustics↗

Integration of microbial ecology and statistics: a test to compare gene libraries.

Libraries of 16S rRNA genes provide insight into the membership of microbial communities. Statistical methods help to determine whether differences in library composition are artifacts of sampling or are due to underlying differences in the communities from which they are derived. To contribute to a growing statistical framework for comparing 16S rRNA libraries, we present a computer program, integral -LIBSHUFF, which calculates the integral form of the Cramér-von Mises statistic. This implementation builds upon the LIBSHUFF program, which uses an approximation of the statistic and makes a number of modifications that improve precision and accuracy. Once integral -LIBSHUFF calculates the P values, when pairwise comparisons are tested at the 0.05 level, the probability of falsely identifying a significant P value is 0.098 for a study with two libraries, 0.265 for three libraries, and 0.460 for four libraries. The potential negative effects of making the multiple pairwise comparisons necessitate correcting for the increased likelihood that differences between treatments are due to chance and do not reflect biological differences. Using integral -LIBSHUFF, we found that previously published 16S rRNA gene libraries constructed from Scottish and Wisconsin soils contained different bacterial lineages. We also analyzed the published libraries constructed for the zebrafish gut microflora and found statistically significant changes in the community during development of the host. These analyses illustrate the power of integral -LIBSHUFF to detect differences between communities, providing the basis for ecological inference about the association of soil productivity or host gene expression and microbial community composition.

Aeromonas↗

Individual survival time prediction using statistical models.

Doctors' survival predictions for terminally ill patients have been shown to be inaccurate and there has been an argument for less guesswork and more use of carefully constructed statistical indices. As statisticians, the authors are less confident in the predictive value of statistical models and indices for individual survival times. This paper discusses and illustrates a variety of measures which can be used to summarise predictive information available from a statistical model. The authors argue that models and statistical indices can be useful at the group or population level, but that human survival is so uncertain that even the best statistical analysis cannot provide single-number predictions of real use for individual patients.

Carcinoma, Non-Small-Cell Lung↗

Prognostic value of adaptive textural features--the effect of standardizing nuclear first-order gray level statistics and mixing information from nuclei having different area.

BACKGROUND: Nuclear texture analysis is a useful method to obtain quantitative information for use in prognosis of cancer. The first-order gray level statistics of a digitized light microscopic nuclear image may be influenced by variations in the image input conditions. Therefore, we have previously standardized the nuclear gray level mean value and standard deviation. However, there is a clear relation between nuclear DNA content, area, first-order statistics, and texture. For nuclei with approximately the same DNA content, the mean gray level increases with an increasing nuclear area. The aims of the present methodical work were to study: (1) whether the prognostic value of adaptive textural features varies with nuclear area, and (2) the effect of standardizing nuclear first-order statistics. METHODS: Nuclei from 134 cases of ovarian cancer were grouped into intervals according to nuclear area. Adaptive features were extracted from two different image sets, i.e., standardized and non-standardized nuclear images. RESULTS: The prognostic value of adaptive textural features varied strongly with nuclear area. A standardization of the first-order statistics significantly reduced this prognostic information. Several single features discriminated the two classes of cancer with a correct classification rate of 70%. CONCLUSION: Nuclei having an area between 2000-4999 pixels contained most of the class distance information between the good and poor prognosis classes of cancer. By considering the relation between nuclear area and texture, we avoided a loss of information caused by standardizing the first-order statistics and mixing data from cells having different nuclear area.

Algorithms↗

Statistical versus fuzzy measures of variable interaction in patients with stroke.

UNLABELLED: Evidence-based medicine, founded in probability-based statistics, applies what is the case for the collective to the individual patient. An intuitive approach, however, would define structure in the (physiologic) system of interest, the human being, directly relevant to other systems (patients) composed of similar variables. A difference in measure of variable interaction in the patient from that in the collective would show how extrapolation of information from the latter to the single patient is counterintuitive. METHODS: We compare statistical to 'fuzzy' measures of variable interaction. Three diagnostic variables are considered in 30 stroke patients who underwent the same diagnostic tests. 'Fit' (fuzzy information) values [0, 1] for degree of variable severity were expertly assigned by 2 blinded raters for real and fabricated patients. Fabricated patients were composed of real-patient 'fit' values after shuffling. Real and fabricated patients were each numerically represented as a set. Three groups of fabricated patients and the real patient group were studied. Statistical [Pearson's product-moment (regression analysis) and Spearman's rank correlation] and three different fuzzy measures of variable interaction were applied to patient data. RESULTS: Interaction for blood-vessel measured strong in real patients, and weak after one shuffle, using all fuzzy measures. By comparison, the same interaction was found in real patients by only 1 rater (Rater 2) using 1 statistical technique (Spearman's rank correlation) which, as did Pearson product-moment correlation, found a 'significant' interaction between blood-heart in fabricated patients. CONCLUSION: Our study suggests that the measure of variable interaction in nature - as combined in the individual (real) patient - is captured robustly by fuzzy measures and not so by standard statistical measures.

Evidence-Based Medicine↗

Phonetic diversity, statistical learning, and acquisition of phonology.

In learning to perceive and produce speech, children master complex language-specific patterns. Daunting language-specific variation is found both in the segmental domain and in the domain of prosody and intonation. This article reviews the challenges posed by results in phonetic typology and sociolinguistics for the theory of language acquisition. It argues that categories are initiated bottom-up from statistical modes in use of the phonetic space, and sketches how exemplar theory can be used to model the updating of categories once they are initiated. It also argues that bottom-up initiation of categories is successful thanks to the perception-production loop operating in the speech community. The behavior of this loop means that the superficial statistical properties of speech available to the infant indirectly reflect the contrastiveness and discriminability of categories in the adult grammar. The article also argues that the developing system is refined using internal feedback from type statistics over the lexicon, once the lexicon is well-developed. The application of type statistics to a system initiated with surface statistics does not cause a fundamental reorganization of the system. Instead, it exploits confluences across levels of representation which characterize human language and make bootstrapping possible.

Child, Preschool↗

Testing statistical significance scores of sequence comparison methods with structure similarity.

BACKGROUND: In the past years the Smith-Waterman sequence comparison algorithm has gained popularity due to improved implementations and rapidly increasing computing power. However, the quality and sensitivity of a database search is not only determined by the algorithm but also by the statistical significance testing for an alignment. The e-value is the most commonly used statistical validation method for sequence database searching. The CluSTr database and the Protein World database have been created using an alternative statistical significance test: a Z-score based on Monte-Carlo statistics. Several papers have described the superiority of the Z-score as compared to the e-value, using simulated data. We were interested if this could be validated when applied to existing, evolutionary related protein sequences. RESULTS: All experiments are performed on the ASTRAL SCOP database. The Smith-Waterman sequence comparison algorithm with both e-value and Z-score statistics is evaluated, using ROC, CVE and AP measures. The BLAST and FASTA algorithms are used as reference. We find that two out of three Smith-Waterman implementations with e-value are better at predicting structural similarities between proteins than the Smith-Waterman implementation with Z-score. SSEARCH especially has very high scores. CONCLUSION: The compute intensive Z-score does not have a clear advantage over the e-value. The Smith-Waterman implementations give generally better results than their heuristic counterparts. We recommend using the SSEARCH algorithm combined with e-values for pairwise sequence comparisons.

Base Sequence↗

Pattern statistics on Markov chains and sensitivity to parameter estimation.

BACKGROUND: In order to compute pattern statistics in computational biology a Markov model is commonly used to take into account the sequence composition. Usually its parameter must be estimated. The aim of this paper is to determine how sensitive these statistics are to parameter estimation, and what are the consequences of this variability on pattern studies (finding the most over-represented words in a genome, the most significant common words to a set of sequences,...). RESULTS: In the particular case where pattern statistics (overlap counting only) computed through binomial approximations we use the delta-method to give an explicit expression of sigma, the standard deviation of a pattern statistic. This result is validated using simulations and a simple pattern study is also considered. CONCLUSION: We establish that the use of high order Markov model could easily lead to major mistakes due to the high sensitivity of pattern statistics to parameter estimation.

Journal Article↗

Statistical methods in personality assessment research.

Emerging models of personality structure and advances in the measurement of personality and psychopathology suggest that research in personality and personality assessment has entered a stage of advanced development, in this article we examine whether researchers in these areas have taken advantage of new and evolving statistical procedures. We conducted a review of articles published in the Journal of Personality, Assessment during the past 5 years. Of the 449 articles that included some form of data analysis, 12.7% used only descriptive statistics, most employed only univariate statistics, and fewer than 10% used multivariate methods of data analysis. We discuss the cost of using limited statistical methods, the possible reasons for the apparent reluctance to employ advanced statistical procedures, and potential solutions to this technical shortcoming.

Journal Article↗

A survey of laboratory and statistical issues related to farmworker exposure studies.

Developing internally valid, and perhaps generalizable, farmworker exposure studies is a complex process that involves many statistical and laboratory considerations. Statistics are an integral component of each study beginning with the design stage and continuing to the final data analysis and interpretation. Similarly, data quality plays a significant role in the overall value of the study. Data quality can be derived from several experimental parameters including statistical design of the study and quality of environmental and biological analytical measurements. We discuss statistical and analytic issues that should be addressed in every farmworker study. These issues include study design and sample size determination, analytical methods and quality control and assurance, treatment of missing data or data below the method's limits of detection, and post-hoc analyses of data from multiple studies. Key words: analytical methodology, biomarkers, laboratory, limit of detection, omics, quality control, sample size, statistics.

Agriculture↗

[Models of causal inference: critical analysis of the use of statistics in epidemiology].

The foundations on which the concept of risk has been constructed are discussed. A description of Rubin's model of causal inference, which was first developed in the domain of applied statistics, and later incorporated into a branch of epidemiology, is taken as the starting point. Analysis of the premisses of causal inference brings to light the logical stages in the construction of the concept of risk, allowing it to be understood "from the inside". The abovementioned branch of statistics and epidemiology seeks to demonstrate that statistics can infer causality instead of simply revealing statistical associations; the model gives the basis for estimating that which way be defined as the effect of a cause. Using this procedural distinction between causal inference and association, the model also seeks to differentiate between the epidemiologial dimension of concepts and the merely statistical dimension. This leads to greater complexity when handing the concepts of interation and coofounding. The redective aspects inherent in this methodological construction of risk are here high lighted. Thus, whether applied to individual or populational inferences, this methodological construction imposes limits that need to be taken into account in its theoretical and practical application to epidemiology.

Causality↗

Incorporation of statistical uncertainty in health economic modelling studies using second-order Monte Carlo simulations.

Health economic modelling studies are of interest to many parties with different responsibilities and diverging interests. Therefore, it is obvious that recognising the relevance of statistical uncertainty and dealing with it appropriately are required to obtain unbiased results from health economic modelling studies, especially when those data are being used for reimbursement decisions. In this manuscript we explore the relevance of the incorporation of statistical uncertainty in a health economic model and identify various types of statistical uncertainty. The concepts were applied to a hypothetical Markov model for a hypothetical antiparkinsonian (AP) product. The method was based on the incorporation of probability distributions in the input variables using a second-order Monte Carlo simulation and the definition of minimum relevant differences for clinical and economic input variables and outcomes. Our paper shows that the outcomes of a health economic model might be severely biased when statistical uncertainty is not taken into account, which justifies the need for the incorporation of statistical uncertainty in a health economic model.

Antiparkinson Agents↗

Systematic evaluation and comparison of statistical tests for publication bias.

BACKGROUND: This study evaluates the statistical and discriminatory powers of three statistical test methods (Begg's, Egger's, and Macaskill's) to detect publication bias in meta-analyses. METHODS: The data sources were 130 reviews from the Cochrane Database of Systematic Reviews 2002 issue, which considered a binary endpoint and contained 10 or more individual studies. Funnel plots with observers'agreements were selected as a reference standard. We evaluated a trade-off between sensitivity and specificity by varying cut-off p-values, power of statistical tests given fixed false positive rates, and area under the receiver operating characteristic curve. RESULTS: In 36 reviews, 733 original studies evaluated 2,874,006 subjects. The number of trials included in each ranged from 10 to 70 (median 14.5). Given that the false positive rate was 0.1, the sensitivity of Egger's method was 0.93, and was larger than that of Begg's method (0.86) and Macaskill's method (0.43). The sensitivities of three statistical tests increased as the cut-off p-values increased without a substantial decrement of specificities. The area under the ROC curve of Egger's method was 0.955 (95% confidence interval, 0.889-1.000) and was not different from that of Begg's method (area=0.913, p=0.2302), but it was larger than that of Macaskill's method (area=0.719, p=0.0116). CONCLUSION: Egger's linear regression method and Begg's method had stronger statistical and discriminatory powers than Macaskill's method for detecting publication bias given the same type I error level. The power of these methods could be improved by increasing the cut-off p-value without a substantial increment of false positive rate.

Linear Models↗

[A statistical evaluation of the variability in the measurements of the resistive index in kidney transplantation].

INTRODUCTION: Doppler ultrasound (US) is a valuable tool to measure blood flow in the transplanted kidney, but its operator-dependence can greatly affect repeatability and reproducibility of measurements. Aim of this work was to evaluate intraobserver and interobserver variability in measuring the resistive index (RI) in renal transplants. PATIENTS AND METHODS: Ten renal transplant recipients were randomly selected among those undergoing follow-up and examined by two operators (FG and LB) with 3.5 MHz and 10 MHz scanheads to assess the variability of RI measurements. Each observer obtained two measurements of the RI with each scanhead within a 10-15 minutes' period. In all, 80 measurements were made, 4 per patient per observer. The statistical analysis included two-tailed Student's t-test for paired data and calculation of repeatability/reproducibility coefficients. RESULTS: Student's t-test analysis demonstrated a statistically significant difference (p = 0.037) between the means of the first and second measurements by FG with the 3.5 MHz scanhead and the first and the second measurements by LB with the same scanhead. Differences between the other means were not statistically significant. Intraobserver variability ranged 0.03 units (or 2.07%) and 0.07 units (or 4.24%), while interobserver variability was 0.04 units with both 3.5 and 10 MHz scanheads, or 3.61 and 3.73%, respectively. CONCLUSIONS: Doppler US of renal transplants has statistically quantifiable operator-dependent variability: the possible evidence of statistically significant differences can be minimized by having the same operator make the measurements. However, RI variations ranging 0.02 to 0.04 units should not be considered significant.

Adult↗

Statistical Inference (part 1): Basic Concepts.

In the most common research situations, the investigator cannot directly assess the whole population of interest. For this reason, a sample is often studied to infer the actual population measures or parameters. Statistical inference comprises the application of methods to analyze the sample data in order to estimate the population parameters. The basic assumption in statistical inference is that each individual within the population of interest has the same probability of being included in a specific sample. When the sample is not randomly selected. the study findings can still be generalized if the sample can be considered representative of the whole population of interest. A set of statistical methods used to infer the population parameters is performed under the assumption that the sample estimates follow a bell-shaped distribution, called normal distribution. This article presents with the help of examples, the logic used in the sampling distribution theory. The concept of normal (also called gaussian) sampling distribution has an important role in statistical inference, even when the population values are not normally distributed. In fact, in the statistical inference process, the form of the distribution of the sample estimates is more important than the distribution of the individual values.

Journal Article↗

Improved statistical characterization of prosthetic heart valve hydrodynamics using a performance index and regression analysis.

BACKGROUND AND AIMS OF THE STUDY: The ISO 5840 Standard (Cardiovascular implants - Cardiac valves) currently requires a minimum of three test samples per size for hydrodynamic testing. Typically, the only statistical analysis performed is a descriptive analysis, with the mean (+/- SE) given for each size and cardiac output (CO). The study aim was to develop better statistical methods, incorporating regression analysis of a performance index, equal to the effective orifice area divided by the tissue annulus area. The analysis is performed on the full dataset, with size and CO as independent variables. METHODS: Hydrodynamic data of Ionescu-Shiley pericardial valves from a published study were used to compare the two analysis methods. Three samples each of size 19, 23 and 27 mm valves were tested at COs of 4.2, 5.6, 7.0 and 8.4 l/min. Descriptive statistics were performed for each size and CO. Regression analysis was also performed on the full dataset. Confidence intervals (CI) were calculated for each statistical method and compared. RESULTS: The regression equation that best fitted the data was: Performance Index (PI) = -1.63 + (0.011 x CO) + (0.167 x size) - (0.0036 x size2). All four parameter estimates were significantly different from zero (p <0.02). The SE of the mean was 0.015 for COs of 4.2 or 8.4 l/min, and 0.013 for COs of 5.6 or 7.0 l/min, less than that of nine of 12 of the individual descriptive analysis. CI for the regression analysis were substantially tighter, averaging one-third the width of those of the descriptive statistics. CONCLUSION: The tighter CI resulting from the regression analysis allows a better comparison of the PI to an objective performance criterion. Such methods should be considered for inclusion in the new version of the ISO 5840 standard for prosthetic heart valves.

Aortic Valve↗