Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Sample-size requirements for developing strategies, based on the pupal/demographic survey, for the targeted control of dengue.

Several methods to determine the sample size required for a reliable and practical assessment of the number of Aedes aegypti pupae in a community in Puerto Rico have been explored. Because the pupae were highly aggregated, the data were fitted to a negative binomial distribution. Classical statistical-inference methods for sample-size determination demanded the sampling of >3,000 premises for a reliable estimation of the mean number of pupae/person (with a 15% error). This number was reduced to 1,000-1,200 premises after applying a finite-population correction. Database sub-sampling simulations, with increasing sample sizes, showed that the variability in the mean relative abundance of container types and in the mean number of pupae/container substantially decreased after sampling 186 and 310 premises, respectively. Sequential sampling was applied to test the hypotheses that the number of female pupae/person was at least 0.19 (considered the dengue epidemic threshold) or no greater than 0.10 (arbitrarily set as the safe level). After sampling only 25 premises in the first survey and 125 in the second, it was determined that the densities of female pupae were above the epidemic threshold. Thus, sequential sampling provided substantial reductions in the sample size required to determine if vector control was needed. Validation of the Ae. aegypti thresholds required for dengue transmission could confer viability and efficiency to dengue-vector surveillance and control programmes.

Aedes↗

Joint analysis of two microarray gene-expression data sets to select lung adenocarcinoma marker genes.

BACKGROUND: Due to the high cost and low reproducibility of many microarray experiments, it is not surprising to find a limited number of patient samples in each study, and very few common identified marker genes among different studies involving patients with the same disease. Therefore, it is of great interest and challenge to merge data sets from multiple studies to increase the sample size, which may in turn increase the power of statistical inferences. In this study, we combined two lung cancer studies using microarray GeneChip, employed two gene shaving methods and a two-step survival test to identify genes with expression patterns that can distinguish diseased from normal samples, and to indicate patient survival, respectively. RESULTS: In addition to common data transformation and normalization procedures, we applied a distribution transformation method to integrate the two data sets. Gene shaving (GS) methods based on Random Forests (RF) and Fisher's Linear Discrimination (FLD) were then applied separately to the joint data set for cancer gene selection. The two methods discovered 13 and 10 marker genes (5 in common), respectively, with expression patterns differentiating diseased from normal samples. Among these marker genes, 8 and 7 were found to be cancer-related in other published reports. Furthermore, based on these marker genes, the classifiers we built from one data set predicted the other data set with more than 98% accuracy. Using the univariate Cox proportional hazard regression model, the expression patterns of 36 genes were found to be significantly correlated with patient survival (p < 0.05). Twenty-six of these 36 genes were reported as survival-related genes from the literature, including 7 known tumor-suppressor genes and 9 oncogenes. Additional principal component regression analysis further reduced the gene list from 36 to 16. CONCLUSION: This study provided a valuable method of integrating microarray data sets with different origins, and new methods of selecting a minimum number of marker genes to aid in cancer diagnosis. After careful data integration, the classification method developed from one data set can be applied to the other with high prediction accuracy.

Adenocarcinoma↗

MultiSeq: unifying sequence and structure data for evolutionary analysis.

BACKGROUND: Since the publication of the first draft of the human genome in 2000, bioinformatic data have been accumulating at an overwhelming pace. Currently, more than 3 million sequences and 35 thousand structures of proteins and nucleic acids are available in public databases. Finding correlations in and between these data to answer critical research questions is extremely challenging. This problem needs to be approached from several directions: information science to organize and search the data; information visualization to assist in recognizing correlations; mathematics to formulate statistical inferences; and biology to analyze chemical and physical properties in terms of sequence and structure changes. RESULTS: Here we present MultiSeq, a unified bioinformatics analysis environment that allows one to organize, display, align and analyze both sequence and structure data for proteins and nucleic acids. While special emphasis is placed on analyzing the data within the framework of evolutionary biology, the environment is also flexible enough to accommodate other usage patterns. The evolutionary approach is supported by the use of predefined metadata, adherence to standard ontological mappings, and the ability for the user to adjust these classifications using an electronic notebook. MultiSeq contains a new algorithm to generate complete evolutionary profiles that represent the topology of the molecular phylogenetic tree of a homologous group of distantly related proteins. The method, based on the multidimensional QR factorization of multiple sequence and structure alignments, removes redundancy from the alignments and orders the protein sequences by increasing linear dependence, resulting in the identification of a minimal basis set of sequences that spans the evolutionary space of the homologous group of proteins. CONCLUSION: MultiSeq is a major extension of the Multiple Alignment tool that is provided as part of VMD, a structural visualization program for analyzing molecular dynamics simulations. Both are freely distributed by the NIH Resource for Macromolecular Modeling and Bioinformatics and MultiSeq is included with VMD starting with version 1.8.5. The MultiSeq website has details on how to download and use the software: http://www.scs.uiuc.edu/~schulten/multiseq/

Algorithms↗

The scientific status of the Rorschach.

Demonstrated that the interpretation of projective test data is semantic, not probabilistic. The clinician does not employ the model of statistical inference in evaluating the meaning of test responses, although he may employ probabilistic rules as guides to interpreation. Procedural rules for interpreting meaning, the nature of clinical diagnosis from psychological tests, and the meaning of prediction as a clinical activity are discussed. It was concluded that clinicians do not make inferences, in the mathematical sense of this term.

Association↗

Unidimensionality and bandwidth in the Center for Epidemiologic Studies Depression (CES-D) Scale.

In this study, we compared classical test theory (CTT) and item response theory (IRT) approaches in analyzing the Center for Epidemiological Studies Depression (CES-D) Scale (Radloff, 1977). Standard item analyses, as well as Rasch (1960) analyses, both revealed item departures from unidimensionality in a sample of 2,455 older persons responding to the CES-D. Positive affect items in the scale performed poorly overall, their removal reducing the scale's bandwidth only slightly. Modeling depression scores derived from Rasch measures and raw totals showed subtle but important differences for statistical inference. The assessment of depressive risk was slightly enhanced by using 16-item scale measures obtained from the results of the Rasch analysis as the dependent variable. Confirmatory factor analysis and parallel analysis verified the advantages of removing positively worded items. IRT and CTT techniques proved to be complementary in this study and can be usefully combined to improve measuring depression.

Aged↗

Pairwise multiple comparison test procedures: an update for clinical child and adolescent psychologists.

Locating pairwise differences among treatment groups is a common practice of applied researchers. Articles published in this journal have addressed the issue of statistical inference within the context of an analysis of variance (ANOVA) framework, describing procedures for comparing means, among other issues. In particular, 1 article (Jaccard & Guilamo-Ramos, 2002b) presented some new methods of performing contrasts of means whereas another presented a framework for obtaining robust tests within this same context (Jaccard & Guilamo-Ramos, 2002a). The purpose of this article is to add to these contributions by presenting some newer methods for conducting pairwise comparisons of means, that is by extending the contributions of the first article and applying the framework of the second article to pairwise multiple comparisons. The newer methods are intended to provide additional sensitivity to detect treatment group differences and provide tests that are robust to the effects of variance heterogeneity, nonnormality, or both.

Adolescent↗

Mental, neuropsychic, and brain patterns of defense. Neuropsychic defense continua from psychopathology to the particularly human parallel networks that are problematic for artificial intelligence.

The synthesis of psychodynamic and neuroscientific data may be advanced by the study of similarly configured specific patterns of defense along continua from the normal and neurotic mental mechanisms of defenses, through what might be termed neuropsychiatric defenses influenced by the neurological state, to the more clearly neurological cortical reactions. These similar patterns and continua are the particularly human way brain adds its flavor of organism to mind, and the way in which all the psychodynamic mechanisms are specifically imbedded in brain. A purely psychological psychodynamics lacking this medical perspective fails to model human mental processes and reflects problems with artificial intelligence models that do not model brain. The pathological reactions that we address in psychiatry and neuropsychiatry interfere with and thereby reveal parallel processing features in (1) analogical and metaphorical thinking; (2) reduplication, redundancy, and repetitiveness; (3) self, person, and environmental recognition, including transference processes; (4) approximation and statistical inference; (5) spatiality and motor control; (6) affectivity, perception of sensations and sexuality; (7) projective mechanisms; and (8) problem solving by optimization, vectorial summation, and the mutual interaction of interconnected agencies. The new concepts in artificial intelligence built upon brainlike structures called neural nets portend great economies, emergent natural features that are more like human processes, and hierarchical and cyclic organizational demands. Medical psychoanalysts who comprehend dynamic and brain mechanisms and can describe them in terms that refer to both domains and their interaction should find increasing theoretical and practical convergence of their work with the emerging study and modeling of these processes by new forms of computers.

Adult↗

Using temporally spaced sequences to simultaneously estimate migration rates, mutation rate and population sizes in measurably evolving populations.

We present a Bayesian statistical inference approach for simultaneously estimating mutation rate, population sizes, and migration rates in an island-structured population, using temporal and spatial sequence data. Markov chain Monte Carlo is used to collect samples from the posterior probability distribution. We demonstrate that this chain implementation successfully reaches equilibrium and recovers truth for simulated data. A real HIV DNA sequence data set with two demes, semen and blood, is used as an example to demonstrate the method by fitting asymmetric migration rates and different population sizes. This data set exhibits a bimodal joint posterior distribution, with modes favoring different preferred migration directions. This full data set was subsequently split temporally for further analysis. Qualitative behavior of one subset was similar to the bimodal distribution observed with the full data set. The temporally split data showed significant differences in the posterior distributions and estimates of parameter values over time.

Biological Evolution↗

Multidimensional protein identification technology: current status and future prospects.

Protein profiling using high-throughput tandem mass spectrometry has become a powerful method for analyzing changes in global protein expression patterns in cells and tissues as a function of developmental, physiologic and disease processes. This review summarizes the utility and practical application of multidimensional protein identification technology as a platform for comprehensive proteomic profiling of complex biologic samples. The strengths and potential problems and limitations associated with this powerful technology are discussed, with an emphasis placed on one of the biggest challenges currently facing large-scale expression profiling projects -- namely, data analysis. Complementary bioinformatic computational data mining strategies, such as clustering, functional annotation and statistical inference, are also discussed as these are increasingly necessary for interpreting the results of global proteomic profiling studies.

Animals↗

[Risk factors and predictors of induced abortion: a population-based study].

This study aimed to identify key risk factors and predictors of induced abortion. A cross-sectional population-based study was conducted with a representative sample of 3,002 women 15 to 49 years of age in southern Brazil, randomly assigned to answer questions on induced abortion using either the ballot-box method or the indirect questioning method. Socioeconomic, demographic, and reproductive data were obtained through a pre-coded questionnaire. Data analysis used epidemiological statistical inferences and Bayes' theorem to calculate a posteriori probability. Induced abortion was strongly associated with fetal loss for all age groups. In adolescents, the main predictors were low socioeconomic level, low schooling, elevated school drop-out, and knowledge of a large number of contraceptive methods. For all other women, socioeconomic characteristics and skin color were not associated with abortion. For women aged 20 to 49 years, marital status and reproductive characteristics, including knowledge of contraceptive methods, were the most frequent risk factors and predictors of induced abortion.

Abortion, Induced↗

Spatial and temporal variation in the mosquitoes (Diptera: Culicidae) inhabiting waste tires in Nicholas County, West Virginia.

Larvae of 12 mosquito species were collected from abandoned tire piles at peridomestic and forested sites in Nicholas County, WV, from March through November of 2001. No larvae were found in March, but the numbers of species increased to 10 by July and remained relatively constant, at 9-11 in any given month, throughout November. Larvae of Ochlerotatus triseriatus (Say), the most commonly encountered species in every month of collection, were significantly more likely to be found in forested tire pile sites. Conversely, Culex restuans Theobald, Anopheles punctipennis (Say), Cx. territans Walker, and Aedes albopictus (Skuse) larvae were significantly more likely to be found in peridomestic tire piles. Larvae of the remaining seven species were either found in equal proportions at peridomestic and woodland sites, or there were too few collections to make statistical inferences. Opportunities for competitive interactions between Ae. albopictus and Oc. triseriatus in Nicholas County would be minimized because the peak occurrence of the two species differ temporally and spatially.

Animals↗

Internal quality control of radioimmunoassays: monitoring of error.

A cumulative sum technique has been specially designed to monitor the error between replicate determinations made on quality control plasma for consecutive batches of assays. This procedure has played a vital role in assessing assay performance. Special consideration has been given to small sample sizes (n = 2 or 3) which is generally the rule rather than the exception in many situations. This technique has been applied to numerous steroid radioimmunoassays and has ensured that both the mean value and the standard error of hormone levels of a quality control pool were under control. Data from routine assays of oestriol and testosterone in plasma from women are presented. Since this technique provides a sensitive measure of monitoring error, it assists the endocrinologist in elucidating statistical inferences which are a manifestation of assay performance.

Estriol↗

Predicting abundance of desert riparian birds: validation and calibration of the Effective Area Model.

Reliable prediction of the effects of landscape change on species abundance is critical to land managers who must make frequent, rapid decisions with long-term consequences. However, due to inherent temporal and spatial variability in ecological systems, previous attempts to predict species abundance in novel locations and/or time frames have been largely unsuccessful. The Effective Area Model (EAM) uses change in habitat composition and geometry coupled with response of animals to habitat edges to predict change in species abundance at a landscape scale. Our research goals were to validate EAM abundance predictions in new locations and to develop a calibration framework that enables absolute abundance predictions in novel regions or time frames. For model validation, we compared the EAM to a null model excluding edge effects in terms of accurate prediction of species abundance. The EAM outperformed the null model for 83.3% of species (N=12) for which it was possible to discern a difference when considering 50 validation sites. Likewise, the EAM outperformed the null model when considering subsets of validation sites categorized on the basis of four variables (isolation, presence of water, region, and focal habitat). Additionally, we explored a framework for producing calibrated models to decrease prediction error given inherent temporal and spatial variability in abundance. We calibrated the EAM to new locations using linear regression between observed and predicted abundance with and without additional habitat covariates. We found that model adjustments for unexplained variability in time and space, as well as variability that can be explained by incorporating additional covariates, improved EAM predictions. Calibrated EAM abundance estimates with additional site-level variables explained a significant amount of variability (P < 0.05) in observed abundance for 17 of 20 species, with R2 values >25% for 12 species, >48% for six species, and >60% for four species when considering all predictive models. The calibration framework described in this paper can be used to predict absolute abundance in sites different from those in which data were collected if the target population of sites to which one would like to statistically infer is sampled in a probabilistic way.

Animals↗

Advantages of using the net-benefit approach for analysing uncertainty in economic evaluation studies.

No consensus has yet been reached on how to analyse uncertainty in economic evaluation studies where individual patient data are available for costs and health effects. This paper summarises the available results regarding the analysis of uncertainty on the cost-effectiveness plane and argues for using the net-benefit approach when analysing uncertainty in cost-effectiveness studies. The net-benefit approach avoids the interpretation and statistical problems related to the incremental cost effectiveness ratio and implies several advantages. First, traditional statistical methods can be used for confidence-interval estimation and hypothesis testing. Second, calculation of the optimal sample size and the power of the study are facilitated allowing the correlation between costs and effects to vary within and between patient groups. Third, the use of a Bayesian approach to cost-effectiveness analysis is facilitated. Fourth, a formal relation between cost-effectiveness acceptability curves and statistical inference is provided. Finally, the net-benefit approach gives the Fieller's limits of the confidence interval for the incremental cost-effectiveness ratio in the cost-effectiveness plane. Based on these advantages the net-benefit approach should strongly be considered when analysing uncertainty in cost-effectiveness analyses.

Bayes Theorem↗

Glycoprotein IIb/IIIa inhibitors in patients undergoing percutaneous mechanical intervention for acute myocardial infarction.

Limitations in study designs and adoption of rigid criteria for randomization in clinical trials on acute myocardial infarction (AMI) may result in the enrollment of artificial populations, and subsequent trial results may be misleading in many ways rendering problematic the generalization of the trial results to a "real world population" of AMI. Furthermore, the "frequentist" approach in study designs with inclusion of thousands of low-risk patients and high statistical inference have produced inconclusive or negative results despite the high potential for a strong impact on outcome of the study drug, or device, or strategy. The investigative approaches to the use of IIb/IIIa inhibitors as adjunctive treatment to primary coronary intervention (PCI) for AMI are a clear instance of this crucial problem in the evidence-based medicine era. Five concluded randomized trials comparing abciximab with placebo in patients undergoing primary PCI for AMI have produced different and conflicting results with a broad spectrum of possibilities. No benefit of the drug in patients receiving infarct artery stenting, benefit of the drug only in patients undergoing conventional balloon angioplasty with provisional stenting and limited to the early phase, and mainly driven by the decrease in the need for urgent target vessel revascularization, benefit in terms of decreased mortality, reinfarction and target vessel revascularization at 1-month but not maintained at 6 months, long-term benefit in the composite of death, reinfarction and target vessel revascularization, improved early and late outcome including long-term survival. This review of the trials of abciximab and other IIb/IIIa inhibitors in patients undergoing PCI for AMI tries to put the studies and their results into a proper perspective for the correct use of adjunctive IIb/IIIa inhibitor use in patients with AMI.

Abciximab↗

Classifying gene expression profiles from pairwise mRNA comparisons.

We present a new approach to molecular classification based on mRNA comparisons. Our method, referred to as the top-scoring pair(s) (TSP) classifier, is motivated by current technical and practical limitations in using gene expression microarray data for class prediction, for example to detect disease, identify tumors or predict treatment response. Accurate statistical inference from such data is difficult due to the small number of observations, typically tens, relative to the large number of genes, typically thousands. Moreover, conventional methods from machine learning lead to decisions which are usually very difficult to interpret in simple or biologically meaningful terms. In contrast, the TSP classifier provides decision rules which i) involve very few genes and only relative expression values (e.g., comparing the mRNA counts within a single pair of genes); ii) are both accurate and transparent; and iii) provide specific hypotheses for follow-up studies. In particular, the TSP classifier achieves prediction rates with standard cancer data that are as high as those of previous studies which use considerably more genes and complex procedures. Finally, the TSP classifier is parameter-free, thus avoiding the type of over-fitting and inflated estimates of performance that result when all aspects of learning a predictor are not properly cross-validated.

Journal Article↗

Behçet's disease: a review and a report of 12 cases from Sweden.

In a retrospective study of 12 patients with Behçet's disease, more than half were found to originate from the Near East, where the prevalence of the disease is known to be high. The immigrant patients were all males, whereas 3 of the 5 patients with Swedish ancestry were females. Certain differences emerged between the two groups, including different sex ratio and absence of HLA B5 association and pathergy skin reaction among the Swedish patients. Moreover, serious neurological and ocular symptoms showing no tendency to recede with age afflicted all the Swedish female patients. Urogenital symptoms were, besides ulcers, common in both groups, including prostatitis, urethritis, orchitis, chronic sterile cystitis and relapsing salpingitis. Although the maternal does not allow statistical inferences, the estimated prevalence was higher than expected among both Swedish and immigrant patients. Recent studies, including the diagnostic criteria proposed by the "International Study Group for Behçet's disease", are discussed in relation to previously used criteria as well as present findings. The sensitivity and specificity of the first mentioned criteria and the ones proposed by Mason & Barnes seemed equal.

Adult↗

Weighted specific-category kappa measure of interobserver agreement.

When two observers classify a sample of items using the same categorical scale, and when different disagreements are differentially weighted, the weighted Kappa (Kw) by Cohen may serve as a measure of interobserver agreement. We propose a Kappa-based weighted measure (K(ws)) of agreement on some specific category s, with Kw being a weighted average of all K(ws)s. Therefore, while Cohen's Kw is a summary measure of the overall agreement, the proposed K(ws) provides a measure of the extent to which the observers agree on the specific categories, with both measures being suitable for ordinal categories because of the weights being used. Statistical inferences for K(ws) and its unweighted counterpart are also discussed. A numerical example is provided.

Humans↗