Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Bayesian hypothesis testing of four-taxon topologies using molecular sequence data.

The reconstruction of phylogenetic trees from molecular sequences presents unusual problems for statistical inference. For example, three possible alternatives must be considered for four taxa when inferring the correct unrooted tree (referred to as a topology). In our view, classical hypothesis testing is poorly suited to this triangular set of alternative hypotheses. In this article, we develop Bayesian inference to determine the posterior probability that a four-taxon topology is correct given the sequence data and the evolutionary parsimony algorithm for phylogenetic reconstruction. We assess the frequency properties of our models in a large simulation study. Bayesian inference under the principles of evolutionary parsimony is shown to be well calibrated with reasonable discriminating power for a wide range of realistic conditions, including conditions that violate the assumptions of evolutionary parsimony.

Base Sequence↗

Testing for treatment effect in the presence of regression toward the mean.

We are often faced with the statistical problem of evaluating the effect of a treatment in the extreme of a population. This requires taking measurements on truncated random variables and, hence, it becomes necessary to take proper account of the effect of regression toward the mean. The usual statistical procedures are inappropriate for testing treatment effect in the presence of regression toward the mean. Likelihood ratio and score tests based on truncated distributions should provide valid statistical inferences in these situations. We conducted simulation studies to investigate the properties of these methods and found that the likelihood ratio test performs well even when the sample size is moderate, whereas the score test does not seem to control the nominal significance level. We compared the likelihood ratio test to a regression-based t-test, assuming the mean of the baseline distribution to be known, and found the likelihood ratio test more powerful. In the case where the baseline mean is unknown, we also investigated Wald's test and compared it with the likelihood ratio test and score test with respect to validity and power using simulation. Wald's test and the score test do not control the nominal significance level unless the sample size is extremely large. Overall, the likelihood ratio test has the best performance among all the methods studied. The proposed likelihood ratio test is illustrated using an example of a cholesterol study.

Biometry↗

Prevalence of rheumatoid arthritis in circumpolar native populations.

OBJECTIVE: To compare the prevalence of rheumatoid arthritis (RA) in related, but geographically separate, indigenous circumpolar populations. METHODS: Cases were identified by community survey in Russia and by examination of cases located through arthritis registries, a computerized patient information database, and query of local health care providers in Alaska. All possible cases were verified by examination and application of the American College of Rheumatology 1987 criteria. RESULTS: The prevalence rates of RA (age standardized to US population of 1980) varied from 0.62% in the Alaskan Yupik to 1.78% in the Alaskan Inupiat. The Russian Chukchi rate was 0.73% and that of the Siberian Eskimo was 1.42%. CONCLUSION: The Alaskan Yupik Eskimo and Chukchi natives had prevalence rates of RA within the usual range of North American Caucasian groups, in contrast to the Russian Siberian Eskimo and the Alaskan Inupiat Eskimo of the Barrow region, whose high rates approached those of unrelated North American native groups living in very different environments. The Alaskan Inupiat rate was significantly higher than that of the Alaskan Yupik (OR = 2.51, 95% CI 1.25-5.07; p = 0.013), but statistical inferences are limited in the Russian study populations by the small case numbers. The high prevalence rates probably have a genetic basis, although an environmental influence cannot be excluded.

Adult↗

Adjusting sample size for anticipated dropouts in clinical trials.

Statistical models for calculating sample sizes for controlled clinical trials often fail to take into account the negative impact that dropouts have on the power of intent-to-treat analyses. Empirically defined dropout correction coefficients are proposed to adjust sample sizes for endpoint analysis of variance (ANOVA) and analysis of covariance (ANCOVA) that have been initially calculated assuming complete data. The implications of type of analysis (change-score ANOVA or ANCOVA), correlational structure of the repeated measurements (compound symmetry or autoregressive), and percentage of dropouts (20% or 30%) are considered, together with other less influential design and data parameters. We recommend the use of ANCOVA to correct for baseline differences and for time-in-study if there is a nonspecific change across time. Given a realistic autoregressive (order 1) correlational structure for the repeated measurements and a proposed endpoint ANCOVA, the empirical results support the common practice of increasing calculated sample size by the anticipated number of dropouts. The previous rationale has been to retain a requisite number of "completers" on which to base statistical inferences. We believe the present results provide the first documentation of the relevance of that strategy for intent-to-treat analyses in which the incomplete data for dropouts must be included. Based on comparative power analyses, the strategy also seems appropriate for maintaining the power of mixed-model regression analyses, simple regression on a normalized time scale, and analyses of trends fitted to imputed scores for dropouts.

Clinical Trials as Topic↗

A novel tumor suppressor locus on chromosome 18q involved in the development of human lung cancer.

The high incidence of loss of heterozygosity (LOH) on chromosome 18q in advanced non-small cell lung carcinomas indicates the presence of tumor suppressor gene(s) on this chromosome arm, which plays an important role in the acquisition of malignant phenotypes in lung cancers. In the present study, we examined 62 lung cancer specimens and 54 lung cancer cell lines for allelic imbalance at 11 microsatellite loci to define common regions of 18q deletions. Allelic imbalance of 18q was detected in 24 (55.8%) non-small cell lung carcinoma specimens and in 6 (31.6%) small cell lung carcinoma specimens, whereas a similar frequency of LOH was statistically inferred to occur in cell lines by analyzing marker homozygosity as an indirect measure of LOH. Five specimens and 11 cell lines showed partial or interstitial deletions of chromosome 18q, and 2 of them had homozygous deletions at the 18q21.1 region. A commonly deleted region was assigned between the D18S46 and y953G12R loci. The size of this region is less than 1 Mb, and the coding exons of three candidate tumor suppressor genes, Smad2, Smad4, and DCC, were mapped outside the region. This result suggests that the common region harbors a novel tumor suppressor gene involved in the progression of lung cancer.

Carcinoma, Non-Small-Cell Lung↗

Assessing the sensitivity of regression results to unmeasured confounders in observational studies.

This paper presents a general approach for assessing the sensitivity of the point and interval estimates of the primary exposure effect in an observational study to the residual confounding effects of unmeasured variable after adjusting for measured covariates. The proposed method assumes that the true exposure effect can be represented in a regression model that includes the exposure indicator as well as the measured and unmeasured confounders. One can use the corresponding reduced model that omits the unmeasured confounder to make statistical inferences about the true exposure effect by specifying the distributions of the unmeasured confounder in the exposed and unexposed groups along with the effects of the unmeasured confounder on the outcome variable. Under certain conditions, there exists a simple algebraic relationship between the true exposure effect in the full model and the apparent exposure effect in the reduced model. One can then estimate the true exposure effect by making a simple adjustment to the point and interval estimates of the apparent exposure effect obtained from standard software or published reports. The proposed method handles both binary response and censored survival time data, accommodates any study design, and allows the unmeasured confounder to be discrete or normally distributed. We describe applications on two major medical studies.

Appetite Depressants↗

A statistical problem for inference to regulatory structure from associations of gene expression measurements with microarrays.

MOTIVATION: One approach to inferring genetic regulatory structure from microarray measurements of mRNA transcript hybridization is to estimate the associations of gene expression levels measured in repeated samples. The associations may be estimated by correlation coefficients or by conditional frequencies (for discretized measurements) or by some other statistic. Although these procedures have been successfully applied to other areas, their validity when applied to microarray measurements has yet to be tested. RESULTS: This paper describes an elementary statistical difficulty for all such procedures, no matter whether based on Bayesian updating, conditional independence testing, or other machine learning procedures such as simulated annealing or neural net pruning. The difficulty obtains if a number of cells from a common population are aggregated in a measurement of expression levels. Although there are special cases where the conditional associations are preserved under aggregation, in general inference of genetic regulatory structure based on conditional association is unwarranted

Algorithms↗

Difference to Inference: teaching logical and statistical reasoning through on-line interactivity.

Difference to Inference is an on-line JAVA program that simulates theory testing and falsification through research design and data collection in a game format. The program, based on cognitive and epistemological principles, is designed to support learning of the thinking skills underlying deductive and inductive logic and statistical reasoning. Difference to Inference has database connectivity so that game scores can be counted as part of course grades.

Decision Making↗

ELISA: structure-function inferences based on statistically significant and evolutionarily inspired observations.

UNLABELLED: The problem of functional annotation based on homology modeling is primary to current bioinformatics research. Researchers have noted regularities in sequence, structure and even chromosome organization that allow valid functional cross-annotation. However, these methods provide a lot of false negatives due to limited specificity inherent in the system. We want to create an evolutionarily inspired organization of data that would approach the issue of structure-function correlation from a new, probabilistic perspective. Such organization has possible applications in phylogeny, modeling of functional evolution and structural determination. ELISA (Evolutionary Lineage Inferred from Structural Analysis, http://romi.bu.edu/elisa) is an online database that combines functional annotation with structure and sequence homology modeling to place proteins into sequence-structure-function "neighborhoods". The atomic unit of the database is a set of sequences and structural templates that those sequences encode. A graph that is built from the structural comparison of these templates is called PDUG (protein domain universe graph). We introduce a method of functional inference through a probabilistic calculation done on an arbitrary set of PDUG nodes. Further, all PDUG structures are mapped onto all fully sequenced proteomes allowing an easy interface for evolutionary analysis and research into comparative proteomics. ELISA is the first database with applicability to evolutionary structural genomics explicitly in mind. AVAILABILITY: The database is available at http://romi.bu.edu/elisa.

Amino Acid Sequence↗

Statistical properties and inference of the antimicrobial MIC test.

A common method for measuring the drug-specific minimum inhibitory concentration (MIC) of an antibacterial agent is via a two-fold broth dilution test known as the MIC test. Because this procedure implicitly rounds data upward, inference based on unadjusted measurements is biased and overestimates bacterial resistance to a drug. We detail this test procedure and its associated bias, which, in many cases, has an expected value of approximately 0.5 on the log(2) scale. In addition, new bias-corrected estimates of resistance are proposed. A numeric example is used to illustrate the extent to which the traditional resistance estimate can overestimate the true proportion of resistant strains, a phenomenon which is remedied by using the proposed estimates.

Bias↗

Genetic Interaction Motif Finding by expectation maximization--a novel statistical model for inferring gene modules from synthetic lethality.

BACKGROUND: Synthetic lethality experiments identify pairs of genes with complementary function. More direct functional associations (for example greater probability of membership in a single protein complex) may be inferred between genes that share synthetic lethal interaction partners than genes that are directly synthetic lethal. Probabilistic algorithms that identify gene modules based on motif discovery are highly appropriate for the analysis of synthetic lethal genetic interaction data and have great potential in integrative analysis of heterogeneous datasets. RESULTS: We have developed Genetic Interaction Motif Finding (GIMF), an algorithm for unsupervised motif discovery from synthetic lethal interaction data. Interaction motifs are characterized by position weight matrices and optimized through expectation maximization. Given a seed gene, GIMF performs a nonlinear transform on the input genetic interaction data and automatically assigns genes to the motif or non-motif category. We demonstrate the capacity to extract known and novel pathways for Saccharomyces cerevisiae (budding yeast). Annotations suggested for several uncharacterized genes are supported by recent experimental evidence. GIMF is efficient in computation, requires no training and automatically down-weights promiscuous genes with high degrees. CONCLUSION: GIMF effectively identifies pathways from synthetic lethality data with several unique features. It is mostly suitable for building gene modules around seed genes. Optimal choice of one single model parameter allows construction of gene networks with different levels of confidence. The impact of hub genes the generic probabilistic framework of GIMF may be used to group other types of biological entities such as proteins based on stochastic motifs. Analysis of the strongest motifs discovered by the algorithm indicates that synthetic lethal interactions are depleted between genes within a motif, suggesting that synthetic lethality occurs between-pathway rather than within-pathway.

Algorithms↗

Structure and energetics of channel-forming protein-polysaccharide complexes inferred via computational statistical thermodynamics.

The ion channel protein alpha-hemolysin (alphaHL) forms supramolecular complexes with the polysaccharide beta-cyclodextrin (betaCD). This system has potential uses in nanoscale device engineering. It has been found recently that betaCD formed longer- or shorter-lived complexes with some engineered alphaHL mutants then with a wild type protein (Gu et al. J. Gen. Physiol. 2001, 118, 481-493). However, how changes in the protein sequence affect complex lifetime was not completely understood in part due to the lack of knowledge of structures of these metastable complexes. In this paper, we present an extensive molecular modeling study of the betaCD-alphaHL and selected mutant complexes to gain insights into the betaCD-alphaHL interaction mechanisms and to predict possible structures and energetics of the complexes. Thermodynamic integration (TI) and umbrella sampling (US) techniques (with the weighted histogram analysis method (WHAM)) were used to calculate the relative binding affinities of the complexes formed with the wild type alphaHL and the M113N, M113E, M113A, and M113V mutants. Our results are in excellent agreement with experiment. While betaCD-M113N and betaCD-M113A complexes were stable in the configuration of the wild type complex, the equilibrium configuration of the betaCD-M113V and betaCD-M113E complexes was significantly different. In these cases, TI alone was insufficient to accurately calculate the corresponding free energy differences. By utilizing a TI/US combination in a novel manner, we were able to accurately calculate free energy changes in these flexible systems. The betaCD-M113A and betaCD-M113E complexes, which exhibited shorter lifetimes than other complexes in an experiment, in simulations exhibited greater flexibility and higher water solvation of the betaCD adapter. MD simulations of the betaCD-M113N complex with betaCD in a downward orientation were also performed.

Bacterial Proteins↗

Advantages of permutation (randomization) tests in clinical and experimental pharmacology and physiology.

1. The statistical procedures that are used most commonly in clinical and experimental pharmacology and physiology are designed to test for differences between two means. 2. The classical procedures for detecting such differences are those in which, under the population model of inference, the test statistic is referred to the t- or F-distributions. The validity of statistical inferences from these tests depends on a number of assumptions. Foremost among these is that the experimental groups have been constructed by taking random samples from defined populations. The statistical inferences then apply to the sampled populations. 3. In biomedical research this sampling process is seldom followed. Instead, samples are usually acquired by non-random selection, and are then divided by randomization into experimental groups. This being the case, it is theoretically invalid to use the classical t- or F-tests to analyse the experimental results. 4. The validity of inferences from the classical tests also depends on other assumptions, such as that the sampled populations are normal in form and of equal variance. It is difficult to be certain that these assumptions are fulfilled when group sizes are small, as they usually are in pharmacology and physiology. Breach of them, especially if the groups are unequal in size, can lead to serious statistical errors. 5. Exact permutation tests are designed to make statistical inferences under the randomization model. These conclusions apply only to the results of experiments actually performed. By permuting the statistic of interest, such as the difference between arithmetic means, geometric means, medians, mid-ranges or mean-ranks of randomized groups of observations, the probability is calculated that the observed difference or a more extreme one could have occurred by chance. This inferential process is consistent with the way most biomedical experiments are designed and conducted. 6. Exact permutation tests, or sampled permutation tests based on Monte Carlo random sampling of all possible permutations, can now be performed on personal computers. They are commended to biomedical investigators as being superior to the classical tests for analysing their experimental results when the central tendencies of two independent groups, or of two sets of measurements on the same group, are compared. 7. When there is doubt that the assumptions for t-tests are satisfied, investigators sometimes use non-parametric rank-order procedures such as the Wilcoxon-Mann-Whitney rank-sum test for independent groups or the Wilcoxon signed rank-sum test for paired observations.(ABSTRACT TRUNCATED AT 400 WORDS)

Clinical Trials as Topic↗

Stereological analysis of three-dimensional structure organization of surfaces in multiphase specimens: statistical methods and model-inferences.

In a multiphase material the structural components or phases are everywhere in contact with each other. The relative area of surface contact between various phases is an important aspect of the short-range ordering or organization of the structure. The stereological quantitation of such specific interfaces is a simple and well-known technique. The proper statistical definition of realistic models for the frequency of contact and the quantitative estimation of phase-specific affinities is studied. The meaningful interpretation of sets of estimated affinities poses a major problem of statistical inference which is dealt with in detail and illustrated by a worked-out biological example.

Animals↗

Mixed-effects models for the evaluation of long-term trends in exposure levels with an example from the nickel industry.

Longitudinal studies play an important role in evaluating the temporal behavior of occupational exposures. The purpose of this paper is to examine certain features of longitudinal data and to present a general conceptual framework by which these features may be taken into account so that statistically valid inferences can be made. Statistical methods that rely on the application of mixed-effects models are proposed for evaluating long-term trends in exposures to workplace contaminants. The mixed-effects model presented herein has fixed effects for trend components and random effects for workers, job groups, buildings and plants. These models differ from conventional techniques in that they accommodate hierarchically structured data and account for the correlation that may arise due to the clustering of measurements based on when and where the data were collected. While primary interest is focused on determining the magnitude of trends in exposure levels over time, the model also provides information about the magnitude of the sources of variation associated with different groupings of workers. Application of the mixed-effects model is illustrated with a large database of shift-long personal exposure measurements collected on workers exposed to nickel aerosols in the nickel-producing industry.

Cluster Analysis↗

Comparison of statistical models for analyzing genotype, inferred haplotype, and molecular haplotype data.

This report compares statistical models based on molecular and inferred haplotypes of the human paraoxonase-1 gene (PON1). In a study of 402 women comprising three race/ethnicities, 137 women had ambiguous inferred haplotypes. The inferred haplotypes (the one with highest posterior probability) for 20 of these women differed from molecular haplotypes, while based on the posterior distribution from the imputation method, 30 discrepancies were expected. We examined the proportion of the variance in PON1 enzymatic activity (phenotype) explained by genotype, and by inferred and molecular haplotype information. For Caucasians, there was an improvement in adjusted R(2) from 16% for the genotype count model, to 29% for imputed haplotypes, and a further improvement to 33% for molecular haplotypes. For Hispanics and African-Americans, there was no indication that haplotypes helped in explaining PON1 activity, and the imputed model gave essentially the same R(2) as the molecular model. For African-Americans, none of the models had adjusted R(2) that exceeded 4%, while for Hispanics they were all about 21-22%. We propose a new parsimonious model which uses all the genotype information and selected haplotype information. For PON1, this model achieves essentially the same adjusted R(2) as the all-haplotype model, with a potential cost savings and without giving the extreme predictions for uncommon haplotype combinations that the all-haplotype models provides.

Aryldialkylphosphatase↗