Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

The analysis of placement values for evaluating discriminatory measures.

The idea of using measurements such as biomarkers, clinical data, or molecular biology assays for classification and prediction is popular in modern medicine. The scientific evaluation of such measures includes assessing the accuracy with which they predict the outcome of interest. Receiver operating characteristic curves are commonly used for evaluating the accuracy of diagnostic tests. They can be applied more broadly, indeed to any problem involving classification to two states or populations (D= 0 or 1). We show that the ROC curve can be interpreted as a cumulative distribution function for the discriminatory measure Y in the affected population (D= 1) after Y has been standardized to the distribution in the reference population (D= 0). The standardized values are called placement values. If the placement values have a uniform(0, 1) distribution, then Y is not discriminatory, because its distribution in the affected population is the same as that in the reference population. The degree to which the distribution of the standardized measure differs from uniform(0, 1) is a natural way to characterize the discriminatory capacity of Y and provides a nontraditional interpretation for the ROC curve. Statistical methods for making inference about distribution functions therefore motivate new approaches to making inference about ROC curves. We demonstrate this by considering the ROC-GLM regression model and observing that it is equivalent to a regression model for the distribution of placement values. The likelihood of the placement values provides a new approach to ROC parameter estimation that appears to be more efficient than previously proposed methods. The method is applied to evaluate a pulmonary function measure in cystic fibrosis patients as a predictor of future occurrence of severe acute pulmonary infection requiring hospitalization. Finally, we note the relationship between regression models for the mean placement value and recently proposed models for the area under the ROC curve which is the classic summary index of discrimination.

Biometry↗

The complete nucleotide sequence of the chicken ovotransferrin mRNA.

The nucleotide sequence of an almost double-stranded cDNA copy [Cochet, M., Perrin, F., Gannon, F., Krust, A., Chambon, P., McKnight, G. S., Lee, D. C., Mayo, K. E., and Palmiter, R. D. (1979) Nucleic Acids Res. 6, 2435-2452] of chicken ovotransferrin (conalbumin) mRNA has been determined. Taking into account the previously reported 5'-end sequence [Cochet, M., Gannon, F., Hen, R., Maroteaux, L., Perrin, F., and Chambon, P. (1979) Nature (Lond.) 282, 567-574] we present the complete nucleotide sequence of the ovotransferrin mRNA from which the amino acid sequence of the protein is inferred. A computer and statistical analysis of the nucleotide sequence reveals a pattern of internal homology which confirms that the present-day chicken ovotransferrin gene (and by extrapolation the transferrin genes of other species) has evolved by duplication and gives some support to the quadruplication hypothesis of transferrin evolution.

Animals↗

Challenges and Opportunities in Analyzing Cancer-Associated Microbiomes.

The study of cancer-associated microbiomes has gained significant attention in recent years, spurred by advances in high-throughput sequencing and metagenomic analysis. Microbiome research holds promise for identifying noninvasive biomarkers and possibly new paradigms for cancer treatment. In this review, we explore the key computational challenges and opportunities in analyzing cancer-associated microbiomes (in tumor/normal tissues and other body sites, e.g., gut, oral, and skin), focusing on sequencing-driven strategies and associated considerations for taxonomic and functional characterization. The discussion covers the strengths and limitations of current analysis tools for identifying contamination, determining compositional bias, and resolving species and strains, as well as the statistical, metabolic, and network inferences that are essential to uncover host-microbiome interactions. Several key considerations are required to guide the choice of databases used for metagenomic analysis in such studies. Recent advances in spatial and single-cell technologies have provided insights into cancer-associated microbiomes, and Artificial Intelligence-driven protein function prediction might enable rapid advances in this field. Finally, we provide a perspective on how the field can evolve to manage the ever-growing size of datasets and generate robust and testable hypotheses. This article is part of a special series: Driving Cancer Discoveries with Computational Research, Data Science, and Machine Learning/AI .

Humans↗

Impact of missing genotype data on Monte-Carlo simulation based haplotype analysis.

In the context of haplotype association analysis of unphased genotype data, methods based on Monte-Carlo simulations are often used to compensate for missing or inappropriate asymptotic theory. Moreover, such methods are an indispensable means to deal with multiple testing problems. We want to call attention to a potential trap in this usually useful approach: The simulation approach may lead to strongly inflated type I errors in the presence of different missing rates between cases and controls, depending on the chosen test statistic. Here, we consider four different testing strategies for haplotype analysis of case-control data. We recommend to interpret results for data sets with non-comparable distributions of missing genotypes with special caution, in case the test statistic is based on inferred haplotypes per individual. Moreover, our results are important for the conduction and interpretation of genome-wide association studies.

Case-Control Studies↗

Identifying rate-limiting nodes in large-scale cortical networks for visuospatial processing: an illustration using fMRI.

With the advent of functional neuroimaging techniques, in particular functional magnetic resonance imaging (fMRI), we have gained greater insight into the neural correlates of visuospatial function. However, it may not always be easy to identify the cerebral regions most specifically associated with performance on a given task. One approach is to examine the quantitative relationships between regional activation and behavioral performance measures. In the present study, we investigated the functional neuroanatomy of two different visuospatial processing tasks, judgement of line orientation and mental rotation. Twenty-four normal participants were scanned with fMRI using blocked periodic designs for experimental task presentation. Accuracy and reaction time (RT) to each trial of both activation and baseline conditions in each experiment was recorded. Both experiments activated dorsal and ventral visual cortical areas as well as dorsolateral prefrontal cortex. More regionally specific associations with task performance were identified by estimating the association between (sinusoidal) power of functional response and mean RT to the activation condition; a permutation test based on spatial statistics was used for inference. There was significant behavioral-physiological association in right ventral extrastriate cortex for the line orientation task and in bilateral (predominantly right) superior parietal lobule for the mental rotation task. Comparable associations were not found between power of response and RT to the baseline conditions of the tasks. These data suggest that one region in a neurocognitive network may be most strongly associated with behavioral performance and this may be regarded as the computationally least efficient or rate-limiting node of the network.

Adolescent↗

[Linking quantitation of electrophoresis pattern and data analysis in AFLP for Oncomelania hupensis].

OBJECTIVE: To search into a method for analyzing the quantitative data in amplified fragment length polymorphism (AFLP) electrophoresis. METHODS: Oncomelania snails collected from the field were screened. Forty snails found uninfected with schistosomiasis were divided randomly into two groups and used to isolate genomic DNA. AFLP electrophoresis pattern was first transformed into quantitative data by Glyko BandScan software, and the bands were read according to different standards of band-reading to acquire the corresponding data. These data sets were analyzed by genetic statistics to get an inference set, and the analysis of this inference set was performed to reach a summary description. RESULTS: The results of genetic variation from different standards of band-reading were different With the increase of the standard value of band-reading, the indices indicating the genetic polymorphism of Oncomelania hupensis population (e.g. Shannon's information index) also increased. When the standard value reached at certain level, the values of these indices began to decrease. Compared with the above indices, the change for gene flow turned out contrary to the genetic identity. The distributions of inference results from different standards of band-reading all showed significant normal distribution. The mean value of genetic variation based on total grey was very close to that on the proportion of total grey. The average genetic identity between the "subpopulations" was 0.956 according to proportion of total grey or 0.958 from the total grey with an average genetic distance between the "subpopulations" of 0.045 and 0.043 respectively. CONCLUSION: It seems to be a reasonable and accurate method by quantifying the AFLP electrophoresis pattern followed by analyzing the data through the use of the different standards of band-reading.

Animals↗

[Algebraic regularities of the ratio of heterozygosity with a mean and variance of quantitative character].

The relationships of heterozygosity with the mean and variance of quantitative character were considered under neutrality, additivity and overdominance of polymorphic loci. Attention was drawn to dependence of the patterns of relationships on the number of polymorphic loci (which varied from 1 to 10) and on the type of polymorphic loci, both homogeneous (polymorphic loci are of the same type) and heterogeneous (polymorphic loci are of the two types) samples of 10 polymorphic loci and their combination. It is shown that increase in the number of polymorphic loci is accompanied with extension of the limits of corresponding relations, whereas the patterns of these relations depend on the type of connection of separate polymorphic locus with the quantitative character and on the ratio of different loci in the set of polymorphic loci. It is assumed that the relationship of heterozygosity with quantitative characters, displayed in the number of experimental works, may contain a component mediated by similar statistical effects. It is inferred that the discrepancy between different authors' experimental data on existence or lack of some relationships between multiplicative heterozygosity and morphological variability of quantitative characters can be explained by different types of relations of polymorphic loci to quantitative characters encountered in their works.

Alleles↗

Glucose metabolism, lipidic pattern, apolipoprotein A and B during hemodialysis with cuprophan and hemodialysis-hemoperfusion.

We studied differences in the glucose metabolism, lipid pattern and apolipoprotein A and B after one month of combined hemodialysis-hemoperfusion (HD-HP) in eight regular maintenance hemodialysis (HD) patients. After one month of HD-HP predialytic serum creatinine and phosphate were lower than during the preceding HD period; the lipidic pattern and apolipoproteins were unchanged. A single HD-HP session significantly reduced triglycerides; this does not occur during HD. Postdialytic changes of all other lipoprotein metabolism parameters are identical in the two techniques. As regards glucose metabolism, after one month of regular HD-HP treatment glycemic curves during the intravenous glucose tolerance test (IVGTT) were perfectly matched with those in HD. However, insulin production was lower, with a reduction of the insulin resistance index. Differences did not reach statistical significance. The Authors infer that activated charcoal, even though it achieves better blood purification in uremia, may be unable to remove specific substances which hamper some key enzyme activities of the carbohydrate and lipid metabolism.

Adult↗

New inferences from tree shape: numbers of missing taxa and population growth rates.

The relative positions of branching events in a phylogeny contain information about evolutionary and population dynamic processes. We provide new summary statistics of branching event times and describe how these statistics can be used to infer rates of species diversification from interspecies trees or rates of population growth from intraspecies trees. We also introduce a phylogenetic method for estimating the level of taxon sampling in a clade. Different evolutionary models and different sampling regimes can produce similar patterns of branching events, so it is important to consider explicitly the model assumptions involved when making evolutionary inferences. Results of an analysis of the phylogeny of the mosquito-borne flaviviruses suggest that there could be several thousand currently unidentified viruses in this clade.

Animals↗

An empirical Bayesian significance test of cDNA library data.

Automated high-throughput sequencing of cDNA clones from numerous libraries has generated a wealth of information about both genome sequence and relative transcript abundances. A common statistical challenge in the analysis of library sequences is to infer whether there is differential expression for the same transcript under two different conditions, such as normal and diseased tissue. In contrast to the continuously variable intensity measurements from microarray experiments, data from cDNA library sequencing presents itself as a discrete count of the incidence of some clone or transcript in a finite sample. In this paper, we first propose a statistical model for data generated from cDNA library sequencing efforts. The model is based on the Poisson mixed with generalized inverse Gaussian (PGIG), introduced by Sichel (1971, 1975). PGIG has been used in modeling population abundance, ecological studies, word frequencies in publications, etc. Using data from the literature, we show that the proposed model provides a good fit to the observed data. Using this new model for cDNA library data, we developed an empirical Bayesian significance test (EBST) for inferring the statistical significance of differential gene expression from discrete data.

Bayes Theorem↗

Linear regression modeling to compare fluoride release profiles of various restorative materials.

OBJECTIVES: The aim of this study was to compare the released fluoride profiles of various restorative materials by using linear regression analysis. METHODS: Specimens were prepared using a cylindrical Teflon mold with a height of 2 mm and a radius of 8 mm. After being prepared, specimens were immediately placed into artificial saliva which was replaced at various times during 6 weeks. These released intrinsic fluoride amounts were measured by using an ion selective electrode. Then, data obtained cumulatively were statistically analyzed, and the released profiles were compared. RESULTS: It was observed that the materials released fluoride at different levels of concentration and the largest fluoride release was obtained from the conventional glass ionomer cement. This was followed by resin modified glass ionomer cement, polyacid modified composite resin, and fluoride releasing composite resin, respectively. Although the released fluoride amounts of the materials were different, their release profiles were found to be similar in that the release was initially fast and then it became steady as time passed. SIGNIFICANCE: The statistical modeling of the release profiles helps to compare the fluoride release behavior of materials and also to predict fluoride release amounts for the future. In literature, for these purposes, separate nonlinear statistical models have extensively been utilized. However, the single linear statistical modeling approach has numerous advantages such as providing estimators having good statistical properties, exact results, precise inference and simplicity in calculation. Therefore, this study was conducted to introduce the use of single linear regression modeling to compare release profiles statistically.

Compomers↗

[The measurement parameters and statistical analysis of therapeutic study].

We can draw a correct conclusion unless we select appropriate measurement parameters and properly do the statistical analysis including description and inference, after checking and processing the data of therapeutic study. A comprehensive understanding and correct application of the measurement parameters as well as statistical processing is the key to improve our scientific research. In this article, some common mistakes were addressed in the light of basic concepts in therapeutic study, focusing on the measurement parameters (survival rate, negative conversion rate or positive conversion rate, relative benefit increase, relative risk reduction, relative risk increase, number needed to treat, likelihood of being helped vs. harmed, index of effectiveness) and statistical processing (data type, distribution character, intention-to-treat analysis, efficacy analysis, treatment received analysis, hypothesis testing).

Clinical Trials as Topic↗

Bayesian processing of vestibular information.

Complex self-motion stimulations in the dark can be powerfully disorienting and can create illusory motion percepts. In the absence of visual cues, the brain has to use angular and linear acceleration information provided by the vestibular canals and the otoliths, respectively. However, these sensors are inaccurate and ambiguous. We propose that the brain processes these signals in a statistically optimal fashion, reproducing the rules of Bayesian inference. We also suggest that this processing is related to the statistics of natural head movements. This would create a perceptual bias in favour of low velocity and acceleration. We have constructed a Bayesian model of self-motion perception based on these assumptions. Using this model, we have simulated perceptual responses to centrifugation and off-vertical axis rotation and obtained close agreement with experimental findings. This demonstrates how Bayesian inference allows to make a quantitative link between sensor noise and ambiguities, statistics of head movement, and the perception of self-motion.

Acceleration↗

An empirical Bayes approach to inferring large-scale gene association networks.

MOTIVATION: Genetic networks are often described statistically using graphical models (e.g. Bayesian networks). However, inferring the network structure offers a serious challenge in microarray analysis where the sample size is small compared to the number of considered genes. This renders many standard algorithms for graphical models inapplicable, and inferring genetic networks an 'ill-posed' inverse problem. METHODS: We introduce a novel framework for small-sample inference of graphical models from gene expression data. Specifically, we focus on the so-called graphical Gaussian models (GGMs) that are now frequently used to describe gene association networks and to detect conditionally dependent genes. Our new approach is based on (1) improved (regularized) small-sample point estimates of partial correlation, (2) an exact test of edge inclusion with adaptive estimation of the degree of freedom and (3) a heuristic network search based on false discovery rate multiple testing. Steps (2) and (3) correspond to an empirical Bayes estimate of the network topology. RESULTS: Using computer simulations, we investigate the sensitivity (power) and specificity (true negative rate) of the proposed framework to estimate GGMs from microarray data. This shows that it is possible to recover the true network topology with high accuracy even for small-sample datasets. Subsequently, we analyze gene expression data from a breast cancer tumor study and illustrate our approach by inferring a corresponding large-scale gene association network for 3883 genes.

Algorithms↗

RAxML-III: a fast program for maximum likelihood-based inference of large phylogenetic trees.

MOTIVATION: The computation of large phylogenetic trees with statistical models such as maximum likelihood or bayesian inference is computationally extremely intensive. It has repeatedly been demonstrated that these models are able to recover the true tree or a tree which is topologically closer to the true tree more frequently than less elaborate methods such as parsimony or neighbor joining. Due to the combinatorial and computational complexity the size of trees which can be computed on a Biologist's PC workstation within reasonable time is limited to trees containing approximately 100 taxa. RESULTS: In this paper we present the latest release of our program RAxML-III for rapid maximum likelihood-based inference of large evolutionary trees which allows for computation of 1.000-taxon trees in less than 24 hours on a single PC processor. We compare RAxML-III to the currently fastest implementations for maximum likelihood and bayesian inference: PHYML and MrBayes. Whereas RAxML-III performs worse than PHYML and MrBayes on synthetic data it clearly outperforms both programs on all real data alignments used in terms of speed and final likelihood values. Availability SUPPLEMENTARY INFORMATION: RAxML-III including all alignments and final trees mentioned in this paper is freely available as open source code at http://wwwbode.cs.tum/~stamatak CONTACT: stamatak@cs.tum.edu.

Algorithms↗

Bayesian inference of phylogeny and its impact on evolutionary biology.

As a discipline, phylogenetics is becoming transformed by a flood of molecular data. These data allow broad questions to be asked about the history of life, but also present difficult statistical and computational problems. Bayesian inference of phylogeny brings a new perspective to a number of outstanding issues in evolutionary biology, including the analysis of large phylogenetic trees and complex evolutionary models and the detection of the footprint of natural selection in DNA sequences.

Algorithms↗

Exact unconditional tables for significance testing in the 2 x 2 multinomial trial.

This paper presents tables analogous to T-tables for use in the 2 x 2 multinomial trial, where the continuity corrected Z-statistic is used to make exact unconditional inference. This is the first solution of a discrete exact unconditional inference problem involving a multivariate nuisance parameter for which no ancillary statistic exists.

Bias↗