Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

IsoBayes: a Bayesian approach for single-isoform proteomics inference.

MOTIVATION: Studying protein isoforms is an essential step in biomedical research; at present, the main approach for analyzing proteins is via bottom-up mass spectrometry proteomics, which return peptide identifications, that are indirectly used to infer the presence of protein isoforms. However, the detection and quantification processes are noisy; in particular, peptides may be erroneously detected, and most peptides, known as shared peptides, are associated to multiple protein isoforms. As a consequence, studying individual protein isoforms is challenging, and inferred protein results are often abstracted to the gene-level or to groups of protein isoforms. RESULTS: Here, we introduce IsoBayes, a novel statistical method to perform inference at the isoform level. Our method enhances the information available, by integrating mass spectrometry proteomics and transcriptomics data in a Bayesian probabilistic framework. To account for the uncertainty in the measurement process, we propose a two-layer latent variable approach: first, we sample if a peptide has been correctly detected (or, alternatively filter peptides); second, we allocate the abundance of such selected peptides across the protein(s) they are compatible with. This enables us, starting from peptide-level data, to recover protein-level data; in particular, we: (i) infer the presence/absence of each protein isoform (via a posterior probability), (ii) estimate its abundance (and credible interval), and (iii) target isoforms where transcript and protein relative abundances significantly differ. We benchmarked our approach in simulations, and in two multi-protease real datasets: our method displays good sensitivity and specificity when detecting protein isoforms, its estimated abundances highly correlate with the ground truth, and can detect changes between protein and transcript relative abundances. AVAILABILITY AND IMPLEMENTATION: IsoBayes is freely distributed as a Bioconductor R package, and is accompanied by an example usage vignette.

Proteomics↗

Overall concordance correlation coefficient for evaluating agreement among multiple observers.

Accurate and precise measurement is an important component of any proper study design. As elaborated by Lin (1989, Biometrics 45, 255-268), the concordance correlation coefficient (CCC) is more appropriate than other indices for measuring agreement when the variable of interest is continuous. However, this agreement index is defined in the context of comparing two fixed observers. In order to use multiple observers in a study involving large numbers of subjects, there is a need to assess agreement among these multiple observers. In this article, we present an overall CCC (OCCC) in terms of the interobserver variability for assessing agreement among multiple fixed observers. The OCCC turns out to be equivalent to the generalized CCC (King and Chinchilli, 2001, Statistics in Medicine 20, 2131-2147; Lin, 1989; Lin, 2000, Biometrics 56, 324-325) when the squared distance function is used. We evaluated the OCCC through generalized estimating equations (Barnhart and Williamson, 2001, Biometrics 57, 931-940) and U-statistics (King and Chinchilli, 2001) for inference. This article offers the following important points. First, it addresses the precision and accuracy indices as components of the OCCC. Second, it clarifies that the OCCC is the weighted average of all pairwise CCCs. Third, it is intuitively defined in terms of interobserver variability. Fourth, the inference approaches of GEE and the U-statistics are compared via simulations for small samples. Fifth, we illustrate the use of the OCCC by two medical examples with the GEE, U-statistics, and bootstrap approaches.

Adult↗

The problem of multiple inference in identifying point-source environmental hazards.

Point-source environmental hazards are often identified by examination of unusual clusters of disease cases. The very large number of potential clusters give rise to the statistical problem of "multiple inference," i.e., the more clusters examined, the greater the risk of "false-positive" associations emerging by chance alone. This paper first distinguishes the situation of clusters identified by anecdotal observation from those that emerge from systematic searches. The latter may or may not include a systematic enumeration of potential causal factors associated with each potential disease cluster. If exposure information is not systematically available, empirical Bayes procedures are suggested as a basis for ranking the observed clusters in order of priority for further investigation. If exposure information is systematically available, empirical Bayes procedures can be used to select associations to report or to rank them in order of priority for confirmation. In addition, procedures are described for testing the global null hypothesis of no exposure-disease associations and for estimating the number of true-positive associations. These approaches are advocated in preference to classical frequentist approaches of multiplying p values by the number of tests performed.

Environmental Pollution↗

Mapping a quantitative trait locus via the EM algorithm and Bayesian classification.

Mapping a locus controlling a quantitative genetic trait (e.g., blood pressure) to a specific genomic region is of considerable interest. Data on the quantitative trait under consideration and several codominant genetic markers with known genomic locations are collected from members of families and statistically analyzed to draw inferences on the genomic position of the trait locus. The vector of parameters of interest comprises the pairwise recombination fractions, theta, between the putative quantitative trait locus and the marker loci. One of the major complications in estimating theta for a quantitative trait in humans is the lack of haplotype information on members of families. The purpose of this study was to devise a computationally simple and efficient method of estimation of theta in the absence of haplotype information. We have proposed a two-stage estimation procedure using the expectation-maximization (EM) algorithm. In the first stage, parameters of the QTL are estimated based on data of a sample of unrelated individuals. From estimates thus obtained, we have used a Bayes' rule to infer QTL genotypes of parents in families. Finally, in the second stage of the procedure, we have proposed an EM algorithm for obtaining the maximum likelihood estimate of theta based on data of informative families (which are identified upon inferring parental QTL genotypes performed in the first stage). We have shown, using simulated data, that the proposed procedure is cost-effective, computationally simple, and statistically efficient. As expected, analysis of data on multiple markers jointly is more efficient than the analysis based on single markers.

Algorithms↗

Estimation of auditory brainstem response, ABR, by means of Bayesian inference.

The present paper describes a new method to estimate the auditory brainstem response when the electrical activity from the recording electrodes displays non-stationarity, i.e. varies between low and high levels. The method is based on a statistical approach called Bayesian inference and weights the individual components (here blocks of 250 sweeps) inversely proportional to the level of the noise activity during the recording. Fifty sets of data from 10 consecutive patients obtained during stimulation at high intensity are used to evaluate the difference between the classic averaging and the present method which is called Bayes estimation. In approximately 30% of the cases, a significant all-over improvement is obtained by the new method. The classic averaging technique would here require 50% more sweeps to be taken to obtain the same precision of the ABR estimate, on average. Also the latency and amplitude parameters of the Jv wave complex are evaluated and it is shown that the parameter variance decreases by a factor of approximately 2 by using the Bayes estimation. The new technique is compared with a similar technique recently presented by Hoke et al. (1984) and the differences and similarities are discussed.

Bayes Theorem↗

Testing mutual independence between two discrete-valued spatial processes: a correction to pearson chi-squared.

A common feature of data collected in environmental and earth sciences is that they typically exhibit spatial autocorrelation. Violating the assumption of independent observations can have dramatic effects on inferences derived from standard statistical methods. In this article, we examine the consequences of spatial autocorrelation on Pearson's chi-squared test of mutual independence between two categorical responses with a general number of classes. Correspondingly, we suggest a simple modification to the standard test statistic that allows for spatial autocorrelation. Our modified statistic is based on a first-order correction factor and thus provides only an approximate test. However, we show by Monte Carlo simulation that this approximation results in satisfactory inferences in several situations of practical interest. The usefulness of the method is displayed through an application to categorical data arising in the study of the relationship between the distribution pattern of plant species and woodland age in a forest in northern Belgium.

Belgium↗

Inferring the fitness effects of DNA mutations from polymorphism and divergence data: statistical power to detect directional selection under stationarity and free recombination.

The fitness effects of classes of DNA mutations can be inferred from patterns of nucleotide variation. A number of studies have attributed differences in levels of polymorphism and divergence between silent and replacement mutations to the action of natural selection. Here, I investigate the statistical power to detect directional selection through contrasts of DNA variation among functional categories of mutations. A variety of statistical approaches are applied to DNA data simulated under Sawyer and Hartl's Poisson random field model. Under assumptions of free recombination and stationarity, comparisons that include both the frequency distributions of mutations segregating within populations and the numbers of mutations fixed between populations have substantial power to detect even very weak selection. Frequency distribution and divergence tests are applied to silent and replacement mutations among five alleles of each of eight Drosophila simulans genes. Putatively "preferred" silent mutations segregate at higher frequencies and are more often fixed between species than "unpreferred" silent changes, suggesting fitness differences among synonymous codons. Amino acid changes tend to be either rare polymorphisms or fixed differences, consistent with a combination of deleterious and adaptive protein evolution. In these data, a substantial fraction of both silent and replacement DNA mutations appear to affect fitness.

Adaptation, Biological↗

Probability and the patient state space.

This paper describes work to develop a model-based system to support clinical decision-making. In previous articles, we have developed (from 695 measurement sets obtained from 148 patients) a physiologic state classification based on a set of 11 cardiovascular and metabolic measurements. There is an R or reference state, for stable ICU patients. Patients under (operative, traumatic, or compensated septic) stress, or with (septic or hepatic) metabolic, respiratory, or cardiac insufficiency are in the A, B, C, or D states, respectively. We wished to make the state easier to measure and eventually available continuously, automatically, and noninvasively, as well as reflecting a wider group of bodily systems. The 5 centers define a 4 dimensional affine subspace, designated the cardiovascular state space. Using eigenvector analysis, we have found four new derived physiologic variables CV1, CV2, CV3, and CV4 that span the state space. We have fit sets of linear regression equations that allow the patient's position in the state space, and therefore his state, to be determined from more easily obtainable sets of measurements. Further, we selected 1966 measurement sets from 512 patients at two hospitals. We used the data from 250 of these patients to define 13 prototypical types, namely survivors and deaths from various combinations of sepsis, cardiogenic decompensation, cirrhosis, and pneumonitis, following trauma or general surgery. For any future patient, the statistical theory of Bayesian inference allows one to infer back from the measurements observed to the probability of his being of any of these types and of surviving or dying. We used this method to predict the outcome of the other 262 patients, prospectively. Statistically, the predictions of survival or death were not significantly different from the actual. For individual patients, the method predicts a clinical course that closely follows the actual episodes in their history. These results confirm and explain the validity of the concept of the patient state and make the state easier to compute. The patient state and the probability plot together help to stage, select, and evaluate therapy. They do not replace the clinician's judgement, but rather are tools that help the clinician to exercise judgement.

Adult↗

Bioconductor: an open source framework for bioinformatics and computational biology.

This chapter describes the Bioconductor project and details of its open source facilities for analysis of microarray and other high-throughput biological experiments. Particular attention is paid to concepts of container and workflow design, connections of biological metadata to statistical analysis products, support for statistical quality assessment, and calibration of inference uncertainty measures when tens of thousands of simultaneous statistical tests are performed.

Animals↗

Calcium regulation of single ryanodine receptor channel gating analyzed using HMM/MCMC statistical methods.

Type-II ryanodine receptor channels (RYRs) play a fundamental role in intracellular Ca(2+) dynamics in heart. The processes of activation, inactivation, and regulation of these channels have been the subject of intensive research and the focus of recent debates. Typically, approaches to understand these processes involve statistical analysis of single RYRs, involving signal restoration, model estimation, and selection. These tasks are usually performed by following rather phenomenological criteria that turn models into self-fulfilling prophecies. Here, a thorough statistical treatment is applied by modeling single RYRs using aggregated hidden Markov models. Inferences are made using Bayesian statistics and stochastic search methods known as Markov chain Monte Carlo. These methods allow extension of the temporal resolution of the analysis far beyond the limits of previous approaches and provide a direct measure of the uncertainties associated with every estimation step, together with a direct assessment of why and where a particular model fails. Analyses of single RYRs at several Ca(2+) concentrations are made by considering 16 models, some of them previously reported in the literature. Results clearly show that single RYRs have Ca(2+)-dependent gating modes. Moreover, our results demonstrate that single RYRs responding to a sudden change in Ca(2+) display adaptation kinetics. Interestingly, best ranked models predict microscopic reversibility when monovalent cations are used as the main permeating species. Finally, the extended bandwidth revealed the existence of novel fast buzz-mode at low Ca(2+) concentrations.

Algorithms↗

Statistical methods for analysis of time course gene expression data.

Since many biological systems or regulatory networks are dynamic systems, gene expression levels measured over different time points during a given biological process can often provide more insights about the underlying system. These gene expression data measured over time are often called the time-course gene expression data. One unique feature of such data is the time dependency of the gene expression levels for a given gene at different times or between two different genes. Statistical analysis needs to account for such dependency in order to make valid inferences. This paper presents several statistical methods for analyzing such time-course gene expression data, including the time-lagged correlation coefficient for analyzing the relationship between genes, a mixed-effects model with splines for clustering genes and for estimating missing gene expression data, and a new method for aligning gene expression profiles obtained under two experimental conditions and for identifying gene clusters that show significant changes between two experimental conditions. We used the yeast cell cycle gene expression data sets to illustrate these methods and obtained the biologically meaningful conclusions from these analyses.

Cell Cycle Proteins↗

Classification by multiple-resolution statistical analysis with application to automated recognition of marine mammal sounds.

A multiple-resolution statistical pattern recognition technique for classification by supervised learning is developed and then applied to automated recognition of marine mammal sounds. The data to be classified may be either unprocessed or transformed, e.g., time series or time-frequency distributions of acoustic transients. Training data consist of samples previously grouped by a human expert into labeled sets; these sets are presumed to be associated with different "classes." The labeled sets are then characterized by occupancy statistics associated with a multiple-resolution, binary partition of the (unreduced) sample space. Classification of a new sample is performed by calculating a posteriori probabilities of membership of the new sample in each class, computed by Bayesian inference from the occupancy statistics of the associated labeled set. These a posteriori probabilities are calculated by a recursive algorithm that progresses from coarse to fine resolution in the sample space. The algorithm is implemented in a simple, highly efficient computer program. Automated classification of both time series and time-frequency distributions of marine-mammal vocalizations is demonstrated using a small number of labeled samples (approximately ten samples per class).

Algorithms↗

Toward evidence-based medical statistics. 2: The Bayes factor.

Bayesian inference is usually presented as a method for determining how scientific belief should be modified by data. Although Bayesian methodology has been one of the most active areas of statistical development in the past 20 years, medical researchers have been reluctant to embrace what they perceive as a subjective approach to data analysis. It is little understood that Bayesian methods have a data-based core, which can be used as a calculus of evidence. This core is the Bayes factor, which in its simplest form is also called a likelihood ratio. The minimum Bayes factor is objective and can be used in lieu of the P value as a measure of the evidential strength. Unlike P values, Bayes factors have a sound theoretical foundation and an interpretation that allows their use in both inference and decision making. Bayes factors show that P values greatly overstate the evidence against the null hypothesis. Most important, Bayes factors require the addition of background knowledge to be transformed into inferences--probabilities that a given conclusion is right or wrong. They make the distinction clear between experimental evidence and inferential conclusions while providing a framework in which to combine prior with current evidence.

Bayes Theorem↗

The sensitivity and specificity of markers for event times.

The statistical literature on assessing the accuracy of risk factors or disease markers as diagnostic tests deals almost exclusively with settings where the test, Y, is measured concurrently with disease status D. In practice, however, disease status may vary over time and there is often a time lag between when the marker is measured and the occurrence of disease. One example concerns the Framingham risk score (FR-score) as a marker for the future risk of cardiovascular events, events that occur after the score is ascertained. To evaluate such a marker, one needs to take the time lag into account since the predictive accuracy may be higher when the marker is measured closer to the time of disease occurrence. We therefore consider inference for sensitivity and specificity functions that are defined as functions of time. Semiparametric regression models are proposed. Data from a cohort study are used to estimate model parameters. One issue that arises in practice is that event times may be censored. In this research, we extend in several respects the work by Leisenring et al. (1997) that dealt only with parametric models for binary tests and uncensored data. We propose semiparametric models that accommodate continuous tests and censoring. Asymptotic distribution theory for parameter estimates is developed and procedures for making statistical inference are evaluated with simulation studies. We illustrate our methods with data from the Cardiovascular Health Study, relating the FR-score measured at enrollment to subsequent risk of cardiovascular events.

Aged↗

On analytical methods and inferences for 2 x 2 contingency table data from medical studies.

Analysis of 2 x 2 contingency tables is not as trivial as it appears. The choice of the statistical test can affect the inferences resulting from data analysis, especially at small sample sizes. Canned statistical programs do not necessarily lead to an appropriate test. These points are demonstrated using examples from the literature.

Data Interpretation, Statistical↗

A retrospective assessment of the accuracy of the paternity inference program CERVUS.

CERVUS is a Windows-based software package written to infer paternity in natural populations. It offers advantages over exclusionary-based methods of paternity inference in that multiple nonexcluded males can be statistically distinguished, laboratory typing error is considered and statistical confidence is determined for assigned paternities through simulation. In this study we use a panel of 84 microsatellite markers to retrospectively determine the accuracy of statistical confidence when CERVUS was used to infer paternity in a population of red deer (Cervus elaphus). The actual confidence of CERVUS-assigned paternities was not significantly different from that predicted by simulation.

Animals↗

Joint linkage of multiple loci for a complex disorder.

Many investigators who have been searching for linkage to complex diseases have by now accumulated a drawer full of negative results. If disease is actually caused by genes at several loci, these data might contain multiple-locus system (MLS) information that the investigator does not realize. Trying to obtain this information formally, through the MLS likelihood, leads to severe computational and statistical difficulties. Therefore, we propose a scheme of inference based on single-locus (SL) statistics, considered jointly. By simulation, we find that the MLS lod score is closely approximated by the sum of SL lod scores. However, we also find that for moderately large systems, say three of four loci, both MLS and SL lod scores are likely to be inconclusive. Nonetheless, MLS can often be detected through the correlation of individual pedigree SL lod scores. Significant correlation is itself evidence of an MLS, because, in the absence of linkage, false-positive lod scores are necessarily random. Under epistasis SL lod scores tend to be positively correlated among pedigrees, while under independent action SL lod scores from high-density samples tend to be negatively correlated.

Chi-Square Distribution↗

Detection of disease genes by use of family data. II. Application to nuclear families.

Two likelihood-based score statistics are used to detect association between a disease and a single diallelic polymorphism, on the basis of data from arbitrary types of nuclear families. The first statistic, the nonfounder statistic, extends the transmission/disequilibrium test to accommodate affected and unaffected offspring and missing parental genotypes. The second statistic, the founder statistic, compares observed or inferred parental genotypes with those of some reference population. In this comparison, the genotypes of affected parents or of those with many affected offspring are weighted more heavily than are the genotypes of unaffected parents or of those with few affected offspring. Genotypes of single unrelated cases and controls can be included in this analysis. We illustrate the two statistics by applying them to data on a polymorphism of the SDR5A2 gene in nuclear families with multiple cases of prostate cancer. We also use simulations to compare the power of the nonfounder statistic with that of the score statistic, on the basis of the conditional logistic regression of offspring genotypes.

Alleles↗