Stopping rules, Bayesian reconstructions and sieves.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Families in which a single male is affected with a disease which might be either X linked recessive or autosomal recessive present problems in counselling. Before female relatives can be counselled, the probabilities of each mode of inheritance must be assessed, taking into account the prior probabilities, the pedigree structure, any DNA probe data, and any carrier testing data. The widely used linkage analysis package LINKAGE can be used to do the calculation, which is much simpler than the conventional Bayesian method.
Inferring coalescent trees from genomic data has become a major subject in population genetics, particularly with the recent advances in tree sequence reconstruction methods. However, it remains unclear how well these methods perform for imbalanced genealogies. Such imbalances can arise from processes such as cultural transmission of reproductive success (CTRS) or positive selection. Using simulated genomic data, we benchmarked three major software packages, SINGER, Relate, and tsinfer, by comparing the imbalance of reconstructed trees by these methods with that of the true simulated trees, for three indices that quantify this imbalance. The three methods performed well under scenarios yielding balanced trees. However, their accuracy declined as imbalance increased. Performances also varied with mutation rate, recombination rate, and sample size. This study opens possibilities for applying these methods to infer CTRS or positive selection in large-scale genomic datasets, using simulation-based inference such as approximate Bayesian computation.
A compound sampling model, where a unit-specific parameter is sampled from a prior distribution and then observed are generated by a sampling distribution depending on the parameter, underlies a wide variety of biopharmaceutical data. For example, in a multi-centre clinical trial the true treatment effect varies from centre to centre. Observed treatment effects deviate from these true effects through sampling variation. Knowledge of the prior distribution allows use of Bayesian analysis to compute the posterior distribution of clinic-specific treatment effects (frequently summarized by the posterior mean and variance). More commonly, with the prior not completely specified, observed data can be used to estimate the prior and use it to produce the posterior distribution: an empirical Bayes (or variance component) analysis. In the empirical Bayes model the estimated prior mean gives the typical treatment effect and the estimated prior standard deviation indicates the heterogeneity of treatment effects. In both the Bayes and empirical Bayes approaches, estimated clinic effects are shrunken towards a common value from estimates based on single clinics. This shrinkage produces more efficient estimates. In addition, the compound model helps structure approaches to ranking and selection, provides adjustments for multiplicity, allows estimation of the histogram of clinic-specific effects, and structures incorporation of external information. This paper outlines the empirical Bayes approach. Coverage will include development and comparison of approaches based on parametric priors (for example, a Gaussian prior with unknown mean and variance) and non-parametric priors, discussion of the importance of accounting for uncertainty in the estimated prior, comparison of the output and interpretation of fixed and random effects approaches to estimating population values, estimating histograms, and identification of key considerations in the use and interpretation of empirical Bayes methods.
MOTIVATION: Despite the widely recognized importance of replicability in biological research, computational methods to quantify irreplicability and identify irreplicable instances remain underdeveloped. This article presents an efficient and robust computational framework to address this gap. RESULTS: To tackle the challenge of defining an acceptable level of intrinsic heterogeneity among replicable studies, we introduce a distinguishability criterion, ensuring that replicable effects, while potentially heterogeneous, can be distinguished from zero effects and maintain consistent directions with high probability. We implement a Bayesian model criticism approach, reporting a Bayesian P-value to identify potential irreplicable instances. Through numerical experiments, we demonstrate the efficacy of the proposed methods in detecting batch effects in high-throughput experiments and identifying instances of the publication bias. Finally, we apply the framework to multi-tissue eQTL data from the GTEx consortium, uncovering tissue-specific eQTLs that represent biological heterogeneity across tissues. AVAILABILITY AND IMPLEMENTATION: An R package DiscRep implementing our method is available on GitHub (https://github.com/PengWang96/DiscRep).
The predictive performance of a nomogram for dosing warfarin was compared with that of a computer program. The nomogram and the computer program were developed from the log-linear model describing warfarin pharmacodynamics at steady state. The nomogram's dose-response curves were generated by using previously reported pharmacodynamic and pharmacokinetic values for an outpatient population receiving warfarin. The series of dose-response curves were plotted by altering the pharmacodynamic values over a range of 3 standard deviations. The ability of the nomogram to predict the steady-state prothrombin time ratio (PTR) after an adjustment in the dosage of warfarin was evaluated, and the results were compared with those of a commercially available program involving Bayesian regression. Data for 65 outpatients were evaluated. The mean +/- S.D. nomogram-predicted, computer-predicted, and measured PTRs were 1.63 +/- 0.27, 1.64 +/- 0.24, and 1.66 +/- 0.23, respectively. The mean prediction errors for the nomogram and the computer program were -0.037 and -0.026, respectively, and the mean percent absolute prediction errors were 11.6% and 11.0%, respectively. Neither method was biased, and differences between the results for the two methods were not significant. The predictive performance of the warfarin dosing nomogram was comparable to that of the computer program.
Existing computer-based decision aids in the areas of psychiatric diagnosis and consultation are reviewed, and the prospects for expert system development within the mental health field are discussed. Emphasis is placed upon the decision-making models used in these systems rather than on their particular application area. The decision-making paradigms discussed are (1) data bank analysis, (2) statistical pattern recognition, (3) Bayesian analysis, (4) logical flow chart method, and (5) knowledge-based (expert system) approaches. For each paradigm, its essential features, its strengths and weaknesses, and some example applications are presented.
We have developed a computer-administered history designed to directly interview hospitalized patients with pulmonary disease. A frame-based decision system is used to direct the history and to generate a one- to five-member differential diagnostic list based on this history. This system incorporates a cognitive model of question selection and a Bayesian scoring algorithm. Structures to control the choice of questions are embedded in the diagnostic frames and in a QUERY program that makes the final choice of questions. We have compared the behavior of this decision-driven approach with a history taken using a paper questionnaire. The paper-based history presents 182 questions to every patient and captured 75% of 85 pulmonary diseases in its differential lists. The decision-driven system asks 50.7 +/- 31.0 (mean +/- standard deviation) and captured 74% of 61 pulmonary diseases. Our experience suggests that the use of a computerized diagnostic knowledge base to direct the selection of pertinent questions can substantially reduce the number of questions necessary to collect a diagnostically useful patient history.
Explore the source record for details and available documents.
A pharmacokinetic program that allows individualization of drug dosage regimens through the Bayesian method is described. The program, which is designed for the Hewlett-Packard HP-41 CV calculator, is based upon the one-compartment open model with either instantaneous or zero-order absorption. Individualized estimation of the patient's kinetic parameters (clearance and volume of distribution) is performed by analyzing the plasma levels measured in the patient as well as considering the population data of the drug. After estimating the individual kinetic parameters by the Bayesian method, the program predicts the dosage regimen that will elicit the desired peak and trough plasma levels at steady state. For comparison purposes, the least-squares estimates for clearance and volume of distribution are calculated, and dosage prediction can also be made on the basis of the least-squares estimates. The least-squares estimates can be used to calculate population pharmacokinetic parameters according to the Standard Two-Stage method. Several examples of clinical use of the program are presented. The examples refer to patients with classic hemophilia who were treated with Factor VIII concentrates. In these patients, the Bayesian kinetic parameters of Factor VIII have been estimated through the calculator program. The Bayesian parameter estimates generated by the HP-41 have been compared with those determined by a Bayesian program (ADVISE) designed for microcomputers.
The predictive performance of two computer programs for lidocaine dosing were evaluated. Two-compartment Bayesian and nonlinear least-squares regression programs were used in two groups of patients (15 acute arrhythmia patients and 14 chronic arrhythmia patients). Lidocaine was given as a 1.5 mg/kg bolus and a 2.8 mg/min infusion for 48 h. A second bolus (0.5 mg/kg) was given 10 min after the first bolus over 2 min. Serum samples of the patients receiving lidocaine were drawn at 2, 15, 30 min and 1, 2, and 4 h and were used in forecasting the serum concentrations at 6, 8, 12, and 48 h. Predictive performance was assessed by mean error and mean-squared error. The results (mean +/- 95% confidence intervals) demonstrated the Bayesian program predicted a significant (p less than 0.05) difference at 12 h between the two arrhythmia groups (acute 0.52 [-0.95; -0.09] and chronic 0.28 [0.12; 0.44]). The results also demonstrated the Bayesian method was significantly more precise compared to the nonlinear least-squares regression program at 8, 12, and 48 h for the acute group. While caution is warranted, this study demonstrated that the predictive performance by a two-compartment Bayesian model is more accurate in predicting future lidocaine serum concentrations than that by nonlinear least-squares regression.
A new method for estimating audiograms using behavioral responses is presented. The method is based upon a modification of the Bayesian probability formula in which an outcome is predicted from a static set of events. In the new method, classification of audiograms by sequential testing (CAST), the probabilities of occurrence of audiogram patterns are dynamically updated according to the outcome of each test trial. Computer simulation using an infant response model suggests that the procedure is efficient, sensitive, and specific.
MOTIVATION: Accurately predicting complex protein-protein interactions (PPIs) is crucial for decoding biological processes, from cellular functioning to disease mechanisms. However, experimental methods for determining PPIs are computationally expensive. Thus, attention has been recently drawn to machine learning approaches. Furthermore, insufficient effort has been made toward analyzing signed PPI networks, which capture both activating (positive) and inhibitory (negative) interactions. To accurately represent biological relationships, we present the Signed Two-Space Proximity Model (S2-SPM) for signed PPI networks, which explicitly incorporates both types of interactions, reflecting the complex regulatory mechanisms within biological systems. This is achieved by leveraging two independent latent spaces to differentiate between positive and negative interactions while representing protein similarity through proximity in these spaces. Our approach also enables the identification of archetypes representing extreme protein profiles. RESULTS: S2-SPM's superior performance in predicting the presence and sign of interactions in SPPI networks is demonstrated in link prediction tasks against relevant baseline methods. Additionally, the biological prevalence of the identified archetypes is confirmed by an enrichment analysis of Gene Ontology (GO) terms, which reveals that distinct biological tasks are associated with archetypal groups formed by both interactions. This study is also validated regarding statistical significance and sensitivity analysis, providing insights into the functional roles of different interaction types. Finally, the robustness and consistency of the extracted archetype structures are confirmed using the Bayesian Normalized Mutual Information (BNMI) metric, proving the model's reliability in capturing meaningful SPPI patterns. AVAILABILITY: S2-SPM is implemented and freely available under the MIT license at https://github.com/Nicknakis/S2SPM.
Differential analysis is a routine procedure in the statistical analysis toolbox across many applied fields, including quantitative proteomics, the main illustration of the present paper. The state-of-the-art limma approach uses a hierarchical formulation with moderated-variance estimators for each analyte directly injected into the t-statistic. While standard hypothesis testing strategies are recognised for their low computational cost, allowing for quick extraction of the most differential among thousands of elements, they generally overlook key aspects such as handling missing values, inter-element correlations, and uncertainty quantification. The present paper proposes a fully Bayesian framework for differential analysis, leveraging a conjugate hierarchical formulation for both the mean and the variance. Inference is performed by computing the posterior distribution of compared experimental conditions and sampling from the distribution of differences. This approach provides well-calibrated uncertainty quantification at a similar computational cost as hypothesis testing by leveraging closed-form equations. Furthermore, a natural extension enables multivariate differential analysis that accounts for possible inter-element correlations. We also demonstrate that, in this Bayesian treatment, missing at random data should generally be ignored in univariate settings, and further derive a tailored approximation that handles multiple imputation for the multivariate setting. We argue that probabilistic statements in terms of effect size and associated uncertainty are better suited to practical decision-making. Therefore, we finally propose simple and intuitive inference criteria, such as the overlap coefficient, which express group similarity as a probability rather than traditional, and often misleading, p-values. The performance of this approach is evaluated through an extensive empirical study using both synthetic and controlled real-world proteomics datasets. Overall, we believe that this Bayesian framework for (multivariate) differential analysis provides a valuable and intuitive counterpart to standard methods at a comparable computational cost.
The use of Bayes' theorem as a diagnostic tool in clinical medicine normally requires an input of exact probability estimates. However, humans tend to think in categories ("likely," "unlikely," etc.) rather than in terms of exact probability. A computer simulation of the presenting features of a case of pelvic infection has been used to compare the effects of quantitative and qualitative probability estimates on the diagnostic accuracy of Bayes' theorem. For the commoner conditions (prior probability greater than or equal to 0.2) the use of a two- or three-category system is virtually equivalent to the use of exact probability. However, uncommon conditions (prior probability less than or equal to 0.03) are completely ignored by the qualitative system. It is concluded that the use of simple categories of probability is acceptable for a Bayesian diagnostic system provided that the target conditions have a relatively high prior probability.
We present a general spreadsheet model for evaluating diagnostic performance of clinical tests. Our model depicts test results as an r X c matrix, with r possible test results and c possible clinical states. Analysis of this matrix is based on the Ri/Cj ratio, calculated as a number of subjects having a specified result Ri within a given clinical state Cj, divided by total subjects within this clinical state. From this model, we can identify three special cases: (1) a 2 X c matrix, with two possible test results of T+ or T-, over c possible clinical states; (2) an r X 2 matrix, with r possible test results, over two possible clinical states of D+ or D-; and (3) a 2 X 2 matrix, with two possible test results over two possible clinical states. Application of the Ri/Cj ratio to the r X c matrix provides a useful approach to graphic analysis of multiple test results over multiple clinical states. The Ri/Cj ratio also provides a general approach to Bayesian analysis, in which likelihood ratio, relative operating characteristic analysis, sensitivity, and specificity represent special cases or special applications.
The emergency department (ED) is a unique setting for pharmacokinetic-guided drug administration because of the need to rapidly optimize therapy. We compared outcomes in patients receiving intravenous aminophylline according to population-based ED guidelines (group 1) or Bayesian-derived pharmacokinetic estimates (group 2), we determined predictors for admission or discharge in our study group, and we assessed the ability of a Bayesian pharmacokinetic model to estimate theophylline requirements in the ED. The study population was composed of 82 patients (42 males, 40 females) with a mean age of 43 +/- 15.5 years. Fifteen patients were excluded because of protocol violations. Of the 67 cases studied, 30 were assigned to group 1, and 37 were assigned to group 2. Patient demographics, baseline theophylline concentration, and theophylline loading dose did not differ significantly between treatment groups. The aminophylline maintenance infusion was significantly (P less than .001) lower in group 1 (0.4 +/- 0.2 mg/kg/h) than in group 2 (0.6 +/- 0.2 mg/kg/h). Serum theophylline concentrations at one hour post-loading-dose did not differ significantly between treatment groups; however, significant differences were observed at two hours post-load (P less than .002) and four hours post-load (P less than .001). Baseline peak flow rate (PFR) was significantly (P less than .03) higher in group 1 (170 +/- 85 L/min) than in group 2 (132 +/- 62 L/min), but did not differ significantly at any other times throughout the study. The PFR one hour post-load (PFR-1) was the strongest (P less than .003) predictor of outcome.(ABSTRACT TRUNCATED AT 250 WORDS)
An experimental computer system was developed to support diagnosis of rheumatic disorders by computing diagnostic probabilities using modified likelihood ratios. The authors examined whether the performance of the model was affected by the settings in which the data used to derive the likelihood ratios were collected. The sensitivities and specificities of various clinical features for diagnosing rheumatoid arthritis (RA) were obtained from: 1) a study of 1,570 consecutive outpatients at a rheumatology clinic; 2) a review of the literature; 3) estimates by rheumatologists; and 4) a population study. Considerable variations in sensitivity and specificity but satisfactory agreement in likelihood ratios were found across the four data sets. The likelihood ratios were then used to compute the probabilities of RA in a test series of 570 of the rheumatology clinic outpatients. The model's diagnoses with likelihood ratios from the other sources were adequate. When the likelihood ratios from these sources were combined, discrimination came close to what could be achieved by using the likelihood ratios based on the data from the clinic. The method applied in the study, which makes use of variation of input data instead of variation of test series, and the results are relevant to assessing the external validity and transferability of Bayesian decision-support systems.