Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

A new summarization method for Affymetrix probe level data.

MOTIVATION: We propose a new model-based technique for summarizing high-density oligonucleotide array data at probe level for Affymetrix GeneChips. The new summarization method is based on a factor analysis model for which a Bayesian maximum a posteriori method optimizes the model parameters under the assumption of Gaussian measurement noise. Thereafter, the RNA concentration is estimated from the model. In contrast to previous methods our new method called 'Factor Analysis for Robust Microarray Summarization (FARMS)' supplies both P-values indicating interesting information and signal intensity values. RESULTS: We compare FARMS on Affymetrix's spike-in and Gene Logic's dilution data to established algorithms like Affymetrix Microarray Suite (MAS) 5.0, Model Based Expression Index (MBEI), Robust Multi-array Average (RMA). Further, we compared FARMS with 43 other methods via the 'Affycomp II' competition. The experimental results show that FARMS with default parameters outperforms previous methods if both sensitivity and specificity are simultaneously considered by the area under the receiver operating curve (AUC). We measured two quantities through the AUC: correctly detected expression changes versus wrongly detected (fold change) and correctly detected significantly different expressed genes in two sets of arrays versus wrongly detected (P-value). Furthermore FARMS is computationally less expensive then RMA, MAS and MBEI. AVAILABILITY: The FARMS R package is available from http://www.bioinf.jku.at/software/farms/farms.html. SUPPLEMENTARY INFORMATION: http://www.bioinf.jku.at/publications/papers/farms/supplementary.ps

Algorithms↗

Identification of land use with water quality data in stormwater using a neural network.

To control stormwater pollution effectively, development of innovative, land-use-related control strategies will be required. An approach that could differentiate land-use types from stormwater quality would be the first step to solving this problem. We propose a neural network approach to examine the relationship between stormwater water quality and various types of land use. The neural network model can be used to identify land-use types for future known and unknown cases. The neural model uses a Bayesian network and has 10 water quality input variables, four neurons in the hidden layer, and five land-use target variables (commercial, industrial, residential, transportation, and vacant). We obtained 92.3 percent of correct classification and 0.157 root-mean-squared error on test files. Based on the neural model, simulations were performed to predict the land-use type of a known data set, which was not used when developing the model. The simulation accurately described the behavior of the new data set. This study demonstrates that a neural network can be effectively used to produce land-use type classification with water quality data.

Agriculture↗

Hierarchical modelling of small area and hospital variation in short-term prognosis after acute myocardial infarction. A longitudinal study of 35- to 74-year-old men in Denmark between 1978 and 1997.

Models for analysis of trends in hospital and small area variation in case fatality after acute myocardial infarction are presented. The data are from administrative registries in Denmark. Hierarchical modelling in a logistic regression with a Bayesian approach is used. Model selection is undertaken using the deviance and the Bayesian information criteria. There is a modest trend for hospital variation in case-fatality rates that coincides with the introduction of new treatment strategies. This hospital variation is considerably larger than the variation at the area level. There is no trend for variation of the case-fatality rates at the area level. Unstructured random effects slightly outperform spatially correlated random effects at the area level. Somewhat high correlations over time within hospitals and within areas were detected for the case-fatality rates. Heavy-tailed distributions (T-distributions) could be an alternative for the random effect distribution in data from administrative registries and compete in the model selection with the normal distribution in this study.

Adult↗

User modeling and adaptation in health promotion dialogs with an animated character.

In this paper, we describe our experience with the design and implementation of an embodied conversational agent (ECA) that converses with users to change their dietary behavior. Our intent is to develop a system that dynamically models the agent and the user and adapts the agent's counseling dialog accordingly. Towards this end, we discuss our efforts to automatically determine the user's dietary behavior stage of change and attitude towards the agent on the basis of unconstrained typed text dialog, first with another person and then with an ECA controlled by an experimenter in a wizard of Oz study. We describe how the results of these studies have been incorporated into an algorithm that combines the results from simple parsing rules together with contextual features using a Bayesian network to determine user stage and attitude automatically.

Artificial Intelligence↗

GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.

MOTIVATION: Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. RESULTS: In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. AVAILABILITY AND IMPLEMENTATION: The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).

Amino Acid Motifs↗

Bioinformatics-driven, rational engineering of protein thermostability.

A longstanding goal in protein engineering is to identify specific sequence changes that endow proteins with desired functional properties. As opposed to traditional rational and random protein engineering techniques, we have employed a bioinformatic approach to identify specific sequence changes that influence key functional properties of a protein within a defined superfamily. Specifically, we have used the Bayesian sequence-based algorithms PROBE and Classifier to identify a strand-turn-strand motif that contributes to thermophilicity among members of the serine protease subtilase superfamily. By replacing a 16 amino acid sequence in the mesophilic subtilisin E (from Bacillus subtilis) with a bioinformatics-generated thermophilic model sequence, the melting temperature of subtilisin E was increased by 13 degrees C. While wild-type subtilisin E was inactive at 90 degrees C, the mutant retained a substantial fraction of its function, with ca. one-third of the activity that it has at 45 degrees C.

Algorithms↗

MRDtarget: A heuristic Gaussian approach for optimizing targeted capture regions to enhance Minimal Residual Disease detection.

Molecular residual disease (MRD) detection, initially developed for hematologic malignancies, has become a critical biomarker for monitoring solid tumors. MRD detection primarily relies on circulating tumor DNA (ctDNA) analysis using next-generation sequencing, offering high sensitivity and broad genomic coverage. However, challenges remain in designing cost-effective panels that maximize mutation detection while maintaining biological relevance. Fixed panels often lack sufficient patient-specific mutation coverage, while WES-based personalized MRD assays, despite their high sensitivity, are costly and less accessible. We developed a tumor comprehensive genomic profiling (CGP)-informed personalized MRD assay to detect tumor-derived mutations, which allowed us to design patient-specific personalized panels and meanwhile, provide a cost-effective alternative to whole exome sequencing (WES). To address these limitations, we developed MRDtarget, a heuristic multivariate Gaussian model-based targeted capture region selection method. By expanding beyond traditional hotspot regions, MRDtarget optimizes variant tracking for MRD detection, significantly improving sensitivity. Using a Bayesian inference-based heuristic approach, MRDtarget integrates multi-feature informativeness rates to identify optimal genomic regions for capture. Experimental results demonstrate that MRDtarget enables the detection of more variants per patient. This study underscores the importance of rational panel design to improve MRD sensitivity and provides a novel approach to enhance precision diagnostics and treatment for solid tumor patients.

Humans↗

Genuine Bayesian multiallelic significance test for the Hardy-Weinberg equilibrium law.

Statistical tests that detect and measure deviation from the Hardy-Weinberg equilibrium (HWE) have been devised but are limited when testing for deviation at multiallelic DNA loci is attempted. Here we present the full Bayesian significance test (FBST) for the HWE. This test depends neither on asymptotic results nor on the number of possible alleles for the particular locus being evaluated. The FBST is based on the computation of an evidence index in favor of the HWE hypothesis. A great deal of forensic inference based on DNA evidence assumes that the HWE is valid for the genetic loci being used. We applied the FBST to genotypes obtained at several multiallelic short tandem repeat loci during routine parentage testing; the locus Penta E exemplifies those clearly in HWE while others such as D10S1214 and D19S253 do not appear to show this.

Alleles↗

Bayesian inference on genetic merit under uncertain paternity.

A hierarchical animal model was developed for inference on genetic merit of livestock with uncertain paternity. Fully conditional posterior distributions for fixed and genetic effects, variance components, sire assignments and their probabilities are derived to facilitate a Bayesian inference strategy using MCMC methods. We compared this model to a model based on the Henderson average numerator relationship (ANRM) in a simulation study with 10 replicated datasets generated for each of two traits. Trait 1 had a medium heritability (h2) for each of direct and maternal genetic effects whereas Trait 2 had a high h2 attributable only to direct effects. The average posterior probabilities inferred on the true sire were between 1 and 10% larger than the corresponding priors (the inverse of the number of candidate sires in a mating pasture) for Trait 1 and between 4 and 13% larger than the corresponding priors for Trait 2. The predicted additive and maternal genetic effects were very similar using both models; however, model choice criteria (Pseudo Bayes Factor and Deviance Information Criterion) decisively favored the proposed hierarchical model over the ANRM model.

Animals↗

Automated global structure extraction for effective local building block processing in XCS.

Learning Classifier Systems (LCSs), such as the accuracy-based XCS, evolve distributed problem solutions represented by a population of rules. During evolution, features are specialized, propagated, and recombined to provide increasingly accurate subsolutions. Recently, it was shown that, as in conventional genetic algorithms (GAs), some problems require efficient processing of subsets of features to find problem solutions efficiently. In such problems, standard variation operators of genetic and evolutionary algorithms used in LCSs suffer from potential disruption of groups of interacting features, resulting in poor performance. This paper introduces efficient crossover operators to XCS by incorporating techniques derived from competent GAs: the extended compact GA (ECGA) and the Bayesian optimization algorithm (BOA). Instead of simple crossover operators such as uniform crossover or one-point crossover, ECGA or BOA-derived mechanisms are used to build a probabilistic model of the global population and to generate offspring classifiers locally using the model. Several offspring generation variations are introduced and evaluated. The results show that it is possible to achieve performance similar to runs with an informed crossover operator that is specifically designed to yield ideal problem-dependent exploration, exploiting provided problem structure information. Thus, we create the first competent LCSs, XCS/ECGA and XCS/BOA, that detect dependency structures online and propagate corresponding lower-level dependency structures effectively without any information about these structures given in advance.

Algorithms↗

Computer-assisted individual estimation of radioiodine thyroid uptake in Grave's disease.

A computer-assisted Bayesian individual estimation of radioiodine thyroid uptake kinetics for patients suffering from Grave's disease is proposed. The program provides a fast computation of the activity to be administered to a given patient to achieve a target thyroid absorbed dose. This determination relies upon the patient biological covariates and upon a small number of measurements performed during a preliminar kinetic study of radioiodine thyroid uptake. Our results indicate that a two-sample Bayesian approach is reliable when external thyroid counts are performed at 2 h and 168 h after a test dose and has advantages over conventional kinetic experiments in terms of patient acceptability. This method is implemented on widespread computers and interfaced with a patient database. An interactive user interface with in-line data checking is provided. The program could be also a tool to better study the relationship between the absorbed dose and the clinical effect.

Adult↗

Modeling of farnesyltransferase inhibition by some thiol and non-thiol peptidomimetic inhibitors using genetic neural networks and RDF approaches.

Inhibition of farnesyltransferase (FT) enzyme by a set of 78 thiol and non-thiol peptidomimetic inhibitors was successfully modeled by a genetic neural network (GNN) approach, using radial distribution function descriptors. A linear model was unable to successfully fit the whole data set; however, the optimum Bayesian regularized neural network model described about 87% inhibitory activity variance with a relevant predictive power measured by q2 values of leave-one-out and leave-group-out cross-validations of about 0.7. According to their activity levels, thiol and non-thiol inhibitors were well-distributed in a topological map, built with the inputs of the optimum non-linear predictor. Furthermore, descriptors in the GNN model suggested the occurrence of a strong dependence of FT inhibition on the molecular shape and size rather than on electronegativity or polarizability characteristics of the studied compounds.

Enzyme Inhibitors↗

Quantification of atherosclerotic plaque components using in vivo MRI and supervised classifiers.

In this work we aimed to study the possibility of using supervised classifiers to quantify the main components of carotid atherosclerotic plaque in vivo on the basis of multisequence MRI data. MRI data consisting of five MR weightings were obtained from 25 symptomatic subjects. Histological micrographs of endarterectomy specimens from the 25 carotids were used as a standard of reference for training and evaluation. The set of subjects was divided in a training set (12 subjects) and an evaluation set (13 subjects). Four different classifiers and two human MRI readers determined the percentages of calcified tissue, fibrous tissue, lipid core, and intraplaque hemorrhage on the subject level for all subjects in the evaluation set. Quantification of the relatively small amounts of calcium could not be done with statistical significance by either the classifiers or the MRI readers. For the other tissues a simple Bayesian classifier (Bayes) performed better than the other classifiers and the MRI readers. All classifiers performed better than the MRI readers in quantifying the sum of hemorrhage and lipid proportions. The MRI readers overestimated the hemorrhage proportions and tended to underestimate the lipid proportions. In conclusion, this pilot study demonstrates the benefits of algorithmic classifiers for quantifying plaque components.

Algorithms↗

Phylogenomics and molecular evolution of polyomaviruses.

We provide in this chapter an overview of the basic steps to reconstruct evolutionary relationships through standard phylogeny estimation approaches as well as network approaches for sequences more closely related. We discuss the importance of sequence alignment, selecting models of evolution, and confidence assessment in phylogenetic inference. We also introduce the reader to a variety of software packages used for such studies. Finally, we demonstrate these approaches throughout using a data set of 33 whole genomes of polyomaviruses. A robust phylogeny of these genomes is estimated and phylogenetic relationships among the polyomaviruses determined using Bayesian and maximum likelihood approaches. Furthermore, population samples of SV40 are used to demonstrate the utility of network approaches for closely related sequences. The phylogenetic analysis suggested a close relationship among the BK viruses, JC viruses, and SV40 with a more distant association with mouse polyomavirus, monkey polymavirus (LPV) and then avian polyomavirus (BFDV).

Computational Biology↗

Evolution of homologous recombination rates across bacteria.

Bacteria are nonsexual organisms but are capable of exchanging DNA at diverse degrees through homologous recombination. Intriguingly, the rates of recombination vary immensely across lineages where some species have been described as purely clonal and others as "quasi-sexual." However, estimating recombination rates has proven a difficult endeavor and estimates often vary substantially across studies. It is unclear whether these variations reflect natural variations across populations or are due to differences in methodologies. Consequently, the impact of recombination on bacterial evolution has not been extensively evaluated and the evolution of recombination rate-as a trait-remains to be accurately described. Here, we developed an approach based on Approximate Bayesian Computation that integrates multiple signals of recombination to estimate recombination rates. We inferred the rate of recombination of 162 bacterial species and one archaeon and tested the robustness of our approach. Our results confirm that recombination rates vary drastically across bacteria; however, we found that recombination rate-as a trait-is conserved in several lineages but evolves rapidly in others. Although some traits are thought to be associated with recombination rate (e.g., GC-content), we found no clear association between genomic or phenotypic traits and recombination rate. Overall, our results provide an overview of recombination rate, its evolution, and its impact on bacterial evolution.

Bacteria↗

Computer-assisted drug assay interpretation based on Bayesian estimation of individual pharmacokinetics: application to lidocaine.

A microcomputer program for individualized drug level prediction based on Bayesian forecasting is presented. It is written so that the clinician can integrate patient demographics and drug levels to design a new dosage regimen tailored to an individual patient. The program's great flexibility and robustness make it appropriate for realistic clinical settings. A validation with a data set of lidocaine concentrations measured in 18 patients revealed that the program can predict serum lidocaine levels accurately enough to enhance individual patient dosage adjustment within a few hours after a dosage regimen is started.

Bayes Theorem↗

Stochastic search variable selection for identifying multiple quantitative trait loci.

In this article, we utilize stochastic search variable selection methodology to develop a Bayesian method for identifying multiple quantitative trait loci (QTL) for complex traits in experimental designs. The proposed procedure entails embedding multiple regression in a hierarchical normal mixture model, where latent indicators for all markers are used to identify the multiple markers. The markers with significant effects can be identified as those with higher posterior probability included in the model. A simple and easy-to-use Gibbs sampler is employed to generate samples from the joint posterior distribution of all unknowns including the latent indicators, genetic effects for all markers, and other model parameters. The proposed method was evaluated using simulated data and illustrated using a real data set. The results demonstrate that the proposed method works well under typical situations of most QTL studies in terms of number of markers and marker density.

Bayes Theorem↗

Learning yeast gene functions from heterogeneous sources of data using hybrid weighted Bayesian networks.

We developed a machine learning system for determining gene functions from heterogeneous sources of data sets using a Weighted Naive Bayesian Network (WNB). The knowledge of gene functions is crucial for understanding many fundamental biological mechanisms such as regulatory pathways, cell cycles and diseases. Our major goal is to accurately infer functions of putative genes or ORFs (Open Reading Frames) from existing databases using computational methods. However, this task is intrinsically difficult since the underlying biological processes represent complex interactions of multiple entities. Therefore many functional links would be missing when only one or two source of data is used in the prediction. Our hypothesis is that integrating evidence from multiple and complementary sources could significantly improve the prediction accuracy. In this paper, our experimental results not only suggest that the above hypothesis is valid, but also provide guidelines for using the WNB system for data collection, training and predictions. The combined training data sets contain information from gene annotations, gene expressions, clustering outputs, keyword annotations and sequence homology from public databases. The current system is trained and tested on the genes of budding yeast Saccharomyces cerevisiae. Our WNB model can also be used to analyze the contribution of each source of information toward the prediction performance through the weight training process. The contribution analysis could potentially lead to significant scientific discovery by facilitating the interpretation and understanding of the complex relationships between biological entities.

Artificial Intelligence↗