Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,549 records · Page 86Linked to original sources

Comparison of linear and nonlinear classification algorithms for the prediction of drug and chemical metabolism by human UDP-glucuronosyltransferase isoforms.

Partial least squares discriminant analysis (PLSDA), Bayesian regularized artificial neural network (BRANN), and support vector machine (SVM) methodologies were compared by their ability to classify substrates and nonsubstrates of 12 isoforms of human UDP-glucuronosyltransferase (UGT), an enzyme "superfamily" involved in the metabolism of drugs, nondrug xenobiotics, and endogenous compounds. Simple two-dimensional descriptors were used to capture chemical information. For each data set, 70% of the data were used for training, and the remainder were used to assess the generalization performance. In general, the SVM methodology was able to produce models with the best predictive performance, followed by BRANN and then PLSDA. However, a small number of data sets showed either equivalent or better predictability using PLSDA, which may indicate relatively linear relationships in these data sets. All SVM models showed predictive ability (>60% of test set predicted correctly) and five out of the 12 test sets showed excellent prediction (>80% prediction accuracy). These models represent the first use of pattern recognition methods to discriminate between substrates and nonsubstrates of human drug metabolizing enzymes and the first thorough assessment of three classification algorithms using multiple metabolic data sets.

Algorithms↗

Enhancing detection of polygenic adaptation: a comparative study of machine learning and statistical approaches using simulated evolve-and-resequence data.

BACKGROUND: Detecting signals of polygenic adaptation remains a significant challenge in population genomics, as traditional methods often struggle to identify the associated subtle, multi-locus allele-frequency shifts. Here, we introduced and tested several novel approaches combining machine learning techniques with traditional statistical tests to detect polygenic adaptation patterns in time-series of allele frequency changes from whole genome data. We implemented a Naive Bayesian Classifier (NBC) and One-Class Support Vector Machines (OCSVM), and compared their performance against the classical Fisher's Exact Test (FET). Furthermore, we combined machine learning and statistical models (OCSVM-FET and NBC-FET), resulting in 5 competing approaches. The framework is mainly designed and validated for evolve-and-resequence (EaR) experimental designs, where defined selection pressures and temporal sampling are feasible, but might be applicable for certain natural experiments as well. RESULTS: Using a simulated dataset based on empirical C. riparius Pool-Seq data, we evaluated methods across evolutionary scenarios varying in generation, selection strength, and number of loci under selection. Our results demonstrate that the combined OCSVM-FET approach consistently outperformed competing methods, achieving the lowest false positive rate, highest area under the curve, and high accuracy. The performance peak aligned with what we term the 'late dynamic phase' of adaptation - the period after initial selection has occurred but before fixation - highlighting the method's sensitivity to ongoing selective processes. CONCLUSIONS: Furthermore, we emphasize the critical role of parameter tuning, balancing biological assumptions with methodological rigor. While broader applicability remains an important direction for future work, the present benchmarking is intentionally scoped to EaR experimental contexts.

Machine Learning↗

Linear and nonlinear modeling of antifungal activity of some heterocyclic ring derivatives using multiple linear regression and Bayesian-regularized neural networks.

Antifungal activity was modeled for a set of 96 heterocyclic ring derivatives (2,5,6-trisubstituted benzoxazoles, 2,5-disubstituted benzimidazoles, 2-substituted benzothiazoles and 2-substituted oxazolo(4,5-b)pyridines) using multiple linear regression (MLR) and Bayesian-regularized artificial neural network (BRANN) techniques. Inhibitory activity against Candida albicans (log(1/C)) was correlated with 3D descriptors encoding the chemical structures of the heterocyclic compounds. Training and test sets were chosen by means of k-Means Clustering. The most appropriate variables for linear and nonlinear modeling were selected using a genetic algorithm (GA) approach. In addition to the MLR equation (MLR-GA), two nonlinear models were built, model BRANN employing the linear variable subset and an optimum model BRANN-GA obtained by a hybrid method that combined BRANN and GA approaches (BRANN-GA). The linear model fit the training set (n = 80) with r2 = 0.746, while BRANN and BRANN-GA gave higher values of r2 = 0.889 and r2 = 0.937, respectively. Beyond the improvement of training set fitting, the BRANN-GA model was superior to the others by being able to describe 87% of test set (n = 16) variance in comparison with 78 and 81% the MLR-GA and BRANN models, respectively. Our quantitative structure-activity relationship study suggests that the distributions of atomic mass, volume and polarizability have relevant relationships with the antifungal potency of the compounds studied. Furthermore, the ability of the six variables selected nonlinearly to differentiate the data was demonstrated when the total data set was well distributed in a Kohonen self-organizing neural network (KNN).

Antifungal Agents↗

Circular clustering of protein dihedral angles by Minimum Message Length.

Early work on proteins identified the existence of helices and extended sheets in protein secondary structures, a high-level classification which remains popular today. Using the Snob program for information-theoretic Minimum Message Length (MML) classification, we are able to take the protein dihedral angles as determined by X-ray crystallography, and cluster sets of dihedral angles into groups. Previous work by Hunter and States has applied a similar Bayesian classification method, AutoClass, to protein data with site position represented by 3 Cartesian co-ordinates for each of the alpha-Carbon, beta-Carbon and Nitrogen, totalling 9 co-ordinates. By using the von Mises circular distribution in the Snob program, we are instead able to represent local site properties by the two dihedral angles, phi and psi. Since each site can be modelled as having 2 degrees of freedom, this orientation-invariant dihedral angle representation of the data is more compact than that of nine highly-correlated Cartesian co-ordinates. Using the information-theoretic message length concepts discussed in the paper, such a more concise model is more likely to represent the underlying generating process from which the data came. We report on the results of our classification, plotting the classes in (phi, psi) space; and introducing a symmetric information-theoretic distance measure to build a minimum spanning tree between the classes. We also give a transition matrix between the classes and note the existence of three classes in the region phi approximately -1.09 rad and psi approximately -0.75 rad which are close on the spanning tree and have high inter-transition probabilities. This gives rise to a tight, abundant and self-perpetuating structure.

Computational Biology↗

Relationship of P300 single-trial responses with reaction time and preceding stimulus sequence.

Variation of single-trial P300 responses was studied both in relation to reaction times and to the preceding stimulus sequence in an auditory oddball paradigm. Single-trial responses were estimated with the Subspace regularization method that is based on Bayesian estimation and linear modeling. The results of the single-trial method were compared to those of averaging. Both methods showed that the latency of the P300 was shorter and its amplitude larger for faster than slower reaction times. The P300 latency was shorter for target tones that were preceded by a large number of standard tones compared to those preceded by a small number of standard tones. The P300 amplitude was statistically significantly affected by the stimulus sequence only when analyzed with conventional averaging. In-depth analysis of standard deviations showed that the variability of the P300 single-trial latencies could explain the differences between the two methods. Specifically, the regression analysis showed that the latency correlated negatively with the number of preceding standard tones and positively with the reaction time, whereas the P300 amplitude correlated positively with the number of the preceding standard stimuli and negatively with the reaction time. The analysis of the single-trial responses gives information about the behavior of the P300 component that is lost with conventional averaging. The method used in this study is independent of subjective decision making and can be used to model changes in the dynamical behavior of the P300 component objectively.

Adult↗

Comparison of four single-point phenytoin dosage prediction techniques using computer-simulated pharmacokinetic values.

The predictive abilities of the following four single-point phenytoin dosage adjustment methods were compared using computer-simulated data: the Bayesian feedback method of Vozeh et al. (B), a linearized version of the Bayesian method (LB), the population-clearance method of Graves et al. (G), and the Rambeck nomogram (R). A series of 512 "subjects" with normally distributed values for volume of distribution, weight, and the Michaelis-Menten variables Vmax and Km were simulated. The steady-state serum concentration (SSSC) resulting from the administration of a standard dose of phenytoin sodium (5 mg/kg/day) was calculated, and "subjects" with SSSCs less than or equal to 12 mg/L or greater than or equal to 17 mg/L were entered in the study. If the concentration was greater than 50 mg/L or the standard dosage exceeded Vmax, the dosage was reduced empirically by 25%. Normally distributed random errors were introduced into the SSSC values to simulate actual patient data. The pharmacokinetic values, dosages, and SSSCs were used for predicting the dosage required to attain an SSSC of 14.9 mg/L. In the unstratified population, the mean error and mean-squared error were lowest for methods G and B, followed by methods LB and R. Methods B and LB gave the highest percentages of satisfactory dosage predictions based on the resultant SSSC value. The performance of all methods was superior at initial SSSCs greater than 8 mg/L.(ABSTRACT TRUNCATED AT 250 WORDS)

Humans↗

Bayesian neural networks for aroma classification.

Bayesian Neural Networks (BNNs) are investigated to test their potential to distinguish between different aroma impressions. Special attention is thereby drawn on mixed aroma impressions, resulting from the flavor description of a single compound with more than one aroma quality. The structures of 133 pyrazine-derived aroma compounds as well as their aroma descriptions are selected for comparison. The information fed into the neural networks is based on molecular descriptors calculated from the geometrically optimized chemical structures. While in the case of the Probabilistic Neural Network (PNN) the networks' output consists of a categorical variable, the output for the General Regression Neural Network (GRNN) is defined in a numerical way. The best models attain comparable performance with a correct prediction of 90.8% of the cases for PNN and 89.9% for GRNN, respectively. Comparison of the BNN results to those obtained by Multiple Linear Regression (MLR) points out that the nonlinear methods work significantly better on the studied problem and that BNNs can be applied to multiple-category problems in structure-flavor relationships with good accuracy.

Bayes Theorem↗

Bayesian segregation analysis of somatic cell scores of Ontario Holstein cattle.

Bayesian segregation analysis using a Gibbs sampling approach was applied to four sets of simulated data and one set of field data to detect evidence of major genes affecting the evaluated trait. The substitution effect of a major gene and its allelic frequency were estimated for each set of data. For two datasets simulated with a model with no major gene effect, the resulting estimates of polygenic variance and heritability agreed with the simulated values and tests for the presence of a major gene were not significant. Analyses of two sets of data simulated with a major gene produced posterior distributions that gave significant evidence of major gene effects but underestimated the substitution values of the major gene. The segregation analysis of field data suggested that a major gene significantly affected somatic cell score (SCS) in the population of Ontario Holstein cattle. The estimated heritability of SCS was approximately 0.16. The major gene variance accounted for about 17% of the total genetic variance and the point estimate of the frequency of the allele having a positive effect on SCS was 0.30. However, the precision of these estimates is questionable based on the simulation results. The effect of the major gene may be underestimated.

Animals↗

Comparing a mass-balance algorithm with a Bayesian regression analysis computer program for predicting serum phenytoin concentrations.

The ability of a mass-balance algorithm to predict non-steady-state phenytoin concentrations in neurosurgery patients was compared with that of Phenda, a computerized Bayesian regression analysis program. Fifty neurosurgery patients who had had two or more initial phenytoin serum concentrations measured at least 60 hours apart and at least 1 hour after any i.v. doses, with the second concentration being not more than twice and not less than half of the first, and who had had a third or final phenytoin measurement (for use in a prediction analysis) were evaluated. The patients' maximum rates of metabolism were calculated by using the two initial phenytoin concentrations and a mass-balance algorithm, and the third phenytoin concentration was predicted. The patients' demographics and phenytoin dosages and concentrations were entered into Phenda, which was used to predict the third phenytoin concentration. The ability of the two methods to predict the third concentration was evaluated by the method of Sheiner and Beal. Fifty observations from 48 patients were evaluated. The mass-balance algorithm had a positive prediction bias of 2.52 mg/L and a precision error of 5.08 mg/L, compared with 2.30 and 5.30, respectively, for Phenda. The difference in the results between the two methods was not significant. There was no significant difference between the mass-balance algorithm and Phenda in the ability to predict phenytoin concentrations.

Algorithms↗

Bayesian reanalysis of a quantitative trait locus accounting for multiple environments by scaling in broilers.

A Bayesian method was developed to handle QTL analyses of multiple experimental data of outbred populations with heterogeneity of variance between sexes for all random effects. The method employed a scaled reduced animal model with random polygenic and QTL allelic effects. A parsimonious model specification was applied by choosing assumptions regarding the covariance structure to limit the number of parameters to estimate. Markov chain Monte Carlo algorithms were applied to obtain marginal posterior densities. Simulation demonstrated that joint analysis of multiple environments is more powerful than separate single trait analyses of each environment. Measurements on broiler BW obtained from 2 experiments concerning growth efficiency and carcass traits were used to illustrate the method. The population consisted of 10 full-sib families from a cross between 2 broiler lines. Microsatellite genotypes were determined on generations 1 and 2, and phenotypes were collected on groups of generation 3 animals. The model included a polygenic correlation, which had a posterior mean of 0.70 in the analyses. The reanalysis agreed on the presence of a QTL in marker bracket MCW0058-LEI0071 accounting for 34% of the genetic variation in males and 24% in females in the growth efficiency experiment. In the carcass experiment, this QTL accounted for 19% of the genetic variation in males and 6% in females.

Animal Husbandry↗

A new dynamic Bayesian network (DBN) approach for identifying gene regulatory networks from time course microarray data.

MOTIVATION: Signaling pathways are dynamic events that take place over a given period of time. In order to identify these pathways, expression data over time are required. Dynamic Bayesian network (DBN) is an important approach for predicting the gene regulatory networks from time course expression data. However, two fundamental problems greatly reduce the effectiveness of current DBN methods. The first problem is the relatively low accuracy of prediction, and the second is the excessive computational time. RESULTS: In this paper, we present a DBN-based approach with increased accuracy and reduced computational time compared with existing DBN methods. Unlike previous methods, our approach limits potential regulators to those genes with either earlier or simultaneous expression changes (up- or down-regulation) in relation to their target genes. This allows us to limit the number of potential regulators and consequently reduce the search space. Furthermore, we use the time difference between the initial change in the expression of a given regulator gene and its potential target gene to estimate the transcriptional time lag between these two genes. This method of time lag estimation increases the accuracy of predicting gene regulatory networks. Our approach is evaluated using time-series expression data measured during the yeast cell cycle. The results demonstrate that this approach can predict regulatory networks with significantly improved accuracy and reduced computational time compared with existing DBN approaches.

Algorithms↗

Estimation of single-trial multicomponent ERPs: differentially variable component analysis (dVCA).

A Bayesian inference framework for estimating the parameters of single-trial, multicomponent, event-related potentials is presented. Single-trial recordings are modeled as the linear combination of ongoing activity and multicomponent waveforms that are relatively phase-locked to certain sensory or motor events. Each component is assumed to have a trial-invariant waveform with trial-dependent amplitude scaling factors and latency shifts. A Maximum a Posteriori solution of this model is implemented via an iterative algorithm from which the component's waveform, single-trial amplitude scaling factors and latency shifts are estimated. Multiple components can be derived from a single-channel recording based on their differential variability, an aspect in contrast with other component analysis techniques (e.g., independent component analysis) where the number of components estimated is equal to or smaller than the number of recording channels. Furthermore, we show that, by subtracting out the estimated single-trial components from each of the single-trial recordings, one can estimate the ongoing activity, thus providing additional information concerning task-related brain dynamics. We test this approach, which we name differentially variable component analysis (dVCA), on simulated data and apply it to an experimental dataset consisting of intracortically recorded local field potentials from monkeys performing a visuomotor pattern discrimination task.

Algorithms↗

Assembly of protein tertiary structures from fragments with similar local sequences using simulated annealing and Bayesian scoring functions.

We explore the ability of a simple simulated annealing procedure to assemble native-like structures from fragments of unrelated protein structures with similar local sequences using Bayesian scoring functions. Environment and residue pair specific contributions to the scoring functions appear as the first two terms in a series expansion for the residue probability distributions in the protein database; the decoupling of the distance and environment dependencies of the distributions resolves the major problems with current database-derived scoring functions noted by Thomas and Dill. The simulated annealing procedure rapidly and frequently generates native-like structures for small helical proteins and better than random structures for small beta sheet containing proteins. Most of the simulated structures have native-like solvent accessibility and secondary structure patterns, and thus ensembles of these structures provide a particularly challenging set of decoys for evaluating scoring functions. We investigate the effects of multiple sequence information and different types of conformational constraints on the overall performance of the method, and the ability of a variety of recently developed scoring functions to recognize the native-like conformations in the ensembles of simulated structures.

Bayes Theorem↗

Predicting disease outcome of non-invasive transitional cell carcinoma of the urinary bladder using an artificial neural network model: results of patient follow-up for 15 years or longer.

BACKGROUND: Patients with non-invasive (Ta/T1) transitional cell carcinoma (TCC) of the urinary bladder are often observed without progression in the long-term follow-up period, although many of them experience recurrence of disease. It is difficult to accurately predict the disease outcome of each patient with Ta/T1 TCC using conventional prognostic criteria. In this study, we examined the usefulness of artificial neural networks (ANNs) to predict the long-term disease outcome of patients with TCC of the urinary bladder. METHODS: A retrospective, prognostic study of 90 patients with Ta/T1 TCC of the urinary bladder, diagnosed by transurethral resection of the bladder tumor between April 1981 and March 1985, and then followed up for 15 years or longer, was carried out. Data were analyzed using the Bayesian network tool of SPSS Neural Connection 2.1. The input neural data consisted of tumor stage, grade, tumor number, age, gender, tumor architecture and estimates of mean nuclear volume. The data set was randomly divided into 68 training and 22 testing examples for the prediction of disease progression and tumor recurrence within 15 years. RESULTS: During 15 years follow-up, tumor recurrence was noted in 42/90 (47%) Ta/T1 tumors. The ANN model could not predict tumor recurrence. Conversely, disease progression was noted in 17/90 (19%) Ta/T1 tumors, and, in the test set, 4/22 (18%) Ta/T1 tumors underwent disease progression. The sensitivity of the ANN model to predict progression was 100% (specificity 67%; positive predictive value 40%; negative predictive value 100%). Patients who were judged to have a favorable prognosis using ANN analysis did not progress within the 15-year follow-up period. CONCLUSION: The results of the ANN study indicate that long-term progression-free survival of patients with non-invasive TCC of the urinary bladder can be precisely predicted. A favorable prognosis using ANNs would be one of the exclusion criteria for immediate or future total cystectomy.

Adult↗

Dealing with gene expression missing data.

Compared evaluation of different methods is presented for estimating missing values in microarray data: weighted K-nearest neighbours imputation (KNNimpute), regression-based methods such as local least squares imputation (LLSimpute) and partial least squares imputation (PLSimpute) and Bayesian principal component analysis (BPCA). The influence in prediction accuracy of some factors, such as methods' parameters, type of data relationships used in the estimation process (i.e. row-wise, column-wise or both), missing rate and pattern and type of experiment [time series (TS), non-time series (NTS) or mixed (MIX) experiments] is elucidated. Improvements based on the iterative use of data (iterative LLS and PLS imputation--ILLSimpute and IPLSimpute), the need to perform initial imputations (modified PLS and Helland PLS imputation--MPLSimpute and HPLSimpute) and the type of relationships employed (KNNarray, LLSarray, HPLSarray and alternating PLS--APLSimpute) are proposed. Overall, it is shown that data set properties (type of experiment, missing rate and pattern) affect the data similarity structure, therefore influencing the methods' performance. LLSimpute and ILLSimpute are preferable in the presence of data with a stronger similarity structure (TS and MIX experiments), whereas PLS-based methods (MPLSimpute, IPLSimpute and APLSimpute) are preferable when estimating NTS missing data.

Algorithms↗

Bayesian adaptive sequence alignment algorithms.

The selection of a scoring matrix and gap penalty parameters continues to be an important problem in sequence alignment. We describe here an algorithm, the 'Bayes block aligner, which bypasses this requirement. Instead of requiring a fixed set of parameter settings, this algorithm returns the Bayesian posterior probability for the number of gaps and for the scoring matrices in any series of interest. Furthermore, instead of returning the single best alignment for the chosen parameter settings, this algorithm returns the posterior distribution of all alignments considering the full range of gapping and scoring matrices selected, weighing each in proportion to its probability based on the data. We compared the Bayes aligner with the popular Smith-Waterman algorithm with parameter settings from the literature which had been optimized for the identification of structural neighbors, and found that the Bayes aligner correctly identified more structural neighbors. In a detailed examination of the alignment of a pair of kinase and a pair of GTPase sequences, we illustrate the algorithm's potential to identify subsequences that are conserved to different degrees. In addition, this example shows that the Bayes aligner returns an alignment-free assessment of the distance between a pair of sequences.

Adenylate Kinase↗

Timing of human immunodeficiency virus type 1 (HIV-1) transmission from mother to child: bayesian estimation using a mixture.

The timing of mother-to-child HIV transmission is not directly observable but influences the infected child's viral and immune status in the neonatal period. A hierarchical model was developed in a Bayesian framework to 'back-calculate' the timing of HIV-1 transmission from mother to child from the virological and immunological kinetics in the infected infant. Joint evolution of viral markers and immune response was modelled as a continuous time Markov process. The modelling of the period from infection to birth was based on a mixture of three distributions taking into account the various mother-to-child transmission pathways: In utero (early or late in gestation) and intrapartum (during the delivery process), integrating the fact that transmission is a continuum during the pregnancy. Gibbs sampling was used to estimate the marginal posterior distributions of the transition intensities between stages of HIV infection and those of the individual times from infection to birth. We applied our model to data on 135 perinatally HIV-1-infected children included in the French Prospective Study on Pediatric HIV infection. The model suggested that transmission occurred late in utero during the last month of pregnancy and that the day of delivery was a particularly critical time in HIV-1 transmission from mother to child. The paper ends with a discussion of model assumptions and a comparison with results obtained using a non-parametric method.

Antibodies, Viral↗

Mixture models and subpopulation classification: a pharmacokinetic simulation study and application to metoprolol CYP2D6 phenotype.

Mixture models are applied in population pharmacometrics to characterize underlying population distributions that are not adequately approximated by a single normal or lognormal distribution. In addition to obtaining individualized maximum a posteriori Bayesian post hoc parameter estimates, the subpopulation to which an individual was classified can be determined. However, the accuracy of the classification of subjects to subpopulations is not well studied. We investigated the impact of several factors on the accuracy of classification in mixture models applied to pharmacokinetics using a simulation strategy. The availability of actual subject data allowed us to evaluate mixture model classification in a potentially common application, namely, the classification of clearance into poor metabolizer (PM) or extensive metabolizer (EM) subgroups with the known phenotype status in subjects receiving metoprolol. The factors explored in the simulation study were the magnitude of difference between the clearances in two subpopulations, the between subject variability in clearance, the mixing-fraction, and the population sample size. Populations were simulated at various levels of the above factors and analyzed with a mixture model using NONMEM. The population pharmacokinetics of metoprolol were modeled with the EM/PM phenotype as a known covariate, and without the phenotype covariate using a mixture model. Within the range of scenarios studied, the proportion of subjects classified into the correct subpopulation was high. The simulation-estimation study suggests that a greater separation between two subpopulations, a smaller variability in the parameter distribution, a larger sample size, and a smaller size subpopulation tend to be associated with a greater accuracy of subpopulation classification when a mixture model is applied to pharmacokinetic data. In a population pharmacokinetic analysis of metoprolol, a drug that undergoes polymorphic metabolism, it was possible to correctly identify phenotype status using a mixture model.

Administration, Oral↗