Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian networks”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

V1 neurons signal acquisition of an internal representation of stimulus location.

A fundamental aspect of visuomotor behavior is deciding where to look or move next. Under certain conditions, the brain constructs an internal representation of stimulus location on the basis of previous knowledge and uses it to move the eyes or to make other movements. Neuronal responses in primary visual cortex were modulated when such an internal representation was acquired: Responses to a stimulus were affected progressively by sequential presentation of the stimulus at one location but not when the location was varied randomly. Responses of individual neurons were spatially tuned for gaze direction and tracked the Bayesian probability of stimulus appearance. We propose that the representation arises in a distributed cortical network and is associated with systematic changes in response selectivity and dynamics at the earliest stages of cortical visual processing.

Analysis of Variance↗

Statistical approach to neural network model building for gentamicin peak predictions.

Feed forward neural networks are flexible, nonlinear modeling tools that are an extension of traditional statistical techniques. The hypothesis that feed forward neural network models can be built in a similar fashion as a statistical model was tested. Feed forward neural network models were built using forward and backward variable selection, and zero to five hidden nodes, and tanh and linear transfer functions were used. Gentamicin serum concentrations were predicted as a model drug for testing these methods. Peak observations from 392 patients were used to train, test, and validate the feed forward neural network. Inputs were demographic and drug dosing information. Model selection was performed using the Akaike information criteria (AIC), Bayesian information criteria (BIC), and a method of stopped training. The models with lowest root mean square (rms) error were those with all 10 inputs and five hidden nodes. Average rms error in the validation set was lowest for stopped training (1.46), then AIC (1.51), and finally BIC (1.56). Larger models tended to result in the best predictions. Overfitting can occur in models that are too large, either by using too many nodes in the hidden layer (rms = 1.49) or by using too many inputs with little information associated with them (rms = 1.70). We conclude that neural networks can be built using a large number of parameters that have good predictive performance. Care must be used during training to avoid overfitting the data. A stopped training method resulted in the network with the lowest rms error.

Adult↗

Emergence of two novel HIV-1 Circulating Recombinant Forms (CRF190_0708 and CRF191_0708): molecular characterization and clinical insights from a five-year study in Yunnan, China.

BACKGROUND: To characterize HIV-1 molecular epidemiology and identify novel circulating recombinant forms (CRFs) among antiretroviral therapy (ART)-naïve heterosexuals in Yunnan, China, and evaluate their clinical impact. METHODS: This study examined 636 HIV-1 pol sequences to analyze genetic diversity, pretreatment drug resistance (PDR), and transmission networks. Near full-length genomes were obtained to identify and characterize novel recombinants, with their evolutionary history inferred by Bayesian analysis. Co-receptor tropism was predicted, and the five-year clinical outcomes (including immune reconstitution and virologic response) of patients infected with the novel CRFs were compared. RESULTS: The most prevalent type identified was CRF08_BC, accounting for 50.16% of cases. The prevalence of drug resistance was 5.97% (38/636), with the K103N mutation being the most common. An analysis of transmission networks revealed that 52.2% (272/521) of clusters were associated with CRF07_BC and CRF08_BC. Two novel second-generation CRFs were identified: CRF190_0708, with an estimated time to the most recent common ancestor (tMRCA) of 1998.9, and CRF191_0708, with a more recent tMRCA ranging from 2009.5 to 2011.6. During the five-year follow-up period, viral rebound was observed in 7 patients in the CRF190_0708 group and in 1 patient in the CRF191_0708 group. Drug-resistance mutations (M184V and K103N) were detected in a subset of rebound cases in the CRF190_0708 group. CONCLUSIONS: This study identifies two novel HIV-1 recombinants, CRF190_0708 and CRF191_0708, highlighting ongoing viral evolution in Yunnan. Preliminary findings suggest possible clinical differences, warranting further investigation. Continued molecular surveillance is needed. TRIAL REGISTRATION: The clinical study was registered at ClinicalTrials.gov under the identifier NCT03852849. The date of registration was March 22, 2019.

Adult↗

Coronary artery bypass risk prediction using neural networks.

BACKGROUND: Neural networks are nonparametric, robust, pattern recognition techniques that can be used to model complex relationships. METHODS: The applicability of multilayer perceptron neural networks (MLP) to coronary artery bypass grafting risk prediction was assessed using The Society of Thoracic Surgeons database of 80,606 patients who underwent coronary artery bypass grafting in 1993. The results of traditional logistic regression and Bayesian analysis were compared with single-layer (no hidden layer), two-layer (one hidden layer), and three-layer (two hidden layer) MLP neural networks. These networks were trained using stochastic gradient descent with early stopping. All prediction models used the same variables and were evaluated by training on 40,480 patients and cross-validation testing on a separate group of 40,126 patients. Techniques were also developed to calculate effective odds ratios for MLP networks and to generate confidence intervals for MLP risk predictions using an auxiliary "confidence MLP." RESULTS: Receiver operating characteristic curve areas for predicting mortality were approximately 76% for all classifiers, including neural networks. Calibration (accuracy of posterior probability prediction) was slightly better with a two-member committee classifier that averaged the outputs of a MLP network and a logistic regression model. Unlike the individual methods, the committee classifier did not overestimate or underestimate risk for high-risk patients. CONCLUSIONS: A committee classifier combining the best neural network and logistic regression provided the best model calibration, but the receiver operating characteristic curve area was only 76% irrespective of which predictive model was used.

Bayes Theorem↗

Automatic basis selection techniques for RBF networks.

This paper proposes a generic criterion that defines the optimum number of basis functions for radial basis function (RBF) neural networks. The generalization performance of an RBF network relates to its prediction capability on independent test data. This performance gives a measure of the quality of the chosen model. An RBF network with an overly restricted basis gives poor predictions on new data, since the model has too little flexibility (yielding high bias and low variance). By contrast, an RBF network with too many basis functions also gives poor generalization performance since it is too flexible and fits too much of the noise on the training data (yielding low bias but high variance). Bias and variance are complementary quantities, and it is necessary to assign the number of basis function optimally in order to achieve the best compromise between them. In this paper we use Stein's unbiased risk estimator to derive an analytical criterion for assigning the appropriate number of basis functions. Two cases of known and unknown noise have been considered and the efficacy of this criterion in both situations is illustrated experimentally. The paper also shows an empirical comparison between this method and two well known classical methods, cross validation and the Bayesian information criterion, BIC.

Electronic Data Processing↗

An integrated comprehensive workbench for inferring genetic networks: voyagene.

We propose an integrated, comprehensive network-inferring system for genetic interactions, named VoyaGene, which can analyze experimentally observed expression profiles by using and combining the following five independent inferring models: Clustering, Threshold-Test, Bayesian, multi-level digraph and S-system models. Since VoyaGene also has effective tools for visualizing the inferred results, researchers may evaluate the combination of appropriate inferring models, and can construct a genetic network to an accuracy that is beyond the reach of a single inferring model. Through the use of VoyaGene, the present study demonstrates the effectiveness of combining different inferring models.

Algorithms↗

Assessing different classification methods for virtual screening.

How well do different classification methods perform in selecting the ligands of a protein target out of large compound collections not used to train the model? Support vector machines, random forest, artificial neural networks, k-nearest-neighbor classification with genetic-algorithm-optimized feature selection, trend vectors, naïve Bayesian classification, and decision tree were used to divide databases into molecules predicted to be active and those predicted to be inactive. Training and predicted activities were treated as binary. The database was generated for the ligands of five different biological targets which have been the object of intense drug discovery efforts: HIV-reverse transcriptase, COX2, dihydrofolate reductase, estrogen receptor, and thrombin. We report significant differences in the performance of the methods independent of the biological target and compound class. Different methods can have different applications; some provide particularly high enrichment, others are strong in retrieving the maximum number of actives. We also show that these methods do surprisingly well in predicting recently published ligands of a target on the basis of initial leads and that a combination of the results of different methods in certain cases can improve results compared to the most consistent method.

Algorithms↗

EDECS: the Emergency Department Expert Charting System.

EDECS, the Emergency Department Expert Charting System, integrates clinical guidelines into the everyday practice of medicine. By generating the medical record and patient aftercare instructions, it facilitates patient care. For this reason, doctors are willing to use it. While using it, the doctors are continually presented with advice regarding documentation, testing, and treatment. Unlike guidelines that attempt to modify behavior through traditional educational methods, these computerized guidelines are seen by the physician every time she sees a patient. We have demonstrated this by directly integrating the guidelines into the process of patient care; we can increase compliance with the guidelines [1]. At present EDECS exists for the chief complaints of occupational exposure to body fluids, acute low back pain, recurrent seizure, fever in children, and males with penile discharge or dysuria. Upon examining the patient, the physician proceeds to the computer, which prompts him for essential information regarding the history and physical examination. Certain items are required for all patients with the chief complaint, others are required based on the answers to these items. Data is analyzed by the computer, which provides advice regarding testing and treatment. Once testing is completed, the system suggests a probable diagnosis and aids in patient disposition and discharge planning. Finally, EDECS prints the medical record as well as patient-specific aftercare instructions. EDECS is a user friendly system; most data is entered via mouse. It is written in the OS-based expert system shell AM(TM) and can be run on an IBM compatible PC or PC network. Rules are generally written in an "if...then" format, but more sophisticated rule structures, including Bayesian models, are used when needed. Each module contains separate subroutines for the history, physical, laboratory ordering, treatment, and disposition. These modules call each other in a dynamic fashion. The system is currently being evaluated for its effect on documentation, appropriateness of use of ancillary tests, appropriateness of use of treatments, physician satisfaction, patient satisfaction, and patient outcomes. Initial results of the system's effect on documentation and use of ancillary tests and treatments show much promise [1]. The Occupational Exposure to Body Fluids module has shown a statistically significant, and sometimes rather dramatic, increase in the level of documentation for nearly all items. Further, the advice given by EDECS has caused an increase in the appropriate use of testing and treatments. For example, unnecessary ancillary tests dropped from 1.5 per patient without the computer to 0.1 per patient with the computer's aid. EDECS facilitates quality management activities and research since it collects standardized information and stores it in an easily retrievable database format. It can also be used to educate medical students and residents about the proper care of patients with a given chief complaint. At this session, EDECS will be demonstrated, and issues regarding the development of guidelines, the encoding of guidelines in rules, and the organizational structure of the software will be presented and discussed.

Child↗

Mapping determinants of variation in energy metabolism, respiration and flight in Drosophila.

We employed quantitative trait locus (QTL) mapping to dissect the genetic architecture of a hierarchy of functionally related physiological traits, including metabolic enzyme activity, metabolite storage, metabolic rate, and free-flight performance in recombinant inbred lines of Drosophila melanogaster. We identified QTL underlying variation in glycogen synthase, hexokinase, phosphoglucomutase, and trehalase activity. In each case variation mapped away from the enzyme-encoding loci, indicating that trans-acting regions of the genome are important sources of variation within the metabolic network. Individual QTL associated with variation in metabolic rate and flight performance explained between 9 and 35% of the phenotypic variance. Bayesian QTL analysis identified epistatic effects underlying variation in flight velocity, metabolic rate, glycogen content, and several metabolic enzyme activities. A region on the third chromosome was associated with expression of the glucose-6-phosphate branchpoint enzymes and with metabolic rate and flight performance. These genomic regions are of special interest as they may coordinately regulate components of energy metabolism with effects on whole-organism physiological performance. The complex biochemical network is encoded by an equally complex network of interacting genetic elements with potentially pleiotropic effects. This has important consequences for the evolution of performance traits that depend upon these metabolic networks.

Animals↗

Endocrine-disrupting chemical-induced gene networks confer coronary heart disease risk revealed by causal inference and single-cell analyses.

BACKGROUND: Endocrine-disrupting chemicals (EDCs) are linked to coronary heart disease (CHD), but underlying mechanisms remain unclear. We aimed to identify EDC-related genes and evaluate their causal roles in CHD. METHODS: We curated EDC-related genes from a compound-gene interaction database and integrated them with CHD genome-wide association study (GWAS) summary statistics and tissue-specific expression quantitative trait loci (eQTL) data. Two-sample Mendelian randomization (MR) and Bayesian colocalization were applied to infer causality. Functional enrichment, single-cell RNA sequencing of human coronary arteries, and EDC-gene networks were further analyzed. RESULTS: After FDR correction, 39 genes were significantly associated with CHD risk via MR. Four genes-ZNF827, FCHO1, IPO9 (protective), and RPL13 (risk-increasing)-showed strong colocalization (PPH4 > 0.9). Pathway and single-cell analyses of coronary artery tissue indicated that vascular and immune pathways mediate these effects. An interaction network highlighted associations between specific EDCs and candidate genes implicated in CHD susceptibility. CONCLUSION: This integrative genomic study provides evidence that EDCs influence CHD susceptibility through distinct gene networks, revealing potential mechanisms and molecular targets for prevention and therapy.

Humans↗

Missing-value estimation using linear and non-linear regression with Bayesian gene selection.

MOTIVATION: Data from microarray experiments are usually in the form of large matrices of expression levels of genes under different experimental conditions. Owing to various reasons, there are frequently missing values. Estimating these missing values is important because they affect downstream analysis, such as clustering, classification and network design. Several methods of missing-value estimation are in use. The problem has two parts: (1) selection of genes for estimation and (2) design of an estimation rule. RESULTS: We propose Bayesian variable selection to obtain genes to be used for estimation, and employ both linear and nonlinear regression for the estimation rule itself. Fast implementation issues for these methods are discussed, including the use of QR decomposition for parameter estimation. The proposed methods are tested on data sets arising from hereditary breast cancer and small round blue-cell tumors. The results compare very favorably with currently used methods based on the normalized root-mean-square error. AVAILABILITY: The appendix is available from http://gspsnap.tamu.edu/gspweb/zxb/missing_zxb/ (user: gspweb; passwd: gsplab).

Algorithms↗

Visualizing conflicting evolutionary hypotheses in large collections of trees: using consensus networks to study the origins of placentals and hexapods.

Many phylogenetic methods produce large collections of trees as opposed to a single tree, which allows the exploration of support for various evolutionary hypotheses. However, to be useful, the information contained in large collections of trees should be summarized; frequently this is achieved by constructing a consensus tree. Consensus trees display only those signals that are present in a large proportion of the trees. However, by their very nature consensus trees require that any conflicts between the trees are necessarily disregarded. We present a method that extends the notion of consensus trees to allow the visualization of conflicting hypotheses in a consensus network. We demonstrate the utility of this method in highlighting differences amongst maximum likelihood bootstrap values and Bayesian posterior probabilities in the placental mammal phylogeny, and also in comparing the phylogenetic signal contained in amino acid versus nucleotide characters for hexapod monophyly.

Algorithms↗

A regularized discriminative model for the prediction of protein-peptide interactions.

MOTIVATION: Short well-defined domains known as peptide recognition modules (PRMs) regulate many important protein-protein interactions involved in the formation of macromolecular complexes and biochemical pathways. Since high-throughput experiments like yeast two-hybrid and phage display are expensive and intrinsically noisy, it would be desirable to more specifically target or partially bypass them with complementary in silico approaches. In the present paper, we present a probabilistic discriminative approach to predicting PRM-mediated protein-protein interactions from sequence data. The model is motivated by the discriminative model of Segal and Sharan as an alternative to the generative approach of Reiss and Schwikowski. In our evaluation, we focus on predicting the interaction network. As proposed by Williams, we overcome the problem of susceptibility to over-fitting by adopting a Bayesian a posteriori approach based on a Laplacian prior in parameter space. RESULTS: The proposed method was tested on two datasets of protein-protein interactions involving 28 SH3 domain proteins in Saccharmomyces cerevisiae, where the datasets were obtained with different experimental techniques. The predictions were evaluated with out-of-sample receiver operator characteristic (ROC) curves. In both cases, Laplacian regularization turned out to be crucial for achieving a reasonable generalization performance. The Laplacian-regularized discriminative model outperformed the generative model of Reiss and Schwikowski in terms of the area under the ROC curve on both datasets. The performance was further improved with a hybrid approach, in which our model was initialized with the motifs obtained with the method of Reiss and Schwikowski. AVAILABILITY: Software and supplementary material is available from http://lehrach.com/wolfgang/dmf.

Algorithms↗

Phylogenetics and evolution of nematode-trapping fungi (Orbiliales) estimated from nuclear and protein coding genes.

The systematic classification of nematode-trapping fungi is redefined based on phylogenies inferred from sequence analyses of 28S rDNA, 5.8S rDNA and beta-tubulin genes. Molecular data were analyzed with maximum parsimony, maximum likelihood and Bayesian analysis. An emended generic concept of nematode-trapping fungi is provided. Arthrobotrys is characterized by adhesive networks, Dactylellina by adhesive knobs, and Drechslerella by constricting-rings. Phylogenetic placement of taxa characterized by stalked adhesive knobs and non-constricting rings also is confirmed in Dactylellina. Species that produce unstalked adhesive knobs that grow out to form loops are transferred from Gamsylella to Dactylellina, and those that produce unstalked adhesive knobs that grow out to form networks are transferred from Gamsylella to Arthrobotrys. Gamsylella as currently circumscribed cannot be treated as a valid genus. A hypothesis for the evolution of trapping-devices is presented based on multiple gene data and morphological studies. Predatory and nonpredatory fungi appear to have been derived from nonpredatory members of Orbilia. The adhesive knob is considered to be the ancestral type of trapping device from which constricting rings and networks were derived via two pathways. In the first pathway adhesive knobs retained their adhesive material forming simple two-dimension networks, eventually forming complex three-dimension networks. In the second pathway adhesive knobs lost their adhesive materials, with their ends meeting to form nonconstricting rings and they in turn formed constricting rings with three inflated-cells.

Animals↗

Neonatal intensive care unit characteristics affect the incidence of severe intraventricular hemorrhage.

OBJECTIVES: The incidence of intraventricular hemorrhage (IVH), adjusted for known risk factors, varies across neonatal intensive care units (NICU)s. The effect of NICU characteristics on this variation is unknown. The objective was to assess IVH attributable risks at both patient and NICU levels. STUDY DESIGN: Subjects were <33 weeks' gestation, <4 days old on admission in the Canadian Neonatal Network database (all infants admitted in 1996-97 to 17 NICUs). The variation in severe IVH rates was analyzed using Bayesian hierarchical modeling for patient level and NICU level factors. RESULTS: Of 3772 eligible subjects, the overall crude incidence rates of grade 3-4 IVH was 8.3% (NICU range 2.0-20.5%). Male gender, extreme preterm birth, low Apgar score, vaginal birth, outborn birth, and high admission severity of illness accounted for 30% of the severe IVH rate variation; admission day therapy-related variables (treatment of acidosis and hypotension) accounted for an additional 14%. NICU characteristics, independent of patient level risk factors, accounted for 31% of the variation. NICUs with high patient volume and high neonatologist/staff ratio had lower rates of severe IVH. CONCLUSIONS: The incidence of severe IVH is affected by NICU characteristics, suggesting important new strategies to reduce this important adverse outcome.

Acute Disease↗

A novel neural network-based survival analysis model.

A feedforward neural network architecture aimed at survival probability estimation is presented which generalizes the standard, usually linear, models described in literature. The network builds an approximation to the survival probability of a system at a given time, conditional on the system features. The resulting model is described in a hierarchical Bayesian framework. Experiments with synthetic and real world data compare the performance of this model with the commonly used standard ones.

Bayes Theorem↗

A two-stage classifier for identification of protein-protein interface residues.

MOTIVATION: The ability to identify protein-protein interaction sites and to detect specific amino acid residues that contribute to the specificity and affinity of protein interactions has important implications for problems ranging from rational drug design to analysis of metabolic and signal transduction networks. RESULTS: We have developed a two-stage method consisting of a support vector machine (SVM) and a Bayesian classifier for predicting surface residues of a protein that participate in protein-protein interactions. This approach exploits the fact that interface residues tend to form clusters in the primary amino acid sequence. Our results show that the proposed two-stage classifier outperforms previously published sequence-based methods for predicting interface residues. We also present results obtained using the two-stage classifier on an independent test set of seven CAPRI (Critical Assessment of PRedicted Interactions) targets. The success of the predictions is validated by examining the predictions in the context of the three-dimensional structures of protein complexes.

Amino Acid Sequence↗

Establishing glucose- and ABA-regulated transcription networks in Arabidopsis by microarray analysis and promoter classification using a Relevance Vector Machine.

Establishing transcriptional regulatory networks by analysis of gene expression data and promoter sequences shows great promise. We developed a novel promoter classification method using a Relevance Vector Machine (RVM) and Bayesian statistical principles to identify discriminatory features in the promoter sequences of genes that can correctly classify transcriptional responses. The method was applied to microarray data obtained from Arabidopsis seedlings treated with glucose or abscisic acid (ABA). Of those genes showing >2.5-fold changes in expression level, approximately 70% were correctly predicted as being up- or down-regulated (under 10-fold cross-validation), based on the presence or absence of a small set of discriminative promoter motifs. Many of these motifs have known regulatory functions in sugar- and ABA-mediated gene expression. One promoter motif that was not known to be involved in glucose-responsive gene expression was identified as the strongest classifier of glucose-up-regulated gene expression. We show it confers glucose-responsive gene expression in conjunction with another promoter motif, thus validating the classification method. We were able to establish a detailed model of glucose and ABA transcriptional regulatory networks and their interactions, which will help us to understand the mechanisms linking metabolism with growth in Arabidopsis. This study shows that machine learning strategies coupled to Bayesian statistical methods hold significant promise for identifying functionally significant promoter sequences.

Abscisic Acid↗