Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Predicting the prognosis of breast cancer by integrating clinical and microarray data with Bayesian networks.

MOTIVATION: Clinical data, such as patient history, laboratory analysis, ultrasound parameters--which are the basis of day-to-day clinical decision support--are often underused to guide the clinical management of cancer in the presence of microarray data. We propose a strategy based on Bayesian networks to treat clinical and microarray data on an equal footing. The main advantage of this probabilistic model is that it allows to integrate these data sources in several ways and that it allows to investigate and understand the model structure and parameters. Furthermore using the concept of a Markov Blanket we can identify all the variables that shield off the class variable from the influence of the remaining network. Therefore Bayesian networks automatically perform feature selection by identifying the (in)dependency relationships with the class variable. RESULTS: We evaluated three methods for integrating clinical and microarray data: decision integration, partial integration and full integration and used them to classify publicly available data on breast cancer patients into a poor and a good prognosis group. The partial integration method is most promising and has an independent test set area under the ROC curve of 0.845. After choosing an operating point the classification performance is better than frequently used indices.

Bayes Theorem↗

Analysis of HIV-1 pol sequences using Bayesian Networks: implications for drug resistance.

Human Immunodeficiency Virus-1 (HIV-1) antiviral resistance is a major cause of antiviral therapy failure and compromises future treatment options. As a consequence, resistance testing is the standard of care. Because of the high degree of HIV-1 natural variation and complex interactions, the role of resistance mutations is in many cases insufficiently understood. We applied a probabilistic model, Bayesian networks, to analyze direct influences between protein residues and exposure to treatment in clinical HIV-1 protease sequences from diverse subtypes. We can determine the specific role of many resistance mutations against the protease inhibitor nelfinavir, and determine relationships between resistance mutations and polymorphisms. We can show for example that in addition to the well-known major mutations 90M and 30N for nelfinavir resistance, 88S should not be treated as 88D but instead considered as a major mutation and explain the subtype-dependent prevalence of the 30N resistance pathway.

Amino Acid Sequence↗

Integrated monitoring and analysis for early warning of patient deterioration.

Recently there has been an upsurge of interest in strategies for detecting at-risk patients in order to trigger the timely intervention of a Medical Emergency Team (MET), also known as a Rapid Response Team (RRT). We review a real-time automated system, BioSign, which tracks patient status by combining information from vital signs monitored non-invasively on the general ward. BioSign fuses the vital signs in order to produce a single-parameter representation of patient status, the Patient Status Index. The data fusion method adopted in BioSign is a probabilistic model of normality in five dimensions, previously learnt from the vital sign data acquired from a representative sample of patients. BioSign alerts occur either when a single vital sign deviates by close to +/-3 standard deviations from its normal value or when two or more vital signs depart from normality, but by a smaller amount. In a trial with high-risk elective/emergency surgery or medical patients, BioSign alerts were generated, on average, every 8 hours; 95% of these were classified as 'True' by clinical experts. Retrospective analysis has also shown that the data fusion algorithm in BioSign is capable of detecting critical events in advance of single-channel alerts.

Critical Care↗

The evolutionary gain of spliceosomal introns: sequence and phase preferences.

Theories regarding the evolution of spliceosomal introns differ in the extent to which the distribution of introns reflects either a formative role in the evolution of protein-coding genes or the adventitious gain of genetic elements. Here, systematic methods are used to assess the causes of the present-day distribution of introns in 10 families of eukaryotic protein-coding genes comprising 1,868 introns in 488 distinct alignment positions. The history of intron evolution inferred using a probabilistic model that allows ancestral inheritance of introns, gain of introns, and loss of introns reveals that the vast majority of introns in these eukaryotic gene families were not inherited from the most recent common ancestral genes, but were gained subsequently. Furthermore, among inferred events of intron gain that meet strict criteria of reliability, the distribution of sites of gain with respect to reading-frame phase shows a 5:3:2 ratio of phases 0, 1 and 2, respectively, and exhibits a nucleotide preference for MAG GT (positions -3 to +2 relative to the site of gain). The nucleotide preferences of intron gain may prove to be the ultimate cause for the phase bias. The phase bias of intron gain is sufficient to account quantitatively for the well-known 5:3:2 bias in phase frequencies among extant introns, a conclusion that holds even when taxonomic heterogeneity in phase patterns is considered. Thus, intron gain accounts for the vast majority of extant introns and for the bias toward phase 0 introns that previously was interpreted as evidence for ancient formative introns.

Animals↗

Inference and analysis of the relative stability of bacterial chromosomes.

The stability of genomes is highly variable, both in terms of gene content and gene order. Here I calibrate the loss of gene order conservation (GOC) through time by fitting a simple probabilistic model on pairwise comparisons involving 126 bacterial genomes. The model computes the probability of separation of pairs of contiguous genes per unit of time and fits the data better than previous ones while allowing a mechanistic interpretation for the loss of GOC with time. Although the information on operons is not used in the model, I observe, as expected, that most highly conserved pairs of genes are indeed within operons. However, even the other pairs are much more conserved than expected given the observed experimental rearrangement rates. After 500 Myr, about 50% of the originally contiguous orthologues remain so in the average genome. Hence, the large majority of rearrangements must be deleterious and random genome rearrangements are unlikely to provide for positively selected structural changes. I then use the deviations from the model to define an intrinsic measure of genome stability that allowed the comparison of distantly related genomes and the inference of ancestral states. This shows that clades differ in genome stability, with cyanobacteria being the least stable and gamma-proteobacteria the most stable. Without correction for phylogeny, free-living bacteria are the least stable group of genomes, followed by pathogens, and then endomutualists. However, after correction for phylogenetic inertia (or the removal of cyanobacteria from the analysis), there is no significant association between genome stability and lifestyle or genome size. Hence, although this method has allowed uncovering some of mechanisms leading to rearrangements, we still ignore the forces that differentially shape selection upon genome stability in different species.

Bacteria↗

Dependence among sites in RNA evolution.

Although probabilistic models of genotype (e.g., DNA sequence) evolution have been greatly elaborated, less attention has been paid to the effect of phenotype on the evolution of the genotype. Here we propose an evolutionary model and a Bayesian inference procedure that are aimed at filling this gap. In the model, RNA secondary structure links genotype and phenotype by treating the approximate free energy of a sequence folded into a secondary structure as a surrogate for fitness. The underlying idea is that a nucleotide substitution resulting in a more stable secondary structure should have a higher rate than a substitution that yields a less stable secondary structure. This free energy approach incorporates evolutionary dependencies among sequence positions beyond those that are reflected simply by jointly modeling change at paired positions in an RNA helix. Although there is not a formal requirement with this approach that secondary structure be known and nearly invariant over evolutionary time, computational considerations make these assumptions attractive and they have been adopted in a software program that permits statistical analysis of multiple homologous sequences that are related via a known phylogenetic tree topology. Analyses of 5S ribosomal RNA sequences are presented to illustrate and quantify the strong impact that RNA secondary structure has on substitution rates. Analyses on simulated sequences show that the new inference procedure has reasonable statistical properties. Potential applications of this procedure, including improved ancestral sequence inference and location of functionally interesting sites, are discussed.

Animals↗

Toucan: deciphering the cis-regulatory logic of coregulated genes.

TOUCAN is a Java application for the rapid discovery of significant cis-regulatory elements from sets of coexpressed or coregulated genes. Biologists can automatically (i) retrieve genes and intergenic regions, (ii) identify putative regulatory regions, (iii) score sequences for known transcription factor binding sites, (iv) identify candidate motifs for unknown binding sites, and (v) detect those statistically over-represented sites that are characteristic for a gene set. Genes or intergenic regions are retrieved from Ensembl or EMBL, together with orthologs and supporting information. Orthologs are aligned and syntenic regions are selected as candidate regulatory regions. Putative sites for known transcription factors are detected using our MotifScanner, which scores position weight matrices using a probabilistic model. New motifs are detected using our MotifSampler based on Gibbs sampling. Binding sites characteristic for a gene set--and thus statistically over-represented with respect to a reference sequence set--are found using a binomial test. We have validated Toucan by analyzing muscle-specific genes, liver-specific genes and E2F target genes; we have easily detected many known binding sites within intergenic DNA and identified new biologically plausible sites for known and unknown transcription factors. Software available at http://www.esat.kuleuven.ac. be/ approximately dna/BioI/Software.html.

Algorithms↗

EUGENE'HOM: A generic similarity-based gene finder using multiple homologous sequences.

EUGENE'HOM is a gene prediction software for eukaryotic organisms based on comparative analysis. EUGENE'HOM is able to take into account multiple homologous sequences from more or less closely related organisms. It integrates the results of TBLASTX analysis, splice site and start codon prediction and a robust coding/non-coding probabilistic model which allows EUGENE'HOM to handle sequences from a variety of organisms. The current target of EUGENE'HOM is plant sequences. The EUGENE'HOM web site is available at http://genopole.toulouse.inra.fr/bioinfo/eugene/EuGeneHom/cgi-bin/EuGeneHom.pl.

Algorithms↗

Genome-wide searching for pseudouridylation guide snoRNAs: analysis of the Saccharomyces cerevisiae genome.

One of the largest families of small RNAs in eukaryotes is the H/ACA small nucleolar RNAs (snoRNAs), most of which guide RNA pseudouridine formation. So far, an effective computational method specifically for identifying H/ACA snoRNA gene sequences has not been established. We have developed snoGPS, a program for computationally screening genomic sequences for H/ACA guide snoRNAs. The program implements a deterministic screening algorithm combined with a probabilistic model to score gene candidates. We report here the results of testing snoGPS on the budding yeast Saccharomyces cerevisiae. Six candidate snoRNAs were verified as novel RNA transcripts, and five of these were verified as guides for pseudouridine formation at specific sites in ribosomal RNA. We also predicted 14 new base-pairings between snoRNAs and known pseudouridine sites in S.cerevisiae rRNA, 12 of which were verified by gene disruption and loss of the cognate pseudouridine site. Our findings include the first prediction and verification of snoRNAs that guide pseudouridine modification at more than two sites. With this work, 41 of the 44 known pseudouridine modifications in S.cerevisiae rRNA have been linked with a verified snoRNA, providing the most complete accounting of the H/ACA snoRNAs that guide pseudouridylation in any species.

Algorithms↗

Plasmodium interspersed repeats: the major multigene superfamily of malaria parasites.

Functionally related homologues of known genes can be difficult to identify in divergent species. In this paper, we show how multi-character analysis can be used to elucidate the relationships among divergent members of gene superfamilies. We used probabilistic modelling in conjunction with protein structural predictions and gene-structure analyses on a whole-genome scale to find gene homologies that are missed by conventional similarity-search strategies and identified a variant gene superfamily in six species of malaria (Plasmodium interspersed repeats, pir). The superfamily includes rif in P.falciparum, vir in P.vivax, a novel family kir in P.knowlesi and the cir/bir/yir family in three rodent malarias. Our data indicate that this is the major multi-gene family in malaria parasites. Protein localization of products from pir members to the infected erythrocyte membrane in the rodent malaria parasite P.chabaudi, demonstrates phenotypic similarity to the products of pir in other malaria species. The results give critical insight into the evolutionary adaptation of malaria parasites to their host and provide important data for comparative immunology between malaria parasites obtained from laboratory models and their human counterparts.

Amino Acid Motifs↗

Comparative genomics analysis of NtcA regulons in cyanobacteria: regulation of nitrogen assimilation and its coupling to photosynthesis.

We have developed a new method for prediction of cis-regulatory binding sites and applied it to predicting NtcA regulated genes in cyanobacteria. The algorithm rigorously utilizes concurrence information of multiple binding sites in the upstream region of a gene and that in the upstream regions of its orthologues in related genomes. A probabilistic model was developed for the evaluation of prediction reliability so that the prediction false positive rate could be well controlled. Using this method, we have predicted multiple new members of the NtcA regulons in nine sequenced cyanobacterial genomes, and showed that the false positive rates of the predictions have been reduced on an average of 40-fold compared to the conventional methods. A detailed analysis of the predictions in each genome showed that a significant portion of our predictions are consistent with previously published results about individual genes. Intriguingly, NtcA promoters are found for many genes involved in various stages of photosynthesis. Although photosynthesis is known to be tightly coordinated with nitrogen assimilation, very little is known about the underlying mechanism. We postulate for the fist time that these genes serve as the regulatory points to orchestrate these two important processes in a cyanobacterial cell.

Algorithms↗

The Mouse Functional Genome Database (MfunGD): functional annotation of proteins in the light of their cellular context.

MfunGD (http://mips.gsf.de/genre/proj/mfungd/) provides a resource for annotated mouse proteins and their occurrence in protein networks. Manual annotation concentrates on proteins which are found to interact physically with other proteins. Accordingly, manually curated information from a protein-protein interaction database (MPPI) and a database of mammalian protein complexes is interconnected with MfunGD. Protein function annotation is performed using the Functional Catalogue (FunCat) annotation scheme which is widely used for the analysis of protein networks. The dataset is also supplemented with information about the literature that was used in the annotation process as well as links to the SIMAP Fasta database, the Pedant protein analysis system and cross-references to external resources. Proteins that so far were not manually inspected are annotated automatically by a graphical probabilistic model and/or superparamagnetic clustering. The database is continuously expanding to include the rapidly growing amount of functional information about gene products from mouse. MfunGD is implemented in GenRE, a J2EE-based component-oriented multi-tier architecture following the separation of concern principle.

Animals↗

Stubb: a program for discovery and analysis of cis-regulatory modules.

Given the DNA-binding specificities (motifs) of one or more transcription factors, an important bioinformatics problem is to discover significant clusters of binding sites for the transcription factors(s). Such clusters often correspond to cis-regulatory modules mediating regulation of an adjacent gene. In earlier work, we developed the Stubb program that uses a probabilistic model and a maximum likelihood approach to efficiently detect cis-regulatory modules over genomic scales. It may optionally exploit a second related genome to improve module prediction accuracy. We describe here the use of a web-based interface for the Stubb program. The interface is equipped with a special post-processing step for in-depth analysis of specific modules, in order to reveal individual binding sites predicted in the module. The web server may be accessed at the URL http://stubb.rockefeller.edu/.

Algorithms↗

Gene loss rate: a probabilistic measure for the conservation of eukaryotic genes.

The rate of conservation of a gene in evolution is believed to be correlated with its biological importance. Recent studies have devised various conservation measures for genes and have shown that they are correlated with several biological characteristics of functional importance. Specifically, the state-of-the-art propensity for gene loss (PGL) measure was shown to be strongly correlated with gene essentiality and its number of protein-protein interactions (PPIs). The observed correlation between conservation and functional importance varies however between conservation measures, underscoring the need for accurate and general measures for the rate of gene conservation. Here we develop a novel maximum-likelihood approach to computing the rate in which a gene is lost in evolution, motivated by the same principles as those underlying PGL. However, in difference to PGL which considers only the most parsimonious ancestral states of the internal nodes of the phylogenetic tree relating the species, our approach weighs in a probabilistic manner all possible ancestral states, and includes the branch length information as part of the probabilistic model. In application to data of 16 eukaryotic genomes, our approach shows higher correlations with experimental data than PGL, including data on gene lethality, level of connectivity in a PPI network and coherence within functionally related genes.

Algorithms↗

Heterosexual transmission of human immunodeficiency virus: variability of infectivity throughout the course of infection. European Study Group on Heterosexual Transmission of HIV.

Although individuals infected with human immunodeficiency virus (HIV) seem to be more infectious in the late stages of HIV infection and possibly also during the seroconversion period, most estimates of per-sexual-contact infectivity have been obtained without allowing for variability over the course of infection. In this analysis, a probabilistic model was fitted to data from a European study carried out between 1987 and 1992 that involved 499 (359 males and 140 females) HIV-infected subjects (index cases) and their regular heterosexual partners. The model used allowed infectivity (the per-sexual-contact HIV transmission probability, mu) to vary through three stages: the first 3 months following infection, the subsequent asymptomatic period, and the advanced stage (HIV-related clinical symptoms or a CD4-positive T lymphocyte count less than 200/mm3). Male-to-female infectivity through penile-anal sex was found to be higher in both the early and advanced stages of infection (mu=0.183) than in the longer intermediate period (mu=0.014) (p < 0.03). Failure to demonstrate significant differences between stages for other types of contact (male-to-female penile-vaginal contacts: mu=0.0007; female-to-male transmission: mu=0.0005) may reflect insufficient power rather than a true lack of variability. Indeed, the results for penile-anal sex suggest that persons who are in the process of seroconverting may be much more infectious than asymptomatic infected persons, whatever the type of contact. Prevention education should stress the risk of HIV transmission from subjects who may be unaware of their infection.

Acquired Immunodeficiency Syndrome↗

Control for environmental risk factors in assessing genetic effects on disease familial aggregation.

A probabilistic model was developed to assess the impact of two independent dichotomous familial risk factors on familial aggregation of a disorder in pairs of relatives where one member was ascertained as a proband or index subject (i.e., a case or control). Under this model, one risk factor is of primary interest (i.e., a susceptibility gene), while the effect of the other is to be controlled (i.e., an environmental risk factor). Familial aggregation was examined within strata defined by the status of proband and relative with respect to the environmental factor. The findings suggest that for proband-relative pairs, under both the additive and multiplicative models, an environmental factor can be controlled in the analysis based solely on the status of the proband. If the relation between the genetic and environmental factors is neither additive nor multiplicative, however, the analysis must take account of environmental risk factors in both proband and relative.

Bias↗

Mean dose to lymphocytes during radiotherapy treatments.

Using a probabilistic model with parameters from four radiotherapy protocols used in Mexican hospitals for the treatment of cervical cancer, we have calculated the distribution of dose to cells in peripheral blood of patients. Values of the mean dose to the lymphocytes during and after a 60Co treatment are compared to estimates from an in vivo chromosome aberration study performed on five patients. Calculations indicate that the mean dose to the circulating blood is about 2% of the tumor dose, while the mean dose to recirculating lymphocytes may reach up to 7% of the tumor dose. Differences up to a factor of two in the dose to the blood are predicted for different protocols delivering equal tumor doses. The data suggest mean doses higher than the predictions of the model.

Chromosome Aberrations↗

Validation of a formula that calculates the estimated risk of respiratory distress syndrome.

OBJECTIVE: Several groups, including ours, have developed probabilistic models that incorporate both the surfactant-to-albumin ratio (TDx-FLM II) and gestational age to more accurately predict the risk of neonatal respiratory distress syndrome (RDS) and eliminate the current categorical "immature"/"indeterminate"/"mature" interpretation. We validate our model using a separate data set, with the goal of providing the clinician with a risk score. METHODS: The medical records of all women who had TDx-FLM II testing performed at Brigham and Women's Hospital between January 1, 2003, and December 31, 2005, were reviewed to gather a population upon which to validate our previous logistic regression model. Receiver operating characteristic curve and Hosmer-Lemeshow analysis was conducted to determine the performance of our model and another model in this new population. RESULTS: A total of 233 mother-neonate pairs (21 RDS, 212 non-RDS) met criteria for analysis. The receiver operating characteristic analysis illustrated that our previous formula was a strong predictor of the risk of RDS with an area under the curve of 0.902 (95% confidence interval 0.849-0.955). In addition, using the Hosmer-Lemeshow analysis, our formula produced an excellent overall fit (P=.95), whereas another published model was a poor fit to our data (P=.002). CONCLUSION: Our previously derived logistic regression model formula incorporating TDx-FLM II results and gestational age to predict risk of neonatal respiratory distress syndrome was robust and stable over time in an independent data set. The results suggest that the equation can be implemented clinically to assist physicians and patients and used by other institutions after their own internal validation. LEVEL OF EVIDENCE: III.

Humans↗