Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Spontaneous quantal currents in a central neuron match predictions from binomial analysis of evoked responses.

Inhibitory postsynaptic currents occurring spontaneously in the teleost Mauthner cell were analyzed with the single-electrode voltage-clamp technique. They were collected during depolarizing steps and were outward-going; this procedure allowed them to be isolated from possible excitatory currents flowing in the opposite direction. Their amplitude histograms were found to exhibit regularly spaced multiple peaks, each of which had a Gaussian distribution of the same width. These compound inhibitory postsynaptic currents represent responses evoked by background firing of presynaptic neurons, and when tetrodotoxin was applied topically, only the first peak in the frequency histogram, which can be attributed to single exocytotic events, remained. The mean conductance of this quantal unit equalled 46.0 nS, which corresponds to the opening of 1000-2000 Cl- channels activated by glycine--the transmitter at these synapses. Its waveform and those of the larger units were essentially the same. Furthermore, in each set of data provided by a given Mauthner cell, the size of a quantum was quite constant, with its variance and those of further peaks being equivalent to that of background noise. These properties, which characterize the quantal events on the basis of spontaneous synaptic activity, were strikingly similar to those of the basic units derived by the simple binomial to those of the basic units derived by the simple binomial analysis of unitary postsynaptic potentials, thus validating the use of this statistical model to quantify the quantal nature of release at central connections. The quite straightforward method used here to extract single miniature currents from complex signals should be applicable to the other systems of the central nervous system, where the pertinence of this probabilistic model of release has yet to be demonstrated.

Animals↗

Drug resistance as a dynamic process in a model for multistep gene amplification under various levels of selection stringency.

Resistance to antineoplastic drugs has been a major impediment to the successful treatment of cancer. Recent studies suggest that several mechanisms are responsible for the emergence of drug resistance but that high levels of resistance and poor prognosis are strongly associated with gene or oncogene amplification. In this report we describe a probabilistic model for gene amplification in a tumor that grows under various drug protocols. The model is new in that it treats drug resistance as a dynamic process and examines specific assumptions about the underlying molecular events. Using this model, we specify the conditions for the emergence of drug-resistant mutants prior to selection as well as the relationship between the stringency of the selecting environment and the characteristics of the resultant cellular phenotype.

Antineoplastic Agents↗

The variation of pesticide residues in fruits and vegetables and the associated assessment of risk.

UNLABELLED: High levels of triazophos residues detected in carrots during routine monitoring led to the discovery of a wide variability between levels in individual roots. Conventional point estimates of consumer exposure were carried out. Due to the assumptions used, these calculations were likely to give rise to gross overestimates. In 1997, data were obtained for individual apples, pears, peaches, nectarines, oranges, bananas, and tomatoes that showed similar levels of variability in a range of organophosphate and carbamate residues. Point estimate models that had previously been used for intake estimates for carrots were not appropriate since it was necessary to take account of not only the variation of residue levels from crop item to crop item but also the variation in eating patterns in individual consumers. Probabilistic modeling was identified as a suitable way to produce multifactorial submodels and address some of the problems of combining distributions of consumption and residues. Consumption data from 1675 toddlers were linked with residue distributions from individual crop items not only to allow combinations of fruit consumed but also to allow for the variability in residue levels that occur between individual crop items. The model was also capable of taking account of the percentages of crops that did not contain any detectable residues; this information was available from initial screens of bulked samples and percentage of crop not treated in the case of carrots. The outputs from the models were given as percentages of consumers that could exceed a toxicological end point; this could be the acute reference dose or a factor of the no-observable-adverse-effect level. Modeling in this way was considered to give a realistic view of the likely short-term exposure and the output was used as an aid to decision making in terms of necessary regulatory action. BACKGROUND: As a result of high levels of triazophos detected in carrots during routine monitoring, studies were carried out to determine the variability of organophosphate residue levels in individual roots. Results obtained indicated that the highest residue levels could be 25 times the mean level in bulked samples (which were used in routine monitoring). Since sufficient levels of organophosphate compounds can give rise to toxicological effects after a single exposure, it was considered necessary to carry out assessments of short-term or acute consumer risk. At that time, models available worldwide were designed only to carry out point estimates of long-term exposure. From consumption data, it was possible to derive the levels of carrot consumption during a single day and calculations were carried out assuming all carrots contained the highest levels of residues found in trials. This led to a gross overestimate of likely exposure but was considered to give to intakes that eroded margins of safety; these were not a cause for extreme regulatory action. Further studies were carried out on other crops that may be eaten whole, at one sitting, and without processing to consider whether the large variability of organophosphate residues was a phenomenon that was common to other fruits and vegetables.

Animals↗

Learning overcomplete representations.

In an overcomplete basis, the number of basis vectors is greater than the dimensionality of the input, and the representation of an input is not a unique combination of basis vectors. Overcomplete representations have been advocated because they have greater robustness in the presence of noise, can be sparser, and can have greater flexibility in matching structure in the data. Overcomplete codes have also been proposed as a model of some of the response properties of neurons in primary visual cortex. Previous work has focused on finding the best representation of a signal using a fixed overcomplete basis (or dictionary). We present an algorithm for learning an overcomplete basis by viewing it as probabilistic model of the observed data. We show that overcomplete bases can yield a better approximation of the underlying statistical distribution of the data and can thus lead to greater coding efficiency. This can be viewed as a generalization of the technique of independent component analysis and provides a method for Bayesian reconstruction of signals in the presence of noise and for blind source separation when there are more sources than mixtures.

Algorithms↗

Sexual risk and HIV acquisition among men who have sex with men travelers to Key West, Florida: a mathematical modeling analysis.

The present study investigated the sexual risk behaviors of men who have sex with men (MSM) traveling to a popular gay tourist destination in the United States. In 2004, a brief survey was administered to 247 MSM tourists recruited from gay-oriented venues in Key West, Florida. Data collected included demographics, HIV status, length of stay, substance use, and sexual risk behaviors. A probabilistic model of HIV transmission was used to translate participants' reports of their sexual behaviors while in Key West into estimates of their risk of acquiring HIV. Twenty-two percent of participants reported anal sex with multiple partners over a relatively brief period (M = 4.1 days), and approximately one third reported having sex with a partner met during the vacation period. Modeling analyses suggested that sexual activity among vacationing MSM would account for approximately 201 new HIV infections among MSM visitors to Key West each year. Although previous studies have documented sexual risk behavior in travelers, quantitative estimates of the impact of these behaviors on the spread of HIV are lacking. Findings suggest that the risk-taking behavior of MSM on vacation may play an important role in the dissemination of HIV and other sexually transmitted diseases (STDs). Future research should assess additional factors (e.g., use of highly active antiretroviral therapy) that may affect HIV transmission in MSM travelers. In addition, efforts are needed to develop effective risk-reduction interventions for this population.

Adult↗

Probabilistic neural network model for the in silico evaluation of anti-HIV activity and mechanism of action.

A theoretical model has been developed that discriminates between active and nonactive drugs against HIV-1 with four different mechanisms of action for the active drugs. The model was built up using a probabilistic neural network (PNN) algorithm and a database of 2720 compounds. The model showed an overall accuracy of 97.34% in the training series, 85.12% in the selection series, and 84.78% in an external prediction series. The model not only correctly classified a very heterogeneous series of organic compounds but also discriminated between very similar active/nonactive chemicals that belong to the same family of compounds. More specifically, the model recognized 96.02% of nonactive compounds, 94.24% of active compounds that inhibited reverse transcriptase, 97.24% of protease inhibitors, 97.14% of virus uncoating inhibitors, and 90.32% of integrase inhibitors. The results indicate that this approach may represent a powerful tool for modeling large databases in QSAR with applications in medicinal chemistry.

Algorithms↗

Comparison of likelihood and Bayesian methods for estimating divergence times using multiple gene Loci and calibration points, with application to a radiation of cute-looking mouse lemur species.

Divergence time and substitution rate are seriously confounded in phylogenetic analysis, making it difficult to estimate divergence times when the molecular clock (rate constancy among lineages) is violated. This problem can be alleviated to some extent by analyzing multiple gene loci simultaneously and by using multiple calibration points. While different genes may have different patterns of evolutionary rate change, they share the same divergence times. Indeed, the fact that each gene may violate the molecular clock differently leads to the advantage of simultaneous analysis of multiple loci. Multiple calibration points provide the means for characterizing the local evolutionary rates on the phylogeny. In this paper, we extend previous likelihood models of local molecular clock for estimating species divergence times to accommodate multiple calibration points and multiple genes. Heterogeneity among different genes in evolutionary rate and in substitution process is accounted for by the models. We apply the likelihood models to analyze two mitochondrial protein-coding genes, cytochrome oxidase II and cytochrome b, to estimate divergence times of Malagasy mouse lemurs and related outgroups. The likelihood method is compared with the Bayes method of Thorne et al. (1998, Mol. Biol. Evol. 15:1647-1657), which uses a probabilistic model to describe the change in evolutionary rate over time and uses the Markov chain Monte Carlo procedure to derive the posterior distribution of rates and times. Our likelihood implementation has the drawbacks of failing to accommodate uncertainties in fossil calibrations and of requiring the researcher to classify branches on the tree into different rate groups. Both problems are avoided in the Bayes method. Despite the differences in the two methods, however, data partitions and model assumptions had the greatest impact on date estimation. The three codon positions have very different substitution rates and evolutionary dynamics, and assumptions in the substitution model affect date estimation in both likelihood and Bayes analyses. The results demonstrate that the separate analysis is unreliable, with dates variable among codon positions and between methods, and that the combined analysis is much more reliable. When the three codon positions were analyzed simultaneously under the most realistic models using all available calibration information, the two methods produced similar results. The divergence of the mouse lemurs is dated to be around 7-10 million years ago, indicating a surprisingly early species radiation for such a morphologically uniform group of primates.

Animals↗

The incomplete perfect phylogeny haplotype problem.

The problem of resolving genotypes into haplotypes, under the perfect phylogeny model, has been under intensive study recently. All studies so far handled missing data entries in a heuristic manner. We prove that the perfect phylogeny haplotype problem is NP-complete when some of the data entries are missing, even when the phylogeny is rooted. We define a biologically motivated probabilistic model for genotype generation and for the way missing data occur. Under this model, we provide an algorithm, which takes an expected polynomial time. In tests on simulated data, our algorithm quickly resolves the genotypes under high rates of missing entries.

Algorithms↗

Model-based particle picking for cryo-electron microscopy.

We describe an algorithm for finding particle images in cryo-EM micrographs. The algorithm starts from a crude 3D map of the target particle, computed from a relatively small number of manually picked images, and then projects the map in many different directions to give synthetic 2D templates. The templates are clustered and averaged and then cross-correlated with the micrographs. A probabilistic model of the imaging process then scores cross-correlation peaks to produce the final picks. We give quantitative results on two quite different target particles: keyhole limpet hemocyanin and p97 AAA ATPase. On these particles our automatic particle picker shows human performance level, as measured by the Fourier shell correlations of 3D reconstructions.

Adenosine Triphosphatases↗

Probability distributions and compartment boundaries in the development of Drosophila.

Clone analysis and fate mapping probe several properties of development. Here it is shown that data on fate mapping support a probabilistic model of cell commitment in Drosophila blastoderms. Adult cells have a distribution of possible ancestors, as W. Baker (1978b) inferred from the theory of compartment-boundary development (Garcia-Bellido, Ripoll and Morata 1973). Fate-map data are used here to describe quantitatively the ancestry distributions on the blastoderm fate map. The properties of the distributions are sensitive to, and probes of, developmental events, such as relative time of cellularization and time of commitment. The theory of this analysis shows first how the meaningful interpretation of the stage represented by a fate map depends on the assumptions made in mapping. A general mapping model described below makes it possible to evaluate several interpretations. Interestingly, the data require a 3-dimensional map, and it is argued that this must be due to an effect of the preblastoderm nuclear synctial stage. Second, the theory shows how compartment boundaries affect ancestry distribution and why they have no observable effect on mapping. Third, the variability implied by ancestry variance does not create too much "noise" to make meaningful maps of small areas; rather, oddly enough, it tends to magnify the apparent distances within small areas to make them more resolvable. Empirical results include probabilistic maps of the Drosophila blastoderm. These results argue that time of commitment varies even for cells in the same compartment, demonstrating the need for a more complex model of early development than that proposed in the compartment model. The results also help to evaluate the significance of compartment boundaries in respect to developmental commitment.

Animals↗

Probabilistic approach to determining unbiased random-coil carbon-13 chemical shift values from the protein chemical shift database.

We describe a probabilistic model for deriving, from the database of assigned chemical shifts, a set of random coil chemical shift values that are "unbiased" insofar as contributions from detectable secondary structure have been minimized (RCCSu). We have used this approach to derive a set of RCCSu values for 13Calpha and 13Cbeta for 17 of the 20 standard amino acid residue types by taking advantage of the known opposite conformational dependence of these parameters. We present a second probabilistic approach that utilizes the maximum entropy principle to analyze the database of 13Calpha and 13Cbeta chemical shifts considered separately; this approach yielded a second set of random coil chemical shifts (RCCSmax-ent). Both new approaches analyze the chemical shift database without reference to known structure. Prior approaches have used either the chemical shifts of small peptides assumed to model the random coil state (RCCSpeptide) or statistical analysis of chemical shifts associated with structure not in helical or strand conformation (RCCSstruct-stat). We show that the RCCSmax-ent values are strikingly similar to published RCCSpeptide and RCCSstruct-stat values. By contrast, the RCCSu values differ significantly from both published types of random coil chemical shift values. The differences (RCCSpeptide - RCCSu) for individual residue types show a correlation with known intrinsic conformational propensities. These results suggest that random coil chemical shift values from both prior approaches are biased by conformational preferences. RCCSu values appear to be consistent with the current concept of the "random coil" as the state in which the geometry of the polypeptide ensemble samples the allowed region of (phi, psi)-space in the absence of any dominant stabilizing interactions and thus represent an improved basis for the detection of secondary structure. Coupled with the growing database of chemical shifts, this probabilistic approach makes it possible to refine relationships among chemical shifts, their conformational propensities, and their dependence on pH, temperature, or neighboring residue type.

Carbon Isotopes↗

Focus on success: using a probabilistic approach to achieve an optimal balance of compound properties in drug discovery.

The success of any drug will depend on how closely it achieves an ideal combination of potency, selectivity, pharmacokinetics and safety. The key to achieving this success efficiently is to consider the overall balance of molecular properties of compounds against the ideal profile for the therapeutic indication from the earliest stages of a drug discovery project. The use of in silico predictive models of absorption, distribution, metabolism and elimination (ADME) and physicochemical properties is a major aid in this exercise, as it enables virtual molecules to be assessed across a broad range of properties from initial library generation, through to candidate selection. Of course, no measurement, whether in silico, in vitro or in vivo, is perfect and the uncertainties in any data should be explicitly taken into account when basing conclusions on test results. In addition, in the early stages of drug discovery, when designing a library that is lead seeking or building compound structure-activity relationships, the quality of any set of molecules should also be balanced against the chemical diversity covered. Here, a scheme is presented for achieving these goals based on a suite of predictive ADME models, probabilistic scoring and multiobjective optimisation for library design. The use of this platform for applications in lead identification and optimisation is illustrated.

Animals↗

Neonatal tolerance revisited by mathematical modelling.

The classical view of a neonatal tolerance window has recently been challenged by several studies showing that neonates can evoke normal immune responses. For example, neonatal immunity against the male antigen H-Y can be induced in female recipients by inoculation of male donor spleen cells enriched for professional antigen-presenting cells (APCs). In the same set-up, adult female recipients become tolerant by giving large doses of spleen cells. Using a probabilistic model, we here show how the number of T cells and endogenous APCs in the recipient, and the dose and quality of donor APCs can explain all observed phenomena. We thus reconcile the classical neonatal tolerance window with the recent data on neonatal immunity and adult tolerance.

Animals↗

Bayesian infinite mixture model based clustering of gene expression profiles.

MOTIVATION: The biologic significance of results obtained through cluster analyses of gene expression data generated in microarray experiments have been demonstrated in many studies. In this article we focus on the development of a clustering procedure based on the concept of Bayesian model-averaging and a precise statistical model of expression data. RESULTS: We developed a clustering procedure based on the Bayesian infinite mixture model and applied it to clustering gene expression profiles. Clusters of genes with similar expression patterns are identified from the posterior distribution of clusterings defined implicitly by the stochastic data-generation model. The posterior distribution of clusterings is estimated by a Gibbs sampler. We summarized the posterior distribution of clusterings by calculating posterior pairwise probabilities of co-expression and used the complete linkage principle to create clusters. This approach has several advantages over usual clustering procedures. The analysis allows for incorporation of a reasonable probabilistic model for generating data. The method does not require specifying the number of clusters and resulting optimal clustering is obtained by averaging over models with all possible numbers of clusters. Expression profiles that are not similar to any other profile are automatically detected, the method incorporates experimental replicates, and it can be extended to accommodate missing data. This approach represents a qualitative shift in the model-based cluster analysis of expression data because it allows for incorporation of uncertainties involved in the model selection in the final assessment of confidence in similarities of expression profiles. We also demonstrated the importance of incorporating the information on experimental variability into the clustering model. AVAILABILITY: The MS Windows(TM) based program implementing the Gibbs sampler and supplemental material is available at http://homepages.uc.edu/~medvedm/BioinformaticsSupplement.htm CONTACT: medvedm@email.uc.edu

Bayes Theorem↗

Integrated assessment and prediction of transcription factor binding.

Systematic chromatin immunoprecipitation (chIP-chip) experiments have become a central technique for mapping transcriptional interactions in model organisms and humans. However, measurement of chromatin binding does not necessarily imply regulation, and binding may be difficult to detect if it is condition or cofactor dependent. To address these challenges, we present an approach for reliably assigning transcription factors (TFs) to target genes that integrates many lines of direct and indirect evidence into a single probabilistic model. Using this approach, we analyze publicly available chIP-chip binding profiles measured for yeast TFs in standard conditions, showing that our model interprets these data with significantly higher accuracy than previous methods. Pooling the high-confidence interactions reveals a large network containing 363 significant sets of factors (TF modules) that cooperate to regulate common target genes. In addition, the method predicts 980 novel binding interactions with high confidence that are likely to occur in so-far untested conditions. Indeed, using new chIP-chip experiments we show that predicted interactions for the factors Rpn4p and Pdr1p are observed only after treatment of cells with methyl-methanesulfonate, a DNA-damaging agent. We outline the first approach for consistently integrating all available evidences for TF-target interactions and we comprehensively identify the resulting TF module hierarchy. Prioritizing experimental conditions for each factor will be especially important as increasing numbers of chIP-chip assays are performed in complex organisms such as humans, for which "standard conditions" are ill defined.

Algorithms↗

Profile-profile methods provide improved fold-recognition: a study of different profile-profile alignment methods.

To improve the detection of related proteins, it is often useful to include evolutionary information for both the query and target proteins. One method to include this information is by the use of profile-profile alignments, where a profile from the query protein is compared with the profiles from the target proteins. Profile-profile alignments can be implemented in several fundamentally different ways. The similarity between two positions can be calculated using a dot-product, a probabilistic model, or an information theoretical measure. Here, we present a large-scale comparison of different profile-profile alignment methods. We show that the profile-profile methods perform at least 30% better than standard sequence-profile methods both in their ability to recognize superfamily-related proteins and in the quality of the obtained alignments. Although the performance of all methods is quite similar, profile-profile methods that use a probabilistic scoring function have an advantage as they can create good alignments and show a good fold recognition capacity using the same gap-penalties, while the other methods need to use different parameters to obtain comparable performances.

Models, Chemical↗

Clustering ensembles: models of consensus and weak partitions.

Clustering ensembles have emerged as a powerful method for improving both the robustness as well as the stability of unsupervised classification solutions. However, finding a consensus clustering from multiple partitions is a difficult problem that can be approached from graph-based, combinatorial, or statistical perspectives. This study extends previous research on clustering ensembles in several respects. First, we introduce a unified representation for multiple clusterings and formulate the corresponding categorical clustering problem. Second, we propose a probabilistic model of consensus using a finite mixture of multinomial distributions in a space of clusterings. A combined partition is found as a solution to the corresponding maximum-likelihood problem using the EM algorithm. Third, we define a new consensus function that is related to the classical intraclass variance criterion using the generalized mutual information definition. Finally, we demonstrate the efficacy of combining partitions generated by weak clustering algorithms that use data projections and random data splits. A simple explanatory model is offered for the behavior of combinations of such weak clustering components. Combination accuracy is analyzed as a function of several parameters that control the power and resolution of component partitions as well as the number of partitions. We also analyze clustering ensembles with incomplete information and the effect of missing cluster labels on the quality of overall consensus. Experimental results demonstrate the effectiveness of the proposed methods on several real-world data sets.

Algorithms↗

High-resolution computational models of genome binding events.

Direct physical information that describes where transcription factors, nucleosomes, modified histones, RNA polymerase II and other key proteins interact with the genome provides an invaluable mechanistic foundation for understanding complex programs of gene regulation. We present a method, joint binding deconvolution (JBD), which uses additional easily obtainable experimental data about chromatin immunoprecipitation (ChIP) to improve the spatial resolution of the transcription factor binding locations inferred from ChIP followed by DNA microarray hybridization (ChIP-Chip) data. Based on this probabilistic model of binding data, we further pursue improved spatial resolution by using sequence information. We produce positional priors that link ChIP-Chip data to sequence data by guiding motif discovery to inferred protein-DNA binding sites. We present results on the yeast transcription factors Gcn4 and Mig2 to demonstrate JBD's spatial resolution capabilities and show that positional priors allow computational discovery of the Mig2 motif when a standard approach fails.

Base Sequence↗