Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Models of protein sequence evolution and their applications.

Homologous sequences are correlated due to their common ancestry. Probabilistic models of sequence evolution are employed routinely to properly account for these phylogenetic correlations. These increasingly realistic models provide a basis for studying evolution and for exploiting it to better understand protein structure and function. Notable recent advances have been made in the treatment of insertion and deletion events, the estimation of amino-acid replacement rates, and the detection of positive selection.

Amino Acid Sequence↗

Modeling splice sites with Bayes networks.

MOTIVATION: The main goal in this paper is to develop accurate probabilistic models for important functional regions in DNA sequences (e.g. splice junctions that signal the beginning and end of transcription in human DNA). These methods can subsequently be utilized to improve the performance of gene-finding systems. The models built here attempt to model long-distance dependencies between non-adjacent bases. RESULTS: An efficient modeling method is described which models biological data more accurately than a first-order Markov model without increasing the number of parameters. Intuitively, a small number of parameters helps a learning system to avoid overfitting. Several experiments with the model are presented, which show a small improvement in the average accuracy as compared with a simple Markov model. These experiments suggest that single long distance dependencies do not help the recognition problem, thus confirming several previous studies which have used more heuristic modeling techniques. AVAILABILITY: This software is available for downloaded and as a web resource at http://www.ai.uic.edu/software CONTACT: kasif@eecs.uic.edu

Bayes Theorem↗

Entropic analysis of biological growth models.

Using the concept of isomorphism, we provide a quantitative description of the thermodynamic characteristics of the Logistic, Bertalanffy, and Gompertz biological growth models. With the help of the entropy expressions derived from the isomorphic probabilistic models, we show that the Logistic, Bertalanffy, and Gompertz growth models have distinct thermodynamic characteristics, though they have similar sigmoid trajectories in the time domain. The entropic analysis further reveals that the Gompertz model corresponds to a completely open system, in which cells receive adequate nutrition and competition among cells can be neglected. We have also applied the present entropic analysis to the modeling of tumor growth. Our results suggest that the entropic expression provides a theoretical justification for using the Gompertz model for describing the development of tumor cell populations. The entropic analysis may play an important role in studying the mechanisms of different biological growth processes.

Animals↗

Approach to stochastic modelling of consumer exposure for any substance from canned foods using simulant migration data.

A two-dimensional probabilistic model was constructed to estimate the short-term dietary exposure of UK consumers to any generalized migrant from coated light metal food packaging. Using three UK National Dietary and Nutrition Surveys (NDNS) comprising 4-7-day dietary surveys for different age and gender groups, actual body weights and survey years, a sample representative of the dietary consumption of the UK population was obtained comprising around 4,200 food items. Interrogation of the raw data showed that the per capita consumption of food and beverage for an adult was 2.9 kg per person day(-1), which is comparable with the US FDA value of 3.0 kg. The packaging type of each food item was assigned from the survey descriptions or by sampling from distributions based upon market share information and expert judgement. Each food item was assigned to the relevant food simulant: A (aqueous), B (acidic) or D (fatty), so that simulant migration data could be used. The exposure model was used to evaluate exposure for a given level of migration and, conversely, the level of migration that could be tolerated whilst keeping within a target threshold exposure level. As examples, migration at 10 microg dm(-2) into fatty foods only resulted in an exposure ranging from 0.06 to 0.22 microg kg(-1) body (actual) weight day(-1) depending on the scenario. The model revealed that if migration from metal coatings was only into fatty foods, migration in the range 1.83-4.95 microg dm(2) (97.5th percentile, depending on the scenario) would give an exposure of less than 1.5 microg per person day(-1). This is a toxicological threshold limit used in the USA. If migration into simulants A and B is also considered to be at the same level as that for simulant D, then the level of migration for the threshold to be reached is, not surprisingly, lower (0.64-0.87 microg dm(-2)) than that if migration were only into fatty foods. In this case, clearly the main contributors to the exposure were foodstuffs represented by simulants A and B because of their importance in the diet. These estimates are based on 4- or 7-day food diaries and chronic exposure over the long-term would be expected to be lower.

Adolescent↗

A model for evaluating the effect of son or daughter preference on population size.

This paper develops a micro probabilistic model to describe the family extension process. The parameters are: the probability that a newborn is a boy (P), the number of desired boys (B), the number of desired girls (G), and the maximum possible number of children (N). This maximum is a stopping rule rather than a biological maximum. The variables are the ultimate number of boys and girls. According to this model, each couple determines B, G, and N, at the beginning of the reproductive period and continues to reproduce until at least B and at least G are achieved, or until the total number of children reaches N. The probability distribution of the ultimate number of boys and girls in the population is derived for this model. Simulation techniques are used to generate offspring. The results showed that the population size increases with the absolute differences, [B-G] for fixed N. They also suggested that son or daughter preference may be an important factor in fertility determinants, which may have important implications in population policies.

Computer Simulation↗

Predictors of nonsentinel lymph node positivity in patients with a positive sentinel node for melanoma.

BACKGROUND: Patients found to harbor melanoma micrometastases in the sentinel lymph node (SLN) are recommended to proceed to complete lymph node dissection (CLND), although the majority of patients will have no additional disease identified in the nonsentinel lymph nodes (NSLNs). We sought to assess predictive factors associated with finding positive NSLNs, and identify a subset of patients with low likelihood of finding additional disease on CLND. STUDY DESIGN: We queried our prospective melanoma database for patients from January 1996 to August 2003 with a positive SLN. Univariable logistic regression models were fit for multiple factors and a positive NSLN. To derive a probabilistic model for occurrence of one or more positive NSLN(s), a multivariable logistic model was fit using a stepwise variable selection method. RESULTS: Of 980 patients who underwent SLN biopsy for cutaneous melanoma, 232 (24%) had a positive SLN; 221 (23%) followed by CLND. Of these patients, 34 (15%) had one or more positive NSLN(s). In multivariable analysis, male gender (odds ratio [OR] 3.6 [95% CI 1.33, 9.71]; p = 0.01), Breslow thickness (OR 4.58 [95% CI 1.28, 16.36]; p = 0.019), extranodal extension (OR 3.2 [95% CI 1.0, 10.5]; p = 0.05), and three or more positive sentinel nodes (OR 65.81 [95% CI 5.2, 825.7]; p = 0.001) were all associated with the likelihood of finding additional positive nodes on CLND. Of 47 patients with minimal tumor burden in the SLN, only 1 (2%) had additional disease in the NSLN. CONCLUSIONS: These results provide additional data to plan clinical trials to answer the question of who can safely avoid CLND after a positive SLN. Patients with minimal tumor burden in the SLN might be the most likely group, although defining "minimal tumor burden" must be standardized. Serial sectioning and immunohistochemistry on the NSLN in any "low-risk" group must be performed in a clinical trial to confirm that residual disease is unlikely before avoiding CLND can be recommended.

Adolescent↗

MONKEY: identifying conserved transcription-factor binding sites in multiple alignments using a binding site-specific evolutionary model.

We introduce a method (MONKEY) to identify conserved transcription-factor binding sites in multispecies alignments. MONKEY employs probabilistic models of factor specificity and binding-site evolution, on which basis we compute the likelihood that putative sites are conserved and assign statistical significance to each hit. Using genomes from the genus Saccharomyces, we illustrate how the significance of real sites increases with evolutionary distance and explore the relationship between conservation and function.

Binding Sites↗

Composite dependency-reflecting model for core promoter recognition in vertebrate genomic DNA sequences.

This paper deals with the development of a predictive probabilistic model, a composite dependency-reflecting model (CDRM), which was designed to detect core promoter regions and transcription start sites (TSS) in vertebrate genomic DNA sequences, an issue of some importance for genome annotation. The model actually represents a combination of first-, second-, third- and much higher order or long-range dependencies obtained using the expanded maximal dependency decomposition (EMDD) procedure, which iteratively decomposes data sets into subsets on the basis of dependency degree and patterns inherent in the target promoter region to be modeled. In addition, decomposed subsets are modeled by using a first-order Markov model, allowing the predictive model to reflect dependency between adjacent positions explicitly. In this way, the CDRM allows for potentially complex dependencies between positions in the core promoter region. Such complex dependencies may be closely related to the biological and structural contexts since promoter elements are present in various combinations separated by various distances in the sequence. Thus, CDRM may be appropriate for recognizing core promoter regions and TSSs in vertebrate genomic contig. To demonstrate the effectiveness of our algorithm, we tested it using standardized data and real core promoters, and compared it with some current representative promoter-finding algorithms. The developed algorithm showed better accuracy in terms of specificity and sensitivity than the promoter-finding ones used in performance comparison.

Algorithms↗

Chromatography as Lévy stochastic process.

The Stochastic Theory of Chromatography has been revised in light of some of the most relevant Lévy's findings in Theory of Probability, including the so-called Lévy's distance, the characteristic function and the theory of infinitesimally divisible distributions. These concepts represent the key to exploit and understand, at a molecular basis, phenomena typical of chromatographic separations under linear conditions, such as peak tailing and splitting. In particular, Lévy's distance has been used to quantify the degree of convergence of real peaks towards an ideal Gaussian shape; the characteristic function properties, introduced by Lévy to deal with the problem of the addition of independent random variables, have been employed to solve a wide variety of chromatographic models (including adsorption on heterogeneous surfaces) and to interpret mobile phase dispersion from a probabilistic point of view. Finally, Lévy's studies concerning infinitesimally divisible distributions have allowed to introduce in the stochastic description of chromatography, effects associated to dispersion in mobile phase. It has been demonstrated that, according to Lévy's canonical representation of stochastic processes, the basis of chromatography is a mobile phase Poisson Process. Represented as a Lévy's process, the microscopic-probabilistic model of chromatography permits the establishment of a connection between single-molecule properties and their statistical fluctuations and shapes of real chromatographic peaks allowing, at the same time, for the constitution of a link between different branches of physical sciences.

Chromatography↗

Biclustering microarray data by Gibbs sampling.

MOTIVATION: Gibbs sampling has become a method of choice for the discovery of noisy patterns, known as motifs, in DNA and protein sequences. Because handling noise in microarray data presents similar challenges, we have adapted this strategy to the biclustering of discretized microarray data. RESULTS: In contrast with standard clustering that reveals genes that behave similarly over all the conditions, biclustering groups genes over only a subset of conditions for which those genes have a sharp probability distribution. We have opted for a simple probabilistic model of the biclusters because it has the key advantage of providing a transparent probabilistic interpretation of the biclusters in the form of an easily interpretable fingerprint. Furthermore, Gibbs sampling does not suffer from the problem of local minima that often characterizes Expectation-Maximization. We demonstrate the effectiveness of our approach on two synthetic data sets as well as a data set from leukemia patients.

Biomarkers, Tumor↗

Weighted neighbor joining: a likelihood-based approach to distance-based phylogeny reconstruction.

We introduce a distance-based phylogeny reconstruction method called "weighted neighbor joining," or "Weighbor" for short. As in neighbor joining, two taxa are joined in each iteration; however, the Weighbor criterion for choosing a pair of taxa to join takes into account that errors in distance estimates are exponentially larger for longer distances. The criterion embodies a likelihood function on the distances, which are modeled as correlated Gaussian random variables with different means and variances, computed under a probabilistic model for sequence evolution. The Weighbor criterion consists of two terms, an additivity term and a positivity term, that quantify the implications of joining the pair. The first term evaluates deviations from additivity of the implied external branches, while the second term evaluates confidence that the implied internal branch has a positive branch length. Compared with maximum-likelihood phylogeny reconstruction, Weighbor is much faster, while building trees that are qualitatively and quantitatively similar. Weighbor appears to be relatively immune to the "long branches attract" and "long branch distracts" drawbacks observed with neighbor joining, BIONJ, and parsimony.

Animals↗

Prototype of simulation models for epizootics in domestic animals.

Based on the Reed-Frost model (Model I), the authors conducted computer simulation of an epizootic model (Model II) constructed on the assumption that any infected animal in a group, after a given time-period of infectivity, would be removed from the group at the beginning of the next time-period. Models I and II were simulated 100 times for each of the different conditions, viz. the initial size of group, 100 and 1,000, the five steps of contact rate or contact size, and the five more steps of contact rate for the group of 1,000 animals in Model I. From the results obtained, it is believed that as a constant parameter, contact size may be preferably used instead of contact rate in these models. Model II mostly gave higher morbidities than Model I, and earlier termination of epizootics, except the simulation with the smallest contact size. This fact may be due to the effect of herd immunity involved only in Model I. The long duration of epizootic was demonstrated in two of the 100 simulations of Model II with 1,000 individuals and contact size 1. This is characteristic of probabilistic models which are really instructive to studying the flow of epizootic.

Animals↗

Brain mechanisms of selective learning: event-related potentials provide evidence for error-driven learning in humans.

Selective learning has been observed in Pavlovian conditioning in animals and in judgements of event contingencies in humans. This analogy led to the suggestion that the formation of associations underlies both types of learning. An alternative theory proposes that both tasks involve the computation of event contingencies as prescribed by probability theory. Error-driven models of learning incorporate trial-by-trial error-correction mechanisms during training whereas probabilistic models view learning merely as the storage of frequency information for later use during judgement of event contingencies. Competitive interaction between cues was observed in a contingency judgement task. Event-related brain potentials (ERPs) provided evidence for brain events related to the discrepancy between actual and expected outcomes during training thus supporting error-driven accounts of selective learning.

Adult↗

Markov modeling of contaminant concentrations in indoor air.

Most models for contaminant dispersion in indoor air are deterministic and do not account for the probabilistic nature of the pollutant concentration at a given room position and time. Such variability can be important when estimating concentrations involving small numbers of contaminant particles. This article describes the use of probabilistic models termed Markov chains to account for a portion of this variability. The deterministic and Markov models are related in that the former provide the expected concentration values. To explain this relationship, a single-zone (well-mixed room) scenario is described as a Markov chain. Subsequently, a two-zone room is cast as a Markov model, and the latter is applied to assessing a health care worker's risk of tuberculosis infection. Airborne particles carrying Mycobacterium tuberculosis bacilli are usually present in small numbers in a room occupied by an infectious tuberculosis patient. For a given scenario, the Markov model permits estimates of variability in exposure intensity and the resulting variability in infection risk.

Air Microbiology↗

Gene finding with a hidden Markov model of genome structure and evolution.

MOTIVATION: A growing number of genomes are sequenced. The differences in evolutionary pattern between functional regions can thus be observed genome-wide in a whole set of organisms. The diverse evolutionary pattern of different functional regions can be exploited in the process of genomic annotation. The modelling of evolution by the existing comparative gene finders leaves room for improvement. RESULTS: A probabilistic model of both genome structure and evolution is designed. This type of model is called an Evolutionary Hidden Markov Model (EHMM), being composed of an HMM and a set of region-specific evolutionary models based on a phylogenetic tree. All parameters can be estimated by maximum likelihood, including the phylogenetic tree. It can handle any number of aligned genomes, using their phylogenetic tree to model the evolutionary correlations. The time complexity of all algorithms used for handling the model are linear in alignment length and genome number. The model is applied to the problem of gene finding. The benefit of modelling sequence evolution is demonstrated both in a range of simulations and on a set of orthologous human/mouse gene pairs. AVAILABILITY: Free availability over the Internet on www server: http://www.birc.dk/Software/evogene.

Algorithms↗

Understanding actin organization in cell structure through lattice based Monte Carlo simulations.

Understanding the connection between mechanics and cell structure requires the exploration of the key molecular constituents responsible for cell shape and motility. One of these molecular bridges is the cytoskeleton, which is involved with intracellular organization and mechanotransduction. In order to examine the structure in cells, we have developed a computational technique that is able to probe the self-assembly of actin filaments through a lattice based Monte Carlo method. We have modeled the polymerization of these filaments based upon the interactions of globular actin through a probabilistic model encompassing both inert and active proteins. The results show similar response to classic ordinary differential equations at low molecular concentrations, but a bi-phasic divergence at realistic concentrations for living mammalian cells. Further, by introducing localized mobility parameters, we are able to simulate molecular gradients that are observed in nonhomogeneous protein distributions in vivo. The method and results have potential applications in cell and molecular biology as well as self-assembly for organic and inorganic systems.

Actins↗

Statistical models for protein validation using tandem mass spectral data and protein amino acid sequence databases.

The purpose of this work is to develop and verify statistical models for protein identification using peptide identifications derived from the results of tandem mass spectral database searches. Recently we have presented a probabilistic model for peptide identification that uses hypergeometric distribution to approximate fragment ion matches of database peptide sequences to experimental tandem mass spectra. Here we apply statistical models to the database search results to validate protein identifications. For this we formulate the protein identification problem in terms of two independent models, two-hypothesis binomial and multinomial models, which use the hypergeometric probabilities and cross-correlation scores, respectively. Each database search result is assumed to be a probabilistic event. The Bernoulli event has two outcomes: a protein is either identified or not. The probability of identifying a protein at each Bernoulli event is determined from relative length of the protein in the database (the null hypothesis) or the hypergeometric probability scores of the protein's peptides (the alternative hypothesis). We then calculate the binomial probability that the protein will be observed a certain number of times (number of database matches to its peptides) given the size of the data set (number of spectra) and the probability of protein identification at each Bernoulli event. The ratio of the probabilities from these two hypotheses (maximum likelihood ratio) is used as a test statistic to discriminate between true and false identifications. The significance and confidence levels of protein identifications are calculated from the model distributions. The multinomial model combines the database search results and generates an observed frequency distribution of cross-correlation scores (grouped into bins) between experimental spectra and identified amino acid sequences. The frequency distribution is used to generate p-value probabilities of each score bin. The probabilities are then normalized with respect to score bins to generate normalized probabilities of all score bins. A protein identification probability is the multinomial probability of observing the given set of peptide scores. To reduce the effect of random matches, we employ a marginalized multinomial model for small values of cross-correlation scores. We demonstrate that the combination of the two independent methods provides a useful tool for protein identification from results of database search using tandem mass spectra. A receiver operating characteristic curve demonstrates the sensitivity and accuracy level of the approach. The shortcomings of the models are related to the cases when protein assignment is based on unusual peptide fragmentation patterns that dominate over the model encoded in the peptide identification process. We have implemented the approach in a program called PROT_PROBE.

Amino Acid Sequence↗

Cost-effectiveness analysis of two strategies for mass screening for colorectal cancer in France.

The implementation of colorectal cancer mass screening is a high public health priority in France, as in most other industrialised countries. Despite evidences that screening using guaiac fecal occult blood test may reduce colorectal cancer mortality, no European country has organised widespread mass screening with this test. The low sensitivity of this test constitutes its main limitation. Immunological tests, which provide higher sensitivity than the guaiac test, may constitute a satisfactory alternative. This study was carried out to compare the costs and the effectiveness of 20 years of biennial colorectal cancer (CRC) screening with an automated reading immunological test (Magstream) with those obtained with a guaiac stool test (Haemoccult). The model used to estimate the costs and effectiveness of successive biennial CRC screening campaigns was a transitional probabilistic model. The parameters used in this model concerning costs and CRC epidemiological data were calculated from results obtained in the screening program run in Calvados or from published results of foreign studies because of the lack of French studies. The use of Magstream for 20 years of biennial screening costs 59 euros more than Haemoccult per target individual, and should lead to a mean increase in individual life expectancy of 0.0198 years (i.e. about one week), which corresponds to an incremental cost-effectiveness ratio of 2980 euros per years of life saved. Our results suggest that using an immunological test could increase the effectiveness of CRC screening at a reasonable cost for society.

Aged↗