Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Mutation and childhood cancer: a probabilistic model for the incidence of retinoblastoma.

The incidences of some childhood cancers have been shown to fit a two-mutation hypothesis for cancer initiation. According to this hypothesis, the first mutation can be either germinal or somatic while the second is always somatic. A probabilistic model involving the mean number of tumors per genetically susceptible individual is developed as a function of age and is compared with age incidence data for retinoblastoma. The change in the mean number of tumors with time is interpreted in terms of the growth of retinal cells. In patients who are not genetically susceptible, the times of occurrence of the first and second somatic mutations can be inferred from a comparison of familial and non-familial unilateral case incidences. The total incidences of hereditary and nonhereditary forms of retinoblastoma are related to germinal and somatic mutation rates. The even distribution of certain childhood cancers throughout the world suggests that their incidences are determined by spontaneous mutation rates rather than by local environmental mutagenic carcinogens.

Child↗

Prediction of desiccation sensitivity in seeds of woody species: a probabilistic model based on two seed traits and 104 species.

BACKGROUND AND AIMS: Seed desiccation sensitivity limits the ex situ conservation of up to 47 % of plant species, dependent on habitat. Whilst desirable, empirically determining desiccation tolerance levels in seeds of all species is unrealistic. A probabilistic model for the rapid identification of woody species at high risk of displaying seed desiccation sensitivity is presented. METHODS: The model was developed using binary logistic regression on seed trait data [seed mass, moisture content, seed coat ratio (SCR) and rainfall in the month of seed dispersal] for 104 species from 37 families from a semi-deciduous tropical forest in Panamá. KEY RESULTS: For the Panamanian species, only seed mass and SCR were significantly related to the response to desiccation, with the desiccation-sensitive seeds being large and having a relatively low SCR (i.e. thin 'seed' coats). Application of this model to a further 38 species, of known seed storage behaviour, from two additional continents and differing vegetation types (dryland Africa and temperate Europe) correctly predicted the response to desiccation in all cases, and resolved conflicting published data for two species (Acer pseudoplatanus and Azadirachta indica). CONCLUSIONS: This model may have application as a decision-making tool in the handling of species of unknown seed storage behaviour in species from three disparate habitats.

Africa↗

Evaluation of a probabilistic model for staging of oesophageal carcinoma.

With the help of two experts in gastrointestinal oncology from the Netherlands Cancer Institute, Antoni van Leeuwenhoekhuis, a decision-support system is being developed for patient-specific therapy selection for oesophageal carcinoma. The kernel of the system is a probabilistic model describing the characteristics of oesophageal carcinoma and the pathophysiological processes of invasion and metastasis. Using data from 185 patients, an evaluation study of the model was conducted. We found that for 86% of the patients, the model established the stage of the patient's carcinoma correctly.

Artificial Intelligence↗

On a probabilistic model for the numerical estimation of nocturnal migration of birds.

The study of nocturnal bird migration by cone methods of observation has a century-long history but has continued to be used up to the present. To describe the flux and estimate the number of passing birds a probabilistic model is proposed. This model is based on the concept of dynamic Poisson ensemble of points in appropriate phase space and has two parameters. One is scalar and the other one is functional. We constructed consistent estimations of these parameters and discuss their use for the numerical estimation of the flux of birds observed in a narrow light cone generated by the bright lunar disk and formed by an open angle of telescope. Selection on the same type of birds was suggested as the necessary condition for the model application. Ground speed of each bird was introduced into the model as a new but obligatory value determining the quantification of the flux of bird.

Animal Migration↗

Predicting stress fractures using a probabilistic model of damage, repair and adaptation.

This paper is concerned with the theoretical prediction of stress fractures in the bones of athletes, soldiers and others during periods of intensive exercise. Previously [J. Orthop. Res. 19 (2001) 919] we showed that test data on the fatigue strength of bone in vitro could be described using Weibull's probabilistic model, allowing predictions to be made of the probability of failure as a function of time under cyclic loading. This paper extends the earlier argument by including two living processes which act to reduce the incidence of failure: (i) repair of damage, and; (ii) adaptation by bone deposition. Having incorporated these aspects into the mathematical model, we applied the theory to a specific case: the human second metatarsal. We predicted a 17% incidence of stress fractures, all occurring within 6 weeks of commencement of the training programme. These predictions agreed well with clinical findings. Interestingly, we concluded that the major effect in preventing stress fractures comes from repair rather than from adaptation, which has a relatively minor role because it acts more slowly.

Adaptation, Physiological↗

Optimisation of pyrogen testing in parenterals according to different pharmacopoeias by probabilistic modelling.

The rabbit test to detect pyrogenic contamination in parenterals is crucial to ensure patient safety. The pharmacopoeial tests in Europe, the US and Japan are based on the fever reaction of rabbits, but differ in their experimental design and in their algorithms to assess contamination. Employing an international reference endotoxin, fever can be induced in rabbits. Data from 171 rabbits built the base for probabilistic modelling of the fever reaction and for the comparison of the pharmacopoeial tests. The rabbit fever reaction could be modelled as a function of the amount of injected endotoxin (per kg body weight) by linear regression. Combining the pharmacopoeial algorithms of the rabbit pyrogen test with the developed model allowed analysis of differences regarding test results and animal consumption. This showed that the assessment of pyrogenic contamination strongly depends on the respective pyrogen test stipulated by regulations. Additionally, the approach was used to develop a new experimental design. Two specific versions of this design resulted in a reduction of the number of animals used by about 30% while the safety of the test was maintained. A need for harmonisation is evident, allowing optimisation of the experimental design, which promotes animal welfare.

Algorithms↗

Mitotic rate and younger age are predictors of sentinel lymph node positivity: lessons learned from the generation of a probabilistic model.

BACKGROUND: Sentinel lymph node (SLN) biopsy allows surgeons to identify patients with subclinical nodal involvement who may benefit from lymphadenectomy and, possibly, adjuvant therapy. Several factors have been variably, and sometimes discordantly, reported to have predictive value for SLN metastasis to best select which patients require SLN biopsy. METHODS: We reviewed 419 patients who underwent SLN biopsy for melanoma from a prospectively collected melanoma database. To derive a probabilistic model for the occurrence of a positive SLN, a multivariate logistic model was fit by using a stepwise variable selection method. The accuracy of each model was evaluated by using receiver operator characteristic curves. RESULTS: On univariate analysis, the number of mitoses per square millimeter, increasing Breslow depth, decreasing age, ulceration, and melanoma on the trunk showed a significant relationship to a positive SLN. Multivariate analysis revealed that once age, mitotic rate, and Breslow thickness were included, no other factor, including ulceration, was significantly associated with a positive SLN. The data suggest that younger patients with tumors <1 mm may still have a substantial risk for a positive SLN, especially if the mitotic rate is high. CONCLUSIONS: In addition to Breslow depth, mitoses per square millimeter and younger age were factors identified as independent predictors of a positive SLN. This model may identify patients with thin melanoma at sufficient risk for metastases to justify SLN biopsy.

Adult↗

Gene selection for classification of cancers using probabilistic model building genetic algorithm.

Recently, DNA microarray-based gene expression profiles have been used to correlate the clinical behavior of cancers with the differential gene expression levels in cancerous and normal tissues. To this end, after selection of some predictive genes based on signal-to-noise (S2N) ratio, unsupervised learning like clustering and supervised learning like k-nearest neighbor (k NN) classifier are widely used. Instead of S2N ratio, adaptive searches like Probabilistic Model Building Genetic Algorithm (PMBGA) can be applied for selection of a smaller size gene subset that would classify patient samples more accurately. In this paper, we propose a new PMBGA-based method for identification of informative genes from microarray data. By applying our proposed method to classification of three microarray data sets of binary and multi-type tumors, we demonstrate that the gene subsets selected with our technique yield better classification accuracy.

Algorithms↗

Rich probabilistic models for gene expression.

Clustering is commonly used for analyzing gene expression data. Despite their successes, clustering methods suffer from a number of limitations. First, these methods reveal similarities that exist over all of the measurements, while obscuring relationships that exist over only a subset of the data. Second, clustering methods cannot readily incorporate additional types of information, such as clinical data or known attributes of genes. To circumvent these shortcomings, we propose the use of a single coherent probabilistic model, that encompasses much of the rich structure in the genomic expression data, while incorporating additional information such as experiment type, putative binding sites, or functional information. We show how this model can be learned from the data, allowing us to discover patterns in the data and dependencies between the gene expression patterns and additional attributes. The learned model reveals context-specific relationships, that exist only over a subset of the experiments in the dataset. We demonstrate the power of our approach on synthetic data and on two real-world gene expression data sets for yeast. For example, we demonstrate a novel functionality that falls naturally out of our framework: predicting the "cluster" of the array resulting from a gene mutation based only on the gene's expression pattern in the context of other mutations.

Algorithms↗

BCI Competition 2003--Data set III: probabilistic modeling of sensorimotor mu rhythms for classification of imaginary hand movements.

Brain-computer interfaces require effective online processing of electroencephalogram (EEG) measurements, e.g., as a part of feedback systems. We present an algorithm for single-trial online classification of imaginary left and right hand movements, based on time-frequency information derived from filtering EEG wideband raw data with causal Morlet wavelets, which are adapted to individual EEG spectra. Since imaginary hand movements lead to perturbations of the ongoing pericentral mu rhythm, we estimate probabilistic models for amplitude modulation in lower (10 Hz) and upper (20 Hz) frequency bands over the sensorimotor hand cortices both contra- and ipsilaterally to the imagined movements (i.e., at EEG channels C3 and C4). We use an integrative approach to accumulate over time evidence for the subject's unknown motor intention. Disclosure of test data labels after the competition showed this approach to succeed with an error rate as low as 10.7%.

Algorithms↗

IntNetDB v1.0: an integrated protein-protein interaction network database generated by a probabilistic model.

BACKGROUND: Although protein-protein interaction (PPI) networks have been explored by various experimental methods, the maps so built are still limited in coverage and accuracy. To further expand the PPI network and to extract more accurate information from existing maps, studies have been carried out to integrate various types of functional relationship data. A frequently updated database of computationally analyzed potential PPIs to provide biological researchers with rapid and easy access to analyze original data as a biological network is still lacking. RESULTS: By applying a probabilistic model, we integrated 27 heterogeneous genomic, proteomic and functional annotation datasets to predict PPI networks in human. In addition to previously studied data types, we show that phenotypic distances and genetic interactions can also be integrated to predict PPIs. We further built an easy-to-use, updatable integrated PPI database, the Integrated Network Database (IntNetDB) online, to provide automatic prediction and visualization of PPI network among genes of interest. The networks can be visualized in SVG (Scalable Vector Graphics) format for zooming in or out. IntNetDB also provides a tool to extract topologically highly connected network neighborhoods from a specific network for further exploration and research. Using the MCODE (Molecular Complex Detections) algorithm, 190 such neighborhoods were detected among all the predicted interactions. The predicted PPIs can also be mapped to worm, fly and mouse interologs. CONCLUSION: IntNetDB includes 180,010 predicted protein-protein interactions among 9,901 human proteins and represents a useful resource for the research community. Our study has increased prediction coverage by five-fold. IntNetDB also provides easy-to-use network visualization and analysis tools that allow biological researchers unfamiliar with computational biology to access and analyze data over the internet. The web interface of IntNetDB is freely accessible at http://hanlab.genetics.ac.cn/IntNetDB.htm. Visualization requires Mozilla version 1.8 (or higher) or Internet Explorer with installation of SVGviewer.

Algorithms↗

A probabilistic model for mining implicit 'chemical compound-gene' relations from literature.

MOTIVATION: The importance of chemical compounds has been emphasized more in molecular biology, and 'chemical genomics' has attracted a great deal of attention in recent years. Thus an important issue in current molecular biology is to identify biological-related chemical compounds (more specifically, drugs) and genes. Co-occurrence of biological entities in the literature is a simple, comprehensive and popular technique to find the association of these entities. Our focus is to mine implicit 'chemical compound and gene' relations from the co-occurrence in the literature. RESULTS: We propose a probabilistic model, called the mixture aspect model (MAM), and an algorithm for estimating its parameters to efficiently handle different types of co-occurrence datasets at once. We examined the performance of our approach not only by a cross-validation using the data generated from the MEDLINE records but also by a test using an independent human-curated dataset of the relationships between chemical compounds and genes in the ChEBI database. We performed experimentation on three different types of co-occurrence datasets (i.e. compound-gene, gene-gene and compound-compound co-occurrences) in both cases. Experimental results have shown that MAM trained by all datasets outperformed any simple model trained by other combinations of datasets with the difference being statistically significant in all cases. In particular, we found that incorporating compound-compound co-occurrences is the most effective in improving the predictive performance. We finally computed the likelihoods of all unknown compound-gene (more specifically, drug-gene) pairs using our approach and selected the top 20 pairs according to the likelihoods. We validated them from biological, medical and pharmaceutical viewpoints.

Artificial Intelligence↗

Probabilistic modeling of rosette formation.

Rosetting, or forming a cell aggregate between a single target nucleated cell and a number of red blood cells (RBCs), is a simple assay for cell adhesion mediated by specific receptor-ligand interaction. For example, rosette formation between sheep RBC and human lymphocytes has been used to differentiate T cells from B cells. Rosetting assay is commonly used to determine the interaction of Fc gamma-receptors (FcgammaR) expressed on inflammatory cells and IgG coated on RBCs. Despite its wide use in measuring cell adhesion, the biophysical parameters of rosette formation have not been well characterized. Here we developed a probabilistic model to describe the distribution of rosette sizes, which is Poissonian. The average rosette size is predicted to be proportional to the apparent two-dimensional binding affinity of the interacting receptor-ligand pair and their site densities. The model has been supported by experiments of rosettes mediated by four molecular interactions: FcgammaRIII interacting with IgG, T cell receptor and coreceptor CD8 interacting with antigen peptide presented by major histocompatibility molecule, P-selectin interacting with P-selectin glycoprotein ligand 1 (PSGL-1), and L-selectin interacting with PSGL-1. The latter two are structurally similar and are different from the former two. Fitting the model to data enabled us to evaluate the apparent effective two-dimensional binding affinity of the interacting molecular pairs: 7.19x10(-5) microm4 for FcgammaRIII-IgG interaction, 4.66x10(-3) microm4 for P-selectin-PSGL-1 interaction, and 0.94x10(-3) microm4 for L-selectin-PSGL-1 interaction. These results elucidate the biophysical mechanism of rosette formation and enable it to become a semiquantitative assay that relates the rosette size to the effective affinity for receptor-ligand binding.

Animals↗

Development of databases for use in validation studies of probabilistic models of dietary exposure to food chemicals and nutrients.

The data currently available in the European Union in terms of food consumption and of food chemical and nutrient concentration data present many limitations when used for estimating intake. The most refined techniques currently available were used within the European Union FP5 Monte Carlo project to estimate, as accurately as possible, the intake of food additives, pesticide residues and nutrients. Databases of 'true' intakes of food additives (based on brand level food consumption records and additive concentration data), pesticide residues (based on duplicate diet studies) and nutrients (based on biomarker studies) have thus been generated. These kind of estimates are rarely repeatable because the databases generated and used to calculate them require an extraordinary expenditure of time and resources. The databases created served the purpose of estimating as accurately as possible 'true' chemical intakes for assessing the validity of additive, nutrient and pesticide probabilistic models.

Databases, Factual↗

Prediction of protein continuum secondary structure with probabilistic models based on NMR solved structures.

BACKGROUND: The structure of proteins may change as a result of the inherent flexibility of some protein regions. We develop and explore probabilistic machine learning methods for predicting a continuum secondary structure, i.e. assigning probabilities to the conformational states of a residue. We train our methods using data derived from high-quality NMR models. RESULTS: Several probabilistic models not only successfully estimate the continuum secondary structure, but also provide a categorical output on par with models directly trained on categorical data. Importantly, models trained on the continuum secondary structure are also better than their categorical counterparts at identifying the conformational state for structurally ambivalent residues. CONCLUSION: Cascaded probabilistic neural networks trained on the continuum secondary structure exhibit better accuracy in structurally ambivalent regions of proteins, while sustaining an overall classification accuracy on par with standard, categorical prediction methods.

Algorithms↗

Probabilistic modelling for estimating gas kinetics and decompression sickness risk in pigs during H2 biochemical decompression.

We modelled the kinetics of H2 flux during gas uptake and elimination in conscious pigs exposed to hyperbaric H2. The model used a physiological description of gas flux fitted to the observed decompression sickness (DCS) incidence in two groups of pigs: untreated controls, and animals that had received intestinal injections of H2-metabolizing microbes that biochemically eliminated some of the H2 stored in the pigs' tissues. To analyse H2 flux during gas uptake, animals were compressed in a dry chamber to 24 atm (ca 88% H2, 9% He, 2% O2, 1% N2) for 30-1440 min and decompressed at 0.9 atm min(-1) (n = 70). To analyse H2 flux during gas elimination, animals were compressed to 24 atm for 3 h and decompressed at 0.45-1.8 atm min(-1) (n = 58). Animals were closely monitored for 1 h post-decompression for signs of DCS. Probabilistic modelling was used to estimate that the exponential time constant during H2 uptake (tau(in)) and H2 elimination (tau(out)) were 79 +/- 25 min and 0.76 +/- 0.14 min, respectively. Thus, the gas kinetics affecting DCS risk appeared to be substantially faster for elimination than uptake, which is contrary to customary assumptions of gas uptake and elimination kinetic symmetry. We discuss the possible reasons for this asymmetry, and why absolute values of H2 kinetics cannot be obtained with this approach.

Animals↗

Population pharmacokinetic-pharmacodynamic model of craving in an enforced smoking cessation population: indirect response and probabilistic modeling.

PURPOSE: A population pharmacokinetic-pharmacodynamic model accounting for placebo effect was used to relate nicotine concentration and enforced smoking cessation craving score measured by the Tiffany rating scale short form. METHODS: Twenty-four smokers were enrolled in a placebo-controlled, randomized, double-blind, three periods, crossover trial. The study objective was to describe the nicotine-induced changes on craving scores. Two modeling strategies based on a mechanistic (indirect response models with drug-related inhibition on the k(in) synthesis rate and with a drug-related stimulation of the k(out) removal rate were evaluated) and a probabilistic (logistic regression) approach were used. RESULTS: Placebo response model properly fitted the circadian changes on craving scores. The analysis revealed that the indirect response model with inhibition on k(in) was the preferred model for the smoking data whereas the preferred model for the Nicotine Replacement Therapy data was the one with stimulation on k(out). The logistic analysis showed that the nicotine concentration was a significant predictor of reduction in craving during the free-smoking period. CONCLUSIONS: Nicotine dosage regimen can influence the nicotine mechanism of action: an instantaneous delivery at an individually selected time seems to inhibit the onset of craving while constant delivery at a pre-defined time seems to attenuate the craving.

Adult↗

Validation analysis of probabilistic models of dietary exposure to food additives.

The validity of a range of simple conceptual models designed specifically for the estimation of food additive intakes using probabilistic analysis was assessed. Modelled intake estimates that fell below traditional conservative point estimates of intake and above 'true' additive intakes (calculated from a reference database at brand level) were considered to be in a valid region. Models were developed for 10 food additives by combining food intake data, the probability of an additive being present in a food group and additive concentration data. Food intake and additive concentration data were entered as raw data or as a lognormal distribution, and the probability of an additive being present was entered based on the per cent brands or the per cent eating occasions within a food group that contained an additive. Since the three model components assumed two possible modes of input, the validity of eight (2(3)) model combinations was assessed. All model inputs were derived from the reference database. An iterative approach was employed in which the validity of individual model components was assessed first, followed by validation of full conceptual models. While the distribution of intake estimates from models fell below conservative intakes, which assume that the additive is present at maximum permitted levels (MPLs) in all foods in which it is permitted, intake estimates were not consistently above 'true' intakes. These analyses indicate the need for more complex models for the estimation of food additive intakes using probabilistic analysis. Such models should incorporate information on market share and/or brand loyalty.

Databases, Factual↗