Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

A memory-efficient dynamic programming algorithm for optimal alignment of a sequence to an RNA secondary structure.

BACKGROUND: Covariance models (CMs) are probabilistic models of RNA secondary structure, analogous to profile hidden Markov models of linear sequence. The dynamic programming algorithm for aligning a CM to an RNA sequence of length N is O(N3) in memory. This is only practical for small RNAs. RESULTS: I describe a divide and conquer variant of the alignment algorithm that is analogous to memory-efficient Myers/Miller dynamic programming algorithms for linear sequence alignment. The new algorithm has an O(N2 log N) memory complexity, at the expense of a small constant factor in time. CONCLUSIONS: Optimal ribosomal RNA structural alignments that previously required up to 150 GB of memory now require less than 270 MB.

Algorithms↗

Cluster-Rasch models for microarray gene expression data.

BACKGROUND: We propose two different formulations of the Rasch statistical models to the problem of relating gene expression profiles to the phenotypes. One formulation allows us to investigate whether a cluster of genes with similar expression profiles is related to the observed phenotypes; this model can also be used for future prediction. The other formulation provides an alternative way of identifying genes that are over- or underexpressed from their expression levels in tissue or cell samples of a given tissue or cell type. RESULTS: We illustrate the methods on available datasets of a classification of acute leukemias and of 60 cancer cell lines. For tumor classification, the results are comparable to those previously obtained. For the cancer cell lines dataset, we found four clusters of genes that are related to drug response for many of the 90 drugs that we considered. In addition, for each type of cell line, we identified genes that are over- or underexpressed relative to other genes. CONCLUSIONS: The cluster-Rasch model provides a probabilistic model for describing gene expression patterns across samples and can be used to relate gene expression profiles to phenotypes.

Acute Disease↗

A Bayesian model comparison approach to inferring positive selection.

A popular approach to detecting positive selection is to estimate the parameters of a probabilistic model of codon evolution and perform inference based on its maximum likelihood parameter values. This approach has been evaluated intensively in a number of simulation studies and found to be robust when the available data set is large. However, uncertainties in the estimated parameter values can lead to errors in the inference, especially when the data set is small or there is insufficient divergence between the sequences. We introduce a Bayesian model comparison approach to infer whether the sequence as a whole contains sites at which the rate of nonsynonymous substitution is greater than the rate of synonymous substitution. We incorporated this probabilistic model comparison into a Bayesian approach to site-specific inference of positive selection. Using simulated sequences, we compared this approach to the commonly used empirical Bayes approach and investigated the effect of tree length on the performance of both methods. We found that the Bayesian approach outperforms the empirical Bayes method when the amount of sequence divergence is small and is less prone to false-positive inference when the sequences are saturated, while the results are indistinguishable for intermediate levels of sequence divergence.

Bayes Theorem↗

PepNovo: de novo peptide sequencing via probabilistic network modeling.

We present a novel scoring method for de novo interpretation of peptides from tandem mass spectrometry data. Our scoring method uses a probabilistic network whose structure reflects the chemical and physical rules that govern the peptide fragmentation. We use a likelihood ratio hypothesis test to determine whether the peaks observed in the mass spectrum are more likely to have been produced under our fragmentation model than under a model that treats peaks as random events. We tested our de novo algorithm PepNovo on ion trap data and achieved results that are superior to popular de novo peptide sequencing algorithms. PepNovo can be accessed via the URL http://www-cse.ucsd.edu/groups/bioinformatics/software.html.

Algorithms↗

Relationships between host viremia and vector susceptibility for arboviruses.

Using a threshold model where a minimum level of host viremia is necessary to infect vectors affects our assessment of the relative importance of different host species in the transmission and spread of these pathogens. Other models may be more accurate descriptions of the relationship between host viremia and vector infection. Under the threshold model, the intensity and duration of the viremia above the threshold level is critical in determining the potential numbers of infected mosquitoes. A probabilistic model relating host viremia to the probability distribution of virions in the mosquito bloodmeal shows that the threshold model will underestimate the significance of hosts with low viremias. A probabilistic model that includes avian mortality shows that the maximum number of mosquitoes is infected by feeding on hosts whose viremia peaks just below the lethal level. The relationship between host viremia and vector infection is complex, and there is little experimental information to determine the most accurate model for different arthropod-vector-host systems. Until there is more information, the ability to distinguish the relative importance of different hosts in infecting vectors will remain problematic. Relying on assumptions with little support may result in erroneous conclusions about the importance of different hosts.

Animals↗

A Class of Probabilistic Unfolding Models for Polytomous Responses.

By revisiting the approaches used to present the Rasch model for polytomous response, this paper uses the principle of the rating formulation (Andrich, 1978) to construct a class of unfolding models for polytomous responses in terms of a set of latent dichotomous unfolding variables. By anchoring the dichotomous unfolding variables involved at the same location, this paper presents a formulation of a very general class of unfolding models for ordered polytomous responses, of which the unfolding models for ordered polytomous responses proposed hitherto are special cases. Within this class, the analytic and measurement properties of the probabilistic functions are well interpreted in terms of the latitudes of acceptance parameters of the dichotomous unfolding models. Based on the general form of this class of unfolding models, some new models are readily specified. Copyright 2001 Academic Press.

Journal Article↗

A model for assessing the sensitivity and specificity of tests subject to selection bias. Application to exercise radionuclide ventriculography for diagnosis of coronary artery disease.

A probabilistic model was developed which allows one to estimate sensitivity and specificity of diagnostic tests for coronary artery disease without reference to angiography. The feasibility of the model was evaluated first in a series of computer simulations, and the model was then applied to the assessment of ejection fraction in 933 patients without prior myocardial infarction who underwent exercise radionuclide ventriculography. In 196 patients who were referred to angiography, the conventional abnormal ejection fraction criterion--an absolute rise of less than 0.05 with exercise--had a sensitivity of 79% and a specificity of 68% when referenced to coronary angiography. In 737 patients who were not referred for angiography and who were analyzed instead by our probabilistic model, sensitivity was 63% (p = 0.004 compared to that in the 119 angiographically diseases patients) and specificity was 79% (p = 0.036 compared to that in the 77 angiographic normals). Both the higher sensitivity and lower specificity in the catheterized patients are consistent with a preferential referral of positive test responders to angiography, and of negative test responders away from angiography. The distortion of test sensitivity and specificity which results from this selection bias can be circumvented by substituting a probabilistic estimate of disease for conventional angiographic ascertainment.

Adult↗

Simulation of working population exposures to carbon monoxide using EXPOLIS-Milan microenvironment concentration and time-activity data.

Current air pollution levels have been shown to affect human health. Probabilistic modeling can be used to assess exposure distributions in selected target populations. Modeling can and should be used to compare exposures in alternative future scenarios to guide society development. Such models, however, must first be validated using existing data for a past situation. This study applied probabilistic modeling to carbon monoxide (CO) exposures using EXPOLIS-Milan data. In the current work, the model performance was evaluated by comparing modeled exposure distributions to observed ones. Model performance was studied in detail in two dimensions; (i) for different averaging times (1, 8 and 24 h) and (ii) using different detail in defining the microenvironments in the model (two, five and 11 microenvironments). (iii) The number of exposure events leading to exceeding the 8-h guideline was estimated. Population time activity was modeled using a fractions-of-time approach assuming that some time is spent in each microenvironment used in the model. This approach is best suited for averaging times from 24 h upwards. In this study, we tested how this approach affects results when used for shorter averaging times, 1 and 8 h. Models for each averaging time were run with two, five and 11 microenvironments. The two-microenvironment models underestimated the means and standard deviations (SDs) slightly for all averaging times. The five- and 11-microenvironment models matched the means quite well but underestimated SDs in several cases. For 1- and 24-h averaging times the simulated SDs are slightly smaller than the corresponding observed values. The 8-h model matched the observed exposure levels best. The results show that for CO (i) the modeling approach can be applied for averaging times from 8 to 24 h and as a screening model even to an averaging time of 1 h; (ii) the number of microenvironments affects only weakly the results and in the studied cases only exposure levels below the 80th percentile; (iii) this kind of model can be used to estimate the number of high-exposure events related to adverse health effects. By extrapolation beyond the observed data, it was shown that Milanese office workers may experience adverse health effects caused by CO.

Air Pollutants, Occupational↗

Support vector machine prediction of signal peptide cleavage site using a new class of kernels for strings.

A new class of kernels for strings is introduced. These kernels can be used by any kernel-based data analysis method, including support vector machines (SVM). They are derived from probabilistic models to integrate biologically relevant information. We show how to compute the kernels corresponding to several classical probabilistic models, and illustrate their use by building a SVM for the problem of predicting the cleavage site of signal peptides from the amino-acid sequence of a protein. At a given rate of false positive this method retrieves up to 47% more true positives than the classical weight matrix method.

Algorithms↗

A probabilistic effect assessment model for hazardous substances at the workplace.

A major problem in risk assessment is the quantification of uncertainties. A probabilistic model was developed to consider uncertainties in the effect assessment of hazardous substances at the workplace. Distributions for extrapolation factors (time extrapolation, inter- and intraspecies extrapolation) were determined on the basis of appropriate empirical data. Together with the distribution for the benchmark dose obtained from substance-specific dose-response modelling for the exemplary substances 2,4,4-trimethylpentene (TMP) and aniline, they represent the input distributions for probabilistic modelling. These distributions were combined by Monte Carlo simulation. The resulting target distribution describes the probability that an aspired protection level for workers is achieved at a certain dose and the uncertainty associated with the assessment. In the case of aniline, substance-specific data on differences in susceptibility (between species; among humans due to genetic polymorphisms of N-acetyltransferase) were integrated in the model. Medians of the obtained target distributions of the basic models for TMP and aniline, but not of the specific aniline model are similar to deterministically derived reference values. Differences of more than one order of magnitude between the medians and the 5th percentile of the target distributions indicate substantial uncertainty associated with the effect assessment of these substances. The probabilistic effect assessment model proves to be a practical tool to integrate quantitative information on uncertainty and variability in hazard characterisation.

Alkenes↗

Classification of burn injuries using near-infrared spectroscopy.

Early surgical management of those burn injuries that will not heal spontaneously is critical. The decision to excise and graft is based on a visual assessment that is often inaccurate but yet continues to be the primary means of grading the injury. Superficial and intermediate partial-thickness injuries generally heal with appropriate wound care while deep partial- and full-thickness injuries generally require surgery. This study explores the possibility of using near-infrared spectroscopy to provide an objective and accurate means of distinguishing shallow injuries from deeper burns that require surgery. Twenty burn injuries are studied in five animals, with burns covering <1% of the total body surface area. Carefully controlled superficial, intermediate, and deep partial-thickness injuries as well as full-thickness injuries could be studied with this model. Near-infrared reflectance spectroscopy was used to evaluate these injuries 1 to 3 hours after the insult. A probabilistic model employing partial least-squares logistic regression was used to determine the degree of injury, shallow (superficial or intermediate partial) from deep (deep partial and full thickness), based on the reflectance spectrum of the wound. A leave-animal-out cross-validation strategy was used to test the predictive ability of a 2-latent variable, partial least-squares logistic regression model to distinguish deep burn injuries from shallow injuries. The model displayed reasonable ranking quality as summarized by the area under the receiver operator characteristics curve, AUC = 0.879. Fixing the threshold for the class boundaries at 0.5 probability, the model sensitivity (true positive fraction) to separate deep from shallow burns was 0.90, while model specificity (true negative fraction) was 0.83. Using an acute porcine model of thermal burn injuries, the potential of near-infrared spectroscopy to distinguish between shallow healing burns and deeper burn injuries was demonstrated. While these results should be considered as preliminary and require clinical validation, a probabilistic model capable of differentiating these classes of burns would be a significant aid to the burn specialist.

Algorithms↗

Exploring the "two-hit hypothesis" in NF2: tests of two-hit and three-hit models of vestibular schwannoma development.

Neurofibromatosis 2 (NF2) is a genetic disease that occurs in approximately 1 in 40,000 live births. Almost all affected individuals develop bilateral tumors of Schwann cells that surround the vestibular nerves; these tumors are known as vestibular schwannomas (VS). Evidence from molecular genetic studies suggests that at least two mutations are involved in formation of VS in patients with NF2. Several authors proposed probabilistic models for this process in other tumors, and showed that such models are consistent with incidence data. We evaluated two different probabilistic models for a "2-hit" hypothesis for VS development in NF2 patients, and we present results from fitting these models to incidence data. Molecular evidence does not exclude the possibility that additional hits are necessary for the development of VS, and we also assessed a "3-hit" model for tumor formation. The "3-hit" model fits the data marginally better than one of the "2-hit" models and much better than the other "2-hit" model. Our findings suggest that more than two mutations may be necessary for VS development in NF2 patients.

Humans↗

Pooled Genomic Indexing (PGI): analysis and design of experiments.

Pooled Genomic Indexing (PGI) is a novel method for physical mapping of clones onto known sequences. PGI is carried out by pooling arrayed clones and generating shotgun sequence reads from the pools. The shotgun sequences are compared to a reference sequence. In the simplest case, clones are placed on an array and are pooled by rows and columns. If a shotgun sequence from a row pool and another shotgun sequence from a column pool match the reference sequence at a close distance, they are both assigned to the clone at the intersection of the two pools. Accordingly, the clone is mapped onto the region of the reference sequence between the two matches. A probabilistic model for PGI is developed, and several pooling designs are described and analyzed, including transversal designs and designs from linear codes. The probabilistic model and the pooling schemes are validated in simulated experiments where 625 rat bacterial artificial chromosome (BAC) clones and 207 mouse BAC clones are mapped onto homologous human sequence.

Animals↗

Bayesian population decoding of motor cortical activity using a Kalman filter.

Effective neural motor prostheses require a method for decoding neural activity representing desired movement. In particular, the accurate reconstruction of a continuous motion signal is necessary for the control of devices such as computer cursors, robots, or a patient's own paralyzed limbs. For such applications, we developed a real-time system that uses Bayesian inference techniques to estimate hand motion from the firing rates of multiple neurons. In this study, we used recordings that were previously made in the arm area of primary motor cortex in awake behaving monkeys using a chronically implanted multielectrode microarray. Bayesian inference involves computing the posterior probability of the hand motion conditioned on a sequence of observed firing rates; this is formulated in terms of the product of a likelihood and a prior. The likelihood term models the probability of firing rates given a particular hand motion. We found that a linear gaussian model could be used to approximate this likelihood and could be readily learned from a small amount of training data. The prior term defines a probabilistic model of hand kinematics and was also taken to be a linear gaussian model. Decoding was performed using a Kalman filter, which gives an efficient recursive method for Bayesian inference when the likelihood and prior are linear and gaussian. In off-line experiments, the Kalman filter reconstructions of hand trajectory were more accurate than previously reported results. The resulting decoding algorithm provides a principled probabilistic model of motor-cortical coding, decodes hand motion in real time, provides an estimate of uncertainty, and is straightforward to implement. Additionally the formulation unifies and extends previous models of neural coding while providing insights into the motor-cortical code.

Action Potentials↗

Diagnostic boundaries, reasoning and depressive disorder, I. Development of a probabilistic morbidity model for public health psychiatry.

BACKGROUND: In recent years diagnostic practice in psychiatry has become increasingly structured in an attempt to standardize definitions of disorders and improve reliability. At the same time there has been an increasing recognition of the need to take account of uncertainty in the process of diagnostic decision making. For the most part, diagnosis is still represented by a binary outcome while this is known to entail a substantial loss of information. Many diagnostic schemes involve, in part, taking thresholds on the numbers of symptoms required from symptom lists. METHODS: A model is proposed here, using ideas derived from latent class analysis to permit generalization from these schemes through moving from a binary to a probabilistic measure of psychiatric case status and replacing thresholds with smoothed transitions. RESULTS: An outcome measure is produced where disorder status is expressed in terms of probabilities without changing the meaning of the original measure. Prevalence estimates (using ICD-10 Depressive Episode criteria) are more stable and can be given with increased precision. CONCLUSIONS: Disorder status when expressed in this way retains more diagnostic information and provides a useful extension to traditional binary analyses when looking at prevalence and risk factor estimation.

Bayes Theorem↗

A statistical model for binocular rivalry.

A probabilistic model is presented for the phase durations in binocular rivalry experiments. The hypothetical construct of inhibition or reaction inhibition is used to account for the length of the successive phases of left-eye dominance and right-eye dominance. In accordance with Hull's Postulate X.B. it is assumed that the inhibition increases linearly at rate a1 during periods of left-eye dominance and decreases linearly at rate a0 during periods of right-eye dominance. Two different versions of the proposed model are presented: the beta and the Bessel inhibition models. Inhibition fluctuates between the boundaries 0 and 1 in the beta inhibition model and between -infinity and +infinity in the Bessel inhibition model. The transition rates lambda1(t) for switches from a state of left-eye dominance to a state of right-eye dominance, and lambda0(t) for switches from a state of right-eye dominance to a state of left-eye dominance depend on inhibition: lambda1(t)=l1 (Y(t)), lambda0(t)=l0(Y(t)), where l1 is a non-decreasing function and l0 is a non-increasing function. In the beta inhibition model l1(y)=c1/(1 - y) and l0(y)=c0/y. In the Bessel inhibition model l1(y)=u1e(y) and l0(y)=u0/e(y). Special attention is given to the derivation of the expectation of the stationary phase durations.

Functional Laterality↗

Inter-document coreference resolution of abnormal findings in radiology documents.

In the clinical environment, it is often necessary to track the progression of a condition or various pertinent findings over time. Establishing automatic mechanisms for tracking pertinent findings can aid in the management of a condition as well as provide feedback for treatment outcomes assessment. This work focuses on the challenge of correlating observation of pertinent findings, specifically lung masses, across documents from serial computed tomography examinations for lung cancer patients. A probabilistic model is presented to characterize the likeliness of two observed findings from different documents referring to the same entity. A greedy algorithm is also presented that utilizes the probabilistic model to establish coreference links between findings. Results from a preliminary evaluation of this methodology show a precision of 72% and a recall of 63% for the described inter-document coreference resolution task.

Algorithms↗

Kinetics of growth inhibition in a neoplastic population after repeated cytostatic treatment.

The analysis of growth kinetics of a neoplastic population after cytostatic treatment (radiation, chemotherapy) has been modeled by means of the cell cycle transition probabilities mu and transition intensities lambda. The transition between phases of the cell cycle and quiescent G0 cells during repeated cytostatic treatment has been exposed by the probabilistic model on the basis of the experimental data. The cytostatic treatment (the most effective time, dose and therapeutic combination) can be scheduled using established size and volume of the tumor and an estimate of cell proportion in various compartments of the neoplastic population. The probabilistic model has been worked out with the object of repeated cytostatic treatments in the optimum time with respect to the reduction of resistant quiescent G0 population. By repeating treatment in these times, the G0 population may be effectively reduced, the tumor is loosing its proliferation (recovery) potential and the further tumor growth may be for a long time delayed or even ceased.

Animals↗