Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Contributions of frequency distribution analysis to the understanding of coronary restenosis. A reappraisal of the gaussian curve.

BACKGROUND: Clinical restenosis after balloon angioplasty can be categorized by use of dichotomous terms based on the presence or absence of recurrent myocardial ischemia. In contrast, recent investigations have concluded that late luminal renarrowing, documented through angiographic imaging, occurs to a variable extent in nearly all stenoses. This process has been characterized by a gaussian or normal frequency distribution, with restenosis simply representing an extreme form of this delayed remodeling. In the current study, frequency distribution analysis was used to examine the process of coronary restenosis in a large cohort of patients at risk. METHODS AND RESULTS: Quantitative coronary angiographic analysis was applied to 9279 cineangiograms obtained in 3093 patients before and immediately after angioplasty and after 6-month follow-up. Late loss, defined as the change in minimum lumen diameter of the target stenosis from postdilation to follow-up, did not statistically conform to a normal distribution (P<.0001 by both chi2 statistic and Kolmogorov-Smirnov test), even after the exclusion of the 236 stenoses that displayed total occlusions at follow-up angiography. Examination of deviation from a normal curve revealed an excessively high frequency of stenoses that experienced either little change (0.0+/-0.3 mm) or marked change (1.0 to 2.0 mm) in late loss, with a low frequency of stenoses with intermediate values (0.3 to 1.0 mm). Similarly, although the distribution of percent diameter stenosis of the target lesion was statistically normal immediately after dilation, this gaussian distribution disappeared during the follow-up period. Other angiographic indexes of restenosis also failed to approximate a normal curve. In an attempt to improve the goodness of fit, a probabilistic model of late loss was created on the basis of deconvolution of the observed data distribution. Two theoretical, discrete populations of stenoses were identified, one with and one without overall late luminal narrowing. Unlike the gaussian distribution, this model provided a good representation of the observed data (P=NS for lack of fit). CONCLUSIONS: The frequency distributions of angiographic indexes of restenosis often superficially resemble a gaussian curve, an appearance that is artifactually enhanced by the measurement imprecision of current quantitative techniques. Nevertheless, standard indexes of coronary restenosis fail to conform statistically to a normal distribution. The pattern of deviations observed supports the possible existence of discrete subpopulations of lesions, each with a different propensity toward the development of restenosis after coronary intervention.

Adult↗

Estimating entropy rates with Bayesian confidence intervals.

The entropy rate quantifies the amount of uncertainty or disorder produced by any dynamical system. In a spiking neuron, this uncertainty translates into the amount of information potentially encoded and thus the subject of intense theoretical and experimental investigation. Estimating this quantity in observed, experimental data is difficult and requires a judicious selection of probabilistic models, balancing between two opposing biases. We use a model weighting principle originally developed for lossless data compression, following the minimum description length principle. This weighting yields a direct estimator of the entropy rate, which, compared to existing methods, exhibits significantly less bias and converges faster in simulation. With Monte Carlo techinques, we estimate a Bayesian confidence interval for the entropy rate. In related work, we apply these ideas to estimate the information rates between sensory stimuli and neural responses in experimental data (Shlens, Kennel, Abarbanel, & Chichilnisky, in preparation).

Bayes Theorem↗

A step forward in studying the compact genetic algorithm.

The compact Genetic Algorithm (cGA) is an Estimation of Distribution Algorithm that generates offspring population according to the estimated probabilistic model of the parent population instead of using traditional recombination and mutation operators. The cGA only needs a small amount of memory; therefore, it may be quite useful in memory-constrained applications. This paper introduces a theoretical framework for studying the cGA from the convergence point of view in which, we model the cGA by a Markov process and approximate its behavior using an Ordinary Differential Equation (ODE). Then, we prove that the corresponding ODE converges to local optima and stays there. Consequently, we conclude that the cGA will converge to the local optima of the function to be optimized.

Algorithms↗

Automated global structure extraction for effective local building block processing in XCS.

Learning Classifier Systems (LCSs), such as the accuracy-based XCS, evolve distributed problem solutions represented by a population of rules. During evolution, features are specialized, propagated, and recombined to provide increasingly accurate subsolutions. Recently, it was shown that, as in conventional genetic algorithms (GAs), some problems require efficient processing of subsets of features to find problem solutions efficiently. In such problems, standard variation operators of genetic and evolutionary algorithms used in LCSs suffer from potential disruption of groups of interacting features, resulting in poor performance. This paper introduces efficient crossover operators to XCS by incorporating techniques derived from competent GAs: the extended compact GA (ECGA) and the Bayesian optimization algorithm (BOA). Instead of simple crossover operators such as uniform crossover or one-point crossover, ECGA or BOA-derived mechanisms are used to build a probabilistic model of the global population and to generate offspring classifiers locally using the model. Several offspring generation variations are introduced and evaluated. The results show that it is possible to achieve performance similar to runs with an informed crossover operator that is specifically designed to yield ideal problem-dependent exploration, exploiting provided problem structure information. Thus, we create the first competent LCSs, XCS/ECGA and XCS/BOA, that detect dependency structures online and propagate corresponding lower-level dependency structures effectively without any information about these structures given in advance.

Algorithms↗

Classification image analysis: estimation and statistical inference for two-alternative forced-choice experiments.

We consider estimation and statistical hypothesis testing on classification images obtained from the two-alternative forced-choice experimental paradigm. We begin with a probabilistic model of task performance for simple forced-choice detection and discrimination tasks. Particular attention is paid to general linear filter models because these models lead to a direct interpretation of the classification image as an estimate of the filter weights. We then describe an estimation procedure for obtaining classification images from observer data. A number of statistical tests are presented for testing various hypotheses from classification images based on some more compact set of features derived from them. As an example of how the methods we describe can be used, we present a case study investigating detection of a Gaussian bump profile.

Choice Behavior↗

Vision and touch are automatically integrated for the perception of sequences of events.

The purpose of the present experiment was to investigate the integration of sequences of visual and tactile events. Subjects were presented with sequences of visual flashes and tactile taps simultaneously and instructed to count either the flashes (Session 1) or the taps (Session 2). The number of flashes could differ from the number of taps by +/-1. For both sessions, the perceived number of events was significantly influenced by the number of events presented in the task-irrelevant modality. Touch had a stronger influence on vision than vision on touch. Interestingly, touch was the more reliable of the two modalities-less variable estimates when presented alone. For both sessions, the perceptual estimates were less variable when stimuli were presented in both modalities than when the task-relevant modality was presented alone. These results indicate that even when one signal is explicitly task irrelevant, sensory information tends to be automatically integrated across modalities. They also suggest that the relative weight of each sensory channel in the integration process depends on its relative reliability. The results are described using a Bayesian probabilistic model for multimodal integration that accounts for the coupling between the sensory estimates.

Adult↗

Maximal predicted duration of viremia in bluetongue virus-infected cattle.

Central to the development of rational trade policies pertaining to bluetongue virus (BTV) infection is determination of the risk posed by ruminants previously exposed to the virus. Precise determination of the maximal duration of infectious viremia is essential to the development of an appropriate quarantine period prior to movement of animals from BTV-endemic to BTV-free regions. The objective of this study was to predict the duration of detectable viremia in BTV-infected cattle using a probabilistic modeling analysis of existing data. Data on the duration of detectable viremia in cattle were obtained from previously published studies. Data sets were created from a large field study of naturally infected cattle in Australia and from experimental infections of cattle with Australian and US serotypes of BTV. Probability distributions were fitted to the pooled empirical data, and the 3 probability distributions that provided the best fit to the data were the gamma, Weibull, and lognormal probability distributions. These asymmetric probability distributions are often well suited for decay processes, such as the time to termination of detectable viremia. The analyses indicated a > 99% probability of detectable BTV viremia ceasing after < or = 9 weeks of infection in adult cattle and after a slightly longer interval in BTV-infected, colostrum-deprived newborn calves.

Animals↗

Detection of transposable elements by their compositional bias.

BACKGROUND: Transposable elements (TE) are mobile genetic entities present in nearly all genomes. Previous work has shown that TEs tend to have a different nucleotide composition than the host genes, either considering codon usage bias or dinucleotide frequencies. We show here how these compositional differences can be used as a tool for detection and analysis of TE sequences. RESULTS: We compared the composition of TE sequences and host gene sequences using probabilistic models of nucleotide sequences. We used hidden Markov models (HMM), which take into account the base composition of the sequences (occurrences of words n nucleotides long, with n ranging here from 1 to 4) and the heterogeneity between coding and non-coding parts of sequences. We analyzed three sets of sequences containing class I TEs, class II TEs and genes respectively in three species: Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana. Each of these sets had a distinct, homogeneous composition, enabling us to distinguish between the two classes of TE and the genes. However the particular base composition of the TEs differed in the three species studied. CONCLUSIONS: This approach can be used to detect and annotate TEs in genomic sequences and complements the current homology-based TE detection methods. Furthermore, the HMM method is able to identify the parts of a sequence in which the nucleotide composition resembles that of a coding region of a TE. This is useful for the detailed annotation of TE sequences, which may contain an ancient, highly diverged coding region that is no longer fully functional.

Animals↗

An interactive visualization tool to explore the biophysical properties of amino acids and their contribution to substitution matrices.

BACKGROUND: Quantitative descriptions of amino acid similarity, expressed as probabilistic models of evolutionary interchangeability, are central to many mainstream bioinformatic procedures such as sequence alignment, homology searching, and protein structural prediction. Here we present a web-based, user-friendly analysis tool that allows any researcher to quickly and easily visualize relationships between these bioinformatic metrics and to explore their relationships to underlying indices of amino acid molecular descriptors. RESULTS: We demonstrate the three fundamental types of question that our software can address by taking as a specific example the connections between 49 measures of amino acid biophysical properties (e.g., size, charge and hydrophobicity), a generalized model of amino acid substitution (as represented by the PAM74-100 matrix), and the mutational distance that separates amino acids within the standard genetic code (i.e., the number of point mutations required for interconversion during protein evolution). We show that our software allows a user to recapture the insights from several key publications on these topics in just a few minutes. CONCLUSION: Our software facilitates rapid, interactive exploration of three interconnected topics: (i) the multidimensional molecular descriptors of the twenty proteinaceous amino acids, (ii) the correlation of these biophysical measurements with observed patterns of amino acid substitution, and (iii) the causal basis for differences between any two observed patterns of amino acid substitution. This software acts as an intuitive bioinformatic exploration tool that can guide more comprehensive statistical analyses relating to a diverse array of specific research questions.

Amino Acid Sequence↗

XRate: a fast prototyping, training and annotation tool for phylo-grammars.

BACKGROUND: Recent years have seen the emergence of genome annotation methods based on the phylo-grammar, a probabilistic model combining continuous-time Markov chains and stochastic grammars. Previously, phylo-grammars have required considerable effort to implement, limiting their adoption by computational biologists. RESULTS: We have developed an open source software tool, xrate, for working with reversible, irreversible or parametric substitution models combined with stochastic context-free grammars. xrate efficiently estimates maximum-likelihood parameters and phylogenetic trees using a novel "phylo-EM" algorithm that we describe. The grammar is specified in an external configuration file, allowing users to design new grammars, estimate rate parameters from training data and annotate multiple sequence alignments without the need to recompile code from source. We have used xrate to measure codon substitution rates and predict protein and RNA secondary structures. CONCLUSION: Our results demonstrate that xrate estimates biologically meaningful rates and makes predictions whose accuracy is comparable to that of more specialized tools.

Algorithms↗

Automated recognition of malignancy mentions in biomedical literature.

BACKGROUND: The rapid proliferation of biomedical text makes it increasingly difficult for researchers to identify, synthesize, and utilize developed knowledge in their fields of interest. Automated information extraction procedures can assist in the acquisition and management of this knowledge. Previous efforts in biomedical text mining have focused primarily upon named entity recognition of well-defined molecular objects such as genes, but less work has been performed to identify disease-related objects and concepts. Furthermore, promise has been tempered by an inability to efficiently scale approaches in ways that minimize manual efforts and still perform with high accuracy. Here, we have applied a machine-learning approach previously successful for identifying molecular entities to a disease concept to determine if the underlying probabilistic model effectively generalizes to unrelated concepts with minimal manual intervention for model retraining. RESULTS: We developed a named entity recognizer (MTag), an entity tagger for recognizing clinical descriptions of malignancy presented in text. The application uses the machine-learning technique Conditional Random Fields with additional domain-specific features. MTag was tested with 1,010 training and 432 evaluation documents pertaining to cancer genomics. Overall, our experiments resulted in 0.85 precision, 0.83 recall, and 0.84 F-measure on the evaluation set. Compared with a baseline system using string matching of text with a neoplasm term list, MTag performed with a much higher recall rate (92.1% vs. 42.1% recall) and demonstrated the ability to learn new patterns. Application of MTag to all MEDLINE abstracts yielded the identification of 580,002 unique and 9,153,340 overall mentions of malignancy. Significantly, addition of an extensive lexicon of malignancy mentions as a feature set for extraction had minimal impact in performance. CONCLUSION: Together, these results suggest that the identification of disparate biomedical entity classes in free text may be achievable with high accuracy and only moderate additional effort for each new application domain.

Algorithms↗

Exonuclease activity and P nucleotide addition in the generation of the expressed immunoglobulin repertoire.

BACKGROUND: Immunoglobulin rearrangement involves random and imprecise processes that act to both create and constrain diversity. Two such processes are the loss of nucleotides through the action of unknown exonuclease(s) and the addition of P nucleotides. The study of such processes has been compromised by difficulties in reliably aligning immunoglobulin genes and in the partitioning of nucleotides between segment ends, and between N and P nucleotides. RESULTS: A dataset of 294 human IgM sequences was created and partitioned with the aid of a probabilistic model. Non-random removal of nucleotides is seen between the three IGH gene types with the IGHV gene averaging removals of 1.2 nucleotides compared to 4.7 for the other gene ends (p < 0.001). Individual IGHV, IGHD and IGHJ gene subgroups also display statistical differences in the level of nucleotide loss. For example, within the IGHJ group, IGHJ3 has average removals of 1.3 nucleotides compared to 6.4 nucleotides for IGHJ6 genes (p < 0.002). Analysis of putative P nucleotides within the IgM and pooled datasets revealed only a single putative P nucleotide motif (GTT at the 3' D-REGION end) to occur at a frequency significantly higher then would be expected from random N nucleotide addition. CONCLUSIONS: The loss of nucleotides due to the action of exonucleases is not random, but is influenced by the nucleotide composition of the genes. P nucleotides do not make a significant contribution to diversity of immunoglobulin sequences. Although palindromic sequences are present in 10% of immunologlobulin rearrangements, most of the 'palindromic' nucleotides are likely to have been inserted into the junction during the process of N nucleotide addition. P nucleotides can only be stated with confidence to contribute to diversity of less than 1% of sequences. Any attempt to identify P nucleotides in immunoglobulins is therefore likely to introduce errors into the partitioning of such sequences.

Base Composition↗

Uncertainty principle of genetic information in a living cell.

BACKGROUND: Formal description of a cell's genetic information should provide the number of DNA molecules in that cell and their complete nucleotide sequences. We pose the formal problem: can the genome sequence forming the genotype of a given living cell be known with absolute certainty so that the cell's behaviour (phenotype) can be correlated to that genetic information? To answer this question, we propose a series of thought experiments. RESULTS: We show that the genome sequence of any actual living cell cannot physically be known with absolute certainty, independently of the method used. There is an associated uncertainty, in terms of base pairs, equal to or greater than micros (where micro is the mutation rate of the cell type and s is the cell's genome size). CONCLUSION: This finding establishes an "uncertainty principle" in genetics for the first time, and its analogy with the Heisenberg uncertainty principle in physics is discussed. The genetic information that makes living cells work is thus better represented by a probabilistic model rather than as a completely defined object.

Animals↗

Similarity searches in genome-wide numerical data sets.

We present psi-square, a program for searching the space of gene vectors. The program starts with a gene vector, i.e., the set of measurements associated with a gene, and finds similar vectors, derives a probabilistic model of these vectors, then repeats search using this model as a query, and continues to update the model and search again, until convergence. When applied to three different pathway-discovery problems, psi-square was generally more sensitive and sometimes more specific than the ad hoc methods developed for solving each of these problems before.

Journal Article↗

Refining motifs by improving information content scores using neighborhood profile search.

The main goal of the motif finding problem is to detect novel, over-represented unknown signals in a set of sequences (e.g. transcription factor binding sites in a genome). The most widely used algorithms for finding motifs obtain a generative probabilistic representation of these over-represented signals and try to discover profiles that maximize the information content score. Although these profiles form a very powerful representation of the signals, the major difficulty arises from the fact that the best motif corresponds to the global maximum of a non-convex continuous function. Popular algorithms like Expectation Maximization (EM) and Gibbs sampling tend to be very sensitive to the initial guesses and are known to converge to the nearest local maximum very quickly. In order to improve the quality of the results, EM is used with multiple random starts or any other powerful stochastic global methods that might yield promising initial guesses (like projection algorithms). Global methods do not necessarily give initial guesses in the convergence region of the best local maximum but rather suggest that a promising solution is in the neighborhood region. In this paper, we introduce a novel optimization framework that searches the neighborhood regions of the initial alignment in a systematic manner to explore the multiple local optimal solutions. This effective search is achieved by transforming the original optimization problem into its corresponding dynamical system and estimating the practical stability boundary of the local maximum. Our results show that the popularly used EM algorithm often converges to sub-optimal solutions which can be significantly improved by the proposed neighborhood profile search. Based on experiments using both synthetic and real datasets, our method demonstrates significant improvements in the information content scores of the probabilistic models. The proposed method also gives the flexibility in using different local solvers and global methods depending on their suitability for some specific datasets.

Journal Article↗

Exposures to diesel exhaust in the International Brotherhood of Teamsters, 1950-1990.

A prior case-control study found a positive, monotonic exposure-response relationship between exposure to diesel exhaust and lung cancer among decedents of the Central States Conference of the International Brotherhood of Teamsters. In response to critiques of the Teamsters' exposure estimates by the Health Effects Institute's Diesel Epidemiology Panel, historical exposures and associated uncertainties are investigated here. Historic diesel exhaust exposures are predicted as a function of heavy-duty diesel truck emissions, increasing use of diesel engines, and occupational elemental carbon (EC) measurements taken during the late 1980s and early 1990s. EC from diesel and nondiesel sources is distinguished in light of recent studies indicating a substantial contribution of gasoline vehicles to ambient EC. Monte Carlo sampling is used to characterize exposure distributions. The methodology used in this article-a probabilistic model for historical exposure assessment-is novel.

Air Pollutants, Occupational↗

Patterns of intron gain and loss in fungi.

Little is known about the patterns of intron gain and loss or the relative contributions of these two processes to gene evolution. To investigate the dynamics of intron evolution, we analyzed orthologous genes from four filamentous fungal genomes and determined the pattern of intron conservation. We developed a probabilistic model to estimate the most likely rates of intron gain and loss giving rise to these observed conservation patterns. Our data reveal the surprising importance of intron gain. Between about 150 and 250 gains and between 150 and 350 losses were inferred in each lineage. We discuss one gene in particular (encoding 1-phosphoribosyl-5-pyrophosphate synthetase) that displays an unusually high rate of intron gain in multiple lineages. It has been recognized that introns are biased towards the 5' ends of genes in intron-poor genomes but are evenly distributed in intron-rich genomes. Current models attribute this bias to 3' intron loss through a poly-adenosine-primed reverse transcription mechanism. Contrary to standard models, we find no increased frequency of intron loss toward the 3' ends of genes. Thus, recent intron dynamics do not support a model whereby 5' intron positional bias is generated solely by 3'-biased intron loss.

Adenosine↗

Simple scaling laws control the genetic architectures of human complex traits.

Genome-wide association studies have revealed that the genetic architectures of complex traits vary widely, including in terms of the numbers, effect sizes, and allele frequencies of significant hits. However, at present we lack a principled way of understanding the similarities and differences among traits. Here, we describe a probabilistic model that combines the effects of mutation, drift, and stabilizing selection at individual sites with a genome-scale model of phenotypic variation. In this model, the architecture of a trait arises from the distribution of selection coefficients of mutations and from two scaling parameters. We fit this model for 95 highly polygenic quantitative traits of different kinds from the UK Biobank. Notably, we infer that all these traits have fairly similar, though not identical, distributions of selection coefficients. This similarity suggests that differences in architectures of highly polygenic traits arise mainly from the two scaling parameters: the mutational target size and heritability per site, which vary by orders of magnitude among traits. When these two scale factors are accounted for, we find that the architectures of all 95 traits are very similar.

Humans↗