Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Automatic analysis and classification of surface electromyography.

In this paper, parametric modeling of surface electromyography (EMG) algorithms that facilitates automatic SEMG feature extraction and artificial neural networks (ANN) are combined for providing an integrated system for the automatic analysis and diagnosis of myopathic disorders. Three paradigms of ANN were investigated: the multilayer backpropagation algorithm, the self-organizing feature map algorithm and a probabilistic neural network model. The performance of the three classifiers was compared with that of the old Fisher linear discriminant (FLD) classifiers. The results have shown that the three ANN models give higher performance. The percentage of correct classification reaches 90%. Poorer diagnostic performance was obtained from the FLD classifier. The system presented here indicates that surface EMG, when properly processed, can be used to provide the physician with a diagnostic assist device.

Algorithms↗

Mixture models for linkage analysis of affected sibling pairs and covariates.

To determine the genetic etiology of complex diseases, a common study design is to recruit affected sib/relative pairs (ASP/ARP) and evaluate their genome-wide distribution of identical by descent (IBD) sharing using a set of highly polymorphic markers. Other attributes or environmental exposures of the ASP/ARP, which are thought to affect liability to disease, are sometimes collected. Conceivably, these covariates could refine the linkage analysis. Most published methods for ASP/ARP linkage with covariates can be conceptualized as logistic models in which IBD status of the ASP is predicted by pair-specific covariates. We develop a different approach to the problem of ASP analysis in the presence of covariates, one that extends naturally to ARP under certain conditions. For ASP linkage analysis, we formulate a mixture model in which a disease mutation is segregating in only a fraction alpha of the sibships, with 1 - alpha sibships being unlinked. Covariate information is used to predict membership within groups; in this report, the two groups correspond to the linked and unlinked sibships. For an ASP with covariate(s) Z = z and multilocus genotype X = x, the mixture model is alpha(z)g(x; lambda) + [1 - alpha(z)]g(0)(x), in which g(0)(x) follows the distribution of genotypes under the null IBD distribution and g(x; lambda) allows for increased IBD sharing. Two mixture models are developed. The pre-clustering model uses covariate information to form probabilistic clusters and then tests for excess IBD sharing independent of the covariates. The Cov-IBD model determines probabilistic group membership by joint consideration of covariate and IBD values. Simulations show that incorporating covariates into linkage analysis can enhance power substantially. A feature of our conceptualization of ASP linkage analysis, with covariates, is that it is apparent how data analysis might evaluate covariates prior to the linkage analysis, thus avoiding the loss of power described by Leal and Ott [2000] when data are stratified.

Algorithms↗

Everyday conditional reasoning: a working memory-dependent tradeoff between counterexample and likelihood use.

Considerable evidence has revealed that working memory capacity is an important determinant of conditional reasoning performance. There are two accounts describing the conditional inference process, the probabilistic and the mental models accounts. According to the mental models account, reasoners retrieve and integrate counterexample information to attain a conclusion. According to the probabilistic account, reasoners base their judgments on probabilistic information. It can be assumed that reasoning according to the mental models process would require more working memory resources than would solving the inference on the basis of probabilistic information. By means of a verbal report study, we showed that participants with low working memory capacity more often use probabilistic information, whereas participants with higher working memory capacity are more likely to use counterexample information. Working memory capacity thus not only relates to reasoning performance, it also determines which process reasoners will engage in.

Cognition↗

Hidden Markov models from molecular dynamics simulations on DNA.

An enhanced bioinformatics tool incorporating the participation of molecular structure as well as sequence in protein DNA recognition is proposed and tested. Boltzmann probability models of sequence-dependent DNA structure from all-atom molecular dynamics simulations were obtained and incorporated into hidden Markov models (HMMs) that can recognize molecular structural signals as well as sequence in protein-DNA binding sites on a genome. The binding of catabolite activator protein (CAP) to cognate DNA sequences was used as a prototype case for implementation and testing of the method. The results indicate that even HMMs based on probabilistic roll/tilt dinucleotide models of sequence-dependent DNA structure have some capability to discriminate between known CAP binding and nonbinding sites and to predict putative CAP binding sites in unknowns. Restricting HMMs to sequence only in regions of strong consensus in which the protein makes base specific contacts with the cognate DNA further improved the discriminatory capabilities of the HMMs. Comparison of results with controls based on sequence only indicates that extending the definition of consensus from sequence to structure improves the transferability of the HMMs, and provides further supportive evidence of a role for dynamical molecular structure as well as sequence in genomic regulatory mechanisms.

DNA↗

Classification in well-defined and ill-defined categories: evidence for common processing strategies.

Early work in perceptual and conceptual categorization assumed that categories had criterial features and that category membership could be determined by logical rules for the combination of features. More recent theories have assumed that categories have an ill-defined structure and have prosposed probabilistic or global similarity models for the verification of category membership. In the experiments reported here, several models of categorization were compared, using one set of categories having criterial features and another set having an ill-defined structure. Schematic faces were used as exemplars in both cases. Because many models depend on distance in a multidimensional space for their predictions, in Experiment 1 a multidimensional scaling study was performed using the faces of both sets as stimuli, In Experiment 2, subjects learned the category membership of faces for the categories having criterial features. After learning, reaction times for category verification and typicality judgments were obtained. Subjects also judged the similarity of pairs of faces. Since these categories had characteristic as well as defining features, it was possible to test the predictions of the feature comparison model (Smith et al.), which asserts that reaction times and typicalities are affected by characteristic features. Only weak support for this model was obtained. Instead, it appeared that subjects developed logical rules for the classification of faces. A characteristic feature affected reaction times only when it was part of the rule system devised by the subject. The procedure for Experiment 3 was like that for Experiment 2, but with ill-defined rather than well-defined categories. The obtained reaction times had high correlations with some of the models for ill-defined categories. However, subjects' performance could best be described as one of feature testing based on a logical rule system for classification. These experiments indicate that whether or not categories have criterial features, subjects attempt to develop a set of feature tests that allow for exemplar classification. Previous evidence supporting probabilistic or similarity models may be interpreted as resulting from subjects' use of the most efficient rules for classification and the averaging of responses for subjects using different sets of rules.

Concept Formation↗

Positional entropy during pigeon homing I: application of Bayesian latent state modelling.

In these two companion papers, we introduce a new approach to the analysis of bird navigation which brings together several novel mathematical and technical applications. Miniaturized GPS logging devices provide track data of sufficiently high spatial and temporal resolution that considerable variation in flight behaviour can be observed remotely from the form of the track alone. We analyse a fundamental measure of bird flight track complexity, spatio-temporal entropy, and explore its state-like structure using a probabilistic hidden Markov model. The emergence of a robust three-state structure proves that the technique has analytical power, since this structure was not obvious in the tracks alone. We propose the hypothesis that positional entropy is indicative of underlying navigational uncertainty, and that familiar area navigation may break down into three states of navigational confidence. By interpreting the relationship between these putative states and features on the map, we are able to propose a number of hypothetical navigational strategies feeding into these states. The first of these two papers details the novel technical developments associated with this work and the second paper contains a navigational interpretation of the results particularly with respect to visual features of the landscape.

Animals↗

Positional entropy during pigeon homing II: navigational interpretation of Bayesian latent state models.

In these two companion papers we introduce a new approach to the analysis of bird navigation which brings together several novel mathematical and technical applications. Miniaturized GPS logging devices provide track data of sufficiently high spatial and temporal resolution that considerable variation in flight behaviour can be observed remotely from the form of the track alone. We analyse a fundamental measure of bird flight track complexity, spatio-temporal entropy, and explore its state-like structure using a probabilistic hidden Markov model. The emergence of a robust three-state structure proves that the technique has analytical power, since this structure was not obvious in the tracks alone. We propose the hypothesis that positional entropy is indicative of underlying navigational uncertainty, and that familiar area navigation may break down into three states of navigational confidence. By interpreting the relationship between these putative states and features on the map, we are able to propose a number of hypothetical navigational strategies feeding into these states. The first of these two papers details the novel technical developments associated with this work and the second paper contains a navigational interpretation of the results particularly with respect to visual features of the landscape.

Animals↗

Reference biospheres for post-closure performance assessment: inter-comparison of SHETRAN simulations and BIOMASS results.

Example Reference Biosphere 2B (ERB2B) is a hypothetical river catchment, described in the IAEA-sponsored BIOMASS study on biosphere aspects of post-closure radiological safety assessments for repositories for solid radioactive wastes. In ERB2B, a radioactively contaminated aquifer interacts with the soils and sediments of the river catchment. A 'semi-distributed', lumped-parameter model (SDLP) was set up for the site as part of the BIOMASS study. In the model, empirically derived transfer functions are used to reduce the complexity of real hydrological transport systems to readily calculable mass-balance accounting routines. In this work, a physically based, spatially distributed modelling system SHETRAN was set up for the site and comparison made with the existing SDLP model. The work has shown that, using standard soil properties in SHETRAN, the soil rapidly saturates and much of the hydrologically effective rainfall (precipitation less evapotranspiration) is lost as saturation-excess surface runoff. This is contrary to the assumptions in the SDLP model. The difficulty arose from the original formulation of catchment characteristics in BIOMASS. Specifically, there was a large water volume entering the soils from precipitation together with an upward flux of groundwater across the lower boundary of a substantial part of the catchment. This water had to be lost from the catchment in some way and the thinness of the soil zone precluded dominance of subsurface, lateral flow over surface runoff. Increasing the saturated conductivity from 1 to 20 m d(-1) reduced the surface flows to similar values to those assumed in the SDLP model (this could also have been achieved by increasing the soil depth). Even with the high saturated conductivity there were still major differences between the two representations. In the woodland on the upper slopes of the valley, the SHETRAN simulation was slightly wetter than the SDLP model, whereas in the shrubland and marshland near the river it was drier than the SDLP model. In the SDLP model, subsurface lateral flows are ignored if there is surface flow, and deep subsurface flows are ignored if there are shallow subsurface flows. In the SDLP model, there is a major assumed change in flow regime between summer and winter. This is not the case in the SHETRAN simulation. Overall, this work illustrates the problems of using 'semi-distributed', lumped-parameter models without prior calibration against a physically based model and the potential for implying unexpected and possibly implausible hydrological characteristics through the specification of flows without considering whether they could occur for realistic soil depths and properties. As there is a need for application of such SDLP models, particularly when undertaking probabilistic calculations, it is suggested that, in future, explicit hydrological modelling should be undertaken first, so that a physically realistic representation can be produced as a basis for assessment studies of the migration of radionuclides or other contaminants.

Fresh Water↗

Predictive Bayesian microbial dose-response assessment based on suggested self-organization in primary illness response: Cryptosporidium parvum.

The probability of illness caused by very low doses of pathogens cannot generally be tested due to the numbers of subjects that would be needed, though such assessments of illness dose response are needed to evaluate drinking water standards. A predictive Bayesian dose-response assessment method was proposed previously to assess the unconditional probability of illness from available information and avoid the inconsistencies of confidence-based approaches. However, the method uses knowledge of the conditional dose-response form, and this form is not well established for the illness endpoint. A conditional parametric dose-response function for gastroenteric illness is proposed here based on simple numerical models of self-organized host-pathogen systems and probabilistic arguments. In the models, illnesses terminate when the host evolves by processes of natural selection to a self-organized critical value of wellness. A generalized beta-Poisson illness dose-response form emerges for the population as a whole. Use of this form is demonstrated in a predictive Bayesian dose-response assessment for cryptosporidiosis. Results suggest that a maximum allowable dose of 5.0 x 10(-7) oocysts/exposure (e.g., 2.5 x 10(-7) oocysts/L water) would correspond with the original goals of the U.S. Environmental Protection Agency Surface Water Treatment Rule, considering only primary illnesses resulting from Poisson-distributed pathogen counts. This estimate should be revised to account for non-Poisson distributions of Cryptosporidium parvum in drinking water and total response, considering secondary illness propagation in the population.

Animals↗

Probabilistic methods for addressing uncertainty and variability in biological models: application to a toxicokinetic model.

Population variability and uncertainty are important features of biological systems that must be considered when developing mathematical models for these systems. In this paper we present probability-based parameter estimation methods that account for such variability and uncertainty. Theoretical results that establish well-posedness and stability for these methods are discussed. A probabilistic parameter estimation technique is then applied to a toxicokinetic model for trichloroethylene using several types of simulated data. Comparison with results obtained using a standard, deterministic parameter estimation method suggests that the probabilistic methods are better able to capture population variability and uncertainty in model parameters.

Animals↗

Analysis and synthesis of textured motion: particles and waves.

Natural scenes contain a wide range of textured motion phenomena which are characterized by the movement of a large amount of particle and wave elements, such as falling snow, wavy water, and dancing grass. In this paper, we present a generative model for representing these motion patterns and study a Markov chain Monte Carlo algorithm for inferring the generative representation from observed video sequences. Our generative model consists of three components. The first is a photometric model which represents an image as a linear superposition of image bases selected from a generic and overcomplete dictionary. The dictionary contains Gabor and LoG bases for point/particle elements and Fourier bases for wave elements. These bases compete to explain the input images and transfer them to a token (base) representation with an O(10(2))-fold dimension reduction. The second component is a geometric model which groups spatially adjacent tokens (bases) and their motion trajectories into a number of moving elements--called "motons." A moton is a deformable template in time-space representing a moving element, such as a falling snowflake or a flying bird. The third component is a dynamic model which characterizes the motion of particles, waves, and their interactions. For example, the motion of particle objects floating in a river, such as leaves and balls, should be coupled with the motion of waves. The trajectories of these moving elements are represented by coupled Markov chains. The dynamic model also includes probabilistic representations for the birth/death (source/sink) of the motons. We adopt a stochastic gradient algorithm for learning and inference. Given an input video sequence, the algorithm iterates two steps: 1) computing the motons and their trajectories by a number of reversible Markov chain jumps, and 2) learning the parameters that govern the geometric deformations and motion dynamics. Novel video sequences are synthesized from the learned models and, by editing the model parameters, we demonstrate the controllability of the generative model.

Algorithms↗

Warren K. Sinclair keynote address: contemporary issues in risk-informed decision making on the disposition of radioactive waste.

A consistent and transparent risk-informed approach to managing nuclear waste is plagued with different regulators, different rules and regulations for different waste types, different compliance requirements, and indecisions about probabilistic vs. deterministic models. Low-activity waste management is particularly void of a path forward with respect to being risk-informed. Risk assessment is not referenced in the statutes on low-activity waste even though both the U.S. Environmental Protection Agency and U.S. Nuclear Regulatory Commission (U.S. NRC) have policies to apply consistent risk management approaches to all of their programs. The U.S. NRC has developed guidance on the preparation of probabilistic performance assessments for low-activity waste facilities, but there have been no serious takers and a lack of initiative on the part of licensees. Thus, little to no experience exists on risk-informing low-activity waste. The missed opportunities include establishing a risk basis that would allow for simpler, safer, and much less costly alternatives for low-activity waste disposal while enabling society to have the full benefit of radiation technologies. There is hope that congressional action or regulatory rule making will address some of these issues with the result being the adoption of a more general and unified approach to risk-informed regulation of all types of waste. Just as much of the initiative for risk-informed nuclear power came from industry, it must also be the case for nuclear waste. A start would be the adoption of a basic framework of risk assessment in waste management applicable to all types of waste--radioactive and nonradioactive. The "set of triplets" risk assessment framework that is applicable to any kind of risk is an established alternative. It is believed that such a framework with the support of a regulatory structure made compatible through appropriate rulemaking or congressional action, and the experience of the probabilistic performance assessments for the Waste Isolation Pilot Plant and the proposed Yucca Mountain high-level waste repository, could result in the right path forward for the regulation and management of low-activity waste.

Decision Making↗

Preparation of name and address data for record linkage using hidden Markov models.

BACKGROUND: Record linkage refers to the process of joining records that relate to the same entity or event in one or more data collections. In the absence of a shared, unique key, record linkage involves the comparison of ensembles of partially-identifying, non-unique data items between pairs of records. Data items with variable formats, such as names and addresses, need to be transformed and normalised in order to validly carry out these comparisons. Traditionally, deterministic rule-based data processing systems have been used to carry out this pre-processing, which is commonly referred to as "standardisation". This paper describes an alternative approach to standardisation, using a combination of lexicon-based tokenisation and probabilistic hidden Markov models (HMMs). METHODS: HMMs were trained to standardise typical Australian name and address data drawn from a range of health data collections. The accuracy of the results was compared to that produced by rule-based systems. RESULTS: Training of HMMs was found to be quick and did not require any specialised skills. For addresses, HMMs produced equal or better standardisation accuracy than a widely-used rule-based system. However, accuracy was worse when used with simpler name data. Possible reasons for this poorer performance are discussed. CONCLUSION: Lexicon-based tokenisation and HMMs provide a viable and effort-effective alternative to rule-based systems for pre-processing more complex variably formatted data such as addresses. Further work is required to improve the performance of this approach with simpler data such as names. Software which implements the methods described in this paper is freely available under an open source license for other researchers to use and improve.

Data Collection↗

Using a state-space model with hidden variables to infer transcription factor activities.

MOTIVATION: In a gene regulatory network, genes are typically regulated by transcription factors (TFs). Transcription factor activity (TFA) is more difficult to measure than gene expression levels are. Other models have extracted information about TFA from gene expression data, but without explicitly modeling feedback from the genes. We present a state-space model (SSM) with hidden variables. The hidden variables include regulatory motifs in the gene network, such as feedback loops and auto-regulation, making SSM a useful complement to existing models. RESULTS: A gene regulatory network incorporating, for example, feed-forward loops, auto-regulation and multiple-inputs was constructed with an SSM model. First, the gene expression data were simulated by SSM and used to infer the TFAs. The ability of SSM to infer TFAs was evaluated by comparing the profiles of the inferred and simulated TFAs. Second, SSM was applied to gene expression data obtained from Escherichia coli K12 undergoing a carbon source transition and from the Saccharomyces cerevisiae cell cycle. The inferred activity profile for each TF was validated either by measurement or by activity information from the literature. The SSM model provides a probabilistic framework to simulate gene regulatory networks and to infer activity profiles of hidden variables. AVAILABILITY: Supplementary data and Matlab code will be made available at the URL below. SUPPLEMENTARY INFORMATION: http://www.chems.msu.edu/groups/chan/ssm.zip.

Algorithms↗

A pilot study on the use of decision theory and value of information analysis as part of the NHS Health Technology Assessment programme.

OBJECTIVES: To demonstrate the benefits of using appropriate decision-analytic methods and value of information analysis (DA-VOI). Also to establish the feasibility and implications of applying these methods to inform the prioritisation process of the NHS Health Technology Assessment (HTA) programme, and possibly extending their use therein. DATA SOURCES: Three research topics that were considered by the HTA panels in the September 2002 and February 2003 prioritisation rounds. REVIEW METHODS: A brief and non-technical overview of DA-VOI methods was circulated to the panels and Prioritisation Strategy Group (PSG). For each case study the results were presented to the panels and the PSG in the form of brief case-study reports. Feedback on the DA-VOI analysis and its presentation was obtained in the form of completed questionnaires from panel members, and reports from panel senior lecturers and PSG members. RESULTS: Although none of the research topics identified met all of the original selection criteria for inclusion as case studies in the pilot, it was possible to construct appropriate decision-analytic models and conduct probabilistic analysis for each topic. In each case, the tasks were completed within the time-frame required by the existing HTA research prioritisation process. The brief case-study reports provided a description of the decision problem, a summary of the current evidence base and a characterisation of decision uncertainty in the form of cost-effectiveness acceptability curves. Estimates of value of information for the decision problem were presented for relevant patient groups and clinical settings, as well as the value of information associated with particular model inputs. The implications for the value of research in each of the areas were presented in general terms. Details were also provided on what the analysis suggested regarding the design of any future research in terms of features such as the relevant patient groups and comparators, and whether experimental design was likely to be required. CONCLUSIONS: The pilot study showed that, even with very short timelines, it is possible to undertake DA-VOI that can feed into the priority-setting process that has been developed for the HTA programme. There are however a number of areas that need to be established at the beginning of the process, such as clarification of the nature of the decision problem for which additional research is being considered, explicitness about which existing data should be used and how data that exhibit particular weaknesses should be down-weighted in the analysis. Other areas, including optimum application of researcher time, integrating the vignette (a summary of the clinical problem and existing evidence) and the use of DA-VOI, training, use of sensitivity analyses, and deployment of clinical expertise, are also considered in terms of the potential implementation of DA-VOI within the HTA programme. Recommendations for further research include how literature searching should focus on those variables to which the model's results are most sensitive and with the highest expected value of perfect information; methods of evidence synthesis (multiple parameter synthesis) to consider the evidence surrounding multiple comparators and networks of evidence; and ways in which the value of sample information can be used by the NHS HTA programme and other research funders to decide on the most efficient design of new evaluative research. There is also a need for an analytical framework to be developed that can jointly address the question of whether additional resources would better be devoted to additional research or interventions to change clinical practice.

Biomedical Technology↗

Using hidden Markov models and observed evolution to annotate viral genomes.

MOTIVATION: ssRNA (single stranded) viral genomes are generally constrained in length and utilize overlapping reading frames to maximally exploit the coding potential within the genome length restrictions. This overlapping coding phenomenon leads to complex evolutionary constraints operating on the genome. In regions which code for more than one protein, silent mutations in one reading frame generally have a protein coding effect in another. To maximize coding flexibility in all reading frames, overlapping regions are often compositionally biased towards amino acids which are 6-fold degenerate with respect to the 64 codon alphabet. Previous methodologies have used this fact in an ad hoc manner to look for overlapping genes by motif matching. In this paper differentiated nucleotide compositional patterns in overlapping regions are incorporated into a probabilistic hidden Markov model (HMM) framework which is used to annotate ssRNA viral genomes. This work focuses on single sequence annotation and applies an HMM framework to ssRNA viral annotation. A description of how the HMM is parameterized, whilst annotating within a missing data framework is given. A Phylogenetic HMM (Phylo-HMM) extension, as applied to 14 aligned HIV2 sequences is also presented. This evolutionary extension serves as an illustration of the potential of the Phylo-HMM framework for ssRNA viral genomic annotation. RESULTS: The single sequence annotation procedure (SSA) is applied to 14 different strains of the HIV2 virus. Further results on alternative ssRNA viral genomes are presented to illustrate more generally the performance of the method. The results of the SSA method are encouraging however there is still room for improvement, and since there is overwhelming evidence to indicate that comparative methods can improve coding sequence (CDS) annotation, the SSA method is extended to a Phylo-HMM to incorporate evolutionary information. The Phylo-HMM extension is applied to the same set of 14 HIV2 sequences which are pre-aligned. The performance improvement that results from including the evolutionary information in the analysis is illustrated.

Algorithms↗

The quantitative evaluation of functional neuroimaging experiments: mutual information learning curves.

Learning curves are presented as an unbiased means for evaluating the performance of models for neuroimaging data analysis. The learning curve measures the predictive performance in terms of the generalization or prediction error as a function of the number of independent examples (e.g., subjects) used to determine the parameters in the model. Cross-validation resampling is used to obtain unbiased estimates of a generic multivariate Gaussian classifier, for training set sizes from 2 to 16 subjects. We apply the framework to four different activation experiments, in this case [(15)O]water data sets, although the framework is equally valid for multisubject fMRI studies. We demonstrate how the prediction error can be expressed as the mutual information between the scan and the scan label, measured in units of bits. The mutual information learning curve can be used to evaluate the impact of different methodological choices, e.g., classification label schemes, preprocessing choices. Another application for the learning curve is to examine the model performance using bias/variance considerations enabling the researcher to determine if the model performance is limited by statistical bias or variance. We furthermore present the sensitivity map as a general method for extracting activation maps from statistical models within the probabilistic framework and illustrate relationships between mutual information and pattern reproducibility as derived in the NPAIRS framework described in a companion paper.

Adult↗

Heterotachy and long-branch attraction in phylogenetics.

BACKGROUND: Probabilistic methods have progressively supplanted the Maximum Parsimony (MP) method for inferring phylogenetic trees. One of the major reasons for this shift was that MP is much more sensitive to the Long Branch Attraction (LBA) artefact than is Maximum Likelihood (ML). However, recent work by Kolaczkowski and Thornton suggested, on the basis of simulations, that MP is less sensitive than ML to tree reconstruction artefacts generated by heterotachy, a phenomenon that corresponds to shifts in site-specific evolutionary rates over time. These results led these authors to recommend that the results of ML and MP analyses should be both reported and interpreted with the same caution. This specific conclusion revived the debate on the choice of the most accurate phylogenetic method for analysing real data in which various types of heterogeneities occur. However, variation of evolutionary rates across species was not explicitly incorporated in the original study of Kolaczkowski and Thornton, and in most of the subsequent heterotachous simulations published to date, where all terminal branch lengths were kept equal, an assumption that is biologically unrealistic. RESULTS: In this report, we performed more realistic simulations to evaluate the relative performance of MP and ML methods when two kinds of heterogeneities are considered: (i) within-site rate variation (heterotachy), and (ii) rate variation across lineages. Using a similar protocol as Kolaczkowski and Thornton to generate heterotachous datasets, we found that heterotachy, which constitutes a serious violation of existing models, decreases the accuracy of ML whatever the level of rate variation across lineages. In contrast, the accuracy of MP can either increase or decrease when the level of heterotachy increases, depending on the relative branch lengths. This result demonstrates that MP is not insensitive to heterotachy, contrary to the report of Kolaczkowski and Thornton. Finally, in the case of LBA (i.e. when two non-sister lineages evolved faster than the others), ML outperforms MP over a wide range of conditions, except for unrealistic levels of heterotachy. CONCLUSION: For realistic combinations of both heterotachy and variation of evolutionary rates across lineages, ML is always more accurate than MP. Therefore, ML should be preferred over MP for analysing real data, all the more so since parametric methods also allow one to handle other types of biological heterogeneities much better, such as among sites rate variation. The confounding effects of heterotachy on tree reconstruction methods do exist, but can be eschewed by the development of mixture models in a probabilistic framework, as proposed by Kolaczkowski and Thornton themselves.

Animals↗