Search PubMed⌕ Search

PubMed · 10786294

Solving large scale phylogenetic problems using DCM2.

Abstract

In an earlier paper, we described a new method for phylogenetic tree reconstruction called the Disk Covering Method, or DCM. This is a general method which can be used with any existing phylogenetic method in order to improve its performance. We showed analytically and experimentally that when DCM is used in conjunction with polynomial time distance-based methods, it improves the accuracy of the trees reconstructed. In this paper, we discuss a variant on DCM, that we call DCM2. DCM2 is designed to be used with phylogenetic methods whose objective is the solution of NP-hard optimization problems. We show that DCM2 can be used to accelerate searches for Maximum Parsimony trees. We also motivate the need for solutions to NP-hard optimization problems by showing that on some very large and important datasets, the most popular (and presumably best performing) polynomial time distance methods have poor accuracy.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

D H Huson, L Vawter, T J Warnow. 1999. Solving large scale phylogenetic problems using DCM2.. https://pubmed.ncbi.nlm.nih.gov/10786294/

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Identification of protein domains on topological basis.

A theoretical method is proposed to identify structural domains in proteins of known structures. It is based on the distribution of the local axes of the polypeptide chain. In particular, a statistical analysis is applied to the contributions of the local axes to the absolute writhing number, a topological property of a space curve resulting from the number of self-crossings in the curve projections onto a unit sphere. This finding supports the hypothesis that topological requirements should be satisfied in the process of protein folding and in the final organization of the tertiary structures.

Databases, Factual↗

Using patient-reportable clinical history factors to predict myocardial infarction.

Using a derivation data set of 1253 patients, we built several logistic regression and neural network models to estimate the likelihood of myocardial infarction based upon patient-reportable clinical history factors only. The best performing logistic regression model and neural network model had C-indices of 0.8444 and 0.8503, respectively, when validated on an independent data set of 500 patients. We conclude that both logistic regression and neural network models can be built that successfully predict the probability of myocardial infarction based on patient-reportable history factors alone. These models could have important utility in applications outside of a hospital setting when objective diagnostic test information is not yet be available.

Databases, Factual↗

Why are "natively unfolded" proteins unstructured under physiologic conditions?

"Natively unfolded" proteins occupy a unique niche within the protein kingdom in that they lack ordered structure under conditions of neutral pH in vitro. Analysis of amino acid sequences, based on the normalized net charge and mean hydrophobicity, has been applied to two sets of proteins: small globular folded proteins and "natively unfolded" ones. The results show that "natively unfolded" proteins are specifically localized within a unique region of charge-hydrophobicity phase space and indicate that a combination of low overall hydrophobicity and large net charge represent a unique structural feature of "natively unfolded" proteins.

Databases, Factual↗