Loss of information in genetic distances.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to D Penny.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
A branch and bound algorithm is described for searching rapidly for minimal length trees from biological data. The algorithm adds characters one at a time, rather than adding taxa, as in previous branch and bound methods. The algorithm has been programmed and is available from the authors. A worked example is given with 33 characters and 15 taxa. About 8 x 10(12) binary trees are possible with 15 taxa but the branch and bound program finds the minimal tree in less than 5 min on an IBM PC.
Six protein sequences from the same 11 mammalian taxa were used to estimate the accuracy and reliability of phylogenetic trees using real, rather than simulated, data. A tree comparison metric was used to measure the increase in similarity of minimal trees as larger, randomly selected subsets of nucleotide positions were taken. The ratio of the observed to the expected number of incompatibilities for each nucleotide position (character) is a good predictor of the number of changes required at that position on the minimal (most-parsimonious) tree. This allows a higher weighting of nucleotide positions that have changed more slowly and should result in the minimal length tree converging to the correct tree as more sequences are obtained. An estimate was made of the smallest subset of trees that need to be considered to include the actual historical tree for a given set of data. It was concluded that it is possible to give a reasonable estimate of the reliability of the final tree, at least when several sequences are combined. With the present data, resolving the rodent-primate-lagomorph (rabbit) trichotomy is the least certain aspect of the final tree, followed then by establishing the position of dog. In our opinion, it is unreasonable to publish an evolutionary tree derived from sequence data without giving an idea of the reliability of the tree.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
We have recently reported a method to identify the shortest possible phylogenetic tree for a set of protein sequences [Foulds Hendy & Penny (1979) J. Mol. Evol. 13. 127--150; Foulds, Penny & Hendy (1979) J. Mol. Evol. 13, 151--166]. The present paper discusses issues that arise during the construction of minimal phylogenetic trees from protein-sequence data. The conversion of the data from amino acid sequences into nucleotide sequences is shown to be advantageous. A new variation of a method for constructing a minimal tree is presented. Our previous methods have involved first constructing a tree and then either proving that it is minimal or transforming it into a minimal tree. The approach presented in the present paper progressively builds up a tree, taxon by taxon. We illustrate this approach by using it to construct a minimal tree for ten mammalian haemoglobin alpha-chain sequences. Finally we define a measure of the complexity of the data and illustrate a method to derive a directed phylogenetic tree from the minimal tree.
The problem of determining the minimal phylogenetic tree is discussed in relation to graph theory. It is shown that this problem is an example of the Steiner problem in graphs which is to connect a set of points by a minimal length network where new points can be added. There is no reported method of solving realistically-sized Steiner problems in reasonable computing time. A heuristic method of approaching the phylogenetic problem is presented, together with a worked example with 7 mammalian cytochrome c sequences. It is shown in this case that the method develops a phylogenetic tree that has the smallest possible number of amino acid replacements. The potential and limitations of the method are discussed. It is stressed that objective methods must be used for comparing different trees. In particular it should be determined how close a given tree is to a mathematically determined lower bound. A theorem is proved which is used to establish a lower bound on the lenghtof any tree and if a tree is found with a length equal to the lower bound, then no shorter tree can exist.
We have recently described a method of building phylogenetic trees and have outlined an approach for proving whether a particular tree is optimal for the data used. In this paper we describe in detail the method of establishing lower bounds on the length of a minimal tree by partitioning the data set into subsets. All characters that could be involved in duplications in the data are paired with all other such characters. A matching algorithm is then used to obtain the pairing of characters that reveals the most duplications in the data. This matching may still not account for all nucleotide substitutions on the tree. The structure of the tree is then used to help select subsets of three or more characters until the lower bound found by partitioning is equal to the length of the tree. The tree must then be a minimal tree since no tree can exist with a length less than that of the lower bound. The method is demonstrated using a set of 23 vertebrate cytochrome c sequences with the criterion of minimizing the total number of nucleotide substitutions. There are 131130 7045768798 96033440625 topologically distinct trees that can be constructed from this data set. The method described in this paper does identify 144 minimal tree variants. The method is general in the sense that it can be used for other data and other criteria of length. It need not however always be possible to prove a treee minimal but the method will give an upper and lower bound on the length of minimal trees.
Explore the source record for details and available documents.
The process of determining the optimal phylogenetic tree from amino acid sequences or comparable data is divided into six stages. Particular attention is given both to the criteria that are used when testing for the optimal tree and the problem of determining the position of the original ancestor. Four types of criteria for evaluating the optimal tree are considered: 1. parsimony (fewest total changes), 2. path lengths from an ancestor to existing species, 3. subtracting the difference between each pair of species as measured on the tree and as compared directly with the data ("excess differences"), 4. Moore Residual Coefficient. These criteria are examined on a set of test data and some of the reasons for the differences among them are discussed. For example, the "average percent standard deviation" weights excess differences unequally in inverse proportion to the square of the observed differences. The Moore Residual Coefficient and the "excess differences" will not necessarily give a value of zero when there are no duplicated changes unless there can only be two states for each character (i.e. binary data). The path length and difference criteria (as well as the Moore Residual Coefficient) give unequal weighting to the individual branches of the tree by counting some branches more times than others. Particularly because of this some criteria will reject trees that are equally parsimonious and the criteria are said to be invalid. However the criterion of parsimony is insensitive in that it can give the same value for several basic networks and it does not specify the position of the original ancestor, the root of the tree. The importance is emphasised of stating a model and examining its predictions before a criterion is chosen to select the best network. The number of rooted trees that can be derived from a basic network (or unrooted tree) is described in relation to how detailed a description of the original ancestor is required. Four methods are described for determining the position of the root of the tree or original ancestor. Each method depends upon some additional information to that used in constructing the basic network and the method chosen will depend on this additional knowledge.
Explore the source record for details and available documents.
Biologically active lipids increase the growth of pea stem sections within 3 hours at the same time their respiration is increased and their growth rate is more than that of the intact plant. The greater final length of the intact internode is due to a longer growth period.BOTH ACTIVE AND INACTIVE LIPIDS ARE RAPIDLY TAKEN UP AND ENTER ALL MAJOR METABOLIC FRACTIONS: among centrifugal fractions methyl oleate tends to label those that contain metabolically active membranes. It is concluded that lipids active in the bioassay are probably the effective molecules at the subcellular site of action.No direct effect of lipids on isolated mitochondria could be shown. The respiration of stem tissue was not influenced by dinitrophenol and carbonyl cyano m-chlorophenyl hydrazone although dinitrophenol inhibited growth. Lipid-induced respiration was sensitive to these agents as well as to cyanide, indicating cytochrome oxidase is probably involved.The promotion of growth and respiration by lipids is not linked to protein synthesis, since actinomycin D, puromycin and cycloheximide failed to inhibit the respiratory increase even though strongly limiting amino acid incorporation into protein. It is most likely that the effect of lipids on growth is due to their promotion of respiration.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Many different programs have been developed for the prediction of the secondary structure of an RNA sequence. Some of these programs generate an ensemble of structures, all of which have free energy close to that of the optimal structure, making it important to be able to quantify how similar these different structures are. To deal with this problem, we define a new class of metrics, the mountain metrics, on the set of RNA secondary structures of a fixed length. We compare properties of these metrics with other well known metrics on RNA secondary structures. We also study some global and local properties of these metrics.