Search PubMed⌕ Search

Biomedical subjects

R Holmquist

Publications and source records attributed to R Holmquist.

At least 19 recordsLinked to original sources

Higher-primate phylogeny--why can't we decide?

At present, no definitive agreement on either the correct branching order or differential rates of evolution among the higher primates exists, despite the accumulated integration of decades of morphological, immunological, protein and nucleic acid sequence data, and numerous reasonable theoretical models for the analysis, interpretation, and understanding of those data. Of the three distinct unrooted phylogenetic trees, that joining human with chimpanzee and the gorilla with the orangutan is currently favored, but the two alternatives that group humans with either gorillas or the orangutan rather than with chimpanzees also have support. This paper is a synthetic and critical review of the methodological literature and isolates some 20 specific reasons why uncertainty in the evolutionary understanding of our closest living relatives persists. Many of the difficulties are eliminated or ameliorated by Lake's new methods of phylogenetic invariants and operator metrics. In the companion paper these new methods are used to analyze both the nuclear and mitochondrial DNA of the higher primates.

Animals↗

Analysis of higher-primate phylogeny from transversion differences in nuclear and mitochondrial DNA by Lake's methods of evolutionary parsimony and operator metrics.

In the companion paper (Holmquist et al. 1988), we concluded that there is no agreement on either the correct branching order or differential rates of evolution among the higher primates, and we examined in depth why this uncertainty in the evolutionary understanding of our closest living relatives persists. Recently, Lake developed two novel methods, based on group properties of transition and transversion operators, that (a) permit, in principle, objective resolution of problems of the above type and (b) attach a statistical significance level to the conclusions drawn. In the present paper, we develop formulas for using these two methods in tandem and apply them to study transversion differences in (1) nuclear DNA for a 7-kb segment of the psi eta-globin locus and a 3-kb intergenic region between the psi beta- and delta-globin loci and (2) mitochondrial DNA for the 896-bp fragment of Brown et al. Although each of these nucleotide sequence regions has its characteristic tempo and mode of evolution, the nuclear and mitochondrial data together, comprising a total of 10,939 base positions, support a Homo/Pan clade at the 97% confidence level. If we calibrate the divergence point for humans and chimpanzees at 5 Myr, consideration of the transversion branch lengths for the combined nuclear data indicates that the gorilla lineage branched off 600,000-900,000 years prior to that, although the 2 sigma sampling errors do not preclude either a temporal trifurcation for the three species or a considerably more ancient branch point for the gorilla. To resolve the length of this central branch to a relative accuracy of 25% and 30% will require a factor of 16 and nine times more data, respectively--i.e., in excess of 100,000 homologous nucleotides for each of the four primates. For the nuclear genes, heterogeneity in evolutionary rates between different parts of the genome is mostly restricted to the human lineage for these two segments. The lineage leading to chimpanzees has evolved 0.4 (3-kb fragment) to 3.5 (7-kb segment) times as rapidly as the lineage leading to humans, and that leading to the gorilla has evolved approximately one-fifth to one-half as rapidly as that leading to chimpanzees. Thus, even local molecular clocks can "tick" badly. As significant is the fact that virtually contiguous parts of the genome tick at markedly different rates.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals↗

The spatial distribution of fixed mutations within genes coding for proteins.

We have examined the extensive amino acid sequence data now available for five protein families - the alpha crystallin A chain, myoglobin, alpha and beta hemoglobin, and the cytochromes c - with the goal of estimating the true spatial distribution of base substitutions within genes that code for proteins. In every case the commonly used Poisson density failed to even approximate the experimental pattern of base substitution. For the 87 species of beta hemoglobin examined, for example, the probability that the observed results were from a Poisson process was the minuscule 10(-44). Analogous results were obtained for the other functional families. All the data were reasonably, but not perfectly, described by the negative binomial density. In particular, most of the data were described by one of the very simple limiting forms of this density, the geometric density. The implications of this for evolutionary inference are discussed. It is evident that most estimates of total base substitutions between genes are badly in need of revision.

Amino Acid Sequence↗

Transitions and transversions in evolutionary descent: an approach to understanding.

In this paper I lay a quantitative theoretical groundwork for understanding the proportions of the possible types of base substitutions observed between 12 genes sharing a common ancestor and isolated from extant species. The experimentally observed types of base substitution between two sequenced genes do not give a direct measure of the types of base substitutions that occur during evolutionary descent. However, by use of a statistical assemblage of these observations, we can recover, without the assumption of parsimony, the conditional base substitution probabilities that determine this descent. Three methods - direct count, regression, and informational entropy maximization - are described by which these probabilities can be estimated from experimental data. The methods are complementary in that each is most useful for somewhat different types of experimental data. These methods are used to study the ratio of transversions to transitions during gene divergence. Though this ratio is not constant during divergence, it does approach a stable limiting value that in principle can vary from zero, corresponding to 100% transition differences, to infinity, corresponding to 0% transition differences. In practice the limiting ratio tends to hover around a value of two, which is expected on a random basis. However, base substitution pathways that are very nonrandom also may lead to a limiting ratio of exactly two, so that such a value is not diagnostic for random pathways. The limiting ratio can be directly calculated from a knowledge of the twelve conditional probabilities for each type of base substitution, or from a knowledge of the equilibrium base composition of the DNAs compared. An expression is given for this calculation. Fifteen years ago Jean Derancourt, Andrew Lebor and Emile Zuckerkandl (1967), analyzing the amino acid sequence of globin chains coded by nuclear genes, made the original observation that the proportion of transition differences decreases with increasing evolutionary time. Recently Brown et al. (1982) and Brown and Simpson (1982) have reported a decrease in the observed proportion of transition differences in mitochondrial DNA with increasing evolutionary divergence. The conditions that must be satisfied for this type of behavior to occur at stable base composition and with stable base substitution probabilities are defined. Multiple substitutions per se do not lead to a decrease in transition differences with increasing evolutionary divergence.

Animals↗

The current status of REH theory.

The recent evaluation by Fitch (1980) of REH theory for macromolecular divergence is a severely erroneous and distorted analysis of our work over the past decade. We reply to those distortions here. At present, there is no factual basis for believing Fitch's assessment that corrections which move evolutionary estimates of total mutations fixed closer to the true distance must do so at the expense of an increased variance sufficient to compromise the value of the improvement. By direct calculation the variance in the estimates of total mutations fixed given by REH theory is comparable to that of other models now in the literature for the case in which genetic events are equiprobable. A general argument is given that suggests that, as we consider more and more carefully the selective, functional, and structural constraints on the evolution of genes and proteins, this variance may be expected to decrease toward a lower bound.

Amino Acid Sequence↗

The estimation of genetic divergence.

We have independently repeated the computer simulations on which Nei and Tateno (1978) base their criticism of REH theory and have extended the analysis to include mRNAs as well as proteins. The simulation data confirm the correctness of the REH method. The high average value of the fixation intensity mu 2 found by Nei and Tateno is due to two factors: 1) they reported only the five replications in which mu 2 was high, excluding the forty-five replications containing the more representative data; and 2) the lack of information, inherent to protein sequence data, about fixed mutations at the third nucleotide position within codons, as the values are lower when the estimate is made from the mRNAs that code for the proteins. REH values calculated from protein or nucleic acid data on the basis of the equiprobability of genetic events underestimate, not overestimate, the total fixed mutations. In REH theory the experimental data determine the estimate T2 of the time average number of codons that have been free to fix mutations during a given period of divergence. In the method of Nei and Tateno it is assumed, despite evidence to the contrary, that every amino acid position may fix a mutation. Under the latter assumption, the measure X2 of genetic divergence suggested by Nei and Tateno is not tenable: values of X2 for a alpha hemoglobin divergences are less than the minimum number of fixed substitutions known to have occurred. Within the context of REH theory, a paradox, first posed by Zuckerkandl, with respect to the high rate of covarion turnover and the nature of general function sites in proteins is resolved.

Amino Acids↗

Evolutionary analysis of alpha and beta hemoglobin genes by REH theory under the assumption of the equiprobability of genetic events.

It is shown how REH theory in conjunction with mRNA or gene sequence data can be used to obtain estimates of the fixation intensity, the number of varions, and the total mutations fixed between homologous pairs of nucleic acids. These estimates are more accurate than those that can be derived from amino acid sequence data. The method is illustrated for alpha and beta hemoglobin genes and these improved estimates are compared with those made from the amino acid sequences for which those genes code. Significant differences are found between the estimates made by these two methods. For the beta hemoglobin gene sequences examined here, the fixation intensity is somewhat less than the protein data had suggested, and the number of varions is considerably greater. Depending on the gene sequences examined, between 62 and 83% of the codons appear able to fix mutations during the divergences considered. This reflects the constraints of natural selection on acceptable mutations. The total number of base replacements separating the genes for human, mouse, and rabbit beta hemoglobin varies from 61 to 105 depending on the pair examined. Rabbit alpha and beta hemoglobin are separated by at least 290 fixed mutations. For such distantly related sequences estimates made from protein and mRNA data differ less, reflecting the higher quality of information from the many observed changes in primary structure. The effects of nonrandom gene structure on these evolutionary estimates and the fact that various genetic events are not equiprobable are discussed.

Animals↗

A general method for biological inference: illustrated by the estimation of gene nucleotide transition probabilities.

The method of maximum entropy inference developed by Jaynes can be a particularly useful method for obtaining unbiased estimates of biological parameters when the experimental knowledge about a system can be explicitly formulated. Base transition probabilities between genes, though central to evolutionary theory and understanding, present a difficult estimation problem because the ancestral genes are not experimentally accessible. The necessary estimates must therefore be made on the basis of experimental knowledge other than a direct frequency count of base replacements (A leads to C, for example) between contemporary genes. It is shown how maximum entropy inference together with the experimentally observed fact of compositional fidelity in a given gene family can be used to obtain meaningful gene base transition probabilities at each of the three nucleotide positions within codons. Both symmetric and asymmetric transition probabilities are considered. Tables of these probabilities are given for each codon position for the alpha-hemoglobin, beta-hemoglobin, myoglobin, cytochrome c, and the parvalbumin group genes. Tabular values of the average amino acid composition of these five protein families and the average nucleotide composition of their coding genes at varied codon loci are given. It is thus no longer necessary to assume in theories of evolutionary divergence equimolar base ratios A:C:G:U::1:1:1:1 or that each base has an equal chance of mutating to and being fixed as any one of the other three bases.

Amino Acids↗

beta-Galactosidase and selective neutrality.

Three hypotheses to explain the amino acid composition of proteins are inconsistent (P congruent to 10(-9) with the experimental data for beta-galactosidase from Escherichia coli. The exceptional length of this protein, 1021 residues, permits rigorous tests of these hypotheses without complication from statistical artifacts. Either this protein is not at compositional equilibrium, which is unlikely from knowledge about other proteins, or the evolution of this protein and its coding gene have not been selectively neutral. However, the composition of approximately 60 percent of the molecule is consistent with either a selectively neutral or nonneutral evolutionary process.

Amino Acid Sequence↗

Chromatography of a Triton extracted beef cardiac cytochrome C1 preparation on diethylaminoethyl cellulose.

During the chromatography of a Triton X-100 extracted preparation of the mitochondrial membrane proteins on diethylaminoethyl cellulose, we have observed two chromatographic fractions containing cytochrome c1. One elutes from diethylaminoethyl cellulose with aqueous buffers alone, and the other elutes with those buffers after the addition of the nonionic detergent Triton X-100. The two forms occur in equimolar ratios and each retains its original chromatographic character on rechromatography, with no conversion of one form into the other.

Animals↗

Evaluation of compositional nonrandomness in proteins.

Cornish-Bowden and Marson have recently suggested that the finite sampling component of Q, a measure of nonrandomness in the amino acid composition of proteins, may have been underestimated because it was calculated on the basis of the genetic code table frequencies rather than on the basis of the average natural abundance with which the twenty amino acids actually occur in proteins. This underestimate would lead to an overestimate of Qc a measure of selective effects above and beyond those imposed by the average natural abundance of the amino acids. In this paper the finite sampling component of Q is quantitatively estimated on the basis of these natural abundances and found to reduce Qc from its previous average value of 24.3 to the lower value of 9.7, with the standard deviation of the population of Qc values being 12.5. Individual Qc values are given for 81 protein families of mean composition per 61 codons of Ala5.3Arg2.4Asn3.0Asp3.6Cys1.5Gln2.6Glu3.5Gly4.7His1.3Ile3.4Leu4.5Lys4.2Met1.0Phe2.3Pro2.3Ser4.2Thr3.6Trp0.8Tyr2.6Val4.2. The mean Qc value of 9.7 is notably small, and indicates that quantitatively minimal adjustments away from the average protein composition are necessary to maintain many different biological functions. This small value, however, is shown to differ significantly from the value of zero expected were the natural abundances of the amino acids the only selective constraint. These small deviations from the natural abundances are thus effectively selected for in the Darwinian sense.

Amino Acid Sequence↗