Search PubMedSearch

Biomedical subjects

M O Dayhoff

Publications and source records attributed to M O Dayhoff.

8 recordsLinked to original sources

Evolution of homologous physiological mechanisms based on protein sequence data.

1. Genetic duplications can give rise to homologous physiological mechanisms that include structurally related protein components. There are many such examples of related proteins within the human body. 2. Evolutionary histories showing the origins and subsequent divergences of these distantly related proteins can be derived from the protein sequences and correlated with the functional characteristics of these proteins. 3. The hormones related to glucagon provide an example of homology of physiological mechanisms and emergence of new functions subsequent to gene duplications. 4. The proteins related to troponin C illustrate the participation of distantly related proteins in the same mechanism (muscle contraction), the relationship of proteins characteristic of a specialized tissue to proteins found in all eukaryote cells, and the correlation of genetic duplications with the evolutionary appearance of different types of muscle.

Amino Acid Sequence

Methods for identifying proteins by using partial sequences.

Methods for the identification of a protein segment by using the information from partial sequence analyses are described. If the protein sequence is known, a segment can usually be identified with confidence through comparison with the data file of all known sequences when the identity and position of only seven amino acid residues (not necessarily contiguous) are known. Partial sequences are obtained from extremely sensitive microsequencing procedures. Tissue is incubated with amino acids, one or more of which are distinctively radiolabeled. Proteins of interest are isolated and a sequenator experiment performed to locate the positions of radioactivity in an NH(2)-terminal segment of approximately 30 residues. We derive and investigate an equation for the probability of finding a unique match to any pattern of radioactivity. From this we suggest a new strategy. In one incubation, several amino acids are labeled with each kind of isotope. The most information is contained in patterns in which approximately equal numbers of positions are occupied by residues distinguished by different labels (including no label). The amino acid composition of the segment will typically not be known in advance. Labeling residues expected to occupy 36% of the positions suffices for a 98% chance of success in uniquely characterizing any human segment. Such a strategy will permit the identification of most proteins from a single tissue incubation. The mathematical discussion is general and applies to any segment from a sequence and to sequences obtained by any method. Improved identification procedures should expedite the accumulation of information on the expression and function of proteins.

Amino Acid Sequence

A comprehensive examination of protein sequences for evidence of internal gene duplication.

We have implemented a routine procedure for screening protein sequences for evidence of intragenic duplications. We tested 163 protein sequences representing 116 superfamilies of unrelated proteins. Twenty superfamilies contain proteins with internal gene duplications. The intragenic duplications detected can be divided into two major types. (1) One or more duplications of all or part of a gene produce a protein with two or several detectable regions of sequence homology. Sequences from 18 superfamilies contained this type of duplication. (2) Repeated reduplication of a small DNA segment can produce a protein that is repetitive over most of its length. Three superfamilies contain such repetitive sequences. We also investigated the limits of detection of ancient duplications using sequences derived by random mutation of a model sequence consisting of ten 10-residue repeats. The original repetitive nature of the sequence was usually detected after 250 point mutations even though the ancestral segment could not be accurately reconstructed.

Amino Acid Sequence

Evolution of lipoproteins deduced from protein sequence data.

1. Human serum apolipoprotein A-I contains a prominent 11-residue sequence periodicity. 2. Similar 11-residue segments occur in the other sequenced human apolipoproteins, C-I, C-III, and A-II. 3. Computer analyses of the sequences support the hypothesis that they evolved from a common ancestor. 4. An evolutionary history of these proteins is proposed. 5. The estimated rate of change of these proteins indicates that all four types will be found throughout the vertebrates and that related proteins will also be found in invertebrates.

Amino Acid Sequence

The origin and evolution of protein superfamilies.

The organization of proteins into superfamilies based primarily on their sequences is introduced: examples are given of the methods used to cluster the related sequences and to elucidate the evolutionary history of the corresponding genes within each superfamily. Within the framework of this organization, the amount of sequence information currently and potentially available in all living forms can be discussed. The 116 superfamilies already sampled reflect possibly 10% of the total number. There are related proteins from many species in all of these superfamilies, suggesting that the origin of a new superfamily is rare indeed. The proteins so far sequenced are so rigorously conserved by the evolutionary process that we would expect to recognize as related descendants of any protein found in the ancestral vertebrate. The evolutionary history of the thyrotropin-gonadotropin beta chain superfamily is discussed in detail as an example. Some proteins are so constrained in structure that related forms can be recognized in prokaryotes and eukaryotes. Evolution in these superfamilies can be traced back close to the origin of life itself. From the evolutionary tree of the c-type cytochromes the identity of the prokaryote types involved in the symbiotic origin of mitochondria and chloroplasts begins to emerge.

Amino Acid Sequence