Search PubMed⌕ Search

Biomedical subjects

R F Doolittle

Publications and source records attributed to R F Doolittle.

At least 91 records · Page 5Linked to original sources

Construction of a facsimile data set for large genome sequence analysis.

A test was devised for exploring the question of whether it will be possible to identify genes in large-scale genome studies solely by sequence comparison with current sequence collections. To this end, a facsimile data set was constructed by dividing GenBank Release 56 randomly into two halves, one to serve as a reference set and the other intended to simulate raw data anticipated from large genome sequence projects. All supplementary information and identifying marks were removed from the test set after assignment of random identification numbers to each entry and their encryption. Because noncoding intervening sequences (introns) are underrepresented in GenBank, a program that introduced (simulated) introns into mRNA and prokaryotic sequences was devised. In a further attempt to make the problem of identification more realistic, random base substitutions and single-base deletions were also incorporated. The randomly ordered entries were concatenated, along with random intergenic flanking sequences, into a single long "chromosome" 33 Mb in length and then cut into "cosmids" 50-100 kb long. The chopping process was conducted in such a way that terminal overlaps would allow the order of the entries in the chromosome to be reconstituted. Finally, the sequences of a substantial fraction of the cosmids were converted to their complements. Preliminary searching of 10 test cosmids revealed that more than two-thirds of the entries in the test set should be readily identifiable by type of gene product solely on the basis of comparison with the reference set. These preliminary results suggest that existing computer regimens and sequence collections would be able to identify the majority of eukaryotic genes in any new raw data set, the existence of introns not withstanding. Moreover, the analysis can be conducted in pace with the data collection so that the search results and summary identifications will be instantly available to the research community at large.

Animals↗

Presence of a vertebrate fibrinogen-like sequence in an echinoderm.

Sequence comparisons of the three homologous polypeptide chains that compose vertebrate fibrinogens imply that the molecule evolved before the divergence of vertebrates and invertebrates, but, to our knowledge, no protein resembling vertebrate fibrinogen has even been reported from an invertebrate. We used primers based on sequences conserved between lamprey and human fibrinogens and applied the polymerase chain reaction (PCR) to cDNA preparations from various invertebrates. A fibrinogen-like sequence was identified in cDNA prepared from the soft tissues of a sea cucumber, parastichopus parvimensis. The PCR-prepared material was then used to clone two closely related mRNA sequences from a sea cucumber soft tissue cDNA library. The putative fibrinogen-related proteins, FReP-A and FReP-B, correspond to the carboxyl-terminal two-thirds of vertebrate fibrinogen beta and gamma chains. Computer comparisons of various fibrinogen-related sequences indicate that the sea cucumber proteins diverged before the beta-gamma gene duplication.

Amino Acid Sequence↗

Complete sequence of the lamprey fibrinogen alpha chain.

The complete amino acid sequence of the lamprey fibrinogen alpha chain has been determined by a combination of peptide sequencing and cDNA and genomic cloning. The chain, which has an apparent molecular weight by dodecyl sulfate-polyacrylamide gel electrophoresis of ca. 100,000, is composed of 961 amino acid residues and has a calculated molecular weight of 96,722. It is distinguished by a large number of 18-residue repeats in a region where mammalian fibrinogens have 13-residue repeats. The data are in accord with our previous finding that the lamprey alpha chain has a distinctive amino acid composition, almost half the residues being glycine, serine, or threonine. The chain differs from mammalian alpha chains in that there are no cysteines in the carboxy-terminal half, and thus no intrachain loop, nor are there any RGD sequences in the lamprey alpha chain. Taken together with previous data on the sequences of the beta and gamma chains, the findings bear significantly on our understanding of fibrin formation. The alpha chain also provides an interesting case of structural convergence during evolution.

Amino Acid Sequence↗

Similar amino acid sequences revisited.

The rapid accumulation of protein sequences, many bearing unexpected resemblances to each other, is providing a new perspective on evolution.

Amino Acid Sequence↗

Origins and evolutionary relationships of retroviruses.

As is the case for some other RNA viruses, the amino acid sequences of retroviral proteins change at an astonishing rate. For example, the proteases of the human immunodeficiency virus (HIV) and the visna lentivirus with which it is often compared are as different as the proteases of fungi and mammals, and those of the human type I leukemia virus are as different from HIV or visna as are the proteins of humans and bacteria. That the sequences of retrovirus proteins can be recognized as sharing common ancestry with non-retroviral proteins implies that the vastly accelerated change has begun only recently or occurs very sporadically. Only a scheme whereby exogenous retroviruses exist as short-lived bursts upon a backdrop of germline-encoded endogenous viruses is consistent with the sequence data. Retroviruses are related to many other reverse transcriptase-bearing entities present in the genomes of eukaryotes. They also have proteins that are homologous with those of some plant and animal DNA viruses, and their reverse transcriptase is recognizably similar to sequences found in the introns of some fungal mitochondria. Computer alignment of all these sequences allows an overall phylogeny to be constructed that chronicles the history of events leading to infectious retroviruses.

Animals↗

Sequence comparisons of retroviral proteins: relative rates of change and general phylogeny.

The inferred amino acid sequences of 10 specific gene products from nine retroviruses were aligned by computer, all evolutionary distances between them calculated, and evolutionary trees constructed. Not unexpectedly, the various gene products are changing at different rates, the reverse transcriptase being the least and the envelope proteins the most different from one retrovirus to another. For the most part, trees based on the retroviral enzyme sequences are congruent, indicating that extensive genetic recombination has not been a major factor in the evolution of the central part of the genome. In the case of envelope protein sequences, however, the sequences clearly exhibit evidence of multiple cross-over events between quite distantly related retroviruses. A composite phylogenetic tree was constructed from the four retroviral enzyme sequences, and a number of important historical happenings were interpreted in the light of the time scale it affords.

Amino Acid Sequence↗

cDNA sequences of two apolipoproteins from lamprey.

The messages for two small but abundant apolipoproteins found in lamprey blood plasma were cloned with the aid of oligonucleotide probes based on amino-terminal sequences. In both cases, numerous clones were identified in a lamprey liver cDNA library, consistent with the great abundance of these proteins in lamprey blood. One of the cDNAs (LAL1) has a coding region of 105 amino acids that corresponds to a 21-residue signal peptide, a putative 8-residue propeptide, and the 76-residue mature protein found in blood. The other cDNA (LAL2) codes for a total of 191 residues, the first 23 of which constitute a signal peptide. The two proteins, which occur in the "high-density lipoprotein fraction" of ultracentrifuged plasma, have amino acid compositions similar to those of apolipoproteins found in mammalian blood; computer analysis indicates that the sequences are largely helix-permissive. When the sequences were searched against an amino acid sequence data base, rat apolipoprotein IV was the best matching candidate in both cases. Although a reasonable alignment can be made with that sequence and LAL1, definitive assignment of the two lamprey proteins to typical mammalian classes cannot be made at this point.

Amino Acid Sequence↗

Progressive sequence alignment as a prerequisite to correct phylogenetic trees.

A progressive alignment method is described that utilizes the Needleman and Wunsch pairwise alignment algorithm iteratively to achieve the multiple alignment of a set of protein sequences and to construct an evolutionary tree depicting their relationship. The sequences are assumed a priori to share a common ancestor, and the trees are constructed from difference matrices derived directly from the multiple alignment. The thrust of the method involves putting more trust in the comparison of recently diverged sequences than in those evolved in the distant past. In particular, this rule is followed: "once a gap, always a gap." The method has been applied to three sets of protein sequences: 7 superoxide dismutases, 11 globins, and 9 tyrosine kinase-like sequences. Multiple alignments and phylogenetic trees for these sets of sequences were determined and compared with trees derived by conventional pairwise treatments. In several instances, the progressive method led to trees that appeared to be more in line with biological expectations than were trees obtained by more commonly used methods.

Algorithms↗

Homology between the DNA-binding domain of the GCN4 regulatory protein of yeast and the carboxyl-terminal region of a protein coded for by the oncogene jun.

The product of the recently described oncogene jun shows significant amino acid sequence homology with the GCN4 yeast transcriptional activator protein. The similarity is restricted to the 66 carboxyl-terminal amino acids, thought to be the DNA-binding domain of the GCN4 protein. In these alpha-helix-permissive regions of the jun and GCN4 products there is also a lesser but still significant amino acid resemblance to the fos protein and a marginal degree of similarity to myc proteins. The amino acid sequence homology between GCN4 and jun gene products suggests that the jun protein may bind to DNA in a sequence-specific way and exert a regulatory function.

Amino Acid Sequence↗

Relocation of a protease-like gene segment between two retroviruses.

An anomalous sequence in certain lentiviruses was found to be related to a region in a completely different part of the simian retrovirus type I (SRV-I) and its close relative, the hamster intracisternal A particle (IAP-H18). The segment is not present in the human immunodeficiency virus (HIV), which is also a lentivirus, nor is it found in any one of a dozen other retroviruses whose sequences have been reported. These observations imply that a horizontal transfer of newly acquired genetic information has taken place between an SRV-I-type virus and one of the lentivirus type, and that this event occurred more recently than did the divergence of members of this latter group and HIV. Comparison of the viral nucleic acid sequences that encode these segments revealed the presence of imperfect direct nucleotide repeats resembling the retroviral endonuclease cleavage sites at the 5' and 3' ends of these regions.

Amino Acid Sequence↗

URF6, last unidentified reading frame of human mtDNA, codes for an NADH dehydrogenase subunit.

The polypeptide encoded in URF6, the last unassigned reading frame of human mitochondrial DNA, has been identified with antibodies to peptides predicted from the DNA sequence. Antibodies prepared against highly purified respiratory chain NADH dehydrogenase from beef heart or against the cytoplasmically synthesized 49-kilodalton iron-sulfur subunit isolated from this enzyme complex, when added to a deoxycholate or a Triton X-100 mitochondrial lysate of HeLa cells, specifically precipitated the URF6 product together with the six other URF products previously identified as subunits of NADH dehydrogenase. These results strongly point to the URF6 product as being another subunit of this enzyme complex. Thus, almost 60% of the protein coding capacity of mammalian mitochondrial DNA is utilized for the assembly of the first enzyme complex of the respiratory chain. The absence of such information in yeast mitochondrial DNA dramatizes the variability in gene content of different mitochondrial genomes.

Amino Acid Sequence↗

Complementary DNA sequence of lamprey fibrinogen beta chain.

The cDNA sequence of the beta chain of lamprey fibrinogen has been determined. To that end, an oligonucleotide probe was synthesized that corresponded to an amino acid sequence from the carboxy-terminal region of the lamprey fibrinogen beta chain. The insert actually began with residue 3 of the fibrin beta chain; it ran through to a terminator codon following the carboxy-terminal residue at position 443 and then continued for an additional 606 nucleotides of noncoding sequence to its 3' end. The inferred amino acid sequence was verified by comparison with assorted cyanogen bromide fragments isolated from the beta-chain protein, including two carbohydrate-containing peptides that corresponded to segments containing the carbohydrate-attachment consensus sequence. Overall, the lamprey chain is 49% identical with the beta chain from human fibrinogen. This is the same degree of resemblance as was found for the lamprey and human gamma chains. Moreover, the principal regions of conservation are the same in both the beta and gamma chains. Differences and similarities in the physiological behavior of the two fibrinogens are assessed in terms of the observed amino acid replacements.

Amino Acid Sequence↗

Antibodies against the COOH-terminal undecapeptide of subunit II, but not those against the NH2-terminal decapeptide, immunoprecipitate the whole human cytochrome c oxidase complex.

Antibodies against synthetic peptides derived from the DNA sequence of human cytochrome c oxidase subunit II (COII) have been tested for their capacity to immunoprecipitate the whole enzyme complex. Antibodies against the COOH-terminal undecapeptide of COII (anti-COII-C), when incubated with a Triton X-100 mitochondrial lysate from HeLa cells pulse-labeled with [35S]methionine under conditions selective for mitochondrial protein synthesis and chased for 18 h in unlabeled medium, precipitated the pulse-labeled three largest subunits (mitochondrially synthesized) of cytochrome c oxidase in proportions close to equimolarity. Antibodies against the NH2-terminal decapeptide of COII (anti-COII-N), although equally reactive as the anti-COII-C antibodies with the sodium dodecyl sulfate-solubilized COII, did not precipitate any of the three labeled subunits from the Triton X-100 mitochondrial lysate. In other experiments, all the 13 subunits which have been identified in the mammalian cytochrome c oxidase were immunoprecipitated from a Triton X-100 mitochondrial lysate of cells long-term labeled with [35S]methionine by anti-COII-C antibodies, but not by anti-COII-N antibodies. By contrast, in immunoblots of total mitochondrial proteins dissociated with sodium dodecyl sulfate, the anti-COII-C antibodies reacted specifically only with COII. These results strongly suggest that, in the native cytochrome c oxidase complex, the epitope recognized by the anti-COII-C antibodies is in the COII subunit and that, therefore, in such complex, the COOH-terminal peptide of COII is exposed to antibodies, whereas the NH2-terminal peptide is not accessible.

Amino Acid Sequence↗