Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Polymerase chain reaction and direct sequencing of Neisseria gonorrhoeae protein IB gene: partial nucleotide and amino acid sequence analysis of strains S4, S11, S48 (serovar IB4) and S34 (serovar IB5).

A pair of primers were designed for the polymerase chain reaction (PCR) to amplify a 341-base pair fragment of the gene encoding the outer membrane protein IB (PIB) of Neisseria gonorrhoeae. This PCR technique is specific and sensitive, being able to detect gonococcal strains belonging to ten different PIB serovars, but not PIA gonococcus nor other negative control bacteria. PCR products of four representative PIB strains were directly sequenced. Of the three strains belonging to serovar IB4, two (S11 and S48) shared identical nucleotide and amino acid sequences in the PIB region examined. The third IB4 strain (S4) revealed sequences identical to the published IB26 strain (P9). The sequences of strains P9, S4, S11 and S48 were found to differ from those of strain S34 (serovar IB5). The PCR sequencing technique can further differentiate strains belonging to a common serovar and establish clonal relationships among strains. As a molecular epidemiological tool, the PCR-sequencing strategy can augment existing typing methods including serotyping.

Amino Acid Sequence↗

Sequence analysis of the gene coding for glyceraldehyde-3-phosphate dehydrogenase (gpd) of Podospora anserina: use of homologous regulatory sequences to improve transformation efficiency.

The glyceraldehyde-3-phosphate dehydrogenase (gpd) gene of Podospora anserina has been isolated from a genomic library by heterologous hybridization with the corresponding gene of Curvularia lunata. The coding region consists of 1014 nucleotides and is interrupted by a single intron. The amino-acid sequence encoded by the gpd gene shows a high degree of sequence identity with the corresponding gene products of various fungi. Multiple alignments of all fungal GPD sequences so far available resulted in the construction of a phylogenetic tree. The evolutionary relationships of the various fungi belonging to different taxa will be discussed on the basis of these data. Sequence analysis of 1.9 kbp of the 5' non-coding region revealed the presence of typical fungal promoter elements. Utilizing different parts of the 5' regulatory sequence of the Podospora gpd gene, expression vectors containing a dominant selectable marker gene (hygromycin B phosphotransferase) have been constructed for the transformation of P. anserina protoplasts. The use of these homologous gpd regulatory sequences resulted in a significant increase in transformation efficiencies compared to those obtained with vectors in which the selectable marker gene is under the control of the corresponding heterologous promoter of Aspergillus nidulans.

Amino Acid Sequence↗

Nucleotide sequence of human papillomavirus (HPV) type 41: an unusual HPV type without a typical E2 binding site consensus sequence.

The complete nucleotide sequence of human papillomavirus type 41 (HPV-41) has been determined. HPV-41 was originally isolated from a facial wart, but its DNA has subsequently been detected in some skin carcinomas and premalignant keratoses (Grimmel et al., Int. J. Cancer, 1988, 41, 5-9; de Villiers, Grimmel and Neumann, unpublished results). The analysis of the cloned HPV-41 nucleic acid reveals that its genome organisation is characteristic as for other papillomavirus types. Yet, the analysis indicates at the same time that this virus is most distantly related to all other types of human-pathogenic papillomaviruses sequenced thus far and appears to identify HPV-41 as the first member of a new subgroup of HPV. The overall nucleotide homology to other sequenced HPV types is below 50%. The closest other HPV type is represented by HPV-18, sharing 49% identical nucleotides. The typical E2 binding sequence ACCN6GGT, found in all papillomaviruses analyzed to date, does not occur in the URR of the HPV-41 genome. Modified E2 binding sequences, as described for BPV 1 (Li et al., Genes Dev. 1989, 3, 510-526), are located in the domain proximal to the E6 ORF. These are ACCN6GTT, AACN6GGT and the two perfect palindromic sequences AACGAATTCGTT.

Amino Acid Sequence↗

Sequence polymorphisms in the long terminal repeat of bovine leukemia virus: evidence for selection pressures in regulatory sequences.

Bovine leukemia virus (BLV) is an oncogenic virus widespread in cattle. It belongs to the genus Deltaretrovirus of the family Retroviridae along with human and simian T-lymphotropic viruses. The BLV transcriptional promoter is located in the proviral 5' long terminal repeat (LTR), composed of U3, R, and U5 regions. BLV LTR contains multiple cis-acting elements important for promoter activity, a short coding sequence (encoding the NH(2) terminus of the G4 regulatory protein), and non-regulatory/non-coding regions. Variation in coding sequences of BLV structural proteins has been studied extensively, but little work has been done on sequence variability of non-coding regions, mostly located in LTR. Here, we report the first study on the natural diversity of the BLV LTR, using viral isolates from 52 cattle in several different areas worldwide. Nucleotide variations from the consensus sequence were observed in most isolates and clustered phylogenetically, corresponding to the geographic distribution of donor cattle. Overall, regulatory regions were significantly more conserved than non-regulatory regions in the BLV LTR, as well as in LTR sub-regions (U3, R, and U5). Evidence of selection pressures in BLV LTR suggests that selection occurs not only in coding sequences, but may also involve regulatory sequences.

5' Untranslated Regions↗

Automated protein sequence database classification. II. Delineation Of domain boundaries from sequence similarities.

MOTIVATION: Decomposing each protein into modular domains is a basic prerequisite to classify accurately structural units in biological molecules. Boundaries between domains are indicated by two similar amino acid sequence segments located within the same protein (repeats) or within homologous proteins at notably different distances from their respective N- or C-termini. RESULTS: We have developed an automated method that combines such positional constraints derived from various detected pairwise sequence similarities to delineate the modular organization of proteins. The procedure has been applied to a non-redundant data set of 26 990 proteins whose sequences were taken from the PIR and SWISS-PROT databanks and shared <60% sequence identity amongst pairs. The resultant clustering, delineation and multiple alignment of 24 380 sequence fragments yielded a new database of 4364 domain families. Comparison of the domain collection with that of PRODOM indicates a clear improvement in the number and size of domain families, domain boundaries and multiple sequence alignments. The accuracy and sensitivity of the method are illustrated by results obtained for ankyrin-like repeats and EGF-like modules. AVAILABILITY: The resulting database, called DOMO, is available through the database search routine SRS at Infobiogen (http://www.infobiogen.fr/srs5/), EBI (http://srs.ebi.ac.uk:5000/) and EMBL (http://www.embl-heidelberg.de/srs5/) World Wide Web sites. CONTACT: gracy@infobiogen.fr

Algorithms↗

Purification, cloning and sequence analysis of RsrI DNA methyltransferase: lack of homology between two enzymes, RsrI and EcoRI, that methylate the same nucleotide in identical recognition sequences.

RsrI DNA methyltransferase (M-RsrI) from Rhodobacter sphaeroides has been purified to homogeneity, and its gene cloned and sequenced. This enzyme catalyzes methylation of the same central adenine residue in the duplex recognition sequence d(GAATTC) as does M-EcoRI. The reduced and denatured molecular weight of the RsrI methyltransferase (MTase) is 33,600 Da. A fragment of R. sphaeroides chromosomal DNA exhibited M.RsrI activity in E. coli and was used to sequence the rsrIM gene. The deduced amino acid sequence of M.RsrI shows partial homology to those of the type II adenine MTases HinfI and DpnA and N4-cytosine MTases BamHI and PvuII, and to the type III adenine MTases EcoP1 and EcoP15. In contrast to their corresponding isoschizomeric endonucleases, the deduced amino acid sequences of the RsrI and EcoRI MTases show very little homology. Either the EcoRI and RsrI restriction-modification systems assembled independently from closely related endonuclease and more distantly related MTase genes, or the MTase genes diverged more than their partner endonuclease genes. The rsrIM gene sequence has also been determined by Stephenson and Greene (Nucl. Acids Res. (1989) 17, this issue).

Amino Acid Sequence↗

Conserved noncoding sequences among cultivated cereal genomes identify candidate regulatory sequence elements and patterns of promoter evolution.

Surveys for conserved noncoding sequences (CNS) among genes from monocot cereal species were conducted to assess the general properties of CNS in grass genomes and their correlation with known promoter regulatory elements. Initial comparisons of 11 orthologous maize-rice gene pairs found that previously defined regulatory motifs could be identified within short CNS but could not be distinguished reliably from random sequence matches. Among the different phylogenetic footprinting algorithms tested, the VISTA tool yielded the most informative alignments of noncoding sequence. VISTA was used to survey for CNS among all publicly available genomic sequences from maize, rice, wheat, barley, and sorghum, representing >300 gene comparisons. Comparisons of orthologous maize-rice and maize-sorghum gene pairs identified 20 bp as a minimal length criterion for a significant CNS among grass genes, with few such CNS found to be conserved across rice, maize, sorghum, and barley. The frequency and length of cereal CNS as well as nucleotide substitution rates within CNS were consistent with the known phylogenetic distances among the species compared. The implications of these findings for the evolution of cereal gene promoter sequences and the utility of using the nearly completed rice genome sequence to predict candidate regulatory elements in other cereal genes by phylogenetic footprinting are discussed.

Algorithms↗

cDNA nucleotide sequence and primary structure of mouse uterine peptidylarginine deiminase. Detection of a 3'-untranslated nucleotide sequence common to the mRNA of transiently expressed genes and rapid turnover of this enzyme's mRNA in the estrous cycle.

Peptidylarginine deiminase is a protein-modulating enzyme which converts the arginine residues in proteins to citrulline residues. This study describes the complete primary structure of mouse peptidylarginine deiminase, which was deduced from nucleotide sequence analysis of cDNA clones plus proteochemical analysis of the purified enzyme. The composite cDNA sequence contained a 5' untranslated region of 7 bases, an open reading frame of 2019 bases that encoded 673 amino acids, a 3' untranslated region of 2662 bases, and part of a poly(A) tail. The N-terminal and C-terminal sequences of the enzyme matched the sequences deduced from nucleotide analysis. Furthermore, we determined that the N-terminal sequence was N alpha-acetyl-Met-Gln-, a sequence which has never previously been reported among N alpha-acetyl-Met proteins. The Arg 352 of the enzyme was converted to a citrulline residue and the potential Asn-linked glycosylation site (Asn542-Glu543-Ser544) had no carbohydrate moiety. Thus, mouse peptidylarginine deiminase consists of 673 amino acids with a molecular mass of 76,260. Mouse peptidylarginine deiminase mRNA has two AU-rich structures in the 3' untranslated region which exhibit a high degree of similarity to those in lymphokine, cytokine and proto-oncogene mRNA species. Since the rat enzyme (previously reported) does not possess these characteristic structures, we compared the levels of enzyme activity and mRNA in the mouse and rat uterus at four defined phases of the estrous cycle. The degradation of peptidylarginine deiminase and its mRNA proceeded significantly faster in the mouse than in the rat. We speculate that the unusual structure of the mouse enzyme and its mRNA be involved in this species-specific rapid degradation.

Amino Acid Sequence↗

Cloning and sequencing of the gene encoding glutamine synthetase I from the archaeum Pyrococcus woesei: anomalous phylogenies inferred from analysis of archaeal and bacterial glutamine synthetase I sequences.

The gene glnA encoding glutamine synthetase I (GSI) from the archaeum Pyrococcus woesei was cloned and sequenced with the Sulfolobus solfataricus glnA gene as the probe. An operon reading frame of 448 amino acids was identified within a DNA segment of 1,528 bp. The encoded protein was 49% identical with the GSI of Methanococcus voltae and exhibited conserved regions characteristic of the GSI family. The P. woesei GSI was aligned with available homologs from other archaea (S. solfataricus, M. voltae) and with representative sequences from cyanobacteria, proteobacteria, and gram-positive bacteria. Phylogenetic trees were constructed from both the amino acid and the nucleotide sequence alignments. In accordance with the sequence similarities, archaeal and bacterial sequences did not segregate on a phylogeny. On the basis of sequence signatures, the GSI trees could be subdivided into two ensembles. One encompassed the GSI of cyanobacteria and proteobacteria, but also that of the high-G + C gram-positive bacterium Streptomyces coelicolor (all of which are regulated by the reversible adenylylation of the enzyme subunits); the other embraced the GSI of the three archaea as well as that of the low-G + C gram-positive bacteria (Clostridium acetobutilycum, Bacillus subtilis) and Thermotoga maritima (none of which are regulated by subunit adenylylation). The GSIs of the Thermotoga and the Bacillus-Clostridium lineages shared a direct common ancestor with that of P. woesei and the methanogens and were unrelated to their homologs from cyanobacteria, proteobacteria, and S. coelicolor. The possibility is presented that the GSI gene arose among the archaea and was then laterally transferred from some early methanogen to a Thermotoga-like organism. However, the relationship of the cyanobacterial-proteobacterial GSIs to the Thermotoga GSI and the GSI of low-G+C gram-positive bacteria remains unexplained.

Amino Acid Sequence↗

Human interleukin-9: genomic sequence, chromosomal location, and sequences essential for its expression in human T-cell leukemia virus (HTLV)-I-transformed human T cells.

We have isolated the genomic sequence of human interleukin-9 (IL-9) based on its sequence homology with a human IL-9 cDNA isolated from human T-cell leukemia virus (HTLV)-I-transformed T cells by expression cloning. The entire genomic sequence has been determined and the gene consists of five exons and four introns. The human IL-9 gene is mapped to the long arm of human chromosome 5 at band 5q31-32, a region found to be deleted in a number of patients with acquired 5q- abnormalities and hematologic disorders. Several blocks of transcriptional control sequences have been identified at the 5'-flanking region of the human IL-9 gene that may play an important role in the control of IL-9 gene expression. The 5'-regulatory region of the human IL-9 gene also contains sequences identified in the 5'-flanking regions of other cytokine genes mapped to the long arm of human chromosome 5, including IL-3, IL-4, IL-5, and granulocyte-macrophage colony-stimulating factor and other T-cell growth factor genes including IL-2 and IL-6. The IL-9 gene is constitutively expressed in the HTLV-I-transformed human T cells and the expression of IL-9 in these cells can be further induced by 12-O-tetradecanoyl phorbol 13-acetate. Transient transfection analysis using the plasmid containing the 5'-flanking region of IL-9 gene upstream from the firefly luciferase ciferase report gene indicated that the 0.9-kb Smal-Sacl fragment of the IL-9 gene contains sequences required for the constitutive and activated expression of IL-9 gene in HTLV-I-transformed cells. These results will now allow us to study the regulatory mechanism of IL-9 gene expression in normal and leukemic human T cells.

Amino Acid Sequence↗

Subunit 1 of cytochrome oxidase from Neurospora crassa: nucleotide sequence of the coding gene and partial amino acid sequence of the protein.

A partial protein sequence (223 residues) of cytochrome oxidase subunit 1 from Neurospora crassa has been established. The nucleotide sequence of a cloned mitochondrial DNA segment, including the structural gene coding for the mature subunit 1 (CO I locus) was determined. In contrast to the situation in yeast, the CO I locus in N. crassa is not interrupted by long intervening sequences. A polypeptide of 555 residues with a mol. wt. of 61 000 has been deduced from the reading frame established by protein sequencing. With the exception of the C-terminal part of the polypeptide, the proposed sequences for subunit 1 of N. crassa, yeast, and man are largely homologous. Protein sequencing reveals that a region of low homology close to the C-terminal portion belongs to the structural gene in N. crassa. The DNA sequence coding for the prepiece , which characterizes the polypeptide precursor of the N. crassa subunit 1, has not yet been localized. A RNA species of approximately 6.8 kb has been identified as the CO I transcript. There is no indication of splicing of this large transcript.

Amino Acid Sequence↗

Defining the sequence specificity of the Saccharomyces cerevisiae DNA binding protein REB1p by selecting binding sites from random-sequence oligonucleotides.

We have used a random selection protocol to define the consensus and range of binding sites for the Saccharomyces cerevisiae REB1 protein. Thirty-five elements were sequenced which bound specifically to a GST-REB1p fusion protein coupled to glutathione-Sepharose under conditions in which more than 99.9% of the random sequences were not retained. Twenty-two of the elements contained the core sequence CGGGTRR, with all but one of the remaining elements containing only one deviation from the core. Of the core sequence, the only residues that were absolutely conserved were the three consecutive G residues. Statistical analysis of a nucleotide-use matrix suggested that the REB1p binding site also extends into flanking sequences with the optimal sequence for REB1p binding being GNGCCGGGGTAACNC. There was a positive correlation between the ability of the sites to bind in vitro and activate transcription in vivo; however, the presence of non-conformants suggests that the binding site may contribute more to transcriptional activation than simply allowing protein binding. Interestingly, one of the REB1p binding elements had a DNAse 1 footprint appreciably longer than other elements with similar affinity. Analysis of its sequence indicated the potential for a second REB1p binding site on the opposite strand. This suggests that two closely positioned low-affinity sites can function together as a highly active site. In addition, database searches with some of the randomly defined REB1p binding sites suggest that related elements are commonly found within 'TATA-less' promoters.

Base Sequence↗

An expressed sequence tag database of T-cell-enriched activated chicken splenocytes: sequence analysis of 5251 clones.

The cDNA and gene sequences of many mammalian cytokines and their receptors are known. However, corresponding information on avian cytokines is limited due to the lack of cross-species activity at the functional level or strong homology at the molecular level. To improve the efficiency of identifying cytokines and novel chicken genes, a directionally cloned cDNA library from T-cell-enriched activated chicken splenocytes was constructed, and the partial sequence of 5251 clones was obtained. Sequence clustering indicates that 2357 (42%) of the clones are present as a single copy, and 2961 are distinct clones, demonstrating the high level of complexity of this library. Comparisons of the sequence data with known DNA sequences in GenBank indicate that approximately 25% of the clones match known chicken genes, 39% have similarity to known genes in other species, and 11% had no match to any sequence in the database. Several previously uncharacterized chicken cytokines and their receptors were present in our library. This collection provides a useful database for cataloging genes expressed in T cells and a valuable resource for future investigations of gene expression in avian immunology. A chicken EST Web site (http://udgenome. ags.udel. edu/chickest/chick.htm) has been created to provide access to the data, and a set of unique sequences has been deposited with GenBank (Accession Nos. AI979741-AI982511). Our new Web site (http://www. chickest.udel.edu) will be active as of March 3, 2000, and will also provide keyword-searching capabilities for BLASTX and BLASTN hits of all our clones.

Animals↗

Human cellular sequences detectable with adenovirus probes. I. Evidence for novel repeat sequences and a possible E1a-like cellular "gene".

Previous studies suggesting homology between human cellular DNA and the DNAs from adenovirus types 2 and 5 are extended in the present paper. A clone (ChAdh), isolated from a human genomic DNA library using an adenovirus probe, hybridized to discrete regions of adenovirus 2 DNA, including part of the transforming genes E1a and E1b, as well as to repeated sequences within human DNA. The E1a and E1b genes both hybridize to the same 300 base pair Sau3AI fragment within ChAdh although there is no obvious homology between E1a and E1b. The Ad 2 E1a gene was also used as a probe to screen other cellular DNAs to determine whether repeated sequences detectable with Ad2 DNA probes were conserved over long evolutionary periods. Hybridization was detected to the genomes of man, rat, mouse and fruit fly, but not to those of yeast and bacteria. In addition to a "smear" hybridization, discrete fragments were detected in both rodent and fruit fly DNAs. The experiments reported suggest the existence of two different types of cellular sequences detected by Ad 2 DNA: (1) repeated sequences conserved in a variety of eukaryote genomes and (2) a possible unique sequence detected with an E1a probe different from that responsible for hybridization to repeated sequences. This unique sequence was detected as an EcoRI fragment in mouse DNA and had a molecular size of about 8.8 kb.

Adenoviruses, Human↗

Rigorous pattern-recognition methods for DNA sequences. Analysis of promoter sequences from Escherichia coli.

The basic nature of the sequence features that define a promoter sequence for Escherichia coli RNA polymerase have been established by a variety of biochemical and genetic methods. We have developed rigorous analytical methods for finding unknown patterns that occur imperfectly in a set of several sequences, and have used them to examine a set of bacterial promoters. The algorithm easily discovers the "consensus" sequences for the -10 and -35 regions, which are essentially identical to the results of previous analyses, but requires no prior assumptions about the common patterns. By explicitly specifying the nature of the search for consensus sequences, we give a rigorous definition to this concept that should be widely applicable. We also have provided estimates for the statistical significance of common patterns discovered in sets of sequences. In addition to providing a rigorous basis for defining known consensus regions, we have found additional features in these promoters that may have functional significance. These added features were located on either side of the -35 region. The pattern 5', or upstream, from the -35 region was found using the standard alphabet (A, G, C and T), but the pattern between the -10 and the -35 regions was detectable only in a sub-alphabet. Recent results relating DNA sequence to helix conformation suggest that the former (upstream) pattern may have a functional significance. Possible roles in promoter function are discussed in this light, and an observation of altered promoter function involving the upstream region is reported that appears to support the suggestion of function in at least one case.

Base Sequence↗

Sequence analysis of the viral core protein and the membrane-associated proteins V1 and NV2 of the flavivirus West Nile virus and of the genome sequence for these proteins.

Cell-associated flaviviruses contain the two membrane proteins V3 and NV2 besides the viral core protein V2 whereas extracellular viruses do contain V2 protein and the two membrane proteins V3 and V1. Since the V1 protein could not be detected in infected cells it has been suggested that V1 is generated from NV2 by proteolytic cleavage during the release of virus from cells (D. Shapiro, W. E. Brandt, and P. K. Russell (1972), Virology 50, 906-911). We have isolated the viral structural proteins V1, V2, and NV2 from the flavivirus West Nile virus and determined their amino-terminal amino acid sequences and amino acid sequences of peptides derived from these proteins. We have also transcribed parts of the viral genome into cDNA and cloned and sequenced this cDNA. The analyses of the protein structure of V1, V2, and NV2 together with the determination of the amino-terminal sequence of V3 (data not shown) have allowed us to identify the nucleotide region coding for the structural proteins V2, NV2, and V1. The primary structure of this nucleotide sequence is presented in this report. The data show that the amino terminus of the viral core protein V2 is followed by the amino termini of the proteins NV2, V1, and V3, respectively. These data for the first time identify the exact order of all structural proteins of a flavivirus identified so far. Our data strongly support the above-mentioned hypothesis that V1 is derived from NV2 by proteolytic cleavage and furthermore indicate that V1 represents the nonglycosylated carboxy-terminal part of NV2 which contains those sequences which anchor NV2 in the viral membrane. A working hypothesis is presented in which two species of cellular enzymes, signalase(s) removing signal sequences and enzymes involved in cleaving polyproteins after a pair of basic amino acids, do generate the proteins V2, NV2, and V1 from the growing peptide chain synthesized during translation of the 42 S genome RNA which functions as mRNA for these proteins.

Amino Acid Sequence↗

Rapid sequencing of the Sendai virus 6.8 kb large (L) gene through primer walking with an automated DNA sequencer.

The determination of the complete DNA sequence of the large (L) polymerase gene of Sendai virus strain Fushimi was used to explore the potential and feasibility of primer walking with fluorescent dye-labelled dideoxynucleotide terminators on an automated ABI DNA sequencer. The rapid identification of the complete sequence demonstrated that this approach is a time- and cost-saving alternative to classical sequencing techniques. Analysis of the data revealed that the L gene of Sendai virus strain Fushimi consists of exactly 6800 nucleotides and that the deduced amino acid sequence identifies a single open reading frame encoding a protein of 252.876 kDa. In contrast to Sendai virus strain Enders, the L mRNA of strain Fushimi is monocistronic. The comparison of the deduced amino acid sequences of the L genes of three different Sendai virus strains confirmed the existence of conserved as well as variable regions in the L protein and revealed a high grade of conservation in the carboxyterminal third. Furthermore, functional amino acid sequence motifs, like elements of RNA-dependent RNA polymerases and ATP-binding sites as postulated previously, were identified.

Base Sequence↗

Amino acid sequence of phosvitin derived from the nucleotide sequence of part of the chicken vitellogenin gene.

The amino acid sequence of the egg yolk storage protein phosvitin has been deduced from the nucleotide sequence of part of the chicken vitellogenin gene. Of the phosvitin sequence, 210 amino acids including the N-terminal residue are contained on one large exon, whereas the remaining six amino acids are encoded on the next exon. Phosvitin contains a core region of 99 amino acids, consisting of 80 serines, grouped in runs of maximally 14 residues interspersed by arginines, lysines, and asparagines. The serines of the core region are encoded by AGC and AGT codons exclusively and the arginines by AGA and AGG, which results in a continuous stretch of 99 codons with adenine in the first position. The N-terminal quarter of the phosvitin sequence contains 16 serines grouped in a cluster with alanines and threonines and coded mainly by TCX triplets. The C-terminal part includes 27 serines, preferentially coded by AGC and AGT, 13 histidine residues, and the sequence ...Asn-Gly-Ser... at which the carbohydrate moiety of phosvitin is attached. Heteroduplex formation between cloned DNAs from chicken and Xenopus vitellogenin genes shows that the phosvitin sequence contains a stretch of highly conserved sequence.

Amino Acid Sequence↗