Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

The sialidase gene from Clostridium septicum: cloning, sequencing, expression in Escherichia coli and identification of conserved sequences in sialidases and other proteins.

An oligonucleotide mixture corresponding to the codons for conserved and repeated amino acid sequences of bacterial sialidases (Roggentin et al. 1989) was used to clone a 4.3 kb PstI restriction fragment of Clostridium septicum DNA in Escherichia coli. The complete nucleotide sequence of the sialidase gene was determined from this fragment. The derived amino acid sequence corresponds to a protein of 110,000 Da. The ribosomal binding site and promoter-like consensus sequences were identified upstream from the putative ATG initiation codon. The molecular and immunological properties of the sialidase expressed by E. coli are similar to those of the sialidase as isolated from C. septicum. The newly synthesized protein is assumed to include a leader peptide of 26 amino acids. On sequence alignment, the sialidases from C. septicum, C. sordellii and C. perfringens show significant homologies. As in other bacterial sialidases, conserved amino acid sequences occur at four positions in the protein. Aside from the consensus sequences, only poor homology to other bacterial and viral sialidases was found. The consensus sequence could be identified even in other, non-sialidase proteins, indicating a common function or the evolutionary relatedness of these proteins.

Amino Acid Sequence↗

Expression of two vacuolar-type ATPase B subunit isoforms in swimbladder gas gland cells of the European eel: nucleotide sequences and deduced amino acid sequences.

The poly(A)(+) RNA of swimbladder gas gland cells of the European eel Anguilla anguilla was isolated and used for cDNA synthesis. Using a pair of degenerate PCR primers directed towards the evolutionary highly conserved central part of the B subunit of vacuolar type H(+)-ATPase (V-ATPase) a fragment of 388 bp was amplified. By sequencing the cloned PCR products two different amplicons with a sequence identity of about 86% were obtained. BLASTN searches revealed a high degree of similarity of both to V-ATPase B subunits of other species. The sequences were completed by performing rapid amplification of cDNA ends PCR, subsequent cloning, and sequencing of the obtained products. The expression of two different isoforms of the V-ATPase B subunit is already demonstrated for Homo sapiens and Bos taurus. This is the first report that attributes the same phenomenon to a non-mammalian species, A. anguilla. The first isoform found in eel (vatB2) shows the highest degree of amino acid sequence homology with the human brain isoform (98.2%), the second one (vatB1) with the B subunit sequence of rainbow trout (Oncorhynchus mykiss) gill and kidney (98, 6%). The alignment of the deduced amino acid sequences of vatB1 and vatB2 shows that the highest sequence variation between these two isoforms is found at the amino-terminus, where vatB1 is nine amino acids shorter than vatB2, while at the carboxy-terminus it is two amino acids longer than vatB2. This has also been reported for the human and bovine kidney isoforms when compared with the brain isoforms. Northern blot analysis using specific hybridization probes revealed the expression of two mRNA's with lengths of about 2.9 kb and 3.5 kb for vatB1 and vatB2, respectively. For mammals, it is well known that V-ATPases containing the kidney isoforms of the B subunit are responsible for the extrusion of protons across the plasma membranes of several cell types. The fact that eel vatB1 seems to share structural features with the kidney isoforms in mammals supports the hypothesis that in gas gland cells a V-ATPase contributes to the acidification of the blood in the swimbladder.

Adenosine Triphosphatases↗

Nucleotide sequence heterogeneity of alpha satellite repetitive DNA: a survey of alphoid sequences from different human chromosomes.

The human alpha satellite DNA family is composed of diverse, tandemly reiterated monomer units of approximately 171 basepairs localized to the centromeric region of each chromosome. These sequences are organized in a highly chromosome-specific manner with many, if not all human chromosomes being characterized by individually distinct alphoid subsets. Here, we compare the nucleotide sequences of 153 monomer units, representing alphoid components of at least 12 different human chromosomes. Based on the analysis of sequence variation at each position within the 171 basepair monomer, we have derived a consensus sequence for the monomer unit of human alpha satellite DNA which we suggest may reflect the monomer sequence from which different chromosomal subsets have evolved. Sequence heterogeneity is evident at each position within the consensus monomer unit and there are no positions of strict nucleotide sequence conservation, although some regions are more variable than others. A substantial proportion of the overall sequence variation may be accounted for by nucleotide changes which are characteristic of monomer components of individual chromosomal subsets or groups of subsets which have a common evolutionary history.

Base Sequence↗

The presence of five nifH-like sequences in Clostridium pasteurianum: sequence divergence and transcription properties.

The nifH gene encodes the iron protein (component II) of the nitrogenase complex. We have previously shown the presence in Clostridium pasteurianum of two nifH-like sequences in addition to the nifH1 gene which codes for a protein identical to the isolated iron protein. In the present study, we report that there are at least five nifH-like sequences in C. pasteurianum. DNA sequencing data indicate that the six nifH (nifH1) and nifH-like (nifH2, nifH3, nifH4, nifH5 and nifH6) sequences are not identical and vary from each other to different extents with sequence identity ranging between 68 to 99.9% within the nifH coding regions. Under normal N2-fixing growth conditions (molybdenum-containing medium), transcripts of nifH1 and most of the nifH-like sequences accumulate. The above results suggest the functioning of more than one "nifH" gene under N2-fixing growth conditions for C. pasteurianum. A common sequence was found around the -100 regions of all nif or nif-like transcription units. Sequences identical to or very similar to the consensus Escherichia coli promoter were found in the -35 and -10 regions.

Amino Acid Sequence↗

Histone Sequence Database: a compilation of highly-conserved nucleoprotein sequences.

By searching the current protein sequence databases using sequences from human and chicken histones H1/H5, H2A, H2B, H3 and H4, a database of aligned histone protein sequences with statistically significant sequence similarity to the search sequence was constructed. In addition, a nucleotide sequence database of the corresponding coding regions for these proteins has been assembled. The region of each of the core histones containing the histone fold motif is identified in the protein alignments. The database contains >1300 protein and nucleotide sequences. All sequences and alignments in this database are available through the World Wide Web at http://www.ncbi.nlm.nih.gov/Baxevani/HISTO NES.

Amino Acid Sequence↗

A unique amino acid sequence involved in the putative carbohydrate-binding domain of a legume lectin specific for sialylated carbohydrate chains: primary sequence determination of Maackia amurensis hemagglutinin (MAH).

The primary sequence of 247 amino acids of Maackia amurensis hemagglutinin (MAH) was determined using a protein sequencer. After digestion with endoproteinase Lys-C, Asp-N, Arg-C, or Glu-C of MAH, the resulting peptides were purified by reversed phase high performance liquid chromatography (HPLC) and then subjected to sequence analysis. The primary sequence of MAH was compared with those of several legume lectins, and it was found that the amino acid sequence of the putative carbohydrate-binding domain of MAH exhibited a high degree of homology with those of di-N-acetylchitobiose-binding Cytisus sessilifolius lectin I (CSA-I), Laburnum alpinum lectin I (LAA-I), and Ulex europaeus lectin II (UEA-II). In the legume lectins whose primary sequences have already been determined several amino acid residues involved in carbohydrate-binding were found to be conserved. Very interestingly, in the primary sequence of MAH, one amino acid residue corresponding to the conserved amino acid, asparagine, in the primary sequences of all other legume lectins was shown to be substituted by aspartic acid. This is the first report of the occurrence of an exceptional amino acid residue among the conserved amino acid residues in the carbohydrate-binding domain of the legume lectins.

Amino Acid Sequence↗

Identifying sequence-structure pairs undetected by sequence alignments.

We examine how effectively simple potential functions previously developed can identify compatibilities between sequences and structures of proteins for database searches. The potential function consists of pairwise contact energies, repulsive packing potentials of residues for overly dense arrangement and short-range potentials for secondary structures, all of which were estimated from statistical preferences observed in known protein structures. Each potential energy term was modified to represent compatibilities between sequences and structures for globular proteins. Pairwise contact interactions in a sequence-structure alignment are evaluated in a mean field approximation on the basis of probabilities of site pairs to be aligned. Gap penalties are assumed to be proportional to the number of contacts at each residue position, and as a result gaps will be more frequently placed on protein surfaces than in cores. In addition to minimum energy alignments, we use probability alignments made by successively aligning site pairs in order by pairwise alignment probabilities. The results show that the present energy function and alignment method can detect well both folds compatible with a given sequence and, inversely, sequences compatible with a given fold, and yield mostly similar alignments for these two types of sequence and structure pairs. Probability alignments consisting of most reliable site pairs only can yield extremely small root mean square deviations, and including less reliable pairs increases the deviations. Also, it is observed that secondary structure potentials are usefully complementary to yield improved alignments with this method. Remarkably, by this method some individual sequence-structure pairs are detected having only 5-20% sequence identity.

Algorithms↗

On combining protein sequences and nucleic acid sequences in phylogenetic analysis: the homeobox protein case.

Amino acid encoding genes contain character state information that may be useful for phylogenetic analysis on at least two levels. The nucleotide sequence and the translated amino acid sequences have both been employed separately as character states for cladistic studies of various taxa, including studies of the genealogy of genes in multigene families. In essence, amino acid sequences and nucleic acid sequences are two different ways of character coding the information in a gene. Silent positions in the nucleotide sequence (first or third positions in codons that can accrue change without changing the identity of the amino acid that the triplet codes for) may accrue change relatively rapidly and become saturated, losing the pattern of historical divergence. On the other hand, non-silent nucleotide alterations and their accompanying amino acid changes may evolve too slowly to reveal relationships among closely related taxa. In general, the dynamics of sequence change in silent and non-silent positions in protein coding genes result in homoplasy and lack of resolution, respectively. We suggest that the combination of nucleic acid and the translated amino acid coded character states into the same data matrix for phylogenetic analysis addresses some of the problems caused by the rapid change of silent nucleotide positions and overall slow rate of change of non-silent nucleotide positions and slowly changing amino acid positions. One major theoretical problem with this approach is the apparent non-independence of the two sources of characters. However, there are at least three possible outcomes when comparing protein coding nucleic acid sequences with their translated amino acids in a phylogenetic context on a codon by codon basis. First, the two character sets for a codon may be entirely congruent with respect to the information they convey about the relationships of a certain set of taxa. Second, one character set may display no information concerning a phylogenetic hypothesis while the other character set may impact information to a hypothesis. These two possibilities are cases of non-independence, however, we argue that congruence in such cases can be thought of as increasing the weight of the particular phylogenetic hypothesis that is supported by those characters. In the third case, the two sources of character information for a particular codon may be entirely incongruent with respect to phylogenetic hypotheses concerning the taxa examined. In this last case the two character sets are independent in that information from neither can predict the character states of the other. Examples of these possibilities are discussed and the general applicability of combining these two sources of information for protein coding genes is presented using sequences from the homeobox region of 46 homeobox genes from Drosophila melanogaster to develop a hypothesis of genealogical relationship of these genes in this large multigene family.

Algorithms↗

Sequence and properties of beta-xylosidase from Bacillus pumilus IPO. Contradiction of the previous nucleotide sequence.

The nucleotide sequence of the beta-xylosidase (xynB) gene from Bacillus pumilus has been reported previously [Moriyama, H., Fukusaki, E., Crespo, J.C., Shinmyo, A. & Okada, H. (1987) Eur. J. Biochem. 166, 539-545]. However, the sequence identified in the present study is quite different from the previously reported one. The total length of the PstI--EcoRI fragment of a plasmid pOXN295 containing the xynB gene is 2201 bp from our sequencing, while the length of the fragment in the previous data was 2466 bp. The sequences are similar in the N-terminal (500 bp) and C-terminal (260 bp) regions, but those in the central region are completely different. From the following observations, the previous sequence seems to have no reliable experimental basis. First, the restriction sites observed for pOXN295 are quite different from the sites deduced from the sequence. Second, the amino acid composition deduced from the sequence and the composition identified by amino acid analysis of the purified beta-xylosidase are very different. It is confirmed, on the other hand, that our new sequence agrees well with these experimental data. The enzyme was purified to homogeneity from Bacillus pumilus and Escherichia coli harboring a hybrid plasmid which highly expresses the xynB gene. The molecular mass of the enzyme was estimated to be 190 kDa by high performance gel filtration chromatography using TSK-G3000SW and 56 kDa by SDS/polyacrylamide gel electrophoresis. The pH optimum was 7.0, and the optimum temperature was 40 degrees C. The Vm value was estimated to be 1.23 +/- 0.14 mukat/mg (or p-nitrophenyl beta-D-xyloside) and 0.14 +/- 0.011 mukat/mg (for xylobiose), while Km was estimated to be 3.9 +/- 0.59 mM (for p-nitrophenyl beta-D-xyloside) and 8.9 +/- 1.19 mM (for xylobiose).

Amino Acid Sequence↗

Cloning, sequencing, and characterization of genomic subtracted sequences from Listeria monocytogenes.

Individual sequences of a genomic subtracted, PCR-amplified, mixed-sequence probe (GS probe) were cloned and sequenced. The GS probe differentiated restriction fragment length polymorphism patterns for Listeria monocytogenes but did not hybridize with members of other bacterial genera. Sequence analysis identified several L. monocytogenes sequences already present in the GenBank database; the putative identities of other sequences were inferred from homology data, and still other sequences did not exhibit significant levels of homology with any GenBank sequences.

Amino Acid Sequence↗

L-, P-, and M-ring proteins of the flagellar basal body of Salmonella typhimurium: gene sequences and deduced protein sequences.

The flgH, flgI, and fliF genes of Salmonella typhimurium encode the major proteins for the L, P, and M rings of the flagellar basal body. We have determined the sequences of these genes and the flgJ gene and examined the deduced amino acid sequences of their products. FlgH and FlgI, which are exported across the cell membrane to their destinations in the outer membrane and periplasmic space, respectively, both had typical N-terminal cleaved signal-peptide sequences. FlgH is predicted to have a considerable amount of beta-sheet structure, as has been noted for other outer membrane proteins. FlgI is predicted to have an even greater amount of beta-structure. FliF, as is usual for a cytoplasmic membrane protein of a procaryote, lacked a signal peptide; it is predicted to have considerable alpha-helical structure, including an N-terminal sequence that is likely to be membrane-spanning. However, it had overall a quite hydrophilic sequence with a high charge density, especially towards its C terminus. The flgJ gene, immediately adjacent to flgI and the last gene of the flgB operon, encodes a flagellar protein of unknown function whose deduced sequence was hydrophilic and may correspond to a cytoplasmic protein. Several aspects of the DNA sequence of these genes and their surrounds suggest complex regulation of the flagellar gene system. A notable example occurs within the flgB operon, where between the end of flgG (encoding the distal rod protein of the basal body) and the start of flgH (encoding the L-ring protein) there was an unusually long noncoding region containing a potential stem-loop sequence, which could attenuate termination of transcription or stabilize part of the transcript against degradation. Another example is the interface between the flgB and flgK operons, where transcription termination of the former may occur within the coding region of the latter.

Amino Acid Sequence↗

Signal sequence analysis of expressed sequence tags from the nematode Nippostrongylus brasiliensis and the evolution of secreted proteins in parasites.

BACKGROUND: Parasitism is a highly successful mode of life and one that requires suites of gene adaptations to permit survival within a potentially hostile host. Among such adaptations is the secretion of proteins capable of modifying or manipulating the host environment. Nippostrongylus brasiliensis is a well-studied model nematode parasite of rodents, which secretes products known to modulate host immunity. RESULTS: Taking a genomic approach to characterize potential secreted products, we analyzed expressed sequence tag (EST) sequences for putative amino-terminal secretory signals. We sequenced ESTs from a cDNA library constructed by oligo-capping to select full-length cDNAs, as well as from conventional cDNA libraries. SignalP analysis was applied to predicted open reading frames, to identify potential signal peptides and anchors. Among 1,234 ESTs, 197 (~16%) contain predicted 5' signal sequences, with 176 classified as conventional signal peptides and 21 as signal anchors. ESTs cluster into 742 distinct genes, of which 135 (18%) bear predicted signal-sequence coding regions. Comparisons of clusters with homologs from Caenorhabditis elegans and more distantly related organisms reveal that the majority (65% at P < e-10) of signal peptide-bearing sequences from N. brasiliensis show no similarity to previously reported genes, and less than 10% align to conserved genes recorded outside the phylum Nematoda. Of all novel sequences identified, 32% contained predicted signal peptides, whereas this was the case for only 3.4% of conserved genes with sequence homologies beyond the Nematoda. CONCLUSIONS: These results indicate that secreted proteins may be undergoing accelerated evolution, either because of relaxed functional constraints, or in response to stronger selective pressure from host immunity.

Animals↗

An Ambystoma mexicanum EST sequencing project: analysis of 17,352 expressed sequence tags from embryonic and regenerating blastema cDNA libraries.

BACKGROUND: The ambystomatid salamander, Ambystoma mexicanum (axolotl), is an important model organism in evolutionary and regeneration research but relatively little sequence information has so far been available. This is a major limitation for molecular studies on caudate development, regeneration and evolution. To address this lack of sequence information we have generated an expressed sequence tag (EST) database for A. mexicanum. RESULTS: Two cDNA libraries, one made from stage 18-22 embryos and the other from day-6 regenerating tail blastemas, generated 17,352 sequences. From the sequenced ESTs, 6,377 contigs were assembled that probably represent 25% of the expressed genes in this organism. Sequence comparison revealed significant homology to entries in the NCBI non-redundant database. Further examination of this gene set revealed the presence of genes involved in important cell and developmental processes, including cell proliferation, cell differentiation and cell-cell communication. On the basis of these data, we have performed phylogenetic analysis of key cell-cycle regulators. Interestingly, while cell-cycle proteins such as the cyclin B family display expected evolutionary relationships, the cyclin-dependent kinase inhibitor 1 gene family shows an unusual evolutionary behavior among the amphibians. CONCLUSIONS: Our analysis reveals the importance of a comprehensive sequence set from a representative of the Caudata and illustrates that the EST sequence database is a rich source of molecular, developmental and regeneration studies. To aid in data mining, the ESTs have been organized into an easily searchable database that is freely available online.

Ambystoma↗

Searching sequence databases via de novo peptide sequencing by tandem mass spectrometry.

There are many computer programs that can match tandem mass spectra of peptides to database-derived sequences; however, situations can arise where mass spectral data cannot be correlated with any database sequence. In such cases, sequences can be automatically deduced de novo, without recourse to sequence databases, and the resulting peptide sequences can be used to perform homologous nonexact searches of sequence databases. This article describes details on how to implement both a de novo sequencing program called "Lutefisk," and a version of FASTA that has been modified to account for sequence ambiguities inherent in tandem mass spectrometry data.

Algorithms↗

The complete amino acid sequence of the human transglutaminase K enzyme deduced from the nucleic acid sequences of cDNA clones.

In order to study the expression and role of transglutaminases in the formation of the cross-linked cell envelope of human epidermis, we have used a synthetic oligonucleotide encoding the consensual active site sequence of known transglutaminase sequences. By Northern blot analysis, newborn foreskin epidermis expresses three different mRNA species of about 3.7, 3.3, and 2.9 kilobases while normal cultured epidermal keratinocytes express only the 3.7- and 2.9-kilobase species. The largest species corresponds to a known ubiquitous tissue type II or transglutaminase C activity, the smallest corresponds to a known type I or transglutaminase K activity, and the mid-sized component apparently encodes a transglutaminase E activity that has recently been shown to be expressed in terminally differentiating epidermis (Kim, H. C., Lewis, M. S., Gorman, J. L., Park, S. C., Girard, J. E., Folk, J. E. & Chung, S. I. (1990) J. Biol. Chem., in press). Using the active site oligonucleotide as a probe, we have isolated and sequenced cDNA clones encoding the transglutaminase K enzyme. The deduced complete protein sequence has 813-amino acid residues of 89.3 kDa, has a pl of 5.7, and is likely to be an essentially globular protein, which are properties expected from the partially purified enzyme. It shares 49-53% sequence homology with the other transglutaminases of known sequence, especially in regions carboxyl-terminal to the active site, and possesses sequences likely to confer its Ca2+ dependence. Interestingly, its larger size is due to extended sequences on its amino and carboxyl termini, absent on the other transglutaminases, that may define its unique properties.

Amino Acid Sequence↗

[Research on the recombinant plasmid pDJH2 of L. interrogans serovar lai: sequencing and alignment with other known bacterial Omp sequence].

The Leptospira whole cell vaccine (LWCV) currently used in China is safe and effective, out the immunity following vaccination with two doses of the fluid medium vaccine is of low order. The duration of immunity conferred by this vaccine is rather short, six months or at most one year. Therefore, it is necessary to develop new generation vaccines against Leptospirosis for the developing world. In this paper we report the sequencing of the insert fragment of pDJH2 from genomic DNA of L. interrogans sevovar lai strain 017 and its alignment with other bacterial omp sequences. A genomic library of Leptospira interrogaans serovar lai strain 017 was constructed with the plasmid vector pUC18. A recombinant plasmid designated pJDH2 was screened from the genomic library. Inserted fragment of pDH2 is 1.9 kb by gel electrophoresis. Immunization/protection was studied in BALB/c mice model. The results showed highly significant difference between pDJH2 and pUC18 (control). Inserted fragment of pDJH2 DNA sequencing was performed by Dr Yan Zhengxin (Max-Planck-Institut for Biology. Tubingen, Germany). Insert fragment was cloned into pBluescript II KS-(stratagene) and sequenced by using AB1 (Applied Bio Systems, Model 373A). Two open reading frames of 565 and 662 nucleotides were identified. There were identifiable initiation codons, terminators, Shine-Dalgano ribosome combining site, Pribnow boxes and Sextama boxes within the 2 sequenced regions. Nucleotide sequences were analysed using Gene Work, a suit of computer program developed by Department of Biochemistry St. Jude Children's Research Hospital Memphis. U.S.A. The results of formatted alignment showed the predicted nucleotide sequence of ORF1 of the serovar lai had significant similarity with ORF2 (49.36%). L. kirschneri ompL1 (49.26%), Borrelia burgdoferi omp (48.97%), Treponema phagedenis omp (47.3%); Salmonella typhimurium ompC(46.87%), Yersinia enterocolitica ompH (46.7%), Leptospira borgpeterseni pfap (46.3%), and Serratia marcescens omp (43.3%). The close relationship of the pDJH2 ORF1 and ORF2 nucleotide sequences from Leptospira kirschneri ompL 1 is apparent. Whether the recombinant pDJH2 will prove useful for vaccine development remains to be tested.

Animals↗

Complete nucleotide sequence of an Indian strain of Japanese encephalitis virus: sequence comparison with other strains and phylogenetic analysis.

The RNA genome of an Indian strain of Japanese encephalitis virus (JEV), GP78, was reverse transcribed and the cDNA fragments were cloned in bacterial plasmids. Nucleotide sequencing of the cDNA clones covering the entire genome of the virus established that the GP78 genome was 10,976 nucleotides long. An open reading frame of 10,296 bases, capable of coding for a 3,432 amino acid polyprotein, was flanked by 95- and 585-base long 5'- and 3'-non-coding regions, respectively. When compared with the nucleotide sequence of the JaOArS982 strain, the JEV GP78 genome had a number of nucleotide substitutions that were scattered throughout the genome except for the 5'-noncoding region, the sequence of which was fully conserved. Comparison of the complete genome sequences of different JEV isolates showed a 1.3-4.1% nucleotide sequence divergence among them, which resulted in 0.6-1.8% amino acid sequence divergence. Analysis based on the complete genome sequences of different JEV isolates showed that the GP78 isolate from India was phylogenetically closer to the Chinese SA14 isolate.

Adult↗

Determination of the primary sequence of the duck alpha D globin mRNA and comparison of all adult duck and chick globin mRNA sequences.

The nucleotide sequence of the duck alpha D globin mRNA was determined. Its main feature is an exceptionally short 3' non-coding segment of only 46 nucleotides, placed after the coding sequence of 141 codons. The last of the 6 adult globin mRNA of duck and chicken being thus sequenced, a comparison of all their features has become possible. Comparing the duck alpha D mRNA to the related sequence in the chicken, we found greater homology than comparing it to the linked alpha A globin sequence in the same species. Extensive homology can be found for a same globin chain alpha A, alpha D or beta in between different avian species including also the goose and the ostrich; the avian alpha globin chains show a lower degree of sequence conservation in between species than the beta chains. In contrast, within one species the three globin sequences have further diverged. The divergence between the alpha A and alpha D globin within a same species point to individual functional specificity and hence independent evolution and suggest that a mechanism of 'gene conversion' did not operate in between the avian alpha globin genes. Two segments of the amino acid sequence which we named 'A alpha' and 'B alpha' remain homologous in all avian alpha globins; two other regions 'A beta' and 'B beta' are identical in between the beta globins. Segment A is placed at the 5' end of exon II, and segment B at the 3' end of the same exon; some amino acids in those segments are involved in the Heme binding site. Being almost identical in all know mammalian and avian globins of the alpha respectively the beta type, regions A and B seem to represent the best conserved sequences in adult globin mRNA maintained during the divergence of species.

Animals↗