Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Interaction of the H4 autonomously replicating sequence core consensus sequence and its 3'-flanking domain.

Yeast autonomously replicating sequence (ARS) elements are composed of a conserved 11-base-pair (bp) core consensus sequence and a less well defined 3'-flanking region. We have investigated the relationship between the H4 ARS core consensus sequence and its 3'-flanking domain. The minimal sequences necessary and sufficient for function were determined by combining external 3' and 5' deletions to produce a nested set of ARS fragments. Sequences 5' of the core consensus were dispensable for function, but at least 66 bp of 3'-flanking domain DNA was required for full ARS function. The importance of the relative orientation of the core consensus element with respect to the 3'-flanking domain was tested by precisely inverting 14 bp of DNA including the core consensus sequence by oligonucleotide mutagenesis. This core inversion mutant was defective for all ARS function, showing that a fixed relative orientation of the core consensus and 3'-flanking domain is required for function. The 3'-flanking domain of the minimal functional H4 ARS fragment contains three sequences with a 9-of-11-bp match to the core consensus. The role of these near-match sequences was tested by directed mutagenesis. When all near-match sequences with an 8-of-11-bp match or better were simultaneously disrupted by point mutations, the resulting ARS construct retained full replication function. Therefore, multiple copies of a sequence closely related to the core consensus element are not required for H4 ARS function.

Base Sequence↗

Cloning and nucleotide sequence analysis of the dog insulin gene. Coded amino acid sequence of canine preproinsulin predicts an additional C-peptide fragment.

A 4.0-kilobase HindIII/EcoRI-cleaved dog genomic DNA fragment was shown to contain the dog insulin gene by restriction mapping using a human insulin cDNA probe. This fragment was subsequently cloned in a lambda vector, and the nucleotide sequence of the dog insulin gene was determined. As in several other species, the insulin gene of the dog is interrupted by two intervening sequences, one of 151 base pairs located in the 5' untranslated region and the other of 264 base pairs occurring within the codon of the 7th amino acid of the C-peptide. Translation of the nucleotide sequence in one frame revealed the primary structure of canine preproinsulin. An interesting feature of the coded amino acid sequence is that it predicts a C-peptide of 31 amino acids, 8 residues longer than that reported by Peterson et al. (Peterson, J. D., Nehrlich, S., Oyer, P. E., and Steiner, D. F. (1973) J. Biol. Chem. 247, 4866-4871). The additional octapeptide sequence, Glu-Val-Glu-Asp-Leu-Gln-Val-Arg, is located NH2-terminal to the 23-residue C-peptide sequence described in the earlier report. Its coding sequence is interrupted by the second intervening sequence. The arginine at position 8 suggests that a trypsin-like cleavage may separate the NH2-terminal octapeptide from the remainder of the C-peptide during the post-translational processing of dog proinsulin in the pancreas. The revised C-peptide sequence suggests that the proinsulin C-peptide is more highly conserved in length and overall sequence than was previously supposed.

Amino Acid Sequence↗

Sequence analysis of the right end of chromosome XV in Saccharomyces cerevisiae: an insight into the structural and functional significance of sub-telomeric repeat sequences.

Approximately 3.9 kb of DNA, centromere proximal to the previously sequenced Y' element at the right end of chromosome XV in Saccharomyces cerevisiae strain YP1, has been sequenced. A number of the known sub-telomeric repeat sequences were identified, including Y', core X and STRs A, B. C and D. Several of these repeat elements contain potentially functional sequences. In addition, two other members of repeated gene families were identified. The first of these shows 61% and 60% DNA sequence identity to Enolases 1 and 2 respectively. The Enolase-like sequence appears to be species specific, with three copies being found in all strains of S. cerevisiae studied. The location of the three copies is the same for all strains. The second repeated sequence has homology with known open reading frames on chromosomes III, V and XI. There are five or six copies of this sequence in all S. cerevisiae and S. paradoxus strains studied and three in S. bayanus strains. The analysis of this region and comparison to sub-telomeric regions on other chromosomes gives some indication as to the potential functional and structural significance of sub-telomeric repeat sequences. In addition, these findings are consistent with the idea that sub-telomeric regions may be targets for unusual recombination events.

Amino Acid Sequence↗

An expressed-sequence-tag database of the human prostate: sequence analysis of 1168 cDNA clones.

The human prostate is a complex glandular organ with functional development under hormonal regulation. Diseases of the prostate result in significant morbidity and mortality in the form of benign prostatic hypertrophy and prostate adenocarcinoma. The characterization of the molecular framework of the human prostate at the level of expressed genes will facilitate the understanding of normal and pathological prostate biology. The purposes of this study were to acquire an initial assessment of the qualitative and quantitative diversity of gene expression in the normal human prostate and to determine the extent that genes with prostate-restricted expression can be assessed using an expressed sequence tag approach. We have constructed a directional cDNA library from normal adult human prostate tissue and partially sequenced the 5' end of 1168 randomly selected cDNA clones, resulting in more than 400 kb of DNA sequence. Homology searches of the sequenced cDNAs against the GenBank and dbEST databases revealed that 43% of the sequences are identical to human genes whose functions are known, 5% are similar but not identical to known genes in humans or lower organisms, 5% match the mitochondrial genome, 9% are composed of interspersed DNA repeats, 30% are homologous to sequences in the dbEST database without a described function, and 6% are novel sequences. A total of 780 distinct species were identified. In addition to the 74 novel transcripts, 4 genes, prostate-specific antigen (PSA), prostate secretory protein (PSP), prostate acid phosphatase (PAP), and human glandular kallekrein 2 (HK2), have no homologous sequences in the databases that originate from sources other than prostate and thus may represent genes with prostate-restricted expression. Sequences matching PSA, PSP, and PAP each accounted for > 1% of the total ESTs and represent highly abundant transcripts, correlating with the abundance of these proteins in the prostate gland. No novel transcripts were represented by more than one EST and thus are expressed at levels much lower than the known prostate-specific genes.

Adult↗

Cloning and nucleotide sequence of the Salmonella typhimurium LT2 metF gene and its homology with the corresponding sequence of Escherichia coli.

The Salmonella typhimurium LT2 metF gene, encoding 5,10-methylenetetrahydrofolate reductase, has been cloned. Strains with multicopy plasmids carrying the metF gene overproduce the enzyme 44-fold. The nucleotide sequence of the metF gene was determined, and an open reading frame of 888 nucleotides was identified. The polypeptide deduced from the DNA sequence contains 296 amino acids and has a molecular weight of 33,135 daltons. Mung bean nuclease mapping experiments located the transcription start point and possible transcription termination region for the gene. There is a 25 bp nucleotide sequence between the translation termination site and the possible transcription termination region. This region possesses a GC-rich sequence that could form a stable stem and loop structure once transcribed (delta G = -9 kcal/mol), followed by an AT-rich sequence, both of which are characteristic of rho-independent transcription terminators. The nucleotide and deduced amino acid sequences of the S. typhimurium metF gene are compared with the corresponding sequences of the Escherichia coli metF gene. The nucleotide sequences show 85% homology. Most of the nucleotide differences found do not alter the amino acid sequences, which show 95% homology. The results also show that a change has occurred in the metF region of the S. typhimurium chromosome as compared to the E. coli chromosome.

Amino Acid Sequence↗

Complete nucleotide sequence of the rabbit beta-like globin gene cluster. Analysis of intergenic sequences and comparison with the human beta-like globin gene cluster.

The nucleotide sequence of the entire beta-like globin gene cluster of rabbits has been determined. This sequence of a continuous stretch of 44.5 x 10(3) base-pairs (bp) starts about 6 x 10(3) bp upstream from epsilon (the 5'-most gene) and ends about 12 x 10(3) bp downstream from beta (the 3'-most gene). Analysis of the sequence reveals that: (1) the sequence is relatively A + T rich (about 60%); (2) regions with high G + C content are associated with OcC repeats, a short interspersed repeated DNA in rabbits; (3) the distribution of polypurines, polypyrimidines and alternating purine/pyrimidine tracts is not random within the cluster; (4) most open reading frames are associated with known globin coding regions, OcC repeats or long interspersed repeats (L1 repeats); (5) the most prominent open reading frames are found in the L1 repeats; (6) different strand asymmetries in base composition are associated with embyronic and adult genes as well as the tandem L1 repeats at the 3' end of the cluster; and (7) essentially all the repeats appear to have been inserted by a transposon mechanism. A comparison of the sequence with itself by a dot-plot analysis has revealed nine new members of the OcC family of repeats in addition to the six previously reported. The OcC repeats tend to be clustered, particularly in the epsilon-gamma and gamma-psi delta intergenic regions. Dot-plot comparisons between the rabbit and the human clusters have revealed extensive sequence matches. Homology starts about 6 x 10(3) bp 5' to epsilon or as far upstream as the rabbit sequence is available. It continues throughout the entire cluster and stops about 0.7 x 10(3) bp 3' to beta, at which point several repeats have inserted in both rabbits and humans. Throughout the gene cluster, the homology is interrupted mainly by insertions or deletions in either the rabbit or the human genome. Almost all of the insertions are of known short or long repeated DNAs. The positions of the insertions are different in the two gene clusters, which indicates that both short and long repeats have been transposing throughout the genome for the time since the mammalian radiation. An alignment of rabbit and human sequences allows the calculation of the substitution rate around epsilon. Sequences far removed from the gene are evolving at a rate equivalent to the pseudogene rate, although some short regions show an apparently higher rate.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals↗

Amino acid sequence of guinea pig liver transglutaminase from its cDNA sequence.

Transglutaminases (EC 2.3.2.13) catalyze the formation of epsilon-(gamma-glutamyl)lysine cross-links and the substitution of a variety of primary amines for the gamma-carboxamide groups of protein-bound glutaminyl residues. These enzymes are involved in many biological phenomena. In this paper, the complete amino acid sequence of guinea pig liver transglutaminase, a typical tissue-type nonzymogenic transglutaminase, was predicted by the cloning and sequence analysis of DNA complementary to its mRNA. The cDNA clones carrying the sequences for the 5'- and 3'-end regions of mRNA were obtained by use of the sequence of the partial-length cDNA of guinea pig liver transglutaminase [Ikura, K., Nasu, T., Yokota, H., Sasaki, R., & Chiba, H. (1987) Agric. Biol. Chem. 51, 957-961]. A total of 3695 bases were identified from sequence data of four overlapping cDNA clones. Northern blot analysis of guinea pig liver poly(A+) RNA showed a single species of mRNA with 3.7-3.8 kilobases, indicating that almost all of the mRNA sequence was analyzed. The composite cDNA sequence contained 68 bases of a 5'-untranslated region, 2073 bases of an open reading frame that encoded 691 amino acids, a stop codon (TAA), 1544 bases of a 3'-noncoding region, and a part of a poly(A) tail (7 bases). The molecular weight of guinea pig liver transglutaminase was calculated to be 76,620 from the amino acid sequence deduced, excluding the initiator Met. This enzyme contained no carbohydrate [Folk, J. E., & Chung, S. I. (1973) Adv. Enzymol. Relat. Areas Mol. Biol. 38, 109-191], but six potential Asn-linked glycosylation sites were found in the sequence deduced.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

A protein binding to the J kappa recombination sequence of immunoglobulin genes contains a sequence related to the integrase motif.

Site-specific recombination requires conserved DNA sequences specific to each system, and system-specific proteins that recognize specific DNA sequences. The site-specific recombinases seem to fall into at least two families, based on their protein structure and chemistry of strand breakage. One of these is the resolvase-invertase family, members of which seem to form a serine-phosphate linkage with DNA. Members of the other family, called the integrase family, contain a conserved tyrosine residue that forms a covalent linkage with the 3'-phosphate of DNA at the site of recombination. Structural comparison of integrases shows that these proteins share a highly conserved 40-residue motif. V-(D)-J recombination of the immunoglobulin gene requires conserved recombination signal sequences (RS) of a heptamer CACTGTG and a T-rich nonamer GGTTTTTGT, which are separated by a spacer sequence of either 12 or 23 bases We have recently purified, almost to homogeneity, a protein that specifically binds to the immunoglobulin J kappa RS containing the 23-base-pair spacer sequence. By synthesizing probes on the basis of partial amino-acid sequences of the purified protein, we have now isolated and characterized the complementary DNA of this protein. The amino-acid sequence deduced from the cDNA sequence reveals that the J kappa RS-binding protein has a sequence similar to the 40-residue motif of integrases of phages, bacteria and yeast, indicating that this protein could be involved in V-(D)-J recombination as a recombinase.

Amino Acid Sequence↗

Sequence ordinations: a multivariate analysis approach to analysing large sequence data sets.

Ordination is a powerful method for analysing complex data sets but has been largely ignored in sequence analysis. This paper shows how to use principal coordinates analysis to find low-dimensional representations of distance matrices derived from aligned sets of sequences. The method takes a matrix of Euclidean distances between all pairs of sequence and finds a coordinate space where the distances are exactly preserved. The main problem is to find a measure of distance between aligned sequences that is Euclidean. The simplest distance function is the square root of the percentage difference (as measured by identities) between two sequences, where one ignores any positions in the alignment where there is a gap in any sequence. If one does not ignore positions with a gap, the distances cannot be guaranteed to be Euclidean but the deleterious effects are trivial. Two examples of using the method are shown. A set of 226 aligned globins were analysed and the resulting ordination very successfully represents the known patterns of relationship between the sequences. In the other example, a set of 610 aligned 5S rRNA sequences were analysed. Sequence ordinations complement phylogenetic analyses. They should not be viewed as a complete alternative.

Amino Acid Sequence↗

Prediction of the coding sequences of unidentified human genes. IV. The coding sequences of 40 new genes (KIAA0121-KIAA0160) deduced by analysis of cDNA clones from human cell line KG-1.

In this series of projects regarding the accumulation of sequence information of unidentified human genes, we newly deduced the sequences of 40 full-length cDNA clones of human cell line KG-1, and predicted the coding sequences of the corresponding genes, named KIAA0121 to 0160. The results of a computer search of public databases indicated that the sequences of 13 genes were unrelated to any reported genes, while the remaining 27 genes carried sequences which showed some similarities to known genes. Obvious unique sequences noted were as follows. A stretch of triplet repeats was contained in each of three genes: These were GAG(Glu) in KIAA0122 and KIAA0147, and TCC(Ser) in KIAA0150. A stretch of 10 amino acid-residues was repeated 21 times in KIAA0139, and a homologous sequence of 76-78 nucleotides was found repeated 6 times in the untranslated region of KIAA0125. Northern hybridization analysis demonstrated that 13 genes were expressed in a cell- or tissue-specific manner. Although a vast number of expressed sequence tags (ESTs) have been registered for comprehensive analysis of cDNA clones, our sequence data indicated that their distribution is very unbalanced: e.g. while no EST hit 7 genes, 85 ESTs fell in a single gene.

Amino Acid Sequence↗

Selective cloning and sequence analysis of the human L1 (LINE-1) sequences which transposed in the relatively recent past.

L1 (LINE-1), a long interspersed repetitive DNA family of mammalian genomes, is thought to be a sequence family derived from a retrotransposon-like element(s), but its actively transposable unit(s) has not been identified yet. We developed a novel method for selective isolation of the human L1 sequences which transposed in a relatively recent past and may have still retained a feature of the 'active L1' unit. From the inspection of the nucleotide sequences, we conjectured that the 'active L1' or 'nearly active L1' units should have a high content of the CpG dinucleotide sequence, a mutation hot spot sequence, and contain several sites for rare cutters such as BssH II and Nar I at their 5' terminal regions. Using these rare cutter sites as selection markers, the L1 sequences were isolated, which had the high content of CpG at the 5' terminal regions and over 90% homology to L1 transcripts found in a human teratocarcinoma cell line. These L1s were shown to be 'relatively new L1' units which had integrated into chromosomes within these several million years during evolution. From the sequence data of these L1s and L1 cDNA, a consensus sequence of the 5' terminal region of high CpG L1s were constructed. A region of the consensus sequence showed about 69% homology to the 5' terminal region of Drosophila jockey element.

Animals↗

HLA-DQB1 sequencing-based typing using newly identified conserved nucleotide sequences in introns 1 and 2.

Sequencing-based typing (SBT) human leukocyte antigen (HLA) class I and II genes should examine entire exon sequences where polymorphisms lie. Primers for the amplification of complete exons therefore anneal in introns and their design relies on accurate intron sequences being available. We decided to develop a SBT method for HLA-DQB1 using amplification primers which anneal in introns 1 and 2, yet the amount of intron sequence data previously available in databases was sparse. Therefore, we undertook a systematic sequencing of introns 1 and 2 using DNA from cell lines homozygous for DQB1. This study confirmed an earlier report that the non-coding regions of this gene are the most polymorphic seen in the human genome. Intron sequences within an allele group were largely identical, the exceptions being DQB1*0301 differing from other DQB1*03 allele groups and DQB1*0601 differing from all other DQB1*06 alleles. A retroviral Alu element, related to the AluYa5a2 subfamily, was identified uniquely inserted in intron 2 of DQB1*02 alleles. For the typing approach, six amplification primers were designed based on conserved allele group sequences covering all of the HLA DQB antigens, and two sequencing primers were also designed which anneal in intron 2. This method has proved to be very robust and has been used as part of a referral DNA sequencing service for a number of years.

Alleles↗

The Sinorhizobium meliloti insertion sequence (IS) elements ISRm102F34-1/ISRm7 and ISRm220-13-5 belong to a new family of insertion sequence elements.

The Sinorhizobium meliloti insertion sequence (IS) elements ISRm102F34-1 and ISRm220-13-5 are 1481 and 1550 base pairs (bp) in size, respectively. ISRm102F34-1 is bordered by 15 bp imperfect terminal inverted repeat sequences (two mismatches), whereas the terminal inverted repeat of ISRm220-13-5 has a length of 16 bp (two mismatches). Both insertion sequence elements generate a 6-bp target duplication upon transposition. The putative transposase enzymes of ISRm102F34-1 and ISRm220-13-5 consist of 449 or 448 amino acid residues with predicted molecular weights of 50.7 or 51.3 kDa and theoretical isoelectric points of 10.8 or 11.1, respectively. ISRm102F34-1 is identical in 98.9% of its nucleotide sequence to an apparently inactive copy of an insertion sequence element, designated ISRm7, which flanks the left-end of the nodule formation efficiency (nfe) region of plasmid pRmeGR4b of S. meliloti strain GR4. ISRm102F34-1 and ISRm220-13-5 are closely related since they show an overall identity of 57.0% at the nucleotide sequence level and of 47.3% at the deduced amino acid level of their putative transposases. Both insertion sequence elements displayed significant similarity to the Xanthomonas campestris ISXc6 and its homolog IS1478a. Since none of these insertion sequence elements could be allocated to existing families of insertion sequence elements, a new family is proposed. Analysis of the distribution of ISRm102F34-1/ISRm7 in various local S. meliloti populations sampled from Medicago sativa, Medicago sphaerocarpa and Melilotus alba host plants at different locations in Spain revealed its presence in 35% of the isolates with a copy number ranging from 1 to 5. Furthermore, ISRm102F34-1/ISRm7 homologs were identified in other rhizobial species.

Amino Acid Sequence↗

Sequence analysis of 16S rRNA from mycoplasmas by direct solid-phase DNA sequencing.

Automated solid-phase DNA sequencing was used for determination of partial 16S ribosomal DNA sequences of mycoplasmas. The sequence information was used to establish phylogenetic relationships of 11 different mycoplasmas whose 16S rRNA sequences had not been determined earlier. A biotinylated fragment corresponding to positions 344 to 939 in the Escherichia coli sequence was generated by PCR. The PCR product was immobilized onto streptavidin-coated paramagnetic beads, and direct sequencing was performed in both directions. One previously unclassified avian mycoplasma was found to belong to the Mycoplasma lipophilum cluster of the hominis group. Microheterogeneities were discovered in the rRNA operons of Mycoplasma mycoides subsp. mycoides (SC type), confirming the existence of two different rRNA operons. The 16S rRNA sequence of M. mycoides subsp. capri was identical to that of M. mycoides subsp. mycoides (type SC), except that no microheterogeneities were revealed. Furthermore, automated solid-phase DNA sequencing was used to identify a mycoplasmal contamination of a cell culture as Mycoplasma hyorhinis, which proved to be very difficult by conventional methods. The results suggest that the direct solid-phase DNA sequencing procedure is a powerful tool for identification of mycoplasmas and is also useful in taxonomic studies.

Base Sequence↗

DNA sequences of genes encoding Acinetobacter calcoaceticus protocatechuate 3,4-dioxygenase: evidence indicating shuffling of genes and of DNA sequences within genes during their evolutionary divergence.

The DNA sequence of a 2,391-base-pair HindIII restriction fragment of Acinetobacter calcoaceticus DNA containing the pcaCHG genes is reported. The DNA sequence reveals that A. calcoaceticus pca genes, encoding enzymes required for protocatechuate metabolism, are arranged in a single transcriptional unit, pcaEFDBCHG, whereas homologous genes are arranged differently in Pseudomonas putida. The pcaG and pcaH genes represent separate reading frames respectively encoding the alpha and beta subunits of protocatechuate 3,4-dioxygenase (EC 1.13.1.3); previously a single designation, pcaA, had been used to represent DNA encoding this enzyme. The alpha and beta protein subunits appear to share common ancestry with each other and with catechol 1,2-dioxygenases from A. calcoaceticus and P. putida. Marked conservation of amino acid sequence is observed in a region containing two histidyl residues and two tyrosyl residues that appear to ligate iron within each oxygenase. In some regions within the aligned oxygenase sequences, DNA sequences appear to be conserved at a level beyond the extent that might have been demanded by selection at the level of protein. In other regions, divergence of DNA sequences appears to have been achieved by substitution of DNA sequence from one genetic segment into another. The results are interpreted to be the consequence of sequence exchange by gene conversion between slipped strands of DNA during evolutionary divergence; mismatch repair between slipped strands may contribute to the maintenance of DNA sequence in divergent genes.

Acinetobacter↗

Structure and heterogeneity of the a sequences of human herpesvirus 6 strain variants U1102 and Z29 and identification of human telomeric repeat sequences at the genomic termini.

The unit-length genome of human herpesvirus 6 (HHV-6) consists of a single unique component (U) bounded by direct repeats DRL and DRR and forms head-to-tail concatemers during productive infection. cis-elements which mediate cleavage and packaging of progeny virions (a sequences) are found at the termini of all herpesvirus genomes. In HHV-6, DRL and DRR are identical and a sequences may therefore also occur at the U-DR junctions to give the arrangement aDRLa-U-aDRRa. We have sequenced the genomic termini, the U-DRR junction, and the DRR.DRL junction of HHV-6 strain variants U1102 and Z29. A (GGGTTA)n motif identical to the human telomeric repeat sequence (TRS) was found adjacent to, but did not form, the termini of both strain variants. The DRL terminus and U-DRR junction contained sequences closely related to that of the well-conserved herpesvirus packaging signal Cn-Gn-Nn-Gn (pac-1), followed by tandem arrays of TRSs separated by single copies of a hexanucleotide repeat. HHV-6 strain U1102 contained repeat sequences not found in HHV-6 Z29. In contrast, the DRR terminus of both variants contained a simple tandem array of TRSs and a close homolog of a herpesvirus pac-2 signal (GCn-Tn-GCn). The DRR.DRL junction was formed by simple head-to-tail linkage of the termini, yielding an intact cleavage signal, pac-2.x.pac-1, where x is the putative cleavage site. The left end of DR was the site of intrastrain size heterogeneity which mapped to the putative a sequences. These findings suggest that TRSs form part of the a sequence of HHV-6 and that the arrangement of a sequences in the genome can be represented as aDRLa-U-a-DRRa.

Base Sequence↗

Bacterial repetitive extragenic palindromic sequences are DNA targets for Insertion Sequence elements.

BACKGROUND: Mobile elements are involved in genomic rearrangements and virulence acquisition, and hence, are important elements in bacterial genome evolution. The insertion of some specific Insertion Sequences had been associated with repetitive extragenic palindromic (REP) elements. Considering that there are a sufficient number of available genomes with described REPs, and exploiting the advantage of the traceability of transposition events in genomes, we decided to exhaustively analyze the relationship between REP sequences and mobile elements. RESULTS: This global multigenome study highlights the importance of repetitive extragenic palindromic elements as target sequences for transposases. The study is based on the analysis of the DNA regions surrounding the 981 instances of Insertion Sequence elements with respect to the positioning of REP sequences in the 19 available annotated microbial genomes corresponding to species of bacteria with reported REP sequences. This analysis has allowed the detection of the specific insertion into REP sequences for ISPsy8 in Pseudomonas syringae DC3000, ISPa11 in P. aeruginosa PA01, ISPpu9 and ISPpu10 in P. putida KT2440, and ISRm22 and ISRm19 in Sinorhizobium meliloti 1021 genome. Preference for insertion in extragenic spaces with REP sequences has also been detected for ISPsy7 in P. syringae DC3000, ISRm5 in S. meliloti and ISNm1106 in Neisseria meningitidis MC58 and Z2491 genomes. Probably, the association with REP elements that we have detected analyzing genomes is only the tip of the iceberg, and this association could be even more frequent in natural isolates. CONCLUSION: Our findings characterize REP elements as hot spots for transposition and reinforce the relationship between REP sequences and genomic plasticity mediated by mobile elements. In addition, this study defines a subset of REP-recognizer transposases with high target selectivity that can be useful in the development of new tools for genome manipulation.

Bacteria↗

Comparison of ZP3 protein sequences among vertebrate species: to obtain a consensus sequence for immunocontraception.

The deduced ZP3 amino acid (aa) sequences of 13 vertebrate species namely mouse, hamster, rabbit, pig, porcine, cow, dog, cat, human, bonnet, marmoset, carp, and frog were compared using the PILEUP and PRETTY alignment programs (GCG, Wisconsin, USA). The published aa sequences obtained from 13 vertebrate species indicated the overall evolutionarily conservation in the N-terminus, central region, and C-terminus of the ZP3 polypeptide. More variations of ZP3 polypeptide sequences were seen in the alignments of carp and frog from the 11 mammalian species making the leader sequence more prominent. The canonical furin proteolytic processing signal at the C-terminus was found in all the ZP3 polypeptide sequences except of carp and frog. In the central region, the ZP3 deduced aa sequences of all the 13 vertebrate species aligned well, and six relatively conserved sequences were found. There are 11 conserved cysteine residues in the central region across all species including carp and frog, indicating that these residues have longer evolutionary history. The ZP3 aa sequence similarities were examined using the GAP program (GCG). The highest aa similarities are observed between the members of the same order within the class mammalia, and also (95.4%) between pig (ungulata) and rabbit (lagomorpha). The deduced ZP3 aa sequences per se may not be enough to build a phylogenetic tree.

Amino Acid Sequence↗