Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Leishmania major: molecular cloning, sequencing, and expression of the heat shock protein 60 gene reveals unique carboxy terminal peptide sequences.

Heat shock proteins (HSP) in the size range of M(r) 60,000 are major targets of the immune response in vivo. The leishmania heat-inducible proteins of M(r) 65-67,000 are expressed at relatively high levels in infected macrophages (Infection and Immunity 1993, 61, 3265-3272) and may be important targets of the host response. To facilitate further studies concerned with these proteins, the HSP60 gene of Leishmania major was cloned, sequenced, and expressed. A lambdaEMBL-3 L. major genomic library was screened with a PCR-generated DNA probe derived from a highly conserved region of the leishmania HSP60 gene. A single clone that hybridized strongly was characterized. Sequence analysis revealed an open reading frame of 1770 bp encoding a putative polypeptide of 589 amino acids with a predicted size of M(r) 64,790 and with the highest degree of amino acid sequence similarity (56%) to HSP60 from Trypanosoma cruzi. Less extensive amino acid sequence similarity (48%) was observed between that leishmania HSP60 and the corresponding human protein. Notably, significant regions of sequence dissimilarity between the leishmania and human proteins were identified principally within the carboxy-terminal regions of the proteins. The entire coding region of the leishmania HSP60 gene was subcloned into the pET-3a vector and expressed in Escherichia coli. Purified recombinant protein was used to examine sera from patients with tegumentary leishmaniasis from Colombia for the presence of antibodies to HSP60. Unlike sera from healthy, uninfected controls, sera from patients reacted strongly with recombinant leishmania HSP60. This recognition had specificity in that these same sera showed little or no reactivity with either recombinant mycobacterial HSP65 or recombinant human HSP60. These findings indicate that patients with tegumentary forms of leishmaniasis have humoral responses to leishmania HSP60. Further studies of this protein will clarify its importance as a target of the immune response and as a potential antigen for serodiagnosis.

Amino Acid Sequence↗

Complete nucleotide sequence of HTLV-II isolate NRA: comparison of envelope sequence variation of HTLV-II isolates from U.S. blood donors and U.S. and Italian i.v. drug users.

The entire nucleotide sequence of human T-cell lymphotropic virus type II (HTLV-II) from a previously described isolate of patient NRA (HTLV-IINRA) was determined. Clones encoding the 5' LTR and gag, pol, env and tax/rex open reading frames were subcloned and sequenced on both strands. The provirus consisted of 8957 nucleotides and showed 95.2% homology with the HTLV-IIMo prototype at the nucleotide level. Less than 5% amino acid variation between HTLV-IINRA and HTLV-IIMo was observed for coding regions. Although isolate HTLV-IINRA had an additional 25 amino acids at the 3' end of tax/rex, this region was 96% homologous with the 5' end of HTLV-IIMo 3' LTR. To further investigate HTLV-II variability, a portion of the env gp46 gene derived from 9 HTLV-II infected persons was amplified by polymerase chain reaction and sequenced. Sequence was obtained for 320 nucleotides corresponding to HTLV-IIMo positions 5291 to 5610. Isolates similar to the HTLV-IIMo and HTLV-IINRA prototypes were identified, and sequences were highly conserved.

Amino Acid Sequence↗

Repetitive sequence-mediated rearrangements in Chlorella ellipsoidea chloroplast DNA: completion of nucleotide sequence of the large inverted repeat.

A 3454 base pair (bp) sequence of the large inverted repeat (IR) of chloroplast DNA (cpDNA) from the unicellular green alga Chlorella ellipsoidea has been determined. The sequence includes: (1) the boundaries between the IR and the large single copy (LSC) and the small single copy (SSC) regions, (2) the gene for psbA and (3) an approximately 1.0 kbp region between psbA and the rRNA genes which contains a variety of short dispersed repeats. The total size of the Chlorella IR was determined to be 15243 bp. The junction between the IR and the small single copy region is located close to the putative promoter of the rRNA operon (906 bp upstream of the -35 sequence on each IR). The junction between the IR and the large single copy region is also just upstream of the putative psbA promoter, 218 bp upstream from the ATG initiation codon. A few sets of unique sequences were found repeatedly around both junctions. Some of the sequences flanking the IR-LSC junction suggest a unidirectional and serial expansion of the IR within the genome. The psbA gene is located close to the LSC-side junction and codes for a protein of 352 amino acid residues. A highly conserved C-terminal Gly is absent Unlike the psbA of Chlamydomonas species, which contains 2-4 large introns, the gene of Chlorella has no introns.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Sequence and characterization of an insertion sequence, IS711, from Brucella ovis.

The nucleotide (nt) sequence of a previously discovered insertion in Brucella ovis was determined and found to have the hallmarks of an insertion sequence (IS). The element, designated IS711, of 842 bp, is similar in G + C content to that of the Brucella genome and is bounded by 20-bp imperfect inverted repeats (IR). The element appears to duplicate the nt TA of a consensus target site, YTAR (R, purines; Y, pyrimidines). When the complete nt sequence of four elements and 300 bp of the 3' ends of five other elements were compared to IS711 and to each other, minor nt sequence variations were found amongst most of them. Similar to several other transposable elements, IS711 has overlapping ORFs rather than one long ORF extending the length of the element. Even though only ten B. ovis IS711 elements were characterized, in three cases we found these elements flanked by either identical or similar nt sequences. This suggests that some target sites are hot spots for insertion and that some of the elements may be duplicated by mechanisms other than transposition. No DNA or protein database entries had an obvious resemblance to either IS711 or its deduced gene products.

Amino Acid Sequence↗

Primary structure of the human elafin precursor preproelafin deduced from the nucleotide sequence of its gene and the presence of unique repetitive sequences in the prosegment.

The human elafin gene was cloned and its entire nucleotide sequence was determined to deduce the amino acid sequence for the precursor of elafin, an elastase-specific inhibitor. The gene spans approximately 1.7 kb and is divided into 3 exons. The gene product preproelafin consists of 117 amino acids: the initiator Met, a putative 25-amino acid signal peptide, a pro-sequence of about 34 amino acids, and the C-terminal 57 amino acids for mature elafin. Possible covalent clotting of the prosegment and its physiological significance have been pointed out based on a remarkable sequence similarity between the pro-sequence and the guinea pig seminal clotting protein SVP-1.

Amino Acid Sequence↗

Sequencing of HLA class II genes based on the conserved diversity of the non-coding regions: sequencing based typing of HLA-DRB genes.

In this paper, we present a novel sequencing based typing strategy for the HLA-DRB1, 3, 4 and 5 loci. The new approach is based on a group-specific amplification from intron 1 to intron 2 according to the serologically-defined antigens. For this purpose, we have determined the 3' 500 bp-fragment of intron 1 and the 5' 340 bp-fragment of intron 2 of all serological antigens and their most frequent subtypes. We discovered a remarkably conserved diversity characterized by lineage-specific sequence motifs. This lineage-specificity of non-coding motifs in the 1st and 2nd intron offered the possibility to establish a clear serology-related amplification strategy. The method allows the complete analysis of the 2nd exon and the definition of the cis/trans linkage of sequence motifs by intron-mediated polymerase chain reaction (PCR)-based separation of the haplotypes in nearly all serologically heterozygous samples. In particular, the non-coding variabilities between the DR52-associated DRB1 groups made their independent amplification possible. Thus, compared to the standard procedures using exon-based amplification primers, the groups DR3, DR12, some DR13 alleles (1301, 1302) and the DR14 group could be amplified by specific primer mixes. The DR8 could be amplified with an individual primer mix not co-amplifying the DR12. The DR11 and DR13 did not show any individual motif in intron 1 or intron 2. In order to achieve a separate amplification, they had to be amplified by multispecific primer mixes (DR3/11/13/14; DR3/11/13 or DR11/13/14) excluding the other haplotype. Thus, exclusively the alleles in rare DR11,13 heterozygosities without a DRB1*1301 or 1302 could not be amplified separately. Fourteen primer mixes are used to amplify the specificities DR1-14, and 6 primer mixes for the specificities DR51-53. The sequence homology of the 3' end of intron 1 facilitated the application of only three different sequencing primers for all DRB alleles.

Alleles↗

A conserved sequence motif within the exceptionally diverse telomeric sequences of budding yeasts.

Telomeric DNA sequences have generally been found to be remarkably conserved in evolution, typically consisting of repeated, very short sequence units containing clusters of G residues. Recently however the telomeric DNA of the asexual yeast Candida albicans was shown to consist of much longer repeat units. Here we report the identification of seven additional telomeric sequences from sexual and asexual budding yeast species. The telomeric repeat units from this group of relatively closely related species show more phylogenetic diversity in length (8-25 bp), sequence, and composition than has been seen previously throughout a wide phylogenetic range of other eukaryotes. We also show that certain strains of the asexual diploid species Candida tropicalis have two forms of telomeric repeats, which appear to differ by a single base pair. Despite their great diversity, the telomeric repeat units of C. albicans, Saccharomyces cerevisiae, and all of the species we have examined in this report share a conserved approximately 6-bp motif of T and G residues resembling more typical telomeric sequences.

Base Sequence↗

Bacterial interspersed mosaic elements (BIMEs) are a major source of sequence polymorphism in Escherichia coli intergenic regions including specific associations with a new insertion sequence.

A significant fraction of Escherichia coli intergenic DNA sequences is composed of two families of repeated bacterial interspersed mosaic elements (BIME-1 and BIME-2). In this study, we determined the sequence organization of six intergenic regions in 51 E. coli and Shigella natural isolates. Each region contains a BIME in E. coli K-12. We found that multiple sequence variations are located within or near these BIMEs in the different bacteria. Events included excisions of a whole BIME-1, expansion/deletion within a BIME-2 and insertions of non-BIME sequences like the boxC repeat or a new IS element, named IS 1397. Remarkably, 14 out of IS 1397 integration sites correspond to a BIME sequence, strongly suggesting that this IS element is specifically associated with BIMEs, and thus inserts only in extragenic regions. Unlike BIMEs, IS 1397 is not detected in all E. coli isolates. Possible relationships between the presence of this IS element and the evolution of BIMEs are discussed.

Amino Acid Sequence↗

Conservation of sequence in recombination signal sequence spacers.

The variable domains of immunoglobulins and T cell receptors are assembled through the somatic, site specific recombination of multiple germline segments (V, D, and J segments) or V(D)J rearrangement. The recombination signal sequence (RSS) is necessary and sufficient for cell type specific targeting of the V(D)J rearrangement machinery to these germline segments. Previously, the RSS has been described as possessing both a conserved heptamer and a conserved nonamer motif. The heptamer and nonamer motifs are separated by a 'spacer' that was not thought to possess significant sequence conservation, however the length of the spacer could be either 12 +/- 1 bp or 23 +/- 1 bp long. In this report we have assembled and analyzed an extensive data base of published RSS. We have derived, through extensive consensus comparison, a more detailed description of the RSS than has previously been reported. Our analysis indicates that RSS spacers possess significant conservation of sequence, and that the conserved sequence in 12 bp spacers is similar to the conserved sequence in the first half of 23 bp spacers.

Animals↗

Short sequences define genetic lineages: phylogenetic analysis of group A rotaviruses based on partial sequences of genome segments 4 and 9.

Genetic diversity in strains of human group A rotaviruses was analysed by phylogenetic methods. The study material comprised 109 serotype G1 or G4 rotavirus samples isolated in Finland during 1986-1990. Parts of the coding regions of rotaviral genome segments 4 and 9, which encode proteins with serotype specificity, the spike protein VP4 (P serotype) and the outer capsid protein VP7 (G serotype), respectively, were sequenced. As determined by analysis of segment 4 sequences all G1 strains and all except one G4 strain showed P[8] specificity, the one being of P[6] specificity. The G1P[8] strains could be further differentiated into four groups based on segment 9 sequences, while G4P[8] strains formed only one group. Type P[8] (G1P[8] and G4P[8]) strains formed two main groups based on segment 4 sequences, suggesting free segregation of segment 4 between these G strains. Most global G1, G4 and P[8] strains in GenBank/EMBL originating from the 1970s to the present co-clustered with these groups, suggesting that the groups exist as relatively stable lineages. No linear accumulation of nucleotide substitutions was detected in strains of one serotype during the study period. Also, the deduced amino acids of the antigenic regions A, B and C of VP7 were nearly conserved within the phylogenetic lineages. Interestingly, only short amino acid sequences were necessary to divide the e-types correctly into phylogenetic lineages. These amino acid signature motifs were located in aa 29-68 of VP7 and aa 121-135 of VP4 of the G1 and P[8] lineages, respectively.

Amino Acid Sequence↗

Sequence-specific DNA binding of the proto-oncoprotein ets-1 defines a transcriptional activator sequence within the long terminal repeat of the Moloney murine sarcoma virus.

The ets proto-oncogene family is a group of sequence-related genes whose normal cellular function is unknown. In a study of cellular proteins involved in the transcriptional regulation of murine retroviruses in T lymphocytes, we have discovered that a member of the ets gene family encodes a sequence-specific DNA-binding protein. A mouse ets-1 cDNA clone was obtained by screening a mouse thymus cDNA expression library with a double-stranded oligonucleotide probe representing 20 bp of the Moloney murine sarcoma virus (MSV) long terminal repeat (LTR). The cDNA sequence has an 813-bp open reading frame (ORF) whose predicted amino acid sequence is 97.6% identical to the 272 carboxy-terminal amino acids of the human ets-1 protein. The ORF was expressed in bacteria, and the 30-kD protein product was shown to bind DNA in a sequence-specific manner by mobility-shift assays, Southwestern blot analysis, and methylation interference. A mutant LTR containing four base pair substitutions in the ets-1 binding site was constructed and was shown to have reduced binding in vitro. Transcriptional efficiency of the MSV LTR promoter containing this disrupted ets-1 binding site was compared to the activity of a wild-type promoter in mouse T lymphocytes in culture, and 15- to 20-fold reduction in expression of a reporter gene was observed. We propose that ets-1 functions as a transcriptional activator of mammalian type-C retroviruses and speculate that ets-related genes constitute a new group of eukaryotic DNA-binding proteins.

Amino Acid Sequence↗

The complete amino acid sequence of the Clostridium botulinum type-E neurotoxin, derived by nucleotide-sequence analysis of the encoding gene.

The entire structural gene of the Clostridium botulinum NCTC 11219 type-E neurotoxin (BoNT/E) has been cloned as five overlapping DNA fragments, generated by polymerase chain reaction (PCR). Analysis of triplicate clones of each fragment, derived from three independent PCR, has allowed the derivation of the entire nucleotide sequence of the BoNT/E gene. Translation of the sequence has shown BoNT/E to consist of 1252 amino acids and, as such, represents the smallest BoNT characterised to date. The light chain of the toxin exhibits the highest level of sequence similarity to tetanus toxin (TeTx, 40%). The light chains of BoNT/A and BoNT/D share 33% similarity with BoNT/E, while BoNT/C exhibits 32% similarity. In contrast, the TeTx heavy chain exhibits the lowest degree of similarity (35%) with BoNT/E, with the BoNT heavy chains sharing 46%, 36% and 37%, for neurotoxin types A, C and D, respectively. Comparisons with partial amino acid sequences of the light chain of BoNT/E from C. botulinum strain Beluga and that from the strains Mashike, Iwanai and Otaru, indicate single amino acid differences in each case. Alignment of all characterised neurotoxin sequences (BoNT/A, BoNT/C, BoNT/D, BoNT/E and TeTx) shows them to be composed of highly conserved amino acid domains interspersed with amino acid tracts exhibiting little overall similarity. The most divergent region corresponds to the extreme COOH-terminus of each toxin, which may reflect differences in specificity of binding to neurone acceptor sites.

Amino Acid Sequence↗

Isolation, sequence analysis and characterization of cDNA clones coding for the C chain of mouse C1q. Sequence similarity of complement subcomponent C1q, collagen type VIII and type X and precerebellin.

A mouse macrophage lambda gt11 cDNA library was screened using a genomic DNA clone coding for the C-chain gene of human C1q. Approximately 600,000 recombinant phage plaques were hybridized with peroxidase-labeled human C-chain probe and detected by enhanced chemiluminescence. Five positive clones were obtained. The size of the full-length cDNA is 1019 bp. The sequence identity of the nucleotide sequence with human C1q C chain is 79%, the identity of the deduced amino acid sequences is 73%. The mouse C1q C chain exhibits the same structural features as the human C chain, e.g. conservation of the cysteine residues. Like the mouse A chain, the mouse C chain has an RGD sequence that may be recognized by receptors of the integrin family. No RGD sequences have been found in any of the human C1q chains. The size of the C-chain mRNA (1.2 kb) and its tissue distribution (macrophages being the cell type with the highest mRNA concentration) are identical to the mRNA of the mouse A and B chains. Alignment of human and mouse C1q A, B and C chains exhibits two blocks of highly conserved residues within the C-terminal globular regions. Three other proteins, collagen type VIII and type X and precerebellin share this similarity with C1q, indicating the structural and probably functional importance of these regions within the non-collagenous domains of the molecules.

Amino Acid Sequence↗

Complete nucleotide sequence of the Actinomyces viscosus T14V sialidase gene: presence of a conserved repeating sequence among strains of Actinomyces spp.

The nucleotide sequence of the Actinomyces viscosus T14V sialidase gene (nanH) and flanking regions was determined. An open reading frame of 2,703 nucleotides that encodes a predominately hydrophobic protein of 901 amino acids (M(r), 92,871) was identified. The amino acid sequence at the amino terminus of the predicted protein exhibited properties characteristic of a typical leader peptide. Five 12-amino-acid units that shared between 33 and 67% sequence identity were noted within the central domain of the protein. Each unit contained the sequence Ser-X-Asp-X-Gly-X-Thr-Trp, which is conserved among other bacterial and trypanosoma sp. sialidases. Thus, the A. viscosus T14V nanH gene and the other prokaryotic and eukaryotic sialidase genes evolved from a common ancestor. Southern hybridization analyses under conditions of high stringency revealed the existence of DNA sequences homologous to A. viscosus T14V nanH in the genomes of 18 strains of five Actinomyces species that expressed various levels of sialidase activity. The data demonstrate that the sialidase genes from divergent groups of Actinomyces spp. are highly conserved.

Actinomycetaceae↗

Human endogenous retroviruslike genome with type C pol sequences and gag sequences related to human T-cell lymphotropic viruses.

We have cloned several prototypic members of the family of human endogenous retroviruslike elements having a histidine tRNA primer-binding site (RTVL-H) and have determined the nucleotide sequence of one of these clones (RTVL-H2). The RTVL-H2 sequence is 5,813 nucleotides long, with long terminal repeats of 450 nucleotides. Although this particular sequence contains no long open reading frames, computer searches have revealed several segments of amino acid homology with known retroviral gene products. In the gag region of RTVL-H2, there is a segment with significant homology to a region of the gag protein p30 of type C baboon endogenous virus. In the pol region of RTVL-H2, three segments similar to the Moloney leukemia virus (MLV) pol polyprotein were detected. These correspond to parts of the protease, reverse transcriptase, and endonuclease domains of the MLV pol gene. Interestingly, the last two pol domains are equidistant in RTVL-H2 and the type C murine retroviruslike DNA sequence (MuRRS), both having deletions of equal sizes relative to the MLV pol gene. One other segment similar to a retroviral gene product was identified in the RTVL-H2 gag region. This segment has 55 to 60% amino acid homology to a 50-amino-acid region of the gag nucleic acid-binding proteins encoded by human T-cell lymphotropic viruses types I and II and bovine leukemia virus. Thus, the RTVL-H2 genome harbors sequences related to evolutionarily distant retroviruses.

Amino Acid Sequence↗

Cloning, sequencing, and expression of Rhodococcus L-phenylalanine dehydrogenase. Sequence comparisons to amino-acid dehydrogenases.

L-Phenylalanine dehydrogenase catalyzes the NAD(+)-dependent, reversible, oxidative deamination of L-phenylalanine to form ammonia, phenyl pyruvate, and NADH. The enzyme has been purified to homogeneity from Rhodococcus sp. M4, and a partial amino acid sequence was obtained. A cosmid library of Rhodococcus sp. M4 genomic DNA was prepared and used to isolate a 2.5-kilobase PstI fragment that contained the pdh gene. The open reading frame of 1068 nucleotides encodes a polypeptide of 356 amino acids, portions of which match the amino acid sequence determined for the purified enzyme. Expression of the Rhodococcus pdh gene in Escherichia coli, which does not contain a phenylalanine dehydrogenase activity, yields a soluble enzyme exhibiting phenylalanine dehydrogenase activity. Both the enzyme purified from Rhodococcus and the enzyme expressed in E. coli are post-translationally modified by removal of the amino-terminal methionine. The overall amino acid sequence is homologous to previously reported sequences of leucine and phenylalanine dehydrogenases as well as several glutamate dehydrogenases. The amino-terminal portion of the enzyme contains residues involved in L-amino acid binding and catalysis, while the carboxyl-terminal portion contains the presumptive dinucleotide-binding domain. A detailed sequence comparison of Rhodococcus phenylalanine dehydrogenase with leucine, phenylalanine, and glutamate dehydrogenases suggests residues involved in general amino acid binding and others that provide for amino acid discrimination.

Amino Acid Oxidoreductases↗

Deletion analysis of minimal sequence requirements for autonomous replication of ors8, a monkey early-replicating DNA sequence.

We have generated a panel of deletion mutants of ors8 (483 bp), a mammalian autonomously replicating DNA sequence, previously isolated by extrusion of nascent monkey (CV-1) DNA from replication bubbles active at the onset of S phase. The deletion mutants were tested for replication function by the DpnI resistance assay, in vivo, after transfection into HeLa cells, and in vitro. An internal fragment of 186-bp that is required for autonomous replication function of ors8 was identified. This fragment, when subcloned into pBR322 and similarly tested, was capable of autonomous replication in vivo and in vitro. The 186-bp fragment contains several repeated sequence motifs, such as the ATTA and ATTTAT motifs, occurring three and five times, respectively, the sequences TAGG and TAGA, occurring three and seven times, respectively, two 5'-ATT-3' repeats, a 44-bp imperfect inverted repeat (IR) sequence, and an imperfect consensus binding element for the transcription factor Oct-1. A measurable sequence-directed DNA curvature was also detected, coinciding with the AT-rich regions of the 186-bp fragment.

Animals↗

Oligopeptide biases in protein sequences and their use in predicting protein coding regions in nucleotide sequences.

We have examined oligopeptides with lengths ranging from 2 to 11 residues in protein sequences that show no obvious evolutionary relationship. All sequences in the Protein Identification Resource database were carefully classified by sensitive homology searches into superfamilies to obtain unbiased oligopeptide counts. The results, contrary to previous studies, show clear prejudices in protein sequences. The oligopeptide preferences were used to help decide the significance of sequence homologies and to improve the more general methods for detecting protein coding regions within nucleotide sequences.

Amino Acid Sequence↗