Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

DNA sequence determination by hybridization: a strategy for efficient large-scale sequencing.

The concept of sequencing by hybridization (SBH) makes use of an array of all possible n-nucleotide oligomers (n-mers) to identify n-mers present in an unknown DNA sequence. Computational approaches can then be used to assemble the complete sequence. As a validation of this concept, the sequences of three DNA fragments, 343 base pairs in length, were determined with octamer oligonucleotides. Possible applications of SBH include physical mapping (ordering) of overlapping DNA clones, sequence checking, DNA fingerprinting comparisons of normal and disease-causing genes, and the identification of DNA fragments with particular sequence motifs in complementary DNA and genomic libraries. The SBH techniques may accelerate the mapping and sequencing phases of the human genome project.

Animals↗

Proteus mirabilis fimbriae: N-terminal amino acid sequence of a major fimbrial subunit and nucleotide sequences of the genes from two strains.

Proteus mirabilis, a common cause of urinary tract infection in hospitalized and catheterized patients, produces mannose-resistant/klebsiella-like (MR/K) and mannose-resistant/proteus-like (MR/P) hemagglutinins. The gene encoding the major structural subunit of a fimbria, possibly MR/K, was identified in two strains. A degenerate oligonucleotide probe based on the N terminus of the Proteus uroepithelial cell adhesin and antiserum raised against the denatured polypeptide were used to screen a cosmid gene bank of strain HU1069. A cosmid clone that reacted with the probe and antiserum was identified, and a fimbria-like open reading frame was determined by nucleotide sequencing. The predicted N-terminal amino acid sequence of the processed polypeptide, ENETPAPKVSSTKGEIQLKG (residues 23 to 42), did not match the uroepithelial cell adhesin N terminus but, rather, matched exactly the N-terminal amino acid sequence of a polypeptide with an apparent molecular size of 19.5 kDa isolated by sodium dodecyl sulfate-polyacrylamide gel electrophoresis of a fimbrial preparation from strain HI4320 expressing MR/K hemagglutinin. By using an oligonucleotide from the HU1069 open reading frame, the fimbrial gene was isolated and sequenced from a cosmid gene bank clone of strain HI4320. A 552-bp open reading frame predicts a 184-amino-acid polypeptide including a 22-amino-acid hydrophobic leader sequence. The unprocessed polypeptide is predicted to be 18,921 Da; the processed polypeptide is predicted to be 16,749 Da. The predicted amino acid sequence of the polypeptide encoded by the gene, designated pmfA, displayed 36% exact matches with the mannose-resistant fimbrial subunit encoded by smfA of Serratia marcescens but only 15% exact matches with the predicted sequence encoded by mrkA of Klebsiella pneumoniae.

Amino Acid Sequence↗

Upstream induction sequence, the cis-acting element required for response to the allantoin pathway inducer and enhancement of operation of the nitrogen-regulated upstream activation sequence in Saccharomyces cerevisiae.

Expression of the DAL2, DAL4, DAL7, DUR1,2, and DUR3 genes in Saccharomyces cerevisiae is induced by the presence of allophanate, the last intermediate of the allantoin degradative pathway. Analysis of the DAL7 5'-flanking region identified an element, designated the DAL upstream induction sequence (DAL UIS), required for response to inducer. The operation of this cis-acting element requires functional DAL81 and DAL82 gene products. We determined the DAL UIS structure by using saturation mutagenesis. A specific dodecanucleotide sequence is the minimum required for response of reporter gene transcription to inducer. There are two copies of the sequence in the 5'-flanking region of the DAL7 gene. There are one or more copies of the sequence upstream of each allantoin pathway gene that responds to inducer. The sequence is also found 5' of the allophanate-inducible CAR2 gene as well. No such sequences were detected upstream of allantoin pathway genes that do not respond to the presence of inducer. We also demonstrated that the presence of a UIS element adjacent to the nitrogen-regulated upstream activation sequence significantly enhances its operation.

Allantoin↗

Multilocus short sequence repeat sequencing approach for differentiating among Mycobacterium avium subsp. paratuberculosis strains.

We describe a multilocus short sequence repeat (MLSSR) sequencing approach for the genotyping of Mycobacterium avium subsp. paratuberculosis (M. paratuberculosis) strains. Preliminary analysis identified 185 mono-, di-, and trinucleotide repeat sequences dispersed throughout the M. paratuberculosis genome, of which 78 were perfect repeats. Comparative nucleotide sequencing of the 78 loci of six M. paratuberculosis isolates from different host species and geographic locations identified a subset of 11 polymorphic short sequence repeats (SSRs), with an average of 3.2 alleles per locus. Comparative sequencing of these 11 loci was used to genotype a collection of 33 M. paratuberculosis isolates representing different multiplex PCR for IS900 loci (MPIL) or amplified fragment length polymorphism (AFLP) types. The analysis differentiated the 33 M. paratuberculosis isolates into 20 distinct MLSSR types, consistent with geographic and epidemiologic correlates and with an index of discrimination of 0.96. MLSSR analysis was also clearly able to distinguish between sheep and cattle isolates of M. paratuberculosis and easily and reproducibly differentiated strains representing the predominant MPIL genotype (genotype A18) and AFLP genotypes (genotypes Z1 and Z2) of M. paratuberculosis described previously. Taken together, the results of our studies suggest that MLSSR sequencing enables facile and reproducible high-resolution subtyping of M. paratuberculosis isolates for molecular epidemiologic and population genetic analyses.

Animals↗

Matrix genes of measles virus and canine distemper virus: cloning, nucleotide sequences, and deduced amino acid sequences.

The nucleotide sequences encoding the matrix (M) proteins of measles virus (MV) and canine distemper virus (CDV) were determined from cDNA clones containing these genes in their entirety. In both cases, single open reading frames specifying basic proteins of 335 amino acid residues were predicted from the nucleotide sequences. Both viral messages were composed of approximately 1,450 nucleotides and contained 400 nucleotides of presumptive noncoding sequences at their respective 3' ends. MV and CDV M-protein-coding regions were 67% homologous at the nucleotide level and 76% homologous at the amino acid level. Only chance homology was observed in the 400-nucleotide trailer sequences. Comparisons of the M protein sequences of MV and CDV with the sequence reported for Sendai virus (B. M. Blumberg, K. Rose, M. G. Simona, L. Roux, C. Giorgi, and D. Kolakofsky, J. Virol. 52:656-663; Y. Hidaka, T. Kanda, K. Iwasaki, A. Nomoto, T. Shioda, and H. Shibuta, Nucleic Acids Res. 12:7965-7973) indicated the greatest homology among these M proteins in the carboxyterminal third of the molecule. Secondary-structure analyses of this shared region indicated a structurally conserved, hydrophobic sequence which possibly interacted with the lipid bilayer.

Amino Acid Sequence↗

Nucleotide sequence of the tail sheath gene of bacteriophage T4 and amino acid sequence of its product.

The nucleotide sequence of gene 18 of bacteriophage T4 was determined by the Maxam-Gilbert method, partially aided by the dideoxy method. To confirm the deduced amino acid sequence of the tail sheath protein (gp18) that is encoded by gene 18, gp18 was extensively digested by trypsin or lysyl endopeptidase and subjected to reverse-phase high-performance liquid chromatography. Approximately 40 peptides, which cover 88% of the primary structure, were fractionated, the amino acid compositions were determined, and the corresponding sequences in DNA were identified. Furthermore, the amino acid sequences of 10 of the 40 peptides were determined by a gas phase protein sequencer, including N- and C-terminal sequences. Thus, the complete amino acid sequence of gp18, which consists of 658 amino acids with a molecular weight of 71,160, was determined.

Amino Acid Sequence↗

Mutations in signal sequence cleavage domain of preproparathyroid hormone alter protein translocation, signal sequence cleavage, and membrane-binding properties.

Signal sequences, known to mediate the targeting of nascent secreted proteins to membranes, share common structural domains: a positively charged amino-terminus, a hydrophobic core, and a signal cleavage domain. Mutations have been introduced into the cDNA encoding the signal sequence of the mammalian protein preproparathyroid hormone to analyze the roles played by the signal cleavage domain in secretion. Two mutant genes were constructed missing the entire six-residue propeptide sequence and several residues of the signal cleavage domain. The effects of these mutations on signal function were assessed after expression in clonal cell lines and in a transcription-linked translation system. Alterations in the signal cleavage domain resulted in reduced translocation and signal cleavage. Furthermore, in one mutant, the removal of the signal cleavage domain converted the signal into a membrane anchor sequence. The nonhydrophobic sequences at the end of the signal sequence thus crucially affect the translocation, cleavage, and membrane-binding properties of signal sequences.

Amino Acid Sequence↗

Identification of individual barley chromosomes based on repetitive sequences: conservative distribution of Afa-family repetitive sequences on the chromosomes of barley and wheat.

The Afa-family repetitive sequences were isolated from barley (Hordeum vulgare, 2n = 14) and cloned as pHvA14. This sequence distinguished each barely chromosome by in situ hybridization. Double color fluorescence in situ hybridization using pHvA14 and 5S rDNA or HvRT-family sequence (subtelomeric sequence of barley) allocated individual barley chromosomes showing a specific pattern of pHvA14 to chromosome 1H to 7H. As the case of the D genome chromosomes of Aegilops squarrosa and common wheat (Triticum aestivum) hybridized by its Afa-family sequences, the signals of pHvA14 in barley chromosomes tended to appear in the distal regions that do not carry many chromosome band markers. In the telomeric regions these signals always placed in more proximal portions than those of HvRT-family. Based on the distribution patterns of Afa-family sequences in the chromosomes of barley and D genome chromosomes of wheat, we discuss a possible mechanism of amplification of the repetitive sequences during the evolution of Triticeae. In addition, we show here that HvRT-family also could be used to distinguish individual barley chromosomes from the patterns of in situ hybridization.

Biological Evolution↗

The complete amino acid sequence of the Clostridium botulinum type D neurotoxin, deduced by nucleotide sequence analysis of the encoding phage d-16 phi genome.

The complete nucleotide sequence of Clostridium botulinum type D strain CB16 neurotoxin was determined and the deduced amino acid sequence is reported here for the first time. The structure and function of botulinum type D neurotoxin is discussed from a molecular biological viewpoint. DNA was extracted from toxin-converting phage d-16 phi of C. botulinum type D strain CB16, and a fragment (about 10 kbp) coding for the neurotoxin was cloned into Escherichia coli using lambda gt11. A 21-mer oligonucleotide which corresponds to Phe7 to Val13 of the partial amino acid sequence near the N-terminus of the type D neurotoxin was synthesized and used as a probe to identify the gene encoding type D neurotoxin. The nucleotide sequence contained a single open reading frame coding for 1,275 amino acids (molecular weight of 146,785) and the deduced amino acid sequence corresponded exactly to the partial amino acid sequences determined by direct microsequencing of the neurotoxin fragments. In the dichain molecule of the neurotoxin, Thr2 and Asn443 formed the N-termini of the light chain (M.W. 50,410) and heavy chain (M.W. 96,394) respectively, and these two chains were linked with a disulfide bond between Cys437 on the light chain and Cys450 on the heavy chain. The nucleotide sequence of the D-CB16 neurotoxin differed from that previously reported for type D neurotoxin by three nucleotides.

Amino Acid Sequence↗

Generating unigene collections of expressed sequence tag sequences for use in mass spectrometry identification.

Expressed sequence tag sequences remain the largest resource of DNA sequence for most organisms despite recent advances in genome sequencing. These sequences are short, fragmented versions of the expressed genes. By DNA sequence assembly, the fragments can be assembled into contiguous DNA sequences that are better suited for protein identification by mass spectrometry.

Cluster Analysis↗

Repeated immunogenic amino acid sequences of Plasmodium species share sequence homologies with proteins from humans and human viruses.

The use of recombinant peptides based upon the repeated amino acid sequences of Plasmodium has been proposed for malaria vaccines. By reducing homologies of such peptide vaccines to host proteins, the possibility of autoimmune complications may be reduced, and the effective immune response may be enhanced. The Wilbur and Lipman Wordsearch algorithm was used to identify homologous amino acid sequences between tandemly repeated Plasmodium amino acid sequences and the human and human viral sequences compiled in the National Biomedical Research Foundation database. Six published repetitive immunogenic amino acid sequences from the circumsporozoite (CS) antigen, ring-infected erythrocyte surface antigen (RESA), soluble (S) antigen, and falciparum interspersed repetitive antigen (FIRA) of P. falciparum, and the CS protein of P. vivax, were analyzed by computer. Matches of at least 4 amino acids were found for all sequences. In the database, 29 matches were found for human proteins and 26 matches were found for human viruses with the 6 antigen sequences. Most of the matched proteins, and many of the matched human viruses, are found in blood. The biological significance of these matches remains to be clarified.

Amino Acid Sequence↗

The complete sequence of the human intermediate filament chain keratin 10. Subdomainal divisions and model for folding of end domain sequences.

We present the complete amino acid sequence of the human keratin 10 (type I) intermediate filament chain expressed in terminally differentiated epidermal cells. Comparisons of this sequence with its mouse and bovine counterparts allow us to describe structural features of the functional end domains. First, sections of their respective end domains are highly conserved and permit a redefinition of earlier models for their subdomainal organization. The amino-terminal end domain consists of El, the first 57-58 residues that are basic, glycine-rich, and have been highly conserved among the three species; V1, a region of well-defined quasi repeats of the motif aliphatic-serine/glycinen; and H1, a newly recognized short acidic sequence that has been conserved among the type I keratin family. The carboxyl-terminal end consists of V2 and E2 whose properties but not sequence resemble V1 and E1, respectively. Second, since the E1, H1, and E2 sequences have been highly conserved between the three species, we suggest they are critical elements in defining intermediate filament function. Third, we note that the E and V sequences of the keratin 10 (and other keratin) chains share many properties in common with protein chain turns found in globular proteins. We therefore propose a model in which these sequences form omega loop-like structures (Leszczynski, J. N. & Rose, G. D. (1986) Science 234, 849-855) on the surface of keratin intermediate filaments. This represents the first specific proposal for the end domain structure of any intermediate filament chain.

Amino Acid Sequence↗

Nucleotide sequence of maize chloroplast rpS11 with conserved amino acid sequence between eukaryotes, bacteria and plastids.

Nucleotide sequence of a 721 base pair segment of maize chloroplast DNA, encoding the putative chloroplast ribosomal protein S11 at physical map position 33.1-33.5 Kbp, is described. A Shine-Dalgarno sequence and computer-derived stem-loop structures of dyad symmetry are present in the spacer region between rpS11 and its 5' upstream gene rpL36. The deduced amino acid sequence of maize chloroplast S11 shows 69%, 66%, 62%, 57%, 48% and 45% sequence identity to the corresponding sequences of tobacco, spinach, pea, liverwort, Escherichia coli and Bacillus subtilis, respectively, and 41% sequence identity to three eukaryotic cytoplasmic ribosomal proteins, S14 of Chinese hamster and of human and rp59 of yeast. Maize chloroplast r-protein S11 is larger than the other published S11s of plants and bacteria, due to the apparent tandem introduction of a short sequence stretch of internal homology.

Amino Acid Sequence↗

A LINE2 repetitive DNA sequence from the cichlid fish, Oreochromis niloticus: sequence analysis and chromosomal distribution.

We report the cloning and characterization of a long interspersed nucleotide element (LINE) from a cichlid fish, Oreochromis niloticus, and show the distribution of this element, called CiLINE2 for cichlid LINE2, in the chromosomes of this species. The identification of an open reading frame in CiLINE2 with amino acid sequence similarity to reverse transcriptases encoded by LINE-like elements in Caenorhabditis elegans, Platemys spixii, Schistosoma mansoni, Gallus gallus (CRI), Drosophila melanogaster (I factor), and Homo sapiens (LINE2), as well as the structure of the element, suggest it is a member of this family of non-long terminal repeat-containing retrotransposons. Search of a DNA sequence database identified sequences similar to CiLINE2 in four other fish species (Haplotaxodon microlepis, Oreochromis mossambicus, Pseudotropheus zebra, and Fugu rubripes). Southern blot hybridization experiments revealed the presence of sequences similar to CiLINE2 in all Tilapiini species analyzed from the genera Oreochromis, Tilapia, and Sarotherodon, and gave an estimated copy number of about 5500 for the haploid genome of O. niloticus. Fluorescent in situ hybridization showed that CiLINE2 sequences were organized in small clusters dispersed over all chromosomes of O. niloticus, with a higher concentration near chromosome ends. Furthermore, the long arm of chromosome 1 was strikingly enriched with this sequence. The distribution of LINE2-related elements might underlie the difference in chromosome banding patterns observed between cold-blooded vertebrates and mammals.

Amino Acid Sequence↗

Completion of molecular characterization of Toscana phlebovirus genome: nucleotide sequence, coding strategy of M genomic segment and its amino acid sequence comparison to other phleboviruses.

The M RNA segment of Toscana (TOS) phlebovirus was cloned and the complete nucleotide sequence determined. The M RNA segment is 4215 nucleotides in length, and it contains a single major open reading frame (ORF) in the viral-complementary sequence, between nucleotides 18 and 4034, which can encode for a polyprotein of 1339 amino acids (Mr 149 kDa). The viral segment is expressed via a unique mRNA containing 10-14 non-templated nucleotides at the 5' end and it is truncated at the 3' end by about 140 nucleotides in a purine-rich region. In M predicted amino acid sequences, several hydrophobic regions have been identified. They could function as a signal sequence or a transmembrane region for the different proteins. Comparison of the deduced amino acid sequence of M precursor product revealed 38, 36, and 25% identity and 58, 56, and 47% similarity with those of Rift Valley fever (RVF), Punta Toro (PT) and Unkuniemi (UUK) viruses, respectively. Residues conserved among the proteins are mainly located at the COOH-portion of the precursor, while the major divergence is in the NSm coding regions. Based on sequence comparison and similarity of hydropathic pattern of TOS M segment with other phleboviruses the N-termini of TOS GN and GC glycoproteins were placed at residues 297 and 936 of the precursor.

Amino Acid Sequence↗

Prediction of the coding sequences of unidentified human genes. II. The coding sequences of 40 new genes (KIAA0041-KIAA0080) deduced by analysis of cDNA clones from human cell line KG-1.

By applying the protocol previously established, we isolated and sequenced full-length cDNA clones longer than 2 kb from cDNA library of human immature myeloid cell line KG-1, and the coding sequences of 40 new genes were predicted. A computer search of the sequences indicated that 29 genes contained sequences with similarities to reported genes in the GenBank/EMBL databases. Significant transmembrane domains were identified in 9 genes, 5 of which harbored multiple hydrophobic regions. Protein motifs that matched those in the PROSITE motif database were identified in 13 genes. In terms of sequence similarities and protein motifs, 5 genes were related to transcriptional factors. Repetitive sequences were found in the 3'-untranslated region of 8 genes. Northern hybridization demonstrated that the expression of 9 genes was tissue-specific, while the remaining 31 genes were expressed ubiquitously. It was also noted that 17 genes yielded different sizes of bands possibly due to either alternative splicing or alternative initiation. The chromosomal location of these genes has been determined.

Amino Acid Sequence↗

The chromosomal organization of the human endogenous retrovirus-like sequence HERV-H: clustering of the HERV-H sequences in a 300-kb region close to the GRPR locus on the X chromosome.

Within the haploid genome there are approximately 1,000 copies of the human endogenous retrovirus-like sequence, HERV-H. Although these sequences are scattered throughout the entire genome, in situ hybridization experiments revealed that there are discrete clusters positioned on chromosomes 1 p and 7 q. In this study, we have located three HERV-H sequences which were unexpectedly clustered within a 300-kilobase region close to the GRPR locus on the X chromosome. In previous studies, no clustering of this sequence has been reported at this locus. Our finding demonstrates that, like other repetitive sequences, clustering of HERV-H occurs in the human genome, although these sequences may not always be detected by in situ hybridization methods.

Base Sequence↗

Sequence determination of three variable surface glycoproteins from Trypanosoma congolense. Conserved sequence and structural motifs.

The full-length cDNA sequences of three variable surface glycoproteins from bloodstream forms of Trypanosoma congolense have been determined. They encode preproteins of 429, 449, and 428 amino acids. These proteins contain the typical N-terminal leader sequences of secreted eukaryotic proteins, and display hydrophobic amino acids at their C-termini characteristic of variable surface glycoproteins; these leader sequences serve as transient membrane anchors after protein synthesis. By performing sequence comparisons of all currently known variable surface glycoproteins from T. congolense, several conserved elements could be identified. These elements included positional conservation of most of the cysteine residues, conservation of the flanking sequences surrounding these cysteine residues, clustering of proline residues near the C-termini, and a hydrophobic heptad motif near the end of the N-terminal domains. The N-terminal domains seem to be closely related to the B domains of Trypanosoma brucei variable surface glycoproteins, whereas the C domains have up to now only been identified in T. congolense variable surface glycoproteins. The data suggest that T. congolense variable surface glycoproteins, despite low sequence similarities, could have conserved tertiary structures.

Amino Acid Sequence↗