Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Sequence errors described in GenBank: a means to determine the accuracy of DNA sequence interpretation.

The accuracy of nucleic acid sequence data interpretation was determined by assessing and quantifying the discrepancies reported in the GenBank database. This permitted the calculation of an Error Rate (ER) for nucleic acid sequence determination. If one assumes that most entries (TB, Total Bases) were independently verified or those without reported discrepancies were correct, the ER is 0.368 errors per 1000 bases. However, if one assumes that only those sequences with reported discrepancies (TBIQ, Total Bases from entries In Question) were verified and are thus correct, the ER is 2.887 errors per 1000 bases. This establishes the first set of limit boundaries of the ER for sequence interpretation and sequence errors within the GenBank database and provides the foundation for future assessments and the monitoring of sequence data accumulation. In addition, the ER measure provides a basis to evaluate the efficiency and merit of present and future automated nucleic acid sequencing technologies which will have a direct impact upon the final outcome of the "Human Genome Initiative".

Base Sequence↗

Sequences, sequence clusters and bacterial species.

Whatever else they should share, strains of bacteria assigned to the same species should have house-keeping genes that are similar in sequence. Single gene sequences (or rRNA gene sequences) have very few informative sites to resolve the strains of closely related species, and relationships among similar species may be confounded by interspecies recombination. A more promising approach (multilocus sequence analysis, MLSA) is to concatenate the sequences of multiple house-keeping loci and to observe the patterns of clustering among large populations of strains of closely related named bacterial species. Recent studies have shown that large populations can be resolved into non-overlapping sequence clusters that agree well with species assigned by the standard microbiological methods. The use of clustering patterns to inform the division of closely related populations into species has many advantages for poorly studied bacteria (or to re-evaluate well-studied species), as it provides a way of recognizing natural discontinuities in the distribution of similar genotypes. Clustering patterns can be used by expert groups as the basis of a pragmatic approach to assigning species, taking into account whatever additional data are available (e.g. similarities in ecology, phenotype and gene content). The development of large MLSA Internet databases provides the ability to assign new strains to previously defined species clusters and an electronic taxonomy. The advantages and problems in using sequence clusters as the basis of species assignments are discussed.

Bacteria↗

Perspectives: sequence data base searching in the era of large-scale genomic sequencing.

Large-scale sequencing of human and model organism genomes will have a profound impact on our ability to use sequence data base searching to predict the biochemical functions of sequences of interest. Despite the great value of more sequences in the data bases, a huge increase in data base size will also have adverse effects on data base searches. Upcoming problems will include (1) greatly increased search times, (2) an increase in background noise of high-scoring but biologically irrelevant matches, (3) inaccurate coding region prediction, leading to problems in protein data base searching, and (4) limited first-pass sequence annotation, making it difficult to determine the biological relevance of data base hits. Improved data base annotation tools and construction of smaller data bases of representative and highly-annotated sequences for first-pass analyses will be essential to deal with the impending flood of new genomic sequence.

Animals↗

Sequencing a genome by walking with clone-end sequences: a mathematical analysis.

One approach to sequencing a large genome is (1) to sequence a collection of nonoverlapping "seeds" chosen from a genomic library of large-insert clones [such as bacterial artificial chromosomes (BACs)] and then (2) to take successive "walking" steps by selecting and sequencing minimally overlapping clones, using information such as clone-end sequences to identify the overlaps. In this paper we analyze the strategic issues involved in using this approach. We derive formulas showing how two key factors, the initial density of seed clones and the depth of the genomic library used for walking, affect the cost and time of a sequencing project-that is, the amount of redundant sequencing and the number of steps to cover the vast majority of the genome. We also discuss a variant strategy in which a second genomic library with clones having a somewhat smaller insert size is used to close gaps. This approach can dramatically decrease the amount of redundant sequencing, without affecting the rate at which the genome is covered.

Chromosome Walking↗

Terminal-sequence analysis of bacterial ribosomal RNA. Correlation between the 3'-terminal-polypyrimidine sequence of 16-S RNA and translational specificity of the ribosome.

The 3'-terminal sequences of 16-S ribosomal RNA from a number of bacteria have been determined by a stepwise degradation and 3'-terminal labelling procedure. The sequences obtained were: Bacillus stearothermophilus, -G(Z)approximately 5 Y-U-C-C-U-U-U-C-U (A); B. subtilis, -G(Z)approximately 7 Y-C-U-U-U-C-U; Caulobacter crescentus, -G(Z)3 Y-U-C-C-U-U-U-C-U; Pseudomonas aerugionosa, -G-Z-Z-Y-C-U-C-U-C-C-U-U(A), where Z is any nucleotide other than G. Thus, as previously found in Escherichia coli, all bacterial 16-S rRNAs contain a pyrimidine-rich tract at the 3'-terminus. In B. stearothermophilus and Ps. aeruginosa this region shows substantial heterogeneity involving the 3'-terminal adenylic acid. A low level of 3'-terminal heterogeneity cannot be excluded for the other bacterial 16-S rRNAs examined. The 3'-termini of bacterial 16-S rRNA can be divided into two groups on the basis of sequence homology. The first group comprises E. coli and Ps. aeruginosa; the second, B. stearothermophilus, B. subtilis and C. crescentus. This division correlates with a previous separation of bacterial ribosomes into two categories based on ability to translate different mRNA preparations [Stallcup, Sharrock & Rabinowitz (1974) Biochem. Biophys. Res. Commun. 58, 92-98]. We have previously proposed that the precise base sequence at the 3'-terminus of 16-S rRNA determines the intrinsic capacity of bacterial ribosomes to translate a particular cistron [Shine & Dalgarno (1975) Nature (Lond.) 254, 34-38]. No difference was found in the 3'-terminal heptanucleotide sequence of 16-S rRNA from bacteriophage T7-infected E. coli, as compared to that in uninfected cells. Thus, the T7-induced alteration in translational specificity of E. coli ribosomes is probably not mediated by modification of the terminal seven nucleotides of the smaller rRNA. The 3'-terminal sequences of the 23-S rRNA species were also determined. The sequences obtained were: B stearothermophilus and B. subtilis, -Y-C; C. crescentus, -Y-C-U; Ps. aeruginosa, -Y-C-A; E. coli, -G-Y-U-U-A-A-C-C-U-U. No evidence for 3'-terminal heterogeneity was found. The results obtained are discussed in relation to possible base-pairing roles for the 3'-end of 16-S rRNA in bacterial protein synthesis.

Bacillus subtilis↗

Many random sequences functionally replace the secretion signal sequence of yeast invertase.

In the process of protein secretion, amino-terminal signal sequences are key recognition elements; however, the relation between the primary sequence of an amino-terminal peptide and its ability to function as an export signal remains obscure. The limits of variation permitted for functional signal sequences were determined by replacement of the normal signal sequence of Saccharomyces cerevisiae invertase with essentially random peptide sequences. Since about one-fifth of these sequences can function as an export signal the specificity with which signal sequences are recognized must be very low.

Amino Acid Sequence↗

Electron microscope heteroduplex studies of sequence relations among bacterial plasmids: identification and mapping of the insertion sequences IS1 and IS2 in F and R plasmids.

Heteroduplex experiments between the plasmid R6 and one strand of the deoxyribonucleic acid (DNA) of a lambda phage carrying the insertion sequence IS1 show that IS1 occurs on R6 at the two previously mapped junctions of resistance transfer factor (RTF) DNA with R-determinant DNA. From previous heteroduplex experiments, it then follows that IS1 occurs at the same junctions in R6-5, R100-1, and R1 plasmids. Heteroduplex experiments with the DNA from a lambda phage carrying the insertion sequence IS2 show that one copy of IS2 occurs in R6, R6-5, and R100-1 (but not R1) at a point within the RTF with coordinates 67.5 TO 68.9 kilobase units (kb). In an accompanying paper, Ptashne and Cohen (1975) show that the insertion sequence IS3 occurs on R6 and R6-5. R100-25, a traC mutant, differs from its parent R100-1 only in that it contains an additional copy of IS1 inserted within the tra gene region of 82.1 kb. R100-31, atraX, TC-s mutant of R100-1, is deleted in R100-1 sequences starting at one of the IS3 termini (46.9 kb) and extending with RTF to 61.0 kb. Heteroduplex studies of F plasmids with the DNA of a lambda phage bearing insertion sequence IS2 show that the sequence of F with coordinates 16.3-17.6F is IS2. The occurrence of IS1 at the two junctions of R-determinant DNA and RTF DNA in R plasmids provides a structural basis to explain the mechanism of the previously observed formation of molecules containing one RTF unit and several tandem copies of the R-determinant unit, when R plasmids in Proteus mirabilis are grown in the presence of antibiotics, and the segregation of an R plasmid into an RTF unit and an R-determinant unit. In general, correlation of our results with previous studies shows that insertion sequences play a role in a variety of F- and R-related intra- and intermolecular recombination phenomena.

Base Sequence↗

Identification of Mycobacterium spp. by using a commercial 16S ribosomal DNA sequencing kit and additional sequencing libraries.

Current methods for identification of Mycobacterium spp. rely upon time-consuming phenotypic tests, mycolic acid analysis, and narrow-spectrum nucleic acid probes. Newer approaches include PCR and sequencing technologies. We evaluated the MicroSeq 500 16S ribosomal DNA (rDNA) bacterial sequencing kit (Applied Biosystems, Foster City, Calif.) for its ability to identify Mycobacterium isolates. The kit is based on PCR and sequencing of the first 500 bp of the bacterial rRNA gene. One hundred nineteen mycobacterial isolates (94 clinical isolates and 25 reference strains) were identified using traditional phenotypic methods and the MicroSeq system in conjunction with separate databases. The sequencing system gave 87% (104 of 119) concordant results when compared with traditional phenotypic methods. An independent laboratory using a separate database analyzed the sequences of the 15 discordant samples and confirmed the results. The use of 16S rDNA sequencing technology for identification of Mycobacterium spp. provides more rapid and more accurate characterization than do phenotypic methods. The MicroSeq 500 system simplifies the sequencing process but, in its present form, requires use of additional databases such as the Ribosomal Differentiation of Medical Microorganisms (RIDOM) to precisely identify subtypes of type strains and species not currently in the MicroSeq library.

DNA, Bacterial↗

Endogenous oncornaviral DNA sequences: evidence for two classes of viral DNA sequences in guinea pig cells.

The nature of the endogenous viral DNA sequences in guinea pig cells was studied by hybridization. A segment of the viral RNA (r-VRNA) hybridizing to abundant (or reiterated) DNA sequences (R-VDNA) was isolated by recycling to a Cot of 300. The hybridization of the recycled VRNA, as well as the total VRNA, was followed by determining their kinetics and by Wetmur-Davidson analysis. The kinetics of hybridization of total VRNA were complex, did not follow a second-order kinetics, and revealed two slopes by Wetmur-Davidson analysis. The recycled RNA, on the other hand, had a second-order reaction rate expected of the hybridization between a single species of RNA and DNA sequences and yielded a single straight line in a Wetmur-Davidson plot. The Cot1/2 and slope of the recycled r-VRNA was almost identical to that of the abundant VDNA sequences obtained from the hybridization data of the total VRNA. Guinea pig 28S rRNA with or without recycling was used in monitoring hybridization rate. The kinetics of hybridization of 28S RNA followed a second-order reaction and produced a single straight line by Wetmur-Davidson plot, with a second-order reassociation rate constant of 9.6 x 10(-3) liters/mol-s, a Cot1/2 of 104 mol-s/liter, and reiteration frequency of 146. There was no difference in the kinetics of hybridization of 28S RNA before and after recycling. These experiments showed that guinea pig cells contain two classes of VDNA sequences. (i) R-VDNA sequences with a second-order reassociation rate constant of 8.2 x 10(-4) liters/mol-s, a Cot1/2 of 1,219 mol-s/liter, and a reiteration frequency of 12 represent 37.5% of the viral genome. (ii) Unique VDNA sequences with a second-order reassociation rate constant of 1.2 x 10(-4) liters/mol-s, a Cot1/2 of 7,692 mol-s/liter, and a reiteration frequency of 2 represent 62.5% of the viral genome.

Animals↗

[Sequence complexity of transcribed unique DNA sequences in genome of mouse P815 mastocytoma cells (author's transl)].

The sequence complexity of nuclear RNA from mouse liver, mouse spleen and highly malignant P815 mastocytoma was measured by nRNA driven hybridization to unique DNA sequences of P815 cells. The unique DNA sequences represent 63% of the total nuclear DNA of P815 cells and their availibility in hybridization experiments was found to be 76%. Of these sequences 7.8% formed hybrids with nuclear RNA of this cell, about 11.5% with mouse spleen and about 14.5% with mouse liver nuclear RNA. Assuming an asymmetrical transcription, the complexities of these transcripts are 2.8 X 10(8) nucleotides for mouse P815 mastocytomas, 4.3 X 10(8) for mouse spleen and about 5.3 X 10(8) nucleotides for mouse liver. Cellular specifity of the transcribed information was analyzed in additivity experiments, in which unique DNA sequences, not complementary to the nuclear RNA of one cell were annealed to the nuclear RNAs of the two other tissues/cells. In these experiments most of the nuclear RNA sequences of P815 cells were found to be also present in the nucleus of mouse liver and spleen. Only a small portion of the unique DNA sequences of P815 mastocytoma (about 1.2% corresponding to 4.4 X 10(7) nucleotides) was found to be complementary only to P815 mastocytoma nuclear RNA.

Animals↗

Sequencing of simple sequence repeat anchored polymerase chain reaction amplification products of Biomphalaria glabrata.

Simple sequence repeat anchored polymerase chain reaction amplification (SSR-PCR) is a genetic typing technique based on primers anchored at the 5' or 3' ends of microsatellites, at high primer annealing temperatures. This technique has already been used in studies of genetic variability of several organisms, using different primer designs. In order to conduct a detailed study of the SSR-PCR genomic targets, we cloned and sequenced 20 unique amplification products of two commonly used primers, CAA(CT)6 and (CA)8RY, using Biomphalaria glabrata genomic DNA as template. The sequences obtained were novel B. glabrata genomic sequences. It was observed that 15 clones contained microsatellites between priming sites. Out of 40 clones, seven contained complex sequence repetitions. One of the repeats that appeared in six of the amplified fragments generated a single band in Southern analysis, indicating that the sequence was not widespread in the genome. Most of the annealing sites for the CAA(CT)6 primer contained only the six repeats found within the primer sequence. In conclusion, SSR-PCR is a useful genotyping technique. However, the premise of the SSR-PCR technique, verified with the CAA(CT)6 primer, could not be supported since the amplification products did not result necessarily from microsatellite loci amplification.

Animals↗

Capillary DNA sequencing: maximizing the sequence output.

Like most other DNA sequencing core facilities, one of our continuing goals is to improve our sequence output without substantially adding to cost. To minimize sample-to-sample variability in template DNA concentration, we implemented the rolling circle amplification (RCA) procedure for preparing our DNA templates. In addition to saving time and reducing the number of steps in template DNA preparation, the RCA method has the potential to normalize the DNA concentration in samples that can be sequenced directly without additional purification. In the present study, we used RCA-generated templates to test a recently reported procedure that increased sequence quality by resuspending the sequenced products in low concentrations of agarose before capillary electrophoresis (CE) on a MegaBACE 1000 platform. Although we did not obtain the expected result using the specified procedure, a modification resulted in up to 60% increase in total sequence yield per sample plate. A combination of agarose and formamide-EDTA in the resuspension solution enabled us to generate long-read and high-quality sequences for more than 38,000 templates with minimal additional cost.

DNA, Bacterial↗

Sequence of active site peptides from the penicillin-sensitive D-alanine carboxypeptidase of Bacillus subtilis. Mechanism of penicillin action and sequence homology to beta-lactamases.

It has been proposed that penicillin and other beta-lactam antibiotics are substrate analogs which inactivate certain essential enzymes of bacterial cell wall biosynthesis by acylating a catalytic site amino acid residue (Tipper, D.J., and Strominger, J.L. (1965) Proc. Natl. Acad. Sci. U.S.A. 54, 1133-1141). A key prediction of this hypothesis, that the penicilloyl moiety and an acyl moiety derived from substrate both bind to the same active site residue, has been examined. D-Alanine carboxypeptidase, a penicillin-sensitive membrane enzyme, was purified from Bacillus subtilis and labeled covalently at the antibiotic binding site with [14C]penicillin G or with the cephalosporin [14C]cefoxitin. Alternatively, an acyl moiety derived from the depsipeptide substrate [14C]diacetyl L-Lys-D-Ala-D-lactate was trapped at the catalytic site in near-stoichiometric amounts by rapid denaturation of an acyl-enzyme intermediate. Radiolabeled peptides were purified from a pepsin digest of each of the 14C-labeled D-alanine carboxypeptidases and their amino acid sequences determined. Antibiotic- and substrate-labeled peptic peptides had the same sequence: Tyr-Ser-Lys-Asn-Ala-Asp-Lys-Arg-Leu-Pro-Ile-Ala-Ser-Met. Acyl moieties derived from antibiotic and from substrate were shown to be bound covalently in ester linkage to the identical amino acid residue, a serine at the penultimate position of the peptic peptide. These studies establish that beta-lactam antibiotics are indeed active site-directed acylating agents. Additional amino acid sequence data were obtained by isolating and sequencing [14C]penicilloyl peptides after digestion of [14C]penicilloyl D-alanine carboxypeptidase with either trypsin or cyanogen bromide and by NH2-terminal sequencing of the uncleaved protein. The sequence of the NH2-terminal 64 amino acids was thus determined and the active site serine then identified as residue 36. A computer search for homologous proteins indicated significant sequence homology between the active site of D-alanine carboxypeptidase and the NH2-terminal portion of beta-lactamases. Maximum homology was obtained when the active site serine of D-alanine carboxypeptidase was aligned correctly with a serine likely to be involved in beta-lactamase catalysis. These findings provide strong evidence that penicillin-sensitive D-alanine carboxypeptidases and penicillin-inactivating beta-lactamases are related evolutionarily.

Amino Acid Sequence↗

[Single copy and repetitive DNA sequences of Echinodermata. I. Arrangement of DNA sequences of different multiplicity].

Arrangement of repetitive and single copy DNA sequences in the DNA of 8 Echinodermata species (sea urchins, starfishes and sea-cucumber) has been studied. Comparison of the reassociation kinetics of short and long DNA fragments assayed by hydroxyapatite binding indicates that the pattern of DNA sequence organization of all these species is similar to the so called Xenopus pattern found in genomes of most animals and plants. Interspecies differences consist mainly in the quantities of sequences of various repetition degrees and their interspersion with each other and with single copy sequences. Measurements of the size of S1 nuclease resistant reassociated repetitive sequences show variability in the relative quantities of long and short repetitive sequences of different species. Difference in the arrangement of single copy and repetitive sequences between Echinodermata species are not related to their evolutionary proximity.

Animals↗

Dihydrofolate reductase from amethopterin-resistant Lactobacillus casei. Sequences of the cyanogen bromide peptides and complete sequences of the enzyme.

The complete amino acid sequence of dihydrofolate reductase from an amethopterin-resistant strain of Lactobacillus casei has been determined by sequence analysis of peptides produced by cleavage with cyanogen bromide, trypsin, staphylococcal protease, and myxobacter protease. Comparison of this sequence with those of reductases from other bacterial sources shows that the enzymes are homologous. The Lactobacillus casei reductase sequences shows a 29% sequence identity with that of the Escherichia coli enzyme and a 34% identity with the sequence of the enzyme from Streptococcus faecium. The NH2-terminal 68 residues of the L. casei reductase show a 54% sequence identity with that of the enzyme from S. faecium.

Amino Acid Sequence↗

Comparative nucleotide and amino acid sequence analysis of the sequence-specific RNA-binding rotavirus nonstructural protein NSP3.

NSP3, an acidic nonstructural protein, encoded by gene 7 has been implicated as the key player in the assembly of the 11 viral plus-strand RNAs into the early replication intermediates during rotavirus morphogenesis. To date, the sequence of NSP3 from only three animal rotaviruses (SA11, SA114F, and bovine UK) has been determined and that from a human strain has not been reported. To determine the genetic diversity among gene 7 alleles from group A rotaviruses, the nucleotide sequence of the NSP3 gene from 13 strains belonging to nine different G serotypes, from both humans and animals, has been determined. Based on the amino acid sequence identity as well as phylogenetic analysis, NSP3 from group A rotaviruses falls into three evolutionarily related groups, i.e., the SA11 group, the Wa group, and the S2 group. The SA11/SA114F gene appears to have a distant ancestral origin from that of the others and codes for a polypeptide of 315 amino acids (aa) in length. NSP3 from all other group A rotaviruses is only 313 aa in length because of a 2-amino-acid deletion near the carboxy-terminus. While the SA114F gene has the longest 3' untranslated region (UTR) of 132 nucleotides, that from other strains suffered deletions of varying lengths at two positions downstream of the translational termination codon. In spite of the divergence of the nucleotide (nt) sequence in the protein coding region, a stretch of about 80 nt in the 3' UTR is highly conserved in the NSP3 gene from all the strains. This conserved sequence in the 3' UTR might play an important role in the regulation of expression of the NSP3 gene.

Amino Acid Sequence↗

Sequence and transcriptional analysis of an orf virus gene encoding ankyrin-like repeat sequences.

A 1608 bp region located approximately 5.0 kb from the left end of the orf virus (OV) genome (strain NZ2) was sequenced. The sequence revealed a single open reading frame designated G1L. The predicted amino acid sequence of G1L contained eight ankyrinlike repeat sequences. Transcriptional analysis of G1L showed it was transcribed towards the genome terminus during the early phase of infection. S1 nuclease and primer extension analyses showed that the transcriptional start site of the gene was located a short distance downstream from an A + T-rich sequence similar to a vaccinia virus early promoter.

Amino Acid Sequence↗

AF4/FEL, a gene involved in infant leukemia: sequence variations, gene structure, and possible homology with a genomic sequence on 5q31.

The most common chromosome abnormality among infants with acute lymphoblastic leukemia is a t(4;11)(q2l;q23) and patients with this 4;11 translocation have a very poor prognosis. This unique genetic rearrangement fuses the MLL/ALL-1/HRX-Htrx gene at 11q23 with the AF4/FEL gene at 4q21. The resulting chimeric mRNAs presumably encode chimeric proteins which contribute to the leukemogenic state. The AF4 gene remains poorly understood with an unknown function. In this report, we describe the cDNA sequence information from human placental tissue where AF4 mRNA is highly expressed. We identified six intron-exon boundaries in the AF4 genomic structure and discussed more than 30 AF4 cDNA sequence variations reported in the literature. In addition, we identified three overlapping genomic sequences in GenBank entitled the "interleukin growth hormone cluster on chromosome 5q31," which, when aligned and translated, had three regions that suggested homology to the predicted AF4 protein sequence (32% amino acid sequence identity over 314 amino acids, 43% over 63 amino acids, and 50% over 40 amino acids). Of interest, this same chromosome 5q31 region has also been implicated in MLL gene rearrangements in human leukemia.

Amino Acid Sequence↗