Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

Site-specific mutagenesis on cloned DNAs: generation of a mutant of Escherichia coli tyrosine suppressor tRNA in which the sequence G-T-T-C corresponding to the universal G-T-pseudouracil-C sequence of tRNAs is changed to G-A-T-C.

We have cloned the Escherichia coli tyrosine-inserting amber suppressor tRNA gene into the recombinant single-strand phage M12mp3. By using the M13mp3SuIII+ recombinant phage DNA as template and an oligonucleotide bearing a mismatch as primer, we have synthesized in vitro an M13mp3SuIII heteroduplex DNA that has a single mismatch at a predetermined site in the tRNA gene. Transformation of E. coli with the heteroduplex DNA yielded M13 recombinant phages carrying a mutant suppressor tRNA gene in which the sequence G-T-T-C, corresponding to the universal G-T-pseudouracil-C sequence in E. coli tRNAs, is changed to G-A-T-C. The mutant DNA has been characterized by restriction mapping and by sequence analysis. In contrast to results with the wild-type suppressor tRNA gene, cells transformed with recombinant plasmids carrying the mutant tRNA gene are phenotypically Su-. Thus, the single nucleotide change introduced has inactivated the function of the tRNA gene. By using E. coli minicells for studying the expression in vivo of cloned tRNA genes, we have found that cells transformed with recombinant plasmids carrying the mutant tRNA gene contain very little, if any, mature mutant suppressor tRNA. In contrast, the predominant low molecular weight RNA in cells transformed with recombinant plasmids carrying the wild-type suppressor tRNA gene is the mature tyrosine suppressor tRNA. Thus, while our results imply an important role for the G-T-pseudouracil-C sequence common to all E. coli tRNAs, whether this sequence is essential for tRNA biosynthesis, tRNA stability in vivo, or tRNA function remains to be determined. The procedures used to generate the mutant should be of general application toward site-specific mutagenesis on cloned DNAs, including regions that possess high degrees of secondary structure. In addition, the frequency of mutants among the progeny is high enough to enable one to identify and isolate site-specific mutants on any cloned DNA without requiring phenotypic selection.

Base Sequence↗

Secondary structure of the Tetrahymena ribosomal RNA intervening sequence: structural homology with fungal mitochondrial intervening sequences.

Splicing of the ribosomal RNA precursor of Tetrahymena is an autocatalytic reaction, requiring no enzyme or other protein in vitro. The structure of the intervening sequence (IVS) appears to direct the cleavage/ligation reactions involved in pre-rRNA splicing and IVS cyclization. We have probed this structure by treating the linear excised IVS RNA under nondenaturing conditions with various single- and double-strand-specific nucleases and then mapping the cleavage sites by using sequencing gel electrophoresis. A computer program was then used to predict the lowest-free-energy secondary structure consistent with the nuclease cleavage data. The resulting structure is appealing in that the ends of the IVS are in proximity; thus, the IVS can help align the adjacent coding regions (exons) for ligation, and IVS cyclization can occur. The Tetrahymena IVS has several sequences in common with those of fungal mitochondrial mRNA and rRNA IVSs, sequences that by genetic analysis are known to be important cis-acting elements for splicing of the mitochondrial RNAs. In the predicted structure of the Tetrahymena IVS, these sequences interact in a pairwise manner similar to that postulated for the mitochondrial IVSs. These findings suggest a common origin of some nuclear and mitochondrial introns and common elements in the mechanism of their splicing.

Base Sequence↗

A fidelity assay using "dideoxy" DNA sequencing: a measurement of sequence dependence and frequency of forming 5-bromouracil X guanine base mispairs.

DNA replication fidelity has been assayed by using a modified DNA sequencing reaction. In one experimental approach, dideoxycytidine 5'-triphosphate (ddCTP) was used as a chain terminator during replication of M13 phage DNA by the large fragment of DNA polymerase I. The deoxyribonucleotide analogue BrdUTP was used to compete against ddCTP-induced chain terminations as an assay for B X G base mispairing (B represents bromodeoxyuridine when the analogue is present as a base pair or base mispair). By comparing BrdUTP to dCTP for competition against ddCTP, an average misincorporation frequency for BrdUMP of 0.2% was found. A similar average misincorporation frequency has been measured previously for the incorporation of radioactively labeled BrdUMP and dCMP into the synthetic template-primer poly-[d(G,T)] X oligo(dA). The advantage of the sequencing method is that an error frequency is determined for each template guanine in a defined DNA sequence, thus providing information on the effect of neighboring base sequences on fidelity. Misincorporation frequencies varied no more than 5-fold among 50 template guanines tested. The approach used here is not limited for use with nucleotide analogues but is generally applicable in determining misincorporation frequencies and sequence specificities for any deoxynucleoside triphosphate substrate. In a second experimental approach, base mispairing between bromouracil and guanine was demonstrated directly by using 5-bromodideoxyuridine 5'-triphosphate (BrddUTP). A comparison of chain terminations attributable to BrddUTP and to dideoxythymidine 5'-triphosphate (ddTTP) revealed that B X A and T X A base pairs formed at about the same rate, whereas B X G mispairs occurred 4-10 times more frequently than T X G. The elevation in the frequency of B X G over T X G mispairs is consistent with the mutagenic behavior of the base analogue.

Base Sequence↗

Complete sequence of a gene encoding a human type I keratin: sequences homologous to enhancer elements in the regulatory region of the gene.

We report here the complete nucleotide sequence of a gene encoding the 50-kDa keratin expressed in abundance in human epidermal cells. According to its sequence, this gene has a single transcriptional initiation site and a single polyadenylylation signal. Nuclease S1 mapping of this gene with total human epidermal mRNA confirmed the presence of a single initiation site for the 50-kDa keratin gene. When the regulatory sequences 5' upstream from this gene were examined, three sequences that share significant homology with viral and immunoglobulin enhancer elements were found. In comparison, the sequence of the regulatory region of vimentin, a structurally similar intermediate filament gene, was highly divergent [Quax, W., Egberts, W. V., Hendriks, W., Quax-Jeuken, Y. & Bloemendal, H. (1983) Cell 35, 215-223]. This finding may provide a clue to understanding the molecular mechanisms underlying the widely varying levels of expression of different intermediate filament genes in different tissues.

Base Sequence↗

Statistical geometry in sequence space: a method of quantitative comparative sequence analysis.

A statistical method of comparative sequence analysis that combines horizontal and vertical correlations among aligned sequences is introduced. It is based on the analysis mainly of quartet combinations of sequences considered as geometric configurations in sequence space. Numerical invariants related to relative internal segment lengths are assigned to each such configuration and statistical averages of these invariants are established. They are used for internal calibration of the topology of divergence and for quantitative determination of the noise level. Comparison of computer simulations with experimental data reveals the high sensitivity of assignment of basic topologies even if much randomized. In addition, these procedures are checked by vertical analysis of the aligned sequences to allow the study of divergences with positionally varying substitution probabilities.

Base Sequence↗

Transcriptional sequencing: A method for DNA sequencing using RNA polymerase.

We have developed a sequencing method based on the RNA polymerase chain termination reaction with rhodamine dye attached to 3'-deoxynucleoside triphosphate (3'-dNTP). This method enables us to conduct a rapid isothermal sequencing reaction in <30 min, to reduce the amount of template required, and to do PCR direct sequencing without the elimination of primers and 2'-dNTP, which disturbs the Sanger sequencing reaction. An accurate and longer read length was made possible by newly designed four-color dye-3'-dNTPs and mutated RNA polymerase with an improved incorporation rate of 3'-dNTP. This method should be useful for large-scale sequencing in genome projects and clinical diagnosis.

DNA↗

HIV-1 integrase interaction with U3 and U5 terminal sequences in vitro defined using substrates with random sequences.

Successful integration of viral genome into a host chromosome depends on interaction between viral integrase and its recognition sequences. We have used a reconstituted concerted human immunodeficiency virus, type 1 (HIV-1), integration system to analyze the role of integrase (IN) recognition sequences in formation of the IN-viral DNA complex capable of concerted integration. HIV-1 integrase was presented with substrates that contained all 4 bases at 8 mismatched positions that define the inverted repeat relationship between U3 and U5 long terminal repeats (LTR) termini and at positions 17-19, which are conserved in the termini. Evidence presented indicates that positions 17-20 of the IN recognition sequences are needed for a concerted DNA integration mechanism. All 4 bases were found at each randomized position in sequenced concerted DNA integrants, although in some instances there were preferences for specific bases. These results indicate that integrase tolerates a significant amount of plasticity as to what constitutes an IN recognition sequence. By having several positions randomized, the concerted integrants were examined for statistically significant relationships between selections of bases at different positions. The results of this analysis show not only relationships between different positions within the same LTR end but also between different positions belonging to opposite DNA termini.

Base Sequence↗

GEOMETRY: a software package for nucleotide sequence analysis using statistical geometry in sequence space.

GEOMETRY is a software package for the analysis of nucleotide sequences using the method of statistical geometry in sequence space. The package consists of programs performing estimation of the average geometry of sequence quartets, analysis of positional variability and computer simulation of parallel and tree-like sequence divergence with user-defined parameters. It provides an independent tool for evaluation of the reliability of conventional phylogenetic trees and calibration of the time of sequence divergence. GEOMETRY may be of interest for all scientists engaged in the study of molecular phylogeny. The package is available by anonymous FTP from ftp.bionet.nsk.su, directory /incoming/molevol/geom.exe, and will be available from EMBL file server (URL: http://@www.ebi.ac.uk)

Algorithms↗

H3 and H4 histone cDNA sequences from Xenopus: a sequence comparison of H4 genes.

Ovarian poly (A) + RNA from Xenopus laevis and Xenopus borealis was used to construct two cDNA libraries which were screened for histone sequences. cDNA clones to H4 mRNA were obtained from both species and an H3 cDNA clone from Xenopus laevis. The complete DNA sequences of these clones have been determined and are presented. These new sequences are compared with other H3 and H4 DNA sequences both in the coding and 3' noncoding regions. We find that there is considerable non-random codon usage in ten H4 genes. In addition there are some sequence similarities in the 3' noncoding regions of H3 and H4 genes.

Animals↗

Retrovirus-related sequences in human DNA: detection and cloning of sequences which hybridize with the long terminal repeat of baboon endogenous virus.

Human DNA sequences which hybridized with the long terminal repeats (LTR) of baboon type C virus M7 were detected by non-stringent blot hybridization. About 7 to 10 discrete bands of the LTR-related sequences were commonly observed in the DNAs from four independent human cell lines after digestion with either Eco RI, Hind III or Bam HI. The amounts of these sequences were more abundant in tumor cell lines than in a non-malignant cell line. The human sequences related to the M7 LTR seemed to be located at relatively specific sites on the cell DNA. The human DNA clones which hybridized with M7 LTR were detected in the human DNA library described by Lawn et al. (Cell 15, 1157-1174, 1978), at a frequency of about 300 per haploid genome. Five clones were isolated which shared different extent of homology with M7 LTR and whose restriction maps were totally different one another. The DNA structures of two of them resembled the genome of retroviruses. These results suggest the presence of various types of the LTR-related sequences in human DNA: some of them might represent endogenous virus genomes of human cells.

Animals↗

Kilo-sequencing: an ordered strategy for rapid DNA sequence data acquisition.

A strategy for rapid DNA sequence acquisition in an ordered, nonrandom manner, while retaining all of the conveniences of the dideoxy method with M13 transducing phage DNA template, is described. Target DNA 3 to 14 kb in size can be stably carried by our M13 vectors. Suitable targets are stretches of DNA which lack an enzyme recognition site which is unique on our cloning vectors and adjacent to the sequencing primer; current sites that are so useful when lacking are Pst, Xba, HindIII, BglII, EcoRI. By an in vitro procedure, we cut RF DNA once randomly and once specifically, to create thousands of deletions which start at the unique restriction site adjacent to the dideoxy sequencing primer and extend various distances across the target DNA. Phage carrying a desired size of deletions, whose DNA as template will give rise to DNA sequence data in a desired location along the target DNA, may be purified by electrophoresis alive on agarose gels. Phage running in the same location on the agarose gel thus conveniently give rise to nucleotide sequence data from the same kilobase of target DNA.

Base Sequence↗

Using iodinated single-stranded M13 probes to facilitate rapid DNA sequence analysis--nucleotide sequence of a mouse lysine tRNA gene.

From a recombinant lambda phage, we have determined a 387 bp sequence containing a mouse lysine tRNA gene. The putative lys tRNA (anticodon UUU) differs from rabbit liver lys tRNA at five positions. The flanking regions of the mouse gene are not generally homologous to published human and Drosophila lys tRNA genes. However, the mouse gene contains a 14 bp region comprising 13 A-T base pairs, 30-44 bp from the 5' end of the coding region. Cognate A-T rich regions are present in human and Drosophila genes. The coding region is flanked by two 11 bp direct repeats, similar to those associated with alu family sequences. The sequence was determined by a "walking" protocol that employs, as a novel feature, iodinated single-stranded M13 probes to identify M13 subclones which contain sequences partially overlapping and contiguous to an initially determined sequence. The probes can also be used to screen lambda phage and in Southern and dot blot experiments.

Animals↗

Eleven new sequence variants of citrus exocortis viroid and the correlation of sequence with pathogenicity.

Full-length double-stranded cDNA was prepared from purified circular RNA of two new Australian field isolates of citrus exocortis viroid (CEV) using two synthetic oligodeoxynucleotide primers. The cDNA was then cloned into the phage vector M13mp9 for sequence analysis. Sequencing of nine cDNA clones of isolate CEV-DE30 and eleven cDNA clones of isolate CEV-J indicated that both isolates consisted of a mixture of viroid species and led to the discovery of eleven new sequence variants of CEV. These new variants, together with the six reported previously, form two classes of sequence which differ by a minimum of 26 nucleotides in a total of 370 to 375 residues. These two classes correlate with two biologically distinct groups when propagated on tomato plants where one produces severe symptoms and the other gives rise to mild symptoms. Two regions of the native structure of CEV, comprising 18% of the total residues, differ between the sequence variants of mild and severe isolates. Whether or not both of these regions are essential for the variation in pathogenicity has yet to be determined.

Base Sequence↗

Finding the most significant common sequence and structure motifs in a set of RNA sequences.

We present a computational scheme to locally align a collection of RNA sequences using sequence and structure constraints. In addition, the method searches for the resulting alignments with the most significant common motifs, among all possible collections. The first part utilizes a simplified version of the Sankoff algorithm for simultaneous folding and alignment of RNA sequences, but maintains tractability by constructing multi-sequence alignments from pairwise comparisons. The algorithm finds the multiple alignments using a greedy approach and has similarities to both CLUSTAL and CONSENSUS, but the core algorithm assures that the pairwise alignments are optimized for both sequence and structure conservation. The choice of scoring system and the method of progressively constructing the final solution are important considerations that are discussed. Example solutions, and comparisons with other approaches, are provided. The solutions include finding consensus structures identical to published ones.

Algorithms↗

Active Sequences Collection (ASC) database: a new tool to assign functions to protein sequences.

Active Sequences Collection (ASC) is a collection of amino acid sequences, with an unique feature: only short sequences are collected, with a demonstrated biological activity. The current version of ASC consists of three sections: DORRS, a collection of active RGD-containing peptides; TRANSIT, a collection of protein regions active as substrates of transglutaminase enzyme (TGase), and BAC, a collection of short peptides with demonstrated biological activity. Literature references for each entry are reported, as well as cross references to other databases, when available. The current version of ASC includes more than 800 different entries. The main scope of this collection is to offer a new tool to investigate the structural features of protein active sites, additionally to similarity searches against large protein databases or searching for known functional patterns. ASC database is available at the web address http://crisceb.unina2.it/ASC/ which also offers a dedicated query interface to compare user-defined protein sequences with the database, as well as an updating interface to allow contribution of new referenced active sequences.

Amino Acid Sequence↗

Simultaneous determination of different DNA sequences by mass spectrometric evaluation of Sanger sequencing reactions.

All currently available DNA sequencing protocols rest fundamentally upon the homogeneity of the template. In this paper we describe the parallel DNA sequencing of various templates in one sample by a combination of the Sanger method and MALDI-TOF mass spectrometric analysis of the products. PCR-amplified hypervariable 16S rDNA fragments of the bacterium Escherichia coli DF1020 and cDNA of the 6-phosphofructo-1-kinase isoenzymes (PFK-1, EC 2.7.1.11) in rat brain were chosen as model systems for essentially heterogeneous templates. Avoiding cloning of the inhomogeneous PCR products we were able to read three sequences for both the 16S rDNA fragment of E.coli DF1020 and the cDNA of 6-phosphofructo-1-kinase from the peak lists of the Sanger sequencing reactions. Short sequences with a length between 21 and 25 nt were sufficient to reflect the heterogeneity of the 16S rDNA genes in E.coli and the existence of three isoenzymes of PFK-1 in rat brain.

Animals↗

Types and frequencies of sequencing errors in methyl-filtered and high c0t maize genome survey sequences.

The Maize Genome Sequencing Consortium has deposited into GenBank more than 850,000 maize (Zea mays) genome survey sequences (GSSs) generated via two gene enrichment strategies, methylation filtration and high-C(0)t (HC) fractionation. These GSSs are a valuable resource for generating genome assemblies and the discovery of single nucleotide polymorphisms and nearly identical paralogs. Based on the rate of mismatches between 183 GSSs (105 methylation filtration + 78 HC) and 10 control genes, the rate of sequencing errors in these GSSs is 2.3 x 10(-3). As expected many of these errors were derived from insufficient vector trimming and base-calling errors. Surprisingly, however, some errors were due to cloning artifacts. These G.C to A.T transitions are restricted to HC clones; over 40% of HC clones contain at least one such artifact. Because it is not possible to distinguish the cloning artifacts from biologically relevant polymorphisms, HC sequences should be used with caution for the discovery of single nucleotide polymorphisms or paramorphisms. The average rate of sequencing errors was reduced 6-fold (to 3.6 x 10(-4)) by applying more stringent trimming parameters. This trimming resulted in the loss of only 11% of the bases (15,469/144,968). Due to redundancy among GSSs this more stringent trimming reduced coverage of promoters, exons, and introns by only 0%, 1%, and 4%, respectively. Hence, at the cost of a very modest loss of gene coverage, the quality of these maize GSSs can approach Bermuda standards, even prior to assembly.

DNA, Plant↗

The amino-acid sequence of isoinhibitor K form snails (Helix pomatia). A sequence determination by automated Edman degradation and mass-spectral identification of the phenylthiohydantoins.

Performic-acid-oxidized isoinhibitor K of snails (Helix pomatia) was subjected to arginine-directed tryptic proteolysis. Six peptide fragments including one overlap peptide from limited cleavage of the Arg-3-Pro-4 bond were purified to homogeneity. Four arginine peptides and the C-terminal peptide were sequenced by automatic Edman degradation using a special peptide program. The phenylthiohydantoins were all identified by chemical ionization mass spectrometry, except four cysteic acid residues that were identified on an amino acid analyzer after acid hydrolysis. Quantitative evaluation of the phenylthiohydantoins by chemical ionization mass spectrometry using total molecular-ion beam integration greatly facilitated sequencing. The mass spectrum of the dipeptide less than Glu-Gly revealed that the N-terminus was blocked by pyroglutamic acid. The complete amino acid sequence of isoinhibitor K was determined. An almost 50% homology between the sequences of the snail inhibitor and the bovine trypsin-kallikrein inhibitor (Kunitz) became obvious. A comparison of all homologous sequences of this particular class of proteins known to date is presented.

Amino Acid Sequence↗