Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “oligonucleotide composition”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Changes in primary DNA sequence complexity influence the phenotypic consequences of mutations in human gene regulatory regions.

No general rules have been proposed to account for the functional consequences of gene regulatory mutations. In a first attempt to establish the nature of such rules, an analysis was performed of the DNA sequence context of 153 different single base-pair substitutions in the regulatory regions of 65 different human genes underlying inherited disease. Use of a recently proposed measure of DNA sequence complexity (taking into account the level of structural repetitiveness of a DNA sequence, rather than simply the oligonucleotide composition) has served to demonstrate that the concomitant change in local DNA sequence complexity surrounding a substituted nucleotide is related to the likelihood of a regulatory mutation coming to clinical attention. Mutations that led to an increase in complexity exhibited higher odds ratios in favour of pathological consequences than mutations that led to a decrease or left complexity unchanged. This relationship, however, was discernible only for pyrimidine-to-purine transversions. Odds ratios for other types of substitution were not found to be significantly associated with local changes in sequence complexity, even though a trend similar to that observed for Y-->R transversions was also apparent for transitions. These findings suggest that the maintenance of a defined level of DNA sequence complexity, or at least the avoidance of an increase in sequence complexity, is a critical prerequisite for the function of gene regulatory regions.

Base Sequence↗

Small cytoplasmic RNA associated with polyadenylated RNA is involved in the hormonal regulation of gene expression.

The fraction of small RNA (sacc-RNA) associated with cytoplasmic rat liver poly(A)+ RNA by non-covalent, possibly complementary, interactions has been isolated and studied. Fingerprint analysis and Northern blot hybridization data reveal that the specific changes occur in the population of sacc-RNA in response to glucocorticoid treatment. The close similarity of the oligonucleotide composition of sacc-RNA and RNA-component of small nuclear RNP-acceptor of glucocorticoid hormones has been found. The hypothesis of the involvement of the small RNA in the hormonal regulation of posttranscriptional stages of gene expression in the cytoplasm has been put forward.

Animals↗

Genes for VA-RNA in adenovirus 2.

VA-RNA from adenovirus 2 (Ad2) infected cells is shown to consist of two species. The gene coding for the major species maps at position 30 on the viral DNA, where it spans a site cleaved by the restriction enzyme Bam Hl. The minor species, constituting a few percent of the total VA-RNA, is distantly related in oligonucleotide composition to the major species. Its template maps within 700 base pairs to the right of the gene for the major species. The direction of transcription is from left to right on the conventional Ad2 map. These results, gained with a novel method for blotting (E. M. Southern, manuscript in preparation), have also led to the identification of the Bam Hl recognition sequence.

Adenoviridae↗

Identification of initiation sites for the in vitro transcription of rRNA operons rrnE and rrnA in E. coli.

The transcription initiation sites of E. coli rRNA operons were determined using various DNA fragments derived from transducing phage lambda metA20 carrying rrnE and from hybrid plasmid pLC19-3 carrying rrnA. In vitro transcription products were analyzed for their 5' end sequences and their oligonucleotide compositions. The results are in full agreement with the nuceotide sequences of the DNA templates described in an accompanying paper (de Boer, Gilbert and Nomura, 1979) and allow us to make the following conclusions. First, there are two transcription, start sites on each of the rRNA operons; they are 109 bp apart in the case of rrnE and 117 +/- 1 bp aprart in rrnA. Second, the first start site is 283 bp upstream from the m16S rRNA coding region in the case of rrnE, while is 291 bp upstream in rrnA. Initiation starts with ATP in both cases. Finally, the second start sites are 174 and 174 +/- 1 bp from the m16S rRNA genes in rrnE and rrnA, respectively. Initiation starts with CTP in both cases. We have also shown that in the present in vitro transcription system, guanosine tetraphosphate (ppGpp) inhibits the synthesis of full-sized RNAs from both start sites in each rRNA operon.

Base Sequence↗

Lack of GATC sites in the genome of Streptococcus bovis bacteriophage F4.

A strong bias against GATC sites was observed in the genome of phage F4, a lytic Streptococcus bovis bacteriophage. Only three GATC sites were found within the 60.4-kbp genome of this phage. The comparative lack of GATC sequences within the F4 genome was probably not due to dam methylation, as no modification within this site was detected using methylation-sensitive isoschizomer pair restriction endonuclease analysis. The short oligonucleotide composition of available S. bovis DNA sequences suggested the existence of an unknown mechanism for counterselection of GATC sites in S. bovis bacteriophages.

Animals↗

Theoretical analysis of mutation hotspots and their DNA sequence context specificity.

Mutation frequencies vary significantly along nucleotide sequences such that mutations often concentrate at certain positions called hotspots. Mutation hotspots in DNA reflect intrinsic properties of the mutation process, such as sequence specificity, that manifests itself at the level of interaction between mutagens, DNA, and the action of the repair and replication machineries. The hotspots might also reflect structural and functional features of the respective DNA sequences. When mutations in a gene are identified using a particular experimental system, resulting hotspots could reflect the properties of the gene product and the mutant selection scheme. Analysis of the nucleotide sequence context of hotspots can provide information on the molecular mechanisms of mutagenesis. However, the determinants of mutation frequency and specificity are complex, and there are many analytical methods for their study. Here we review computational approaches for analyzing mutation spectra (distribution of mutations along the target genes) that include many mutable (detectable) positions. The following methods are reviewed: derivation of a consensus sequence, application of regression approaches to correlate nucleotide sequence features with mutation frequency, mutation hotspot prediction, analysis of oligonucleotide composition of regions containing mutations, pairwise comparison of mutation spectra, analysis of multiple spectra, and analysis of "context-free" characteristics. The advantages and pitfalls of these methods are discussed and illustrated by examples from the literature. The most reliable analyses were obtained when several methods were combined and information from theoretical analysis and experimental observations was considered simultaneously. Simple, robust approaches should be used with small samples of mutations, whereas combinations of simple and complex approaches may be required for large samples. We discuss several well-documented studies where analysis of mutation spectra has substantially contributed to the current understanding of molecular mechanisms of mutagenesis. The nucleotide sequence context of mutational hotspots is a fingerprint of interactions between DNA and DNA repair, replication, and modification enzymes, and the analysis of hotspot context provides evidence of such interactions.

Animals↗

GLOBIC: a very fast microcomputer program for fingerprinting, characterization and comparison of long nucleotide sequences.

This paper describes the program GLOBIC, which compares, characterizes and fingerprints even 0.1 Mbase sequences in a few minutes with the aid of an IBM-AT microcomputer. Instead of the nucleotide sequences themselves, GLOBIC compares the local nucleotide or short oligonucleotide compositions. GLOBIC presents two-dimensional maps of contour lines depicting the similarity of two different sequences, a sequence compared to itself, to its complementary sequence or to a random sequence. A vocabulary is presented to translate the typical patterns appearing in the two-dimensional maps into their meanings as relationships between the compared sequences. The application of GLOBIC is demonstrated using several examples from the genomic nucleotide sequences of bacteriophage T7, adenovirus type-2 and Epstein-Barr virus.

Adenoviridae↗

Evolution of polypyrimidines in Drosophila.

We surveyed 101 different Drosophila species for the presence of a particular highly repetitive DNA sequence containing long tracts of polypyrimidine/polypurine DNA, first found in D. melanogaster. Out of 55 tested species in the melanogaster group, only the sibling species D. simulans and D. mauritiana, as well as one distant relative in the ananassae subgroup, D. varians, contained the same sequence. All four of these species have long pyrimidine tracts as shown by acid hydrolysis of labelled DNA. All four species have the same sequence, bu the amount of this polypyrimidine/polypurine DNA varies greatly. Four other species in the hydei subgroup were found to contain a polypyrimidine/polpurine sequence, with an oligonucleotide composition different from that of D. melanogaster. This polypyrimidine DNA varies from as much as 10% of the total DNA in D. nigrohydei, to as little as 0.4% in D. neohydei. The long pyrimidine tracts in the hydei subgroup are often more than a thousand nucleotides in length, representing exceedingly homogeneous repetitious sequences.--These results show a rapid but discontinuous pattern of evolution for polypyrimidine/polypurine DNA . These sequences are not species specific, yet closely related species have greatly different amounts of polypyrimidines. Drastic changes occur in the amounts of these satellite type DNA sequences, as if the sequence had no continuous selective advantage in evolution. The implications of these results with regard to the general function and evolution of satellite DNA are discussed.

Animals↗

A yeast transcription system for the 5S rRNA gene.

A cell-free extract of yeast nuclei that can specifically transcribe cloned yeast 5S rRNA genes has been developed. Optima for transcription of 5S rDNA were determined and conditions of extract preparation leading to reproducible activities and specificities established. The major in vitro product has the same size and oligonucleotide composition as in vivo 5S rRNA. The in vitro transcription extract does not transcribe yeast tRNA genes. The extract does increase the transcription of tRNA genes packaged in chromatin.

Cell Nucleus↗

How to interpret an anonymous bacterial genome: machine learning approach to gene identification.

In this report we address the problem of accurate statistical modeling of DNA sequences, either coding or noncoding, for a bacterial species whose genome (or a large portion) was sequenced but not yet characterized experimentally. Availability of these models is critical for successful solution of the genome annotation task by statistical methods of gene finding. We present the method, GeneMark-Genesis, which learns the parameters of Markov models of protein-coding and noncoding regions from anonymous bacterial genomic sequence. These models are subsequently used in the GeneMark and GeneMark.hmm gene-finding programs. Although there is basically one model of a noncoding region for a given genome, several models of protein-coding region are automatically obtained by GeneMark-Genesis. The diversity of protein-coding models reflects the diversity of oligonucleotide compositions, particularly the diversity of codon usage strategies observed in genes from one and the same genome. In the simplest and the most important case, there are just two gene models-typical and atypical ones. We show that the atypical model allows one to predict genes that escape identification by the typical model. Many genes predicted by the atypical model appear to be horizontally transferred genes. The early versions of GeneMark-Genesis were used for annotating the genomes of Methanoccocus jannaschii and Helicobacter pylori. We report the results of accuracy testing of the full-scale version of GeneMark-Genesis on 10 completely sequenced bacterial genomes. Interestingly, the GeneMark.hmm program that employed the typical and atypical models defined by GeneMark-Genesis was able to predict 683 new atypical genes with 176 of them confirmed by similarity search.

Algorithms↗

Ribosome-protected fragments from sindbis 42-S and 26-S RNAs.

Sindbis virus 42-S and 26-S RNAs labeled with 32P were purified from infected chick embryo fibroblasts. The RNA's were incubated in the presence of a wheat germ cell-free translating system under conditions that yielded 40-S and 80-S initiation complexes. After digestion with RNase A, ribosome-protected fragments were isolated by polyacrylamide gel electrophoresis and compared with respect to number, size, cap content and oligonucleotide composition. The two RNA species yielded several fragments of chain length about 35--40 nucleotides from 80S complexes and up to 60--65 nucleotides from 40-S complexes. The 5'-terminal capped sequence, m7 GpppA-U-G that is present in both Sindbis virus RNA's, was not retained in any of the ribosome-protected fragments. Fingerprint analyses indicated that the fragments derived from 40S and 80-S initiation complexes of each species of RNA were overlapping, but the fragments from 42-S and 26-S RNAs were unrelated. The complexity of the fingerprints were consistent with protection of a single, different initiation site in each Sindbis virus RNA.

Base Sequence↗

Genetic variation and host markers in the src gene of recovered avian sarcoma viruses.

The src genes of three recovered avian sarcoma viruses were compared by RNase T1 oligonucleotide fingerprinting and tryptic peptide analysis. In all three recovered avian sarcoma viruses the oligonucleotide composition of src was different and also distinct from that of the parental Schmidt-Ruppin strain of Rous sarcoma virus. This evidence for genetic variation src was strengthened by two dimensional peptide maps of the src gene products pp60src, translated in a reticulocyte lysate system in vitro. Numerous differences between the peptide patterns of the pp60src proteins produced by the parental and the recovered viruses were detected. No two src proteins were identical, while the tryptic peptide maps of the internal gag proteins synthesized by these viruses were indistinguishable. The src proteins of recovered avian sarcoma viruses also contained peptides that were absent from the src protein of parental Schmidt-Ruppin D virus but were found in the endogenous src protein of normal cells. We conclude that there is considerable genetic variation in the src gene of recovered avian sarcoma viruses and that these recovered src genes contain host cell-derived markers.

Animals↗

UV light-induced crosslinking of the complementary strands of plasmid pUC19 DNA restriction fragments.

Restriction fragments of pUC19 DNA were irradiated by various doses of UV light and analyzed by denaturing (alkaline) agarose gel electrophoresis. The irradiation generated retarded species whose mobility indicated two crosslinked DNA strands. Quantitative analysis of the experimental data provided an empirical equation relating the fraction of crosslinked DNA molecules to their length and to the dose of their irradiation by UV light. This equation can be used to predict the crosslinking behavior of pUC19-like DNA molecules whose primary structures do not much differ from a random nucleotide sequence. The amount of interstrand crosslinks increased with the (A+T) content of the pUC19 DNA fragments but the dependence was not clear-cut to indicate that oligonucleotide composition of DNA played a significant role as well.

DNA Restriction Enzymes↗

Genomic signatures: tracing the origin of retroelements at the nucleotide level.

We investigate the nucleotide sequences of 23 retroelements (4 mammalian retroviruses, 1 human, 3 yeast, 2 plant, and 13 invertebrate retrotransposons) in terms of their oligonucleotide composition in order to address the problem of relationship between retrotransposons and retroviruses, and the coadaptation of these retroelements to their host genomes. We have identified by computer analysis over-represented 3-through 6-mers in each sequence. Our results indicate retrotransposons are heterogeneous in contrast to retroviruses, suggesting different modes of evolution by slippage-like mechanisms. Moreover, we have calculated the Observed/Expected number ratio for each of the 256 tetramers and analysed the data using a multivariate approach. The tetramer composition of retroelement sequences appears to be influenced by host genomic factors like methylase activity.

Animals↗

Oligonucleotide sequence and composition determined by matrix-assisted laser desorption/ionization.

Molecular weight measurements of several oligonucleotides ranging in size from 12 to 60 bases were performed by matrix-assisted laser desorption/ionization with a time-of-flight mass spectrometer (MALDI-TOF). In each case, the mass accuracy was better than 0.1%. Sequences for two 12-base oligonucleotides and a 24-base oligonucleotide were determined using calf spleen phosphodiesterase to sequentially cleave from the 5' end. A MALDI-TOF spectrum of the digest mixture shortly after the addition of the enzyme produced a characteristic oligonucleotide ladder. Molecular ions in the mass spectrum corresponded to the products of enzymatic cleavage, and the mass differences between these peaks identified the individual nucleotides. The resolution and mass accuracy of MALDI-TOF were sufficient to unambiguously identify the individual nucleotides in the 12- and 24-base strands.

Base Sequence↗

Development and validation of a method for routine base composition analysis of phosphorothioate oligonucleotides.

A method for routine base composition determination of phosphorothioate oligonucleotides containing as many as 21 nucleobases has been developed and systematically evaluated in terms of factors contributing to assay precision and accuracy. Phosphorothioate internucleotide linkages were oxidized with a mixture of tetrahydrofuran-water-methylimidazole (16:4:1, v/v/v) which has been shown to be 97.3% effective. This step was followed by enzymolysis and HPLC quantitation of individual nucleobases. RSD for inter-day base composition analysis ranged from 1.1 to 1.3%, and inter-lot variation was 0.6-2.0%. Accuracy of the determined nucleobase ratio was independently confirmed through sequencing of the oxidized oligomer by matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF/MS).

Base Composition↗

Comparative DNA sequence features in two long Escherichia coli contigs.

The recent sequencing of two relatively long (approximately 100 kb) contigs of E.coli presents unique opportunities for investigating heterogeneity and genomic organization of the E.coli chromosome. We have evaluated a number of common and contrasting sequence features in the two new contigs with comparisons to all available E.coli sequences (> 1.6 Mb). Our analyses include assessments of: (i) counts and distributions of restriction sites, special oligonucleotides (e.g., Chi sites, Dam and Dcm methylase targets), and other marker arrays; (ii) significant distant and close direct and inverted repeat sequences; (iii) sequence similarities between the long contigs and other E.coli sequences; (iv) characterization and identification of rare and frequent oligonucleotides; (v) compositional biases in short oligonucleotides; and (vi) position-dependent fluctuations in sequence composition. The two contigs reveal a number of distinctive features, including: a cluster of five repeat/dyad elements with very regular spacings resembling a transcription attenuator in one of the contigs; REP elements, ERICs, and other long repeats; distinction of the Chi sequence as the most frequent oligonucleotide; regions of clustering, overdispersion, and regularity of certain restriction sites and short palindromes; and comparative domains of inhomogeneities in the two long contigs. These and other features are discussed in relation to the organization of the E.coli chromosome.

Base Sequence↗

Effects of oligonucleotide length and atomic composition on stimulation of the ATPase activity of translation initiation factor elF4A.

Eukaryotic translation initiation factor 4A (elF4A) has been proposed to use the energy of ATP hydrolysis to remove RNA structure in the 5' untranslated region (UTR) of mRNAs, helping the 43S ribosomal complex bind to an mRNA and scan to find the 5'-most AUG initiator codon. We have examined the effect of changing the atomic composition and length of single-stranded oligonucleotides on binding to elF4A and on stimulation of its ATPase activity once bound. Substitution of 2'-OH groups with 2'-H or 2'-OCH3 groups reduces ATPase stimulation at least 100-fold, to background levels, without significantly affecting oligonucleotide affinity. These effects suggest that 2'-OH groups participate in an elF4A conformational change that occurs subsequent to oligonucleotide binding and is required for ATPase stimulation. Replacing nonbridging oxygen atoms in phosphodiester linkages with sulfur atoms to make phosphorothioate linkages has no significant effect on stimulation, while substantially increasing affinity. Extending the length of an RNA oligonucleotide from 4 to approximately 15 nt gradually increases oligonucleotide affinity and ATPase stimulation. Consistent with this observation, the increase in affinity and stimulation provided by phosphorothioate linkages and 2'-OH groups is proportional to the number of these groups present within larger oligonucleotides. Further, changing the position of blocks of phosphorothioate linkages or 2'-OH groups within a larger oligonucleotide does not affect affinity and has only a small effect on stimulation. These observations suggest that numerous interactions between the oligonucleotide and elF4A contribute individually to binding and ATPase stimulation. Nevertheless, significant stimulation is observed with as few as four RNA residues. These properties may allow elF4A to operate within regions of 5' UTRs containing only short stretches of exposed single-stranded RNA. As stimulation increases when longer stretches of single-stranded RNA are available, it is possible that the accessibility of single-stranded RNA in a 5' UTR influences translation efficiency.

Adenosine Triphosphatases↗