Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Ovalbumin gene: evidence for a leader sequence in mRNA and DNA sequences at the exon-intron boundaries.

Selected regions of cloned EcoRI fragments of the chicken ovalbumin gene have been sequenced. The positions where the sequences coding for ovalbumin mRNA (ov-mRNA) are interrupted in the genome have been determined, and a previously unreported interruption in the DNA sequences coding for the 5' nontranslated region of the messenger has been discovered. Because directly repeated sequences are found at exon-intron boundaries, the nucleotide sequence alone cannot define unique excision-ligation points for the processing of a possible ov-mRNA precursor. However, the sequences in these boundary regions share common features; this leads to the proposal that there are, in fact, unique excision-ligation points common to all boundaries.

Animals↗

Nucleotide sequence at the junction between the coding region of the adenovirus 2 hexon messenger RNA and its leader sequence.

We have determined a 139-base-pair sequence of adenovirus 2 DNA that is located immediately leftwards of the cleavage site for endonuclease Sma I at position 51.1. The established sequence includes the hexon AUG initiator codon, located 75--77 nucleotides leftwards of this cleavage site, and codons for the first 26 amino acids of the hexon polypeptide. By the use of purified hexon mRNA as a template and separated strands of small restriction enzyme fragments as specific primers, the complete 5' noncoding region of the hexon mRNA was synthesized and part of its sequence was determined. The tripartite leader sequence of the hexon mRNA starts 39 nucleotides upstream from the initiator AUG triplet and the total length of the 5' noncoding part of the hexon mRNA was estimated to be 235 nucleotides. The sequence at the junction of the leader sequence permits the formation of secondary structures that may be of importance for the splicing reaction.

Adenoviruses, Human↗

Sequence-dependent energetics of the B-Z transition in supercoiled DNA containing nonalternating purine-pyrimidine sequences.

The likelihood that a given DNA sequence will adopt the Z conformation in negatively supercoiled DNA depends on the energy difference between the B form and the Z form for that sequence relative to other sequences in the same molecule. This energy can be viewed simply as a sum of energies for the nearest-neighbor interactions within the sequence plus the energy required to stabilize the B-Z boundaries. Knowledge of these energetic terms would be of value in predicting when sequences become left-handed in response to negative superhelicity. Here we present an approach that can be used to determine the free-energy changes associated with all the nearest-neighbor interactions that can occur in Z-DNA. Synthetic stretches of d(C-G)n containing one or two transversions were cloned into plasmids, and the extent of the B-Z transition as a function of negative superhelicity was determined for each insert by two-dimensional agarose gel electrophoresis. By subjecting the data to statistical mechanical analysis, it was possible to evaluate the energetic penalty resulting from each base-pair (bp) substitution. Guanine to cytosine transversions cost 2.4 kcal (1 cal = 4.18 J)/(mol X bp), whereas guanine to thymine transversions cost 3.4 kcal/(mol X bp), to stabilize in the Z conformation. We have used these numbers, along with energetic values determined by others for the B-Z transition, to predict that certain strictly nonalternating purine and pyrimidine sequences may adopt the Z form readily.

Base Sequence↗

DNA sequencing with Thermus aquaticus DNA polymerase and direct sequencing of polymerase chain reaction-amplified DNA.

The highly thermostable DNA polymerase from Thermus aquaticus (Taq) is ideal for both manual and automated DNA sequencing because it is fast, highly processive, has little or no 3'-exonuclease activity, and is active over a broad range of temperatures. Sequencing protocols are presented that produce readable extension products greater than 1000 bases having uniform band intensities. A combination of high reaction temperatures and the base analog 7-deaza-2'-deoxyguanosine was used to sequence through G + C-rich DNA and to resolve gel compressions. We modified the polymerase chain reaction (PCR) conditions for direct DNA sequencing of asymmetric PCR products without intermediate purification by using Taq DNA polymerase. The coupling of template preparation by asymmetric PCR and direct sequencing should facilitate automation for large-scale sequencing projects.

Autoanalysis↗

Phylogenetic affiliation of Aeromonas culicicola MTCC 3249(T) based on gyrB gene sequence and PCR-amplicon sequence analysis of cytolytic enterotoxin gene.

We determined the gyrB gene sequences of all 17 hybridizations groups of Aeromonas. Phylogenetic trees showing the evolutionary relatedness of gyrB and 16S rRNA genes in the type strains of Aeromonas were compared. Using this approach, we determined the phylogenetic position of Aeromonas culicicola MTCC 3249(T), isolated from midgut of Culex quinquefasciatus. In the gyrB based-analysis A. culicicola MTCC 3249(T) grouped with A. veronii whereas, it grouped with A. jandaei in the 16S rRNA based tree. The number of nucleotide differences in 16S rRNA sequences was less than found with the gyrB sequence data. Most of the observed nucleotide differences in the gyrB gene were synonymous. The Cophenetic Correlation Coefficient (CCC) for gyrB sequences was 0.87 indicating this gene to be a better molecular chronometer compared to 16S rRNA for delineation of Aeromonas species. This strain was found to be positive for the cytolytic enterotoxin gene. PCR-Amplicon Sequence Analysis (PCR-ASA) of this gene showed that the isolate is affiliated to type I and is potentially pathogenic. These PCR-ASA results agreed in part with the gyrB sequence results.

Aeromonas↗

CAST: an iterative algorithm for the complexity analysis of sequence tracts. Complexity analysis of sequence tracts.

MOTIVATION: Sensitive detection and masking of low-complexity regions in protein sequences. Filtered sequences can be used in sequence comparison without the risk of matching compositionally biased regions. The main advantage of the method over similar approaches is the selective masking of single residue types without affecting other, possibly important, regions. RESULTS: A novel algorithm for low-complexity region detection and selective masking. The algorithm is based on multiple-pass Smith-Waterman comparison of the query sequence against twenty homopolymers with infinite gap penalties. The output of the algorithm is both the masked query sequence for further analysis, e.g. database searches, as well as the regions of low complexity. The detection of low-complexity regions is highly specific for single residue types. It is shown that this approach is sufficient for masking database query sequences without generating false positives. The algorithm is benchmarked against widely available algorithms using the 210 genes of Plasmodium falciparum chromosome 2, a dataset known to contain a large number of low-complexity regions. AVAILABILITY: CAST (version 1.0) executable binaries are available to academic users free of charge under license. Web site entry point, server and additional material: http://www.ebi.ac.uk/research/cgg/services/cast/

Algorithms↗

Prediction of the coding sequences of unidentified human genes. XI. The complete sequences of 100 new cDNA clones from brain which code for large proteins in vitro.

In our series of projects for accumulating sequence information on the coding sequences of unidentified human genes, we have newly determined the sequences of 100 cDNA clones from a set of size-fractionated human brain cDNA libraries, and predicted the coding sequences of the corresponding genes, named KIAA0711 to KIAA0810. These cDNA clones were selected according to their coding potentials of large proteins (50 kDa and more) in vitro. The average sizes of the inserts and corresponding open reading frames were 4.3 kb and 2.6 kb (869 amino acid residues), respectively. Sequence analyses against the public databases indicated that the predicted coding sequences of 78 genes were similar to those of known genes, 64% of which (50 genes) were categorized as proteins functionally related to cell signaling/communication, cell structure/motility and nucleic acid management. As additional information concerning genes characterized in this study, the chromosomal locations of the clones were determined by using human-rodent hybrid panels and the expression profiles among 10 human tissues were examined by reverse transcription-coupled polymerase chain reaction which was substantially improved by enzyme-linked immunosorbent assay.

Brain Chemistry↗

A simple and efficient method to determine the terminal sequences of restriction fragments containing known sequences.

We present an improvement of the inverse PCR method for the determination of end sequences of restriction fragments containing unknown DNA sequences flanked by known segments. In this approach, a short "bridge" DNA is inserted during the self-ligation step of the inverse PCR technique. This bridge DNA acts as primer annealing sites for amplification and subsequent direct sequencing. Successive PCR amplifications enable selective amplification of the unknown sequences from a complex mixture. Unlike previously described methods, our method does not require special materials, such as synthetic adapters or biotinylated primers that must be prepared each time to adapt the target. Furthermore, no complex steps such as dephosphorylation or purification are needed. Our method can save time and reduce the cost of cloning unknown sequences; it is ideal for routine, rapid gene walking. We applied this method to a GC-rich bacterial genome and succeeded in determining the end sequences of a 4.5-kb fragment.

Chromosome Walking↗

Nucleotide sequence of an external transcribed spacer in Xenopus laevis rDNA: sequences flanking the 5' and 3' ends of 18S rRNA are non-complementary.

We have sequenced the external transcribed spacer (ETS) of a ribosomal transcription unit from Xenopus laevis, together with sections of the preceding non-transcribed spacer. Our analysis was carried out on the same cloned transcription unit as that from which the internal transcribed spacers (ITS) were previously sequenced. The ETS is approximately 712 nucleotides long and, like the ITS regions, is generally very rich in C plus G. Features of the sequence include an excess of oligo-C tracts over oligo-G tracts and a tract of 37 nucleotides consisting almost entirely of G and A residues. Parts of the sequence can give rise to stable internal secondary structures. However, in contrast to Escherichia coli, there is no potential for major base-pairing between the 18S flanking regions of the ETS and ITS. Further findings are that there are no initiation (ATG) codons in the ETS and that, as in other X.laevis rDNA cloned units, the sequence preceding the ETS is duplicated, with a few changes, in the "Bam island" sequence of the non-transcribed spacer.

Animals↗

Vertebrate histone genes: nucleotide sequence of a chicken H2A gene and regulatory flanking sequences.

The DNA sequence of a chicken genomal fragment containing a histone H2A gene has been determined. It contains extensive 5' and 3' flanking regions and encodes a protein identical in sequence to the histone H2A protein isolated from chicken erythrocytes. In the 5' flanking region, a possible "TATA box" and three possible "cap sites" can be recognised upstream from the initiation codon. To the 5' side of the "TATA box" is found an unusual sequence of 21 A's interrupted by a central G residue. It occupies the same relative position as the P. miliaris H2A gene-specific 5' dyad symmetry sequence and the "CCAAT box" seen in other eukaryotic polymerase II genes but is clearly different from both. A significant feature of the 3' non-coding region is the presence of a 23 base-pair sequence that is nearly identical to a conserved region found in sea urchin histone genes. The coding region is extremely GC rich, with strong selection for these bases in the third position of codons. Not a single coding triplet ends in U. No intervening sequences were found in this gene.

Animals↗

Preferential binding and structural distortion by Fe2+ at RGGG-containing DNA sequences correlates with enhanced oxidative cleavage at such sequences.

Certain DNA sequences are known to be unusually sensitive to nicking via the Fe2+-mediated Fenton reaction. Most notable are a purine nucleotide followed by three or more G residues, RGGG, and purine nucleotides flanking a TG combination, RTGR. Our laboratory previously demonstrated that nicking in the RGGG sequences occurs preferentially 5' to a G residue with the nicking probability decreasing from the 5' to 3'end of these sequences. Using 1H NMR to characterize Fe2+ binding within the duplex CGAGTTAGGGTAGC/GCTACCCTAACTCG and 7-deazaguanine-containing (Z) variants of it, we show that Fe2+ binds preferentially at the GGG sequence, most strongly towards its 5' end. Substitutions of individual guanines with Z indicate that the high affinity Fe2+ binding at AGGG involves two adjacent guanine N7 moieties. Binding is accompanied by large changes in specific imino, aromatic and methyl proton chemical shifts, indicating that a locally distorted structure forms at the binding site that affects the conformation of the two base pairs 3' to the GGG sequence. The binding of Fe2+ to RGGG contrasts with that previously observed for the RTGR sequence, which binds Fe2+ with negligible structural rearrangements.

Base Sequence↗

Statistical analysis of DNA sequencing data (1): accuracy test of DNA data by partial re-sequencing.

To qualify DNA data, we have developed a statistical method of deciding whether the DNA data has an acceptable accuracy in sequencing process. The method is to test the probability of sequencing errors, based on partial re-sequencing. The method was successfully applied to a yeast mitochondrial DNA which is previously sequenced (1). The analysis indicates that the entire sequence is very accurate although we found one base change error on the ND1 gene sequence data by a partial re-sampling. This method is applicable to any DNA data.

Chromosome Mapping↗

Amino acid sequence of Japanese horseshoe crab (Tachypleus tridentatus) coagulogen B chain: completion of the coagulogen sequence.

The complete amino acid sequence of the B chain derived from Tachypleus tridentatus coagulogen was determined. It consisted of a total of 129 amino acid residues with a NH2-terminal glycine and COOH-terminal phenylalanine. Sequence studies of the whole B chain and the fragments obtained from the digests with trypsin, alpha-chymotrypsin, thermolysin and Staphylococcal protease V8 showed the following sequence: (sequence; see text) These structural studies of the B chain and the previously established amino acid sequences of the A chain and peptide C derived from T. tridentatus coagulogen, now make it possible to complete the whole sequence of coagulogen consisting of 175 amino acid residues with the molecular weight of 19,723.

Amino Acid Sequence↗

Differential distribution of simple sequence repeats in eukaryotic genome sequences.

Complete chromosome/genome sequences available from humans, Drosophila melanogaster, Caenorhabditis elegans, Arabidopsis thaliana, and Saccharomyces cerevisiae were analyzed for the occurrence of mono-, di-, tri-, and tetranucleotide repeats. In all of the genomes studied, dinucleotide repeat stretches tended to be longer than other repeats. Additionally, tetranucleotide repeats in humans and trinucleotide repeats in Drosophila also seemed to be longer. Although the trends for different repeats are similar between different chromosomes within a genome, the density of repeats may vary between different chromosomes of the same species. The abundance or rarity of various di- and trinucleotide repeats in different genomes cannot be explained by nucleotide composition of a sequence or potential of repeated motifs to form alternative DNA structures. This suggests that in addition to nucleotide composition of repeat motifs, characteristic DNA replication/repair/recombination machinery might play an important role in the genesis of repeats. Moreover, analysis of complete genome coding DNA sequences of Drosophila, C. elegans, and yeast indicated that expansions of codon repeats corresponding to small hydrophilic amino acids are tolerated more, while strong selection pressures probably eliminate codon repeats encoding hydrophobic and basic amino acids. The locations and sequences of all of the repeat loci detected in genome sequences and coding DNA sequences are available at http://www.ncl-india.org/ssr and could be useful for further studies.

Animals↗

Improved serial analysis of V1 ribosomal sequence tags (SARST-V1) provides a rapid, comprehensive, sequence-based characterization of bacterial diversity and community composition.

Serial analysis of ribosomal sequence tags (SARST) is a recently developed technology that can generate large 16S rRNA gene (rrs) sequence data sets from microbiomes, but there are numerous enzymatic and purification steps required to construct the ribosomal sequence tag (RST) clone libraries. We report here an improved SARST method, which still targets the V1 hypervariable region of rrs genes, but reduces the number of enzymes, oligonucleotides, reagents, and technical steps needed to produce the RST clone libraries. The new method, hereafter referred to as SARST-V1, was used to examine the eubacterial diversity present in community DNA recovered from the microbiome resident in the ovine rumen. The 190 sequenced clones contained 1055 RSTs and no less than 236 unique phylotypes (based on > or = 95% sequence identity) that were assigned to eight different eubacterial phyla. Rarefaction and monomolecular curve analyses predicted that the complete RST clone library contains 99% of the 353 unique phylotypes predicted to exist in this microbiome. When compared with ribosomal intergenic spacer analysis (RISA) of the same community DNA sample, as well as a compilation of nine previously published conventional rrs clone libraries prepared from the same type of samples, the RST clone library provided a more comprehensive characterization of the eubacterial diversity present in rumen microbiomes. As such, SARST-V1 should be a useful tool applicable to comprehensive examination of diversity and composition in microbiomes and offers an affordable, sequence-based method for diversity analysis.

Bacteria↗

Quality assessment of DNA sequence data: autopsy of a mis-sequenced mtDNA population sample.

Published DNA data sets constitute a body of sequencing results resting in silico that are supposed to reflect the variation of (once) living cells. In cases where the DNA variation reported is suspected to be fraught with artefacts, an autopsy of the full body of data is needed to clarify the amount and causes of mis-sequencing. In this paper we elaborate on strategies that allow a clear-cut identification of the problems in severely flawed mtDNA data. This approach is applied, by way of example, to a data set of HVS-I sequences from the Caucasus, published by Nasidze & Stoneking in 2001. These data bear numerous ambiguous nucleotide positions and suffer from an even higher number of phantom mutations, indicating that severe biochemical problems adversely influenced those sequencing results at the time. Furthermore, systematic omission of sequences with a long C-stretch (incurred by a transition at position 16189) must have severely biased the data set. Since no complete correction of these data has appeared to date, this example of mis-sequencing necessitates circumstantial evidence that is bullet-proof.

Artifacts↗

Identification of the DNA sequence required for transposition immunity of the gamma delta sequence.

A plasmid containing the transposon gamma delta sequence was immune to further insertion of gamma delta (transposition immunity). Plasmids carrying a fragment containing either 0.2 kilobase pairs of the gamma end or 0.4 kilobase pairs of the delta end of the gamma delta sequence were immune, and other parts of the gamma delta sequence did not confer immunity. The terminal 38-base-pair (bp) sequence of the delta end of the gamma delta was sufficient to confer immunity, the 38-bp sequence of the gamma end conferred only moderate immunity, and the terminal 35-bp sequence, which was completely identical at both the gamma and delta ends, was insufficient to confer immunity.

Base Sequence↗

Friend strain of spleen focus-forming virus: a recombinant between mouse type C ecotropic viral sequences and sequences related to xenotropic virus.

The genome of the Friend strain of the spleen focus-forming virus (SFFV) has been analyzed by molecular hybridization. SFFV is composed of genetic sequences homologous to Friend type C helper virus (F-MuLV) and SFFV-specific sequences not present in F-MuLV. These SFFV-specific sequences are present in both the Friend and Rauscher strains of murine erythroleukemia virus. The SFFV-specific sequences are partially homologous to three separate strains of mouse xenotropic virus but not to several cloned mouse ecotropic viruses. Thus, the Friend strain of SFFV appears to be a recombinant between a portion of the F-MuLV genome and RNA sequences that are highly related to murine xenotropic viruses. The implications of the acquisition of the xenotropic virus-related sequences are discussed in relation to the leukemogenicity of SFFV, and a model for the pathogenicity of other murine leukemia-inducing viruses is proposed.

Animals↗