Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

A sample purification method for rugged and high-performance DNA sequencing by capillary electrophoresis using replaceable polymer solutions. B. Quantitative determination of the role of sample matrix components on sequencing analysis.

In the previous paper, a sample cleanup procedure for DNA sequencing reaction products was developed, in which template DNA was removed by ultrafiltration and the total concentration of salts (chloride and di- and deoxynucleotides) was decreased below 10 microM using gel filtration. In this paper, a quantitative study of the effects of these sample solution components on the injected amount and separation efficiency of the sequencing fragments in capillary electrophoresis is presented. The presence of chloride and deoxynucleotides in a total concentration above 10 microM in the sample solution significantly decreased the amount of DNA sequencing fragments injected into the capillary column. However, the separation efficiency was not affected upon increasing the amount of salt. On the other hand, in the presence of only 0.1 microgram of template in the sample (one-third of the lowest quantity recommended in cycle sequencing) and at very low chloride concentration (approximately 5 microM), the separation efficiency decreased by 70%, and the injected amount of DNA sequencing fragments was 40% lower compared to the sample cleaned by the new purification method. The deleterious effect of template DNA on the separation of sequencing fragments was suppressed in the presence of salt in a concentration above 100 microM in the sample solution. Separately, it was found that both the electric field strength and duration of injection affected the resolution of DNA sequencing fragments when the cleaned up sample solution was used. Separation efficiencies of 15 x 10(6) theoretical plates/m were achieved when the sample was loaded at low electric field, e.g., 25 V/cm for 80 s or less. The results demonstrate that the sample solution components (chloride, deoxynucleotides, template DNA) and injection conditions must be controlled to achieve high performance and rugged DNA sequencing analysis.

Chlorides↗

E. coli ribosomal RNA contains sequences homologous to insertion sequences IS1 and IS2.

The insertion sequence (IS) elements, IS1 and IS2, present in multiple copies in the Escherichia coli chromosome, are transposable genetic elements of known nucleotide sequence. These elements can modulate gene expression, but it is not known whether they normally function in genetic control. To determine whether IS elements could exert control through specific RNA transcripts, we hybridised lambda NNC1857 r14 (carrying IS1) and pBR322 (carrying a portion of IS2) to Northern blots of E. coli RNA. Regions of homology between the IS elements and ribosomal RNA were observed. Computer analysis of reported nucleotide sequences detected large segments of homology between the IS elements and both 23S and 16S rRNA. Additional homologous sequences in phi X174 and a leader region of a ribosomal protein gene cluster were also detected. The homologous sequence between IS2 and 16S rTNA is the same sequence in phi X174 DNA which codes for the ends of the E and D gene and the start of J. The partial IS sequences may represent silent evolutionary remnants or they could modulate the expression of genes carrying these sequences.

Base Sequence↗

Amino acid sequence studies on sheep liver fructose-bisphosphatase. II. The complete sequence.

The cyanogen bromide fragments of S-carboxymethylated fructose-bisphosphatase were purified. The amino acid sequences of the small fragments were determined by the dansyl-Edman method. The large fragments were subjected to proteolytic digestion to give smaller peptides more amenable for purification and sequencing by similar methods. Enzyme digests of the S-carboxymethylated enzyme gave overlap peptides containing the methionine residues. In conjunction with the amino acid sequence of the 60-residue N-terminal fragment previously determined on the S-peptide released by limited proteolysis with subtilisin the complete sequence of 336 residues was deduced. The sequence has been compared with the 335 residue sequence of pig kidney fructose-bisphosphatase and some areas of sequence for rabbit liver enzyme. The strong homology previously noted for the S-peptide sequence is maintained for the complete enzyme with only 34 changes in 336 residues when comparing the pig and sheep enzymes.

Amino Acid Sequence↗

Structure of Moloney murine leukemia viral DNA: nucleotide sequence of the 5' long terminal repeat and adjacent cellular sequences.

Some unintegrated and all integrated forms of murine leukemia viral DNA contain long terminal repeats (LTRs). The entire nucleotide sequence of the LTR and adjacent cellular sequences at the 5' end of a cloned integrated proviral DNA obtained from BALB/Mo mouse has been determined. It was compared to the nucleotide sequence of the LTR at the 3' end. The results indicate: (i) a direct 517-nucleotide repeat at the 5' and 3' termini; (ii) 145 nucleotides out of 517 nucleotides represent sequences between the 5'-CAP nucleotide and 3' end of the primer tRNA (strong-stop DNA); (iii) an 11-nucleotide inverted repeat is present at the ends of the 5'-LTR and a total of 17 out of 21 nucleotides at the termini are inverted repeats; (iv) sequences CAATAAAAG (at positions -24 to -31) and CAATAAAC (at positions +46 to +53) resembling the hypothetical DNA-dependent RNA polymerase II promoter site can be identified in the 5'-LTR; (v) the sequence GAAA appears to be repeated on both sides of the junction of viral and cellular sequences; and (vi) in analogy with the bacterial transposons, the presence of an inverted repeat sequence at the termini of 5'-LTR suggests that M-MLV also has the integration properties of a transposon.

Animals↗

The DNA sequence quality machine at IFOM: a simple Web-based tool for quantitative assessment of sequencing reactions.

DNA sequence quality is a factor of paramount importance in the world of modern genetic and genomics. Both the sequencing of Human Genome in the "post-draft" era [NHGRI Standard for quality of Human Genomic Sequences, Rev. 7 July (2002) where http://www.nhgri.nih.gov/Grant_info/Funding/ Statements/RFA/quality_standard.html is the HTTP address] and recent "high-throughput" approaches to genetic investigation such as SAGE [Velculescu, V.E., Zhang, L., Vogelstein, B. et al. (1995) "Serial analysis of gene expression", Science 270, 484-487] need a reliable, standardized measure of the quality of a sequencing reaction. The increasing importance of SNP studies also requires a stronger quality control on sequencing reactions by the final user. We propose here a simple, web-based tool for integrated sequence quality evaluation, high quality region quantitative value calculation and chromatogram display. This software is aimed at the small to medium DNA sequence laboratory or to the single researcher, interested in getting a quantitative measure of the sequence quality, browsing the chromatogram and checking the quality values base by base. The program is freely available from the IFOM bioinformatics web Server at http://bio.ifom-firc.it/Phred20/index.html.

Algorithms↗

Sequence analysis of the genome of the unicellular cyanobacterium Synechocystis sp. strain PCC6803. II. Sequence determination of the entire genome and assignment of potential protein-coding regions.

The sequence determination of the entire genome of the Synechocystis sp. strain PCC6803 was completed. The total length of the genome finally confirmed was 3,573,470 bp, including the previously reported sequence of 1,003,450 bp from map position 64% to 92% of the genome. The entire sequence was assembled from the sequences of the physical map-based contigs of cosmid clones and of lambda clones and long PCR products which were used for gap-filling. The accuracy of the sequence was guaranteed by analysis of both strands of DNA through the entire genome. The authenticity of the assembled sequence was supported by restriction analysis of long PCR products, which were directly amplified from the genomic DNA using the assembled sequence data. To predict the potential protein-coding regions, analysis of open reading frames (ORFs), analysis by the GeneMark program and similarity search to databases were performed. As a result, a total of 3,168 potential protein genes were assigned on the genome, in which 145 (4.6%) were identical to reported genes and 1,257 (39.6%) and 340 (10.8%) showed similarity to reported and hypothetical genes, respectively. The remaining 1,426 (45.0%) had no apparent similarity to any genes in databases. Among the potential protein genes assigned, 128 were related to the genes participating in photosynthetic reactions. The sum of the sequences coding for potential protein genes occupies 87% of the genome length. By adding rRNA and tRNA genes, therefore, the genome has a very compact arrangement of protein- and RNA-coding regions. A notable feature on the gene organization of the genome was that 99 ORFs, which showed similarity to transposase genes and could be classified into 6 groups, were found spread all over the genome, and at least 26 of them appeared to remain intact. The result implies that rearrangement of the genome occurred frequently during and after establishment of this species.

Bacterial Proteins↗

A convenient method for locating sets of related short sequences in DNA sequences of any length.

In investigating sequence variants in a family of highly repeated rat DNA, we needed to search the consensus sequence of the repeat unit of this family for short sequences which would become, with one base change, recognition sites for various restriction endonucleases. To do this, we have designed a pair of programs to search DNA sequences of any length for sets of related short sequences, allowing user-specified mismatches in the short sequence. Since putative regulatory regions are generally short sequences, these programs are also useful for locating all possible versions of such sequences in any given DNA. We describe the programs, and present results of searches using the programs.

Base Sequence↗

The complete nucleotide sequence of mouse immunoglobin gamma 2a gene and evolution of heavy chain genes: further evidence for intervening sequence-mediated domain transfer.

We have determined the complete nucleotide sequence (1990 base pairs) of mouse immunoglobulin gamma 2a gene, and compared it with the sequences of other gamma subclass genes so far sequenced, i.e. gamma 1 and gamma 2b genes. Divergence of the nucleotide sequence between a compared pair of the gamma genes varies extensively among different segments of the gene. For example, comparison of the gamma 2a and gamma 2b genes has revealed a remarkable homology in a long continuous segment (about 900 bases) that covers from the 3' portion of the first intervening sequence to the third intervening sequence. However, there is no particular segment of the gamma gene that is conserved universally among the three gamma genes. These findings suggest that, during their evolution, segments of the gamma genes had been scrambled between different subclass genes through recombinations within intervening sequences, thus providing further evidence for the intervening sequence-mediated domain transfer hypothesis. We have discussed several possible phylogenic trees which can explain the difference of divergence in various segments of the gamma genes.

Animals↗

DeNovoID: a web-based tool for identifying peptides from sequence and mass tags deduced from de novo peptide sequencing by mass spectroscopy.

One of the core activities of high-throughput proteomics is the identification of peptides from mass spectra. Some peptides can be identified using spectral matching programs like Sequest or Mascot, but many spectra do not produce high quality database matches. De novo peptide sequencing is an approach to determine partial peptide sequences for some of the unidentified spectra. A drawback of de novo peptide sequencing is that it produces a series of ordered and disordered sequence tags and mass tags rather than a complete, non-degenerate peptide amino acid sequence. This incomplete data is difficult to use in conventional search programs such as BLAST or FASTA. DeNovoID is a program that has been specifically designed to use degenerate amino acid sequence and mass data derived from MS experiments to search a peptide database. Since the algorithm employed depends on the amino acid composition of the peptide and not its sequence, DeNovoID does not have to consider all possible sequences, but rather a smaller number of compositions consistent with a spectrum. DeNovoID also uses a geometric indexing scheme that reduces the number of calculations required to determine the best peptide match in the database. DeNovoID is available at http://proteomics.mcw.edu/denovoid.

Algorithms↗

Generation of 919 expressed sequence tags from immature flower buds and gene expression analysis using expressed sequence tags in the model plant Lotus japonicus.

Lotus japonicus has received increased attention as a potential model legume plant. In order to study gene expression in reproductive organs and to identify genes that play a crucial function in sexual reproduction, we constructed a cDNA library from immature flower buds containing anthers at the stage of developing tapetum cells in L. japonicus, and characterized 919 expressed sequence tags (ESTs) randomly selected from a cDNA library of the immature flower buds. The 919 ESTs analyzed were clustered into 821 non-redundant EST groups. As a result of a database search, 436 groups (53%) out of the 821 groups showed sequence similarity to genes registered in the public database. Out of these 436 groups, 109 groups showed similarity to genes encoding hypothetical proteins whose function had not yet been estimated. Three hundred eighty five groups (47%) showed no significant homology to known sequences and were classified as novel sequences. A comparison of 821 non-redundant EST sequences and EST sequences derived from the whole plant L. japonicus revealed that 474 EST sequences derived from immature flower buds were not found in the EST sequences of the whole plant. In order to confirm the expression pattern of potential reproductive-organ specific EST clones, nine clones, which were not matched to ESTs derived from the whole plant, were selected, and RT-PCR analysis was performed on these clones. As a result of RT-PCR, we found two novel anther specific clones. One clone was homologous to a gene encoding human cleft lip and palate associated transmembrane protein (CLPTM1) like protein, and the other clone did not show a significant similarity to any genes deposited in the public database. These results indicate that ESTs analyzed here represent a valuable resource for finding reproductive-organ specific genes in Lotus japonicus.

Expressed Sequence Tags↗

Polymorphic class II sequences linked to the rat major histocompatibility complex (RT1) homologous to human DR and DQ sequences.

Until recently, the analysis of Class II genes linked to the rat major histocompatibility complex, RT1, has been confined to serologic and electrophoretic analysis of their gene products. To obtain a more definitive estimate of the number and relative polymorphism of RT1 Class II sequences, we performed Southern blot analysis of rat genomic DNA employing human cDNA probes specific for Class II heavy and light chain genes. Southern blots of EcoRI and BamHI digests of genomic DNA from ten inbred strains, expressing eight RT1 haplotypes, were hybridized with the human DQ beta or DR beta cDNA that are homologous to Class II light chain sequences. Four to eight bands were observed to hybridize with the light chain cDNA: band sizes ranged from 2.5 to 28 kb. Restriction fragment patterns were polymorphic; the only identical patterns observed were those associated with RT1 haplotypes with identical RT1.B regions. The number and size of bands hybridizing with DQ beta and DR beta suggested a minimum of four light chain sequences in each haplotype. Southern blots of BamHI and EcoRI digests of genomic DNA from the same strains were hybridized with a DR alpha cDNA that is homologous to Class II heavy chain sequences. All RT1 haplotypes expressed either a 10.0-kb or 13.0-kb band when digested with BamHI, and either a 17-kb or 3.7-kb band when digested with EcoRI. Considerably less polymorphism was detected with the DR alpha probe; this observation is consistent with previously reported limited protein polymorphism of the rat equivalent of the I-E alpha subunit. The size and number of bands hybridizing with the DR alpha probe suggests a minimum of two heavy chain sequences. These observations suggest that the RT1 complex includes more Class II sequences than have been observed in serologic and electrophoretic analyses of Class II gene products. Furthermore, the level of polymorphism of RT1 Class II sequences appears to be comparable with mouse and human Class II sequences.

Animals↗

The Euglena gracilis chloroplast ribulose-1,5-bisphosphate carboxylase gene. I. Complete DNA sequence and analysis of the nine intervening sequences.

The nucleotide sequence of 6225 base pairs (bp) of Euglena gracilis chloroplast DNA including the complete DNA sequence of the chloroplast-encoded ribulose-1,5-bisphosphate carboxylase large subunit gene along with the flanking DNA sequences is presented. The gene is greater than 5.5 kilobase pairs in length and is organized as 10 exons coding for 475 amino acids, separated by 9 introns. The exons range in size from 45 to 438 bp, while the introns range in size from 382 to 568 bp. The introns have highly conserved boundary sequences with the consensus, 5'-N GTGTGGATTT...(intron)...TTAATTTTAT N-3'. The introns are 82-85 mol% AT, with a pronounced T greater than A greater than G greater than C base bias in the RNA-like strand. They do not appear to encode any polypeptides. In addition, the introns have a conserved sequence 30-50 bp from their 3'-ends with the consensus, 5'-TACAGTTTGAAAATGA-3'. The 5'-TACA sequence bears some homology to the 5'-end of the TACTAACA sequence found in a similar location in yeast nuclear mRNA introns. The conserved sequences of the Euglena rbcL introns may be indicative of a splicing mechanism similar to that of eucaryotic nuclear mRNA introns and group II mitochondrial introns.

Base Composition↗

[Single copy and repetitive nucleotide sequences in the genome of Echinodermata. II. Nucleotide sequence divergence of DNA of echinoderms].

The degree of divergence of short and long repetitive DNA sequences and single copy DNA of five Echinodermata species (sea urchins, starfish, sea-cucumber) was studied by the method of molecular hybridization. Different fractions of 3H-DNA of the sea urchin Strongylocentrotus intermedius were hybridized with the DNA of other species. Thermal stability of the hybridized DNA molecules was determined. The results obtained suggest that short repetitive sequences were most conservative during the evolution of Echinodermata. Single copy DNA fractions of closely related sea urchin species (S. intermedius and S. nudus) have more homologous sequences than long repetitive DNA fractions of the same species. The DNA of evolutionary distant species (sea urchin and starfish) have more homologous long repetitive sequences than the single copy ones. All DNA fractions of S. intermedius have sequences hybridized with the DNA of all other species studied: short repetitive sequences--55%, long repetitive sequences--20%, single copy sequences--12%.

Animals↗

Nucleotide sequence analysis of an 8887 bp region of the left arm of yeast chromosome XIV, encompassing the centromere sequence.

The nucleotide sequencing of 8887 bp of the left arm of chromosome XIV is described. The sequence includes the centromeric region. Both strands were sequenced with an average redundancy of 5.09 per base pair. The overall G+C content is 37.3% (39.2% for putative coding regions versus 32.5% for non-coding regions). Six open reading frames (ORFs) greater than 100 amino acids were detected, all of which are completely confined to the 8.9 kbp region. Codon frequencies of the six ORFs agree with codon usage in Saccharomyces cerevisiae and all show the characteristics of low-level expressed genes. Comparison of the translated sequences with protein sequences in data bases suggests the presence of two ORFs (N2014 and N2007) encoding ribosomal proteins, the latter of which is the previously sequenced MRP7 gene. Another ORF (N2012) could encode a membrane-associated protein since it contains secretory signal sequence and two presumed transmembrane helices. This protein might be involved in mitochondrial energy transfer. ORF N2016 is immediately adjacent to the centromere, suggesting that it corresponds to the SPO1 gene, which is very tightly linked to the centromere at the left arm side of chromosome XIV (Mortimer et al., 1989).

Base Composition↗

A HSV-1 variant (1720) generates four equimolar isomers despite a 9200-bp deletion from TRL and sequences between 9200 np and 97,000 np in inverted orientation being covalently bound to sequences 94,000-126,372 np.

The genome structure of a spontaneously generated HSV-1 strain 17 variant, 1720, has been determined by restriction endonuclease and Southern blot analysis. The short segment of 1720 is unaltered compared to the parental strain 17 genome, whereas the long segment is extensively rearranged. Almost all of TRL (approximately 9.2 kb) has been deleted and consequently IRL is converted into unique sequence. Sequences from approximately 9200 nucleotide position (np) to 97,000 np are present in inverted orientation, covalently bound to sequences in the prototype orientation from approximately 94,000 np to the L/S junction at 126,372 np. Thus, sequences from 94,000 np to 97,000 np are now diploid, with one copy in the normal orientation and location, and the other at the long terminus as an inverted repeat; no inversion of the intervening unique sequences occurs about this novel inverted repeat. In contrast, normal inversions of the long and short segments occur to give four equimolar genomic isomers, indicating that the novel long terminus has gained an "a" sequence. The duplication of sequences between 94,000 np and 97,000 np results in a genome containing two copies of UL43 and one complete and one partial copy each of genes UL42 and UL44 encoding the 65 kD DNA-binding protein and glycoprotein C, respectively. The variant has been shown to grow normally in vitro following high multiplicity infection.

Base Sequence↗

Frequent deletions and sequence aberrations at the transgene junctions of transgenic mice carrying the papillomavirus regulatory and the SV40 TAg gene sequences.

Exogenous DNA microinjected into one-cell mouse zygotes either integrates into the host genome within a short time span, or is rapidly degraded. On integration, a transgene sequence is frequently reiterated. In this report, we describe the enzymatic amplification analysis of transgene junctions of 12 transgenic mice carrying different copy numbers of the same transgene with dissimilar ends. The transgene was composed of the regulatory sequence of the type 18 human papillomavirus linked to the TAg gene of the SV40 virus. Nucleotide sequences of 36 of these junctions were also determined. Deletions were found in 33 (91.7%) of the junctions analysed. At the crossover regions, 55.6% contained short overlapping sequences of one to six nucleotides. Insertions of 2-6 extraneous nucleotides were also found in 8.3% of the transgene junctions. Within a 10-nucleotide sequence on both sides of the transgene junctions, topoisomerase I (topo I) cleavage sites, runs of homogeneous purines or pyrimidiens, alternating purine-pyrimidine tracks and (A-T)-rich sequences were found frequently. Stringent control experiments were also performed to ascertain that the observations made were not artefacts resulting from the polymerase chain reaction. Our data therefore indicate that damage had occurred quite frequently and extensively in our transgene construct. Such transgene damage may also occur to various extents in mice carrying other transgenes. Primary structure of the nucleotide sequences of the injected DNA seems to influence the process of transgene reiteration and aberration.

Animals↗

Local repeat sequence organization of an intergenic spacer in the chloroplast genome of Chlamydomonas reinhardtii leads to DNA expansion and sequence scrambling: a complex mode of "copy-choice replication"?

Parent-specific, randomly amplified polymorphic DNA (RAPD) markers were obtained from total genomic DNA of Chlamydomonas reinhardtii. Such parent-specific RAPD bands (genomic fingerprints) segregated uniparentally (through mt+) in a cross between a pair of polymorphic interfertile strains of Chlamydomonas (C. reinhardtii and C. minnesotti), suggesting that they originated from the chloroplast genome. Southern analysis mapped the RAPD-markers to the chloroplast genome. One of the RAPD-markers, "P2" (1.6 kb) was cloned, sequenced and was fine mapped to the 3 kb region encompassing 3' end of 23S, full 5S and intergenic region between 5S and psbA. This region seems divergent enough between the two parents, such that a specific PCR designed for a parental specific chloroplast sequence within this region, amplified a marker in that parent only and not in the other, indicating the utility of RAPD-scan for locating the genomic regions of sequence divergence. Remarkably, the RAPD-product, "P2" seems to have originated from a PCR-amplification of a much smaller (about 600 bp), but highly repeat-rich (direct and inverted) domain of the 3 kb region in a manner that yielded no linear sequence alignment with its own template sequence. The amplification yielded the same uniquely "sequence-scrambled" product, whether the template used for PCR was total cellular DNA, chloroplast DNA or a plasmid clone DNA corresponding to that region. The PCR product, a "unique" new sequence, had lost the repetitive organization of the template genome where it had originated from and perhaps represented a "complex path" of copy-choice replication.

Animals↗

Codon-level analysis of histone primary sequence: evidence of a repeat tetrapeptide origin and later inclusion of transcribed sequence.

This work is directed to the question of protein sequence conservation. By reference to the genetic code the aminoacyl sequence of histones H2A, H4, H3, H2B and H1 (fragment) were rewritten as the codon sequences. The N-terminal regions were set aside on the grounds of different composition and sequence. The remainder of the molecule could be referred to simple repeat-tetrapeptide proteins by codon composition (high Gxy, low xGy content) and by sequence. Random segments of three to six residues occur characterized by composition and sequence as originating from the complimentary DNA strand, i.e. as codon "transcript". Ancestral features are probably best seen in H3, point mutations appear to be more extensive in H2B and H1. Segments in reverse order in H2A and in "transcript" in H4 distinguish these two from the other three histones. There is a tenuous possibility the N-terminals also originated as repeat-tetrapeptide now intensively modified. At codon-level the 50S ribosomal protein (L7/L12) of E. coli has features in common with histones (including a palindrome-containing N-terminal). It has the composition and sequence of a well-conserved tetrapeptide-repeat strand (statistical support). If interpretations made here are substantially correct, the 50S r-protein illustrates a significant stage in evolution of histone codon strands.

Amino Acid Sequence↗