Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Estimation of errors in "raw" DNA sequences: a validation study.

As DNA sequencing is performed more and more in a mass-production-like manner, efficient quality control measures become increasingly important for process control, but so also does the ability to compare different methods and projects. One of the fundamental quality measures in sequencing projects is the position-specific error probability at all bases in each individual sequence. Accurate prediction of base-specific error rates from "raw" sequence data would allow immediate quality control as well as benchmarking different methods and projects while avoiding the inefficiencies and time delays associated with resequencing and assessments after "finishing" a sequence. The program PHRED provides base-specific quality scores that are logarythmically related to error probabilities. This study assessed the accuracy of PHRED's error-rate prediction by analyzing sequencing projects from six different large-scale sequencing laboratories. All projects used four-color fluorescent sequencing, but the sequencing methods used varied widely between the different projects. The results indicate that the error-rate predictions such as those given by PHRED can be highly accurate for a large variety of different sequencing methods as well as over a wide range of sequence quality.

Base Sequence↗

Long-range and highly sensitive DNase I footprinting by an automated infrared DNA sequencer.

We have shown that an automated DNA sequencer is applicable to fluorescence-based detection of fragments in DNase I footprinting. We demonstrated the potential of long-range and highly sensitive DNase I footprinting taking advantage of an infrared-fluorescence automated DNA sequencer. Footprints of human transcription factor SpI were reproducibly detected ranging approximately between 100 and 750 bp on both strands of an 895-bp DNA fragment in a single electrophoresis run. We developed techniques in data collection and subsequent image processing for highly sensitive detection. Less than 0.1 footprinting unit (fpu: approximately 4.5 ng) of SpI was detected using 3.1 fmol of a 512-bp DNA fragment. This is greater than 10-fold increase in sensitivity over what has previously been reported by visible dye fluorescence DNA sequencers. This method will be very important in systematic analysis of transcription regulatory regions and in large-scale analysis of the transcription process.

DNA↗

Invariants of DNA sequences based on 2DD-curves.

The leading eigenvalue of the matrix associated with a DNA sequence as a important invariant is effectively used in analysis of DNA sequences. Here, we propose a new invariant base on the 2DD-Curves of DNA sequences which is simple for calculation. We can use it as an alternative invariant to characterize the DNA sequence. The utility of the new parameter is illustrated on the DNA sequences of 11 species.

Animals↗

Phylogenetic position of Taylorella equigenitalis determined by analysis of amplified 16S ribosomal DNA sequences.

The 16S ribosomal DNA sequence of Taylorella equigenitalis (formerly Haemophilus equigenitalis), the causative organism of contagious equine metritis, was determined. A phylogenetic analysis of this sequence revealed a phylogenetic position of T. equigenitalis in the beta subclass of the class Proteobacteria apart from the position of Haemophilus influenzae, which belongs to the gamma subclass of Proteobacteria. A close phylogenetic relationship among T. equigenitalis, Alcaligenes xylosoxidans, and Bordetella bronchiseptica was detected; Spirillum volutans and Chromobacterium fluviatile (Iodobacter fluviatile) were in the same group but slightly removed. This relationship is surprising in view of the considerable differences in the G + C contents of the genomes of these bacteria.

Alcaligenes↗

A computer simulation analysis of the accuracy of partial genome sequencing and restriction fragment analysis in estimating genetic relationships: an application to papillomavirus DNA sequences.

BACKGROUND: Determination of genetic relatedness among microorganisms provides information necessary for making inferences regarding phylogeny. However, there is little information available on how well the genetic relationships inferred from different genotyping methods agree with true genetic relationships. In this report, two genotyping methods - restriction fragment analysis (RFA) and partial genome DNA sequencing - were each compared to complete DNA sequencing as the definitive standard for classification. RESULTS: Using the Genbank database, 16 different types or subtypes of papillomavirus were selected as study samples, because numerous complete genome sequences were available. RFA was achieved by computer-simulated digestion. The genetic similarity of samples, based on RFA, was determined from the proportion of fragments that matched in size. DNA sequences of four specific genes (E1, E6, E7, and L1), representing partial genome sequencing, were also selected for comparison to complete genome sequencing. Laboratory error was not taken into account. Evaluation of the correlation between genetic similarity matrices (Mantel's r) and comparisons of the structure of the derived dendrograms (partition metric) indicated that partial genome sequencing (for single genes) had higher agreement with complete genome sequencing, achieving a maximum Mantel's r = 0.97 and a minimum partition metric = 10. RFA had lower agreement, with a maximum Mantel's r = 0.60 and a minimum partition metric = 18. CONCLUSIONS: This simulation indicated that for smaller genomes, such as papillomavirus, partial genome sequencing is superior to restriction fragment analysis in representing genetic relatedness among isolates. The generalizability of these results to larger genomes, as well as the impact of laboratory error, remains to be demonstrated.

Animals↗

DNA sequence recognition by bispyrazinonaphthalimides antitumor agents.

Bifunctional DNA intercalating agents have long attracted considerable attention as anticancer agents. One of the lead compounds in this category is the dimeric antitumor drug elinafide, composed of two tricyclic naphthalimide chromophores separated by an aminoalkyl linker chain optimally designed to permit bisintercalation of the drug into DNA. In an effort to optimize the DNA recognition capacity, different series of elinafide analogues have been prepared by extending the surface of the planar drug chromophore which is important for DNA sequence recognition. We report here a detailed investigation of the DNA sequence preference of three tetracyclic monomeric or dimeric pyrazinonaphthalimide derivatives. Melting temperature measurements and surface plasmon resonance (SPR) studies indicate that the dimerization of the tetracyclic planar chromophore considerably augments the affinity of the drug for DNA, polynucleotides, or hairpin oligonucleotides and promotes selective interaction with G.C sites. The (CH(2))(2)NH(CH(2))(3)NH(CH(2))(2) connector stabilizes the drug-DNA complexes. The methylation of the two nitrogen atoms of this linker chain reduces the binding affinity and increases the dissociation rates of the drug-DNA complexes by a factor of 10. DNase I footprinting experiments were used to investigate the sequence selectivity of the drugs, demonstrating highly preferential binding to G.C-rich sequences. It also served to select a high-affinity site encompassing the sequence 5'-GACGGCCAG which was then introduced into a biotin-labeled hairpin oligonucleotide to accurately measure the binding parameters by SPR. The affinity constant of the unmethylated dimer for this sequence is 500 times higher than that of the monomer compound and approximately 10 times higher than that of the methylated dimer. The DNA groove accessibility was also probed with three related oligonucleotides carrying G --> c(7)G, G --> I, and C --> M substitutions. The level of drug binding to the two hairpin oligonucleotides containing 7-deazaguanine (c(7)G) or 5-methylcytosine (M) residues is unchanged or only slightly reduced compared to that of the unmodified target. In contrast, incorporation of inosine (I) residues considerably decreases the extent of drug binding or even abolishes the interaction as is the case with the monomer. The pyrazinonaphthalimide derivatives are thus much more sensitive to the deletion of the exocyclic guanine 2-amino group exposed in the minor groove of the duplex than to the modification of the major groove elements. The complementary SPR footprinting methodology combining site selection and quantitative DNA affinity analysis constitutes a reliable method for dissecting the DNA sequence selectivity profile of reversible DNA binding small molecules.

Adenine↗

Nucleosome DNA sequence pattern revealed by multiple alignment of experimentally mapped sequences.

Five different algorithms have been applied for detecting DNA sequence pattern hidden in 204 DNA sequences collected from the literature which are experimentally found to be involved in nucleosome formation. Each algorithm was used to perform a multiple alignment of the nucleosome DNA sequences within the window 145 nt, the size of a nucleosome core DNA. From these alignments five pairs of AA and TT dinucleotide positional frequency distributions have been computed. The frequency profiles calculated by different algorithms are rather different due to substantial noise. They, however, share several important features. Both AA and TT dinucleotide positional frequencies display periodicity with the period of 10.3(+/- 0.2) bases. TT dinucleotides appear to be distributed symmetrically relative to AA dinucleotides of the same DNA strand, with the center of symmetry at the midpoint of the nucleosome core DNA. The phase shift between the AA and TT patterns is about 6 bp. Superposition of the five pairs of the AA (TT) positional frequency profiles has produced the refined pattern, with the above features well pronounced. An interesting novel feature of the pattern is an absence of central peaks in the periodical AA and TT distributions. This may indicate that the central section of nucleosome DNA, 15 bp around the dyad axis of the nucleosome, is not bent. Positional distributions of other dinucleotides were not found in this study to be as informative as the ones for AA and TT.

Algorithms↗

Resuspension of DNA sequencing reaction products in agarose increases sequence quality on an automated sequencer.

We are investigating approaches to increase DNA sequencing quality. Since a majorfactor in sequence generation is the cost of reagents and sample preparations, we have developed and optimized methods to sequence directly plasmid DNA isolated from alkaline lysis preparations. These methods remove the costly PCR and post-sequencing purification steps but can result in low sequence quality when using standard resuspension protocols on some sequencing platforms. This work outlines a simple, robust, and inexpensive resuspension protocol for DNA sequencing to correct this shortcoming. Resuspending the sequenced products in agarose before electrophoresis results in a substantial and reproducible increase in sequence quality and read length over resuspension in deionized water and has allowed us to use the aforementioned sample preparation methods to cut considerably the overall sequencing costs without sacrificing sequence quality. We demonstrate that resuspension of unpurified sequence products generated from template DNA isolated by a modified alkaline lysis technique in low concentrations of agarose yields a 384% improvement in sequence quality compared to resuspension in deionized water. Utilizing this protocol, we have produced more than 74,000 high-quality, long-read-length sequences from plasmid DNA template on the MegaBACET 1000 platform.

Cell Fractionation↗

A new repetitive DNA sequence from Trypanosoma cruzi.

Tandemly repeated DNA sequences are found in the genome of higher eukaryotes, and have also been demonstrated in Trypanosoma cruzi. Repeated DNA sequences are potentially useful for the diagnostic detection of T. cruzi (A. Gonzales et al., 1984, Proc. Natl. Acad. Sci. USA, 81:3356-3360). We have isolated two clones from a genomic library of T. cruzi (Y strain) that contain, in one clone a family of at least seven copies of a repetitive sequence of approximately 600 base pairs, and in the other an independent copy of the same sequence. One copy of the repetition (HSP) and the independent clone (HCR) were sequenced by the Sanger procedure (Fig.). This sequence hybridized to four strains of T. cruzi tested and did not hybridize to eleven species of trypanosomatids from five different Genera, being a good candidate for diagnostic assays.

Amino Acid Sequence↗

Multiplex DNA sequencing.

The increasing demand for DNA sequences can be met by replacement of each DNA sample in a device with a mixture of N samples so that the normal throughput is increased by a factor of N. Such a method is described. In order to separate the sequence information at the end of the processing, the DNA molecules of interest are ligated to a set of oligonucleotide "tags" at the beginning. The tagged DNA molecules are pooled, amplified, and chemically fragmented in 96-well plates. The resulting reaction products are fractionated by size on sequencing gels and transferred to nylon membranes. These membranes are then probed as many times as there are types of tags in the original pools, producing, in each cycle of probing, autoradiographs similar to those from standard DNA sequencing methods. Thus, each reaction and gel yields a quantity of data equivalent to that obtained from conventional reactions and gels multiplied by the number of probes used. To date, even after 50 successive probings, the original signal strength and the image quality are retained, an indication that the upper limit for the number of reprobings may be considerably higher.

Automation↗

Complicated organization of a single repeated DNA sequence in the chicken genome is revealed by cloning.

The structural organization of a family of repeated DNA sequences in the chicken genome has been determined by hybridization of a cloned repeated DNA sequence to Southern blots of total DNA. The length of the cloned DNA fragment is 3600 nucleotide pairs. This fragment consists principally, if not entirely, of a single repeated DNA sequence occurring only once within the cloned fragment. In the chicken genome, the family of repeated DNA sequences homologous to the cloned sequence has a limited number of alternative forms. Some of the restriction fragments of total DNA to which the cloned sequence hybridizes correspond to those expected from the location of restriction endonuclease cleavage sites within the cloned sequence. There are also a limited number of other genomic restriction fragments, each present in multiple copies, to which the cloned sequence hybridizes but which do not relate in any obvious way to the length of the cloned sequence. These various restriction fragments differ from one another in that they appear to be present in unequal amounts in total DNA, and many of them do not contain the entire cloned sequence. This study provides some new information about the structure of repeated DNA sequences in the chicken genome. The copies of a repeated DNA sequence may differ from one another both by minor variations in nucleotide sequence (divergence) and in more substantial ways as would be expected to arise from processes such as insertion, deletion, and translocation. In addition to this description of a single cloned repeated DNA sequence from the chicken genome, this paper reports the cloning of more than 100 different restriction fragments of chicken DNA, each of which contains one or more repeated DNA sequences.

Animals↗

DNA sequencing using 96-capillary array electrophoresis.

Practical DNA sequencing in a rugged capillary array electrophoresis system coupled directly to 96-well microtiter plates is demonstrated. A CCD detector was used to monitor all capillaries simultaneously with laser-induced fluorescence at 1.75 frames per second. The reconstructed electropherograms show good signal-to-noise ratios and resolution for the entire capillary array. The system used standard dye labeling and image splitting to obtain fluorescence intensities in two wavelength regions to allow calling up to 410 bases for the DNA sequence. The use of a replaceable poly(ethylene oxide) matrix and a protective poly(vinylpyrrolidone) coating allows high separation speed and short turnaround time for high throughput DNA sequencing. Critical evaluation of the system performance over repeated runs with base calling is presented.

Automation↗

Applications of recursive segmentation to the analysis of DNA sequences.

Recursive segmentation is a procedure that partitions a DNA sequence into domains with a homogeneous composition of the four nucleotides A, C, G and T. This procedure can also be applied to any sequence converted from a DNA sequence, such as to a binary strong(G + C)/weak(A + T) sequence, to a binary sequence indicating the presence or absence of the dinucleotide CpG, or to a sequence indicating both the base and the codon position information. We apply various conversion schemes in order to address the following five DNA sequence analysis problems: isochore mapping, CpG island detection, locating the origin and terminus of replication in bacterial genomes, finding complex repeats in telomere sequences, and delineating coding and noncoding regions. We find that the recursive segmentation procedure can successfully detect isochore borders, CpG islands, and the origin and terminus of replication, but it needs improvement for detecting complex repeats as well as borders between coding and noncoding regions.

Algorithms↗

Fast comparison of a DNA sequence with a protein sequence database.

We describe a computer program, named DNA-Protein Search (DPS), for comparing a megabase DNA sequence with a protein sequence database. The DPS program addresses the problems of frameshifts and introns in the DNA sequence. The DPS program was used to compare each of the following sequences with the Swiss-Prot database: the 1.8-megabase sequence of the Haemophilus influenzae Rd genome, the 0.58-megabase sequence of the Mycoplasma genitalium genome, and the 0.56-megabase sequence of Saccharomyces cerevisiae chromosome VIII. The comparisons found new regions that are similar to protein sequences. The sensitivity of DPS was evaluated using as test data the known coding regions of the three DNA sequences. The results demonstrate that the DPS program is a useful tool for finding the coding regions of the DNA sequence. The DPS program uses an order of magnitude less computer memory and is several times faster than the BLASTX program.

Amino Acid Sequence↗

Isolation and regional mapping of DNA sequences unique to human chromosome 21.

To isolate DNA sequences unique to chromosome 21 we have used a recombinant-DNA library, constructed from a mouse-human somatic-cell hybrid line containing chromosome 21 as the only human chromosome. Individual recombinant phage containing human DNA inserts were identified by their hybridization to total human DNA sequences and by their failure to hybridize to total mouse DNA sequences. A repeat-free human DNA fragment was then subcloned from each of 14 such recombinant phage. An independent somatic-cell hybrid was used to assign all 14 subcloned fragments to chromosome 21. Thirteen of the fragments have been regionally mapped using a somatic-cell hybrid containing a human 21 translocation chromosome. Two probes map proximal to the 21q21.2 translocation breakpoint, and 11 probes map distal to this breakpoint, placing them in the region 21q21.2-21q22. One of seven probes used to screen for restriction-fragment-length polymorphisms recognized polymorphic DNA fragments when hybridized to genomic DNA from unrelated individuals. These 14 unique probes provide useful tools for studying the structure and function of human chromosome 21 as well as for investigating the molecular biology of Down syndrome.

Animals↗

Evolutionary conservation of sex specific DNA sequences.

A family of DNA sequences which appears to be limited to eukaryotes is concentrated on the sex-determining chromosomes of species as widely separated evolutionarily as snakes and mammals. The significance of this distribution is presently seen in terms of the function of these sequences in cycles of chromosome condensation and decondensation involved in the control of gene expression. Thus, their concentration on the sex chromosomes is interpreted in the context of the evolution of such chromosomes which, it is hypothesised, involves the superimposition of the controls of the dominant sex-determining gene(s) upon the entire linkage group. This would result in the prevention of the large majority of its genes from having effects upon the phenotype and, in consequence, lead to their mutation to functionlessness at the maximum rate.

Animals↗

Effects of X-irradiation on the hybridization of rat thymus nuclear RNA with repeated and unique DNA sequences.

The kinetics of DNA hybridization with heterogeneous nuclear RNA (hnRNA) from normal and phytohaemagglutinin (PHA)-stimulated thymocytes has been studied in control rats and in animals 30 min after exposure to whole-body X-radiation with 400 rad. Irradiation results in a diminished ability of hnRNA to form hybrids with DNA at all C0t values ranging from 10(-3) to 10(4). Since this effect is most pronounced in the regions of low repetitive and unique DNA sequences, it is concluded that whole-body X-irradiation of animals may also suppress the transcription, particularly, in these regions.

Animals↗