Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Enzyme-linked fluorescent detection for automated multiplex DNA sequencing.

Initiatives to sequence DNA on a large scale have created a need for increased throughput and decreased costs. One scheme for increasing throughput, multiplex sequencing, involves the processing of a mixture of sequencing templates followed by sequential hybridization to reveal the individual sequence ladders on a membrane. Because multiplex sequencing has not been fully automated, and has not seemed automatable, few sequencing efforts have attempted to exploit it. We describe here a scheme for the automation of multiplex sequencing. Probe hybridized to target DNA is detected via spatially localized enzyme-linked fluorescence. Light output is high enough that imaging is possible with simple instrumentation. Direct imaging within an automated hybridization apparatus is made feasible so that the entire process will be automatic once a multiplex membrane is produced. The technique has the potential to increase severalfold the throughput of automated sequencing instruments required for sequencing the human genome.

Base Sequence↗

Human 18 S ribosomal RNA sequence inferred from DNA sequence. Variations in 18 S sequences and secondary modification patterns between vertebrates.

We have determined the DNA sequences encoding 18 S ribosomal RNA in man and in the frog, Xenopus borealis. We have also corrected the Xenopus laevis 18 S sequence: an A residue follows G-684 in the sequence. These and other available data provide a number of representative examples of variation in primary structure and secondary modification of 18 S ribosomal RNA between different groups of vertebrates. First, Xenopus laevis and Xenopus borealis 18 S ribosomal genes differ from each other by only two base substitutions, and we have found no evidence of intraspecies heterogeneity within the 18 S ribosomal DNA of Xenopus (in contrast to the Xenopus transcribed spacers). Second, the human 18 S sequence differs from that of Xenopus by approx. 6.5%. About 4% of the differences are single base changes; the remainder comprise insertions in the human sequence and other changes affecting several nucleotides. Most of these more extensive changes are clustered in a relatively short region between nucleotides 190 and 280 in the human sequence. Third, the human 18 S sequence differs from non-primate mammalian sequences by only about 1%. Fourth, nearly all of the 47 methyl groups in mammalian 18 S ribosomal RNA can be located in the sequence. The methyl group distribution corresponds closely to that in Xenopus, but there are several extra methyl groups in mammalian 18 S ribosomal RNA. Finally, minor revisions are made to the estimated numbers of pseudouridines in human and Xenopus 18 S ribosomal RNA.

Animals↗

Significantly lower entropy estimates for natural DNA sequences.

If DNA were a random string over its alphabet {A, C, G, T}, an optimal code would assign two bits to each nucleotide. DNA may be imagined to be a highly ordered, purposeful molecule, and one might therefore reasonably expect statistical models of its string representation to produce much lower entropy estimates. Surprisingly, this has not been the case for many natural DNA sequences, including portions of the human genome. We introduce a new statistical model (compression algorithm), the strongest reported to date, for naturally occurring DNA sequences. Conventional techniques code a nucleotide using only slightly fewer bits (1.90) than one obtains by relying only on the frequency statistics of individual nucleotides (1.95). Our method in some cases increases this gap by more than fivefold (1.66) and may lead to better performance in microbiological pattern recognition applications. One of our main contributions, and the principle source of these improvements, is the formal inclusion of inexact match information in the model. The existence of matches at various distances forms a panel of experts which are then combined into a single prediction. The structure of this combination is novel and its parameters are learned using Expectation Maximization (EM). Experiments are reported using a wide variety of DNA sequences and compared whenever possible with earlier work. Four reasonable notions for the string distance function used to identify near matches, are implemented and experimentally compared. We also report lower entropy estimates for coding regions extracted from a large collection of nonredundant human genes. The conventional estimate is 1.92 bits. Our model produces only slightly better results (1.91 bits) when considering nucleotides, but achieves 1.84-1.87 bits when the prediction problem is divided into two stages: (i) predict the next amino acid-based on inexact polypeptide matches, and (ii) predict the particular codon. Our results suggest that matches at the amino acid level play some role, but a small one, in determining the statistical structure of nonredundant coding sequences.

Algorithms↗

Assignment of position-specific error probability to primary DNA sequence data.

DNA sequence predicted from polyacrylamide gel-based technologies is inaccurate because of variations in the quality of the primary data due to limitations of the technology, and to sequence-specific variations due to nucleotide interactions within the DNA molecule and with the gel. The ability to recognize the probability of error in the primary data will be useful in reconstructing the target sequence of a DNA sequencing project, and in estimating the accuracy of the final sequence. This paper describes the use of linear discriminant analysis to assign position-specific probabilities of incorrect, over- and under-prediction of nucleotides for each predicted nucleotide position in primary sequence data generated by a gel-based DNA sequencing technology. Using this method, most of the error potential in primary sequence data can be assigned to a limited number of discrete positions. The use of probability values in the sequence reconstruction process, and in estimating the accuracy of consensus sequence determination is described.

Base Sequence↗

Cloning and characterization of an amplified DNA sequence in chromosomal DNA of Streptomyces aureofaciens 2201.

In strain 2201 of Streptomyces aureofaciens, a high copy number amplified DNA sequence (ADS-Sa2201) was found and characterized. The amplified sequence in the chromosomal DNA of this strain forms a stretch of about 10 kb tandemly repeated 100-500 times. In this strain also, extrachromosomal copies of the repeated unit of the ADS-Sa2201 were found. In cloning experiments any autonomous replicon was found on ADS-Sa2201 and it thus can be presumed that the presence of the extrachromosomal copies of the repeated unit of ADS-Sa2201 is only a result of its excision from the chromosome. A part of the repeated unit of ADS-Sa2201 was also found in chromosomal DNA of strain 13 of S. aureofaciens. In this strain no part of ADS-Sa2201 is present extrachromosomally.

Cloning, Molecular↗

Orphan peak analysis: a novel method for detection of point mutations using an automated fluorescence DNA sequencer.

Automated DNA sequencers draw the four-base profiles of a sample with four different colors, but it is also possible to draw the profiles of a base-specific reaction of four different samples with four colors. PCR-amplified DNAs from four individuals were subjected to a single base-specific sequencing reaction and the products were applied to a set of four lanes of an automated DNA sequencer. A base substitution in an individual was clearly identified as an individual-specific peak with a color specific for the individual. In this way, we analyzed more than 50 individuals and identified several polymorphic base substitutions. The sensitivity of this method was high enough to allow detection of the mutation/polymorphism even if samples from several individuals were applied to one lane. Thus, our method is applicable to screening of a large number of samples in an automated manner.

Amyloid beta-Protein Precursor↗

DNA sequence variation as a clue to the phylogenesis of orthopoxviruses.

We have sequenced DNA equivalent to the E5R ORF of Copenhagen vaccinia virus from an additional strain of vaccinia and from cowpox (three strains), camelpox (two strains), taterapox and ectromelia viruses. None of these showed the disruptions previously reported in the equivalent region of monkeypox virus. We also constructed a viable recombinant of vaccinia virus strain Dairen in which the E5R sequence was disrupted by a 436 bp deletion and substitution of the E. coli gpt gene. Quantitative analysis of the sequences, including available sequences from monkeypox, variola and vaccinia viruses revealed four main groupings, namely cowpox, ectromelia, monkeypox and a cluster which includes variola, camelpox, taterapox and vaccinia viruses. It was noted that, at over 75 % of the positions which differentiated species. all species but one had a common nucleotide. Although the analysis covers one single gene only, the results accord with what is known of the biology of the viruses.

Base Sequence↗

A subtelomeric DNA sequence is required for correct processing of the macronuclear DNA sequences during macronuclear development in the hypotrichous ciliate Stylonychia lemnae.

During macronuclear differentiation in ciliated protozoa a series of programed DNA reorganization processes occur. These include the elimination of micronuclear-specific DNA sequences, the specific fragmentation of the genome into small gene-sized DNA molecules, the de novo addition of telomeric sequences to these DNA molecules and the specific amplification of the remaining DNA molecules. Recently we constructed a vector containing the modified micronuclear version of macronuclear destined DNA sequences that was correctly fragmented and telomeres were added de novo after injection into the developing macronucleus. It therefore must contain all the cis- acting sequences required for these processes. We made a series of vectors deleting different sequences from the original vector. It could be shown that at least in the case studied here no micronuclear-specific sequences are required for specific fragmentation of the genome and telomere addition. However, a short subtelomeric sequence at the 3[prime]-end is essential for these processes, whereas no specific cut seems to occur at the 5[prime]-end. In addition, we can show that the processing activity is restricted to a short period of time during macronuclear differentiation and that a preceding transcription is required for correct processing of macronuclear-destined DNA sequences. Possible mechanisms of these processes will be discussed.

Animals↗

Efficient preparation of short DNA sequence ladders potentially suitable for MALDI-TOF DNA sequencing.

Duplex probes with five-base single-stranded overhangs were developed for positional sequencing by hybridization [Broude et al., Proc Natl Acad Sci USA 91:3072-3076, 1994]. The partially duplex probes can be employed to capture single-stranded oligonucleotide targets and form primer-template complexes. Recently we showed that partially duplex probes can prime Sanger sequencing reactions on immobilized, but non-ligated long single-stranded targets (approximately 500 nucleotide) [Fu et al., Proc Natl Acad Sci, in press]. Here immobilized, non-ligated partially duplex probes were used to capture and sequence short single-stranded targets. This strategy is capable of rapidly preparing large numbers of samples for future mass spectrometric DNA sequencing.

Base Sequence↗

Prediction of function in DNA sequence analysis.

Recognition of function of newly sequenced DNA fragments is an important area of computational molecular biology. Here we present an extensive review of methods for prediction of functional sites, tRNA, and protein-coding genes and discuss possible further directions of research in this area.

Algorithms↗

The estimation of the number and the length distribution of gene conversion tracts from population DNA sequence data.

DNA sequence variation studies report the transfer of small segments of DNA among different sequences caused by gene conversion events. Here, we provide an algorithm to detect gene conversion tracts and a statistical model to estimate the number and the length distribution of conversion tracts for population DNA sequence data. Two length distributions are defined in the model: (1) that of the observed tract lengths and (2) that of the true tract lengths. If the latter follows a geometric distribution, the relationship between both distributions depends on two basic parameters: psi, which measures the probability of detecting a converted site, and phi, the parameter of the geometric distribution, from which the average true tract length, 1/(1-phi), can be estimated. Expressions are provided for estimating phi by the method of the moments and that of the maximum likelihood. The robustness of the model is examined by computer simulation. The present methods have been applied to the published rp49 sequences of Drosophila subobscura. Maximum likelihood estimate of phi for this data set is 0.9918, which represents an average conversion tract length of 122 bp. Only a small percentage of extant conversion events is detected.

Animals↗

DNA sequences of genes encoding Acinetobacter calcoaceticus protocatechuate 3,4-dioxygenase: evidence indicating shuffling of genes and of DNA sequences within genes during their evolutionary divergence.

The DNA sequence of a 2,391-base-pair HindIII restriction fragment of Acinetobacter calcoaceticus DNA containing the pcaCHG genes is reported. The DNA sequence reveals that A. calcoaceticus pca genes, encoding enzymes required for protocatechuate metabolism, are arranged in a single transcriptional unit, pcaEFDBCHG, whereas homologous genes are arranged differently in Pseudomonas putida. The pcaG and pcaH genes represent separate reading frames respectively encoding the alpha and beta subunits of protocatechuate 3,4-dioxygenase (EC 1.13.1.3); previously a single designation, pcaA, had been used to represent DNA encoding this enzyme. The alpha and beta protein subunits appear to share common ancestry with each other and with catechol 1,2-dioxygenases from A. calcoaceticus and P. putida. Marked conservation of amino acid sequence is observed in a region containing two histidyl residues and two tyrosyl residues that appear to ligate iron within each oxygenase. In some regions within the aligned oxygenase sequences, DNA sequences appear to be conserved at a level beyond the extent that might have been demanded by selection at the level of protein. In other regions, divergence of DNA sequences appears to have been achieved by substitution of DNA sequence from one genetic segment into another. The results are interpreted to be the consequence of sequence exchange by gene conversion between slipped strands of DNA during evolutionary divergence; mismatch repair between slipped strands may contribute to the maintenance of DNA sequence in divergent genes.

Acinetobacter↗

Molecular phylogenetics of gadidae and related gadiformes based on mitochondrial DNA sequences.

Mitochondrial DNA sequences of selected regions of the small subunit and large subunit ribosomal RNAs and cytochrome b genes were analyzed for 10 gadid species, representing 8 genera within Gadidae, and 10 species representing 5 other gadiform families. Phylogenetic analyses revealed that Gadiculus is the most basal gadid genus, and that Trisopterus and Micromesistius constitute a relatively basal clade. Lotidae was identified as the family most closely related to Gadidae. Estimation of divergence times indicated that the most ancient Gadidae split between Gadiculus and the remaining gadid genera occurred about 20 million years ago. The clade including the most recent species (Gadus, Boreogadus, Merlangius, Melanogrammus, and Pollachius) diverged from the Trisopterus/Micromesistius clade approximately 12 million years ago.

Animals↗

Sequential DEXAS: a method for obtaining DNA sequences from genomic DNA and blood in one reaction.

Sequential DEXAS (direct exponential amplification and sequencing), a one step amplification and sequencing procedure that allows accurate, inexpensive and rapid DNA sequence determination directly from genomic DNA, is described. This method relies on the simultaneous use of two DNA polymerases that differ both in their ability to incorporate dideoxynucleotides and in the time at which they are activated during the reaction. One enzyme, which incorporates deoxynucleotides and performs amplification of the target DNA sequence, is supplied in an active state whereas the other enzyme, which incorporates dideoxynucleotides and performs the sequencing reaction, is supplied in an inactive state but becomes activated by a temperature step during the thermocycling. Thus, in the initial stage of the reaction, target amplification occurs, while in the second stage the sequencing reaction takes place. We show that Sequential DEXAS yields high quality sequencing results directly from genomic DNA as well as directly from human blood without any prior isolation or purification of DNA.

Base Sequence↗

A comparison of yeast ribosomal protein gene DNA sequences.

The DNA sequences of eight yeast ribosomal protein genes have been compared for the purpose of identifying homologous regions which may be involved in the coordinate regulation of ribosomal protein synthesis. A 12 bp homology was identified in the 5' DNA sequence preceding the structural gene for 6 out of 8 yeast ribosomal protein genes. In each case the homologous sequence was found at a position approximately 300 bp preceding the transcription start of the ribosomal protein gene. This homology was not identified in any non-ribosomal protein gene examined. Additional homologies between ribosomal protein genes were identified in the transcribed regions, including the untranslated 5' and 3' DNA regions flanking the coding regions.

Amino Acid Sequence↗

The DNA sequence quality machine at IFOM: a simple Web-based tool for quantitative assessment of sequencing reactions.

DNA sequence quality is a factor of paramount importance in the world of modern genetic and genomics. Both the sequencing of Human Genome in the "post-draft" era [NHGRI Standard for quality of Human Genomic Sequences, Rev. 7 July (2002) where http://www.nhgri.nih.gov/Grant_info/Funding/ Statements/RFA/quality_standard.html is the HTTP address] and recent "high-throughput" approaches to genetic investigation such as SAGE [Velculescu, V.E., Zhang, L., Vogelstein, B. et al. (1995) "Serial analysis of gene expression", Science 270, 484-487] need a reliable, standardized measure of the quality of a sequencing reaction. The increasing importance of SNP studies also requires a stronger quality control on sequencing reactions by the final user. We propose here a simple, web-based tool for integrated sequence quality evaluation, high quality region quantitative value calculation and chromatogram display. This software is aimed at the small to medium DNA sequence laboratory or to the single researcher, interested in getting a quantitative measure of the sequence quality, browsing the chromatogram and checking the quality values base by base. The program is freely available from the IFOM bioinformatics web Server at http://bio.ifom-firc.it/Phred20/index.html.

Algorithms↗