Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Repseek, a tool to retrieve approximate repeats from large DNA sequences.

UNLABELLED: Chromosomes or other long DNA sequences contain many highly similar repeated sub-sequences. While there are efficient methods for detecting strict repeats or detecting already characterized repeats, there is no software available for detecting approximate repeats in large DNA sequences allowing for weighted substitutions and indels in a coherent statistical framework. Here, we present an implementation of a two-steps method (seed detection followed by their extension) that detects those approximate repeats. Our method is computationally efficient enough to handle large sequences and is flexible enough to account for influencing factors, such as sequence-composition biases both at the seed detection and alignment levels. AVAILABILITY: http://wwwabi.snv.jussieu.fr/public/RepSeek/

Algorithms↗

Dynamic organization of DNA replication in mammalian cell nuclei: spatially and temporally defined replication of chromosome-specific alpha-satellite DNA sequences.

Five distinct patterns of DNA replication have been identified during S-phase in asynchronous and synchronous cultures of mammalian cells by conventional fluorescence microscopy, confocal laser scanning microscopy, and immunoelectron microscopy. During early S-phase, replicating DNA (as identified by 5-bromodeoxyuridine incorporation) appears to be distributed at sites throughout the nucleoplasm, excluding the nucleolus. In CHO cells, this pattern of replication peaks at 30 min into S-phase and is consistent with the localization of euchromatin. As S-phase continues, replication of euchromatin decreases and the peripheral regions of heterochromatin begin to replicate. This pattern of replication peaks at 2 h into S-phase. At 5 h, perinucleolar chromatin as well as peripheral areas of heterochromatin peak in replication. 7 h into S-phase interconnecting patches of electron-dense chromatin replicate. At the end of S-phase (9 h), replication occurs at a few large regions of electron-dense chromatin. Similar or identical patterns have been identified in a variety of mammalian cell types. The replication of specific chromosomal regions within the context of the BrdU-labeling patterns has been examined on an hourly basis in synchronized HeLa cells. Double labeling of DNA replication sites and chromosome-specific alpha-satellite DNA sequences indicates that the alpha-satellite DNA replicates during mid S-phase (characterized by the third pattern of replication) in a variety of human cell types. Our data demonstrates that specific DNA sequences replicate at spatially and temporally defined points during the cell cycle and supports a spatially dynamic model of DNA replication.

Animals↗

Characterisation of a simple, highly repetitive DNA sequence from the parasite Leishmania donovani.

Repetitive DNA sequences of the Leishmania donovani genome have been identified by screening a recombinant DNA library made by cloning sheared genomic DNA into the vector pAT153. Bacterial clones containing a highly repetitive DNA sequence have been isolated. DNA sequencing has shown that this sequence is composed of tandem repeats of the sequence 5'-CCCTAA-3'. This sequence is identical to the telomeric repeats found in Trypanosoma brucei and hybridizes to all Leishmania chromosomes. In this study we show that there is considerable heterogeneity in the distribution and copy number of this repeat and associated hybridising sequences throughout the genomes of different Leishmania species.

Animals↗

Stochastic models for heterogeneous DNA sequences.

The composition of naturally occurring DNA sequences is often strikingly heterogeneous. In this paper, the DNA sequence is viewed as a stochastic process with local compositional properties determined by the states of a hidden Markov chain. The model used is a discrete-state, discrete-outcome version of a general model for non-stationary time series proposed by Kitagawa (1987). A smoothing algorithm is described which can be used to reconstruct the hidden process and produce graphic displays of the compositional structure of a sequence. The problem of parameter estimation is approached using likelihood methods and an EM algorithm for approximating the maximum likelihood estimate is derived. The methods are applied to sequences from yeast mitochondrial DNA, human and mouse mitochondrial DNAs, a human X chromosomal fragment and the complete genome of bacteriophage lambda.

Base Sequence↗

Human Rad51 protein displays enhanced homologous pairing of DNA sequences resembling those at genetically unstable loci.

DNA strand exchange, the central step of homologous recombination, is considered to occur approximately independently of DNA sequence content. However, certain prokaryotic and eukaryotic genomic loci display either an enhanced or reduced frequency of genetic exchange. Here we show that the Homo sapiens DNA strand exchange protein, HsRad51, shows a preference for binding to single-stranded DNA sequences primarily rich in G-residues and poor in A- and C-residues, and that these DNA sequences manifest enhanced HsRad51 protein-dependent homologous pairing. Both of these properties are common to all DNA strand exchange proteins examined thus far. These preferred DNA pairing sequences resemble those found at genetic loci in human cells that cause genomic instability and lead to genetic diseases.

Base Pairing↗

Isolation of human DNA sequences that bind to nuclear factor I, a host protein involved in adenovirus DNA replication.

Nuclear factor I is a 47,000-dalton protein isolated from human HeLa cells that is required for the in vitro replication of adenovirus DNA. This protein was previously shown to bind specifically to nucleotides 17-48 of the left-hand terminus of cloned adenovirus serotype 5 DNA. An in vitro assay for DNA sequences that compete with adenovirus DNA for the binding of nuclear factor I has been developed. With this assay, we have shown specific binding of human DNA sequences to nuclear factor I. Using the DNA binding activity of nuclear factor I, we have isolated and cloned segments of human DNA that bind tightly to this protein. One nuclear factor I binding site is present about every 100,000 base pairs in the HeLa cell genome. The binding of these DNA molecules to nuclear factor I resembles the binding of cloned adenovirus DNA to the protein and is resistant to high ionic strength. The isolation of DNA sequences from HeLa cells that bind specifically to nuclear factor I suggests that this protein interacts with host DNA in vivo.

Adenoviruses, Human↗

Microfluidic devices for DNA sequencing: sample preparation and electrophoretic analysis.

Modern DNA sequencing 'factories' have revolutionized biology by completing the human genome sequence, but in the race to completion we are left with inefficient, cumbersome, and costly macroscale processes and supporting facilities. During the same period, microfabricated DNA sequencing, sample processing and analysis devices have advanced rapidly toward the goal of a 'sequencing lab-on-a-chip'. Integrated microfluidic processing dramatically reduces analysis time and reagent consumption, and eliminates costly and unreliable macroscale robotics and laboratory apparatus. A microfabricated device for high-throughput DNA sequencing that couples clone isolation, template amplification, Sanger extension, purification, and electrophoretic analysis in a single microfluidic circuit is now attainable.

Base Sequence↗

Phasing of RecA monomers on quasi-random DNA sequences.

We show that some arbitrarily chosen DNA sequences have the ability to influence the positioning of RecA monomers in RecA-DNA complexes. The preferential phase of binding of RecA monomers is shown to depend on the DNA sequence and its nucleotide composition. A simple rearrangement of bases in a limited DNA stretch influences the phasing of RecA monomers. On the other hand, that some features of DNA sequences interfere with the phasing on specific DNA sites demonstrates the existence of mechanisms for both positive and negative regulation of phasing on natural DNAs. The possible role of phasing of RecA monomers on DNA is discussed.

Base Sequence↗

Sub-microliter DNA sequencing for capillary array electrophoresis.

DNA sequencing from sub-microliter samples was demonstrated for capillary array electrophoresis by optimizing the analysis of 500 nl reaction aliquots of full-volume reactions and by preparing 500 nl reactions within fused-silica capillaries. Sub-microliter aliquots were removed from the pooled reaction products of 10 microl dye-primer cycle-sequencing reactions and analyzed without modifying either the reagent concentrations or instrument workflow. The impact of precipitation methods, resuspension buffers, and injection times on electrokinetic injection efficiency for 500 nl aliquots were determined by peak heights, signal-to-noise ratios, and changes in base-called readlengths. For 500 nl aliquots diluted to 5 microl in 60% formamide-1 mM EDTA and directly injected, a five-fold increase in signal-to-noise ratios was obtained by increasing injection times from 10 to 80 s without a corresponding increase in peak widths or reduction in readlengths. For 500 nl aliquots precipitated in alcohol, 80 +/- 5% template recovery and a two-fold decrease in conductivity was obtained, resulting in a two-fold increase in peak heights and 50 to 100 bases increase in readlengths. In a comparison of aliquot volumes and precipitation methods, equivalent readlengths were obtained for 500 nl, 4 microl, and 8 microl aliquots by simply adjusting the electrokinetic injection conditions. To ascertain the robustness of this methodology for genomic sequencing, 96 Arabidopsis thaliana subclones were sequenced, with a yield of 38 624 bases obtained from 500 nl aliquots versus 30 764 bases from standard scale reactions. To demonstrate 500 nl sample preparation, reactions were performed in fused-silica capillary reaction chambers using air-based thermal cycling. A readlength of 690 bases was obtained for the polymerase chain reaction product of an Arabidopsis subclone without modifying the reagent concentrations, post-reaction processing or electrokinetic injection workflow. These results demonstrated the fundamental feasibility of small-volume DNA sequencing for high-throughput capillary electrophoresis.

Arabidopsis↗

Slipped-strand mispairing: a major mechanism for DNA sequence evolution.

Simple repetitive DNA sequences are a widespread and abundant feature of genomic DNA. The following several features characterize such sequences: (1) they typically consist of a variety of repeated motifs of 1-10 bases--but may include much larger repeats as well; (2) larger repeat units often include shorter ones within them; (3) long polypyrimidine and poly-CA tracts are often found; and (4) tandem arrangements of closely related motifs are often found. We propose that slipped-strand mispairing events, in concert with unequal crossing-over, can readily account for all of these features. The frequent occurrence of long tandem repeats of particular motifs (polypyrimidine and poly-CA tracts) appears to result from nonrandom patterns of nucleotide substitution. We argue that the intrahelical process of slipped-strand mispairing is much more likely to be the major factor in the initial expansion of short repeated motifs and that, after initial expansion, simple tandem repeats may be predisposed to further expansion by unequal crossing-over or other interhelical events because of their propensity to mispair. Evidence is presented that single-base repeats (the shortest possible motifs) are represented by longer runs in mammalian introns than would be expected on a random basis, supporting the idea that SSM may be a ubiquitous force in the evolution of the eukaryotic genome. Simple repetitive sequences may therefore represent a natural ground state of DNA unselected for coding functions.

Base Composition↗

Sequential and parallel algorithms for DNA sequencing.

MOTIVATION: Reconstruction of the original DNA sequence in the sequencing by the hybridization approach (SBH) requires computational support due to a large number of possible combinations. One can notice a lack of algorithms admitting false-negative data and giving in addition all possible solutions. RESULTS: In this paper, a new method of sequencing has been proposed. An algorithm based on its idea (for the general case, when some data are missing, like in the real experiment) has been implemented and tested. Authentic DNA sequences have been used for testing. A parallel version of the algorithm has also been implemented and tested. The quality of the reconstruction is satisfactory for the library of oligonucleotides of length between 8 and 12, and 100, 200 and 300 bp long sequences. A way to a further decrease in the computation time is also suggested.

Algorithms↗

Estimating the entropy of DNA sequences.

The Shannon entropy is a standard measure for the order state of symbol sequences, such as, for example, DNA sequences. In order to incorporate correlations between symbols, the entropy of n-mers (consecutive strands of n symbols) has to be determined. Here, an assay is presented to estimate such higher order entropies (block entropies) for DNA sequences when the actual number of observations is small compared with the number of possible outcomes. The n-mer probability distribution underlying the dynamical process is reconstructed using elementary statistical principles: The theorem of asymptotic equi-distribution and the Maximum Entropy Principle. Constraints are set to force the constructed distributions to adopt features which are characteristic for the real probability distribution. From the many solutions compatible with these constraints the one with the highest entropy is the most likely one according to the Maximum Entropy Principle. An algorithm performing this procedure is expounded. It is tested by applying it to various DNA model sequences whose exact entropies are known. Finally, results for a real DNA sequence, the complete genome of the Epstein Barr virus, are presented and compared with those of other information carriers (texts, computer source code, music). It seems as if DNA sequences possess much more freedom in the combination of the symbols of their alphabet than written language or computer source codes.

Algorithms↗

Cloning DNA sequences from influenza viral RNA segments.

DNA sequences corresponding to gene segments that code for the nonstructural protein, the matrix protein, and the hemagglutinin of influenza A virus [strain A/Udorn/72 (H3N2)] were cloned in Escherichia coli pBR 322. Initially, positive and negative cDNA strands were prepared separately by reverse transcription. The positive strands of cDNA were transcribed from genomic RNA segments by using a specific dodecamer DNA sequence as a primer; the negative strands of cDNA were transcribed from cytoplasmic viral mRNA segments by using an oligo(dT) primer. DNA duplexes corresponding in size to the virus RNA segments were then purified, inserted into the plasmid DNA, and used for transformation of E. coli. The influenza virus-specific DNA sequences isolated from recombinant plasmid molecules were characterized by mapping restriction enzyme cleavage sites. In addition, the orientation of cloned DNA was determined with reference to the 3' terminus of viral RNA.

Base Sequence↗

Methodologic European external quality assurance for DNA sequencing: the EQUALseq program.

BACKGROUND: DNA sequencing is a key technique in molecular diagnostics, but to date no comprehensive methodologic external quality assessment (EQA) programs have been instituted. Between 2003 and 2005, the European Union funded, as specific support actions, the EQUAL initiative to develop methodologic EQA schemes for genotyping (EQUALqual), quantitative PCR (EQUALquant), and sequencing (EQUALseq). Here we report on the results of the EQUALseq program. METHODS: The participating laboratories received a 4-sample set comprising 2 DNA plasmids, a PCR product, and a finished sequencing reaction to be analyzed. Data and information from detailed questionnaires were uploaded online and evaluated by use of a scoring system for technical skills and proficiency of data interpretation. RESULTS: Sixty laboratories from 21 European countries registered, and 43 participants (72%) returned data and samples. Capillary electrophoresis was the predominant platform (n = 39; 91%). The median contiguous correct sequence stretch was 527 nucleotides with considerable variation in quality of both primary data and data evaluation. The association between laboratory performance and the number of sequencing assays/year was statistically significant (P <0.05). Interestingly, more than 30% of participants neither added comments to their data nor made efforts to identify the gene sequences or mutational positions. CONCLUSIONS: Considerable variations exist even in a highly standardized methodology such as DNA sequencing. Methodologic EQAs are appropriate tools to uncover strengths and weaknesses in both technique and proficiency, and our results emphasize the need for mandatory EQAs. The results of EQUALseq should help improve the overall quality of molecular genetics findings obtained by DNA sequencing.

European Union↗

Simultaneous determination of different DNA sequences by mass spectrometric evaluation of Sanger sequencing reactions.

All currently available DNA sequencing protocols rest fundamentally upon the homogeneity of the template. In this paper we describe the parallel DNA sequencing of various templates in one sample by a combination of the Sanger method and MALDI-TOF mass spectrometric analysis of the products. PCR-amplified hypervariable 16S rDNA fragments of the bacterium Escherichia coli DF1020 and cDNA of the 6-phosphofructo-1-kinase isoenzymes (PFK-1, EC 2.7.1.11) in rat brain were chosen as model systems for essentially heterogeneous templates. Avoiding cloning of the inhomogeneous PCR products we were able to read three sequences for both the 16S rDNA fragment of E.coli DF1020 and the cDNA of 6-phosphofructo-1-kinase from the peak lists of the Sanger sequencing reactions. Short sequences with a length between 21 and 25 nt were sufficient to reflect the heterogeneity of the 16S rDNA genes in E.coli and the existence of three isoenzymes of PFK-1 in rat brain.

Animals↗