Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA sequence analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Analysis of DNA sequences.

Recent developments in the statistical analysis of DNA sequences are reviewed. The pace with which sequence data are being generated and analysed has increased with the growth of the human genome project. Two areas of activity are emphasized: attention to error rates in recorded sequences, and heterogeneity in structure of sequences. There is now empirical evidence suggesting error rates in the range 0.1%-1%, and such rates will affect evolutionary studies since these are about the rates at which DNA sequences from different individuals are expected to differ. Heterogeneity for such quantities as base composition, or lengths between successive subsequences of specified types, may be sufficient to account for observed long-range correlations between bases. The need for statistical models and analyses of DNA sequence data will continue, and will offer interesting challenges.

Algorithms↗

Numerical characterization and similarity analysis of DNA sequences based on 2-D graphical representation of the characteristic sequences.

Based on the classifications of the four nucleic acid bases, He and Wang reduced a DNA sequence to three binary sequences, which are called the characteristic sequences (J. Chem. Inf. Comput. Sci. 42 (2002) 1080). In this paper, we associate each characteristic sequence with a (b)L / (b)L matrix by giving a 2-D 'two horizontal lines' graphical representation, and thus obtain a 3-component vector with entries being the sums of the maximal and minimal eigenvalues of the (b)L / (b)L matrices. The introduced vector results in more simple characterizations and comparisons among the coding sequences of exon 1 of beta-globin gene of eleven different species.

Animals↗

Origin of cell populations after bone marrow transplantation. Analysis using DNA sequence polymorphisms.

After successful bone marrow transplantation, patient hematopoietic and lymphoid cells are replaced by cells derived from the donor marrow. To document and characterize successful engraftment, host and donor cells must be distinguished from each other. We have used DNA sequence polymorphism analysis to determine reliably the host or donor origin of posttransplant cell populations. Using a selected panel of six cloned DNA probes and associated sequence polymorphisms, at least one marker capable of distinguishing between a patient and his sibling donor can be detected in over 95% of cases. Posttransplant patient peripheral leukocytes were examined by DNA restriction enzyme digestion and blot hybridization analysis. We have studied 18 patients at times varying from 13 to 1,365 d after marrow transplantation. Mixed lymphohematopoietic chimerism was detected in 3 patients, with full engraftment documented in 15. One patient with severe combined immunodeficiency syndrome was demonstrated to have T cells of purely donor origin, with granulocytes and B cells remaining of host origin. Posttransplant leukemic relapse was studied in one patient and shown to be of host origin. DNA analysis was of particular clinical value in three cases where failure of engraftment or graft loss was suspected. In two of the three cases, full engraftment was demonstrated and in the third mixed lymphohematopoietic chimerism was detected. DNA sequence polymorphism analysis provides a powerful tool for the documentation of engraftment after bone marrow transplantation, for the evaluation of posttransplant lymphoma or leukemic relapse, and for the comprehensive study of mixed hematopoietic and lymphoid chimeric states.

Adult↗

Distribution of HIV-1 subtypes in female sex workers of Calcutta, India.

BACKGROUND & OBJECTIVES: Different routes for the transmission of HIV-1 in India have been reported and the majority of infections occurred through heterosexual route of transmission. In order to understand the dynamics of HIV-1 transmission, a systematic study was undertaken to determine the viral subtypes circulating among the female sex workers in Calcutta, India. METHODS: Peptide enzyme immunoassay (PEIA), heteroduplex mobility assay (HMA) and DNA sequence analysis were used to ascertain the HIV-1 subtypes. RESULTS: V3 serotyping of 52 HIV-1 seropositive samples identified 33 (60%) to be subtype C. A DNA fragment within C2-V3/C2-V5 regions of HIV-1 gp120 was amplified directly from the lymphocyte DNA to avoid any bias in selecting viral variants and used in HMA. Of the 40 samples analyzed, 38 (95%) belonged to subtype C and 2 were found to be non-typable. Further analysis of these 38 samples revealed that 26 (68%) had maximum homology to the C3-Indian reference strain (IND868), 11 (29%) were most homologous to C2-Zambian strain (ZM18) and 1 (3%) showed close resemblance to C1-Malawi strain (MA959). Nucleotide sequence of 11 subsamples encompassing about 325 base pairs was aligned for the Indian and other geographically distinct isolates. On distance and parsimony trees, most of the samples (8/11) clustered together as subtype C. INTERPRETATION & CONCLUSIONS: Subtype C was the major circulating HIV-1 strain in this geographical region, although variation within this subtype was also noticed. DNA sequence analysis was found to be the best method in determining the nature of the HIV-1 subtype followed by HMA and peptide enzyme immunoassay. These findings may have important implications for the design of effective vaccines in India and emphasizes the need for constant monitoring of the HIV-1 subtypes in different parts of India.

Amino Acid Sequence↗

Identification of S-sulfonation and S-thiolation of a novel transthyretin Phe33Cys variant from a patient diagnosed with familial transthyretin amyloidosis.

Familial transthyretin amyloidosis (ATTR) is an autosomal dominant disorder associated with a variant form of the plasma carrier protein transthyretin (TTR). Amyloid fibrils consisting of variant TTR, wild-type TTR, and TTR fragments deposit in tissues and organs. The diagnosis of ATTR relies on the identification of pathologic TTR variants in plasma of symptomatic individuals who have biopsy proven amyloid disease. Previously, we have developed a mass spectrometry-based approach, in combination with direct DNA sequence analysis, to fully identify TTR variants. Our methodology uses immunoprecipitation to isolate TTR from serum, and electrospray ionization and matrix-assisted laser desorption/ionization mass spectrometry (MS) peptide mapping to identify TTR variants and posttranslational modifications. Unambiguous identification of the amino acid substitution is performed using tandem MS (MS/MS) analysis and confirmed by direct DNA sequence analysis. The MS and MS/MS analyses also yield information about posttranslational modifications. Using this approach, we have recently identified a novel pathologic TTR variant. This variant has an amino acid substitution (Phe --> Cys) at position 33. In addition, like the Cys10 present in the wild type and in this variant, the Cys33 residue was both S-sulfonated and S-thiolated (conjugated to cysteine, cysteinylglycine, and glutathione). These adducts may play a role in the TTR fibrillogenesis.

Amino Acid Sequence↗

Statistical analysis of DNA sequencing data (1): accuracy test of DNA data by partial re-sequencing.

To qualify DNA data, we have developed a statistical method of deciding whether the DNA data has an acceptable accuracy in sequencing process. The method is to test the probability of sequencing errors, based on partial re-sequencing. The method was successfully applied to a yeast mitochondrial DNA which is previously sequenced (1). The analysis indicates that the entire sequence is very accurate although we found one base change error on the ND1 gene sequence data by a partial re-sampling. This method is applicable to any DNA data.

Chromosome Mapping↗

[Application of DNA-related techniques in avian molecular phylogeny].

The DNA techniques most commonly used in avian molecular phylogeny include DNA hybridization, RFLP and DNA sequence analysis, among which DNA sequence analysis is supposed to be the most effective and reliable. DNA hybridization techniques have been widely used in aves, based on which a new avian classification system was born. In avian RFLP analyses, mtDNA are widely used as target sequences. Mitochondrial DNA genes are the most frequently used in avian molecular phylogeny. Although mtDNA phylogenies are likely to be correct in many cases, use of mtDNA sequences can be problematic with such constraints as unilateral inheritance, multiple substitutions, saturations at the third-coded sites, strong bias in base composition and probable nuclear pseudogenes of mtDNA sequences. Although bias are still on the mtDNA sequences, more and more authors turn to nuclear DNA sequences and prefer to a combination of mtDNA and nuclear DNA sequences. And single-copy nuclear DNA receives the most favor. scnDNA introns can perform well in recovering relationships among intermediate to even distantly related congeneric species. scnDNA exons can be used in avian higher ranks. With the exception of molecular markers' own problems including variable rates of nucleotide site evolution, gene hybridization, gene horizontal transfer and lineage sorting, avian molecular phylogeny also faces methodological problems, such as molecular markers selection, taxon sampling and data processing. More attention should be paid to the standardization of methods, not to the new molecular markers.

Animals↗

DNA nucleotide sequence analysis of the PvuII DNA fragment L of the genome of insect iridescent virus type 6 reveals a complex cluster of multiple tandem, overlapping, and interdigitated repetitive DNA elements.

The DNA nucleotide sequence of the PvuII DNA fragment L (0.920 to 0.944 map units (m.u.] of the genome (209 kbp) of insect iridescent virus type 6 was determined. The size of this DNA fragment was 5064 bp with a base composition of 39.79% G + C and 60.21% A + T. The DNA sequence contained many perfect direct repeats of sizes up to 145 bp. In addition to these repetitions, a cluster of four imperfect repetitive DNA elements (R1 to R4) with a complex structural arrangement was detected. R1, R2, and R3 existed in duplicate (two boxes (B] between nucleotide positions 271 and 3466) and their size were as follows: R1-B1/B2 (567/568 bp), R2-B1/B2 (917/931 bp), and R3-B1/B2 (92/88 bp). The R4 repetitive element was found in 12 boxes (between bases 1301 and 4417), which were interrupted at nucleotide positions 1883 to 2236 and 3341 to 3587. These interruptions define three segments (S) harboring boxes B1 to B3 (S1), B4 to B8 (S2), and B9 to B12 (S3). The size of the individual boxes was found to be 239, 233, 107, 244, 222, 242, 242, 148, 240, 242, 242, and 102 bp for R4-B1 to B12, respectively. Five open reading frames (ORFs of 118 to 333 amino acid (AA) residues) were detected. The analysis of the amino acid sequences of the largest ORF revealed that the deduced amino acid sequence of the putative gene product contained two repetitions TR1 (three domains of 50 AA) and TR2 (two domains of 74 AA). Sequences of 43 amino acid residues of ORF 5 (160 to 202 AA) were homologous within the majority of ORFs. A consensus sequence-MANL(X)6 IGSSST(X)6 L(X)1 LGS(X)1 LQISG(X)2 L(X)1 VN- was found in all five ORFs. Although classical canonical and noncanonical transcriptional start signals were detectable, polyadenylation signals were not observed.

Amino Acid Sequence↗

Systematic use of automated fluorescence-based sequence analysis of amplified genomic DNA for rapid detection of point mutations.

Several approaches are now available for screening populations for known mutations in a given gene. However, for detection of multiple mutations in a population that has not been characterized or for detection of new mutations, the value and efficiency of these screening procedures decreases. Although more than 100 different beta-thalassemia mutations have so far been described, the spectrum of mutations in the Eastern Mediterranean and Israel has not been defined in detail. We have used automated fluorescence-based DNA sequence analysis of PCR-amplified genomic DNA employing a cycle-sequencing strategy coupled with advanced analysis software to rapidly detect beta-thalassemia mutations in Israeli patients. This method enabled rapid identification of eight different mutations in 10 patients, including two rare mutations, one of which has never been described in this geographic region. Our results show that automated fluorescence-based DNA sequence analysis of amplified genomic DNA is a rapid and reliable method for detection of point mutations and small deletions or insertions in both heterozygous and homozygous states. This approach is particularly effective for a relatively small gene such as beta-globin, but it can also be used for rapid detection of mutations in large genes by first sequencing clusters of exons and intron/exon borders.

Automation↗

Direct comparison of selected methods for genetic categorisation of Cryptosporidium parvum and Cryptosporidium hominis species.

A study was undertaken to compare the performance of five different molecular methods (available in four different laboratories) for the identification of Cryptosporidium parvum and Cryptosporidium hominis and the detection of genetic variation within each of these species. The same panel of oocyst DNA samples derived from faeces (n=54; coded blindly) was sent for analysis by: (i) DNA sequence analysis of a fragment of the HSP70 gene; (ii) DNA sequence analysis and the ssrRNA gene in laboratory 1; (iii) single-strand conformation polymorphism analysis of part of the ssrRNA; (iv) SSCP analysis of the second internal transcribed spacer (ITS-2) of nuclear ribosomal DNA region in laboratory 2; (v) 60 kDa glycoprotein (gp60) gene sequencing with prior species determination using PCR with restriction fragment length polymorphism analysis of the ssrRNA gene in laboratory 3; and (vi) multilocus genotyping at three microsatellite markers in laboratory 4. For detecting variation within C. parvum and C. hominis, SSCP analysis of ITS-2 was considered to have superior utility and determined 'subgenotypes' in samples containing DNA from both species. SSCP was also most cost effective in terms of time, cost and consumables. Sequence analysis of gp60 and microsatellite markers ML1, ML2 and 'gp15' provided good comparators for the SSCP of ITS-2. However, applicability of these methods to other Cryptosporidium species or genotypes and to environmental samples needs to be evaluated. This trial provided, for the first time, a direct comparison of multiple methods for the genetic characterisation of C. parvum and C. hominis samples. A protocol has been established for the international distribution of samples for the characterisation of Cryptosporidium. This can be applied in further evaluation of molecular methods by investigation of a larger number of unrelated samples to establish sensitivity, typability, reproducibility and discriminatory power based on internationally accepted methods for evaluation of microbial typing schemes.

Adolescent↗

Characterisation of the yenI/yenR locus from Yersinia enterocolitica mediating the synthesis of two N-acylhomoserine lactone signal molecules.

Yersinia enterocolitica produces compounds capable of transcriptionally activating the Photobacterium fischeri bioluminescence (lux) operon. Using high-performance liquid chromatography, high resolution tandem mass spectrometry in conjunction with chemical synthesis, two signal molecules were identified and shown to be N-hexanoyl-L-homoserine lactone (HHL) and N-(3-oxohexanoyl)-L-homoserine lactone (OHHL). A gene (yenI) was isolated from Y. enterocolitica and demonstrated to direct the synthesis of both HHL and OHHL. DNA sequence analysis revealed an open reading frame (ORF) of 642 bp encoding a protein (YenI) of 24.6 kDa with approximately 20% identity to the LuxI family of proteins. Northern blot analysis of yenI expression indicated yenI is transcribed as a single gene and 5' transcript mapping of yenI identified a transcriptional start site 89 bp upstream of the ORF. DNA sequence analysis of the region downstream of yenI located a second ORF, termed yenR, with significant homology to the LuxR family of transcriptional activators. An insertion mutation of yenI abolishes HHL and OHHL production, indicating its central role in N-acylhomoserine lactone synthesis in Y. enterocolitica. Transcriptional analysis using a chromosomal yenI::luxAB fusion has demonstrated that yenI is not subject to autoinduction but is expressed constitutively. Whilst production of the Yop proteins in the wild type and in yenI mutants is indistinguishable, two-dimensional SDS-PAGE analysis of total cell proteins indicated that a number of proteins lack the yenI mutant.

4-Butyrolactone↗

Advances in the analysis of DNA sequence variations using oligonucleotide microchip technology.

The analysis of DNA variation (polymorphisms and mutations) on a genome-wide scale is becoming both increasingly important and technically challenging. An integration of a growing number of molecular biological methods of DNA-sequence analysis with the high-throughput feature of oligonucleotide microarray-based technologies is one of the most promising current directions of research and development.

Base Sequence↗

Compilation and analysis of Bacillus subtilis sigma A-dependent promoter sequences: evidence for extended contact between RNA polymerase and upstream promoter DNA.

Sequence analysis of 236 promoters recognized by the Bacillus subtilis sigma A-RNA polymerase reveals an extended promoter structure. The most highly conserved bases include the -35 and -10 hexanucleotide core elements and a TG dinucleotide at position -15, -14. In addition, several weakly conserved A and T residues are present upstream of the -35 region. Analysis of dinucleotide composition reveals A2- and T2-rich sequences in the upstream promoter region (-36 to -70) which are phased with the DNA helix: An tracts are common near -43, -54 and -65; Tn tracts predominate at the intervening positions. When compared with larger regions of the genome, upstream promoter regions have an excess of An and Tn sequences for n > 4. These data indicate that an RNA polymerase binding site affects DNA sequence as far upstream as -70. This sequence conservation is discussed in light of recent evidence that the alpha subunits of the polymerase core bind DNA and that the promoter may wrap around RNA polymerase.

Adenine↗

Statistical analysis of DNA sequences.

Developments in the statistical analysis of DNA sequence data since 1984 are reviewed. Mathematical methods employing dynamic programming or incorporating Markov chain theory have been developed to search sequences for regions of similarity and to align sequences. When the biological forces of mutation and genetic drift are included in models, distances between aligned sequences allow the construction of evolutionary trees. Theory based on models may lead to estimates of variation of parameter estimates and so give a means of assessing the statistical significance of observed patterns and relationships. The complexity of DNA sequences, however, suggests that most statistical inferences will rest on random permutations of sequences.

Base Sequence↗

Sequence analysis of proviral DNA of porcine endogenous retroviruses.

Among all species analyzed, the domestic pig seems to be the most appropriate organ donor for xenotransplantation. Porcine endogenous retroviruses (PERVs) are present in genomes of all pigs and are capable of infecting human cells in vitro thus posing a serious threat for xenotransplantation procedures. Despite the abundant distribution of PERVs integrated with porcine genome, the majority of PERV proviral DNA is not capable of expressing viral proteins unless seriously mutated. The aim of the study was to analyze PERV genome for mutations. The study was performed on blood samples from 146 pigs. Long-range polymerase chain reaction (Long-PCR) was performed with primer sets designed within long terminal repeats (LTRs). Long-PCR products of different molecular weights were obtained: 530 bp (33.1% of individuals), 580 bp (76.7%), 933 bp (100%), and 2900 bp (59.8%). Amplimers of 7200 bp were absent in 12.8% of individuals, indicating the lack of intact proviral DNA. Sequence analysis showed that most PERV proviral DNA was significantly mutated, thus suggesting the inability to express functional viral RNA; however, it cannot be ruled out that compensatory recombination processes could occur enabling replication of defective proviruses.

Animals↗

Complete DNA sequence and analysis of an emerging cryptic plasmid isolated from Yersinia pestis.

A 6-kb cryptic plasmid (pYC; 5919 bp) has been recovered from Yersinia pestis isolates originating from regions of Yunnan province in China. The sequence of pYC was determined, and analysis of the sequence has revealed that two of the plasmid DNA regions (ORFs 10 and 11) are similar to the DinJ1 and DinJ2 gene products encoded by Escherichia coli chromosomal DNA. This plasmid is increasingly harbored by Y. pestis isolates recovered from a domestic rodent cycle in the southern regions of the province. Further studies will determine the origin and function of pYC.

Base Sequence↗

High-speed conversion of cytosine to uracil in bisulfite genomic sequencing analysis of DNA methylation.

Bisulfite genomic sequencing is a widely used technique for analyzing cytosine-methylation of DNA. By treating DNA with bisulfite, cytosine residues are deaminated to uracil, while leaving 5-methylcytosine largely intact. Subsequent PCR and nucleotide sequence analysis permit unequivocal determination of the methylation status at cytosine residues. A major caveat associated with the currently practiced procedure is that it takes 16-20 hr for completion of the conversion of cytosine to uracil. Here we report that a complete deamination of cytosine to uracil can be achieved in shorter periods by using a highly concentrated bisulfite solution at an elevated temperature. Time course experiments demonstrated that treating DNA with 9 M bisulfite for 20 min at 90 degrees C or 40 min at 70 degrees C all cytosine residues in the DNA were converted to uracil. Under these conditions, the majority of 5-methylcytosines remained intact. When a high molecular weight DNA derived from a cell line (containing a number of genes whose methylation status was known) was treated with bisulfite under the above conditions and amplified and sequenced, the results obtained were consistent with those reported in the literature. Although some degradation of DNA occurred during this process, the amount of treated DNA required for the amplification was nearly equal to that required for the conventional bisulfite genomic sequencing procedure. The increased speed of DNA methylation analysis with this novel procedure is expected to advance various aspects of DNA sciences.

Base Sequence↗

Identification and structural analysis of a MDV gene encoding a protein kinase.

DNA sequence analysis of the BamHI-C fragment of Marek's Disease Virus (MDV) reveals the presence of a 513 amino acid open reading frame (ORF). This ORF codes for a protein with an estimated M(r) of 58,901. Comparison of the amino acid sequence with those available in the Swiss-Prot database indicates extensive homology with a protein kinase (PK) of herpes simplex virus (HSV) and varicella-zoster virus (VZV). In Northern blot hybridization, a transcript of 2.0 kb was detected in MDV (GA strain) infected duck embryo fibroblasts (DEFs). A portion of the ORF was expressed in Escherichia coli as a trpE-fusion protein and used to generate antiserum in New Zealand rabbits. This antiserum specifically detected a protein of 60 kDa in MDV serotype 1, 2 and 3 infected DEFs or chicken embryo fibroblasts (CEFs) by Western blot analysis. This ORF codes for a functional PK.

Amino Acid Sequence↗