Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “DNA sequence analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Genetic characterization of wild-type genotype VII hepatitis A virus.

The complete genome sequence of the only identified genotype VII hepatitis A virus (HAV), strain SLF88, was obtained from PCR amplicons generated by a modified long PCR approach. There was 90% nucleotide identity in the 5' untranslated region compared to other known HAV sequences. In the remainder of the genome containing the long open reading frame, there was about 85% nucleotide identity to human HAV genotypes IA and IB and 80% identity to simian HAV genotype V. Compared to HAV strain HM-175, the capsid amino acids were highly conserved, with only four homologous amino acid changes, while an increasing number of amino acid differences was seen in the P2 and P3 genome regions. While nucleotide variability within the three functional coding regions did not differ, the P3D region was found to have the largest number of amino acid changes compared to HM-175.

Base Sequence↗

Gene cloning of an endoglucanase from the basidiomycete Irpex lacteus and its cDNA expression in Saccharomyces cerevisiae.

A gene (cen1) coding for an endoglucanase I (En-1) was isolated from white rot fungus Irpex lacteus strain MC-2. The cen1 ORF was comprised of 399 amino acid residues and interrupted by 14 introns. The deduced amino acid sequence of the cen1 ORF revealed a multi-domain structure composed of a cellulose-binding domain, a Ser-/Thr-rich linker, and a catalytic domain from the N-terminus. It showed a significant similarity to those of other endoglucanases that belong to family 5 of glycosyl hydrolases. cen1 cDNA was inserted into a yeast expression vector, YEpFLAG-1, and introduced into Saccharomyces cerevesiae. The resulting S. cerevisiae transformant secreted a recombinant En-1 that had enzymatic properties similar to the original En-1. A strong synergistic effect for a degradation of Avicel and phosphoric acid swollen cellulose was observed when recombinant En-1 was used together with a major exo-type cellobiohydrolase I of I. lacteus MC-2.

Amino Acid Sequence↗

Methods for the molecular analysis of cancer. An overview.

Cancer arises as a result of complex and interacting abnormalities. However, over the past 20-25 yr numerous technological advances in molecular biology have led to a dramatic increase in the identification of the molecular processes involved in the development of cancers. The analysis of these molecular changes can be done on different levels: at the DNA or RNA level, or by assessment of posttranscriptional events. This overview discusses the merits of different methods for the analysis of DNA such as SSCP or DGCE. The exciting methods of RNA expression analysis using oligo or cDNA gene chips are also discussed. The importance of methods to analyze posttranscriptional or the effect that altered telomerase activity or methylation status contributes to the phenotype of the cancer cell is emphasized. These techniques will contribute to a better understanding of cancer initiation and progression and will eventually lead toward the development of new molecular-targeted drug therapies.

Animals↗

Genomic organization of human MXI1, a putative tumor suppressor gene.

MXI1, a member of the MYC family of transcription factors, is thought to negatively regulate MYC function and may therefore be a potential tumor suppressor gene. Using detailed restriction mapping and partial DNA sequencing analysis, we have determined the genomic organization of the human MXI1 gene to facilitate a search for mutations that affect MXI1 function. The gene spans a region of approximately 60 kb on chromosome 10q24-q25 and comprises six exons. The correspondence of these exons to previously identified Mxi1 functional domains suggests that alternatively spliced transcripts may regulate Mxi1 functional activity. The presence of a cryptic ATG start codon in exon 2 suggests that a functional protein missing the SIN3-interacting domain (exon 1) may be generated by alternative splicing. Finally, we have identified two polymorphic regions within the MXI1 locus: a polymorphic CA repeat in the third intron and an AAAAC polymorphism in the noncoding region of exon 6. These findings will facilitate the analysis of tumors for the presence of inactivating mutations in MXI1 coding and regulatory sequences.

Alternative Splicing↗

The bldB gene encodes a small protein required for morphogenesis, antibiotic production, and catabolite control in Streptomyces coelicolor.

Mutants blocked at the earliest stage of morphological development in Streptomyces species are called bld mutants. These mutants are pleiotropically defective in the initiation of development, the ability to produce antibiotics, the ability to regulate carbon utilization, and the ability to send and/or respond to extracellular signals. Here we report the identification and partial characterization of a 99-amino-acid open reading frame (ORF99) that is capable of restoring morphogenesis, antibiotic production, and catabolite control to all of the bldB mutants. Of the existing bld mutants, bldB is of special interest because the phenotype of this mutant is the most pleiotropic. DNA sequence analysis of ORF99 from each of the existing bldB mutants identified base changes either within the coding region of the predicted protein or in the regulatory region of the gene. Primer extension analysis identified an apparent transcription start site. A promoter fusion to the xylE reporter gene showed that expression of bldB is apparently temporally regulated and that the bldB gene product is involved in the regulation of its own expression.

Amino Acid Sequence↗

Biological sequence compression algorithms.

Today, more and more DNA sequences are becoming available. The information about DNA sequences are stored in molecular biology databases. The size and importance of these databases will be bigger and bigger in the future, therefore this information must be stored or communicated efficiently. Furthermore, sequence compression can be used to define similarities between biological sequences. The standard compression algorithms such as gzip or compress cannot compress DNA sequences, but only expand them in size. On the other hand, CTW (Context Tree Weighting Method) can compress DNA sequences less than two bits per symbol. These algorithms do not use special structures of biological sequences. Two characteristic structures of DNA sequences are known. One is called palindromes or reverse complements and the other structure is approximate repeats. Several specific algorithms for DNA sequences that use these structures can compress them less than two bits per symbol. In this paper, we improve the CTW so that characteristic structures of DNA sequences are available. Before encoding the next symbol, the algorithm searches an approximate repeat and palindrome using hash and dynamic programming. If there is a palindrome or an approximate repeat with enough length then our algorithm represents it with length and distance. By using this preprocessing, a new program achieves a little higher compression ratio than that of existing DNA-oriented compression algorithms. We also describe new compression algorithm for protein sequences.

Algorithms↗

Method for phosphorothioate antisense DNA sequencing by capillary electrophoresis with UV detection.

The progress of antisense DNA therapy demands development of reliable and convenient methods for sequencing short single-stranded oligonucleotides. A method of phosphorothioate antisense DNA sequencing analysis using UV detection coupled to capillary electrophoresis (CE) has been developed based on a modified chain termination sequencing method. The proposed method reduces the sequencing cost since it uses affordable CE-UV instrumentation and requires no labeling with minimal sample processing before analysis. Cycle sequencing with ThermoSequenase generates quantities of sequencing products that are readily detectable by UV. Discrimination of undesired components from sequencing products in the reaction mixture, previously accomplished by fluorescent or radioactive labeling, is now achieved by bringing concentrations of undesired components below the UV detection range which yields a 'clean', well defined sequence. UV detection coupled with CE offers additional conveniences for sequencing since it can be accomplished with commercially available CE-UV equipment and is readily amenable to automation.

DNA Primers↗

Novel DAX1 mutations in X-linked adrenal hypoplasia congenita and hypogonadotrophic hypogonadism.

OBJECTIVE: Mutations of the DAX1 gene (Dosage-sensitive sex reversal-Adrenal hypoplasia congenita critical region on the X chromosome gene 1), which encodes a novel orphan nuclear receptor, have been identified in patients with X-linked adrenal hypoplasia congenita (AHC) and hypogonadotrophic hypogonadism (HHG). We have investigated two kindreds with AHC and HHG for DAX1 mutations. METHODS: Two kindreds with five affected males, four carrier females and four unaffected males were investigated. The gonadotrophin deficiency in three of the boys was observed to be partial until mid-puberty. DAX1 mutations in the entire 1413 bp coding region were sought by DNA sequence analysis. RESULTS: Two DAX1 mutations, situated within exon 1, were detected. These consisted of an insertional mutation at codon 183 that led to a frameshift and a premature Stop at codon 184, and a missense mutation Leu278Pro that involved a highly conserved leucine residue within the proposed ligand binding domain. Co-segregation of these mutations with the disease in each family, and their absence from 107 alleles in 73 (39 males and 34 females) unrelated control individuals, was demonstrated by allele specific oligonucleotide hybridization (ASO) analysis for the insertional mutation, and by Ban I restriction endonuclease analysis for the missense mutation. CONCLUSIONS: Two novel DAX1 mutations have been detected in two families with adrenal hypoplasia and hypogonadotrophic hypogonadism. The finding of partial gonadotrophin deficiency in the affected males from these families is notable and an early recognition of such a possibility in a patient, which may be facilitated by DAX1 mutational analysis, may help to prevent the sequelae of delayed androgen replacement therapy.

Adolescent↗

Accuracy of automated DNA sequencing: a multi-laboratory comparison of sequencing results.

A double-stranded (ds)DNA template of "unknown" sequence was distributed to approximately 80 core DNA sequencing laboratories by the Association of Biomolecular Resource Facilities (ABRF) for automated DNA sequence analysis. Forty-four different facilities responded with 83 usable sequence submissions. These sequences were grouped by both sequencing protocol (dye-primer or dye-terminator) and whether manually edited or not. The sequences were aligned with the known sequence, and the number of correct base calls, insertions, deletions, no-calls and miscalls were determined for each group. The dye-primer sequencing protocol provided the longest and most accurate sequence. The edited dye-primer data were > 95% accurate out to 400-450 bp, while the edited dye-terminator data could call only 300-350 bases at this accuracy. However, 75% of the laboratories in this sampling preferred the dye-terminator protocol, presumably because of its versatility and convenience. Laboratories that manually edited the automatically called data were able to obtain an additional 100 bases of good sequence when the dye-primer protocol was used. Surprisingly though, editing of dye-terminator results did not increase the amount of good sequence, although the dye-terminator protocol had a superior base-calling ability within the first 100 bases of called sequence.

Autoanalysis↗

Microarray analysis of erythromycin resistance determinants.

AIMS: To develop a DNA microarray for analysis of genes encoding resistance determinants to erythromycin and the related macrolide, lincosamide and streptogramin B (MLS) compounds. METHODS AND RESULTS: We developed an oligonucleotide microarray containing seven oligonucleotide probes (oligoprobes) for each of the six genes (ermA, ermB, ermC, ereA, ereB and msrA/B) that account for more than 98% of MLS resistance in Staphylococcus aureus clinical isolates. The microarray was used to test reference and clinical S. aureus and Streptococcus pyrogenes strains. Target genes from clinical strains were amplified and fluorescently labelled using multiplex PCR target amplification. The microarray assay correctly identified the MLS resistance genes in the reference strains and clinical isolates of S. aureus, and the results were confirmed by direct DNA sequence analysis. Of 18 S. aureus clinical strains tested, 11 isolates carry MLS determinants. One gene (ermC) was found in all 11 clinical isolates tested, and two others, ermA and msrA/B, were found in five or more isolates. Indeed, eight (72%) of 11 clinical isolate strains contained two or three MLS resistance genes, in one of the three combinations (ermA with ermC, ermC with msrA/B, ermA with ermC and msrA/B). CONCLUSIONS: Oligonucleotide microarray can detect and identify the six MLS resistance determinants analysed in this study. SIGNIFICANCE AND IMPACT OF THE STUDY: Our results suggest that microarray-based detection of microbial antibiotic resistance genes might be a useful tool for identifying antibiotic resistance determinants in a wide range of bacterial strains, given the high homology among microbial MLS resistance genes.

Anti-Bacterial Agents↗

Simultaneous targeted alteration of the tyrosinase and c-kit genes by single-stranded oligonucleotides.

We have shown that various forms of oligonucleotides, chimeric RNA-DNA oligonucleotide (RDO) and single-stranded oligodeoxynucleotide (ODN), are capable of chromosomal gene alterations in mammalian cells. Using two ODNs we corrected an inactivating mutation in the tyrosinase gene and introduced an activating mutation into the c-kit gene in a single albino mouse melanocyte. Relying on a pigmentation change caused by tyrosinase gene correction, we determined the frequency of gene targeting events ranging from 2 x 10(-4) to 1 x 10(-3), which is comparable to our previously published data using RDO. However, ODN showed more reproducible gene correction than RDO and produced pigmented cells among 60% of experiments, in comparison with 10% by RDO. DNA sequence analysis of the converted cells revealed that two out of eight individual pigmented clones harbored the mutated c-kit gene. Targeted modification of both genes resulted in the ability of the tyrosinase to convert tyrosine to melanin, and in the constitutive activation of the Kit receptor kinase. Thus, for the first time, we demonstrate the feasibility of simultaneous targeting of two genes in a single cell and show that a selection strategy to identify cells that have undergone a gene modification can enrich the targeted cells with the desired gene alteration.

Amino Acid Sequence↗

Structure and expression of the promoter for the human type II transforming growth factor-beta receptor.

The type II TGF-beta receptor is a serine/threonine kinase whose expression is essential for the action of TGF-beta. In this paper, we describe the cloning and expression of the human type II TGF-beta promoter. DNA sequence analysis indicates that the region near the transcription initiation site lacks a TATA and CAAT box but contains Sp1 binding sites. A similar promoter type has been identified in the 5' flanking regions of genes that code for certain other growth factor receptors. Transfection of the promoter (888 bp fragment) mediated transcriptional activity in bovine vascular smooth muscle cells and human lung fibroblasts.

Animals↗

Beet yellows closterovirus: complete genome structure and identification of a leader papain-like thiol protease.

The sequence of 8734 nucleotides (nt) from the 5'-end of the beet yellows closterovirus (BYV) RNA was determined to complete the 15,480-nt sequence of the virus genome. The 5'-terminal two-thirds of the sequence are occupied by two overlapping open reading frames (ORFs) 1a and 1b, encoding products with calculated M(r) of 295K and 48K, respectively. The RNA sequence surrounding the stop codon in ORF 1a shows structural elements typical of ribosomal frameshifting signals in a number of animal and plant viruses. It is predicted that the ORF 1b product is expressed via a +1 ribosomal frameshifting as the 348K ORF 1a/1b fusion protein. This putative protein contains the array of methyltransferase, RNA helicase, and RNA-dependent RNA polymerase domains that is conserved in the Sindbis-like supergroup of positive-strand RNA viruses. The 348K protein of BYV is longer than the putative replicases of the most closely related viruses (tobra- and tobamoviruses) by about 1300 amino acids distributed between two unique regions, one at the N-terminus, and the other in the central portion. The N-terminal domain showed sequence similarity to the helper component papain-like protease of potyviruses. By using in vitro translation of the T7 transcripts encoding the N-terminal 92K peptide of the BYV ORF 1a product, we found that the N-terminal fragment of 588 amino acids is released from the translation product by cleavage at the Gly-Gly dipeptide. Site-directed mutagenesis of either of the predicted catalytic residues Cys-509 and His-569 or of the Gly-588 at the cleavage site completely abolished the cleavage. The central unique region of the 348K protein contains a domain distantly resembling the aspartic protease of HIV and other lentiviruses. As shown previously, the 3'-terminal portion of the BYV genome encompasses seven more ORFs, one of which codes for a protein related to the HSP70 cell heat shock proteins, whereas two others encode the capsid protein and its diverged copy. Thus, despite the apparent evolutionary relationship with Sindbis-like viruses, BYV comprises a collection of genomic modules absorbed from different sources and has a unique expression strategy.

Amino Acid Sequence↗

A neutral protease from Bacillus nematocida, another potential virulence factor in the infection against nematodes.

A neutral protease (npr) (designated Bae16) toxic to nematodes was purified to homogeneity from the strain Bacillus nematocida. The purified protease showed a molecular mass of approximately 40 kDa and displayed optimal activity at 55 degrees C, pH 6.5. Bioassay experiments demonstrated that this purified protease could destroy the nematode cuticle and its hydrolytic substrates included gelatin and collagen. The gene encoding Bae16 was cloned, and the deduced amino acid sequence showed 94% sequence identity with npr gene from B. amyloliquefaciens, but had low similarity (13-43%) with the previously reported virulence serine proteases from fungi or bacteria, which reflected their differences. Recombinant mature Bae16 (rm-Bae16) was expressed in Escherichia coli BL21 using pET30 vector system, and its nematicidal activity confirmed that Bae16 could be involved in the infection process. Our present study revealed that the npr besides the known alkaline serine protease could serve as a potential virulence factor in the infection against nematodes, furthermore, the two proteases with different characteristics produced by the same strain co-ordinated efforts to kill nematodes. These data helped to understand the interaction between this bacterial pathogen and its host.

Animals↗

Cloning of novel laccase isozyme genes from Trametes sp. AH28-2 and analyses of their differential expression.

Three novel laccase isozyme genes, lacA, lacB, and lacC, have been identified from basidiomycete Trametes sp. AH28-2. These genes display a high similarity with other basidiomycete laccases at the amino acid level. An inferred TATA box and several putative CAAT, MRE, XRE, and CreA consensus sequences were identified in the lacA, lacB, and lacC promoter regions. Different from the TATA boxes of lacA and lacB at about -100, the TATA box of lacC is located at -172. For all the isozymes, copper ion is essential for laccase synthesis in Trametes sp. AH28-2. More interestingly, different aromatic compounds can selectively induce the production of distinct laccase isozymes, with o-toluidine inducing the expression of laccase A (LacA) while 3,5-dihydroxytoluene mainly stimulating the production of laccase B (LacB). Quantitative reverse transcriptase-polymerase chain reaction showed that the accumulation of laccase messenger RNA transcripts is accompanied by the increase of corresponding enzyme activity in cultures. The glucose-repression effect on laccase expression in Trametes sp. AH28-2 was also observed. Furthermore, lower Cu2+ concentration (lower than 0.5 mM) can induce LacA and a novel laccase (LacC), and the latter will disappear when Cu2+ concentration is increased up to 1-2 mM. Upon induction by 3,5-dihydroxytoluene, the ratio of LacA to LacB decreased in the later phase of induction.

Amino Acid Sequence↗

Two novel asparaginyl endopeptidase-like cysteine proteinases from the protist Trichomonas vaginalis: their evolutionary relationship within the clan CD cysteine proteinases.

Cysteine proteinases (CPs) are important virulence factors of the protozoan parasite Trichomonas vaginalis. A total of six genes coding for cathepsin L-like CPs belonging to clan CA have been identified in T. vaginalis. At least 23 distinct spots with proteolytic activity have been detected by two-dimensional (2-D) substrate gel electrophoresis from in vitro grown parasites; however, only few of them have been characterized. In this work, we detected six spots with proteolytic activity and molecular weights between 25 and 35 kDa. The six proteinases correspond to two distinct CP families: the papain-like family, represented by four spots with pIs between 4.5 and 5.5; and the legumain-like family represented by two spots with pI 6.3 and 6.5. Next, we obtained two cDNAs encoding for legumain-like CPs from T. vaginalis, which were named Tvlegu-1 and Tvlegu-2. The size of these cDNA clones were 1225 and 1364 bp, which encoded for 388 and 415 amino acids, respectively. Their putative translation products have molecular masses of 42.8 and 47.2 kDa, corresponding to inactive legumain-like CP precursors. The two sequences share approximately 40% identity at the amino acid level. These protein products can be classified within a branch of the legumain-like family in clan CD cysteine proteinases due to their sensitivity to specific proteinases inhibitors, their DNA sequences, and phylogenetic reconstruction. However, they do not correspond either to the typical asparaginyl endopeptidase or the glycosylphosphatidylinositol (GPI): protein transamidase subfamilies. These results suggest that the TVLEGU-1 and TVLEGU-2 peptidases are likely to be part of a new subfamily within the legumain-like family of clan CD cysteine proteinases. Furthermore, they could be one of the missing links between prokaryotic and eukaryotic CPs in clan CD enzymes.

Amino Acid Sequence↗

CD109 represents a novel branch of the alpha2-macroglobulin/complement gene family.

We report here the genomic organization and phylogenic relationships of CD109, a member of the the alpha2-macroglobulin/complement (AMCOM) gene family. CD109 is a GPI-linked glycoprotein expressed on endothelial cells, platelets, activated T-cells, and a wide variety of tumors. We cloned full-length CD109 cDNA from the mammalian U373 cell line by RT-PCR and performed analysis of its corresponding genomic sequence. The CD109 cDNA spans 128 kb of chromosome 6q with its 33 exons constituting approximately 3.3% of the total CD109 genomic sequence. Sequence analysis revealed that CD109 contains specific motifs in its N-terminus, that are highly conserved in all AMCOM members. CD109 also shares motifs with certain other AMCOM members including: (1) a thioester 'GCGEQ" motif, (2) a furin site of four positively charged amino acids, and (3) a double tyrosine near the C-terminus. Based on a phylogenic analysis of human CD109 with other human homologs as well as orthologs from other mammalian species, C. elegans (ZK337.1) and E. coli homologs, we propose CD109 represents a novel and independent branch of the alpha2-macroglobulin/complement gene family (AMCOM) and may be its oldest member.

Amino Acid Sequence↗

Correcting sequencing errors in DNA coding regions using a dynamic programming approach.

This paper presents an algorithm for detecting and 'correcting' sequencing errors that occur in DNA coding regions. The types of sequencing errors addressed are insertions and deletions (indels) of DNA bases. The goal is to provide a capability which makes single-pass or low-redundancy sequence data more informative, reducing the need for high-redundancy sequencing for gene identification and characterization purposes. This would permit improved sequencing efficiency and reduce genome sequencing costs. The algorithm detects sequencing errors by discovering changes in the statistically preferred reading frame within a putative coding region and then inserts a number of 'neutral' bases at a perceived reading frame transition point to make the putative exon candidate frame consistent. We have implemented the algorithm as a front-end subsystem of the GRAIL DNA sequence analysis system to construct a version which is very error tolerant and also intend to use this as a testbed for further development of sequencing error-correction technology. Preliminary test results have shown the usefulness of this algorithm and also exhibited some of its weakness, providing possible directions for further improvement. On a test set consisting of 68 human DNA sequences with 1% randomly generated indels in coding regions, the algorithm detected and corrected 76% of the indels. The average distance between the position of an indel and the predicted one was 9.4 bases. With this subsystem in place, GRAIL correctly predicted 89% of the coding messages with 10% false message on the 'corrected' sequences, compared to 69% correctly predicted coding messages and 11% falsely predicted messages on the 'corrupted' sequences using standard GRAIL II method (version 1.2).(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms↗