Search PubMedSearch

SEARCH · Search PubMed

Results for “DNA sequence analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

DNA sequence analysis of the type-common glycoprotein-D genes of herpes simplex virus types 1 and 2.

The DNA sequences for the coding and flanking regions of the type-common glycoprotein-D (gD) genes of herpes simplex virus (HSV) types 1 and 2 have been determined. The resultant protein sequences are approximately 80% homologous. Both gD proteins are 393 amino acids long and both have maintained three identical potential glycosylation sites. Amino acid changes are found throughout the proteins, with the majority of changes located in the amino and carboxyl/termini. Most of the amino acid differences were found to be conservative. Hydropathy analysis, which determines hydrophobic and hydrophilic regions, reveals a remarkable structural similarity between the proteins. Examination of 5' flanking sequences demonstrates extensive DNA sequence homology adjacent to the start of gD gene transcription. In addition, another homologous noncoding region was found 3' to the gD gene. This second homologous sequence is 5' to a 1.6-kb transcription unit.

Amino Acid Sequence

Location by DNA sequence analysis of cop mutations affecting the number of plasmid copies of prophage P1.

The DNA sequence of six P1 cop mutants, which are altered in the control of copy number of the plasmid prophage, was compared to that of P1 wild type. Each cop mutant differs from the wild type by a single base substitution. All of these substitutions are located within a 400 base pair region of P1 DNA that also encodes rep, a gene whose product is required for P1 replication.

Amino Acid Sequence

Microcomputer programs for back translation of protein to DNA sequences and analysis of ambiguous DNA sequences.

Three computer programs are described which may be used to translate a DNA sequence into a protein sequence, back translate the protein sequence into an ambiguous DNA sequence, and then do pattern searching in the ambiguous sequence. The programs are written in the C programming language, have been compiled to run on a microcomputer under the CP/M 80 operating system, and may be copied in binary format through a modem. They are also to become available for the IBM/PC.

Amino Acid Sequence

RNA splicing in yeast mitochondria: DNA sequence analysis of mit- mutants deficient in the excision of introns aI1 and aI2 of the gene for subunit I of cytochrome c oxidase.

We have characterized two yeast mutants deficient in the splicing of transcripts of the mitochondrial gene for cytochrome c oxidase subunit I (coxI). Both map to the first intron (aI1). RNA blot analysis shows that in addition to a reduced (mutant M15-190) or blocked (mutant M12-193) excision of the mutated intron aI1, the mutants are unable to excise the adjacent aI2 intron, the reading frame of which displays an amino acid sequence similarity to aI1. Splicing of the downstream introns is not affected, however. Sequence analysis of the first mutant DNA (M12-193) reveals a premature termination of the intron-encoded open reading frame, followed by two alterations at a short distance downstream. The other (M15-190) contains 11 separate changes. Although these occur in the intron reading frame, their main effect on RNA splicing may be exerted through the disturbance of intron secondary structure proposed for the 5' end of several group II introns. The implications of these findings in relation to maturase function and structure of intron aI1 are discussed.

Base Sequence

A DNA sequence analysis program for the Apple Macintosh.

This paper describes a new set of programs for analyzing DNA sequences using the Apple Macintosh computer, a computer ideally suited for this kind of analysis. Because of the Macintosh interface and the availability of high quality software-only speech synthesis, these programs are truly easy to use. Instead of typing in commands, the user directs the program by making selections with the mouse, thereby eliminating most typographical and syntax errors. Output options are selected by "pressing buttons" and then clicking "OK" with the mouse. DNA sequences are confirmed by having the program speak them. The high resolution graphics on the Macintosh not only allow for explanatory diagrams to be used to aid in deciding on input parameters, but can be used to produce slides for presentations and figures for papers. Because of the clipboard and the ability of the Macintosh to readily share data among different applications, data can be saved for use directly in word processing documents (e.g. manuscripts).

Animals

A convenient and adaptable package of DNA sequence analysis programs for microcomputers.

We describe a package of DNA data handling and analysis programs designed for microcomputers. The package is convenient for immediate use by persons with little or no computer experience, and has been optimized by trial in our group for a year. By typing a single command, the user enters a system which asks questions or gives instructions in English. The system will enter, alter, and manage sequence files or a restriction enzyme library. It generates the reverse complement, translates, calculates codon usage, finds restriction sites, finds homologies with various degrees of mismatch, and graphs amino acid composition or base frequencies. A number of options for data handling and printing can be used to produce figures for publication. The package will be available in ANSI Standard FORTRAN for use with virtually any FORTRAN compiler.

Amino Acid Sequence

gm: a practical tool for automating DNA sequence analysis.

The gm (gene modeler) program automates the identification of candidate genes in anonymous, genomic DNA sequence data. gm accepts sequence data, organism-specific consensus matrices and codon asymmetry tables, and a set of parameters as input; it returns a set of models describing the structures of candidate genes in the sequence and a corresponding set of predicted amino acid sequences as output, gm is implemented in C, and has been tested on Sun, VAX, Sequent, MIPS and Cray computers. It is capable of analyzing sequences of several kilobases containing multi-exon genes in less than 1 min execution time on a Sun 4/60.

Algorithms

Osteonectin promoter. DNA sequence analysis and S1 endonuclease site potentially associated with transcriptional control in bone cells.

To understand the basis of osteonectin (SPARC) transcriptional regulation, we have isolated a bovine genomic clone (lambda Og15) encoding exon 1 and 15 kilobase pairs (kb) of flanking DNA. Direct RNA sequencing of the 5' end of the osteonectin message showed it contained a sequence identical to that of a 2.4-kb EcoRI-BamHI fragment located midway in the clone lambda Og15. The results indicate exon 1 is located 10 kb away from exon 2 in the bovine genome. The DNA sequence unit CCTG is repeated five times in exon 1 which is composed exclusively of untranslated sequence. Sequence analysis of the 5'-flanking DNA revealed the presence of many regulatory motifs including a "GC" box with four overlapping SP1 consensus sequences. Immediately downstream from the GC box is a 72-base pair purine-rich stretch composed primarily of direct repeats of the sequence motifs GGGGA and GGA (GAGA box). Digestion of the flanking DNA in vitro with S1 endonuclease showed a site for the enzyme at position -55 which is just 3' to the GAGA box. Chimeric chloramphenicol acetyltransferase constructs were prepared containing the S1-sensitive site and showed substantial transcriptional activity in UMR-106 and fetal and adult human bone cells which are known to be high producers of the protein. The results indicate a potential regulatory activity of the S1 site in osteonectin gene activation.

Animals

DNA sequence analysis of the Hind III M fragment from Chinese vaccine strain of vaccinia virus.

The complete DNA sequence of the Hind III M fragment of vaccinia virus (VV) Tian Tan strain genome was determined by the dideoxynucleotide chain termination method. Three open reading frames (ORFs) were identified in the complementary strand of the sequence, comprised of 2218bp. Among them, ORF K1 initiates its transcription at -45 of the Hind III K fragment. The deduced peptide encoded by K1 contains 284 amino acids with a calculated molecular weight of 32.48 KDa. Its sequence is homologous to the host range protein of VV Copenhagen strain; the variation is only 2.46% at the amino acid level. ORF M2 could encode a peptide of 21.94 KDa with 196 amino acids. This gene was shown to be homologous to that of the 23 KDa peptide of herpes simplex virus type I. A non-coding region of 204bp located between K1 and M2 is rich in palindromic structures. ORF M1 extends its 3' terminus into the Hind III N fragment. Within the M fragment, M1 can only encode 212 amino acids. The major part of ORF M1 is very similar to the M portion of a possible alpha-amanitin resistance gene isolated from VV-WR strain. This work provides a molecular foundation in the construction of a new insertion vector for the preparation of a recombinant vaccinia virus to be used as a polyvalent live vaccine.

Base Sequence

DNA sequence analysis of the 24.5 kilobase pair cytochrome oxidase subunit I mitochondrial gene from Podospora anserina: a gene with sixteen introns.

The DNA sequence of a 26.7 Kilobase pair (10(3) base pairs = 1 Kb) region of the mitochondrial genomes of races s and A from Podospora anserina was determined. Within this region, the 24.5 Kb cytochrome oxidase subunit I gene was located and its exon sequences determined by computer analysis comparisons with other fungal genes. The Podospora COI gene was interrupted by two group II introns (one in race s) and fourteen group I introns ranging in size from about 2.2 Kb to 404 bp. Earlier studies on secondary structure analysis, as well as comparison of their open reading frames (ORFs), showed that the two group II introns were closely related. The fourteen group I introns were representatives of three subgroupings (IB, C and a new category, subgroup ID). Two of these group I introns were separated by just a single exon codon. The analysis of all these introns is discussed in comparison with other fungal introns as well as with the known Podospora anserina introns.

Amino Acid Sequence

DNA sequence analysis of the apocytochrome b gene of Podospora anserina: a new family of intronic open reading frame.

The 5,969 bp (base pair) DNA sequence of the apocytochrome b mitochondrial (mt) gene of race A Podospora anserina was located in a 8.5 Kbp region. This gene contained a 2,499 bp subgroup IB and a 1,306 bp subgroup ID intron as well as a 990 bp subgroup IB intron which is present in race A but not race s. The large subgroup IB intron and the race A specific IB intron both contained potential alternate splice sites which brought their open reading frames into phase with their upstream exon sequences. All three introns were compared with regard to their secondary structures and open reading frames to the other 30 group I introns in Podospora anserina, as well as to other fungal introns. We detected a new family of intronic ORFs comprising seven P. anserina introns, several N. crassa introns, as well as the T4td bacteriophage intron. Sequence similarities to intron-encoded endonucleases were noteworthy. The DNA sequences reported here and in the accompanying paper complete the analysis of race s and race A mitochondrial DNA.

Amino Acid Sequence

Physical mapping and DNA sequence analysis of the rifampicin resistance locus in vaccinia virus.

Rifampicin has been shown to inhibit the maturation of poxviruses at a discrete step in envelope formation (Moss et al., 1969; Pennington et al., 1970; Nagayama et al., 1970; Grimley et al., 1970). A rifampicin-resistant vaccinia virus mutant (RifR) was selected for its ability to grow in the presence of 100 micrograms/ml of rifampicin. Utilizing intact DNA or endonuclease restricted cloned DNA subfragments derived from the RifR mutant virus, the locus specifying rifampicin resistance was physically mapped by marker rescue analysis leftward of the unique XhoI site within the HindIII D fragment. DNA sequencing of a 445 bp fragment encompassing this region revealed an AT to GC transition when compared with the equivalent wild-type DNA fragment. Analysis of the six potential open reading frames within the 445-bp fragment indicated only one available open reading frame. On this basis, the rifampicin-resistant vaccinia virus mutant was shown to have a codon transition from asparagine to aspartic acid.

Amino Acid Sequence

DNA sequence analysis of two bovine immunoglobulin CH gamma pseudogenes.

A bovine calf liver DNA library in lambda 2001 bacteriophage has been screened with a human Ig gamma 4 heavy-chain constant-region gene probe. Four hybridizing clones have been identified, and the DNA sequences in two of these, which have high homology with CH gamma genes, are reported here. Within the bovine sequences, four separate exons can be identified, corresponding to the three CH domains and the hinge of gamma heavy-chain genes. Both of these genes contain atypical sequences around one or more of their exon/intron boundaries with consequent loss of splice sites, indicating that these are probably gamma pseudogenes. One sequence codes for a C-terminal peptide which matches the 18-mer C-terminal heavy-chain peptide of bovine serum IgG2, the other encodes a C-terminal peptide unknown in the bovine. These results suggest that evolutionary duplication of CH gamma genes has occurred in the bovine.

Amino Acid Sequence

Random subcloning of sonicated DNA: application to shotgun DNA sequence analysis.

A method for producing random subclones using sonication to fragment the DNA is presented. The sonication is combined with enzymatic repair of the fragment ends and a rigorous size fractionation step to prepare subclones of relatively homogeneous and specific size. Under some conditions sonication is shown to shear A + T-rich sequences preferentially, although under most conditions it will create a random subclone library. The use of these subclone libraries for an improved "shotgun" DNA sequencing strategy is tested on a 17.2-kb (kilobase) fragment of Epstein-Barr virus.

Base Sequence

Mutational specificities of environmental carcinogens in the lacl gene of Escherichia coli H. V: DNA sequence analysis of mutations in bacteria recovered from the liver of Swiss mice exposed to 1,2-dimethylhydrazine, azoxymethane, and methylazoxymethanolacetate.

The host-mediated assay (HMA) was used to determine the spectra of mutations induced in the lacl gene of Escherichia coli cells recovered from the livers of Swiss mice exposed to the carcinogens 1,2-dimethylhydrazine (SDMH), azoxymethane (AOM), and methylazoxymethanolacetate (MAMA). These spectra were further compared with changes induced by dimethylnitrosamine (DMNA) in the HMA methodology. A total of 177 independent lacl mutations arising in the HMA following exposure to SDMH, AOM, and MAMA were analyzed. Single-base substitutions accounted for 97% of all mutations analyzed. The vast majority of the single-base substitutions consisted of G:C----A:T transitions (94% of all mutations). The remaining mutations consisted of A:T----G:C transitions (3% of all mutations) while non-base substitutions accounted for only 3% of the total mutagenesis. The latter mutations consisted of one frameshift mutation and four lacO deletions. The distribution of G:C----A:T transitions induced by the three chemicals in the first 200 bp of the lacl gene was not random, but rather clustered at sites where a target guanine was flanked at the 5' site by a purine residue.

1,2-Dimethylhydrazine

DNA sequence analysis of a 5.27-kb direct repeat occurring adjacent to the regions of S-episome homology in maize mitochondria.

The DNA sequence of the 5270-bp repeated DNA element from the mitochondrial genome of the fertile cytoplasm of maize has been determined. The repeat is a major site of recombination within the mitochondrial genome and sequences related to the R1(S1) and R2(S2) linear episomes reside immediately adjacent to the repeat. The terminal inverted repeats of the R1 and R2 homologous sequences form one of the two boundaries of the repeat. Frame-shift mutations have introduced 11 translation termination codons into the transcribed S2/R2 URFI gene. The repeated sequence, though recombinantly active, appears to serve no biological function.

Amino Acid Sequence

DNA sequence analysis of endoglucanase genes from Pseudomonas fluorescens subsp. cellulosa and Pseudomonas sp. NCIB 8634.

The DNA of two previously isolated recombinant clones, one from Pseudomonas sp. NCIB 8634 (= Cellvibrio mixtus) (pPC71) and another from Pseudomonas fluorescens subsp. cellulosa (pPFC4) that express endoglucanase activity in E. coli was sequenced. Plasmid pPC71 had three open reading frames, two of which include portions of plasmid pBR322. The third open reading frame occurs entirely within the Pseudomonas DNA insert and encodes a protein with a molecular mass of 5845 Da. The DNA insert in pPFC4 was found to contain an open reading frame (PFC-ORF) that encodes a protein of 32189 Da. The major endoglucanase produced in E. coli cells carrying pPFC4 is about 30,000 Da. It is concluded that PFC-ORF encodes this endoglucanase. Both ribosome and catabolite gene activator protein binding sites lie upstream from the initiating codon of PFC-ORF. An interesting feature of the PFC-ORF protein is the presence of amino acid motifs Val-Ser-Ser-Ser-Ser and Val-Val-Ser-Ser-Ser-Ser-Ser that occur within a 25 amino acid span.

Amino Acid Sequence