Search PubMedSearch

SEARCH · Search PubMed

Results for “Sequence Analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Identification and location of nine T5 bacteriophage tRNA genes by DNA sequence analysis.

Sequence analysis of two DNA fragments generated from bacteriophage T5 DNA by restriction with Hpa I and Hae III has resulted in the detection and localization of nine tRNA genes (His, two Ser genes, Leu, Val, Lys, fMet, Pro, and Ile). The genes which code for tRNAs His and Leu are partials, whereas the remaining genes are complete. A majority of the tRNA genes are located in close proximity to one another. A unique feature of the Pro and Ile genes is that their DNA sequence overlap.

Bacteriophages

A controversy concerning the structure of the mitochondrial genome of a yeast petite mutant: resolution by sequence analysis.

Sequence analysis was used to define the repeat unit that constitutes the mitochondrial genome of a petite (rho-) mutant of the yeast Saccharomyces cerevisiae. This mutant has retained and amplified in tandem a 2,547 bp segment encompassing the second exon of the oxi3 gene excised from wild-type mtDNA between two direct repeats of 11 nucleotides. The identity of the mtDNA segment retained in this petite has recently been questioned (van der Veen et al., 1988). The results presented here confirm the identity of this mtDNA segment to be that determined previously by restriction mapping (Carignani et al., 1983).

Base Sequence

Primary structure of the porin protein of Haemophilus influenzae type b determined by nucleotide sequence analysis.

Sequencing techniques for single- and double-stranded DNA were used to determine the nucleotide sequence of the gene encoding P2, the major outer membrane (porin) protein of Haemophilus influenzae type b (Hib). The open reading frame encoding the P2 protein comprised 361 amino acid codons. Comparison of the inferred amino acid sequence with data obtained by amino acid sequencing of the N terminus of the mature or fully processed P2 protein revealed that this protein has a signal peptide composed of 20 amino acids. N-terminal amino acid sequencing of tryptic peptides derived from purified P2 allowed direct identification of 158 of the 341 amino acids in the fully processed P2 protein; there was 100% correlation between these amino acid sequences and that inferred from the nucleotide sequence. The amino acid sequence of Hib P2 protein had 23 to 25% homology with the sequence of the OmpF porin of Escherichia coli and with that of the Neisseria gonorrhoeae porin P.IA. Codon usage in the Hib P2 gene was significantly different from that observed for a gene encoding a porin of E. coli. DNA hybridization studies indicated that there is a single copy of the P2 gene in the Hib chromosome. The availability of the nucleotide and amino acid sequences for the Hib P2 protein will facilitate investigation of the antigenic characteristics and structure-function relationship of this porin.

Amino Acid Sequence

Cloning and sequence analysis of the human parathyroid hormone gene region.

A region of 50 kb around the human PTH gene was cloned and mapped by restriction analysis. Sequence analysis was performed and 3270bp determined, completing the sequence of the gene. The nucleotide sequence was analysed with regard to homology between human, bovine and rat PTH genes, and various potential cis-acting regulatory elements were identified. The gene region lacks an obvious CpG island. The PTH gene region in patients suffering from (pseudo)-hypoparathyroidism was investigated by Southern blotting. No detectable alteration in the fragment patterns was observed. Results of segregation analysis in families with affected individuals was inconclusive.

Animals

Amino-acid sequence of lac repressor from Escherichia coli. Isolation, sequence analysis and sequence assembly of tryptic peptides and cyanogen-bromide fragments.

The lac repressor from Escherichia coli, composed of four identical subunits with a molecular weight of 37160, was carboxymethylated and fragmented by tryptic digestion and cyanogen bromide treatment. Using ion-exchange chromatography, gel filtration and preparative thin-layer electrophoresis and chromatography 29 of the 30 tryptic peptides were isolated in pure form. Direct Edman degradation and the dansyl-Edman technique were used to determine the sequence of the small tryptic peptides. Special emphasis was put on the sequence determination of the six large tryptic fragments which together account for 177 residues, corresponding to 51% of the repressor subunit with its 347 residues. The large tryptic fragments were analyzed after fragmentation with chymotrypsin, thermolysin and dipeptidyl aminopeptidase I. Thus the sequence of all 30 tryptic peptides could be deduced. The complete sequences of all cyanogen bromide fragments were deduced from peptides obtained by tryptic, chymotryptic and thermolytic digestion of the individual fragments and by automated stepwise Edman degradation of lac repressor and of the large cyanogen bromide fragments. The order of the cyanogen bromide fragments was given by overlapping tryptic peptides. The resulting amino acid composition of the monomer is Asp15, Asn11, Thr18, Ser30, Glu14, Gln27, Pro13, Gly22, Ala44, Cys3, Val33, Met9, Ile17, Leu40, Tyr8, Phe4, Trp2, Lys11, His7, Arg19. The sequence of lac repressor shows no similarities with that of other proteins known to bind to DNA or RNA. The N-terminal 55 residues contain two homologous regions. This part of the sequence which is involved in lac operator binding might have been formed by gene duplication.

Amino Acid Sequence

DNA sequence analysis. Terminal sequences of bacteriophage phi80.

Sequences of the cohesive ends and the 3'-terminal regions of phi80 DNA have been determined. Sequences of the cohesive ends were obtained through the use of two standard methods. The first method involved the incorporation of all four labeled deoxyribonucleotides into the phi80 cohesive ends using DNA polymerase I. The DNA was then partially digested with micrococcal nuclease or pancreatic DNase. The products were separated by two-dimensional electrophoresis and characterized by composition, 3'-terminal, and nearest neighbor analyses. The second method involved partial incorporation using one, two, or three labeled deoxyribonucleotides followed by similar analyses. Sequences of the double-stranded regions adjacent to the cohesive ends were determined by three new methods. These methods were: (a) the DNA was specifically labeled at the 3' terminus and then partially degraded. Labeled oligonucleotide products were sequenced by their mobilities on various separation systems. (b) The cohesive ends were enlarged by limited degradation with exonuclease III. After this treatment, the DNA was partially repaired with labeled nucleotides, digested, and the products were analyzed. (c) A synthetic ologonucleotide primer was bound to phi80 DNA which had been repaired with DNA polymerase I, and then partially digested with lambda-exonuclease. The primer was extended into the region of interest by partial repair with labeled nucleotides. The extended primer was isolated and analyzed.

Base Sequence

Epidemiologic and historical relationships among 87 rabies virus isolates as determined by limited sequence analysis.

Nucleotide sequence analysis of a 200-bp region of the nucleoprotein (N) gene of rabies virus differentiated unique genetic groups of rabies virus from samples collected in areas where dog rabies is enzootic in Asia, Africa, Europe, and the Americas. Patterns of nucleotide sequence identified for an outbreak area were conserved in samples collected over three decades. Epidemiologic relationships among isolates were determined by patterns of conserved nucleotide sequence, and the degree of sequence divergence between samples from separate outbreak areas were measured. This approach suggested that a historical reconstruction of events leading to the introduction of rabies into an area would be possible. In this broader view of rabies epidemiology, the cultural legacy of European exploration and colonization may have also included zoonotic disease.

Amino Acid Sequence

Rigorous pattern-recognition methods for DNA sequences. Analysis of promoter sequences from Escherichia coli.

The basic nature of the sequence features that define a promoter sequence for Escherichia coli RNA polymerase have been established by a variety of biochemical and genetic methods. We have developed rigorous analytical methods for finding unknown patterns that occur imperfectly in a set of several sequences, and have used them to examine a set of bacterial promoters. The algorithm easily discovers the "consensus" sequences for the -10 and -35 regions, which are essentially identical to the results of previous analyses, but requires no prior assumptions about the common patterns. By explicitly specifying the nature of the search for consensus sequences, we give a rigorous definition to this concept that should be widely applicable. We also have provided estimates for the statistical significance of common patterns discovered in sets of sequences. In addition to providing a rigorous basis for defining known consensus regions, we have found additional features in these promoters that may have functional significance. These added features were located on either side of the -35 region. The pattern 5', or upstream, from the -35 region was found using the standard alphabet (A, G, C and T), but the pattern between the -10 and the -35 regions was detectable only in a sub-alphabet. Recent results relating DNA sequence to helix conformation suggest that the former (upstream) pattern may have a functional significance. Possible roles in promoter function are discussed in this light, and an observation of altered promoter function involving the upstream region is reported that appears to support the suggestion of function in at least one case.

Base Sequence

PEPPLOT, a protein secondary structure analysis program for the UWGCG sequence analysis software package.

We describe a program for the analysis of protein secondary structure that operates with the Sequence Analysis Software Package of the University of Wisconsin Genetics Computer Group (UWGCG). The program produces both graphic and printed output. Structure prediction using the Chou and Fasman and Robson et al methods, and hydropathy analysis by the method of Kyte and Doolittle are included along with a simplified method of hydrophobic moment analysis. The power of the program is the coordinated presentation of many different kinds of structural information on the same plot.

Amino Acid Sequence

Identifying constraints on the higher-order structure of RNA: continued development and application of comparative sequence analysis methods.

Comparative sequence analysis addresses the problem of RNA folding and RNA structural diversity, and is responsible for determining the folding of many RNA molecules, including 5S, 16S, and 23S rRNAs, tRNA, RNAse P RNA, and Group I and II introns. Initially this method was utilized to fold these sequences into their secondary structures. More recently, this method has revealed numerous tertiary correlations, elucidating novel RNA structural motifs, several of which have been experimentally tested and verified, substantiating the general application of this approach. As successful as the comparative methods have been in elucidating higher-order structure, it is clear that additional structure constraints remain to be found. Deciphering such constraints requires more sensitive and rigorous protocols, in addition to RNA sequence datasets that contain additional phylogenetic diversity and an overall increase in the number of sequences. Various RNA databases, including the tRNA and rRNA sequence datasets, continue to grow in number as well as diversity. Described herein is the development of more rigorous comparative analysis protocols. Our initial development and applications on different RNA datasets have been very encouraging. Such analyses on tRNA, 16S and 23S rRNA are substantiating previously proposed associations and are now beginning to reveal additional constraints on these molecules. A subset of these involve several positions that correlate simultaneously with one another, implying units larger than a basepair can be under a phylogenetic constraint.

Base Sequence

Direct cloning and sequence analysis of enzymatically amplified genomic sequences.

A method is described for directly cloning enzymatically amplified segments of genomic DNA into an M13 vector for sequence analysis. A 110-base pair fragment of the human beta-globin gene and a 242-base pair fragment of the human leukocyte antigen DQ alpha locus were amplified by the polymerase chain reaction method, a procedure based on repeated cycles of denaturation, primer annealing, and extension by DNA polymerase I. Oligonucleotide primers with restriction endonuclease sites added to their 5' ends were used to facilitate the cloning of the amplified DNA. The analysis of cloned products allowed the quantitative evaluation of the amplification method's specificity and fidelity. Given the low frequency of sequence errors observed, this approach promises to be a rapid method for obtaining reliable genomic sequences from nanogram amounts of DNA.

Base Sequence

ADSP--a new package for computational sequence analysis.

A new protein sequence analysis package, ADSP, is described, of which the SOMAP Screen-Oriented Multiple Alignment Procedure forms an integral part. ADSP (Algorithms and Data Structures for Protein sequence analysis) incorporates facilities to generate potent pattern-recognition discriminators and offers four algorithms with which to scan any NBRF format sequence database: the package has been designed, in particular, to interface with the OWL composite sequence database, one of the largest, distributed non-redundant sources of sequence data of its kind. The system incorporates a powerful method for compound feature analysis, which provides the basis for characterizing and predicting the occurrence of complete protein superfamilies and for pinpointing the emergence of related sub-families. Used iteratively, the approach allows diagnostic performance to be rigorously refined and its efficacy to be assessed both qualitatively and quantitatively, and results in the generation of refined structural or functional features suitable for entry into a database: this compilation of characteristic signatures is distinct from, but complementary to, widely used compendia of pattern templates such as PROSITE.

Algorithms

Fluorescent labeling of cysteinyl residues to facilitate electrophoretic isolation of proteins suitable for amino-terminal sequence analysis.

A protein labeling procedure which enables detection of subpicomole quantities of proteins on sodium dodecyl sulfate (SDS)-polyacrylamide gels is described. Proteins are rendered fluorescent by reduction of disulfide bonds with dithiothreitol followed by alkylation with 5-N-[(iodoacetamidoethyl)amino]naphthalene-1-sulfonic acid (5-I-AEDANS) or 5-iodoacetamido-fluorescein. Labeling is performed prior to electrophoresis, thus eliminating the need for staining with dyes and destaining after electrophoresis. As little as 375 fmol (25 ng) of prelabeled bovine serum albumin can be readily visualized after electrophoresis. Bands are still visible after electrophoretic transfer to nitrocellulose. Simultaneous labeling of proteins in complex mixtures is possible using this technique. This includes cysteine containing proteins of disrupted Newcastle disease virus. The magnitudes of the molecular weight increases which occur upon labeling reflect the cysteine contents of proteins. The mode of chemical modification for the prelabeling procedure was chosen because of its compatibility with analytical techniques, such as amino acid analysis, peptide mapping, or sequence analysis, which may be applied to the protein after electroelution from SDS-acrylamide gels. It replaces the need for reduction and carboxymethylation prior to these analytical procedures. Protein-sequence analysis of prelabeled bovine serum albumin, including samples electroeluted from SDS-acrylamide gels, has justified the choice of this method to facilitate isolation of proteins for sequence analysis. Equivalent sequence data were obtained with reduced bovine serum albumin S-alkylated with iodoacetic acid or 5-I-AEDANS.

Amino Acid Sequence

Origin of tetracycline efflux proteins: conclusions from nucleotide sequence analysis.

The sequences of six tetracycline efflux proteins and three transport proteins which have some resemblance to them were compared. The tetracycline efflux proteins fall into three families: (i) those encoded by pBR322, RP1, and Tn10 (Escherichia coli); (ii) pT181 (Staphylococcus aureus) and pTHT15 (Bacillus subtilis); and (iii) tet347 (Streptomyces rimosus). There is global sequence homology within each of the first two families, but there is none between the families. The pT181/pTHT15 family shares close homology with the N-terminal half of the methylenomycin A efflux protein (Streptomyces coelicor), while tet347 resembles the C-terminal half. Portions of the N-terminal half of the Tn10-encoded protein show significant resemblance to portions in the N-terminal half of the pT181/pTHT15 family, but this sometimes occurs among transport proteins which do not have a common substrate. Tetracycline efflux proteins, therefore, appear to have arisen on at least two, or possibly three, separate occasions, probably from other transport proteins.

Amino Acid Sequence

Sequence analysis of adenovirus DNA: complete nucleotide sequence of the spliced 5' noncoding region of adenovirus 2 hexon messenger RNA.

The complete nucleotide sequence of the 5' noncoding region of the adenovirus 2 hexon messenger RNA has been established by sequence analysis of reverse transcripts. Such transcripts were generated by extension of specific single-stranded DNA primers with reverse transcriptase after hybridization to purified hexon mRNA. The total length of the 5' noncoding region was determined to be 240 nucleotides, of which the spliced tripartite leader sequence contributes 202 nucleotides including the terminal m7G. The sizes of the different segments of the tripartite leader were estimated by comparing the established mRNA sequence with the genomic sequences for the first and third leader segments, and were found to be 42 nucleotides for the first segment, 71 nucleotides for the second and 89 nucleotides for the third. The estimates are ambiguous, however, due to the presence of tandemly repeated sequences at both ends of the intervening sequence between the third leader segment and the body of the hexon mRNA. The sequence of the leader allows the formation of hydrogen-bonded interactions with the 3' end of 18S ribosomal RNA near the capped 5' end and also close to the initiator AUG.

Adenoviruses, Human