Search PubMed⌕ Search

Biomedical subjects

S Pietrokovski

Publications and source records attributed to S Pietrokovski.

29 records · Page 2Linked to original sources

Recent enhancements to the Blocks Database servers.

The Blocks Database contains multiple alignments of conserved regions in protein families which can be searched by e-mail (blocks@blocks.fhcrc.org) and World Wide Web (http://blocks.fhcrc.org/ ) servers to classify protein and nucleotide sequences. Recent enhancements to the servers include: (i) improved calculation of position-specific scoring matrices from blocks; (ii) availability of the Prints protein fingerprint database for searching in Blocks format; (iii) a representative sequence biased towards the Blocks of a protein family; (iv) a tree constructed from the Blocks of a protein family; (v) links to related World Wide Web pages for a family; and (vi) the new Local Alignment of Multiple Alignments (LAMA) method to search a block against a database of blocks.

Amino Acid Sequence↗

A helix-turn-helix DNA-binding motif predicted for transposases of DNA transposons.

A helix-turn-helix (HTH) DNA-binding motif is identified in transposase sequences in Tc1, mariner and pogo DNA transposum. The findings are supported by results of various sequence analysis methods. Tc1 transposases are also predicted to contain another DNA-binding region. These findings are in accord with experimental evidence obtained from Tc1A, Tc3A and pogo transposases. The pogo family transposases, but not the pogo-type transcription factors, contain the HTH motif, suggesting that HTH structures are essential for Tc1/mariner/pogo transposition. Analysis of multiple sequence alignments enabled the identification of the HTH motif in distantly related protein sequences.

Amino Acid Sequence↗

Searching databases of conserved sequence regions by aligning protein multiple-alignments.

A general searching method for comparing multiple sequence alignments was developed to detect sequence relationships between conserved protein regions. Multiple alignments are treated as sequences of amino acid distributions and aligned by comparing pairs of such distributions. Four different comparison measures were tested and the Pearson correlation coefficient chosen. The method is sensitive, detecting weak sequence relationships between protein families. Relationships are detected beyond the range of conventional sequence database searches, illustrating the potential usefulness of the method. The previously undetected relation between flavoprotein subunits of two oxidoreductase families points to the potential active site in one of the families. The similarity between the bacterial RecA, DnaA and Rad51 protein families reveals a region in DnaA and Rad51 proteins likely to bind and unstack single-stranded DNA. Helix--turn--helix DNA binding domains from diverse proteins are readily detected and shown to be similar to each other. Glycosylasparaginase and gamma-glutamyltransferase enzymes are found to be similar in their proteolytic cleavage sites. The method has been fully implemented on the World Wide Web at URL: http://blocks.fhcrc.org/blocks-bin/LAMAvsearch.

Amino Acid Sequence↗

The Blocks database--a system for protein classification.

The Blocks Database contains multiple alignments of conserved regions in protein families. The database can be searched by e-mail and World Wide Web(WWW) servers (http://blocks.fhcrc.org/help) to classify protein and nucleotide sequences.

Amino Acid Sequence↗

Automated construction and graphical presentation of protein blocks from unaligned sequences.

Protein blocks consist of multiply aligned sequence segments that correspond to the most highly conserved regions of protein families. Typically, a set of related proteins has more than one region in common and their relationship can be represented as a series of ungapped blocks separated by unaligned regions. Blockmaker is an automated system available by electronic mail (blockmaker@howard.fhcrc.org) and the World Wide Web (http://www.blocks.fhcrc.org4) that finds blocks in a group of related protein sequences submitted by the user. It adapts and extends existing algorithms to make them useful to biologists looking for conserved regions in a group of related proteins sequences. Two sets of blocks are returned, one in which candidate blocks are detected using the MOTIF algorithm and the other using a Gibbs sampler algorithm that has been adapted for full automation. This use of two block-finding methods based on completely different principles provides a 'reality check,' whereby a block detected by both methods is considered to be correct. Resulting blocks can be displayed using the information-based 'sequence logo' method, adapted to incorporate sequence weights, which provides an intuitive visual description of both the residue and the conservation information at each position. Blocks generated by this system are useful in diverse applications, such as searching databases and designing degenerate PCR primers. As an example, blocks made from amino acid sequences related to Caenorhabditis elegans Tc1 transposase were used to search GenBank, revealing that several fish and amphibian genomic sequences harbor previously unreported Tc1 homologs.

Algorithms↗

Comparing nucleotide and protein sequences by linguistic methods.

Nucleotide and amino acid sequences can be analyzed and compared by their oligomer compositions. Such methods are fundamentally different from comparison methods based on sequence alignment. They are analogous to the linguistic analysis of human texts. The methods have a wide range of sensitivity and can identify homologous as well as functionally and taxonomically related sequences. Significant sequence dissimilarity can also be identified enabling detection of foreign DNA sequences in genomes, genetic libraries and databases. The simplicity and speed of linguistic methods make them very suitable for database searching and maintenance and as a preliminary step to more specific and time-consuming analysis methods.

Amino Acid Sequence↗

Conserved sequence features of inteins (protein introns) and their use in identifying new inteins and related proteins.

Inteins (protein introns) are internal portions of protein sequences that are posttranslationally excised while the flanking regions are spliced together, making an additional protein product. Inteins have been found in a number of homologous genes in yeast, mycobacteria, and extreme thermophile archaebacteria. The inteins are probably multifunctional, autocatalyzing their own splicing, and some were also shown to be DNA endonucleases. The splice junction regions and two regions similar to homing endonucleases were thought to be the only common sequence features of inteins. This work analyzed all published intein sequences with recently developed methods for detecting weak, conserved sequence features. The methods complemented each other in the identification and assessment of several patterns characterizing the intein sequences. New intein conserved features are discovered and the known ones are quantitatively described and localized. The general sequence description of all the known inteins is derived from the motifs and their relative positions. The intein sequence description is used to search the sequence databases for intein-like proteins. A sequence region in a mycobacterial open reading frame possessing all of the intein motifs and absent from sequences homologous to both of its flanking sequences is identified as an intein. A newly discovered putative intein in red algae chloroplasts is found not to contain the endonuclease motifs present in all other inteins. The yeast HO endonuclease is found to have an overall intein-like structure and a few viral polyprotein cleavage sites are found to be significantly similar to the inteins amino-end splice junction motif. The intein features described may serve for detection of intein sequences.

Amino Acid Sequence↗

Imported sequences in the mitochondrial yeast genome identified by nucleotide linguistics.

In addition to universally appearing mitochondrial (mt) genes, origins of replication and transcription start regions typical of all mt genome variants of the yeast Saccharomyces cerevisiae, the mt genomes of some of the strains contain variable sequences. These sequences are apparently largely dispensable. They are mainly composed of group-I and -II introns and intergenic open reading frames (ORFs). Many of the introns contain ORFs, some of which were shown by genetic and biochemical means to be involved in splicing and transposition of the mt introns. Some of the optional sequences are hypothesized to be mobile genetic elements. Nucleotide (nt) sequences of the mt genome of S. cerevisiae were examined by analyzing occurrences of oligodeoxyribonucleotide (oligo) 'words'. This linguistic technique had been found to be sensitive to both function and origin of the sequence [Pietrokovski et al., J. Biomol. Struct. Dyn. 7 (1990) 1251-1268]. A clear difference is found between the oligo vocabularies of the optional and basic yeast mt sequences. The difference is mainly located in protein coding segments of the optional sequences which contain conserved amino acid motifs, characteristic of intronic and intergenic ORFs. The use of nt linguistics to detect the sequence dissimilarity and its causes in yeast mitochondria provides fast and straightforward results, identifying the intronic and intergenic ORFs as DNA sequences of foreign, non-mt origin.

Amino Acid Sequence↗

Linguistic measure of taxonomic and functional relatedness of nucleotide sequences.

The frequencies of "words", oligonucleotides within nucleotide sequences, reflect the genetic information contained in the sequence "texts". Nucleotide sequences are characteristically represented by their contrast word vocabularies. Comparison of the sequences by correlating their contrast vocabularies is shown to reflect well the relatedness (unrelatedness) between the sequences. A single value, the linguistic similarity between the sequences, is suggested as a measure of sequence relatedness. Sequences as short as 1000 bases can be characterized and quantitatively related to other sequences by this technique. The linguistic sequence similarity value is used for analysis of taxonomically and functionally diverse nucleotide sequences. The similarity value is shown to be very sensitive to the relatedness of the source species, thus providing a convenient tool for taxonomic classification of species by their sequence vocabularies. Functionally diverse sequences appear distinct by their linguistic similarity values. This can be a basis for a quick screening technique for functional characterization of the sequences and for mapping functionally distinct regions in long sequences.

Animals↗