Characterization of targeting domains by sequence analysis: glycogen-binding domains in protein phosphatases.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to T Dandekar.
Explore the source record for details and available documents.
A systematic comparison of nine bacterial and archaeal genomes reveals a low level of gene-order (and operon architecture) conservation. Nevertheless, a number of gene pairs are conserved. The proteins encoded by conserved gene pairs appear to interact physically. This observation can therefore be used to predict functions of, and interactions between, prokaryotic gene products.
Major pathogenic functions of Entamoeba histolytica involved in destruction of host tissues are the degradation of extracellular matrix proteins mediated by secreted cysteine proteinases and contact-dependent killing of host cells via membrane-active factors. A soluble protein with an affinity for membranes was purified from amoebic extracts to apparent homogeneity. N-terminal sequencing and subsequent molecular cloning of the factor revealed that it is a member of the cysteine proteinase family of E. histolytica, which we termed CP5. Further experiments with the purified protein showed that it has potent proteolytic activity that is abrogated in the presence of inhibitors specific for cysteine proteinases. The enzyme firmly associates with membranes retaining its proteolytic activity and it produces cytopathic effects on cultured monolayers. A model of the three-dimensional structure of CP5 revealed the presence of a hydrophobic patch that may account for the potential of the protein to associate with membranes. Immunocytochemical localization of the enzyme to the surface of the amoeba in combination with the recent finding that the gene encoding CP5 is missing in the closely related but non-pathogenic Entamoeba dispar suggests a potential role of the protein in host tissue destruction of E. histolytica.
MOTIVATION: The untranslated regions (UTRs) of mRNA upstream (5'UTR) and downstream (3'UTR) of the open reading frame, as well as the mRNA precursor, carry important regulatory sequences. To reveal unidentified regulatory signals, we combine information from experiments with computational approaches. Depending on available knowledge, three different strategies are employed. RESULTS: Searching with a consensus template, new RNAs with regulatory RNA elements can be identified in genomic screens. By this approach, we identify new candidate regulatory motifs resembling iron-responsive elements in the 5'UTRs of HemA, FepB and FrdB mRNA from Escherichia coli. If an RNA element is not yet defined, it may be analyzed by combining results from SELEX (selective enrichment of ligands by exponential amplification) and a search of databases from RNA or genomic sequences. A cleavage stimulating factor (CstF) binding element 3 of the polyadenylation site in the mRNA precursor serves as a test example. Alternatively, the regulatory RNA element may be found by studying different RNA foldings and their correlation with simple experimental tests. We delineate a novel instability element in the 3'UTR of the estrogen receptor mRNA in this way. AVAILABILITY: Strategy, methods and programs are available on request from T.Dandekar. CONTACT: dandekar@embl-heidelberg.de
The genes encoding the small nucleolar RNA (snoRNA) species snR190 and U14 are located close together in the genome of Saccharomyces cerevisiae. Here we report that these two snoRNAs are synthesized by processing of a larger common transcript. In strains mutant for two 5'-->3' exonucleases, Xrn1p and Rat1p, families of 5'-extended forms of snR190 and U14 accumulate; these have 5' extensions of up to 42 and 55 nucleotides, respectively. We conclude that the 5' ends of both snR190 and U14 are generated by exonuclease digestion from upstream processing sites. In contrast to snR190 and U14, the snoRNAs U18 and U24 are excised from the introns of pre-mRNAs which encode proteins in their exonic sequences. Analysis of RNA extracted from a dbr1-delta strain, which lacks intron lariat-debranching activity, shows that U24 can be synthesized only from the debranched lariat. In contrast, a substantial level of U18 can be synthesized in the absence of debranching activity. The 5' ends of these snoRNAs are also generated by Xrn1p and Rat1p. The same exonucleases are responsible for the degradation of several excised fragments of the pre-rRNA spacer regions, in addition to generating the 5' end of the 5.8S rRNA. Processing of the pre-rRNA and both intronic and polycistronic snoRNAs therefore involves common components.
Explore the source record for details and available documents.
Critical events in 3'-end processing of pre-mRNA are the recognition of the AAUAAA polyadenylation signal by cleavage and polyadenylation specificity factor (CPSF) and the binding of cleavage stimulation factor (CstF) via its 64-kDa subunit to the downstream element. The stability of this CPSF.CstF.RNA complex is thought to determine the efficiency of 3'-end processing. Since downstream elements reveal high sequence variability, in vitro selection experiments with highly purified CstF were performed to investigate the sequence requirements for CstF-RNA interaction. CstF was purified from calf thymus and from HeLa cells. Surprisingly, calf thymus CstF contained an additional, novel form of the 64-kDa subunit with a molecular mass of 70 kDa. RNA ligands selected by HeLa and calf thymus CstF contained three highly conserved sequence elements as follows: element 1 (AUGCGUUCCUCGUCC) and two closely related elements, element 2a (YGUGUYN0-4UUYAYUGYGU) and element 2b (UUGYUN0-4AUUUACU(U/G)N0-2YCU). All selected sequences tested functioned as downstream elements in 3'-end processing in vitro. A computer survey of the EMBL data library revealed significant homologies to all selected elements in naturally occurring 3'-untranslated regions. The majority of element 2a homologies was found downstream of coding sequences. Therefore, we postulate that this element represents a novel consensus sequence for downstream elements in 3'-end processing of pre-mRNA.
A topological and functional overview of a DNA recognition protein with unknown structure can be achieved by combining three different, but complementary approaches: modeling by the genetic algorithm, functional analysis of mutated variants, and testing the target DNA using non-canonical oligonucleotides. As an example we choose the Flp protein, a site-specific recombinase from Saccharomyces cerevisiae. We derive the topological outline including the DNA binding cleft, examine DNA binding regions by deletional and mutational analysis, and analyze the DNA binding site using 7-deazaadenine, 7-deazaguanine, inosine and 4-O-methylthymine as probes. The combined data offer a comprehensive sketch of a plausible protein architecture for Flp. The structure is detailed enough to verify the prediction accuracy for different peptide regions from pre-existing data and by new experimental design.
BACKGROUND: Amoebapore of the protozoan Entamoeba histolytica and NK-lysin of porcine cytotoxic lymphocytes are effector peptides from organisms separated extremely early in their evolutionary paths. The peptides intrigued us, however, with indications of some functional similarity. We thus wanted to derive and compare predictions for their as yet unknown three-dimensional structures as a guide for and to be tested by further experiments. RESULTS: Molecular models were generated by use of a genetic algorithm that selects according to basic protein structure principles exploiting available information such as the primary structures, secondary structure predictions and positions of disulfide bonds. Topological differences aside, the structural motif of an antiparallel four-alpha-helix bundle with adjacent connections and intramolecular crosslinks is predicted for both types of peptides. It combines the feature of amphipathic alpha-helices with a disulfide-bonded compact structure known from the beta-sheeted defensins and small toxins. CONCLUSIONS: The models presented here strengthen the notion that amoebapore and NK-lysin are particular among cytolytic and antibacterial polypeptides and share a similar function and structural motif. They also allow experimental testing and a better comparison of the two proteins in view of the predicted similarities and differences of their respective folds.
Specific residue interactions as revealed from a few and readily available experiments can be quite important in shaping a protein's tertiary topology by complementing basic and general folding principles. This experimental information is employed in structure prediction (mainchain topology) based on sequence knowledge and the genetic algorithm with its ability to optimize simultaneously many parameters. Examples investigated include the distribution of cysteinyl S-S bonds, protein side-chain ligands to iron-sulfur cages, cofactor-ligands, crosslinks amongst side-chains, and conserved hydrophobic and catalytic residues. Such interactions yield an improvement in the predicted topology (0.4-6.6 A root mean square deviation in the positions of the backbone C alpha-atoms relative to those observed) compared with those resulting from simulations relying only on basic protein folding principles. For several examples the resultant topology depended critically on knowledge of the few and specific interactions such that the relationship between predicted and observed C alpha-positions was near random without their use. The combined methodology (experimental data and the genetic algorithm) should prove helpful in settings where experiment and theory can cooperate in successive steps to elucidate an unknown structure.
The posttranscriptional control of iron uptake, storage, and utilization by iron-responsive elements (IREs) and iron regulatory proteins (IRPs) provides a molecular framework for the regulation of iron homeostasis in many animals. We have identified and characterized IREs in the mRNAs for two different mitochondrial citric acid cycle enzymes. Drosophila melanogaster IRP binds to an IRE in the 5' untranslated region of the mRNA encoding the iron-sulfur protein (Ip) subunit of succinate dehydrogenase (SDH). This interaction is developmentally regulated during Drosophila embryogenesis. In a cell-free translation system, recombinant IRP-1 imposes highly specific translational repression on a reporter mRNA bearing the SDH IRE, and the translation of SDH-Ip mRNA is iron regulated in D. melanogaster Schneider cells. In mammals, an IRE was identified in the 5' untranslated regions of mitochondrial aconitase mRNAs from two species. Recombinant IRP-1 represses aconitase synthesis with similar efficiency as ferritin IRE-controlled translation. The interaction between mammalian IRPs and the aconitase IRE is regulated by iron, nitric oxide, and oxidative stress (H2O2), indicating that these three signals can control the expression of mitochondrial aconitase mRNA. Our results identify a regulatory link between energy and iron metabolism in vertebrates and invertebrates, and suggest biological functions for the IRE/IRP regulatory system in addition to the maintenance of iron homeostasis.
Grid-free protein folding simulations based on sequence and secondary structure knowledge (using mostly experimentally determined secondary structure information but also analysing results from secondary structure predictions) were investigated using the genetic algorithm, a backbone representation, and standard dihedral angular conformations. Optimal structures are selected according to basic protein building principles. Having previously applied this approach to proteins with helical topology, we have now developed additional criteria and weights for beta-strand-containing proteins, validated them on four small beta-strand-rich proteins with different topologies, and tested the general performance of the method on many further examples from known protein structures with mixed secondary structural type and less than 100 amino acid residues. Topology predictions close to the observed experimental structures were obtained in four test cases together with fitness values that correlated with the similarity of the predicted topology to the observed structures. Root-mean-square deviation values of C alpha atoms in the superposed predicted and observed structures, the latter of which had different topologies, were between 4.5 and 5.5 A(2.9 to 5.1 A without loops). Including 15 further protein examples with unique folds, root-mean-square deviation values ranged between 1.8 and 6.9 A with loop regions and averaged 5.3 A and 4.3 A, including and excluding loop regions, respectively.
In this study, protein structures were predicted using a genetic algorithm. Optimal structures were selected according to basic protein-building principles. Having previously applied this approach to proteins with helical topology, this paper reports preliminary results on 3 proteins with mixed topology (helical and non-helical, e.g. sheet). The method seems to be generally applicable to protein-fold prediction since topology predictions close to the observed experimental structures are obtained.
A motif search for DNA and mRNA sequences matching the rev response core element of HIV identified only DNA ligase I to display the element both in DNA and mRNA and to have phylogenetic conservation of it. Genetic impairment of DNA ligase I is already known to lead to severe immunodeficiency in man. A similar impairment may result from HIV rev protein binding to the DNA ligase I mRNA.
A growing list of examples underscores the roles that regulatory RNA motifs play in controlling the genetic repertoire of cells and developing organisms. Once either an RNA-processing signal, a ribozyme, an element that controls translational or mRNA stability or an RNA localization signal has been identified, it is important to search for other RNA sequences that bear similar regulatory signals. While DNA regulatory elements can often be described by a consensus sequence, RNA signals are frequently composed of a combination of sequence and structure motifs. Here, we discuss the approaches that can be used to identify RNA motifs by searching databases.
Grid-free protein folding simulations were effected using the genetic algorithm, a backbone representation and standard dihedral angular conformations. The topological folding of idealized four-helix bundles was investigated in detail to differentiate among the important protein folding forces used as fitness criteria. Hydrophobic interactions were the most significant while local forces and hydrogen bonds were far less effective in promoting folding. Stable secondary structural regions were also important as nucleating centers. Using the fitness parameters optimized in idealized simulations together with standard secondary structure predictions derived from the amino acid sequence alone, the proper main-chain folding of the four-helix bundle proteins cytochrome b562, cytochrome c' and hemerythrin was achieved. In addition the backbone topology as predicted by the genetic algorithm for crambin, a mixed helix/strand protein with known structure, is presented and discussed.
snR31 is a RNA species of 225 nt. which has the trimethyl guanosine cap structure typical of small nuclear RNAs (snRNAs) and yeast small nucleolar RNAs (snoRNAs), and is associated with the nucleolar proteins fibrillarin (NOP1) and GAR1. On sub-nuclear fractionation, snR31 behaves like other snoRNAs, and is enriched in a nucleolar fraction. The SNR31 genomic locus is close to the SNR5 locus, which encodes another snoRNA. The two genes are divergently transcribed with 217 bp separating the transcription start sites. Disruption of the SNR31 gene does not detectably impair growth in a haploid strain. Analyses of pre-rRNA processing in wild-type and snr31- strains shows some accumulation of the 35S primary transcript in the mutant, indicating a mild impairment of the initial steps in pre-rRNA processing.
We have developed a system for testing mutations by plasmid exchange in the fission yeast Schizosaccharomyces pombe. This system has been used to test the requirement for different regions of the small nuclear RNA U4 in S. pombe. Surprisingly, five of seven deletion and substitution mutations tested in different regions of U4 prevent the accumulation of the mutant RNA. Substitution of the U4 sequence in stem 1 of the U4/U6 interaction domain allows accumulation of the mutant U4, but does not support viability. Two sequences with homology to the Sm binding site are found in the 3' region of S. pombe U4; substitution of the 3' sequence of the two does not interfere with accumulation or function of U4, indicating that the 5' sequence is the functional Sm-binding site.