Search PubMed⌕ Search

Biomedical subjects

Robert B Russell

Publications and source records attributed to Robert B Russell.

At least 19 recordsLinked to original sources

Identification of Drosophila MicroRNA targets.

MicroRNAs (miRNAs) are short RNA molecules that regulate gene expression by binding to target messenger RNAs and by controlling protein production or causing RNA cleavage. To date, functions have been assigned to only a few of the hundreds of identified miRNAs, in part because of the difficulty in identifying their targets. The short length of miRNAs and the fact that their complementarity to target sequences is imperfect mean that target identification in animal genomes is not possible by standard sequence comparison methods. Here we screen conserved 3' UTR sequences from the Drosophila melanogaster genome for potential miRNA targets. The screening procedure combines a sequence search with an evaluation of the predicted miRNA-target heteroduplex structures and energies. We show that this approach successfully identifies the five previously validated let-7, lin-4, and bantam targets from a large database and predict new targets for Drosophila miRNAs. Our target predictions reveal striking clusters of functionally related targets among the top predictions for specific miRNAs. These include Notch target genes for miR-7, proapoptotic genes for the miR-2 family, and enzymes from a metabolic pathway for miR-277. We experimentally verified three predicted targets each for miR-7 and the miR-2 family, doubling the number of validated targets for animal miRNAs. Statistical analysis indicates that the best single predicted target sites are at the border of significance; thus, target predictions should be considered as tentative until experimentally validated. We identify features shared by all validated targets that can be used to evaluate target predictions for animal miRNAs. Our initial evaluation and experimental validation of target predictions suggest functions for two miRNAs. For others, the screen suggests plausible functions, such as a role for miR-277 as a metabolic switch controlling amino acid catabolism. Cross-genome comparison proved essential, as it allows reduction of the sequence search space. Improvements in genome annotation and increased availability of cDNA sequences from other genomes will allow more sensitive screens. An increase in the number of confirmed targets is expected to reveal general structural features that can be used to improve their detection. While the screen is likely to miss some targets, our study shows that valid targets can be identified from sequence alone.

3' Untranslated Regions↗

The relationship between sequence and interaction divergence in proteins.

There is currently a gap in knowledge between complexes of known three-dimensional structure and those known from other experimental methods such as affinity purifications or the two-hybrid system. This gap can sometimes be bridged by methods that extrapolate interaction information from one complex structure to homologues of the interacting proteins. To do this, it is important to know if and when proteins of the same type (e.g. family, superfamily or fold) interact in the same way. Here, we study interactions of known structure to address this question. We found all instances within the structural classification of proteins database of the same domain pairs interacting in different complexes, and then compared them with a simple measure (interaction RMSD). When plotted against sequence similarity we find that close homologues (30-40% or higher sequence identity) almost invariably interact the same way. Conversely, similarity only in fold (i.e. without additional evidence for a common ancestor) is only rarely associated with a similarity in interaction. The results suggest that there is a twilight zone of sequence similarity where it is not possible to say whether or not domains will interact similarly. We also discuss the rare instances of fold similarities interacting the same way, and those where obviously homologous proteins interact differently.

Amino Acid Sequence↗

Crystal structure of an archaeal class I aldolase and the evolution of (betaalpha)8 barrel proteins.

Fructose-1,6-bisphosphate aldolase (FBPA) catalyzes the reversible cleavage of fructose 1,6-bisphosphate to glyceraldehyde 3-phosphate and dihydroxyacetone phosphate in the glycolytic pathway. FBPAs from archaeal organisms have recently been identified and characterized as a divergent family of proteins. Here, we report the first crystal structure of an archaeal FBPA at 1.9-A resolution. The structure of this 280-kDa protein complex was determined using single wavelength anomalous dispersion followed by 10-fold non-crystallographic symmetry averaging and refined to an R-factor of 14.9% (Rfree 17.9%). The protein forms a dimer of pentamers, consisting of subunits adopting the ubiquitous (betaalpha)8 barrel fold. Additionally, a crystal structure of the archaeal FBPA covalently bound to dihydroxyacetone phosphate was solved at 2.1-A resolution. Comparison of the active site residues with those of classical FBPAs, which share no significant sequence identity but display the same overall fold, reveals a common ancestry between these two families of FBPAs. Structural comparisons, furthermore, establish an evolutionary link to the triosephosphate isomerases, a superfamily hitherto considered independent from the superfamily of aldolases.

Archaeal Proteins↗

Annotation in three dimensions. PINTS: Patterns in Non-homologous Tertiary Structures.

The detection of local structural patterns in proteins (e.g. active sites) can provide insights into protein function in the absence of sequence or fold similarity. Methods to detect such similarities are key during structural annotation, for example with results from Structural Genomics initiatives. PINTS (Patterns in Non-homologous Tertiary Structures, http://pints.embl.de) performs database searches for such patterns and most importantly provides a measure of statistical significance for any similarity uncovered. To aid functional annotation of proteins, we allow comparisons of pre-defined patterns against databases of complete structures and of entire structures to databases of particular residues likely to be functionally important.

Binding Sites↗

GlobPlot: Exploring protein sequences for globularity and disorder.

A major challenge in the proteomics and structural genomics era is to predict protein structure and function, including identification of those proteins that are partially or wholly unstructured. Non-globular sequence segments often contain short linear peptide motifs (e.g. SH3-binding sites) which are important for protein function. We present here a new tool for discovery of such unstructured, or disordered regions within proteins. GlobPlot (http://globplot.embl.de) is a web service that allows the user to plot the tendency within the query protein for order/globularity and disorder. We show examples with known proteins where it successfully identifies inter-domain segments containing linear motifs, and also apparently ordered regions that do not contain any recognised domain. GlobPlot may be useful in domain hunting efforts. The plots indicate that instances of known domains may often contain additional N- or C-terminal segments that appear ordered. Thus GlobPlot may be of use in the design of constructs corresponding to globular proteins, as needed for many biochemical studies, particularly structural biology. GlobPlot has a pipeline interface--GlobPipe--for the advanced user to do whole proteome analysis. GlobPlot can also be used as a generic infrastructure package for graphical displaying of any possible propensity.

Algorithms↗

bantam encodes a developmentally regulated microRNA that controls cell proliferation and regulates the proapoptotic gene hid in Drosophila.

Cell proliferation, cell death, and pattern formation are coordinated in animal development. Although many proteins that control cell proliferation and apoptosis have been identified, the means by which these effectors are linked to the patterning machinery remain poorly understood. Here, we report that the bantam gene of Drosophila encodes a 21 nucleotide microRNA that promotes tissue growth. bantam expression is temporally and spatially regulated in response to patterning cues. bantam microRNA simultaneously stimulates cell proliferation and prevents apoptosis. We identify the pro-apoptotic gene hid as a target for regulation by bantam miRNA, providing an explanation for bantam's anti-apoptotic activity.

3' Untranslated Regions↗

A model for statistical significance of local similarities in structure.

Structural biology can provide three-dimensional structures for proteins of unknown function. When sequence or structure comparisons fail to suggest a function, insights can come from discovery of functionally important local structural patterns. Existing methods to detect such patterns lack rigorous statistics needed for widespread application. Here, we derive a formula to calculate statistical significance of the root-mean-square deviation between atoms in such patterns. When combined with a database search method, our statistics permit true functional or structural patterns in different folds to be discerned from noise. The approach is highly complementary to fold comparison for providing functional clues for new structures, and is key for the detection of recurrences of any new pattern.

Algorithms↗

Predictions without templates: new folds, secondary structure, and contacts in CASP5.

We present the assessment of CASP5 predictions in the new fold category. For coordinate predictions, we considered five targets with new folds and eight lying on the fold recognition borderline. We performed detailed visual and numerical comparisons between predicted and experimental structures to assess prediction accuracy. The two procedures largely agreed, but the visual inspection identified instances where metrics, such as GDT_TS, ranked what we considered incorrect predictions highly. We found the quality of the best predictions to be very good: for nearly every target at least one group predicted a structure close to the correct one. However, selection of the best of five models is still problematic. The group of David Baker once again proved to be best overall, with many individual highlights. However, high quality and consistency were also seen from others, suggesting that the community is moving toward general procedures to predict accurate structures for proteins showing no resemblance to anything seen before. Predictions for secondary structure showed at best limited progress since CASP4. The number of targets is probably too small to spot differences in performance between methods, suggesting that such predictions might be better evaluated with schemes involving more proteins. For contact predictions, accuracies are still low, although there were several instances of accurate and useful contacts predicted de novo, and new approaches hint at future progress.

Algorithms↗

Protein disorder prediction: implications for structural proteomics.

A great challenge in the proteomics and structural genomics era is to predict protein structure and function, including identification of those proteins that are partially or wholly unstructured. Disordered regions in proteins often contain short linear peptide motifs (e.g., SH3 ligands and targeting signals) that are important for protein function. We present here DisEMBL, a computational tool for prediction of disordered/unstructured regions within a protein sequence. As no clear definition of disorder exists, we have developed parameters based on several alternative definitions and introduced a new one based on the concept of "hot loops," i.e., coils with high temperature factors. Avoiding potentially disordered segments in protein expression constructs can increase expression, foldability, and stability of the expressed protein. DisEMBL is thus useful for target selection and the design of constructs as needed for many biochemical studies, particularly structural biology and structural genomics projects. The tool is freely available via a web interface (http://dis.embl.de) and can be downloaded for use in large-scale studies.

Circular Dichroism↗

InterPreTS: protein interaction prediction through tertiary structure.

SUMMARY: InterPreTS (Interaction Prediction through Tertiary Structure) is a web-based version of our method for predicting protein-protein interactions (Aloy and Russell, 2002, PROC: Natl Acad. Sci. USA, 99, 5896-5901). Given a pair of query sequences, we first search for homologues in a database of interacting domains (DBID) of known three-dimensional complex structures. Pairs of sequences homologous to a known interacting pair are scored for how well they preserve the atomic contacts at the interaction interface. InterPreTS includes a useful interface for visualising molecular details of any predicted interaction. AVAILABILITY: http://www.russell.embl.de/interprets.

Databases, Protein↗

Evolutionary relationship between the bacterial HPr kinase and the ubiquitous PEP-carboxykinase: expanding the P-loop nucleotidyl transferase superfamily.

Similarities between protein three-dimensional structures can reveal evolutionary and functional relationships not apparent from sequence comparison alone. Here we report such a similarity between the metabolic enzymes histidine phosphocarrier protein kinase (HPrK) and phosphoenolpyruvate carboxykinase (PCK), suggesting that they are evolutionarily related. Current structure classifications place PCK and other P-loop containing nucleotidyl-transferases into different folds. Our comparison of both HPrK and PCK to other P-loop containing proteins reveals that all share a common structural motif consisting of an alphabeta segment containing the P-loop flanked by an additional beta-strand that is adjacent in space, but far apart along the sequence. Analysis also shows that HPrK/PCK differ from other P-loop containing structures no more than they differ from each other. We thus suggest that HPrK and PCK should be classified with other P-loop containing proteins, and that all probably share a common ancestor that probably contained a simple P-loop motif with different protein segments being added or lost over the course of evolution. We used the structure-based sequence alignment containing residues specific to HPrK/PCK to identify additional members of this P-loop containing family.

Amino Acid Motifs↗

Interrogating protein interaction networks through structural biology.

Protein-protein interactions are central to most biological processes. Although much recent effort has been put into methods to identify interacting partners, there has been a limited focus on how these interactions compare with those known from three-dimensional (3D) structures. Because comparison of protein interactions often involves considering homologous, but not identical, proteins, a key issue is whether proteins that are homologous to an interacting pair will interact in the same way, or interact at all. Accordingly, we describe a method to test putative interactions on complexes of known 3D structure. Given a 3D complex and alignments of homologues of the interacting proteins, we assess the fit of any possible interacting pair on the complex by using empirical potentials. For studies of interacting protein families that show different specificities, the method provides a ranking of interacting pairs useful for prioritizing experiments. We evaluate the method on interacting families of proteins with multiple complex structures. We then consider the fibroblast growth factor/receptor system and explore the intersection between complexes of known structure and interactions proposed between yeast proteins by methods such as two-hybrids. We provide confirmation for several interactions, in addition to suggesting molecular details of how they occur.

Amino Acid Sequence↗

Structure of the full-length HPr kinase/phosphatase from Staphylococcus xylosus at 1.95 A resolution: Mimicking the product/substrate of the phospho transfer reactions.

The histidine containing phospho carrier protein (HPr) kinase/phosphatase is involved in carbon catabolite repression, mainly in Gram-positive bacteria. It is a bifunctional enzyme that phosphorylates Ser-46-HPr in an ATP-dependent reaction and dephosphorylates P-Ser-46-HPr. X-ray analysis of the full-length crystalline enzyme from Staphylococcus xylosus at a resolution of 1.95 A shows the enzyme to consist of two clearly separated domains that are assembled in a hexameric structure resembling a three-bladed propeller. The N-terminal domain has a betaalphabeta fold similar to a segment from enzyme I of the sugar phosphotransferase system and to the uridyl-binding portion of MurF; it is structurally organized in three dimeric modules exposed to form the propeller blades. Two unexpected phosphate ions associated with highly conserved residues were found in the N-terminal dimeric interface. The C-terminal kinase domain is similar to that of the Lactobacillus casei enzyme and is assembled in six copies to form the compact central hub of the propeller. Beyond previously reported similarity with adenylate kinase, we suggest evolutionary relationship with phosphoenolpyruvate carboxykinase. In addition to a phosphate ion in the phosphate-binding loop of the kinase domain, we have identified a second phosphate-binding site that, by comparison with adenylate kinases, we believe accommodates a product/substrate phosphate, normally covalently linked to Ser-46 of HPr. Thus, we propose that our structure represents a product/substrate mimic of the kinase/phosphatase reaction.

Amino Acid Sequence↗

CASH--a beta-helix domain widespread among carbohydrate-binding proteins.

In this article, we describe a novel, widespread domain (CASH) that is shared by many carbohydrate-binding proteins and sugar hydrolases. This domain occurs in more than 1000 proteins distributed among all three kingdoms of life. The CASH domain is characterized by internal repetitions of glycines and hydrophobic residues that correspond to the repetitive units of a predicted or observed right-handed beta-helix structure of the pectate lyase superfamily.

Amino Acid Motifs↗