Search PubMed⌕ Search

Biomedical subjects

Alfonso Valencia

Publications and source records attributed to Alfonso Valencia.

At least 55 records · Page 3Linked to original sources

SQUARE--determining reliable regions in sequence alignments.

The Server for Quick Alignment Reliability Evaluation (SQUARE) is a Web-based version of the method we developed to predict regions of reliably aligned residues in sequence alignments. Given an alignment between a query sequence and a sequence of known structure, SQUARE is able to predict which residues are reliably aligned. The server accesses a database of profiles of sequences of known three-dimensional structures in order to calculate the scores for each residue in the alignment. SQUARE produces a graphical output of the residue profile-derived alignment scores along with an indication of the reliability of the alignment. In addition, the scores can be compared against template secondary structure, conserved residues and important sites.

Algorithms↗

The small GTP-binding protein, Rhes, regulates signal transduction from G protein-coupled receptors.

The Ras homolog enriched in striatum, Rhes, is the product of a thyroid hormone-regulated gene during brain development. Rhes and the dexamethasone-induced Dexras1 define a novel distinct subfamily of proteins within the Ras family, characterized by an extended variable domain in the carboxyl terminal region. We have carried this study because there is a complete lack of knowledge on Rhes signaling. We show that in PC12 cells, Rhes is targeted to the plasma membrane by farnesylation. We demonstrate that about 30% of the native Rhes protein is bound to GTP and this proportion is unaltered by typical Ras family nucleotide exchange factors. However, Rhes is not transforming in murine fibroblasts. We have also examined the role of Rhes in cell signaling. Rhes does not stimulate the ERK pathway. By contrast, it binds to and activates PI3K. On the other hand, we demonstrate that Rhes impairs the activation of the cAMP/PKA pathway by thyroid-stimulating hormone, and by an activated beta2 adrenergic receptor by a mechanism that suggests uncoupling of the receptor to its cognate heterotrimeric complex. Overall, our results provide the initial insights into the role in signal transduction of this novel Ras family member.

Animals↗

Identification of amino acid residues crucial for chemokine receptor dimerization.

Chemokines coordinate leukocyte trafficking by promoting oligomerization and signaling by G protein-coupled receptors; however, it is not known which amino acid residues of the receptors participate in this process. Bioinformatic analysis predicted that Ile52 in transmembrane region-1 (TM1) and Val150 in TM4 of the chemokine receptor CCR5 are key residues in the interaction surface between CCR5 molecules. Mutation of these residues generated nonfunctional receptors that could not dimerize or trigger signaling. In vitro and in vivo studies in human cell lines and primary T cells showed that synthetic peptides containing these residues blocked responses induced by the CCR5 ligand CCL5. Fluorescence resonance energy transfer showed the presence of preformed, ligand-stabilized chemokine receptor oligomers. This is the first description of the residues involved in chemokine receptor dimerization, and indicates a potential target for the modification of chemokine responses.

Amino Acids↗

IntAct: an open source molecular interaction database.

IntAct provides an open source database and toolkit for the storage, presentation and analysis of protein interactions. The web interface provides both textual and graphical representations of protein interactions, and allows exploring interaction networks in the context of the GO annotations of the interacting proteins. A web service allows direct computational access to retrieve interaction networks in XML format. IntAct currently contains approximately 2200 binary and complex interactions imported from the literature and curated in collaboration with the Swiss-Prot team, making intensive use of controlled vocabularies to ensure data consistency. All IntAct software, data and controlled vocabularies are available at http://www.ebi.ac.uk/intact.

Animals↗

Solution structure of the hypothetical protein Mth677 from Methanobacterium thermoautotrophicum: a novel alpha+beta fold.

The structure of Mth677, a hypothetical protein from Methanobacterium thermoautotrophicum (Mth), has been determined by using heteronuclear nuclear magnetic resonance (NMR) methods on a double-labeled (15)N-(13)C sample. Mth677 adopts a novel alpha+beta fold, consisting of two alpha-helices (one N terminal and one C terminal) packed on the same side of a central beta-hairpin. This structure is likely shared by its three orthologs, detected in three other Archaebacteria. There are no clear features in the sequences of these proteins or in the genome organization of Mth to make a reliable functional assignment to this protein. However, the structural similarity to Escherichia coli MinE, the protein which controls that division occurs at the midcell site, lends support to the proposal that Mth677 might be, in Mth, the counterpart of the topological specificity domain of MinE in E. coli.

Amino Acid Sequence↗

Early bioinformatics: the birth of a discipline--a personal view.

MOTIVATION: The field of bioinformatics has experienced an explosive growth in the last decade, yet this 'new' field has a long history. Some historical perspectives have been previously provided by the founders of this field. Here, we take the opportunity to review the early stages and follow developments of this discipline from a personal perspective. RESULTS: We review the early days of algorithmic questions and answers in biology, the theoretical foundations of bioinformatics, the development of algorithms and database resources and finally provide a realistic picture of what the field looked like from a resources and finally provide a realistic picture of what the field looked like from a practitioner's viewpoint 10 years ago, with a perspective for future developments.

Algorithms↗

Automatic annotation of protein function based on family identification.

Although genomes are being sequenced at an impressive rate, the information generated tells us little about protein function, which is slow to characterize by traditional methods. Automatic protein function annotation based on computational methods has alleviated this imbalance. The most powerful current approach for inferring the function of new proteins is by studying the annotations of their homologues, since their common origin is assumed to be reflected in their structure and function. Unfortunately, as proteins evolve they acquire new functions, so annotation based on homology must be carried out in the context of orthologues or subfamilies. Evolution adds new complications through domain shuffling: homology (or orthology) frequently corresponds to domains rather than complete proteins. Moreover, the function of a protein may be seen as the result of combining the functions of its domains. Additionally, automatic annotation has to deal with problems related to the annotations in the databases: errors (which are likely to be propagated), inconsistencies, or different degrees of function specification. We describe a method that addresses these difficulties for the annotation of protein function. Sequence relationships are detected and measured to obtain a map of the sequence space, which is searched for differentiated groups of proteins (similar to islands on the map), which are expected to have a common function and correspond to groups of orthologues or subfamilies. This mapmaking is done by applying a clustering algorithm based on Normalized cuts in graphs. The domain problem is addressed in a simple way: pairwise local alignments are analyzed to determine the extent to which they cover the entire sequence lengths of the two proteins. This analysis determines both what homologues are preferred for functional inheritance and the level of confidence of the annotation. To alleviate the problems associated with database annotations, the information on all the homologues that are grouped together with the query protein are taken into account to select the most representative functional descriptors. This method has been applied for the annotation of the genome of Buchnera aphidicola (specific host Baizongia pistaciae). Human inspection of the annotations allowed an estimation of accuracy of 94%; the different kinds of error that may appear when using this approach are described. Results can be accessed at http://www.pdg.cnb.uam.es/funcut.html. The programs are available upon request, although installation in other systems may be complicated.

Algorithms↗

The organization of the microbial biodegradation network from a systems-biology perspective.

Microbial biodegradation of environmental pollutants is a field of growing importance because of its potential use in bioremediation and biocatalysis. We have studied the characteristics of the global biodegradation network that is brought about by all the known chemical reactions that are implicated in this process, regardless of their microbial hosts. This combination produces an efficient and integrated suprametabolism, with properties similar to those that define metabolic networks in single organisms. The characteristics of this network support an evolutionary scenario in which the reactions evolved outwards from the central metabolism. The properties of the global biodegradation network have implications for predicting the fate of current and future environmental pollutants.

Bacteria↗

Predicting reliable regions in protein alignments from sequence profiles.

For applications such as comparative modelling one major issue is the reliability of sequence alignments. Reliable regions in alignments can be predicted using sub-optimal alignments of the same pair of sequences. Here we show that reliable regions in alignments can also be predicted from multiple sequence profile information alone. Alignments were created for a set of remotely related pairs of proteins using five different test methods. Structural alignments were used to assess the quality of the alignments and the aligned positions were scored using information from the observed frequencies of amino acid residues in sequence profiles pre-generated for each template structure. High-scoring regions of these profile-derived alignment scores were a good predictor of reliably aligned regions. These profile-derived alignment scores are easy to obtain and are applicable to any alignment method. They can be used to detect those regions of alignments that are reliably aligned and to help predict the quality of an alignment. For those residues within secondary structure elements, the regions predicted as reliably aligned agreed with the structural alignments for between 92% and 97.4% of the residues. In loop regions just under 92% of the residues predicted to be reliable agreed with the structural alignments. The percentage of residues predicted as reliable ranged from 32.1% for helix residues to 52.8% for strand residues. This information could also be used to help predict conserved binding sites from sequence alignments. Residues in the template that were identified as binding sites, that aligned to an identical amino acid residue and where the sequence alignment agreed with the structural alignment were in highly conserved, high scoring regions over 80% of the time. This suggests that many binding sites that are present in both target and template sequences are in sequence-conserved regions and that there is the possibility of translating reliability to binding site prediction.

Algorithms↗

EVA: Evaluation of protein structure prediction servers.

EVA (http://cubic.bioc.columbia.edu/eva/) is a web server for evaluation of the accuracy of automated protein structure prediction methods. The evaluation is updated automatically each week, to cope with the large number of existing prediction servers and the constant changes in the prediction methods. EVA currently assesses servers for secondary structure prediction, contact prediction, comparative protein structure modelling and threading/fold recognition. Every day, sequences of newly available protein structures in the Protein Data Bank (PDB) are sent to the servers and their predictions are collected. The predictions are then compared to the experimental structures once a week; the results are published on the EVA web pages. Over time, EVA has accumulated prediction results for a large number of proteins, ranging from hundreds to thousands, depending on the prediction method. This large sample assures that methods are compared reliably. As a result, EVA provides useful information to developers as well as users of prediction methods.

Automation↗

Structural (betaalpha)8 TIM barrel model of 3-hydroxy-3-methylglutaryl-coenzyme A lyase.

This study describes three novel homozygous missense mutations (S75R, S201Y, and D204N) in the 3-hydroxy-3-methylglutaryl-CoA (HMG-CoA) lyase gene, which caused 3-hydroxy-3-methylglutaric aciduria in patients from Germany, England, and Argentina. Expression studies in Escherichia coli show that S75R and S201Y substitutions completely abolished the HMG-CoA lyase activity, whereas D204N reduced catalytic efficiency to 6.6% of the wild type. We also propose a three-dimensional model for human HMG-CoA lyase containing a (betaalpha)8 (TIM) barrel structure. The model is supported by the similarity with analogous TIM barrel structures of functionally related proteins, by the localization of catalytic amino acids at the active site, and by the coincidence between the shape of the substrate (HMG-CoA) and the predicted inner cavity. The three novel mutations explain the lack of HMG-CoA lyase activity on the basis of the proposed structure: in S75R and S201Y because the new amino acid residues occlude the substrate cavity, and in D204N because the mutation alters the electrochemical environment of the active site. We also report the localization of all missense mutations reported to date and show that these mutations are located in the beta-sheets around the substrate cavity.

Amino Acid Sequence↗

Evaluation of annotation strategies using an entire genome sequence.

MOTIVATION: Genome-wide functional annotation either by manual or automatic means has raised considerable concerns regarding the accuracy of assignments and the reproducibility of methodologies. In addition, a performance evaluation of automated systems that attempt to tackle sequence analyses rapidly and reproducibly is generally missing. In order to quantify the accuracy and reproducibility of function assignments on a genome-wide scale, we have re-annotated the entire genome sequence of Chlamydia trachomatis (serovar D), in a collaborative manner. RESULTS: We have encoded all annotations in a structured format to allow further comparison and data exchange and have used a scale that records the different levels of potential annotation errors according to their propensity to propagate in the database due to transitive function assignments. We conclude that genome annotation may entail a considerable amount of errors, ranging from simple typographical errors to complex sequence analysis problems. The most surprising result of this comparative study is that automatic systems might perform as well as the teams of experts annotating genome sequences.

Amino Acid Sequence↗

Automatic methods for predicting functionally important residues.

Sequence analysis is often the first guide for the prediction of residues in a protein family that may have functional significance. A few methods have been proposed which use the division of protein families into subfamilies in the search for those positions that could have some functional significance for the whole family, but at the same time which exhibit the specificity of each subfamily ("Tree-determinant residues"). However, there are still many unsolved questions like the best division of a protein family into subfamilies, or the accurate detection of sequence variation patterns characteristic of different subfamilies. Here we present a systematic study in a significant number of protein families, testing the statistical meaning of the Tree-determinant residues predicted by three different methods that represent the range of available approaches. The first method takes as a starting point a phylogenetic representation of a protein family and, following the principle of Relative Entropy from Information Theory, automatically searches for the optimal division of the family into subfamilies. The second method looks for positions whose mutational behavior is reminiscent of the mutational behavior of the full-length proteins, by directly comparing the corresponding distance matrices. The third method is an automation of the analysis of distribution of sequences and amino acid positions in the corresponding multidimensional spaces using a vector-based principal component analysis. These three methods have been tested on two non-redundant lists of protein families: one composed by proteins that bind a variety of ligand groups, and the other composed by proteins with annotated functionally relevant sites. In most cases, the residues predicted by the three methods show a clear tendency to be close to bound ligands of biological relevance and to those amino acids described as participants in key aspects of protein function. These three automatic methods provide a wide range of possibilities for biologists to analyze their families of interest, in a similar way to the one presented here for the family of proteins related with ras-p21.

Algorithms↗

Phage-display and correlated mutations identify an essential region of subdomain 1C involved in homodimerization of Escherichia coli FtsA.

FtsA plays an essential role in Escherichia coli cell division and is nearly ubiquitous in eubacteria. Several evidences postulated the ability of FtsA to interact with other septation proteins and with itself. To investigate these binding properties, we screened a phage-display library with FtsA. The isolated peptides defined a degenerate consensus sequence, which in turn displayed a striking similarity with residues 126-133 of FtsA itself. This result suggested that residues 126-133 were involved in homodimerization of FtsA. The hypothesis was supported by the analysis of correlated mutations, which identified a mutual relationship between a group of amino acids encompassing the ATP-binding site and a set of residues immediately downstream to amino acids 126-133. This information was used to assemble a model of a FtsA homodimer, whose accuracy was confirmed by probing multiple alternative docking solutions. Moreover, a prediction of residues responsible for protein-protein interaction validated the proposed model and confirmed once more the importance of residues 126-133 for homodimerization. To functionally characterize this region, we introduced a deletion in ftsA, where residues 126-133 were skipped. This mutant failed to complement conditional lethal alleles of ftsA, demonstrating that amino acids 126-133 play an essential role in E. coli.

Amino Acid Sequence↗

Reductive genome evolution in Buchnera aphidicola.

We have sequenced the genome of the intracellular symbiont Buchnera aphidicola from the aphid Baizongia pistacea. This strain diverged 80-150 million years ago from the common ancestor of two previously sequenced Buchnera strains. Here, a field-collected, nonclonal sample of insects was used as source material for laboratory procedures. As a consequence, the genome assembly unveiled intrapopulational variation, consisting of approximately 1,200 polymorphic sites. Comparison of the 618-kb (kbp) genome with the two other Buchnera genomes revealed a nearly perfect gene-order conservation, indicating that the onset of genomic stasis coincided closely with establishment of the symbiosis with aphids, approximately 200 million years ago. Extensive genome reduction also predates the synchronous diversification of Buchnera and its host; but, at a slower rate, gene loss continues among the extant lineages. A computational study of protein folding predicts that proteins in Buchnera, as well as proteins of other intracellular bacteria, are generally characterized by smaller folding efficiency compared with proteins of free living bacteria. These and other degenerative genomic features are discussed in light of compensatory processes and theoretical predictions on the long-term evolutionary fate of symbionts like Buchnera.

Base Sequence↗

CAFASP3 in the spotlight of EVA.

We have analysed fold recognition, secondary structure and contact prediction servers from CAFASP3. This assessment was carried out in the framework of the fully automated, web-based evaluation server EVA. Detailed results are available at http://cubic.bioc.columbia.edu/eva/cafasp3/. We observed that the sequence-unique targets from CAFASP3/CASP5 were not fully representative for evaluating performance. For all three categories, we showed how careless ranking might be misleading. We compared methods from all categories to experts in secondary structure and contact prediction and homology modellers to fold recognisers. While the secondary structure experts clearly outperformed all others, the contact experts appeared to outperform only novel fold methods. Automatic evaluation servers are good at getting statistics right and at using these to discard misleading ranking schemes. We challenge that to let machines rule where they are best might be the best way for the community to enjoy the tremendous benefit of CASP as a unique opportunity for brainstorming.

Algorithms↗