Search PubMed⌕ Search

Biomedical subjects

Alfonso Valencia

Publications and source records attributed to Alfonso Valencia.

At least 19 recordsLinked to original sources

The Network of National COVID-19 Data Portals: public health equity through collaboration.

The network of the national COVID-19 Data Portals was developed and linked to the COVID-19 Data Portal (https://www.covid19dataportal.org/)inresponsetothe need for rapid data sharing and analysis during the 2020-2022 SARS-CoV-2 pandemic. Built on open-source code developed by the Swedish COVID-19 Data Portal (now the Swedish Pathogens Portal, www.pathogens.se) the network included 12 national portals addressing demand for local open data sharing and access, across data types and resources. It provides a robust case study of national initiatives for FAIR (Findable, Accessible, Interoperable and Reusable) resources and a foundation for future pandemic preparedness across pathogens globally. In this paper we outline the structure of the origins of the network of National COVID-19 Datal Portals, the technical aspects and code originating from the Swedish Portal and provide an overview of the services and tools offered by each Portal. The paper showcases the process and operation of four Portals: Sweden, Poland, Spain, Norway and The Netherlands. In this study, we observe that pandemic response greatly benefits from an established infrastructure that can be quickly mobilised, developed and extended. Collaborations and preparation built on solid foundations over several years, supported by investment in the form of national and international research grants, is key for sustainability, continuation and readiness to deploy such efforts.

COVID-19↗

The organization of the microbial biodegradation network from a systems-biology perspective.

Microbial biodegradation of environmental pollutants is a field of growing importance because of its potential use in bioremediation and biocatalysis. We have studied the characteristics of the global biodegradation network that is brought about by all the known chemical reactions that are implicated in this process, regardless of their microbial hosts. This combination produces an efficient and integrated suprametabolism, with properties similar to those that define metabolic networks in single organisms. The characteristics of this network support an evolutionary scenario in which the reactions evolved outwards from the central metabolism. The properties of the global biodegradation network have implications for predicting the fate of current and future environmental pollutants.

Bacteria↗

Predicting reliable regions in protein alignments from sequence profiles.

For applications such as comparative modelling one major issue is the reliability of sequence alignments. Reliable regions in alignments can be predicted using sub-optimal alignments of the same pair of sequences. Here we show that reliable regions in alignments can also be predicted from multiple sequence profile information alone. Alignments were created for a set of remotely related pairs of proteins using five different test methods. Structural alignments were used to assess the quality of the alignments and the aligned positions were scored using information from the observed frequencies of amino acid residues in sequence profiles pre-generated for each template structure. High-scoring regions of these profile-derived alignment scores were a good predictor of reliably aligned regions. These profile-derived alignment scores are easy to obtain and are applicable to any alignment method. They can be used to detect those regions of alignments that are reliably aligned and to help predict the quality of an alignment. For those residues within secondary structure elements, the regions predicted as reliably aligned agreed with the structural alignments for between 92% and 97.4% of the residues. In loop regions just under 92% of the residues predicted to be reliable agreed with the structural alignments. The percentage of residues predicted as reliable ranged from 32.1% for helix residues to 52.8% for strand residues. This information could also be used to help predict conserved binding sites from sequence alignments. Residues in the template that were identified as binding sites, that aligned to an identical amino acid residue and where the sequence alignment agreed with the structural alignment were in highly conserved, high scoring regions over 80% of the time. This suggests that many binding sites that are present in both target and template sequences are in sequence-conserved regions and that there is the possibility of translating reliability to binding site prediction.

Algorithms↗

EVA: Evaluation of protein structure prediction servers.

EVA (http://cubic.bioc.columbia.edu/eva/) is a web server for evaluation of the accuracy of automated protein structure prediction methods. The evaluation is updated automatically each week, to cope with the large number of existing prediction servers and the constant changes in the prediction methods. EVA currently assesses servers for secondary structure prediction, contact prediction, comparative protein structure modelling and threading/fold recognition. Every day, sequences of newly available protein structures in the Protein Data Bank (PDB) are sent to the servers and their predictions are collected. The predictions are then compared to the experimental structures once a week; the results are published on the EVA web pages. Over time, EVA has accumulated prediction results for a large number of proteins, ranging from hundreds to thousands, depending on the prediction method. This large sample assures that methods are compared reliably. As a result, EVA provides useful information to developers as well as users of prediction methods.

Automation↗

Structural (betaalpha)8 TIM barrel model of 3-hydroxy-3-methylglutaryl-coenzyme A lyase.

This study describes three novel homozygous missense mutations (S75R, S201Y, and D204N) in the 3-hydroxy-3-methylglutaryl-CoA (HMG-CoA) lyase gene, which caused 3-hydroxy-3-methylglutaric aciduria in patients from Germany, England, and Argentina. Expression studies in Escherichia coli show that S75R and S201Y substitutions completely abolished the HMG-CoA lyase activity, whereas D204N reduced catalytic efficiency to 6.6% of the wild type. We also propose a three-dimensional model for human HMG-CoA lyase containing a (betaalpha)8 (TIM) barrel structure. The model is supported by the similarity with analogous TIM barrel structures of functionally related proteins, by the localization of catalytic amino acids at the active site, and by the coincidence between the shape of the substrate (HMG-CoA) and the predicted inner cavity. The three novel mutations explain the lack of HMG-CoA lyase activity on the basis of the proposed structure: in S75R and S201Y because the new amino acid residues occlude the substrate cavity, and in D204N because the mutation alters the electrochemical environment of the active site. We also report the localization of all missense mutations reported to date and show that these mutations are located in the beta-sheets around the substrate cavity.

Amino Acid Sequence↗

Evaluation of annotation strategies using an entire genome sequence.

MOTIVATION: Genome-wide functional annotation either by manual or automatic means has raised considerable concerns regarding the accuracy of assignments and the reproducibility of methodologies. In addition, a performance evaluation of automated systems that attempt to tackle sequence analyses rapidly and reproducibly is generally missing. In order to quantify the accuracy and reproducibility of function assignments on a genome-wide scale, we have re-annotated the entire genome sequence of Chlamydia trachomatis (serovar D), in a collaborative manner. RESULTS: We have encoded all annotations in a structured format to allow further comparison and data exchange and have used a scale that records the different levels of potential annotation errors according to their propensity to propagate in the database due to transitive function assignments. We conclude that genome annotation may entail a considerable amount of errors, ranging from simple typographical errors to complex sequence analysis problems. The most surprising result of this comparative study is that automatic systems might perform as well as the teams of experts annotating genome sequences.

Amino Acid Sequence↗

Automatic methods for predicting functionally important residues.

Sequence analysis is often the first guide for the prediction of residues in a protein family that may have functional significance. A few methods have been proposed which use the division of protein families into subfamilies in the search for those positions that could have some functional significance for the whole family, but at the same time which exhibit the specificity of each subfamily ("Tree-determinant residues"). However, there are still many unsolved questions like the best division of a protein family into subfamilies, or the accurate detection of sequence variation patterns characteristic of different subfamilies. Here we present a systematic study in a significant number of protein families, testing the statistical meaning of the Tree-determinant residues predicted by three different methods that represent the range of available approaches. The first method takes as a starting point a phylogenetic representation of a protein family and, following the principle of Relative Entropy from Information Theory, automatically searches for the optimal division of the family into subfamilies. The second method looks for positions whose mutational behavior is reminiscent of the mutational behavior of the full-length proteins, by directly comparing the corresponding distance matrices. The third method is an automation of the analysis of distribution of sequences and amino acid positions in the corresponding multidimensional spaces using a vector-based principal component analysis. These three methods have been tested on two non-redundant lists of protein families: one composed by proteins that bind a variety of ligand groups, and the other composed by proteins with annotated functionally relevant sites. In most cases, the residues predicted by the three methods show a clear tendency to be close to bound ligands of biological relevance and to those amino acids described as participants in key aspects of protein function. These three automatic methods provide a wide range of possibilities for biologists to analyze their families of interest, in a similar way to the one presented here for the family of proteins related with ras-p21.

Algorithms↗

Phage-display and correlated mutations identify an essential region of subdomain 1C involved in homodimerization of Escherichia coli FtsA.

FtsA plays an essential role in Escherichia coli cell division and is nearly ubiquitous in eubacteria. Several evidences postulated the ability of FtsA to interact with other septation proteins and with itself. To investigate these binding properties, we screened a phage-display library with FtsA. The isolated peptides defined a degenerate consensus sequence, which in turn displayed a striking similarity with residues 126-133 of FtsA itself. This result suggested that residues 126-133 were involved in homodimerization of FtsA. The hypothesis was supported by the analysis of correlated mutations, which identified a mutual relationship between a group of amino acids encompassing the ATP-binding site and a set of residues immediately downstream to amino acids 126-133. This information was used to assemble a model of a FtsA homodimer, whose accuracy was confirmed by probing multiple alternative docking solutions. Moreover, a prediction of residues responsible for protein-protein interaction validated the proposed model and confirmed once more the importance of residues 126-133 for homodimerization. To functionally characterize this region, we introduced a deletion in ftsA, where residues 126-133 were skipped. This mutant failed to complement conditional lethal alleles of ftsA, demonstrating that amino acids 126-133 play an essential role in E. coli.

Amino Acid Sequence↗

Reductive genome evolution in Buchnera aphidicola.

We have sequenced the genome of the intracellular symbiont Buchnera aphidicola from the aphid Baizongia pistacea. This strain diverged 80-150 million years ago from the common ancestor of two previously sequenced Buchnera strains. Here, a field-collected, nonclonal sample of insects was used as source material for laboratory procedures. As a consequence, the genome assembly unveiled intrapopulational variation, consisting of approximately 1,200 polymorphic sites. Comparison of the 618-kb (kbp) genome with the two other Buchnera genomes revealed a nearly perfect gene-order conservation, indicating that the onset of genomic stasis coincided closely with establishment of the symbiosis with aphids, approximately 200 million years ago. Extensive genome reduction also predates the synchronous diversification of Buchnera and its host; but, at a slower rate, gene loss continues among the extant lineages. A computational study of protein folding predicts that proteins in Buchnera, as well as proteins of other intracellular bacteria, are generally characterized by smaller folding efficiency compared with proteins of free living bacteria. These and other degenerative genomic features are discussed in light of compensatory processes and theoretical predictions on the long-term evolutionary fate of symbionts like Buchnera.

Base Sequence↗

Life cycles of successful genes.

By exploring time-series data from MEDLINE abstracts, we observe that only a few genes have been quoted with increasing frequency during the past 25 years. This is probably the result of selective pressure by the scientific community. Over the years, this selection has produced an extreme power law distribution of the information available for individual genes. Interestingly, those genes that are successfully selected are not necessarily the most important genes to the cell. To stress the implication of this finding we show that there is no correlation between a gene's impact in the scientific literature and its centrality in protein-interaction networks.

Cell Cycle↗

Involvement of intramolecular interactions in the regulation of G protein-coupled receptor kinase 2.

The G protein-coupled receptor (GPCR) kinase GRK2 phosphorylates G protein-coupled receptors in an agonist-dependent manner. GRK2 activity is modulated through interactions of diverse domains of the kinase with G protein betagamma subunits, several lipids, anchoring proteins, and activated receptors. We report that kinase activity toward either GPCR (rhodopsin) or a synthetic peptide substrate is enhanced in the presence of GST-GRK2 fusion proteins or peptides corresponding to either N- or C-terminal sequences of GRK2. This direct stimulatory action of intrinsic domains on GRK2 activity does not add to the effect of other regulators, such as Gbetagamma subunits, and strongly suggests the existence of some mode of autoregulation. The existence of regulatory intramolecular interactions in GRK2 is supported by the facts that a C-terminal peptide protects the N-terminal region from proteolytic cleavage and that two domains of GRK2 independently coexpressed in cells associate as assessed by immunoprecipitation. Molecular modeling suggests that intramolecular interactions among the N-terminal, C-terminal and kinase domains would keep GRK2 in a constrained conformation characteristic of an inactive, basal state. Our model proposes that disruption of such intramolecular contacts by intermolecular interactions with regulatory proteins (mimicked by exogenously added kinase fragments in vitro) would promote the conformational changes required to bring about GRK2 translocation and activation.

Animals↗

Beta-propellers: associated functions and their role in human diseases.

The beta-propeller fold appears as a very fascinating architecture based on four-stranded antiparallel and twisted beta-sheets, radially arranged around a central tunnel. Similar to the alpha/beta-barrel (TIM-barrel) fold, the beta-propeller has a wide range of different functions, and is gaining substantial attention. Some proteins containing beta-propeller domains have been implicated in the pathogenesis of a variety of diseases such as cancer, Alzheimer, Huntington, arthritis, familial hypercholesterolemia, retinitis pigmentosa, osteogenesis, hypertension, and microbial and viral infections. This article reviews some aspects of 3D structure, amino acids sequence regularities, and biological functions of the proteins containing beta-propeller domains. Major emphasis has been laid on beta-propellers whose functions are associated to human diseases. Recent research efforts reported in the fields of protein engineering, drug design, and protein structure-function relationship studies, concerning the beta-propeller architecture, have also been discussed.

Disease↗

Identification of conserved amino acid residues in rat liver carnitine palmitoyltransferase I critical for malonyl-CoA inhibition. Mutation of methionine 593 abolishes malonyl-CoA inhibition.

Carnitine palmitoyltransferase (CPT) I, which catalyzes the conversion of palmitoyl-CoA to palmitoylcarnitine facilitating its transport through the mitochondrial membranes, is inhibited by malonyl-CoA. By using the SequenceSpace algorithm program to identify amino acids that participate in malonyl-CoA inhibition in all carnitine acyltransferases, we found 5 conserved amino acids (Thr(314), Asn(464), Ala(478), Met(593), and Cys(608), rat liver CPT I coordinates) common to inhibitable malonyl-CoA acyltransferases (carnitine octanoyltransferase and CPT I), and absent in noninhibitable malonyl-CoA acyltransferases (CPT II, carnitine acetyltransferase (CAT) and choline acetyltransferase (ChAT)). To determine the role of these amino acid residues in malonyl-CoA inhibition, we prepared the quintuple mutant CPT I T314S/N464D/A478G/M593S/C608A as well as five single mutants CPT I T314S, N464D, A478G, M593S, and C608A. In each case the CPT I amino acid selected was mutated to that present in the same homologous position in CPT II, CAT, and ChAT. Because mutant M593S nearly abolished the sensitivity to malonyl-CoA, two other Met(593) mutants were prepared: M593A and M593E. The catalytic efficiency (V(max)/K(m)) of CPT I in mutants A478G and C608A and all Met(593) mutants toward carnitine as substrate was clearly increased. In those CPT I proteins in which Met(593) had been mutated, the malonyl-CoA sensitivity was nearly abolished. Mutations in Ala(478), Cys(608), and Thr(314) to their homologous amino acid residues in CPT II, CAT, and ChAT caused various decreases in malonyl-CoA sensitivity. Ala(478) is located in the structural model of CPT I near the catalytic site and participates in the binding of malonyl-CoA in the low affinity site (Morillas, M., Gómez-Puertas, P., Rubi, B., Clotet, J., Ariño, J., Valencia, A., Hegardt, F. G., Serra, D., and Asins, G. (2002) J. Biol. Chem. 277, 11473-11480). Met(593) may participate in the interaction of malonyl-CoA in the second affinity site, whose location has not been reported.

Alanine↗

p23 and HSP20/alpha-crystallin proteins define a conserved sequence domain present in other eukaryotic protein families.

We identified families of proteins characterized by the presence of a domain similar to human p23 protein, which include proteins such as Sgt1, involved in the yeast kinetochore assembly; melusin, involved in specific interactions with the cytoplasmic integrin beta1 domain; Rar1, related to pathogenic resistance in plants, and to development in animals; B5+B5R flavo-hemo cytochrome NAD(P)H oxidoreductase type B in humans and mice; and NudC, involved in nucleus migration during mitosis. We also found that p23 and the HSP20/alpha-crystallin family of heat shock proteins, which share the same three-dimensional folding, show a pattern of conserved residues that points to a common origin in the evolution of both protein domains. The p23 and HSP20/alpha-crystallin phylogenetic relationship and their similar role in chaperone activity suggest a common function, probably involving protein-protein interaction, for those proteins containing p23-like domains.

Amino Acid Sequence↗

Bioinformatics methods for the analysis of expression arrays: data clustering and information extraction.

Expression arrays facilitate the monitoring of changes in the expression patterns of large collections of genes. The analysis of expression array data has become a computationally-intensive task that requires the development of bioinformatics technology for a number of key stages in the process, such as image analysis, database storage, gene clustering and information extraction. Here, we review the current trends in each of these areas, with particular emphasis on the development of the related technology being carried out within our groups.

Abstracting and Indexing↗

In silico two-hybrid system for the selection of physically interacting protein pairs.

Deciphering the interaction links between proteins has become one of the main tasks of experimental and bioinformatic methodologies. Reconstruction of complex networks of interactions in simple cellular systems by integrating predicted interaction networks with available experimental data is becoming one of the most demanding needs in the postgenomic era. On the basis of the study of correlated mutations in multiple sequence alignments, we propose a new method (in silico two-hybrid, i2h) that directly addresses the detection of physically interacting protein pairs and identifies the most likely sequence regions involved in the interactions. We have applied the system to several test sets, showing that it can discriminate between true and false interactions in a significant number of cases. We have also analyzed a large collection of E. coli protein pairs as a first step toward the virtual reconstruction of its complete interaction network.

Animals↗