Search PubMed⌕ Search

Biomedical subjects

A Valencia

Publications and source records attributed to A Valencia.

At least 37 records · Page 2Linked to original sources

Intrinsic errors in genome annotation.

Genome sequencing is usually followed by routine annotation of protein function based on the assumption that similar sequences will have similar functions. Here, we introduce a simple calculation to estimate the magnitude of any possible annotation errors. We counted the number of discrepancies in the annotation of well-established sets of similar proteins and extrapolated these values to the pairs of similar sequences used for the annotation of different microbial genomes. We conclude that the number of potential errors in the prediction of detailed functions is higher than is usually believed.

Binding Sites↗

Escherichia coli FtsZ polymers contain mostly GTP and have a high nucleotide turnover.

The cell division protein FtsZ is a GTPase structurally related to tubulin and, like tubulin, it assembles in vitro into filaments, sheets and other structures. To study the roles that GTP binding and hydrolysis play in the dynamics of FtsZ polymerization, the nucleotide contents of FtsZ were measured under different polymerizing conditions using a nitrocellulose filter-binding assay, whereas polymerization of the protein was followed in parallel by light scattering. Unpolymerized FtsZ bound 1 mol of GTP mol(-1) protein monomer. At pH 7.5 and in the presence of Mg(2+) and K(+), there was a strong GTPase activity; most of the bound nucleotide was GTP during the first few minutes but, later, the amount of GTP decreased in parallel with depolymerization, whereas the total nucleotide contents remained invariant. These results show that the long FtsZ polymers formed in solution contain mostly GTP. Incorporation of nucleotides into the protein was very fast either when the label was introduced at the onset of the reaction or subsequently during polymerization. Molecular modelling of an FtsZ dimer showed the presence of a cleft between the two subunits maintaining the nucleotide binding site open to the medium. These results show that the FtsZ polymers are highly dynamic structures that quickly exchange the bound nucleotide, and this exchange can occur in all the subunits.

Bacterial Proteins↗

A la carte transcriptional regulators: unlocking responses of the prokaryotic enhancer-binding protein XylR to non-natural effectors.

To investigate the activation mechanism of the enhancer-binding protein XylR encoded by the TOL plasmid of Pseudomonas putida mt-2, a combinatorial library was generated composed of shuffled N-terminal A domains of the homologous regulators DmpR, XylR and TbuT, reassembled within the XylR structure. When the library was screened in vivo for responsiveness to non-effectors bulkier than one aromatic ring (such as biphenyl) or bearing an entirely different distribution of electronegative groups (e.g. nitrotoluenes), protein variants were found that displayed an expanded inducer range including the new effectors. Although the phenotypes endowed with the corresponding changes were largely similar, the modifications involved different sites within the A domain. The positions of the mutations within a structural model of the A domain suggest that expansion of the inducer profile can be brought about not only by changes in the effector pocket of the protein but also by unlocking steps of the signal transmission mechanism that follows effector binding. These results provide a rationale for evolving in vitro regulators à la carte that are responsive to predetermined, natural or xenobiotic chemical species.

Bacterial Proteins↗

EVA: continuous automatic evaluation of protein structure prediction servers.

UNLABELLED: Evaluation of protein structure prediction methods is difficult and time-consuming. Here, we describe EVA, a web server for assessing protein structure prediction methods, in an automated, continuous and large-scale fashion. Currently, EVA evaluates the performance of a variety of prediction methods available through the internet. Every week, the sequences of the latest experimentally determined protein structures are sent to prediction servers, results are collected, performance is evaluated, and a summary is published on the web. EVA has so far collected data for more than 3000 protein chains. These results may provide valuable insight to both developers and users of prediction methods. AVAILABILITY: http://cubic.bioc.columbia.edu/eva. CONTACT: eva@cubic.bioc.columbia.edu

Automation↗

A hierarchical unsupervised growing neural network for clustering gene expression patterns.

MOTIVATION: We describe a new approach to the analysis of gene expression data coming from DNA array experiments, using an unsupervised neural network. DNA array technologies allow monitoring thousands of genes rapidly and efficiently. One of the interests of these studies is the search for correlated gene expression patterns, and this is usually achieved by clustering them. The Self-Organising Tree Algorithm, (SOTA) (Dopazo,J. and Carazo,J.M. (1997) J. Mol. Evol., 44, 226-233), is a neural network that grows adopting the topology of a binary tree. The result of the algorithm is a hierarchical cluster obtained with the accuracy and robustness of a neural network. RESULTS: SOTA clustering confers several advantages over classical hierarchical clustering methods. SOTA is a divisive method: the clustering process is performed from top to bottom, i.e. the highest hierarchical levels are resolved before going to the details of the lowest levels. The growing can be stopped at the desired hierarchical level. Moreover, a criterion to stop the growing of the tree, based on the approximate distribution of probability obtained by randomisation of the original data set, is provided. By means of this criterion, a statistical support for the definition of clusters is proposed. In addition, obtaining average gene expression patterns is a built-in feature of the algorithm. Different neurons defining the different hierarchical levels represent the averages of the gene expression patterns contained in the clusters. Since SOTA runtimes are approximately linear with the number of items to be classified, it is especially suitable for dealing with huge amounts of data. The method proposed is very general and applies to any data providing that they can be coded as a series of numbers and that a computable measure of similarity between data items can be used. AVAILABILITY: A server running the program can be found at: http://bioinfo.cnio.es/sotarray.

Algorithms↗

Prediction of contact maps with neural networks and correlated mutations.

Contact maps of proteins are predicted with neural network-based methods, using as input codings of increasing complexity including evolutionary information, sequence conservation, correlated mutations and predicted secondary structures. Neural networks are trained on a data set comprising the contact maps of 173 non-homologous proteins as computed from their well resolved three-dimensional structures. Proteins are selected from the Protein Data Bank database provided that they align with at least 15 similar sequences in the corresponding families. The predictors are trained to learn the association rules between the covalent structure of each protein and its contact map with a standard back propagation algorithm and tested on the same protein set with a cross-validation procedure. Our results indicate that the method can assign protein contacts with an average accuracy of 0.21 and with an improvement over a random predictor of a factor >6, which is higher than that previously obtained with methods only based either on neural networks or on correlated mutations. Furthermore, filtering the network outputs with a procedure based on the residue coordination numbers, the accuracy of predictions increases up to 0.25 for all the proteins, with an 8-fold deviation from a random predictor. These scores are the highest reported so far for predicting protein contact maps.

Algorithms↗

Similarity of phylogenetic trees as indicator of protein-protein interaction.

Deciphering the network of protein interactions that underlines cellular operations has become one of the main tasks of proteomics and computational biology. Recently, a set of bioinformatics approaches has emerged for the prediction of possible interactions by combining sequence and genomic information. Even though the initial results are very promising, the current methods are still far from perfect. We propose here a new way of discovering possible protein-protein interactions based on the comparison of the evolutionary distances between the sequences of the associated protein families, an idea based on previous observations of correspondence between the phylogenetic trees of associated proteins in systems such as ligands and receptors. Here, we extend the approach to different test sets, including the statistical evaluation of their capacity to predict protein interactions. To demonstrate the possibilities of the system to perform large-scale predictions of interactions, we present the application to a collection of more than 67 000 pairs of E.coli proteins, of which 2742 are predicted to correspond to interacting proteins.

Chaperonin 60↗

Heteromeric amino acid transporters: biochemistry, genetics, and physiology.

The heteromeric amino acid transporters (HATs) are composed of two polypeptides: a heavy subunit (HSHAT) and a light subunit (LSHAT) linked by a disulfide bridge. HSHATs are N-glycosylated type II membrane glycoproteins, whereas LSHATs are nonglycosylated polytopic membrane proteins. The HSHATs have been known since 1992, and the LSHATs have been described in the last three years. HATs represent several of the classic mammalian amino acid transport systems (e.g., L isoforms, y(+)L isoforms, asc, x(c)(-), and b(0,+)). Members of the HAT family are the molecular bases of inherited primary aminoacidurias cystinuria and lysinuric protein intolerance. In addition to the role in amino acid transport, one HSHAT [the heavy subunit of the cell-surface antigen 4F2 (also named CD98)] is involved in other cell functions that might be related to integrin activation. This review covers the biochemistry, human genetics, and cell physiology of HATs, including the multifunctional character of CD98.

Amino Acid Sequence↗

The potential use of SUISEKI as a protein interaction discovery tool.

Relevant information about protein interactions is stored in textual sources. This sources are commonly used not only as archives of what is already known but also as information for generating new knowledge, particularly to pose hypothesis about new possible interactions that can be inferred from the existing ones. This task is the more creative part of scientific work in experimental systems. We present a large-scale analysis for the prediction of new interactions based on the interaction network for the ones already known and detected automatically in the literature. During the last few years it has became clear that part of the information about protein interactions could be extracted with automatic tools, even if these tools are still far from perfect and key problems such as detection of protein names are not completely solved. We have developed a integrated automatic approach, called SUISEKI (System for Information Extraction on Interactions), able to extract protein interactions from collections of Medline abstracts. Previous experiments with the system have shown that it is able to extract almost 70% of the interactions present in relatively large text corpus, with an accuracy of approximately 80% (for the best defined interactions) that makes the system usable in real scenarios, both at the level of extraction of protein names and at the level of extracting interaction between them. With the analysis of the interaction map of Saccharomyces cerevisiae we show that interactions published in the years 2000/2001 frequently correspond to proteins or genes that were already very close in the interaction network deduced from the literature published before these years and that they are often connected to the same proteins. That is, discoveries are commonly done among highly connected entities. Some biologically relevant examples illustrate how interactions described in the year 2000 could have been proposed as reasonable working hypothesis with the information previously available in the automatically extracted network of interactions.

Cell Cycle↗

Genome sequences and great expectations.

To assess how automatic function assignment will contribute to genome annotation in the next five years, we have performed an analysis of 31 available genome sequences. An emerging pattern is that function can be predicted for almost two-thirds of the 73,500 genes that were analyzed. Despite progress in computational biology, there will always be a great need for large-scale experimental determination of protein function.

Animals↗

Three-dimensional view of the surface motif associated with the P-loop structure: cis and trans cases of convergent evolution.

Here we identify the determinants of the nucleotide-binding ability associated with the P-loop-containing proteins, inferring their functional importance from their structural convergence to a unique three- dimensional (3D) motif. (1) A new surface 3D pattern is identified for the P-loop nucleotide-binding region, which is more selective than the corresponding sequence pattern; (2) the signature displays one residue that we propose is the determinant for the guanine-binding ability (the residues aligned to ras D119; this residue is known to be important only in the G-proteins, we extend the prediction to all the other P-loop- containing proteins); and (3) two cases of convergent evolution at the molecular level are highlighted in the analysis of the active site: the positive charge aligned to ras K117 and the arginine residues aligned to the GAP arginine finger. The analysis of the residues conserved on protein surfaces allows one to identify new functional or evolutionary relationships among protein structures that would not be detectable by conventional sequence or structure comparison methods.

Adenosine Triphosphate↗

Practical limits of function prediction.

The widening gap between known protein sequences and their functions has led to the practice of assigning a potential function to a protein on the basis of sequence similarity to proteins whose function has been experimentally investigated. We present here a critical view of the theoretical and practical bases for this approach. The results obtained by analyzing a significant number of true sequence similarities, derived directly from structural alignments, point to the complexity of function prediction. Different aspects of protein function, including (i) enzymatic function classification, (ii) functional annotations in the form of key words, (iii) classes of cellular function, and (iv) conservation of binding sites can only be reliably transferred between similar sequences to a modest degree. The reason for this difficulty is a combination of the unavoidable database inaccuracies and the plasticity of protein function. In addition, analysis of the relationship between sequence and functional descriptions defines an empirical limit for pairwise-based functional annotations, namely, the three first digits of the six numbers used as descriptors of protein folds in the FSSP database can be predicted at an average level as low as 7.5% sequence identity, two of the four EC digits at 15% identity, half of the SWISS-PROT key words related to protein function would require 20% identity, and the prediction of half of the residues in the binding site can be made at the 30% sequence identity level.

Amino Acid Sequence↗

[Compliance with inhalation treatment of patients with chronic obstructive pulmonary disease].

OBJECTIVES: To evaluate patient compliance with inhaled medication therapy in chronic obstructive pulmonary disease (COPD), to identify determining factors and to propose corrective measures to improve compliance. METHODS: This was an open, observational, cross-sectional, non-comparative, single-measurement, non-random study. The inhalers were the Serevent Accuhaler, the Serevent Inhalador and the Flixotide Inhaler. Compliance was measured in four ways: a) difference in weight at the beginning and end of the study for all devices; b) dose counter reading for the Accuhaler; c) information from patient diaries (by days and by applications); and d) information from patient interviews using the Morinsky-Green Test. Compliance was rated as follows: poor: < 50%, fair 51%-79%, good 80%-119%, or "hypercompliant" > 120%. RESULTS: Seventy-two patients (mean age 65 years) were enrolled. Compliance measured by weight was good in 77.1%, fair in 11.5%, poor in 1.4% and hypercompliant in 10%. Compliance was good for the Accuhaler according to both weight (75%) and counted doses (83.3%). According to patient diaries, compliance was good when assessed by applications (98.8%) and by days (98.3%). According to the Morinksky-Green test, compliance was good for 87.9%. CONCLUSIONS: Compliance was good as assessed by the methods used in this study. Patients who live in families, who enjoy a high socioeconomic level, have simple therapeutic regimens and have a good understanding of their disease and inhaler tend to have good compliance. Careful patient follow-up and good patient-physician communication has improved compliance. However, follow-up studies are needed to check these results.

Administration, Inhalation↗

[The internal consistency and content validity of the Spanish version of the Asthma Autonomy Questionnaire].

The Asthma Autonomy Questionnaire (AAQ) was designed to evaluate asthmatics' desire to learn about their disease and to make decisions. The AAQ consists of 26 items distributed in two scales: Preferences in the Search for Information (PSI, 8 items) and Preferences in Decision Making (PDM, 6 general items and 12 related to 3 scenarios depicting asthma in stable phase, during mild exacerbation and during severe exacerbation). The aim of this study was to analyze the internal consistency (Cronbach's-coefficient) and content validity (factorial analysis of principal components) of the AAQ. After translation and back translation, the Spanish version of the AAQ was administered to 115 adult asthmatics of both sexes and differing levels of severity. The alpha coefficients for the two scales and 3 scenarios ranged from 0.42 (PSI) to 0.73 (stable phase scenario); only for the stable-phase scenario were values high or statistically acceptable. Factorial analysis reproduced the content of the scales only approximately, with some items proving to relate to factors that were different from the scale they originally belonged to. These results indicate that, in its current formulation, the AAQ presents important measurement problems and revision is advisable.

Adult↗

Alpha-amylases of the coffee berry borer (Hypothenemus hampei) and their inhibition by two plant amylase inhibitors.

The adult coffee berry borer (Hypothenemus hampei Ferrari [Coleoptera: Scolytidae]), a major insect pest of coffee, has two major digestive alpha-amylases that can be separated by isoelectric focusing. The alpha-amylase activity has a broad pH optimum between 4.0 and 7.0. Using pH indicators, the pH of the midgut was determined to be between 4.5 and 5.2. At pH 5.0, the coffee berry borer alpha-amylase activity is inhibited substantially (80%) by relatively low levels of the amylase inhibitor (alphaAI-1) from the common bean, Phaseolus vulgaris L., and much less so by the amylase inhibitor from Amaranthus. We used an in-gel zymogram assay to demonstrate that seed extracts can be screened to find suitable inhibitors of amylases. The prospect of using the genes that encode these inhibitors to make coffee resistant to the coffee berry borer via genetic engineering is discussed.

Animals↗

Expression profiles and biological function.

Expression arrays facilitate the monitoring of changes in expression patterns of large collections of genes. It is generally expected that genes with similar expression patterns would correspond to proteins of common biological function. We assess this common assumption by comparing levels of similarity of expression patterns and statistical significance of biological terms that describe the corresponding protein functions. Terms are automatically obtained by mining large collections of Medline abstracts. We propose that the combined use of the tools for expression profiles clustering and automatic function retrieval, can be useful tools for the detection of biologically relevant associations between genes in complex gene expression experiments. The results obtained using publicly available experimental data show how, in general, an increase in the similarity of the expression patterns is accompanied by an enhancement of the amount of specific functional information or, in other words, how the selected terms became more specific following an increase in the specificity of the expression patterns. Particularly interesting are the discrepancies from this general trend, i.e. groups of genes with similar expression patterns but very little in common at the functional level. In these cases the similarity of their expression profiles becomes the first link between previously unrelated genes.

Cluster Analysis↗

Effective use of sequence correlation and conservation in fold recognition.

Protein families are a rich source of information; sequence conservation and sequence correlation are two of the main properties that can be derived from the analysis of multiple sequence alignments. Sequence conservation is related to the direct evolutionary pressure to retain the chemical characteristics of some positions in order to maintain a given function. Sequence correlation is attributed to the small sequence adjustments needed to maintain protein stability against constant mutational drift. Here, we showed that sequence conservation and correlation were each frequently informative enough to detect incorrectly folded proteins. Furthermore, combining conservation, correlation, and polarity, we achieved an almost perfect discrimination between native and incorrectly folded proteins. Thus, we made use of this information for threading by evaluating the models suggested by a threading method according to the degree of proximity of the corresponding correlated, conserved, and apolar residues. The results showed that the fold recognition capacity of a given threading approach could be improved almost fourfold by selecting the alignments that score best under the three different sequence-based approaches.

Amino Acid Sequence↗