Search PubMed⌕ Search

Biomedical subjects

Alfonso Valencia

Publications and source records attributed to Alfonso Valencia.

At least 37 records · Page 2Linked to original sources

Overview of BioCreAtIvE: critical assessment of information extraction for biology.

BACKGROUND: The goal of the first BioCreAtIvE challenge (Critical Assessment of Information Extraction in Biology) was to provide a set of common evaluation tasks to assess the state of the art for text mining applied to biological problems. The results were presented in a workshop held in Granada, Spain March 28-31, 2004. The articles collected in this BMC Bioinformatics supplement entitled "A critical assessment of text mining methods in molecular biology" describe the BioCreAtIvE tasks, systems, results and their independent evaluation. RESULTS: BioCreAtIvE focused on two tasks. The first dealt with extraction of gene or protein names from text, and their mapping into standardized gene identifiers for three model organism databases (fly, mouse, yeast). The second task addressed issues of functional annotation, requiring systems to identify specific text passages that supported Gene Ontology annotations for specific proteins, given full text articles. CONCLUSION: The first BioCreAtIvE assessment achieved a high level of international participation (27 groups from 10 countries). The assessment provided state-of-the-art performance results for a basic task (gene name finding and normalization), where the best systems achieved a balanced 80% precision / recall or better, which potentially makes them suitable for real applications in biology. The results for the advanced task (functional annotation from free text) were significantly lower, demonstrating the current limitations of text-mining approaches where knowledge extrapolation and interpretation are required. In addition, an important contribution of BioCreAtIvE has been the creation and release of training and test data sets for both tasks. There are 22 articles in this special issue, including six that provide analyses of results or data quality for the data sets, including a novel inter-annotator consistency assessment for the test set used in task 2.

Computational Biology↗

Evaluation of BioCreAtIvE assessment of task 2.

BACKGROUND: Molecular Biology accumulated substantial amounts of data concerning functions of genes and proteins. Information relating to functional descriptions is generally extracted manually from textual data and stored in biological databases to build up annotations for large collections of gene products. Those annotation databases are crucial for the interpretation of large scale analysis approaches using bioinformatics or experimental techniques. Due to the growing accumulation of functional descriptions in biomedical literature the need for text mining tools to facilitate the extraction of such annotations is urgent. In order to make text mining tools useable in real world scenarios, for instance to assist database curators during annotation of protein function, comparisons and evaluations of different approaches on full text articles are needed. RESULTS: The Critical Assessment for Information Extraction in Biology (BioCreAtIvE) contest consists of a community wide competition aiming to evaluate different strategies for text mining tools, as applied to biomedical literature. We report on task two which addressed the automatic extraction and assignment of Gene Ontology (GO) annotations of human proteins, using full text articles. The predictions of task 2 are based on triplets of protein--GO term--article passage. The annotation-relevant text passages were returned by the participants and evaluated by expert curators of the GO annotation (GOA) team at the European Institute of Bioinformatics (EBI). Each participant could submit up to three results for each sub-task comprising task 2. In total more than 15,000 individual results were provided by the participants. The curators evaluated in addition to the annotation itself, whether the protein and the GO term were correctly predicted and traceable through the submitted text fragment. CONCLUSION: Concepts provided by GO are currently the most extended set of terms used for annotating gene products, thus they were explored to assess how effectively text mining tools are able to extract those annotations automatically. Although the obtained results are promising, they are still far from reaching the required performance demanded by real world applications. Among the principal difficulties encountered to address the proposed task, were the complex nature of the GO terms and protein names (the large range of variants which are used to express proteins and especially GO terms in free text), and the lack of a standard training set. A range of very different strategies were used to tackle this task. The dataset generated in line with the BioCreative challenge is publicly available and will allow new possibilities for training information extraction methods in the domain of molecular biology.

Computational Biology↗

A sentence sliding window approach to extract protein annotations from biomedical articles.

BACKGROUND: Within the emerging field of text mining and statistical natural language processing (NLP) applied to biomedical articles, a broad variety of techniques have been developed during the past years. Nevertheless, there is still a great ned of comparative assessment of the performance of the proposed methods and the development of common evaluation criteria. This issue was addressed by the Critical Assessment of Text Mining Methods in Molecular Biology (BioCreative) contest. The aim of this contest was to assess the performance of text mining systems applied to biomedical texts including tools which recognize named entities such as genes and proteins, and tools which automatically extract protein annotations. RESULTS: The "sentence sliding window" approach proposed here was found to efficiently extract text fragments from full text articles containing annotations on proteins, providing the highest number of correctly predicted annotations. Moreover, the number of correct extractions of individual entities (i.e. proteins and GO terms) involved in the relationships used for the annotations was significantly higher than the correct extractions of the complete annotations (protein-function relations). CONCLUSION: We explored the use of averaging sentence sliding windows for information extraction, especially in a context where conventional training data is unavailable. The combination of our approach with more refined statistical estimators and machine learning techniques might be a way to improve annotation extraction for future biomedical text mining applications.

Biomedical Research↗

Text mining for metabolic pathways, signaling cascades, and protein networks.

The complexity of the information stored in databases and publications on metabolic and signaling pathways, the high throughput of experimental data, and the growing number of publications make it imperative to provide systems to help the researcher navigate through these interrelated information resources. Text-mining methods have started to play a key role in the creation and maintenance of links between the information stored in biological databases and its original sources in the literature. These links will be extremely useful for database updating and curation, especially if a number of technical problems can be solved satisfactorily, including the identification of protein and gene names (entities in general) and the characterization of their types of interactions. The first generation of openly accessible text-mining systems, such as iHOP (Information Hyperlinked over Proteins), provides additional functions to facilitate the reconstruction of protein interaction networks, combine database and text information, and support the scientist in the formulation of novel hypotheses. The next challenge is the generation of comprehensive information regarding the general function of signaling pathways and protein interaction networks.

Animals↗

Structure of the connector of bacteriophage T7 at 8A resolution: structural homologies of a basic component of a DNA translocating machinery.

The three-dimensional structure of the bacteriophage T7 head-to-tail connector has been obtained at 8A resolution using cryo-electron microscopy and single-particle analysis from purified recombinant connectors. The general morphology of the T7 connector is that of a 12-folded toroidal homopolymer with a channel that runs along the longitudinal axis of the particle. The structure of the T7 connector reveals many structural similarities with the connectors from other bacteriophages. Docking of the atomic structure of the varphi29 connector into the three-dimensional reconstruction of T7 connector reveals that the narrow, distal region of the two oligomers are almost identical. This region of the varphi29 connector has been suggested to be involved in DNA translocation, and is composed of an alpha-beta-alpha-beta-beta-alpha motif. A search for alpha-helices in the same region of the T7 three-dimensional map has located three alpha-helices in approximately the same position as those of the varphi29 connector. A comparison of the predicted secondary structure of several bacteriophage connectors, including among others T7, varphi29, P22 and SPP1, reveals that, despite the lack of sequence homology, they seem to contain the same alpha-beta-alpha-beta-beta-alpha motif as that present in the varphi29 connector. These results allow us to suggest a common architecture related to a basic component of the DNA translocating machinery for several viruses.

Bacteriophage T7↗

Text-mining approaches in molecular biology and biomedicine.

Biomedical articles provide functional descriptions of bioentities such as chemical compounds and proteins. To extract relevant information using automatic techniques, text-mining and information-extraction approaches have been developed. These technologies have a key role in integrating biomedical information through analysis of scientific literature. In this article, important applications such as the identification of biologically relevant entities in free text and the construction of literature-based networks of protein-protein interactions will be introduced. Also, the use of text mining to aid the interpretation of microarray data and the analysis of pathology reports will be discussed. Finally, we will consider the recent evolution of this field and the efforts for community-based evaluations.

Biomedical Research↗

HCAD, closing the gap between breakpoints and genes.

Recurrent chromosome aberrations are an important resource when associating human pathologies to specific genes. However, for technical reasons a large number of chromosome breakpoints are defined only at the level of cytobands and many of the genes involved remain unidentified. We developed a web-based information system that mines the scientific literature and generates textual and comprehensive information on all human breakpoints. We show that the statistical analysis of this textual information and its combination with genomic data can identify genes directly involved in DNA rearrangements. The Human Chromosome Aberration Database (HCAD) is publicly accessible at http://www.pdg.cnb.uam.es/UniPub/HCAD/.

Chromosome Breakage↗

MetaRouter: bioinformatics for bioremediation.

Bioremediation, the exploitation of biological catalysts (mostly microorganisms) for removing pollutants from the environment, requires the integration of huge amounts of data from different sources. We have developed MetaRouter, a system for maintaining heterogeneous information related to bioremediation in a framework that allows its query, administration and mining (application of methods for extracting new knowledge). MetaRouter is an application intended for laboratories working in biodegradation and bioremediation, which need to maintain and consult public and private data, linked internally and with external databases, and to extract new information from it. Among the data-mining features is a program included for locating biodegradative pathways for chemical compounds according to a given set of constraints and requirements. The integration of biodegradation information with the corresponding protein and genome data provides a suitable framework for studying the global properties of the bioremediation network. The system can be accessed and administrated through a web interface. The full-featured system (except administration facilities) is freely available at http://pdg.cnb.uam.es/MetaRouter. Additional material: http://www.pdg.cnb.uam.es/biodeg_net/MetaRouter.

Bacteria↗

Domain definition and target classification for CASP6.

Assessment of structure predictions in CASP6 was based on single domains isolated from experimentally determined structures, which were categorized into comparative modeling, fold recognition, and new fold targets. Domain definitions were defined upon visual examination of the structures with the aid of automated domain-parsing programs. Domain categorization was determined by comparison of the target structures with those in the Protein Data Bank at the time each target expired and a variety of sequence and structure-based methods to determine potential homologous relationships.

Amino Acid Sequence↗

Assessment of predictions submitted for the CASP6 comparative modeling category.

Here we present a full overview of the Critical Assessment of Protein Structure Prediction (CASP6) comparative modeling category. Prediction accuracy for the 43 comparative modeling targets was assessed through detailed numerical comparisons between predicted and experimental structures. Assessments using standard measures for model backbone quality and structural alignment accuracy highlighted a small number of groups with stand out predictions and these findings were backed up by statistical comparisons. We were able to carry out evaluations of side-chain contacts predictions and side-chain rotamer accuracy, for which one group turned out to have statistically better predictions. We also assessed the prediction quality of structurally divergent regions and biologically important sites. Interestingly we were able to show that predictors were not predicting these important functional regions with any greater accuracy than the rest of the structure. In addition we investigated the ability of predictors to build models that improve on the structural template and reached some tentative conclusions from comparisons with the previous CASP experiment.

Algorithms↗

CASP6 assessment of contact prediction.

Here we present the evaluation results of the Critical Assessment of Protein Structure Prediction (CASP6) contact prediction category. Contact prediction was assessed with standard measures well known in the field and the performance of specialist groups was evaluated alongside groups that submitted models with 3D coordinates. The evaluation was mainly focused on long range contact predictions for the set of new fold targets, although we analyzed predictions for all targets. Three groups with similar levels of accuracy and coverage performed a little better than the others. Comparisons of the predictions of the three best methods with those of CASP5/CAFASP3 suggested some improvement, although there were not enough targets in the comparisons to make this statistically significant.

Algorithms↗

Automatic annotation of protein function.

The annotation of protein function at genomic scale is essential for day-to-day work in biology and for any systematic approach to the modeling of biological systems. Currently, functional annotation is essentially based on the expansion of the relatively small number of experimentally determined functions to large collections of proteins. The task of systematic annotation faces formidable practical problems related to the accuracy of the input experimental information, the reliability of current systems for transferring information between related sequences, and the reproducibility of the links between database information and the original experiments reported in publications. These technical difficulties merely lie on the surface of the deeper problem of the evolution of protein function in the context of protein sequences and structures. Given the mixture of technical and scientific challenges, it is not surprising that errors are introduced, and expanded, in database annotations. In this situation, a more realistic option is the development of a reliability index for database annotations, instead of depending exclusively on efforts to correct databases. Several groups have attempted to compare the database annotations of similar proteins, which constitutes the first steps toward the calibration of the relationship between sequence and annotation space.

Artificial Intelligence↗

Death inducer obliterator protein 1 in the context of DNA regulation. Sequence analyses of distant homologues point to a novel functional role.

Death inducer obliterator protein 1 [DIDO1; also termed DIO-1 and death-associated transcription factor 1 (DATF-1)] is encoded by a gene thus far described only in higher vertebrates. Current gene ontology descriptions for this gene assign its function to an apoptosis-related process. The protein presents distinct splice variants and is distributed ubiquitously. Exhaustive sequence analyses of all DIDO variants identify distant homologues in yeast and other organisms. These homologues have a role in DNA regulation and chromatin stability, and form part of higher complexes linked to active chromatin. Further domain composition analyses performed in the context of related homologues suggest that DIDO-induced apoptosis is a secondary effect. Gene-targeted mice show alterations that include lagging chromosomes, and overexpression of the gene generates asymmetric nuclear divisions. Here we describe the analysis of these eukaryote-restricted proteins and propose a novel, DNA regulatory function for the DIDO protein in mammals.

Anaphase↗

ACRATA: a novel electron transfer domain associated to apoptosis and cancer.

BACKGROUND: Recently, several members of a vertebrate protein family containing a six trans-membrane (6TM) domain and involved in apoptosis and cancer (e.g. STEAP, STAMP1, TSAP6), have been identified in Golgi and cytoplasmic membranes. The exact function of these proteins remains unknown. METHODS: We related this 6TM domain to distant protein families using intermediate sequences and methods of iterative profile sequence similarity search. RESULTS: Here we show for the first time that this 6TM domain is homolog to the 6TM heme binding domain of both the NADPH oxidase (Nox) family and the YedZ family of bacterial oxidoreductases. CONCLUSIONS: This finding gives novel insights about the existence of a previously undetected electron transfer system involved in apoptosis and cancer, and suggests further steps in the experimental characterization of these evolutionarily related families.

Adaptor Proteins, Signal Transducing↗

SPOC: a widely distributed domain associated with cancer, apoptosis and transcription.

BACKGROUND: The Split ends (Spen) family are large proteins characterised by N-terminal RNA recognition motifs (RRMs) and a conserved SPOC (Spen paralog and ortholog C-terminal) domain. The aim of this study is to characterize the family at the sequence level. RESULTS: We describe undetected members of the Spen family in other lineages (Plasmodium and Plants) and localise SPOC in a new domain context, in a family that is common to all eukaryotes using profile-based sequence searches and structural prediction methods. CONCLUSIONS: The widely distributed DIO (Death inducer-obliterator) family is related to cancer and apoptosis and offers new clues about SPOC domain functionality.

Amino Acid Sequence↗

Structural model of carnitine palmitoyltransferase I based on the carnitine acetyltransferase crystal.

CPT I (carnitine palmitoyltransferase I) catalyses the conversion of palmitoyl-CoA into palmitoylcarnitine in the presence of L-carnitine, facilitating the entry of fatty acids into mitochondria. We propose a 3-D (three-dimensional) structural model for L-CPT I (liver CPT I), based on the similarity of this enzyme to the recently crystallized mouse carnitine acetyltransferase. The model includes 607 of the 773 amino acids of L-CPT I, and the positions of carnitine, CoA and the palmitoyl group were assigned by superposition and docking analysis. Functional analysis of this 3-D model included the mutagenesis of several amino acids in order to identify putative catalytic residues. Mutants D477A, D567A and E590D showed reduced L-CPT I activity. In addition, individual mutation of amino acids forming the conserved Ser685-Thr686-Ser687 motif abolished enzyme activity in mutants T686A and S687A and altered K(m) and the catalytic efficiency for carnitine in mutant S685A. We conclude that the catalytic residues are His473 and Asp477, while Ser687 probably stabilizes the transition state. Several conserved lysines, i.e. Lys455, Lys505, Lys560 and Lys561, were also mutated. Only mutants K455A and K560A showed decreases in activity of 50%. The model rationalizes the finding of nine natural mutations in patients with hereditary L-CPT I deficiencies.

Amino Acid Sequence↗

YAdumper: extracting and translating large information volumes from relational databases to structured flat files.

Downloading the information stored in relational databases into XML and other flat formats is a common task in bioinformatics. This periodical dumping of information requires considerable CPU time, disk and memory resources. YAdumper has been developed as a purpose-specific tool to deal with the integral structured information download of relational databases. YAdumper is a Java application that organizes database extraction following an XML template based on an external Document Type Declaration. Compared with other non-native alternatives, YAdumper substantially reduces memory requirements and considerably improves writing performance.

Algorithms↗