Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Sequence signature analysis of chromosome identity in three Drosophila species.

BACKGROUND: All eukaryotic organisms need to distinguish each of their chromosomes. A few protein complexes have been described that recognise entire, specific chromosomes, for instance dosage compensation complexes and the recently discovered autosome-specific Painting of Fourth (POF) protein in Drosophila. However, no sequences have been found that are chromosome-specific and distributed over the entire length of the respective chromosome. Here, we present a new, unbiased, exhaustive computational method that was used to probe three Drosophila genomes for chromosome-specific sequences. RESULTS: By combining genome annotations and cytological data with multivariate statistics related to three Drosophila genomes we found sequence signatures that distinguish Muller's F-elements (chromosome 4 in D. melanogaster) from all other chromosomes in Drosophila that are not attributable to differences in nucleotide composition, simple sequence repeats or repeated elements. Based on these signatures we identified complex motifs that are strongly overrepresented in the F-elements and found indications that the D. melanogaster motif may be involved in POF-binding to the F-element. In addition, the X-chromosomes of D. melanogaster and D. yakuba can be distinguished from the other chromosomes, albeit to a lesser extent. Surprisingly, the conservation of the F-element sequence signatures extends not only between species separated by approximately 55 Myr, but also linearly along the sequenced part of the F-elements. CONCLUSION: Our results suggest that chromosome-distinguishing features are not exclusive to the sex chromosomes, but are also present on at least one autosome (the F-element) in Drosophila.

Animals↗

Classifying the precancers: a metadata approach.

BACKGROUND: During carcinogenesis, precancers are the morphologically identifiable lesions that precede invasive cancers. In theory, the successful treatment of precancers would result in the eradication of most human cancers. Despite the importance of these lesions, there has been no effort to list and classify all of the precancers. The purpose of this study is to describe the first comprehensive taxonomy and classification of the precancers. As a novel approach to disease classification, terms and classes were annotated with metadata (data that describes the data) so that the classification could be used to link precancer terms to data elements in other biological databases. METHODS: Terms in the UMLS (Unified Medical Language System) related to precancers were extracted. Extracted terms were reviewed and additional terms added. Each precancer was assigned one of six general classes. The entire classification was assembled as an XML (eXtensible Mark-up Language) file. A Perl script converted the XML file into a browser-viewable HTML (HyperText Mark-up Language) file. RESULTS: The classification contained 4700 precancer terms, 568 distinct precancer concepts and six precancer classes: 1) Acquired microscopic precancers; 2) acquired large lesions with microscopic atypia; 3) Precursor lesions occurring with inherited hyperplastic syndromes that progress to cancer; 4) Acquired diffuse hyperplasias and diffuse metaplasias; 5) Currently unclassified entities; and 6) Superclass and modifiers. CONCLUSION: This work represents the first attempt to create a comprehensive listing of the precancers, the first attempt to classify precancers by their biological properties and the first attempt to create a pathologic classification of precancers using standard metadata (XML). The classification is placed in the public domain, and comment is invited by the authors, who are prepared to curate and modify the classification.

Decision Support Systems, Clinical↗

The TREC 2004 genomics track categorization task: classifying full text biomedical documents.

BACKGROUND: The TREC 2004 Genomics Track focused on applying information retrieval and text mining techniques to improve the use of genomic information in biomedicine. The Genomics Track consisted of two main tasks, ad hoc retrieval and document categorization. In this paper, we describe the categorization task, which focused on the classification of full-text documents, simulating the task of curators of the Mouse Genome Informatics (MGI) system and consisting of three subtasks. One subtask of the categorization task required the triage of articles likely to have experimental evidence warranting the assignment of GO terms, while the other two subtasks were concerned with the assignment of the three top-level GO categories to each paper containing evidence for these categories. RESULTS: The track had 33 participating groups. The mean and maximum utility measure for the triage subtask was 0.3303, with a top score of 0.6512. No system was able to substantially improve results over simply using the MeSH term Mice. Analysis of significant feature overlap between the training and test sets was found to be less than expected. Sample coverage of GO terms assigned to papers in the collection was very sparse. Determining papers containing GO term evidence will likely need to be treated as separate tasks for each concept represented in GO, and therefore require much denser sampling than was available in the data sets. The annotation subtask had a mean F-measure of 0.3824, with a top score of 0.5611. The mean F-measure for the annotation plus evidence codes subtask was 0.3676, with a top score of 0.4224. Gene name recognition was found to be of benefit for this task. CONCLUSION: Automated classification of documents for GO annotation is a challenging task, as was the automated extraction of GO code hierarchies and evidence codes. However, automating these tasks would provide substantial benefit to biomedical curation, and therefore work in this area must continue. Additional experience will allow comparison and further analysis about which algorithmic features are most useful in biomedical document classification, and better understanding of the task characteristics that make automated classification feasible and useful for biomedical document curation. The TREC Genomics Track will be continuing in 2005 focusing on a wider range of triage tasks and improving results from 2004.

Journal Article↗

Emergency protocols in prehospital emergency medicine in Slovenia.

Emergency medicine is a branch of medicine in which prompt and accurate information is of crucial importance, because a patient's survival and the final outcome of a treatment very often depends on it. Therefore the standardization of prehospital emergency service administration with uniform documentation, methods of annotation of technical data as well as uniform gathering and data analysis are necessary. Wanting to process a fair amount of data and specially to disburden personnel, a need for computer aided processing is evident. Application that was developed in strong collaboration with EMS project team and information science engineers definitely improves data quality. Firstly when team members, especially physicians, are constrained to fulfill paper forms and secondly when data are entered into the computer, because all inconsistencies are being reconsidered and corrected if necessary. The developed application also gave the opportunity to the emergency team for prompt and accurate information.

Data Collection↗

Evidence-based guidelines for management of nursing home-acquired pneumonia.

We convened a multidisciplinary, multispecialty panel to develop comprehensive evidence and consensus-based guidelines for managing nursing home-acquired pneumonia. The panel began with explicit criteria for process of care quality measures, performed a comprehensive review of the English-language literature, evaluated the quality of the evidence, and drafted a set of proposed guidelines. The panel reviewed the draft, an annotated bibliography, and data from a study of 30-day survival from nursing home-acquired pneumonia, and then participated in an all-day meeting in January 2001. Using a modified Delphi process, the panel refined the guidelines and developed a care pathway. The guidelines recommend a comprehensive approach, including immunization of staff and residents, and communication between nursing staff and the attending physician within 2 hours of symptom onset. Probable pneumonia was defined. An algorithm was delineated for assessing the patientamprsquos wishes for hospitalization and aggressive care, and deciding on hospitalization based on the severity of the illness as well as the capacity of the nursing home to provide acute care. The timing and extent of evaluation in a nursing home relative to the rapid initiation of antibiotics should depend on whether the patient has any unstable vital signs. An antibiotic covering Streptococcus pneumoniae, Haemophilus influenzae, common gram-negative rods, and Staphylococcus aureus should be given for 10 to 14 days, orally if the patient is able to take medications by mouth.

Delphi Technique↗

IMGT, the international ImMunoGeneTics information system, http://imgt.cines.fr: the reference in immunoinformatics.

IMGT, the international ImMunoGeneTics information system (http://imgt.cines.fr), is a high quality integrated information system specializing in immunoglobulins (IG), T cell receptors (TR), major histocompatibility complex (MHC) and related proteins of the immune system of human and other vertebrates, created in 1989, by the Laboratoire d'ImmunoGénétique Moléculaire (LIGM), at the Université Montpellier II, CNRS, Montpellier, France. IMGT is the global reference in immunogenetics and immunoinformatics and provides a common access to standardized data which include nucleotide and protein sequences, oligonucleotide primers, gene maps, genetic polymorphisms, specificities, 2D and 3D structures. IMGT includes three sequence databases (IMGT/LIGM-DB, IMGT/MHC-DB hosted at EBI, IMGT/PRIMER-DB), one genome database (IMGT/GENE-DB), one 3D structure database (IMGT/3Dstructure-DB), Web resources comprising 8000 HTML pages ("IMGT Marie-Paule page") and interactive tools for sequence (IMGT/V-QUEST, IMGT/JunctionAnalysis, IMGT/Allele-Align, IMGT/PhyloGene) and genome (IMGT/GeneSearch, IMGT/GeneView, IMGT/LocusView) analysis. IMGT data are expertly annotated according to the rules of the IMGT Scientific chart, based on the IMGT-ONTOLOGY concepts. IMGT tools are particularly useful for the analysis of the IG and TR repertoires in physiological normal and pathological situations. IMGT has important applications in medical research (repertoire analysis in autoimmune diseases, AIDS, leukemias, lymphomas, myelomas), biotechnology related to antibody engineering (phage displays, combinatorial libraries) and therapeutic approaches (graft, immunotherapy). IMGT is freely available at http://imgt.cines.fr.

Animals↗

An assessment of the resistance gene analogues of Oryza sativa ssp. japonica: their presence and structure.

Rice is the first cereal genome of known draft sequence, and the finished sequence for it is now nearly complete. In this paper, we describe a preliminary analysis of known rice genes aimed to detect resistance gene analogues of known structural classes. Putative resistance genes were identified in a dual approach--by using BLASTP searches to identify candidate sequences and by using Hidden Markov Models to predict domain presence in the candidates. The set of proteins examined was obtained from the publicly available data of TIGR (The Institute for Genomic Research). 1744 distinct RGAs were identified, 597 of which belonged to the NBS-LRR class. Supplementary data (sequences and annotations) is available on the web site http:/gkoczyk.bioinfo.pl/CMBL.

Computational Biology↗

IMGT, the international ImMunoGeneTics information system, http://imgt.cines.fr.

IMGT, the international ImMunoGeneTics information system (http://imgt.cines.fr), is a high quality integrated knowledge resource specializing in immunoglobulins (IG), T cell receptors (TR) and major histocompatibility complexes (MHC) and related proteins of the immune system (RPI) of human and other vertebrates, created in 1989 by LIGM at the Université Montpellier II, CNRS, Montpellier, France. IMGT provides a common access to standardized data which include nucleotide and protein sequences, oligonucleotide primers, gene maps, genetic polymorphisms, specificities, and 2D and 3D structures. IMGT includes five databases (IMGT/LIGM-DB, IMGT/3Dstructure-DB, IMGT/MHC-DB, IMGT/PRIMER-DB, IMGT/GENE-DB) Web resources ('IMGT Marie-Paule page') and interactive tools (IMGT/V-QUEST, IMGT/JunctionAnalysis, IMGT/PhyloGene, IMGT/LocusView, IMGT/GeneView, IMGT/GeneSearch, IMGT/StructureQuery). IMGT data are expertly annotated according to the rules of the IMGT Scientific chart based on IMGTONTOLOGY. IMGT tools are particularly useful for the analysis of the IG and TR repertoires in physiological normal and pathological situations. IMGT has important applications in medical research (autoimmune diseases, AIDS, leukaemias, lymphomas, myelomas), biotechnology related to antibody engineering (phage displays, combinatorial libraries) and therapeutic approaches (graft, immunotherapy). IMGT is freely available at http://imgt.cines.fr.

Animals↗

A specification for defining and annotating regions of macromolecular structures.

We present a program- and machine-independent standard for annotating macromolecular structures. Data encoded by this specification may be used for communicating information about structures and for exchanging it between different computer systems. The format consists of a set of ASN.1 objects which are mechanically straightforward to parse, but are also easy for humans to create and understand. It differs from all other related standards in that it specifies how a molecule should be displayed without requiring a custom format for the coordinate data.

Amino Acid Sequence↗

The Feasibility of Using Proteome Expression Profile for Genome Annotation.

By investigating into the expression data from ECO2DBASE (Edition 6),the feasibility of using proteome expression profile for genome annotation was tested. Based on our newly developed CRC (cellular role cluster) method,79 proteins extracted from ECO2DBASE were clustered into 4 CRCs. Function related proteins tend to be clustered into same CRC. Total 9 aminoacyl-tRNA synthetases were clustered into CRC2, whereas 4 heat-shock proteins into CRC3. These results indicate with enough proteome expression data and the efficient algorithm, proteome expression profile can provide very important information for genome annotation, while this kind of information is sequence-independent.

Journal Article↗

MitoMorphy: an alignment and annotation tool for human mitochondrial DNA polymorphisms.

MitoMorphy uses a number of publicly available human mitochondrial DNA (mtDNA) sequences from different ethnic groups to compare and annotate the associated polymorphic data. It provides an integrated display of mtDNA sequence comparison, sequence variation, and annotation for 695 different mtDNA sequences from many different ethnic groups around the world.

Journal Article↗

Building an automated classification of DNA-binding protein domains.

Intensive growth in 3D structure data on DNA-protein complexes as reflected in the Protein Data Bank (PDB) demands new approaches to the annotation and characterization of these data and will lead to a new understanding of critical biological processes involving these data. These data and those from other protein structure classifications will become increasingly important for the modeling of complete proteomes. We propose a fully automated classification of DNA-binding protein domains based on existing 3D-structures from the PDB. The classification, by domain, relies on the Protein Domain Parser (PDP) and the Combinatorial Extension (CE) algorithm for structural alignment. The approach involves the analysis of 3D-interaction patterns in DNA-protein interfaces, assignment of structural domains interacting with DNA, clustering of domains based on structural similarity and DNA-interacting patterns. Comparison with existing resources on describing structural and functional classifications of DNA-binding proteins was used to validate and improve the approach proposed here. In the course of our study we defined a set of criteria and heuristics allowing us to automatically build a biologically meaningful classification and define classes of functionally related protein domains. It was shown that taking into consideration interactions between protein domains and DNA considerably improves the classification accuracy. Our approach provides a high-throughput and up-to-date annotation of DNA-binding protein families which can be found at http://spdc.sdsc.edu.

Artificial Intelligence↗

Semantic similarity measures as tools for exploring the gene ontology.

Many bioinformatics resources hold data in the form of sequences. Often this sequence data is associated with a large amount of annotation. In many cases this data has been hard to model, and has been represented as scientific natural language, which is not readily computationally amenable. The development of the Gene Ontology provides us with a more accessible representation of some of this data. However it is not clear how this data can best be searched, or queried. Recently we have adapted information content based measures for use with the Gene Ontology (GO). In this paper we present detailed investigation of the properties of these measures, and examine various properties of GO, which may have implications for its future design.

Classification↗

Improved method for predicting linear B-cell epitopes.

BACKGROUND: B-cell epitopes are the sites of molecules that are recognized by antibodies of the immune system. Knowledge of B-cell epitopes may be used in the design of vaccines and diagnostics tests. It is therefore of interest to develop improved methods for predicting B-cell epitopes. In this paper, we describe an improved method for predicting linear B-cell epitopes. RESULTS: In order to do this, three data sets of linear B-cell epitope annotated proteins were constructed. A data set was collected from the literature, another data set was extracted from the AntiJen database and a data sets of epitopes in the proteins of HIV was collected from the Los Alamos HIV database. An unbiased validation of the methods was made by testing on data sets on which they were neither trained nor optimized on. We have measured the performance in a non-parametric way by constructing ROC-curves. CONCLUSION: The best single method for predicting linear B-cell epitopes is the hidden Markov model. Combining the hidden Markov model with one of the best propensity scale methods, we obtained the BepiPred method. When tested on the validation data set this method performs significantly better than any of the other methods tested. The server and data sets are publicly available at http://www.cbs.dtu.dk/services/BepiPred.

Journal Article↗

Evaluation of an ambulatory system for the quantification of cough frequency in patients with chronic obstructive pulmonary disease.

BACKGROUND: To date, methods used to assess cough have been primarily subjective, and only broadly reflect the impact of chronic cough and/or chronic cough therapies on quality of life. Objective assessment of cough has been attempted, but early techniques were neither ambulatory nor feasible for long-term data collection. We evaluated a novel ambulatory cardio-respiratory monitoring system with an integrated unidirectional, contact microphone, and report here the results from a study of patients with COPD who were videotaped in a quasi-controlled environment for 24 continuous hours while wearing the ambulatory system. METHODS: Eight patients with a documented history of COPD with ten or more years of smoking (6 women; age 57.4 +/- 11.8 yrs.; percent predicted FEV1 49.6 +/- 13.7%) who complained of cough were evaluated in a clinical research unit equipped with video and sound recording capabilities. All patients wore the LifeShirt system (LS) while undergoing simultaneous video (with sound) surveillance. Video data were visually inspected and annotated to indicate all cough events. Raw physiologic data records were visually inspected by technicians who remained blinded to the video data. Cough events from LS were analyzed quantitatively with a specialized software algorithm to identify cough. The output of the software algorithm was compared to video records on a per event basis in order to determine the validity of the LS system to detect cough in COPD patients. RESULTS: Video surveillance identified a total of 3,645 coughs, while LS identified 3,363 coughs during the same period. The median cough rate per patient was 21.3 coughs.hr-1 (Range: 10.1 cghs.hr(-1) - 59.9 cghs.hr(-1)). The overall accuracy of the LS system was 99.0%. Overall sensitivity and specificity of LS, when compared to video surveillance, were 0.781 and 0.996, respectively, while positive- and negative-predictive values were 0.846 and 0.994. There was very good agreement between the LS system and video (kappa = 0.807). CONCLUSION: The LS system demonstrated a high level of accuracy and agreement when compared to video surveillance for the measurement of cough in patients with COPD.

Journal Article↗

Integrating biological data through the genome.

Owing to the ongoing success of the genome sequencing and structural genomics projects, the increase in both sequence and structural data is rapid. The development of tools for the annotation of sequence and structural data has become more important in the hope of keeping up with this data explosion. Scientists in this field have addressed these issues over the last 10 years and there now exists a wealth of methods and approaches to help interpret these data. However, there is no current way in which these methods can be incorporated easily so that the resulting annotations can be viewed together. This review discusses the development of these annotation methods and introduces the BioSapiens Network of Excellence, which has been formed in order to integrate the methods which have been developed in Europe.

Animals↗

Threshold protocol for the exchange of confidential medical data.

BACKGROUND: Medical researchers often need to share clinical data without violating patient confidentiality. Threshold cryptographic protocols divide messages into multiple pieces, no single piece containing information that can reconstruct the original message. The author describes and implements a novel threshold protocol that can be used to search, annotate or transform confidential data without breaching patient confidentiality. METHODS: The basic threshold protocol is: 1) Text is divided into short phrases; 2) Each phrase is converted by a one-way hash algorithm into a seemingly-random set of characters; 3) Threshold Piece 1 is composed of the list of all phrases, with each phrase followed by its one-way hash; 4) Threshold Piece 2 is composed of the text with all phrases replaced by their one-way hash values, and with high-frequency words preserved. Neither Piece 1 nor Piece 2 contains information linking patients to their records. The original text can be re-constructed from Piece 1 and Piece 2. RESULTS: The threshold algorithm produces two files (threshold pieces). In typical usage, Piece 2 is held by the data owner, and Piece 1 is freely distributed. Piece 1 can be annotated and returned to the owner of the original data to enhance the complete data set. Collections of Piece 1 files can be merged and distributed without identifying patient records. Variations of the threshold protocol are described. The author's Perl implementation is freely available. CONCLUSIONS: Threshold files are safe in the sense that they are de-identified and can be used for research purposes. The threshold protocol is particularly useful when the receiver of the threshold file needs to obtain certain concepts or data-types found in the original data, but does not need to fully understand the original data set.

Algorithms↗

Zone analysis in biology articles as a basis for information extraction.

In the field of biomedicine, an overwhelming amount of experimental data has become available as a result of the high throughput of research in this domain. The amount of results reported has now grown beyond the limits of what can be managed by manual means. This makes it increasingly difficult for the researchers in this area to keep up with the latest developments. Information extraction (IE) in the biological domain aims to provide an effective automatic means to dynamically manage the information contained in archived journal articles and abstract collections and thus help researchers in their work. However, while considerable advances have been made in certain areas of IE, pinpointing and organizing factual information (such as experimental results) remains a challenge. In this paper we propose tackling this task by incorporating into IE information about rhetorical zones, i.e. classification of spans of text in terms of argumentation and intellectual attribution. As the first step towards this goal, we introduce a scheme for annotating biological texts for rhetorical zones and provide a qualitative and quantitative analysis of the data annotated according to this scheme. We also discuss our preliminary research on automatic zone analysis, and its incorporation into our IE framework.

Abstracting and Indexing↗