Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

The Los Alamos hepatitis C sequence database.

MOTIVATION: The hepatitis C virus (HCV) is a significant threat to public health worldwide. The virus is highly variable and evolves rapidly, making it an elusive target for the immune system and for vaccine and drug design. At present, some 30 000 HCV sequences have been published. A central website that provides annotated sequences and analysis tools will be helpful to HCV scientists worldwide. RESULTS: The HCV sequence database collects and annotates sequence data and provides them to the public via a website that contains a user-friendly search interface and a large number of sequence analysis tools, based on the model of the highly regarded Los Alamos HIV database. The HCV sequence database was officially launched in September 2003. Since then, its usage has steadily increased and is now at an average of approximately 280 visits per day from distinct IP addresses. AVAILABILITY: The HCV website can be accessed via http://hcv.lanl.gov and http://hcv-db.org.

Amino Acid Sequence↗

A fuzzy guided genetic algorithm for operon prediction.

MOTIVATION: The operon structure of the prokaryotic genome is a critical input for the reconstruction of regulatory networks at the whole genome level. As experimental methods for the detection of operons are difficult and time-consuming, efforts are being put into developing computational methods that can use available biological information to predict operons. METHOD: A genetic algorithm is developed to evolve a starting population of putative operon maps of the genome into progressively better predictions. Fuzzy scoring functions based on multiple criteria are used for assessing the 'fitness' of the newly evolved operon maps and guiding their evolution. RESULTS: The algorithm organizes the whole genome into operons. The fuzzy guided genetic algorithm-based approach makes it possible to use diverse biological information like genome sequence data, functional annotations and conservation across multiple genomes, to guide the organization process. This approach does not require any prior training with experimental operons. The predictions from this algorithm for Escherchia coli K12 and Bacillus subtilis are evaluated against experimentally discovered operons for these organisms. The accuracy of the method is evaluated using an ROC (receiver operating characteristic) analysis. The area under the ROC curve is around 0.9, which indicates excellent accuracy. CONTACT: roschen_csir@rediffmail.com.

Algorithms↗

SELEX_DB: a database on in vitro selected oligomers adapted for recognizing natural sites and for analyzing both SNPs and site-directed mutagenesis data.

SELEX_DB is an online resource containing both the experimental data on in vitro selected DNA/RNA oligomers (aptamers) and the applets for recognition of these oligomers. Since in vitro experimental data are evidently system-dependent, the new release of the SELEX_DB has been supplemented by the database SYSTEM storing the experimental design. In addition, the recognition applet package, SELEX_TOOLS, applying in vitro selected data to annotation of the genome DNA, is accompanied by the cross-validation test database CROSS_TEST discriminating the sites (natural or other) related to in vitro selected sites out of random DNA. By cross-validation testing, we have unexpectedly observed that the recognition accuracy increases with the growth of homology between the training and test sets of protein binding sequences. For natural sites, the recognition accuracy was lower than that for the nearest protein homologs and higher than that for distant homologs and non-homologous proteins binding the common site. The current SELEX_DB release is available at http://wwwmgs.bionet.nsc.ru/mgs/systems/selex/.

Binding Sites↗

On the spatial disposition of the fifth transmembrane helix and the structural integrity of the transmembrane binding site in the opioid and ORL1 G protein-coupled receptor family.

Evidence from statistical cluster analyses of a multiple sequence alignment of G protein-coupled receptor seven-helix folds supports the existence of structurally conserved transmembrane (TM) ligand binding sites in the opioid/opioid receptor-like (ORL1) and amine receptor families. Based on the expectation that functionally conserved regions in homologous proteins will display locally higher levels of sequence identity compared with global sequence similarities that pertain to the overall fold, this approach may have wider applications in functional genomics to annotate sequence data. Binding sites in models of the kappa-opioid receptor seven-helix bundle built from the rhodopsin templates of Baldwin et al. (1997) [J. Mol. Biol., 272, 144-164] and Herzyk and Hubbard (1998) [J. Mol. Biol., 281, 742-751] are compared. The Herzyk and Hubbard template is found to be in better accord with experimental studies of amine, opioid and rhodopsin receptors owing to the reduced physical separation of the extracellular parts of TM helices V and VI and differences in the rotational orientation of the N-terminal of helix V that reveal side chain accessibilities in the Baldwin et al. structure to be out of phase with relative alkylation rates of engineered cysteine residues in the TM binding site of the alpha(2A)-adrenergic receptor. TM helix V in the Baldwin et al. template has been remodelled with a different proline kink to satisfy experimental constraints. A recent proposal that rotation of helix V is associated with receptor activation is critically discussed.

Binding Sites↗

Research in physical medicine and rehabilitation. V. Data entry and early exploratory data analysis.

The process of data entry and initial analysis to locate data errors is described. Basic terms are defined and a simple method of entering data by using word processing software is illustrated. Data checking is done by using visual check of the raw data. Statistical programs are then used to locate possible data errors by finding data points (outliers) that are very different from the average. Special graphic output of statistical programs, scatterplots and box and whisker plots can be used to further locate questionable data. Examples of data entry forms and annotated step by step data cleaning with the use of inexpensive programs for personal computers are presented.

Computers↗

A systematic approach to standardizing the visual appearance of endometriotic lesions for artificial intelligence recognition.

INTRODUCTION: Numerous studies have shown that the diagnostic performance and reproducibility of visual recognition of endometriosis during laparoscopy are poor. The use of artificial intelligence (AI) seems relevant for exhaustive lesion recognition. Standardization of the visual classification of lesions, in the form of an ontology, is an essential prerequisite to enable medical experts to annotate surgical data consistently and subsequently allow engineers to train and build an artificial intelligence tool for endometriosis recognition. MATERIAL AND METHODS: A systematic search was conducted in the MEDLINE (via PubMed), EMBASE, and the Cochrane Library databases up to May 2022, aiming to identify studies describing the laparoscopic visual appearance of superficial endometriosis, endometriomas, and deep infiltrating endometriosis. The accumulated data in the literature concerning the visual appearance of the different forms of endometriosis were used to create an ontology that could be used for artificial intelligence applications. RESULTS: Out of 932 articles screened, 35 studies were selected based on the inclusion criteria of human subjects with histologically confirmed endometriosis lesions visualized via laparoscopy. The selected studies were reviewed to develop a visual ontology of endometriosis lesions observed via laparoscopy. The lesions were categorized into 4 classes and further subdivided into 11 subclasses: superficial (black, red, white, or subtle), adhesions (dense or filmy), deep (obliteration, retraction, or deformation), and ovarian (endometrioma or chocolate fluid). The positive predictive value (PPV) varied across lesion types: black lesions (PPV 47%-97%), red lesions (PPV 33%-100%), white lesions (PPV 20%-81%), and ovarian endometriosis (PPV 42%-98%). Nonspecific lesions such as adhesions (PPV 16%-50%) and subtle superficial lesions (PPV 0%-67%) presented lower PPVs. Deep endometriosis lesions, often buried within organs, required indirect signs (obliteration, retraction, deformation) for identification. CONCLUSIONS: The visual ontology proposed in this systematic search could facilitate the detection and classification of endometriosis lesions using artificial intelligence. This study highlights the challenges of reaching a consensus on lesion recognition and classification in AI projects due to the diverse visual presentations of endometriosis.

Humans↗

A first global analysis of plasmid encoded proteins in the ACLAME database.

Many plasmids are mobile genetic elements (MGEs) and, as other members of that group of DNA entities, their genomes display a mosaic and combinatorial structure, making their classification extremely difficult. As other MGEs, plasmids play a major role in horizontal transfer of genetic materials and genome reorganization. Yet, the full impact of such phenomenon on major properties of the host cell, such as pathogenicity, the ability to use new carbon sources or resistance to antibiotics, remains to be fully assessed. More and more complete plasmid genome sequences are available. However, in the absence of standards for storing plasmid sequence data and annotating genes and gene products on sequenced plasmid genomes, the resulting information remains rather limited. Using 503 sequenced plasmids organized in the ACLAME database, we discuss how, by structuring information on the genomes, their host and the proteins they code for, one can gain access to either global or more detailed analysis of the plasmid sequence information, as illustrated by a network representation of the relationships between plasmids.

Bacteria↗

A physical map of 30,000 human genes.

A map of 30,181 human gene-based markers was assembled and integrated with the current genetic map by radiation hybrid mapping. The new gene map contains nearly twice as many genes as the previous release, includes most genes that encode proteins of known function, and is twofold to threefold more accurate than the previous version. A redesigned, more informative and functional World Wide Web site (www.ncbi.nlm.nih.gov/genemap) provides the mapping information and associated data and annotations. This resource constitutes an important infrastructure and tool for the study of complex genetic traits, the positional cloning of disease genes, the cross-referencing of mammalian genomes, and validated human transcribed sequences for large-scale studies of gene expression.

Animals↗

The treatment of attention-deficit hyperactivity disorder: an annotated bibliography and critical appraisal of published systematic reviews and metaanalyses.

CONTEXT: The Agency for Health Care Policy and Research charged the McMaster Evidence-based Practice Center with conducting a comprehensive systematic review of the literature on the treatment of attention-deficit hyperactivity disorder (ADHD), with input from various groups of stakeholders. One strategy used to avoid duplication of work included a critical appraisal of existing systematic reviews and metaanalyses. OBJECTIVE: To identify and appraise published metaanalyses and systematic reviews on the treatment of ADHD and to produce an annotated bibliography. DATA SOURCES: Medline, Cumulative Index in Nursing and Allied Health (CINAHL), Healthstar, Psycinfo, and Embase were searched to September 1998; the Cochrane Database (1998 issue 3), selected Internet sites, and the files of investigators were also reviewed. STUDY SELECTION: Review articles described as systematic reviews or metaanalyses or including a Methods section were identified independently by 3 reviewers. DATA EXTRACTION: Two reviewers extracted, by consensus, relevant information on the name, methodological quality, ADHD-related aspects (comorbid disorders, family characteristics) of those reviews; data on the population, study setting, interventions, and outcomes evaluated by the reviews were also retrieved. RESULTS: Thirteen reviews, published from 1982 to 1998, were included. Eight included metaanalysis and 5 a qualitative review. Nonpharmacological treatments were mentioned in 6 reviews and combination therapies in 3. One review focused on the treatment of adults. Forty-seven drugs and 20 adverse effects were mentioned. Most reviews had major methodological flaws. CONCLUSIONS: Most published systematic reviews and metaanalyses on the treatment of ADHD have limited value for guiding clinical, policy, and research decisions. A rigorous, systematic review following established methodological criteria is warranted.

Attention Deficit Disorder with Hyperactivity↗

Statistical extraction of Drosophila cis-regulatory modules using exhaustive assessment of local word frequency.

BACKGROUND: Transcription regulatory regions in higher eukaryotes are often represented by cis-regulatory modules (CRM) and are responsible for the formation of specific spatial and temporal gene expression patterns. These extended, approximately 1 KB, regions are found far from coding sequences and cannot be extracted from genome on the basis of their relative position to the coding regions. RESULTS: To explore the feasibility of CRM extraction from a genome, we generated an original training set, containing annotated sequence data for most of the known developmental CRMs from Drosophila. Based on this set of experimental data, we developed a strategy for statistical extraction of cis-regulatory modules from the genome, using exhaustive analysis of local word frequency (LWF). To assess the performance of our analysis, we measured the correlation between predictions generated by the LWF algorithm and the distribution of conserved non-coding regions in a number of Drosophila developmental genes. CONCLUSIONS: In most of the cases tested, we observed high correlation (up to 0.6-0.8, measured on the entire gene locus) between the two independent techniques. We discuss computational strategies available for extraction of Drosophila CRMs and possible extensions of these methods.

Animals↗

Database development in toxicogenomics: issues and efforts.

The marriage of toxicology and genomics has created not only opportunities but also novel informatics challenges. As with the larger field of gene expression analysis, toxicogenomics faces the problems of probe annotation and data comparison across different array platforms. Toxicogenomics studies are generally built on standard toxicology studies generating biological end point data, and as such, one goal of toxicogenomics is to detect relationships between changes in gene expression and in those biological parameters. These challenges are best addressed through data collection into a well-designed toxicogenomics database. A successful publicly accessible toxicogenomics database will serve as a repository for data sharing and as a resource for analysis, data mining, and discussion. It will offer a vehicle for harmonizing nomenclature and analytical approaches and serve as a reference for regulatory organizations to evaluate toxicogenomics data submitted as part of registrations. Such a database would capture the experimental context of in vivo studies with great fidelity such that the dynamics of the dose response could be probed statistically with confidence. This review presents the collaborative efforts between the European Molecular Biology Laboratory-European Bioinformatics Institute ArrayExpress, the International Life Sciences Institute Health and Environmental Science Institute, and the National Institute of Environmental Health Sciences National Center for Toxigenomics Chemical Effects in Biological Systems knowledge base. The goal of this collaboration is to establish public infrastructure on an international scale and examine other developments aimed at establishing toxicogenomics databases. In this review we discuss several issues common to such databases: the requirement for identifying minimal descriptors to represent the experiment, the demand for standardizing data storage and exchange formats, the challenge of creating standardized nomenclature and ontologies to describe biological data, the technical problems involved in data upload, the necessity of defining parameters that assess and record data quality, and the development of standardized analytical approaches.

Animals↗

Online genomics facilities in the new millennium.

The review begins by providing a brief typology of biological databases on the Internet, illustrated by examples of the most influential resources of each kind. We then take an insider look at one typical on-line genomic resource -- the yeast genome database hosted at the Munich Information Center for Protein Sequences (MIPS) -- and explain how and why it has evolved from a basic sequence repository to a multidomain knowledge base. The role of community efforts in curating and annotating genome data is discussed. The crucial role of data integration and interoperability in developing next-generation genomic facilities is underscored.

Animals↗

How Not to be Seen: Predicting Unseen Enzyme Functions using Contrastive Learning.

MOTIVATION: Predicting enzyme function from its sequence is still an unsolved problem in the life sciences. Moreover, with the explosion of annotated genome data, we are inundated with potential enzymatic sequences that have not yet been biochemically characterized. While it is not possible to assign a not-yet-existing label to such a sequence, there is high value in placing the sequence as accurately as possible in known function space. Doing so can help provide more accurate falsifiable hypotheses for experimentalists wishing to characterize enzymes from specific functional families. RESULTS: Here we present a contrastive learning algorithm for predicting enzyme function from sequence. Our method, EnzPlacer, predicts the third, second, and first EC numbers for a protein whose fourth EC number is not in the training corpus. This novel prediction mechanism accurately places a protein sequence within a narrowed-down functional context, even if the precise function remains unknown. AVAILABILITY: EnzPlacer is available from https://github.com/drxiangma/EnzPlacer under a GPL3 license.

Contrastive learning↗

YMD: a microarray database for large-scale gene expression analysis.

The use of microarray technology to perform parallel analysis of the expression pattern of a large number of genes in a single experiment has created a new frontier of medical research. The vast amount of gene expression data generated from multiple microarray experiments requires a robust database system that allows efficient data storage, retrieval, secure access, data dissemination, and integrated data analyses. To address the growing needs of microarray researchers at Yale and their collaborators, we have built the Yale Microarray Database (YMD). YMD is Web-accessible with the following features: (i) a Web program that tracks DNA samples between source plates and arrays, (ii) the capability of finding common genes/clones across different array platforms, (iii) an image file server, (iv) laboratory-based user management and access privileges, (v) project management, (vi) template data entry, (vii) linking gene expression data to annotation databases for functional analysis. YMD is currently being used on a pilot basis by several laboratories for different organisms and array platforms.

Databases, Nucleic Acid↗

HIV sequence databases.

Two important databases are often used in HIV genetic research, the HIV Sequence Database in Los Alamos, which collects all sequences and focuses on annotation and data analysis, and the HIV RT/Protease Sequence Database in Stanford, which collects sequences associated with the development of viral resistance against anti-retroviral drugs and focuses on analysis of those sequences. The types of data and services these two databases offer, the tools they provide, and the way they are set up and operated are described in detail.

Anti-HIV Agents↗

Health-e-child: an integrated biomedical platform for grid-based paediatric applications.

There is a compelling demand for the integration and exploitation of heterogeneous biomedical information for improved clinical practice, medical research, and personalised healthcare across the EU. The Health-e-Child project aims at developing an integrated healthcare platform for European Paediatrics, providing seamless integration of traditional and emerging sources of biomedical information. The long-term goal of the project is to provide uninhibited access to universal biomedical knowledge repositories for personalised and preventive healthcare, large-scale information-based biomedical research and training, and informed policy making. The project focus will be on individualized disease prevention, screening, early diagnosis, therapy and follow-up of paediatric heart diseases, inflammatory diseases, and brain tumours. The project will build a Grid-enabled European network of leading clinical centres that will share and annotate biomedical data, validate systems clinically, and diffuse clinical excellence across Europe by setting up new technologies, clinical workflows, and standards. This paper outlines the design approach being adopted in Health-e-Child to enable the delivery of an integrated biomedical information platform.

Databases as Topic↗

Visualization techniques for genomic data.

In order to take full advantage of the newly available public human genome sequence data and associated annotations, biologists require visualization tools that can accommodate the high frequency of alternative splicing in human genes and other complexities. In this article, we describe techniques for presenting human genomic sequence data and annotations in an interactive, graphical format, with the aim of providing developers with a guide to what features are most likely to meet biologists' needs. These techniques include: one-dimensional semantic zooming to show sequence data alongside gene structures; moveable, adjustable tiers; visual encoding of translation frame to show how alternative transcript structure affects encoded proteins; and display of protein domains in the context of genomic sequence to show how alternative splicing impacts protein structure and function.

Algorithms↗

Rapid and selective surveillance of Arabidopsis thaliana genome annotations with Centrifuge.

UNLABELLED: Centrifuge is a user-friendly system to simultaneously access Arabidopsis gene annotations and intra- and inter-organism sequence comparison data. The tool allows rapid retrieval of user-selected data for each annotated Arabidopsis gene providing, in any combination, data on the following features: predicted protein properties such as mass, pI, cellular location and transmembrane domains; SWISS-PROT annotations; Interpro domains; Gene Ontology records; verified transcription; BLAST matches to the proteomes of A.thaliana, Oryza sativa (rice), Caenorhabditis elegans, Drosophila melanogaster and Homo sapiens. The tool lends itself particularly well to the rapid analysis of contigs or of tens or hundreds of genes identified by high-throughput gene expression experiments. In these cases, a summary table of principal predicted protein features for all genes is given followed by more detailed reports for each individual gene. Centrifuge can also be used for single gene analysis or in a word search mode. AVAILABILITY: http://centrifuge.unil.ch/ CONTACT: edward.farmer@unil.ch.

Arabidopsis↗