Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

SCORPION, a molecular database of scorpion toxins.

Increasing interest in the studies of toxins and the requirements for better structural and functional annotations have created a need for improved data management in the field of toxins. The molecular database, SCORPION, contains more than 200 entries of fully referenced scorpion toxin data including primary sequences, three-dimensional structures, structural and functional annotations of scorpion toxins along with relevant literature references. SCORPION has a set of search tools that allow users to extract data and perform specific queries. These entries have been compiled from public databases and literature, cleaned of errors and enriched with additional structural and functional information. The grouping of scorpion toxins provides a basis for extending and clarifying the existing structural and functional classifications. The bioinformatics modules in SCORPION facilitate analyses aimed at classification of scorpion toxins and identification of sequence patterns associated with specific structural or functional properties of scorpion toxins. The SCORPION database is accessible via the Internet at sdmc.krdl.org.sg:8080/scorpion.

Animals↗

MeMo: a hybrid SQL/XML approach to metabolomic data management for functional genomics.

BACKGROUND: The genome sequencing projects have shown our limited knowledge regarding gene function, e.g. S. cerevisiae has 5-6,000 genes of which nearly 1,000 have an uncertain function. Their gross influence on the behaviour of the cell can be observed using large-scale metabolomic studies. The metabolomic data produced need to be structured and annotated in a machine-usable form to facilitate the exploration of the hidden links between the genes and their functions. DESCRIPTION: MeMo is a formal model for representing metabolomic data and the associated metadata. Two predominant platforms (SQL and XML) are used to encode the model. MeMo has been implemented as a relational database using a hybrid approach combining the advantages of the two technologies. It represents a practical solution for handling the sheer volume and complexity of the metabolomic data effectively and efficiently. The MeMo model and the associated software are available at http://dbkgroup.org/memo/. CONCLUSION: The maturity of relational database technology is used to support efficient data processing. The scalability and self-descriptiveness of XML are used to simplify the relational schema and facilitate the extensibility of the model necessitated by the creation of new experimental techniques. Special consideration is given to data integration issues as part of the systems biology agenda. MeMo has been physically integrated and cross-linked to related metabolomic and genomic databases. Semantic integration with other relevant databases has been supported through ontological annotation. Compatibility with other data formats is supported by automatic conversion.

Computer Simulation↗

Protein family classification and functional annotation.

With the accelerated accumulation of genomic sequence data, there is a pressing need to develop computational methods and advanced bioinformatics infrastructure for reliable and large-scale protein annotation and biological knowledge discovery. The Protein Information Resource (PIR) provides an integrated public resource of protein informatics to support genomic and proteomic research. PIR produces the Protein Sequence Database of functionally annotated protein sequences. The annotation problems are addressed by a classification-driven and rule-based method with evidence attribution, coupled with an integrated knowledge base system being developed. The approach allows sensitive identification, consistent and rich annotation, and systematic detection of annotation errors, as well as distinction of experimentally verified and computationally predicted features. The knowledge base consists of two new databases, sequence analysis tools, and graphical interfaces. PIR-NREF, a non-redundant reference database, provides a timely and comprehensive collection of all protein sequences, totaling more than 1,000,000 entries. iProClass, an integrated database of protein family, function, and structure information, provides extensive value-added features for about 830,000 proteins with rich links to over 50 molecular databases. This paper describes our approach to protein functional annotation with case studies and examines common identification errors. It also illustrates that data integration in PIR supports exploration of protein relationships and may reveal protein functional associations beyond sequence homology.

Amino Acid Motifs↗

BioViews: Java-based tools for genomic data visualization.

Visualization tools for bioinformatics ideally should provide universal access to the most current data in an interactive and intuitive graphical user interface. Since the introduction of Java, a language designed for distributed programming over the Web, the technology now exists to build a genomic data visualization tool that meets these requirements. Using Java we have developed a prototype genome browser applet (BioViews) that incorporates a three-level graphical view of genomic data: a physical map, an annotated sequence map, and a DNA sequence display. Annotated biological features are displayed on the physical and sequence-based maps, and the different views are interconnected. The applet is linked to several databases and can retrieve features and display hyperlinked textual data on selected features. In addition to browsing genomic data, different types of analyses can be performed interactively and the results of these analyses visualized alongside prior annotations. Our genome browser is built on top of extensible, reusable graphic components specifically designed for bioinformatics. Other groups can (and do) reuse this work in various ways. Genome centers can reuse large parts of the genome browser with minor modifications, bioinformatics groups working on sequence analysis can reuse components to build front ends for analysis programs, and biology laboratories can reuse components to publish results as dynamic Web documents.

Animals↗

GeneViTo: visualizing gene-product functional and structural features in genomic datasets.

BACKGROUND: The availability of increasing amounts of sequence data from completely sequenced genomes boosts the development of new computational methods for automated genome annotation and comparative genomics. Therefore, there is a need for tools that facilitate the visualization of raw data and results produced by bioinformatics analysis, providing new means for interactive genome exploration. Visual inspection can be used as a basis to assess the quality of various analysis algorithms and to aid in-depth genomic studies. RESULTS: GeneViTo is a JAVA-based computer application that serves as a workbench for genome-wide analysis through visual interaction. The application deals with various experimental information concerning both DNA and protein sequences (derived from public sequence databases or proprietary data sources) and meta-data obtained by various prediction algorithms, classification schemes or user-defined features. Interaction with a Graphical User Interface (GUI) allows easy extraction of genomic and proteomic data referring to the sequence itself, sequence features, or general structural and functional features. Emphasis is laid on the potential comparison between annotation and prediction data in order to offer a supplement to the provided information, especially in cases of "poor" annotation, or an evaluation of available predictions. Moreover, desired information can be output in high quality JPEG image files for further elaboration and scientific use. A compilation of properly formatted GeneViTo input data for demonstration is available to interested readers for two completely sequenced prokaryotes, Chlamydia trachomatis and Methanococcus jannaschii. CONCLUSIONS: GeneViTo offers an inspectional view of genomic functional elements, concerning data stemming both from database annotation and analysis tools for an overall analysis of existing genomes. The application is compatible with Linux or Windows ME-2000-XP operating systems, provided that the appropriate Java Runtime Environment is already installed in the system.

Bacterial Proton-Translocating ATPases↗

Gencube: centralized retrieval and integration of multi-omics resources from leading databases.

MOTIVATION: The volume of multi-omics data for diverse species is growing at an unprecedented rate, with new genome assemblies, related annotations, and high-throughput sequencing resources being submitted daily to various genomic data repositories. In response to this data influx, both existing and new databases are establishing optimized hierarchical structures to manage the vast amount of information. However, the lack of accessible command-line tools, combined with the functional limitations and unintuitive design of existing options, presents significant challenges for researchers. This gap underscores a critical need for a tool that enables streamlined retrieval and integration of omics data across these diverse repositories. RESULTS: We have developed Gencube, a command-line tool that enables centralized retrieval and integration of a comprehensive set of six different data types-genome assemblies, gene sets, annotations, sequences, comparative genomic data, and NGS-based omics resources-from various leading databases. AVAILABILITY AND IMPLEMENTATION: Gencube is a free and open-source tool, with its code available on GitHub: https://github.com/snu-cdrc/gencube and also archived on Zenodo: https://doi.org/10.5281/zenodo.14607649.

Databases, Genetic↗

ColorHOR--novel graphical algorithm for fast scan of alpha satellite higher-order repeats and HOR annotation for GenBank sequence of human genome.

MOTIVATION: GenBank data are at present lacking alpha satellite higher-order repeat (HOR) annotation. Furthermore, exact HOR consensus lengths have not been reported so far. Given the fast growth of sequence databases in the centromeric region, it is of increasing interest to have efficient tools for computational identification and analysis of HORs from known sequences. RESULTS: We develop a graphical user interface method, ColorHOR, for fast computational identification of HORs in a given genomic sequence, without requiring a priori information on the composition of the genomic sequence. ColorHOR is based on an extension of the key-string algorithm and provides a color representation of the order and orientation of HORs. For the key string, we use a robust 6 bp string from a consensus alpha satellite and its representative nature is tested. ColorHOR algorithm provides a direct visual identification of HORs (direct and/or reverse complement). In more detail, we first illustrate the ColorHOR results for human chromosome 1. Using ColorHOR we determine for the first time the HOR annotation of the GenBank sequence of the whole human genome. In addition to some HORs, corresponding to those determined previously biochemically, we find new HORs in chromosomes 4, 8, 9, 10, 11 and 19. For the first time, we determine exact consensus lengths of HORs in 10 chromosomes. We propose that the HOR assignment obtained by using ColorHOR be included into the GenBank database.

Algorithms↗

Avoiding inconsistencies over time and tracking difficulties in Applied Biosystems AB1700/Panther probe-to-gene annotations.

BACKGROUND: Significant inconsistencies between probe-to-gene annotations between different releases of probe set identifiers by commercial microarray platform solutions have been reported. Such inconsistencies lead to misleading or ambiguous interpretation of published gene expression results. RESULTS: We report here similar inconsistencies in the probe-to-gene annotation of Applied Biosystems AB1700 data, demonstrating that this is not an isolated concern. Moreover, the online information source PANTHER does not provide information required to track such inconsistencies, hence, even correctly annotated datasets, when resubmitted after PANTHER was updated to a new probe-to-gene annotation release, will generate differing results without any feedback on the origin of the change. CONCLUSION: The importance of unequivocal annotation of microarray experiments can not be underestimated. Inconsistencies greatly diminish the usefulness of the technology. Novel methods in the analysis of transcriptome profiles often rely on large disparate datasets stemming from multiple sources. The predictive and analytic power of such approaches rapidly diminishes if only least-common subsets can be used for analysis. We present here the information that needs to be provided together with the raw AB1700 data, and the information required together with the biologic interpretation of such data to avoid inconsistencies and tracking difficulties.

Algorithms↗

Annotation of parasite genomes.

Genome annotation is the application of useful biological descriptions to sequence data. Different levels of time and effort can be invested to produce correspondingly different depths of annotation depending on what methods are employed. Researchers using genome data should, therefore, understand how annotations are generated to assess their validity correctly and to determine what level of inferences can be made accordingly. Thorough annotation requires a large range of procedures, most of which involve manual reviews of all available evidence. First, gene structures often are computed algorithmically and edited based on in-depth analyses of the underlying sequence data. Second, functional predictions draw on data from various sources. Finally, the use of structured and controlled descriptions, such as those provided by gene ontology, can be used so that final descriptions are not only consistent and unambiguous, but capable of being used in further downstream analyses such as cross-species comparisons.

Algorithms↗

Frequency, type, distribution and annotation of simple sequence repeats in Rosaceae ESTs.

Genomic resources for peach, a model species for Rosaceae, are being developed to accelerate gene discovery in other Rosaceae species by comparative mapping. Simple sequence repeats (SSRs) are an important tool for comparative mapping because of their high polymorphism and transportability. To accelerate the development of SSR markers, we analyzed publicly available Rosaceae expressed sequence tags (ESTs) for SSRs. A total of 17,284 ESTs from almond, peach and rose were assembled into putatively non-redundant EST sets. For comparison, 179,099 ESTs from Arabidopsis were also used in the analysis. About 4% of the assembled ESTs contained SSRs in Rosaceae, which was higher than the 2.4% found in Arabidopsis. About half of the SSRs were found in the putative UTR, and the estimated average distance between SSRs in the UTR was 5.5 kb in rose, 5.1 kb in almond, 7 kb in peach and 13 kb in Arabidopsis. In the putative coding region, the estimated average distance was two to four times longer than in the UTR. Rosaceae ESTs containing SSRs were functionally annotated using the GenBank nr database and further classified using the gene ontology terms associated with the matching sequences in the SwissProt database. The detailed data including the sequences and annotation results are available from http://www.genome.clemson.edu/gdr/rosaceaessr/.

Arabidopsis↗

OmicsQ: a user-friendly platform for interactive quantitative omics data analysis.

MOTIVATION: High-throughput omics technologies generate complex datasets with thousands of features that are quantified across multiple experimental conditions, but often suffer from incomplete measurements, missing values, and individually fluctuating variances. This requires analytical tools for accurate, deep and insightful biological interpretation, capable of dealing with a large variety of data properties and different amounts of completeness. Software capable of handling such data complexity and integrating with external applications for downstream analysis remains rare and mostly relies on programming-based environments, limiting accessibility for researchers without computational expertise. RESULTS: We present OmicsQ, an interactive, web-based platform designed to streamline quantitative omics data analysis. OmicsQ provides an intuitive, browser-based visualization interface that integrates established statistical processing tools. Those include robust batch correction, automated experimental design annotation, and handling of missing data without imputation, which maintains data integrity and avoids artifacts from a priori assumptions. OmicsQ seamlessly interacts with external applications (e.g. PolySTest, VSClust, ComplexBrowser) for statistical testing, clustering, analysis of protein complex behavior, and pathway enrichment, offering a comprehensive and flexible workflow from data import to biological interpretation that is broadly applicable across domains. AVAILABILITY AND IMPLEMENTATION: OmicsQ is implemented in R and Shiny and is available at https://computproteomics.bmb.sdu.dk/app_direct/OmicsQ. Source code and installation instructions: https://github.com/computproteomics/OmicsQ, DOI: 10.5281/zenodo.17778420.

Software↗

REGANOR: a gene prediction server for prokaryotic genomes and a database of high quality gene predictions for prokaryotes.

UNLABELLED: With >1,000 prokaryotic genome sequencing projects ongoing or already finished, comprehensive comparative analysis of the gene content of these genomes has become viable. To allow for a meaningful comparative analysis, gene prediction of the various genomes should be as accurate as possible. It is clear that improving the state of genome annotation requires automated gene identification methods to cope with the influence of artifacts, such as genomic GC content. There is currently still room for improvement in the state of annotations. We present a web server and a database of high-quality gene predictions. The web server is a resource for gene identification in prokaryote genome sequences. It implements our previously described, accurate gene finding method REGANOR. We also provide novel gene predictions for 241 complete, or almost complete, prokaryotic genomes. We demonstrate how this resource can easily be utilised to identify promising candidates for currently missing genes from genome annotations with several examples. All data sets are available online. AVAILABILITY: The gene finding server is accessible via https://www.cebitec.uni-bielefeld.de/groups/brf/software/reganor/cgi-bin/reganor_upload.cgi. The server software is available with the GenDB genome annotation system (version 2.2.1 onwards) under the GNU general public license. The software can be downloaded from https://sourceforge.net/projects/gendb/. More information on installing GenDB and REGANOR and the system requirements can be found on the GenDB project page http://www.cebitec.uni-bielefeld.de/groups/brf/software/wiki/GenDBWiki/AdministratorDocumentation/GenDBInstallation

Chromosome Mapping↗

Toward automated biochemotype annotation for large compound libraries.

Combinatorial chemistry allows scientists to probe large synthetically accessible chemical space. However, identifying the sub-space which is selectively associated with an interested biological target, is crucial to drug discovery and life sciences. This paper describes a process to automatically annotate biochemotypes of compounds in a library and thus to identify bioactivity related chemotypes (biochemotypes) from a large library of compounds. The process consists of two steps: (1) predicting all possible bioactivities for each compound in a library, and (2) deriving possible biochemotypes based on predictions. The Prediction of Activity Spectra for Substances program (PASS) was used in the first step. In second step, structural similarity and scaffold-hopping technologies are employed. These technologies are used to derive biochemotypes from bioactivity predictions and the corresponding annotated biochemotypes from MDL Drug Data Report (MDDR) database. About a one million (982,889) commercially available compound library (CACL) has been tested using this process. This paper demonstrates the feasibility of automatically annotating biochemotypes for large libraries of compounds. Nevertheless, some issues need to be considered in order to improve the process. First, the prediction accuracy of PASS program has no significant correlation with the number of compounds in a training set. Larger training sets do not necessarily increase the maximal error of prediction (MEP), nor do they increase the hit structural diversity. Smaller training sets do not necessarily decrease MEP, nor do they decrease the hit structural diversity. Second, the success of systematic bioactivity prediction relies on modeling, training data, and the definition of bioactivities (biochemotype ontology). Unfortunately, the biochemotype ontology was not well developed in the PASS program. Consequently, "ill-defined" bioactivities can reduce the quality of predictions. This paper suggests the ways in which the systematic bioactivities prediction program should be improved.

Chemistry, Pharmaceutical↗

pdbFun: mass selection and fast comparison of annotated PDB residues.

pdbFun (http://pdbfun.uniroma2.it) is a web server for structural and functional analysis of proteins at the residue level. pdbFun gives fast access to the whole Protein Data Bank (PDB) organized as a database of annotated residues. The available data (features) range from solvent exposure to ligand binding ability, location in a protein cavity, secondary structure, residue type, sequence functional pattern, protein domain and catalytic activity. Users can select any residue subset (even including any number of PDB structures) by combining the available features. Selections can be used as probe and target in multiple structure comparison searches. For example a search could involve, as a query, all solvent-exposed, hydrophylic residues that are not in alpha-helices and are involved in nucleotide binding. Possible examples of targets are represented by another selection, a single structure or a dataset composed of many structures. The output is a list of aligned structural matches offered in tabular and also graphical format.

Algorithms↗

PFD: a database for the investigation of protein folding kinetics and stability.

We have developed a new database that collects all protein folding data into a single, easily accessible public resource. The Protein Folding Database (PFD) contains annotated structural, methodological, kinetic and thermodynamic data for more than 50 proteins, from 39 families. A user-friendly web interface has been developed that allows powerful searching, browsing and information retrieval, whilst providing links to other protein databases. The database structure allows visualization of folding data in a useful and novel way, with a long-term aim of facilitating data mining and bioinformatics approaches. PFD can be accessed freely at http://pfd.med.monash.edu.au.

Databases, Protein↗

CGHScan: finding variable regions using high-density microarray comparative genomic hybridization data.

BACKGROUND: Comparative genomic hybridization can rapidly identify chromosomal regions that vary between organisms and tissues. This technique has been applied to detecting differences between normal and cancerous tissues in eukaryotes as well as genomic variability in microbial strains and species. The density of oligonucleotide probes available on current microarray platforms is particularly well-suited for comparisons of organisms with smaller genomes like bacteria and yeast where an entire genome can be assayed on a single microarray with high resolution. Available methods for analyzing these experiments typically confine analyses to data from pre-defined annotated genome features, such as entire genes. Many of these methods are ill suited for datasets with the number of measurements typical of high-density microarrays. RESULTS: We present an algorithm for analyzing microarray hybridization data to aid identification of regions that vary between an unsequenced genome and a sequenced reference genome. The program, CGHScan, uses an iterative random walk approach integrating multi-layered significance testing to detect these regions from comparative genomic hybridization data. The algorithm tolerates a high level of noise in measurements of individual probe intensities and is relatively insensitive to the choice of method for normalizing probe intensity values and identifying probes that differ between samples. When applied to comparative genomic hybridization data from a published experiment, CGHScan identified eight of nine known deletions in a Brucella ovis strain as compared to Brucella melitensis. The same result was obtained using two different normalization methods and two different scores to classify data for individual probes as representing conserved or variable genomic regions. The undetected region is a small (58 base pair) deletion that is below the resolution of CGHScan given the array design employed in the study. CONCLUSION: CGHScan is an effective tool for analyzing comparative genomic hybridization data from high-density microarrays. The algorithm is capable of accurately identifying known variable regions and is tolerant of high noise and varying methods of data preprocessing. Statistical analysis is used to define each variable region providing a robust and reliable method for rapid identification of genomic differences independent of annotated gene boundaries.

Algorithms↗

Sequence- and structure-based protein function prediction from genomic information.

Existing functional annotation transfer is fraught with inaccuracies that may hinder forward interpretation and mining of genomic data. Hand-curation of the annotation placed into databases is not practical. In lieu of experimental evidence, computational biological approaches offer high-throughput tools to predict function accurately; however, these methods are still notably deficient in defining and describing the complexity of protein function. Enriching genomic sequences obtained from sequencing efforts and expression array methods with protein function information and classification will be an efficient first step for incorporating genomic data into drug discovery programs.

Computational Biology↗

MOsDB: an integrated information resource for rice genomics.

The MIPS Rice (Oryza sativa) database (MOsDB; http://mips.gsf.de/proj/rice) provides a comprehensive data collection dedicated to the genome information of rice. Rice (O. sativa L.) is one of the most important food crops for over half the world's population and serves as a major model system in cereal genome research. MOsDB integrates data from two publicly available rice genomic sequences, O. sativa L. ssp. indica and O. sativa L. ssp. japonica. Besides regularly updated rice genome sequence information, MOsDB provides an integrated resource for associated analysis data, e.g. internal and external annotation information as well as a complex characterization of all annotated rice genes. The MOsDB web interface supports various search options and allows browsing the database content. MOsDB is continuously expanding to include an increasing range of data type and the growing amount of information on the rice genome.

DNA Transposable Elements↗