Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

The Zebrafish Information Network (ZFIN): the zebrafish model organism database.

The Zebrafish Information Network (ZFIN) is a web based community resource that serves as a centralized location for the curation and integration of zebrafish genetic, genomic and developmental data. ZFIN is publicly accessible at http://zfin.org. ZFIN provides an integrated representation of mutants, genes, genetic markers, mapping panels, publications and community contact data. Recent enhancements to ZFIN include: (i) an anatomical dictionary that provides a controlled vocabulary of anatomical terms, grouped by developmental stages, that may be used to annotate and query gene expression data; (ii) gene expression data; (iii) expanded support for genome sequence; (iv) gene annotation using the standardized vocabulary of Gene Ontology (GO) terms that can be used to elucidate relationships between gene products in zebrafish and other organisms; and (v) collaborations with other databases (NCBI, Sanger Institute and SWISS-PROT) to provide standardization and interconnections based on shared curation.

Animals↗

Analysis of a Bacillus subtilis genome fragment using a co-operative computer system prototype.

Analysis of the huge volume of data generated by large scale sequencing projects requires the construction of new, sophisticated computer systems. These systems should be able to manage the biological data as well as the results of their analysis. They should also help the user to choose the most appropriate methods, and to string them together in order to solve a global analysis task. In this paper we present the prototype of a software system providing an environment for the analysis of large-scale sequence data. As a first step toward this end, this environment has been put to the test within the Bacillus subtilis genome sequencing project. This system integrates both the descriptive knowledge of the entities involved (genes, regulatory signals and the like) and the methodological knowledge comprising an extensible set of analytical methods. A knowledge representation based on two existing object-oriented models is used to implement this integrated system. In addition, the present prototype provides a suitable user interface both for displaying simultaneously the results generated by several methods and for interacting with the objects. We present in this paper the analysis of a B. subtilis genome fragment, present in data libraries but not annotated. Annotation of the genes present in the fragment allowed us to combine the results of several methods used for predicting coding sequences, and to characterize it as comprising a cryptic phage, the skin element. Comparison between the annotation of the skin element and a standard region of the chromosome indicated that local features of the nucleotide sequence could discriminate between phage and non-phage DNA sequence.

Bacillus subtilis↗

Benchmarking ortholog identification methods using functional genomics data.

BACKGROUND: The transfer of functional annotations from model organism proteins to human proteins is one of the main applications of comparative genomics. Various methods are used to analyze cross-species orthologous relationships according to an operational definition of orthology. Often the definition of orthology is incorrectly interpreted as a prediction of proteins that are functionally equivalent across species, while in fact it only defines the existence of a common ancestor for a gene in different species. However, it has been demonstrated that orthologs often reveal significant functional similarity. Therefore, the quality of the orthology prediction is an important factor in the transfer of functional annotations (and other related information). To identify protein pairs with the highest possible functional similarity, it is important to qualify ortholog identification methods. RESULTS: To measure the similarity in function of proteins from different species we used functional genomics data, such as expression data and protein interaction data. We tested several of the most popular ortholog identification methods. In general, we observed a sensitivity/selectivity trade-off: the functional similarity scores per orthologous pair of sequences become higher when the number of proteins included in the ortholog groups decreases. CONCLUSION: By combining the sensitivity and the selectivity into an overall score, we show that the InParanoid program is the best ortholog identification method in terms of identifying functionally equivalent proteins.

Algorithms↗

Multiresolution wavelet analysis for efficient analysis, compression and remote display of long-term physiological signals.

Increased inter-equipment connectivity coupled with advances in Web technology allows ever escalating amounts of physiological data to be produced, far too much to be displayed adequately on a single computer screen. The consequence is that large quantities of insignificant data will be transmitted and reviewed. This carries an increased risk of overlooking vitally important transients. This paper describes a technique to provide an integrated solution based on a single algorithm for the efficient analysis, compression and remote display of long-term physiological signals with infrequent short duration, yet vital events, to effect a reduction in data transmission and display cluttering and to facilitate reliable data interpretation. The algorithm analyses data at the server end and flags significant events. It produces a compressed version of the signal at a lower resolution that can be satisfactorily viewed in a single screen width. This reduced set of data is initially transmitted together with a set of 'flags' indicating where significant events occur. Subsequent transmissions need only involve transmission of flagged data segments of interest at the required resolution. Efficient processing and code protection with decomposition alone is novel. The fixed transmission length method ensures clutter-less display, irrespective of the data length. The flagging of annotated events in arterial oxygen saturation, electroencephalogram and electrocardiogram illustrates the generic property of the algorithm. Data reduction of 87% to 99% and improved displays are demonstrated.

Algorithms↗

GenColors: accelerated comparative analysis and annotation of prokaryotic genomes at various stages of completeness.

SUMMARY: GenColors is a new web-based software/database system aimed at an improved and accelerated annotation of prokaryotic genomes, considering information on related genomes and making extensive use of genome comparison. It offers a seamless integration of data from ongoing sequencing projects and annotated genomic sequences obtained from GenBank. The genome comparison tools determine, for example, best-bidirectional hits, gene conservation, syntenies and gene core sets. Swiss-Prot/TrEMBL hits allow annotations in an effective manner. To further support the annotation base-specific quality data can also be displayed if available. With GenColors dedicated genome browsers containing a group of related genomes can be easily set up and maintained. It has been efficiently used for Borrelia garinii and is currently applied to various ongoing genome projects. AVAILABILITY: Detailed information on GenColors is available at http://gencolors.imb-jena.de. Online usage of GenColors-based genome browsers is the preferred application mode. The system is also available upon request for local installation.

Borrelia↗

TcruziDB: an integrated Trypanosoma cruzi genome resource.

TcruziDB (http://TcruziDB.org) is an integrated genome database for the parasitic organism Trypanosoma cruzi, the causative agent of Chagas' disease. The database currently incorporates all available sequence data (Genomic, BAC, EST) in a single user-friendly location. The database contains a variety of tools specifically designed for searching unannotated draft sequence via BLAST, keyword searches of pre-computed BLAST results, and protein motif searches. Release 1.0 of the database contains nearly 730 million bp of genome sequence from 1.1 million sequence reads generated by the TIGR-Karolinska-SBRI Trypanosoma cruzi Genome Consortium and 15 million bp of clustered EST and genomic sequence obtained from other sources. As annotation, microarray and proteomic data become available, the database will incorporate and integrate these data using the GUS (http://www.gusdb. org) relational framework.

Animals↗

An automated annotation tool for genomic DNA sequences using GeneScan and BLAST.

Genomic sequence data are often available well before the annotated sequence is published. We present a method for analysis of genomic DNA to identify coding sequences using the GeneScan algorithm and characterize these resultant sequences by BLAST. The routines are used to develop a system for automated annotation of genome DNA sequences.

Algorithms↗

The TIGR Rice Genome Annotation Resource: improvements and new features.

In The Institute for Genomic Research Rice Genome Annotation project (http://rice.tigr.org), we have continued to update the rice genome sequence with new data and improve the quality of the annotation. In our current release of annotation (Release 4.0; January 12, 2006), we have identified 42,653 non-transposable element-related genes encoding 49,472 gene models as a result of the detection of alternative splicing. We have refined our identification methods for transposable element-related genes resulting in 13,237 genes that are related to transposable elements. Through incorporation of multiple transcript and proteomic expression data sets, we have been able to annotate 24 799 genes (31,739 gene models), representing approximately 50% of the total gene models, as expressed in the rice genome. All structural and functional annotation is viewable through our Rice Genome Browser which currently supports 59 tracks. Enhanced data access is available through web interfaces, FTP downloads and a Data Extractor tool developed in order to support discrete dataset downloads.

DNA Transposable Elements↗

CardioOp: an integrated approach to teleteaching in cardiac surgery.

INTRODUCTION/PURPOSE: The complexity of cardiac surgery requires continuous training, education and information addressing different individuals: physicians (cardiac surgeons, residents, anaesthesiologists, cardiologists), medical students, perfusionists and patients. Efficacy and efficiency of education and training will likely be improved by the use of multimedia information systems. Nevertheless, computer-based education is facing some serious disadvantages: 1) multimedia productions require tremendous financial and time resources; 2) the obtained multimedia data are only usable for one specific target user group in one specific instructional context; 3) computer based learning programs often show deficiencies in the support of individual learning styles and in providing individual information adjusted to the learner's individual needs. In this paper we describe a computer-system, providing multiple re-use of multimedia-data in different instructional sceneries and providing flexible composition of content to different target user groups. TOOLS AND METHODS: The ZYX document model has been developed, allowing the modelling and flexible on-the-fly composition of multimedia fragments. It has been implemented as a DataBlade module into the object-relational database system Informix Dynamic Server and allows for presentation-neutral storage of multimedia content from the application domain, delivery and presentation of multimedia material, content based retrieval, re-use and composition of multimedia material for different instructional settings. Multimedia data stored in the repository, that can be processed and authored in terms of our identified needs is created by using a next generation authoring environment called CardioOP-Wizard. High-quality intra-operative video is recorded using a video-robot. Difficult surgical procedures are visualized with generic and CT-based 3D-animations. RESULTS: An on-line architecture for multiple re-use and flexible composition of media data has been established. The system contains the following instructional applications (prototypically implemented): a multimedia textbook on operative techniques, an interactive module for problem based-training, a module for creation and presentation of lectures and a module for patient information. Principles of cognitive psychology and knowledge management have been employed in the program. These instructional applications provide information ranging from basic knowledge at the beginner's level, procedural knowledge for the advanced level to implicit knowledge for the professional level. For media-annotation with meta-data a metainformation system, the CardioOP-Clas has been developed. The prototype focuses on aortocoronary bypass grafting and heart transplantation. CONCLUSION: The demonstrated system reflects an integrated approach in terms of information technology and teaching by means of multiple re-use and composition of stored media-items to the individual user and the chosen educational setting on different instructional levels.

Computer Simulation↗

The implementation of telemedicine within a community cancer network.

Telemedicine is being used by physicians at the member hospitals of the Jefferson Cancer Network (JCN) for consultations regarding the diagnosis and management of cancer patients. The technology employed for this telemedicine system was chosen to meet three related specifications: low capital and operating cost, internal maintainability by community hospital data processing staffs, and compatibility with the existing technologic infrastructure. The solution selected is the ubiquitous desktop personal computer and associated software, and Integrated Services Digital Network (ISDN) communications links. The overall performance of this technology has been very satisfactory; ISDN communications has sufficient bandwidth for the transfer of patient data, including text reports, radiographs, and pathology slide images. The presence of the radiologist's interpretation along with the radiographic images allows the presentation of the images on these systems to be acceptable for review purposes. The video frame rates of these systems (12 to 15 frames per second) is adequate, particularly given the "talking heads" nature of the video presentations. Furthermore, the quality of the video image (resolution, size, frame rate) is secondary to the quality of the presentation of the medical information displayed and the capability for mutual annotation of the patient data during the consultation.

Clinical Trials as Topic↗

DNASTAR's Lasergene sequence analysis software.

Lasergene's eight modules provide tools that enable users to accomplish each step of sequence analysis, from trimming and assembly of sequence data, to gene discovery, annotation, gene product analysis, sequence similarity searches, sequence alignment, phylogenetic analysis, oligonucleotide primer design, cloning strategies, and publication of the results. The Lasergene software suite provides the functions and customization tools needed so that users can perform analyses the software writers never imagined.

Base Sequence↗

MGED standards: work in progress.

The Microarray Gene Expression Data (MGED) society is an international organization established in 1999 for facilitating sharing of functional genomics and proteomics array data. To facilitate microarray data sharing, the MGED society has been working in establishing the relevant data standards. The three main components (which will be described in more detail later) of MGED standards are Minimum Information About a Microarray Experiment (MIAME), a document that outlines the minimum information that should be reported about a microarray experiment to enable its unambiguous interpretation and reproduction; MAGE, which consists of three parts, The Microarray Gene Expression Object Model (MAGE-OM), an XML-based document exchange format (MAGE-ML), which is derived directly from the object model, and the supporting tool kit MAGEstk; and MO, or MGED Ontology, which defines sets of common terms and annotation rules for microarray experiments, enabling unambiguous annotation and efficient queries, data analysis and data exchange without loss of meaning. We discuss here how these standards have been established, how they have evolved, and how they are used.

Animals↗

A high-throughput, near-saturating screen for type III effector genes from Pseudomonas syringae.

Pseudomonas syringae strains deliver variable numbers of type III effector proteins into plant cells during infection. These proteins are required for virulence, because strains incapable of delivering them are nonpathogenic. We implemented a whole-genome, high-throughput screen for identifying P. syringae type III effector genes. The screen relied on FACS and an arabinose-inducible hrpL sigma factor to automate the identification and cloning of HrpL-regulated genes. We determined whether candidate genes encode type III effector proteins by creating and testing full-length protein fusions to a reporter called Delta79AvrRpt2 that, when fused to known type III effector proteins, is translocated and elicits a hypersensitive response in leaves of Arabidopsis thaliana expressing the RPS2 plant disease resistance protein. Delta79AvrRpt2 is thus a marker for type III secretion system-dependent translocation, the most critical criterion for defining type III effector proteins. We describe our screen and the collection of type III effector proteins from two pathovars of P. syringae. This stringent functional criteria defined 29 type III proteins from P. syringae pv. tomato, and 19 from P. syringae pv. phaseolicola race 6. Our data provide full functional annotation of the hrpL-dependent type III effector suites from two sequenced P. syringae pathovars and show that type III effector protein suites are highly variable in this pathogen, presumably reflecting the evolutionary selection imposed by the various host plants.

Arabidopsis↗

The Genexpress Index: a resource for gene discovery and the genic map of the human genome.

Detailed analysis of a set of 18,698 sequences derived from both ends of 10,979 human skeletal muscle and brain cDNA clones defined 6676 functional families, characterized by their sequence signatures over 5750 distinct human gene transcripts. About half of these genes have been assigned to specific chromosomes utilizing 2733 eSTS markers, the polymerase chain reaction, and DNA from human-rodent somatic cell hybrids. Sequence and clone clustering and a functional classification together with comprehensive data base searches and annotations made it possible to develop extensive sequence and map cross-indexes, define electronic expression profiles, identify a new set of overlapping genes, and provide numerous new candidate genes for human pathologies.

Amino Acid Sequence↗

Resources and tools for investigating biomolecular networks in mammals.

Molecular databases serve as primary information resources for the analysis of biological networks providing an essential and invaluable treasure for information exploration. Tools for projecting experimental data sets onto known functional information are a major need to support the analysis of samples produced in clinical research. A new concept is the notation of functional modules, i.e. the characterisation of sets of proteins that perform a defined biological function in cooperation. The determination and analysis of functional modules overcome the limitations of the analysis of individual genes and their properties. Although functional modules are not suitable to fully capture systems properties, they have the potential to unify the information generated by different types of experiments. We describe advances related to the problem of integrating heterogeneous data sets into functional modules for mouse and/or human cellular networks based on publicly available data resources, including advances in the design of ontologies for functional classification, problems of automatic protein functional annotation and integration of microarray data.

Algorithms↗

Flnc: Machine Learning Improves the Identification of Novel Long Noncoding RNAs from Stand-Alone RNA-Seq Data.

Long noncoding RNAs (lncRNAs) play critical regulatory roles in human development and disease. Although there are over 100,000 samples with available RNA sequencing (RNA-seq) data, many lncRNAs have yet to be annotated. The conventional approach to identifying novel lncRNAs from RNA-seq data is to find transcripts without coding potential but this approach has a false discovery rate of 30-75%. Other existing methods either identify only multi-exon lncRNAs, missing single-exon lncRNAs, or require transcriptional initiation profiling data (such as H3K4me3 ChIP-seq data), which is unavailable for many samples with RNA-seq data. Because of these limitations, current methods cannot accurately identify novel lncRNAs from existing RNA-seq data. To address this problem, we have developed software, Flnc, to accurately identify both novel and annotated full-length lncRNAs, including single-exon lncRNAs, directly from RNA-seq data without requiring transcriptional initiation profiles. Flnc integrates machine learning models built by incorporating four types of features: transcript length, promoter signature, multiple exons, and genomic location. Flnc achieves state-of-the-art prediction power with an AUROC score over 0.92. Flnc significantly improves the prediction accuracy from less than 50% using the conventional approach to over 85%. Flnc is available via GitHub platform.

RNA-seq↗

A comparative study of cells in inflammation, EAE and MS using biomedical literature data mining.

Biomedical literature and database annotations, available in electronic forms, contain a vast amount of knowledge resulting from global research. Users, attempting to utilize the current state-of-the-art research results are frequently overwhelmed by the volume of such information, making it difficult and time-consuming to locate the relevant knowledge. Literature mining, data mining, and domain specific knowledge integration techniques can be effectively used to provide a user-centric view of the information in a real-world biological problem setting. Bioinformatics tools that are based on real-world problems can provide varying levels of information content, bridging the gap between biomedical and bioinformatics research. We have developed a user-centric bioinformatics research tool, called BioMap, that can provide a customized, adaptive view of the information and knowledge space. BioMap was validated by using inflammatory diseases as a problem domain to identify and elucidate the associations among cells and cellular components involved in multiple sclerosis (MS) and its animal model, experimental allergic encephalomyelitis (EAE). The BioMap system was able to demonstrate the associations between cells directly excavated from biomedical literature for inflammation, EAE and MS. These association graphs followed the scale-free network behavior (average gamma = 2.1) that are commonly found in biological networks.

Animals↗

Diffusion kernel-based logistic regression models for protein function prediction.

Assigning functions to unknown proteins is one of the most important problems in proteomics. Several approaches have used protein-protein interaction data to predict protein functions. We previously developed a Markov random field (MRF) based method to infer a protein's functions using protein-protein interaction data and the functional annotations of its protein interaction partners. In the original model, only direct interactions were considered and each function was considered separately. In this study, we develop a new model which extends direct interactions to all neighboring proteins, and one function to multiple functions. The goal is to understand a protein's function based on information on all the neighboring proteins in the interaction network. We first developed a novel kernel logistic regression (KLR) method based on diffusion kernels for protein interaction networks. The diffusion kernels provide means to incorporate all neighbors of proteins in the network. Second, we identified a set of functions that are highly correlated with the function of interest, referred to as the correlated functions, using the chi-square test. Third, the correlated functions were incorporated into our new KLR model. Fourth, we extended our model by incorporating multiple biological data sources such as protein domains, protein complexes, and gene expressions by converting them into networks. We showed that the KLR approach of incorporating all protein neighbors significantly improved the accuracy of protein function predictions over the MRF model. The incorporation of multiple data sets also improved prediction accuracy. The prediction accuracy is comparable to another protein function classifier based on the support vector machine (SVM), using a diffusion kernel. The advantages of the KLR model include its simplicity as well as its ability to explore the contribution of neighbors to the functions of proteins of interest.

Databases, Protein↗