Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Searching the Porphyromonas gingivalis genome with peptide fragmentation mass spectra.

An approach is described for genomic database searching based on experimentally observed proteolytic fragments, e.g., isolated from 1D or 2D gels or analyzed directly, that can be applied to unfinished prokaryotic genomic data in the absence of annotations or previously assigned open reading frames (ORFs). This variation on the database search is in contrast to the more familiar use of peptide mass spectral fragmentation data to search fully annotated inferred protein databases, e.g., OWL or SWISS-PROT. We compared the SEQUEST search results from a six reading frame translation of the Porphyromonas gingivalis genome DNA sequence with those from computationally derived ORFs created using publicly available genomics software tools. The ORF approach eliminated many of the artifacts present in output from the six reading frame search. The method was applied to uninterpreted tandem mass spectrometric data derived from proteins secreted by the periodontal pathogen Porphyromonas gingivalis in response to the gingival epithelial cell environment, a model system for the study of host-pathogen interactions relevant to human periodontal disease.

Bacterial Proteins↗

The National Microbial Pathogen Database Resource (NMPDR): a genomics platform based on subsystem annotation.

The National Microbial Pathogen Data Resource (NMPDR) (http://www.nmpdr.org) is a National Institute of Allergy and Infections Disease (NIAID)-funded Bioinformatics Resource Center that supports research in selected Category B pathogens. NMPDR contains the complete genomes of approximately 50 strains of pathogenic bacteria that are the focus of our curators, as well as >400 other genomes that provide a broad context for comparative analysis across the three phylogenetic Domains. NMPDR integrates complete, public genomes with expertly curated biological subsystems to provide the most consistent genome annotations. Subsystems are sets of functional roles related by a biologically meaningful organizing principle, which are built over large collections of genomes; they provide researchers with consistent functional assignments in a biologically structured context. Investigators can browse subsystems and reactions to develop accurate reconstructions of the metabolic networks of any sequenced organism. NMPDR provides a comprehensive bioinformatics platform, with tools and viewers for genome analysis. Results of precomputed gene clustering analyses can be retrieved in tabular or graphic format with one-click tools. NMPDR tools include Signature Genes, which finds the set of genes in common or that differentiates two groups of organisms. Essentiality data collated from genome-wide studies have been curated. Drug target identification and high-throughput, in silico, compound screening are in development.

Bacteria↗

The human proteomics initiative (HPI).

The availability of the human genome sequence has enabled the exploration and exploitation of the human genome and proteome to begin. Research has now focussed on the annotation of the genome and in particular of the proteome. With expert annotation extracted from the literature by biologists as the foundation, it has been possible to expand into the areas of data mining and automatic annotation. With further development and integration of pattern recognition methods and the application of alignments clustering, proteome analysis can now be provided in a meaningful way. These various approaches have been integrated to attach, extract and combine as much relevant information as possible to the proteome. This resource should be valuable to users from both research and industry.

Algorithms↗

PHI-base: a new database for pathogen host interactions.

To utilize effectively the growing number of verified genes that mediate an organism's ability to cause disease and/or to trigger host responses, we have developed PHI-base. This is a web-accessible database that currently catalogs 405 experimentally verified pathogenicity, virulence and effector genes from 54 fungal and Oomycete pathogens, of which 176 are from animal pathogens, 227 from plant pathogens and 3 from pathogens with a fungal host. PHI-base is the first on-line resource devoted to the identification and presentation of information on fungal and Oomycete pathogenicity genes and their host interactions. As such, PHI-base is a valuable resource for the discovery of candidate targets in medically and agronomically important fungal and Oomycete pathogens for intervention with synthetic chemistries and natural products. Each entry in PHI-base is curated by domain experts and supported by strong experimental evidence (gene/transcript disruption experiments) as well as literature references in which the experiments are described. Each gene in PHI-base is presented with its nucleotide and deduced amino acid sequence as well as a detailed description of the predicted protein's function during the host infection process. To facilitate data interoperability, we have annotated genes using controlled vocabularies (Gene Ontology terms, Enzyme Commission Numbers and so on), and provide links to other external data sources (e.g. NCBI taxonomy and EMBL). We welcome new data for inclusion in PHI-base, which is freely accessed at www4.rothamsted.bbsrc.ac.uk/phibase/.

Algal Proteins↗

GeneFarm, structural and functional annotation of Arabidopsis gene and protein families by a network of experts.

Genomic projects heavily depend on genome annotations and are limited by the current deficiencies in the published predictions of gene structure and function. It follows that, improved annotation will allow better data mining of genomes, and more secure planning and design of experiments. The purpose of the GeneFarm project is to obtain homogeneous, reliable, documented and traceable annotations for Arabidopsis nuclear genes and gene products, and to enter them into an added-value database. This re-annotation project is being performed exhaustively on every member of each gene family. Performing a family-wide annotation makes the task easier and more efficient than a gene-by-gene approach since many features obtained for one gene can be extrapolated to some or all the other genes of a family. A complete annotation procedure based on the most efficient prediction tools available is being used by 16 partner laboratories, each contributing annotated families from its field of expertise. A database, named GeneFarm, and an associated user-friendly interface to query the annotations have been developed. More than 3000 genes distributed over 300 families have been annotated and are available at http://genoplante-info.infobiogen.fr/Genefarm/. Furthermore, collaboration with the Swiss Institute of Bioinformatics is underway to integrate the GeneFarm data into the protein knowledgebase Swiss-Prot.

Arabidopsis↗

The Sulfolobus database.

The Sulfolobus database (http://www.sulfolobus.org) integrates, for the first time, all currently available Sulfolobus chromosome sequences with annotations. It also includes all the sequence data for the extrachromosomal elements which can propagate in Sulfolobus organisms. All genomes and annotations deposited in GenBank are included in the database and a genefinder has been run on the sequences to ensure that all potential genes are present, and identifiable, in the database. Every month, all genes are searched against a range of external databases and new results are incorporated. The Sulfolobus database was developed as an asset to the rapidly-growing international community working with Sulfolobus as a model organism for the kingdom Crenarchaeota of the Archaea. It was accessed more that 46 000 times in its first year. The database aims to provide researchers easy access to sequence and gene information and the web-interface includes various searches, free text and BLAST, as well as genome browsing and data extraction. Updated annotations are incorporated regularly and the database will continue to expand as new information becomes available. This includes new sequences, newly identified genes, annotations and other related information.

Chromosomes, Archaeal↗

A comprehensive transcript index of the human genome generated using microarrays and computational approaches.

BACKGROUND: Computational and microarray-based experimental approaches were used to generate a comprehensive transcript index for the human genome. Oligonucleotide probes designed from approximately 50,000 known and predicted transcript sequences from the human genome were used to survey transcription from a diverse set of 60 tissues and cell lines using ink-jet microarrays. Further, expression activity over at least six conditions was more generally assessed using genomic tiling arrays consisting of probes tiled through a repeat-masked version of the genomic sequence making up chromosomes 20 and 22. RESULTS: The combination of microarray data with extensive genome annotations resulted in a set of 28,456 experimentally supported transcripts. This set of high-confidence transcripts represents the first experimentally driven annotation of the human genome. In addition, the results from genomic tiling suggest that a large amount of transcription exists outside of annotated regions of the genome and serves as an example of how this activity could be measured on a genome-wide scale. CONCLUSIONS: These data represent one of the most comprehensive assessments of transcriptional activity in the human genome and provide an atlas of human gene expression over a unique set of gene predictions. Before the annotation of the human genome is considered complete, however, the previously unannotated transcriptional activity throughout the genome must be fully characterized.

Chromosomes, Human, Pair 20↗

NCBI GEO: mining tens of millions of expression profiles--database and tools update.

The Gene Expression Omnibus (GEO) repository at the National Center for Biotechnology Information (NCBI) archives and freely disseminates microarray and other forms of high-throughput data generated by the scientific community. The database has a minimum information about a microarray experiment (MIAME)-compliant infrastructure that captures fully annotated raw and processed data. Several data deposit options and formats are supported, including web forms, spreadsheets, XML and Simple Omnibus Format in Text (SOFT). In addition to data storage, a collection of user-friendly web-based interfaces and applications are available to help users effectively explore, visualize and download the thousands of experiments and tens of millions of gene expression patterns stored in GEO. This paper provides a summary of the GEO database structure and user facilities, and describes recent enhancements to database design, performance, submission format options, data query and retrieval utilities. GEO is accessible at http://www.ncbi.nlm.nih.gov/geo/

Computer Graphics↗

A gene expression map for the euchromatic genome of Drosophila melanogaster.

We used a maskless photolithography method to produce DNA oligonucleotide microarrays with unique probe sequences tiled throughout the genome of Drosophila melanogaster and across predicted splice junctions. RNA expression of protein coding and nonprotein coding sequences was determined for each major stage of the life cycle, including adult males and females. We detected transcriptional activity for 93% of annotated genes and RNA expression for 41% of the probes in intronic and intergenic sequences. Comparison to genome-wide RNA interference data and to gene annotations revealed distinguishable levels of expression for different classes of genes and higher levels of expression for genes with essential cellular functions. Differential splicing was observed in about 40% of predicted genes, and 5440 previously unknown splice forms were detected. Genes within conserved regions of synteny with D. pseudoobscura had highly correlated expression; these regions ranged in length from 10 to 900 kilobase pairs. The expressed intergenic and intronic sequences are more likely to be evolutionarily conserved than nonexpressed ones, and about 15% of them appear to be developmentally regulated. Our results provide a draft expression map for the entire nonrepetitive genome, which reveals a much more extensive and diverse set of expressed sequences than was previously predicted.

Algorithms↗

Provenance and annotation for visual exploration systems.

Exploring data using visualization systems has been shown to be an extremely powerful technique. However, one of the challenges with such systems is an inability to completely support the knowledge discovery process. More than simply looking at data, users will make a semipermanent record of their visualizations by printing out a hard copy. Subsequently, users will mark and annotate these static representations, either for dissemination purposes or to augment their personal memory of what was witnessed. In this paper, we present a model for recording the history of user explorations in visualization environments, augmented with the capability for users to annotate their explorations. A prototype system is used to demonstrate how this provenance information can be recalled and shared. The prototype system generates interactive visualizations of the provenance data using a spatio-temporal technique. Beyond the technical details of our model and prototype, results from a controlled experiment that explores how different history mechanisms impact problem solving in visualization environments are presented.

Algorithms↗

GeneExt: a gene model extension tool for enhanced single-cell RNA-seq analysis.

MOTIVATION: Incomplete gene models negatively impact single-cell gene expression quantification. This is particularly true in non-model species where often gene 3' ends are inaccurately annotated, while most scRNA-seq methods only capture the 3' transcript region. This results in many genes being incorrectly quantified or not detected. RESULTS: GeneExt leverages scRNA-seq data to refine gene annotations. We exemplify GeneExt usage and its impact on the gene expression quantification of eight non-model organism single-cell atlases. By extending and homogenizing gene annotations, our tool will help improve biological interpretation and cross-species comparisons of cell type expression atlases. AVAILABILITY: GeneExt is available at https://github.com/sebepedroslab/GeneExt (DOI: https://doi.org/10.5281/zenodo.18712940) under a GNU General Public license, together with test data and usage instructions.

Software↗

Blast2GO goes grid: developing a grid-enabled prototype for functional genomics analysis.

The vast amount in complexity of data generated in Genomic Research implies that new dedicated and powerful computational tools need to be developed to meet their analysis requirements. Blast2GO (B2G) is a bioinformatics tool for Gene Ontology-based DNA or protein sequence annotation and function-based data mining. The application has been developed with the aim of affering an easy-to-use tool for functional genomics research. Typical B2G users are middle size genomics labs carrying out sequencing, ETS and microarray projects, handling datasets up to several thousand sequences. In the current version of B2G. The power and analytical potential of both annotation and function data-mining is somehow restricted to the computational power behind each particular installation. In order to be able to offer the possibility of an enhanced computational capacity within this bioinformatics application, a Grid component is being developed. A prototype has been conceived for the particular problem of speeding up the Blast searches to obtain fast results for large datasets. Many efforts have been done in the literature concerning the speeding up of Blast searches, but few of them deal with the use of large heterogeneous production Grid Infrastructures. These are the infrastructures that could reach the largest number of resources and the best load balancing for data access. The Grid Service under development will analyse requests based on the number of sequences, splitting them accordingly to the available resources. Lower-level computation will be performed through MPIBLAST. The software architecture is based on the WSRF standard.

Computational Biology↗

Protein molecular function prediction by Bayesian phylogenomics.

We present a statistical graphical model to infer specific molecular function for unannotated protein sequences using homology. Based on phylogenomic principles, SIFTER (Statistical Inference of Function Through Evolutionary Relationships) accurately predicts molecular function for members of a protein family given a reconciled phylogeny and available function annotations, even when the data are sparse or noisy. Our method produced specific and consistent molecular function predictions across 100 Pfam families in comparison to the Gene Ontology annotation database, BLAST, GOtcha, and Orthostrapper. We performed a more detailed exploration of functional predictions on the adenosine-5'-monophosphate/adenosine deaminase family and the lactate/malate dehydrogenase family, in the former case comparing the predictions against a gold standard set of published functional characterizations. Given function annotations for 3% of the proteins in the deaminase family, SIFTER achieves 96% accuracy in predicting molecular function for experimentally characterized proteins as reported in the literature. The accuracy of SIFTER on this dataset is a significant improvement over other currently available methods such as BLAST (75%), GeneQuiz (64%), GOtcha (89%), and Orthostrapper (11%). We also experimentally characterized the adenosine deaminase from Plasmodium falciparum, confirming SIFTER's prediction. The results illustrate the predictive power of exploiting a statistical model of function evolution in phylogenomic problems. A software implementation of SIFTER is available from the authors.

Adenosine Deaminase↗

[Pulmonary embolism in medical inpatients. An approach to trends and costs in Spain].

BACKGROUND: Tromboprophylaxis for hopitalized patients has became a widespread practice in the lasts years. However, the trends in pulmonary embolism (EP) are insufficiently known. We analyzed the trends of EP during the 1997-1999 period in Spain at the hospitals of the National Health System (NHS) and theirs costs were also estimated. METHODS: Diagnosis-Related Group (DGR) 78 from the data of the national hospital discharge register was evaluated. Annual trends, age, sex and (for 1999) mortality, some comorbidities and costs according official data from the Spanish Ministry of Health were annotated. Mostly medical patients were included. Data were compared with those from the Servicio Gallego de Saúde (SERGAS) which maintain a high rate of hospitalary discharges declaration. RESULTS: During the 3-year period 14.021 cases of EP were observed. The 47% of the NHS cases were older than 75 and mortality was 6,8%. An annual increment of hospitalized cases was observed at SERGAS and NHS (in the last 5% at the end of the period). In 1999 costs were estimated in 16-20,2 millions euro. That was 0,097-0,12% of global hospitalary budget. CONCLUSIONS: Trends for hospitalizations for EP are increasing in Spain. It remains in doubt if a true increase in incidence or the improving notification and awareness are the responsibles for this increment. However, the data of SERGAS support the first possibility.

Diagnosis-Related Groups↗

Dynamic integration of gene annotation and its application to microarray analysis.

Comprehensive and structured annotations for all genes on a microarray chip are essential for the interpretation of its expression data. Currently, most chip gene annotations are one-line free text descriptions that are often partial, outdated and unsuitable for large-scale data analysis. Therefore the interpretation of microarray gene expression clusters is often limited. Although researchers can manually navigate a collection of databases for better annotations, it is only practical for limited number of genes. Existing meta-databases fail to provide comprehensive categorized annotations for hundreds of genes simultaneously. We have developed an automatic system to address this issue. GeneView system monitors various data sources, extracts gene information from a source whenever it is updated, comprehensively matches genes, and integrates them into a central database by categories, such as pathway, genetic mapping, phenotype, expression profile, domain structure, protein interaction, disease association, and references. The system consists of four major components: (1) relational database; (2) data processing; (3) user curation; (4) data query. We evaluated it by analyzing genes on cDNA and Affymetrix Oligo chips. In both cases, the system provided more accurate and comprehensive information than those provided by the vendors or the chip users, and helped identify new common functions among genes in the same expression clusters.

Cluster Analysis↗

Creating the gene ontology resource: design and implementation.

The exponential growth in the volume of accessible biological information has generated a confusion of voices surrounding the annotation of molecular information about genes and their products. The Gene Ontology (GO) project seeks to provide a set of structured vocabularies for specific biological domains that can be used to describe gene products in any organism. This work includes building three extensive ontologies to describe molecular function, biological process, and cellular component, and providing a community database resource that supports the use of these ontologies. The GO Consortium was initiated by scientists associated with three model organism databases: SGD, the Saccharomyces Genome database; FlyBase, the Drosophila genome database; and MGD/GXD, the Mouse Genome Informatics databases. Additional model organism database groups are joining the project. Each of these model organism information systems is annotating genes and gene products using GO vocabulary terms and incorporating these annotations into their respective model organism databases. Each database contributes its annotation files to a shared GO data resource accessible to the public at http://www.geneontology.org/. The GO site can be used by the community both to recover the GO vocabularies and to access the annotated gene product data sets from the model organism databases. The GO Consortium supports the development of the GO database resource and provides tools enabling curators and researchers to query and manipulate the vocabularies. We believe that the shared development of this molecular annotation resource will contribute to the unification of biological information.

Animals↗

MPromDb: an integrated resource for annotation and visualization of mammalian gene promoters and ChIP-chip experimental data.

We have developed Mammalian Promoter Database (MPromDb), a novel database that integrates gene promoters with experimentally supported annotation of transcription start sites, cis-regulatory elements, CpG islands and chromatin immunoprecipitation microarray (ChIP-chip) experimental results with intuitively designed presentation. Release 1.0 of MPromDb currently contains 36,407 promoters and first exons (19,170 from human, 15,953 from mouse and 1284 from rat), 3739 transcription factor (TF)-binding sites (2027 from human, 1181 mouse and 531 rat) and 224 TFs with links to PubMed and GenBank references. Target promoters of TFs that have been identified by ChIP-chip assay are integrated into the database. MPromDb serves as a portal for genome-wide promoter analysis of data generated by ChIP-chip experimental studies. MPromDb can be accessed from http://bioinformatics.med.ohio-state.edu/MPromDb.

Animals↗

Using the transcriptome to annotate the genome revisited: application of massively parallel signature sequencing (MPSS).

Transcriptome analysis can provide useful data for refining genome sequence annotation. Application of massively parallel signature sequencing (MPSS) revealed reproducible transcription, in multiple MPSS cycles, from 73% of computationally predicted genes in the Theileria parva schizont lifecycle stage. Signatures spanning consecutive exons confirmed 142 predicted introns. MPSS identified 83 putative genes, >100 codons overlooked by annotation software, and 139 potentially incorrect gene models (with either truncated ORFs or overlooked exons) by interfacing signature locations with stop codon maps. Twenty representative models were confirmed as likely to be incorrect using reverse transcription PCR amplification from independent schizont cDNA preparations. More than 50% of the 60 putative single copy genes in T. parva that were absent from the genome of the closely related T. annulata had MPSS signatures. This study illustrates the utility of MPSS for improving annotation of small, gene-rich microbial eukaryotic genomes.

Animals↗