Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

IMGT databases, web resources and tools for immunoglobulin and T cell receptor sequence analysis, http://imgt.cines.fr.

IMGT, the international ImMunoGeneTics database((R)) (http://imgt.cines.fr), is a high-quality integrated information system specializing in immunoglobulins (IG), T cell receptors (TR) and major histocompatibility complex (MHC) of human and other vertebrates, created in 1989, by LIGM, at the Université Montpellier II, CNRS, Montpellier, France. IMGT provides a common access to standardized data which include nucleotide and protein sequences, oligonucleotide primers, gene maps, genetic polymorphisms, specificities, 2D and 3D structures. IMGT includes several databases (IMGT/LIGM-DB, IMGT/3Dstructure-DB, IMGT/HLA-DB), Web resources ('IMGT Marie-Paule page') and interactive tools (IMGT/V-QUEST, IMGT/JunctionAnalysis). IMGT expertly annotated data and tools described in this paper are particularly useful for the analysis of the IG and TR rearrangements in leukemia, lymphoma and myeloma, and in translocations involving the antigen receptor loci. IMGT is freely available at http://imgt.cines.fr.

Amino Acid Sequence↗

SPD--a web-based secreted protein database.

With the improved secreted protein prediction approach and comprehensive data sources, including Swiss-Prot, TrEMBL, RefSeq, Ensembl and CBI-Gene, we have constructed secretomes of human, mouse and rat, with a total of 18 152 secreted proteins. All the entries are ranked according to the prediction confidence. They were further annotated via a proteome annotation pipeline that we developed. We also set up a secreted protein classification pipeline and classified our predicted secreted proteins into different functional categories. To make the dataset more convincing and comprehensive, nine reference datasets are also integrated, such as the secreted proteins from the Gene Ontology Annotation (GOA) system at the European Bioinformatics Institute, and the vertebrate secreted proteins from Swiss-Prot. All these entries were grouped via a TribeMCL based clustering pipeline. We have constructed a web-based secreted protein database, which has been publicly available at http://spd.cbi.pku.edu.cn. Users can browse the database via a GO assignment or chromosomal-location-based interface. Moreover, text query and sequence similarity search are also provided, and the sequence and annotation data can be downloaded freely from the SPD website.

Animals↗

Integrating genomic data to predict transcription factor binding.

Transcription factor binding sites (TFBS) in gene promoter regions are often predicted by using position specific scoring matrices (PSSMs), which summarize sequence patterns of experimentally determined TF binding sites. Although PSSMs are more reliable than simple consensus string matching in predicting a true binding site, they generally result in high numbers of false positive hits. This study attempts to reduce the number of false positive matches and generate new predictions by integrating various types of genomic data by two methods: a Bayesian allocation procedure, and support vector machine classification. Several methods will be explored to strengthen the prediction of a true TFBS in the Saccharomyces cerevisiae genome: binding site degeneracy, binding site conservation, phylogenetic profiling, TF binding site clustering, gene expression profiles, GO functional annotation, and k-mer counts in promoter regions. Binding site degeneracy (or redundancy) refers to the number of times a particular transcription factor's binding motif is discovered in the upstream region of a gene. Phylogenetic conservation takes into account the number of orthologous upstream regions in other genomes that contain a particular binding site. Phylogenetic profiling refers to the presence or absence of a gene across a large set of genomes. Binding site clusters are statistically significant clusters of TF binding sites detected by the algorithm ClusterBuster. Gene expression takes into account the idea that when the gene expression profiles of a transcription factor and a potential target gene are correlated, then it is more likely that the gene is a genuine target. Also, genes with highly correlated expression profiles are often regulated by the same TF(s). The GO annotation data takes advantage of the idea that common transcription targets often have related function. Finally, the distribution of the counts of all k-mers of length 4, 5, and 6 in gene's promoter region were examined as means to predict TF binding. In each case the data are compared to known true positives taken from ChIP-chip data, Transfac, and the Saccharomyces Genome Database. First, degeneracy, conservation, expression, and binding site clusters were examined independently and in combination via Bayesian allocation. Then, binding sites were predicted with a support vector machine (SVM) using all methods alone and in combination. The SVM works best when all genomic data are combined, but can also identify which methods contribute the most to accurate classification. On average, a support vector machine can classify binding sites with high sensitivity and an accuracy of almost 80%.

Algorithms↗

A comprehensive BAC resource.

The Human Genome Project has generated extensive map and sequence data for a large number of Bacterial Artificial Chromosome (BAC) clones. In order to maximize the efficient use of the data and to minimize the redundant work for the research community, The Institute for Genomic Research (TIGR) comprehensive BAC resource (cBACr) (http://www.tigr.org/tdb/BacResource/BAC_resourc e_intro. html) was built as an expansion of the TIGR human BAC ends database. This resource collects, integrates and reports the information on library, maps, sequence, annotation and functions for each human and mouse BAC. The current database contains 635 016 human BACs and 265 617 mouse BACs that were characterized by various approaches, among which 22 705 human clones and 1000 mouse clones have sequence and annotation data.

Animals↗

Identification of a new gene controlling plant height in rice using the candidate-gene strategy.

A gene underlying a quantitative trait locus (QTL) controlling plant height on chromosome 1 (QTLph1) in rice ( Oryza sativa L.) was identified using the candidate-gene strategy. First, the function of a targeted gene was analyzed using near isogenic lines (NILs) in which the chromosomal region of a targeted QTL was substituted with that of another line. Second, for physiological information, the candidate gene was selected in the annotation data by the genome sequencing. Physiological analyses of an NIL-expressing QTLph1 (NIL6) suggested that the targeted gene controls plant height by enabling higher amounts of sucrose to be translocated in leaves. The results indicated that the gene for sucrose phosphate synthase (SPS; EC 2.4.1.14), the major limiting enzyme for sucrose synthesis, is a candidate gene for QTLph1 among the annotation results of the region of QTLph1. The higher level of SPS transcripts and the activity of SPS in NIL6 compared to control plants, and the fact that the relative SPS activity per SPS protein content was almost the same between NIL6 and Nipponbare suggested that the higher plant height in NIL6 compared to Nipponbare was due to the high SPS activity in NIL6. In agreement with this hypothesis, transgenic rice plants with a maize SPS gene that had about 3 times the SPS activity of that in Nipponbare (control plants) were significantly taller than Nipponbare from the early growth stage. From these results and the physiological data from NIL6, we concluded that SPS is the targeted gene underlying QTLph1.

Base Sequence↗

MannDB - a microbial database of automated protein sequence analyses and evidence integration for protein characterization.

BACKGROUND: MannDB was created to meet a need for rapid, comprehensive automated protein sequence analyses to support selection of proteins suitable as targets for driving the development of reagents for pathogen or protein toxin detection. Because a large number of open-source tools were needed, it was necessary to produce a software system to scale the computations for whole-proteome analysis. Thus, we built a fully automated system for executing software tools and for storage, integration, and display of automated protein sequence analysis and annotation data. DESCRIPTION: MannDB is a relational database that organizes data resulting from fully automated, high-throughput protein-sequence analyses using open-source tools. Types of analyses provided include predictions of cleavage, chemical properties, classification, features, functional assignment, post-translational modifications, motifs, antigenicity, and secondary structure. Proteomes (lists of hypothetical and known proteins) are downloaded and parsed from Genbank and then inserted into MannDB, and annotations from SwissProt are downloaded when identifiers are found in the Genbank entry or when identical sequences are identified. Currently 36 open-source tools are run against MannDB protein sequences either on local systems or by means of batch submission to external servers. In addition, BLAST against protein entries in MvirDB, our database of microbial virulence factors, is performed. A web client browser enables viewing of computational results and downloaded annotations, and a query tool enables structured and free-text search capabilities. When available, links to external databases, including MvirDB, are provided. MannDB contains whole-proteome analyses for at least one representative organism from each category of biological threat organism listed by APHIS, CDC, HHS, NIAID, USDA, USFDA, and WHO. CONCLUSION: MannDB comprises a large number of genomes and comprehensive protein sequence analyses representing organisms listed as high-priority agents on the websites of several governmental organizations concerned with bio-terrorism. MannDB provides the user with a BLAST interface for comparison of native and non-native sequences and a query tool for conveniently selecting proteins of interest. In addition, the user has access to a web-based browser that compiles comprehensive and extensive reports. Access to MannDB is freely available at http://manndb.llnl.gov/.

Algorithms↗

IMGT, the international ImMunoGeneTics database: a high-quality information system for comparative immunogenetics and immunology.

IMGT, the international ImMunoGeneTics database (http://imgt.cines.fr), is a high quality integrated information system specializing in Immunoglobulins (IG), T cell Receptors (TR) and Major Histocompatibility Complex (MHC) of human and other vertebrates, created in 1989, by LIGM, at the Université Montpellier II, CNRS, Montpellier, France. IMGT provides a common access to standardized data, which include nucleotide and protein sequences, oligonucleotide primers, gene maps, genetic polymorphisms, specificities, 2D and 3D structures. IMGT includes several databases (IMGT/LIGM-DB, IMGT/HLA-DB, IMGT/3Dstructure-DB), Web resources ('IMGT Marie-Paule page') which comprise IMGT Scientific Chart, IMGT Repertoire, IMGT Bloc-notes, IMGT Education, IMGT Aide-mémoire and IMGT Index, and interactive tools (IMGT/V-QUEST, IMGT/JunctionAnalysis). These expertly annotated data on the genome, proteome, genetics and structure of the IG, TR and MHC are of high value for comparative genome evolution studies of the adaptative immune response.

Animals↗

Protein structures and information extraction from biological texts: the PASTA system.

MOTIVATION: The rapid increase in volume of protein structure literature means useful information may be hidden or lost in the published literature and the process of finding relevant material, sometimes the rate-determining factor in new research, may be arduous and slow. RESULTS: We describe the Protein Active Site Template Acquisition (PASTA) system, which addresses these problems by performing automatic extraction of information relating to the roles of specific amino acid residues in protein molecules from online scientific articles and abstracts. Both the terminology recognition and extraction capabilities of the system have been extensively evaluated against manually annotated data and the results compare favourably with state-of-the-art results obtained in less challenging domains. PASTA is the first information extraction (IE) system developed for the protein structure domain and one of the most thoroughly evaluated IE system operating on biological scientific text to date. AVAILABILITY: PASTA makes its extraction results available via a browser-based front end: http://www.dcs.shef.ac.uk/nlp/pasta/. The evaluation resources (manually annotated corpora) are also available through the website: http://www.dcs.shef.ac.uk/nlp/pasta/results.html.

Abstracting and Indexing↗

Analyzing gene expression time-courses.

Measuring gene expression over time can provide important insights into basic cellular processes. Identifying groups of genes with similar expression time-courses is a crucial first step in the analysis. As biologically relevant groups frequently overlap, due to genes having several distinct roles in those cellular processes, this is a difficult problem for classical clustering methods. We use a mixture model to circumvent this principal problem, with hidden Markov models (HMMs) as effective and flexible components. We show that the ensuing estimation problem can be addressed with additional labeled data-partially supervised learning of mixtures-through a modification of the Expectation-Maximization (EM) algorithm. Good starting points for the mixture estimation are obtained through a modification to Bayesian model merging, which allows us to learn a collection of initial HMMs. We infer groups from mixtures with a simple information-theoretic decoding heuristic, which quantifies the level of ambiguity in group assignment. The effectiveness is shown with high-quality annotation data. As the HMMs we propose capture asynchronous behavior by design, the groups we find are also asynchronous. Synchronous subgroups are obtained from a novel algorithm based on Viterbi paths. We show the suitability of our HMM mixture approach on biological and simulated data and through the favorable comparison with previous approaches. A software implementing the method is freely available under the GPL from http://ghmm.org/gql.

Algorithms↗

Ontological analysis of gene expression data: current tools, limitations, and open problems.

Independent of the platform and the analysis methods used, the result of a microarray experiment is, in most cases, a list of differentially expressed genes. An automatic ontological analysis approach has been recently proposed to help with the biological interpretation of such results. Currently, this approach is the de facto standard for the secondary analysis of high throughput experiments and a large number of tools have been developed for this purpose. We present a detailed comparison of 14 such tools using the following criteria: scope of the analysis, visualization capabilities, statistical model(s) used, correction for multiple comparisons, reference microarrays available, installation issues and sources of annotation data. This detailed analysis of the capabilities of these tools will help researchers choose the most appropriate tool for a given type of analysis. More importantly, in spite of the fact that this type of analysis has been generally adopted, this approach has several important intrinsic drawbacks. These drawbacks are associated with all tools discussed and represent conceptual limitations of the current state-of-the-art in ontological analysis. We propose these as challenges for the next generation of secondary data analysis tools.

Algorithms↗

The genome of M. acetivorans reveals extensive metabolic and physiological diversity.

Methanogenesis, the biological production of methane, plays a pivotal role in the global carbon cycle and contributes significantly to global warming. The majority of methane in nature is derived from acetate. Here we report the complete genome sequence of an acetate-utilizing methanogen, Methanosarcina acetivorans C2A. Methanosarcineae are the most metabolically diverse methanogens, thrive in a broad range of environments, and are unique among the Archaea in forming complex multicellular structures. This diversity is reflected in the genome of M. acetivorans. At 5,751,492 base pairs it is by far the largest known archaeal genome. The 4524 open reading frames code for a strikingly wide and unanticipated variety of metabolic and cellular capabilities. The presence of novel methyltransferases indicates the likelihood of undiscovered natural energy sources for methanogenesis, whereas the presence of single-subunit carbon monoxide dehydrogenases raises the possibility of nonmethanogenic growth. Although motility has not been observed in any Methanosarcineae, a flagellin gene cluster and two complete chemotaxis gene clusters were identified. The availability of genetic methods, coupled with its physiological and metabolic diversity, makes M. acetivorans a powerful model organism for the study of archaeal biology. [Sequence, data, annotations and analyses are available at http://www-genome.wi.mit.edu/.]

Archaeal Proteins↗

The wheat (Triticum aestivum L.) leaf proteome.

The wheat leaf proteome was mapped and partially characterized to function as a comparative template for future wheat research. In total, 404 proteins were visualized, and 277 of these were selected for analysis based on reproducibility and relative quantity. Using a combination of protein and expressed sequence tag database searching, 142 proteins were putatively identified with an identification success rate of 51%. The identified proteins were grouped according to their functional annotations with the majority (40%) being involved in energy production, primary, or secondary metabolism. Only 8% of the protein identifications lacked ascertainable functional annotation. The 51% ratio of successful identification and the 8% unclear functional annotation rate are major improvements over most previous plant proteomic studies. This clearly indicates the advancement of the plant protein and nucleic acid sequence and annotation data available in the databases, and shows the enhanced feasibility of future wheat leaf proteome research.

Computational Biology↗

IMGT unique numbering for immunoglobulin and T cell receptor constant domains and Ig superfamily C-like domains.

IMGT, the international ImMunoGeneTics information system (http://imgt.cines.fr) provides a common access to expertly annotated data on the genome, proteome, genetics and structure of immunoglobulins (IG), T cell receptors (TR), major histocompatibility complex (MHC), and related proteins of the immune system (RPI) of human and other vertebrates. The NUMEROTATION concept of IMGT-ONTOLOGY has allowed to define a unique numbering for the variable domains (V-DOMAINs) and for the V-LIKE-DOMAINs. In this paper, this standardized characterization is extended to the constant domains (C-DOMAINs), and to the C-LIKE-DOMAINs, leading, for the first time, to their standardized description of mutations, allelic polymorphisms, two-dimensional (2D) representations and tridimensional (3D) structures. The IMGT unique numbering is, therefore, highly valuable for the comparative, structural or evolutionary studies of the immunoglobulin superfamily (IgSF) domains, V-DOMAINs and C-DOMAINs of IG and TR in vertebrates, and V-LIKE-DOMAINs and C-LIKE-DOMAINs of proteins other than IG and TR, in any species.

Amino Acid Sequence↗

IMGT unique numbering for MHC groove G-DOMAIN and MHC superfamily (MhcSF) G-LIKE-DOMAIN.

IMGT, the international ImMunoGeneTics information system (http://imgt.cines.fr) provides a common access to expertly annotated data on the genome, proteome, genetics and structure of immunoglobulins (IG), T cell receptors (TR), major histocompatibility complex (MHC), and related proteins of the immune system (RPI) of human and other vertebrates. The NUMEROTATION concept of IMGT-ONTOLOGY has allowed to define a unique numbering for the variable domains (V-DOMAINs) and constant domains (C-DOMAINs) of the IG and TR, which has been extended to the V-LIKE-DOMAINs and C-LIKE-DOMAINs of the immunoglobulin superfamily (IgSF) proteins other than the IG and TR (Dev Comp Immunol 27:55--77, 2003; 29:185--203, 2005). In this paper, we describe the IMGT unique numbering for the groove domains (G-DOMAINs) of the MHC and for the G-LIKE-DOMAINs of the MHC superfamily (MhcSF) proteins other than MHC. This IMGT unique numbering leads, for the first time, to the standardized description of the mutations, allelic polymorphisms, two-dimensional (2D) representations and three-dimensional (3D) structures of the G-DOMAINs and G-LIKE-DOMAINs in any species, and therefore, is highly valuable for their comparative, structural, functional and evolutionary studies.

Amino Acid Sequence↗

IMGT unique numbering for immunoglobulin and T cell receptor variable domains and Ig superfamily V-like domains.

IMGT, the international ImMunoGeneTics database (http://imgt.cines.fr) is a high quality integrated information system specializing in immunoglobulins (IG), T cell receptors (TR) and major histocompatibility complex (MHC) of human and other vertebrates. IMGT provides a common access to expertly annotated data on the genome, proteome, genetics and structure of the IG and TR, based on the IMGT Scientific chart and IMGT-ONTOLOGY. The IMGT unique numbering defined for the IG and TR variable regions and domains of all jawed vertebrates has allowed a redefinition of the limits of the framework (FR-IMGT) and complementarity determining regions (CDR-IMGT), leading, for the first time, to a standardized description of mutations, allelic polymorphisms, 2D representations (Colliers de Perles) and 3D structures, whatever the antigen receptor, the chain type, or the species. The IMGT numbering has been extended to the V-like domain and is, therefore, highly valuable for comparative analysis and evolution studies of proteins belonging to the IG superfamily.

Amino Acid Sequence↗

Enhanced gluconeogenesis and increased energy storage as hallmarks of aging in Saccharomyces cerevisiae.

A relationship between life span and cellular glucose metabolism has been inferred from genetic manipulations and caloric restriction of model organisms. In this report, we have used the Snf1p glucose-sensing pathway of Saccharomyces cerevisiae to explore the genetic and biochemical linkages between glucose metabolism and aging. Snf1p is a serine/threonine kinase that regulates cellular responses to glucose deprivation. Loss of Snf4p, an activator of Snf1p, extends generational life span whereas loss of Sip2p, a presumed repressor of the kinase, causes an accelerated aging phenotype. An annotated data base of global age-associated changes in gene expression in isogenic wild-type, sip2Delta, and snf4Delta strains was generated from DNA microarray studies. The transcriptional responses suggested that gluconeogenesis and glucose storage increase as wild-type cells age, that this metabolic evolution is exaggerated in rapidly aging sip2Delta cells, and that it is attenuated in longer-lived snf4Delta cells. To test this hypothesis directly, we applied microanalytic biochemical methods to generation-matched cells from each strain and measured the activities of enzymes and concentrations of metabolites in the gluconeogenic, glycolytic, and glyoxylate pathways, as well as glycogen, ATP, and NAD(+). The sensitivity of the assays allowed comprehensive biochemical profiling to be performed using aliquots of the same cell populations employed for the transcriptional profiling. The results provided additional evidence that aging in S. cerevisiae is associated with a shift away from glycolysis and toward gluconeogenesis and energy storage. They also disclosed that this shift is forestalled by two manipulations that extend life span, caloric restriction and genetic attenuation of the normal age-associated increase in Snf1p activity. Together, these findings indicate that Snf1p activation is not only a marker of aging but also a candidate mediator, because a shift toward energy storage over expenditure could impact myriad aspects of cellular maintenance and repair.

Culture Media↗

The GeneAround GO viewer.

We have developed a system for visualizing the Gene Ontology((TM)) hierarchy. The graphical browser interactively displays diagrams of the inheritance relationship for each term to help understand the meanings of terms when handling gene annotation data.

Computer Graphics↗

GelScape: a web-based server for interactively annotating, manipulating, comparing and archiving 1D and 2D gel images.

GelScape is a web-based tool that permits facile, interactive annotation, comparison, manipulation and storage of protein gel images. It uses Java applet-servlet technology to allow rapid, remote image handling and image processing in a platform-independent manner. It supports many of the features found in commercial, stand-alone gel analysis software including spot annotation, spot integration, gel warping, image resizing, HTML image mapping, image overlaying as well as the storage of gel image and gel annotation data in compliance with Federated Gel Database requirements.

Computer Graphics↗