Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

A systematic approach to reconstructing transcription networks in Saccharomycescerevisiae.

Decomposing regulatory networks into functional modules is a first step toward deciphering the logical structure of complex networks. We propose a systematic approach to reconstructing transcription modules (defined by a transcription factor and its target genes) and identifying conditionsperturbations under which a particular transcription module is activateddeactivated. Our approach integrates information from regulatory sequences, genome-wide mRNA expression data, and functional annotation. We systematically analyzed gene expression profiling experiments in which the yeast cell was subjected to various environmental or genetic perturbations. We were able to construct transcription modules with high specificity and sensitivity for many transcription factors, and predict the activation of these modules under anticipated as well as unexpected conditions. These findings generate testable hypotheses when combined with existing knowledge on signaling pathways and protein-protein interactions. Correlating the activation of a module to a specific perturbation predicts links in the cell's regulatory networks, and examining coactivated modules suggests specific instances of crosstalk between regulatory pathways.

Gene Expression Profiling↗

Full subunit coverage liquid chromatography electrospray ionization mass spectrometry (LCMS+) of an oligomeric membrane protein: cytochrome b(6)f complex from spinach and the cyanobacterium Mastigocladus laminosus.

Highly active cytochrome b(6)f complexes from spinach and the cyanobacterium Mastigocladus laminosus have been analyzed by liquid chromatography with electrospray ionization mass spectrometry (LCMS+). Both size-exclusion and reverse-phase separations were used to separate protein subunits allowing measurement of their molecular masses to an accuracy exceeding 0.01% (+/-3 Da at 30,000 Da). The products of petA, petB, petC, petD, petG, petL, petM, and petN were detected in complexes from both spinach and M. laminosus, while the spinach complex also contained ferredoxin-NADP(+) oxidoreductase (Zhang, H., Whitelegge, J. P., and Cramer, W. A. (2001) Flavonucleotide:ferredoxin reductase is a subunit of the plant cytochrome b(6)f complex. J. Biol. Chem. 276, 38159-38165). While the measured masses of PetC and PetD (18935.8 and 17311.8 Da, respectively) from spinach are consistent with the published primary structure, the measured masses of cytochrome f (31934.7 Da, PetA) and cytochrome b (24886.9 Da, PetB) modestly deviate from values calculated based upon genomic sequence and known post-translational modifications. The low molecular weight protein subunits have been sequenced using tandem mass spectrometry (MSMS) without prior cleavage. Sequences derived from the MSMS spectra of these intact membrane proteins in the range of 3.2-4.2 kDa were compared with translations of genomic DNA sequence where available. Products of the spinach chloroplast genome, PetG, PetL, and PetN, all retained their initiating formylmethionine, while the nuclear encoded PetM was cleaved after import from the cytoplasm. While the sequences of PetG and PetN revealed no discrepancy with translations of the spinach chloroplast genome, Phe was detected at position 2 of PetL. The spinach chloroplast genome reports a codon for Ser at position 2 implying the presence of a DNA sequencing error or a previously undiscovered RNA editing event. Clearly, complete annotation of genomic data requires detailed expression measurements of primary structure by mass spectrometry. Full subunit coverage of an oligomeric intrinsic membrane protein complex by LCMS+ presents a new facet to intact mass proteomics.

Amino Acid Sequence↗

Comparative microarray analysis.

Microarrays enable high-throughput parallel gene expression analysis, and their use has grown exponentially during the past decade. We are now in a position where individual experiments could benefit from using the swelling public data repositories to allow microarrays to progress from being a hypothesis-generating tool to a powerful resource that can be used to test hypothesis about biology. Comparative microarray analysis could better distinguish phenotypes from associated phenotypes; identify valid differentially expressed genes by combining many studies; test new hypothesis; and discover fundamental patterns of gene regulation. This review aims to describe the additional methodology needed for such comparative microarray analysis, and we identify and discuss a number of problems such as loss of published data, lack of annotations, and variable array quality, which need to be solved before comparative microarray analysis can be used in a more systematic and powerful manner.

Animals↗

Mining sequence annotation databanks for association patterns.

MOTIVATION: Millions of protein sequences currently being deposited to sequence databanks will never be annotated manually. Similarity-based annotation generated by automatic software pipelines unavoidably contains spurious assignments due to the imperfection of bioinformatics methods. Examples of such annotation errors include over- and underpredictions caused by the use of fixed recognition thresholds and incorrect annotations caused by transitivity based information transfer to unrelated proteins or transfer of errors already accumulated in databases. One of the most difficult and timely challenges in bioinformatics is the development of intelligent systems aimed at improving the quality of automatically generated annotation. A possible approach to this problem is to detect anomalies in annotation items based on association rule mining. RESULTS: We present the first large-scale analysis of association rules derived from two large protein annotation databases-Swiss-Prot and PEDANT-and reveal novel, previously unknown tendencies of rule strength distributions. Most of the rules are either very strong or very weak, with rules in the medium strength range being relatively infrequent. Based on dynamics of error correction in subsequent Swiss-Prot releases and on our own manual analysis we demonstrate that exceptions from strong rules are, indeed, significantly enriched in annotation errors and can be used to automatically flag them. We identify different strength dependencies of rules derived from different fields in Swiss-Prot. A compositional breakdown of association rules generated from PEDANT in terms of their constituent items indicates that most of the errors that can be corrected are related to gene functional roles. Swiss-Prot errors are usually caused by under-annotation owing to its conservative approach, whereas automatically generated PEDANT annotation suffers from over-annotation. AVAILABILITY: All data generated in this study are available for download and browsing at http://pedant.gsf.de/ARIA/index.htm.

Conserved Sequence↗

BioIE: extracting informative sentences from the biomedical literature.

SUMMARY: BioIE is a rule-based system that extracts informative sentences relating to protein families, their structures, functions and diseases from the biomedical literaturE. Based on manual definition of templates and rules, it aims at precise sentence extraction rather than wide recall. After uploading source text or retrieving abstracts from MEDLINE, users can extract sentences based on predefined or user-defined template categories. BioIE also provides a brief insight into the syntactic and semantic context of the source-text by looking at word, N-gram and MeSH-term distributions. Important Applications of BioIE are in, for example, annotation of microarray data and of protein databases. AVAILABILITY: http://umber.sbs.man.ac.uk/dbbrowser/bioie/

Database Management Systems↗

ACT: the Artemis Comparison Tool.

The Artemis Comparison Tool (ACT) allows an interactive visualisation of comparisons between complete genome sequences and associated annotations. The comparison data can be generated with several different programs; BLASTN, TBLASTX or Mummer comparisons between genomic DNA sequences, or orthologue tables generated by reciprocal FASTA comparison between protein sets. It is possible to identify regions of similarity, insertions and rearrangements at any level from the whole genome to base-pair differences. ACT uses Artemis components to display the sequences and so inherits powerful searching and analysis tools. ACT is part of the Artemis distribution and is similarly open source, written in Java and can run on any Java enabled platform, including UNIX, Macintosh and Windows.

Algorithms↗

Characterization of long cDNA clones from human adult spleen. II. The complete sequences of 81 cDNA clones.

To accumulate information on the coding sequences (CDSs) of unidentified genes, we have conducted a sequencing project of human long cDNA clones. Both the end sequences of approximately 10,000 cDNA clones from two size-fractionated human spleen cDNA libraries (average sizes of 4.5 kb and 5.6 kb) were determined by single-pass sequencing to select cDNAs with unidentified sequences. We herein present the entire sequences of 81 cDNA clones, most of which were selected by two approaches based on their protein-coding potentialities in silico: Fifty-eight cDNA clones were selected as those having protein-coding potentialities at the 5'-end of single-pass sequences by applying the GeneMark analysis; and 20 cDNA clones were selected as those expected to encode proteins larger than 100 amino acid residues by analysis of the human genome sequences flanked by both the end sequences of cDNAs using the GENSCAN gene prediction program. In addition to these newly identified cDNAs, three cDNA clones were isolated by colony hybridization experiments using probes corresponding to known gene sequences since these cDNAs are likely to contain considerable amounts of new information regarding the genes already annotated. The sequence data indicated that the average sizes of the inserts and corresponding CDSs of cDNA clones analyzed here were 5.0 kb and 2.0 kb (670 amino acid residues), respectively. From the results of homology and motif searches against the public databases, functional categories of the 29 predicted gene products could be assigned; 86% of these predicted gene products (25 gene products) were classified into proteins relating to cell signaling/communication, nucleic acid management, and cell structure/motility.

Adult↗

Biases and complex patterns in the residues flanking protein N-glycosylation sites.

N-Glycosylation, the most common and most versatile protein modification reaction, occurs at the beta-amide of the aspargine of the Asn-Xaa-Ser/Thr sequon. For reasons that are unclear, not all such sequons are glycosylated. To find patterns that affect glycosylation, we examined the amino acid residues from the 20th preceding the sequon to the 20th residue following it, using bioinformatics tools. A clean data set of annotated, experimentally verified, glycosylated and nonglycosylated sequons derived from 617 well-defined nonredundant N- and N-,O-glycoproteins listed in SWISS-PROT (June 2002) was used. NXS and NXT sequons were analyzed separately. Although no overt patterns were found to explain sequon occupancy or nonoccupancy, trends for over- or underrepresentation of certain amino acids at particular positions were statistically significant and different in NXS and NXT sequons. In extension of earlier reports, none of the 80 Asn-Pro-Ser/Thr found were glycosylated, and a markedly low level of glycosylation was seen in sequons with Pro at the position following the Ser/Thr. In addition, a general observation was made that the considerable number of glycosylated sequons in the C-terminal 10 residues of glycoproteins suggests that N-glycosylation in these cases may be posttranslational and not cotranslational, as widely accepted.

Amino Acid Sequence↗

Origin and molecular evolution of receptor tyrosine kinases with immunoglobulin-like domains.

Receptor tyrosine kinases (RTKs) are involved in the control of fundamental cellular processes in metazoans. In vertebrates, RTK could be grouped in distinct classes based on the nature of their cognate ligand and modular composition of their extracellular domain. RTK with immunoglobulin-like domains (IG-like RTK) encompass several RTK classes and have been found in early metazoans, including sponges. Evolution of IG-like RTK is characterized by extended molecular and functional diversification, which prompted us to study their evolutionary history. For that purpose, a nonredundant data set including annotated protein sequences of IG-like RTK (n = 85) was built, representing 19 species ranging from sponges to humans. Phylogenetic trees were generated from alignment of conserved regions using maximum likelihood approach. Molecular phylogeny strongly suggests that IG-like RTK diversification occurred according to a complex scenario. In particular, we propose that specific cis duplications of a common ancestor to both platelet-derived growth factor receptor (class III) and vascular endothelial growth factor receptor (class V) families preceded two trans duplications. In contrast, other IG-like RTK genes, like Musk and PTK7, apparently did not evolve by duplications, whereas fibroblast growth factor receptors (class IV) evolved through two rounds of trans duplications. The proposed model of IG-like RTK evolution is supported by high bootstrap values and by the clustering of genes encoding class III and class V RTKs at specific chromosomal locations in mouse and human genomes.

Animals↗

ExInt: an Exon Intron Database.

The Exon/Intron Database (ExInt) stores information of all GenBank eukaryotic entries containing an annotated intron sequence. Data are available through a retrieval system, as flat-files and as a MySQL dump file. In this report we discuss several implementations added to ExInt, which is accessible at http://intron.bic.nus.edu.sg/exint/newexint/exint.html.

Animals↗

IMGT, the international ImMunoGeneTics database.

The international ImMunoGeneTics database (IMGT) (http://imgt.cines.fr), is a high quality integrated information system specializing in Immunoglobulins (IG), T cell Receptors (TR) and Major Histocompatibility Complex (MHC) of human and other vertebrates, created in 1989, by the Laboratoire d'ImmunoGénétique Moléculaire (LIGM), at the Université Montpellier II, CNRS, Montpellier, France. IMGT provides a common access to standardized data which include nucleotide and protein sequences, oligonucleotide primers, gene maps, genetic polymorphisms, specificities, 2D and 3D structures. IMGT includes three sequence databases (IMGT/LIGM-DB, IMGT/MHC-DB, IMGT/PRIMER-DB), one genome database (IMGT/GENE-DB) with different interfaces (IMGT/GeneSearch, IMGT/GeneView, IMGT/LocusView), one 3D structure database (IMGT/3Dstructure-DB), Web resources comprising 8000 HTML pages ('IMGT Marie-Paule page') and interactive tools for sequence analysis (IMGT/V-QUEST, IMGT/JunctionAnalysis, IMGT/Allele-Align, IMGT/PhyloGene). IMGT data are expertly annotated according to the rules of the IMGT Scientific chart, based on IMGT-ONTOLOGY. IMGT tools are particularly useful for the analysis of the IG and TR repertoires in physiological normal and pathological situations. IMGT has important applications in medical research (autoimmune diseases, AIDS, leukemias, lymphomas, myelomas), biotechnology related to antibody engineering (phage displays, combinatorial libraries) and thera-peutic approaches (graft, immunotherapy). IMGT is freely available at http://imgt.cines.fr.

Animals↗

dictyBase: a new Dictyostelium discoideum genome database.

Dictyostelium discoideum is a powerful and genetically tractable model system used for the study of numerous cellular molecular mechanisms including chemotaxis, phagocytosis and signal transduction. The past 2 years have seen a significant expansion in the scope and accessibility of online resources for Dictyostelium. Recent advances have focused on the development of a new comprehensive online resource called dictyBase (http://dictybase.org). This database not only provides access to genomic data including functional annotation of genes, gene products and chromosomal mapping, but also to extensive biological information such as mutant phenotypes and corresponding reference material. In conjunction with additional sites (http://genome. imb-jena.de/dictyostelium/, http://dictyensembl. bioch.bcm.tmc.edu and http://www.sanger.ac.uk/Projects/D_discoideum/) from the genome sequencing and assembly centers, these improvements have expanded the scope of the Dictyostelium databases making them accessible and useful to any researcher interested in comparative and functional genomics in metazoan organisms.

Animals↗

IMGT, the international ImMunoGeneTics information system.

The international ImMunoGeneTics information system (IMGT) (http://imgt.cines.fr), created in 1989, by the Laboratoire d'ImmunoGenetique Moleculaire LIGM (Universite Montpellier II and CNRS) at Montpellier, France, is a high-quality integrated knowledge resource specializing in the immunoglobulins (IGs), T cell receptors (TRs), major histocompatibility complex (MHC) of human and other vertebrates, and related proteins of the immune systems (RPI) that belong to the immunoglobulin superfamily (IgSF) and to the MHC superfamily (MhcSF). IMGT includes several sequence databases (IMGT/LIGM-DB, IMGT/PRIMER-DB, IMGT/PROTEIN-DB and IMGT/MHC-DB), one genome database (IMGT/GENE-DB) and one three-dimensional (3D) structure database (IMGT/3Dstructure-DB), Web resources comprising 8000 HTML pages (IMGT Marie-Paule page), and interactive tools. IMGT data are expertly annotated according to the rules of the IMGT Scientific chart, based on the IMGT-ONTOLOGY concepts. IMGT tools are particularly useful for the analysis of the IG and TR repertoires in normal physiological and pathological situations. IMGT is used in medical research (autoimmune diseases, infectious diseases, AIDS, leukemias, lymphomas, myelomas), veterinary research, biotechnology related to antibody engineering (phage displays, combinatorial libraries, chimeric, humanized and human antibodies), diagnostics (clonalities, detection and follow up of residual diseases) and therapeutical approaches (graft, immunotherapy and vaccinology). IMGT is freely available at http://imgt.cines.fr.

Animals↗

Reactome: a knowledgebase of biological pathways.

Reactome, located at http://www.reactome.org is a curated, peer-reviewed resource of human biological processes. Given the genetic makeup of an organism, the complete set of possible reactions constitutes its reactome. The basic unit of the Reactome database is a reaction; reactions are then grouped into causal chains to form pathways. The Reactome data model allows us to represent many diverse processes in the human system, including the pathways of intermediary metabolism, regulatory pathways, and signal transduction, and high-level processes, such as the cell cycle. Reactome provides a qualitative framework, on which quantitative data can be superimposed. Tools have been developed to facilitate custom data entry and annotation by expert biologists, and to allow visualization and exploration of the finished dataset as an interactive process map. Although our primary curational domain is pathways from Homo sapiens, we regularly create electronic projections of human pathways onto other organisms via putative orthologs, thus making Reactome relevant to model organism research communities. The database is publicly available under open source terms, which allows both its content and its software infrastructure to be freely used and redistributed.

Animals↗

QTL MatchMaker: a multi-species quantitative trait loci (QTL) database and query system for annotation of genes and QTL.

Identifying genes that underlie quantitative trait loci (QTL) is a challenging task. Here, we present a new QTL software system, named QTL MatchMaker. The system is designed to integrate and mine QTL information across human, mouse and rat genomes and to annotate functional genomic data. It combines and organizes information from relevant public databases and publications and integrates QTL, physical, genetic and cytogenetic maps across human, mouse and rat. To make this application available to the research community we have developed a website for high-throughput mapping of expressed sequences to QTL and for selection of candidate genes in the physiological genomics context of complex traits. QTL MatchMaker is accessible at http://pmrc.med.mssm.edu:9090/QTL/jsp/qtlhome.jsp.

Animals↗

Genomewide function conservation and phylogeny in the Herpesviridae.

The Herpesviridae are a large group of well-characterized double-stranded DNA viruses for which many complete genome sequences have been determined. We have extracted protein sequences from all predicted open reading frames of 19 herpesvirus genomes. Sequence comparison and protein sequence clustering methods have been used to construct herpesvirus protein homologous families. This resulted in 1692 proteins being clustered into 243 multiprotein families and 196 singleton proteins. Predicted functions were assigned to each homologous family based on genome annotation and published data and each family classified into seven broad functional groups. Phylogenetic profiles were constructed for each herpesvirus from the homologous protein families and used to determine conserved functions and genomewide phylogenetic trees. These trees agreed with molecular-sequence-derived trees and allowed greater insight into the phylogeny of ungulate and murine gammaherpesviruses.

Animals↗

Prospecting for pig single nucleotide polymorphisms in the human genome: have we struck gold?

Gene-to-gene variation in the frequency of single nucleotide polymorphisms (SNPs) has been observed in humans, mice, rats, primates and pigs, but a relationship across species in this variation has not been described. Here, the frequency of porcine coding SNPs (cSNPs) identified by in silico methods, and the frequency of murine cSNPs, were compared with the frequency of human cSNPs across homologous genes. From 150,000 porcine expressed sequence tag (EST) sequences, a total of 452 SNP-containing sequence clusters were found, totalling 1394 putative SNPs. All the clustered porcine EST annotations and SNP data have been made publicly available at http://sputnik.btk.fi/project?name=swine. Human and murine cSNPs were identified from dbSNP and were characterized as either validated or total number of cSNPs (validated plus non-validated) for comparison purposes. The correlation between in silico pig cSNP and validated human cSNP densities was found to be 0.77 (p < 0.00001) for a set of 25 homologous genes, while a correlation of 0.48 (p < 0.0005) was found for a primarily random sample of 50 homologous human and mouse genes. This is the first evidence of conserved gene-to-gene variability in cSNP frequency across species and indicates that site-directed screening of porcine genes that are homologous to cSNP-rich human genes may rapidly advance cSNP discovery in pigs.

Animals↗

Bayesian modeling of differential gene expression.

We present a Bayesian hierarchical model for detecting differentially expressing genes that includes simultaneous estimation of array effects, and show how to use the output for choosing lists of genes for further investigation. We give empirical evidence that expression-level dependent array effects are needed, and explore different nonlinear functions as part of our model-based approach to normalization. The model includes gene-specific variances but imposes some necessary shrinkage through a hierarchical structure. Model criticism via posterior predictive checks is discussed. Modeling the array effects (normalization) simultaneously with differential expression gives fewer false positive results. To choose a list of genes, we propose to combine various criteria (for instance, fold change and overall expression) into a single indicator variable for each gene. The posterior distribution of these variables is used to pick the list of genes, thereby taking into account uncertainty in parameter estimates. In an application to mouse knockout data, Gene Ontology annotations over- and underrepresented among the genes on the chosen list are consistent with biological expectations.

Animals↗