Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

A transcript map encompassing a susceptibility locus for bipolar affective disorder on chromosome 4q35.

Bipolar affective disorder is one of the most common mental illnesses with a population prevalence of approximately 1%. The disorder is genetically complex, with an increasing number of loci being implicated through genetic linkage studies. However, the specific genetic variations and molecules involved in bipolar susceptibility and pathogenesis are yet to be identified. Genetic linkage analysis has identified a bipolar disorder susceptibility locus on chromosome 4q35, and the interval harbouring this susceptibility gene has been narrowed to a size that is amenable to positional cloning. We have used the resources of the Human Genome Project (HGP) and Celera Genomics to identify overlapping sequenced BAC clones and sequence contigs that represent the region implicated by linkage analysis. A combination of bioinformatic tools and laboratory techniques have been applied to annotate this DNA sequence data and establish a comprehensive transcript map that spans approximately 5.5 Mb. This map encompasses the chromosome 4q35 bipolar susceptibility locus, which localises to a "most probable" candidate interval of approximately 2.3 Mb, within a more conservative candidate interval of approximately 5 Mb. Localised within this map are 11 characterised genes and eight novel genes of unknown function, which together provide a collection of candidate transcripts that may be investigated for association with bipolar disorder. Overall, this region was shown to be very gene-poor, with a high incidence of pseudogenes, and redundant and novel repetitive elements. Our analysis of the interval has demonstrated a significant difference in the extent to which the current HGP and Celera sequence data sets represent this region.

Bipolar Disorder↗

Support vector machine-based method for subcellular localization of human proteins using amino acid compositions, their order, and similarity search.

Here we report a systematic approach for predicting subcellular localization (cytoplasm, mitochondrial, nuclear, and plasma membrane) of human proteins. First, support vector machine (SVM)-based modules for predicting subcellular localization using traditional amino acid and dipeptide (i + 1) composition achieved overall accuracy of 76.6 and 77.8%, respectively. PSI-BLAST, when carried out using a similarity-based search against a nonredundant data base of experimentally annotated proteins, yielded 73.3% accuracy. To gain further insight, a hybrid module (hybrid1) was developed based on amino acid composition, dipeptide composition, and similarity information and attained better accuracy of 84.9%. In addition, SVM modules based on a different higher order dipeptide i.e. i + 2, i + 3, and i + 4 were also constructed for the prediction of subcellular localization of human proteins, and overall accuracy of 79.7, 77.5, and 77.1% was accomplished, respectively. Furthermore, another SVM module hybrid2 was developed using traditional dipeptide (i + 1) and higher order dipeptide (i + 2, i + 3, and i + 4) compositions, which gave an overall accuracy of 81.3%. We also developed SVM module hybrid3 based on amino acid composition, traditional and higher order dipeptide compositions, and PSI-BLAST output and achieved an overall accuracy of 84.4%. A Web server HSLPred (www.imtech.res.in/raghava/hslpred/ or bioinformatics.uams.edu/raghava/hslpred/) has been designed to predict subcellular localization of human proteins using the above approaches.

Algorithms↗

A genome wide screening approach for membrane-targeted proteins.

Membrane-associated proteins are critical for intra- and intercellular communication. Accordingly approaches are needed for rapid and comprehensive identification of all membrane-targeted gene products in a given cell or tissue. Here we describe a modification of the yeast Ras recruitment system to this end and designate the modified approach the Ras membrane trap (RMT). A pilot RMT screen was carried out on the central nervous system of the mollusk Lymnaea stagnalis, a model organism from a phylum that still lacks a representative with a sequenced genome. 112 gene products were identified in the screen of which 79 lack assignable homologs in available data bases. Currently available annotation tools predicted membrane association of only 45% of the 112 proteins, although experimental verification in mammalian cells confirmed membrane association for all clones tested. Thus, genome annotation using currently available tools is likely to underpredict representation of membrane-associated gene products. The 32 proteins with known homologies include many targeted to the endoplasmic reticulum or the nucleus, thus RMT provides a tool that can cover intracellular membrane proteomes. Two sequences were found to represent gene families not found to date in invertebrate genomes, emphasizing the need for whole genome sequences from mollusks and indeed from representatives of all major invertebrate phyla.

Animals↗

National Center for Biomedical Ontology: advancing biomedicine through structured organization of scientific knowledge.

The National Center for Biomedical Ontology is a consortium that comprises leading informaticians, biologists, clinicians, and ontologists, funded by the National Institutes of Health (NIH) Roadmap, to develop innovative technology and methods that allow scientists to record, manage, and disseminate biomedical information and knowledge in machine-processable form. The goals of the Center are (1) to help unify the divergent and isolated efforts in ontology development by promoting high quality open-source, standards-based tools to create, manage, and use ontologies, (2) to create new software tools so that scientists can use ontologies to annotate and analyze biomedical data, (3) to provide a national resource for the ongoing evaluation, integration, and evolution of biomedical ontologies and associated tools and theories in the context of driving biomedical projects (DBPs), and (4) to disseminate the tools and resources of the Center and to identify, evaluate, and communicate best practices of ontology development to the biomedical community. Through the research activities within the Center, collaborations with the DBPs, and interactions with the biomedical community, our goal is to help scientists to work more effectively in the e-science paradigm, enhancing experiment design, experiment execution, data analysis, information synthesis, hypothesis generation and testing, and understand human disease.

Biomedical Research↗

OLS4: a new Ontology Lookup Service for a growing interdisciplinary knowledge ecosystem.

SUMMARY: The Ontology Lookup Service (OLS) is an open source search engine for ontologies which is used extensively in the bioinformatics and chemistry communities to annotate biological and biomedical data with ontology terms. Recently, there has been a significant increase in the size and complexity of ontologies due to new scales of biological knowledge, such as spatial transcriptomics, new ontology development methodologies, and curation on an increased scale. Existing Web-based tools for ontology browsing such as BioPortal and OntoBee do not support the full range of definitions used by today's ontologies. In order to support the community going forward, we have developed OLS4, implementing the complete OWL2 specification, internationalization support for multiple languages, and a new user interface with UX enhancements such as links out to external databases. OLS4 has replaced OLS3 in production at EMBL-EBI and has a backward compatible API supporting users of OLS3 to transition. AVAILABILITY AND IMPLEMENTATION: The source code of OLS is available at https://github.com/EBISPOT/ols4 and DOI 10.5281/zenodo.14960290 with Apache 2.0 License. A freely available implementation is accessible at https://www.ebi.ac.uk/ols4.

Biological Ontologies↗

GFPE: gene-finding program evaluation.

Gene-finding program evaluation (GFPE) is a set of Java classes for evaluating gene-finding programs. A command-line interface is also provided. Inputs to the program include the sequence data (in FASTA format), annotations of "actual" sequence features, and annotations of "predicted" sequence features. Annotation files are in the General Feature Format promoted by the Sanger center. GFPE calculates a number of metrics of accuracy of predictions at three levels:the coding level, the exon level, and the protein level.

Computational Biology↗

CaSPredictor: a new computer-based tool for caspase substrate prediction.

MOTIVATION: In vitro studies have shown that the most remarkable catalytic features of caspases, a family of cysteineproteases, are their stringent specificity to Asp (D) in the S1 subsite and at least four amino acids to the left of scissile bound. However, there is little information about the substrate recognition patterns in vivo. The prediction and characterization of proteolytic cleavage sites in natural substrates could be useful for uncovering these structural relationships. RESULTS: PEST-like sequences rich in the amino acids Ser (S), Thr (T), Pro (P), Glu or Asp (E/D), including Asn (N) and Gln (Q) are adjacent structural/sequential elements in the majority of cleavage site regions of the natural caspase substrates described in the literature, supporting its possible implication in the substrate selection by caspases. We developed CaSPredictor, a software which incorporated a PEST-like index and the position-dependent amino acid matrices for prediction of caspase cleavage sites in individual proteins and protein datasets. The program predicted successfully 81% (111/137) of the cleavage sites in experimentally verified caspase substrates not annotated in its internal data file. Its accuracy and confidence was estimated as 80% using ROC methodology. The program was much more efficient in predicting caspase substrates when compared with PeptideCutter and PEPS software. Finally, the program detected potential cleavage sites in the primary sequences of 1644 proteins in a dataset containing 9986 protein entries. AVAILABILITY: Requests for software should be made to Dr José E. Belizário SUPPLEMENTARY INFORMATION: Supplementary information is available for academic users at site http://icb.usp.br/~farmaco/Jose/CaSpredictorfiles.

Algorithms↗

Molecular decomposition of complex clinical phenotypes using biologically structured analysis of microarray data.

MOTIVATION: Today, the characterization of clinical phenotypes by gene-expression patterns is widely used in clinical research. If the investigated phenotype is complex from the molecular point of view, new challenges arise and these have not been addressed systematically. For instance, the same clinical phenotype can be caused by various molecular disorders, such that one observes different characteristic expression patterns in different patients. RESULTS: In this paper we describe a novel algorithm called Structured Analysis of Microarrays (StAM), which accounts for molecular heterogeneity of complex clinical phenotypes. Our algorithm goes beyond established methodology in several aspects: in addition to the expression data, it exploits functional annotations from the Gene Ontology database to build biologically focussed classifiers. These are used to uncover potential molecular disease subentities and associate them to biological processes without compromising overall prediction accuracy. AVAILABILITY: Bioconductor compliant R package SUPPLEMENTARY INFORMATION: Complete analyses are available at http://compdiag.molgen.mpg.de/supplements/lottaz05.

Biomarkers, Tumor↗

Personal access to sequence databases on personal computers.

A comprehensive package of software has been developed to access nucleic acid and protein sequence databases on stand-alone IBM personal computers. The software combines keyword search on the annotation fields of the data with pattern matching algorithms on the biological sequences. Sequences containing complex sites like promoters or kink sites can be identified as well as sequences that are similar to a query sequence. Protein sequences with particular patterns of amino acids such as hydrophobic regions can be identified as well. Considering the relatively inexpensive hard disks now available, personal computers have become a cost-effective alternative to mainframe processing for sequence databases.

Amino Acid Sequence↗

aCHEdb: the database system for ESTHER, the alpha/beta fold family of proteins and the Cholinesterase gene server.

Acetylcholinesterase belongs to a family of proteins, the alpha/beta hydrolase fold family, whose constituents evolutionarily diverged from a common ancestor and share a similar structure of a central beta sheet surrounded by alpha helices. These proteins fulfil a wide range of physiological functions (hydrolases, adhesion molecules, hormone precursors) [Krejci,E., Duval,N., Chatonnet,A., Vincens,P. and Massoulié,J. (1991) Proc. Natl. Acad. Sci. USA , 88, 6647-6651]. ESTHER (for esterases, alpha/beta hydrolase enzymes and relatives) is a database aimed at collecting in one information system, sequence data together with biological annotations and experimental biochemical results related to the structure-function analysis of the enzymes of the family. The major upgrade of the database comes from the use of a new database management system: aCHEdb which uses the ACeDB program designed by Richard Durbin and Jean Thierry-Mieg. It can be found at http://www.ensam.inra.fr/cholinesterase

Animals↗

Identification of a novel gene encoding a flavin-dependent tRNA:m5U methyltransferase in bacteria--evolutionary implications.

Formation of 5-methyluridine (ribothymidine) at position 54 of the T-psi loop of tRNA is catalyzed by site-specific tRNA methyltransferases (tRNA:m(5)U-54 MTase). In all Eukarya and many Gram-negative Bacteria, the methyl donor for this reaction is S-adenosyl-l-methionine (S-AdoMet), while in several Gram-positive Bacteria, the source of carbon is N(5), N(10)-methylenetetrahydrofolate (CH(2)H(4)folate). We have identified the gene for Bacillus subtilis tRNA:m(5)U-54 MTase. The encoded recombinant protein contains tightly bound flavin and is active in Escherichia coli mutant lacking m(5)U-54 in tRNAs and in vitro using T7 tRNA transcript as substrate. This gene is currently annotated gid in Genome Data Banks and it is here renamed trmFO. TrmFO (Gid) orthologs have also been identified in many other bacterial genomes and comparison of their amino acid sequences reveals that they are phylogenetically distinct from either ThyA or ThyX class of thymidylate synthases, which catalyze folate-dependent formation of deoxyribothymine monophosphate, the universal DNA precursor.

Bacillus subtilis↗

The Zebrafish Information Network: the zebrafish model organism database.

The Zebrafish Information Network (ZFIN; http://zfin.org) is a web based community resource that implements the curation of zebrafish genetic, genomic and developmental data. ZFIN provides an integrated representation of mutants, genes, genetic markers, mapping panels, publications and community resources such as meeting announcements and contact information. Recent enhancements to ZFIN include (i) comprehensive curation of gene expression data from the literature and from directly submitted data, (ii) increased support and annotation of the genome sequence, (iii) expanded use of ontologies to support curation and query forms, (iv) curation of morpholino data from the literature, and (v) increased versatility of gene pages, with new data types, links and analysis tools.

Animals↗

The Stanford Microarray Database: implementation of new analysis tools and open source release of software.

The Stanford Microarray Database (SMD; http://smd.stanford.edu/) is a research tool and archive that allows hundreds of researchers worldwide to store, annotate, analyze and share data generated by microarray technology. SMD supports most major microarray platforms, and is MIAME-supportive and can export or import MAGE-ML. The primary mission of SMD is to be a research tool that supports researchers from the point of data generation to data publication and dissemination, but it also provides unrestricted access to analysis tools and public data from 300 publications. In addition to supporting ongoing research, SMD makes its source code fully and freely available to others under an Open Source license, enabling other groups to create a local installation of SMD. In this article, we describe several data analysis tools implemented in SMD and we discuss features of our software release.

Animals↗

Global analysis of outer membrane proteins from Leptospira interrogans serovar Lai.

Recombinant leptospiral outer membrane proteins (OMPs) can elicit immunity to leptospirosis in a hamster infection model. Previously characterized OMPs appear highly conserved, and thus their potential to stimulate heterologous immunity is of critical importance. In this study we undertook a global analysis of leptospiral OMPs, which were obtained by Triton X-114 extraction and phase partitioning. Outer membrane fractions were isolated from Leptospira interrogans serovar Lai grown at 20, 30, and 37 degrees C with or without 10% fetal calf serum and, finally, in iron-depleted medium. The OMPs were separated by two-dimensional gel electrophoresis. Gel patterns from each of the five conditions were compared via image analysis, and 37 gel-purified proteins were tryptically digested and characterized by mass spectrometry (MS). Matrix-assisted laser desorption ionization-time-of-flight MS was used to rapidly identify leptospiral OMPs present in sequence databases. Proteins identified by this approach included the outer membrane lipoproteins LipL32, LipL36, LipL41, and LipL48. No known proteins from any cellular location other than the outer membrane were identified. Tandem electrospray MS was used to obtain peptide sequence information from eight novel proteins designated pL18, pL21, pL22, pL24, pL45, pL47/49, pL50, and pL55. The expression of LipL36 and pL50 was not apparent at temperatures above 30 degrees C or under iron-depleted conditions. The expression of pL24 was also downregulated after iron depletion. The leptospiral major OMP LipL32 was observed to undergo substantial cleavage under all conditions except iron depletion. Additionally, significant downregulation of these mass forms was observed under iron limitation at 30 degrees C, but not at 30 degrees C alone, suggesting that LipL32 processing is dependent on iron-regulated extracellular proteases. However, separate cleavage products responded differently to changes in growth temperature and medium constituents, indicating that more than one process may be involved in LipL32 processing. Furthermore, under iron-depleted conditions there was no concomitant increase in the levels of the intact form of LipL32. The temperature- and iron-regulated expression of LipL36 and the iron-dependent cleavage of LipL32 were confirmed by immunoblotting with specific antisera. Global analysis of the cellular location and expression of leptospiral proteins will be useful in the annotation of genomic sequence data and in providing insight into the biology of Leptospira.

Amino Acid Sequence↗

Genome-scale metabolic model of Helicobacter pylori 26695.

A genome-scale metabolic model of Helicobacter pylori 26695 was constructed from genome sequence annotation, biochemical, and physiological data. This represents an in silico model largely derived from genomic information for an organism for which there is substantially less biochemical information available relative to previously modeled organisms such as Escherichia coli. The reconstructed metabolic network contains 388 enzymatic and transport reactions and accounts for 291 open reading frames. Within the paradigm of constraint-based modeling, extreme-pathway analysis and flux balance analysis were used to explore the metabolic capabilities of the in silico model. General network properties were analyzed and compared to similar results previously generated for Haemophilus influenzae. A minimal medium required by the model to generate required biomass constituents was calculated, indicating the requirement of eight amino acids, six of which correspond to essential human amino acids. In addition a list of potential substrates capable of fulfilling the bulk carbon requirements of H. pylori were identified. A deletion study was performed wherein reactions and associated genes in central metabolism were deleted and their effects were simulated under a variety of substrate availability conditions, yielding a number of reactions that are deemed essential. Deletion results were compared to recently published in vitro essentiality determinations for 17 genes. The in silico model accurately predicted 10 of 17 deletion cases, with partial support for additional cases. Collectively, the results presented herein suggest an effective strategy of combining in silico modeling with experimental technologies to enhance biological discovery for less characterized organisms and their genomes.

Amino Acids↗

Expanded metabolic reconstruction of Helicobacter pylori (iIT341 GSM/GPR): an in silico genome-scale characterization of single- and double-deletion mutants.

Helicobacter pylori is a human gastric pathogen infecting almost half of the world population. Herein, we present an updated version of the metabolic reconstruction of H. pylori strain 26695 based on the revised genome annotation and new experimental data. This reconstruction, iIT341 GSM/GPR, represents a detailed review of the current literature about H. pylori as it integrates biochemical and genomic data in a comprehensive framework. In total, it accounts for 341 metabolic genes, 476 intracellular reactions, 78 exchange reactions, and 485 metabolites. Novel features of iIT341 GSM/GPR include (i) gene-protein-reaction associations, (ii) elementally and charge-balanced reactions, (iii) more accurate descriptions of isoprenoid and lipopolysaccharide metabolism, and (iv) quantitative assessments of the supporting data for each reaction. This metabolic reconstruction was used to carry out in silico deletion studies to identify essential and conditionally essential genes in H. pylori. A total of 128 essential and 75 conditionally essential metabolic genes were identified. Predicted growth phenotypes of single knockouts were validated using published experimental data. In addition, in silico double-deletion studies identified a total of 47 synthetic lethal mutants involving 67 different metabolic genes in rich medium.

Biotin↗

Comparative bioinformatic analysis of genes expressed in common bean (Phaseolus vulgaris L.) seedlings.

To rapidly and cost-effectively generate gene expression data, we developed an annotated unigene database of common bean (Phaseolus vulgaris L.). In this study, 3 cDNA libraries were constructed from the bean breeding line SEL1308, 1 from young leaf and 2 from seedlings inoculated or not inoculated with the fungal pathogen Colletotrichum lindemuthianum (Sacc. & Magnus) Briosi & Cavara, which causes anthracnose in common bean. To this date, 5255 single-pass sequences have been included in the database after selection based on sequence quality. These ESTs were trimmed and clustered using the computer programs Phred and CAP3 to form a unigene collection of 3126 unique sequences. Within clusters, 318 single nucleotide polymorphisms (SNPs) and 68 insertions-deletions (indels) were found, indicating the presence of paralogous gene families in our database. Each unigene sequence was analyzed for possible function using their similarity to known genes represented in the GenBank database and classified into 14 categories. Only 314 unigenes showed significant similarities to Phaseolus genomic sequences and P. vulgaris ESTs, which indicates that 90% (2818 unigenes) of our database represent newly discovered common bean genes. In addition, 12% (387 unigenes) were shown to be specific to common bean. This study represents a first step towards the discovery of novel genes in beans and a valuable source of molecular markers for expressed gene tagging and mapping.

Computational Biology↗

Functional network analysis reveals extended gliomagenesis pathway maps and three novel MYC-interacting genes in human gliomas.

Gene expression profiling has proven useful in subclassification and outcome prognostication for human glial brain tumors. The analysis of biological significance of the hundreds or thousands of alterations in gene expression found in genomic profiling remains a major challenge. Moreover, it is increasingly evident that genes do not act as individual units but collaborate in overlapping networks, the deregulation of which is a hallmark of cancer. Thus, we have here applied refined network knowledge to the analysis of key functions and pathways associated with gliomagenesis in a set of 50 human gliomas of various histogenesis, using cDNA microarrays, inferential and descriptive statistics, and dynamic mapping of gene expression data into a functional annotation database. Highest-significance networks were assembled around the myc oncogene in gliomagenesis and around the integrin signaling pathway in the glioblastoma subtype, which is paradigmatic for its strong migratory and invasive behavior. Three novel MYC-interacting genes (UBE2C, EMP1, and FBXW7) with cancer-related functions were identified as network constituents differentially expressed in gliomas, as was CD151 as a new component of a network that mediates glioblastoma cell invasion. Complementary, unsupervised relevance network analysis showed a conserved self-organization of modules of interconnected genes with functions in cell cycle regulation in human gliomas. This approach has extended existing knowledge about the organizational pattern of gene expression in human gliomas and identified potential novel targets for future therapeutic development.

Adult↗