Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Sequence variation database project at the European Bioinformatics Institute.

The sequence variation project at EBI aims to create a unified resource for browsing and searching sequence differences. Technical advances in reading in new data types and in validating and cross-referencing entries are reported. It is suggested that the hardest problems in unifying mutation databases are related to intellectual property rights. The concept of copylefting is introduced as a potential solution to these.

Computational Biology↗

Bioinformatics for glycomics: status, methods, requirements and perspectives.

The term 'glycomics' describes the scientific attempt to identify and study all the glycan molecules - the glycome - synthesised by an organism. The aim is to create a cell-by-cell catalogue of glycosyltransferase expression and detected glycan structures. The current status of databases and bioinformatics tools, which are still in their infancy, is reviewed. The structures of glycans as secondary gene products cannot be easily predicted from the DNA sequence. Glycan sequences cannot be described by a simple linear one-letter code as each pair of monosaccharides can be linked in several ways and branched structures can be formed. Few of the bioinformatics algorithms developed for genomics/proteomics can be directly adapted for glycomics. The development of algorithms, which allow a rapid, automatic interpretation of mass spectra to identify glycan structures is currently the most active field of research. The lack of generally accepted ways to normalise glycan structures and exchange glycan formats hampers an efficient cross-linking and the automatic exchange of distributed data. The upcoming glycomics should accept that unrestricted dissemination of scientific data accelerates scientific findings and initiates a number of new initiatives to explore the data.

Animals↗

FlyRNAi: the Drosophila RNAi screening center database.

RNA interference (RNAi) has become a powerful tool for genetic screening in Drosophila. At the Drosophila RNAi Screening Center (DRSC), we are using a library of over 21,000 double-stranded RNAs targeting known and predicted genes in Drosophila. This library is available for the use of visiting scientists wishing to perform full-genome RNAi screens. The data generated from these screens are collected in the DRSC database (http://flyRNAi.org/cgi-bin/RNAi_screens.pl) in a flexible format for the convenience of the scientist and for archiving data. The long-term goal of this database is to provide annotations for as many of the uncharacterized genes in Drosophila as possible. Data from published screens are available to the public through a highly configurable interface that allows detailed examination of the data and provides access to a number of other databases and bioinformatics tools.

Animals↗

GDR (Genome Database for Rosaceae): integrated web resources for Rosaceae genomics and genetics research.

BACKGROUND: Peach is being developed as a model organism for Rosaceae, an economically important family that includes fruits and ornamental plants such as apple, pear, strawberry, cherry, almond and rose. The genomics and genetics data of peach can play a significant role in the gene discovery and the genetic understanding of related species. The effective utilization of these peach resources, however, requires the development of an integrated and centralized database with associated analysis tools. DESCRIPTION: The Genome Database for Rosaceae (GDR) is a curated and integrated web-based relational database. GDR contains comprehensive data of the genetically anchored peach physical map, an annotated peach EST database, Rosaceae maps and markers and all publicly available Rosaceae sequences. Annotations of ESTs include contig assembly, putative function, simple sequence repeats, and anchored position to the peach physical map where applicable. Our integrated map viewer provides graphical interface to the genetic, transcriptome and physical mapping information. ESTs, BACs and markers can be queried by various categories and the search result sites are linked to the integrated map viewer or to the WebFPC physical map sites. In addition to browsing and querying the database, users can compare their sequences with the annotated GDR sequences via a dedicated sequence similarity server running either the BLAST or FASTA algorithm. To demonstrate the utility of the integrated and fully annotated database and analysis tools, we describe a case study where we anchored Rosaceae sequences to the peach physical and genetic map by sequence similarity. CONCLUSIONS: The GDR has been initiated to meet the major deficiency in Rosaceae genomics and genetics research, namely a centralized web database and bioinformatics tools for data storage, analysis and exchange. GDR can be accessed at http://www.genome.clemson.edu/gdr/.

Computer Graphics↗

2DDB - a bioinformatics solution for analysis of quantitative proteomics data.

BACKGROUND: We present 2DDB, a bioinformatics solution for storage, integration and analysis of quantitative proteomics data. As the data complexity and the rate with which it is produced increases in the proteomics field, the need for flexible analysis software increases. RESULTS: 2DDB is based on a core data model describing fundamentals such as experiment description and identified proteins. The extended data models are built on top of the core data model to capture more specific aspects of the data. A number of public databases and bioinformatical tools have been integrated giving the user access to large amounts of relevant data. A statistical and graphical package, R, is used for statistical and graphical analysis. The current implementation handles quantitative data from 2D gel electrophoresis and multidimensional liquid chromatography/mass spectrometry experiments. CONCLUSION: The software has successfully been employed in a number of projects ranging from quantitative liquid-chromatography-mass spectrometry based analysis of transforming growth factor-beta stimulated fi-broblasts to 2D gel electrophoresis/mass spectrometry analysis of biopsies from human cervix. The software is available for download at SourceForge.

Computational Biology↗

CyanoBase, the genome database for Synechocystis sp. strain PCC6803: status for the year 2000.

CyanoBase provides an online resource for access to data on genomic information about the cyanobacterium Synechocystis sp. strain PCC6803. The database contains annotations for each protein-coding gene deduced from the entire nucleotide sequence of the genome, gene classification lists, and keyword and similarity search engines. Core portions of CyanoBase consist of annotations for each of the 3168 protein genes deduced from the entire nucleotide sequence of this genome. The contents of each gene were improved by updating with the results of similarity searches and by introducing references for analysis in bioinformatics. The database now contains repository facilities that store and provide experimental information, in addition to providing proposals for the function of each gene. This information should help to avoid unnecessary, overlapping experiments and should assist communication between scientists who wish to elucidate the function of putative genes on the cyanobacteria genome. The current URL of CyanoBase is http://www.kazusa.or.jp:8080/cyano/

Cyanobacteria↗

REDfly: a Regulatory Element Database for Drosophila.

Bioinformatics studies of transcriptional regulation in the metazoa are significantly hindered by the absence of readily available data on large numbers of transcriptional cis-regulatory modules (CRMs). Even the richly annotated Drosophila melanogaster genome lacks extensive CRM information. We therefore present here a database of Drosophila CRMs curated from the literature complete with both DNA sequence and a searchable description of the gene expression pattern regulated by each CRM. This resource should greatly facilitate the development of computational approaches to CRM discovery as well as bioinformatics analyses of regulatory sequence properties and evolution.

Animals↗

GLYCOSCIENCES.de: an Internet portal to support glycomics and glycobiology research.

The development of glycan-related databases and bioinformatics applications is considerably lagging behind compared with the wealth of available data and software tools in genomics and proteomics. Because the encoding of glycan structures is more complex, most of the bioinformatics approaches cannot be applied to glycan structures. No standard procedures exist where glycan structures found in various species, organs, tissues or cells can be routinely deposited. In this article the concepts of the GLYCOSCIENCES.de portal are described. It is demonstrated how an efficient structure-based cross-linking of various glycan-related data originating from different resources can be accomplished using a single user interface. The structure oriented retrieval options-exact structure, substructure, motif, composition and sugar components-are discussed. The types of available data-references, composition, spatial structures, nuclear magnetic resonance (NMR) shifts (experimental and estimated), theoretically calculated fragments and Protein Database (PDB) entries-are exemplified for Man(3.) The free availability and unrestricted use of glycan-related data is an absolute prerequisite to efficiently share distributed resources. Additionally, there is an urgent need to agree to a generally accepted exchange format as well as to a common software interface. An open access repository for glyco-related experimental data will secure that the loss of primary data will be considerably reduced.

Computational Biology↗

CBS Genome Atlas Database: a dynamic storage for bioinformatic results and sequence data.

UNLABELLED: Currently, new bacterial genomes are being published on a monthly basis. With the growing amount of genome sequence data, there is a demand for a flexible and easy-to-maintain structure for storing sequence data and results from bioinformatic analysis. More than 150 sequenced bacterial genomes are now available, and comparisons of properties for taxonomically similar organisms are not readily available to many biologists. In addition to the most basic information, such as AT content, chromosome length, tRNA count and rRNA count, a large number of more complex calculations are needed to perform detailed comparative genomics. DNA structural calculations like curvature and stacking energy, DNA compositions like base skews, oligo skews and repeats at the local and global level are just a few of the analysis that are presented on the CBS Genome Atlas Web page. Complex analysis, changing methods and frequent addition of new models are factors that require a dynamic database layout. Using basic tools like the GNU Make system, csh, Perl and MySQL, we have created a flexible database environment for storing and maintaining such results for a collection of complete microbial genomes. Currently, these results counts to more than 220 pieces of information. The backbone of this solution consists of a program package written in Perl, which enables administrators to synchronize and update the database content. The MySQL database has been connected to the CBS web-server via PHP4, to present a dynamic web content for users outside the center. This solution is tightly fitted to existing server infrastructure and the solutions proposed here can perhaps serve as a template for other research groups to solve database issues. AVAILABILITY: A web based user interface which is dynamically linked to the Genome Atlas Database can be accessed via www.cbs.dtu.dk/services/GenomeAtlas/. SUPPLEMENTARY INFORMATION: This paper has a supplemental information page which links to the examples presented: www.cbs.dtu.dk/services/GenomeAtlas/suppl/bioinfdatabase.

Algorithms↗

Cloning, expression and identification of a new trehalose synthase gene from Thermobifida fusca genome.

A new open reading frame in Thermobifida fusca sequenced genome was identified to encode a new trehalose synthase, annotated as "glycosidase" in the GenBank database, by bioinformatics searching and experimental validation. The gene had a length of 1830 bp with about 65% GC content and encoded for a new trehalose synthase with 610 amino acids and deduced molecular weight of 66 kD. The high GC content seemed not to affect its good expression in E. coli BL21 in which the target protein could account for as high as 15% of the total cell proteins. The recombinant enzyme showed its optimal activities at 25 degrees and pH 6.5 when it converted substrate maltose into trehalose. However it would divert a high proportion of its substrate into glucose when the temperature was increased to 37 degrees, or when the enzyme concentration was high Its activity was not inhibited by 5 mM heavy metals such as Cu2+, Mn2+, and Zn2+ but affected by high concentration of glucose. Blasting against the database indicated that amino acid sequence of this protein had maximal 69% homology with the known trehalose synthases, and two highly conserved segments of the protein sequence were identified and their possible linkage with functions was discussed.

Actinomycetales↗

Nonribosomal peptide synthetase genes in the genome of Fusarium graminearum, causative agent of wheat head blight.

Fungal nonribosomal peptide synthetases (NRPSs) are responsible for the biosynthesis of numerous metabolites which serve as virulence factors in several plant-pathogen interactions. The aim of our work was to investigate the diversity of these genes in a Fusarium graminearum sequence database using bioinformatic techniques. Our search identified 15 NRPS sequences, among which two were found to be closely related to peptide synthetases of various fungi taking part in ferrichrome biosynthesis. Another peptide synthetase gene was similar to that identified in Aspergillus oryzae which is possibly responsible for the biosynthesis of fusarinine, an extracellular iron-chelating siderophore. To our knowledge, this is the first report on the identification of a putative NRPS gene possibly responsible for the biosynthesis of fusarinine-type siderophores. The other NRPSs were found to be related to peptide synthetases taking part in the biosynthesis of various peptides in other fungi. Transcription factors carrying ankyrin repeats were observed in the vicinity of four of the identified peptide synthetase genes. Additionally, NRPS related genes similar to putative long-chain fatty acid CoA ligases, acyl CoA ligases, ABC transport proteins, a highly conserved putative transmembrane protein of Aspergillus nidulans, and alpha-aminoadipate reductases have also been identified. Further studies are in progress to clarify the role of some of the identified NRPS genes in plant pathogenesis.

Amino Acid Sequence↗

Identification of a new subfamily of HNH nucleases and experimental characterization of a representative member, HphI restriction endonuclease.

The restriction endonuclease (REase) R. HphI is a Type IIS enzyme that recognizes the asymmetric target DNA sequence 5'-GGTGA-3' and in the presence of Mg(2+) hydrolyzes phosphodiester bonds in both strands of the DNA at a distance of 8 nucleotides towards the 3' side of the target, producing a 1 nucleotide 3'-staggered cut in an unspecified sequence at this position. REases are typically ORFans that exhibit little similarity to each other and to any proteins in the database. However, bioinformatics analyses revealed that R.HphI is a member of a relatively big sequence family with a conserved C-terminal domain and a variable N-terminal domain. We predict that the C-terminal domains of proteins from this family correspond to the nuclease domain of the HNH superfamily rather than to the most common PD-(D/E)XK superfamily of nucleases. We constructed a three-dimensional model of the R.HphI catalytic domain and validated our predictions by site-directed mutagenesis and studies of DNA-binding and catalytic activities of the mutant proteins. We also analyzed the genomic neighborhood of R.HphI homologs and found that putative nucleases accompanied by a DNA methyltransferase (i.e. predicted REases) do not form a single group on a phylogenetic tree, but are dispersed among free-standing putative nucleases. This suggests that nucleases from the HNH superfamily were independently recruited to become REases in the context of RM systems multiple times in the evolution and that members of the HNH superfamily may be much more frequent among the so far unassigned REase sequences than previously thought.

Amino Acid Sequence↗

Functional Analysis of MS-Based Proteomics Data: From Protein Groups to Networks.

Mass spectrometry-based proteomics allows the quantification of thousands of proteins, protein variants, and their modifications, in many biological samples. These are derived from the measurement of peptide relative quantities, and it is not always possible to distinguish proteins with similar sequences due to the absence of protein-specific peptides. In such cases, peptide signals are reported in protein groups that can correspond to several genes. Here, we show that multi-gene protein groups have a limited impact on GO-term enrichment, but selecting only one gene per group affects network analysis. We thus present the Cytoscape app Proteo Visualizer (https://apps.cytoscape.org/apps/ProteoVisualizer) that is designed for retrieving protein interaction networks from STRING using protein groups as input and thus allows visualization and network analysis of bottom-up MS-based proteomics data sets.

Proteomics↗

Targeting SUV4-20H2-mediated H4K20 methylation restrains growth and migration in pediatric high-grade astrocytomas.

Pediatric astrocytomas are characterized by increased molecular and clinical heterogeneity with epigenetic alterations contributing to aggressiveness and therapy resistance. The repressive histone mark H4K20 trimethylation (H4K20me3) and the methyltransferase SUV4-20H2 (KMT5C) are critical regulators of chromatin integrity and genome stability, with limited investigation in pediatric astrocytomas. KMT5C mRNA levels were evaluated in a publicly available pediatric gliomas database using bioinformatic analysis. Investigation of SUV4-20H2 and H4K20me3 expression was performed in a cohort of 43 pediatric astrocytoma tissues by immunohistochemistry. Their functional role and mechanism of action was investigated in pediatric glioma cell lines by using the substrate-competitive inhibitor of SUV4-20, A-196. Cell viability, apoptosis and migration were assessed using XTT, cleaved PARP, and wound healing assays, respectively. Effects of treatment on H4K20 methylation, DNA damage, mitotic stress [Polo-like kinase (PLK1) expression], and invasion markers (N-cadherin, β-catenin expression) were examined by western immunoblotting. KMT5C mRNA was significantly enriched in pediatric high-grade astrocytomas compared to low-grade tumors. A significant elevation of SUV4-20H2 and H4K20me3 expression was detected in astrocytoma tissues indicating epigenetic dysregulation contributing to malignancy. Treatment with A-196 reduced cell proliferation of pediatric glioma cell lines and induced apoptosis in a dose-dependent manner. It further impaired cell migration, accompanied by reduced N-cadherin and β-catenin expression. Mechanistically, inhibition of SUV4-20 depleted H4K20me3, inducing chromatin destabilization, replication-associated DNA damage and was associated with increased PLK1 expression, consistent with activation of a mitotic stress response. Our findings indicate that SUV4-20H2-mediated H4K20 activity in pediatric high-grade astrocytomas maintains their growth and migratory potential by regulating chromatin integrity and may serve as potential therapeutic target.

H4K20me2/3↗

Differential expression of proteins in response to ceramide-mediated stress signal in colon cancer cells by 2-D gel electrophoresis and MALDI-TOF-MS.

Comparative cancer cell proteome analysis is a strategy to study the implication of ceramides in the transmission of stress signals. To better understand the mechanisms by which ceramide regulate some physiological or pathological events and the response to the pharmacological treatment of cancer, we performed a differential analysis of the proteome of HCT-116 (human colon carcinoma) cells in response to these substances. We first established the first 2-dimensional map of the HCT-116 proteome. Then, HCT116 cell proteome treated or not with C6-ceramide have been compared using two-dimensional electrophoresis, matrix-assisted laser desorption/ionization-mass spectrometry and bioinformatic (genomic databases). 2-DE gel analysis revealed more than fourty proteins that were differentially expressed in control cells and cells treated with ceramide. Among them, we confirmed the differential expression of proteins involved in apoptosis and cell adhesion.

Apoptosis↗

Programmatic access to ICTV virus taxonomy through a public ontology API.

BACKGROUND: The International Committee on Taxonomy of Viruses (ICTV) is responsible for developing and maintaining a universal virus taxonomy. As the reference framework for organising the viral world, it is essential for virology and related fields. Despite its widespread use in research and public health, programmatic access to ICTV taxonomy has remained limited, posing challenges for integration, versioning, and interoperability across databases and bioinformatics resources requiring up-to-date virus taxonomy. FINDINGS: To address this, we developed a public and sustainable solution leveraging ontology-based APIs. All available ICTV Master Species List (MSL) releases, from MSL1 to MSL41, were transformed into a unified, semantically structured ontology comprising more than 195,000 current and historical entities and deployed through the Ontology Lookup Service (OLS). The ontology is automatically rebuilt and republished whenever a new MSL release becomes available. Complementary ICTV-NCBI mappings and helper libraries support integration into downstream systems. CONCLUSIONS: Together, these resources enable, for the first time, public programmatic retrieval of current and historical ICTV taxon names, taxonomic relationships, metadata, and persistent identifiers through stable endpoints, including resolution of former taxonomic terms to their current accepted taxon or taxa and retrieval of taxon histories across releases. More broadly, this work illustrates a general strategy for transforming structured biological datasets into semantically enriched graph resources exposed through scalable public APIs. These developments enhance interoperability, reduce manual curation, and support FAIR-aligned taxonomic data management in virology and pandemic preparedness.

API↗

Cambridge Healthtech Institute's Third Annual Conference on human genetic variation. 16-18 October 2000, Philadelphia, Pennsylvania, USA.

A major goal of pharmacogenomics is to identify the human genetic variation that influences susceptibility to complex diseases. Recently, theoretical statistical analyses have suggested that genes for complex diseases may be found by linkage disequilibrium (i.e., association). Single nucleotide polymorphism (SNP) susceptibility alleles for common diseases can occur at high frequencies in various populations and, thus, have a major impact on morbidity and mortality. To be successful, SNP mapping studies require successful teamwork, integrating clinicians, epidemiologists, molecular genetics experts, laboratory automation engineers, bioinformatics and database experts. New statistical methods are also developing rapidly and promise to further increase the power of these studies. A recent conference on human genetic variation provided an opportunity for experts in all of these disciplines to exchange ideas. At present, great technological challenges need to be overcome in order to increase the throughput greatly while lowering cost and still maintaining high accuracy for SNP genotyping. Although this approach is relatively new (at least on the scale now being contemplated), the large payoffs anticipated to accrue from the successful mapping of SNPs in disease genes has led the area to be very strongly supported by both public and private funding sources. The potential payoff for improving disease diagnosis and therapeutic efficacy, with better avoidance of adverse events based on SNP associations, is providing a tremendous incentive to move this effort forward at an ever-accelerating pace.

Genetic Predisposition to Disease↗

Programmatic access to ICTV virus taxonomy through a public ontology API.

The International Committee on Taxonomy of Viruses (ICTV) is responsible for developing and maintaining a universal virus taxonomy. As the reference framework for organising the viral world, it is essential for virology and related fields. Despite its widespread use in research and public health, programmatic access to ICTV taxonomy has remained limited, posing challenges for integration, versioning, and interoperability across databases and bioinformatics resources requiring up-to-date virus taxonomy. To address this, we developed a public and sustainable solution leveraging ontology-based APIs. Successive ICTV Master Species List (MSL) releases were transformed into a structured ontology and deployed as a unified representation through the Ontology Lookup Service (OLS). The framework also provides ICTV-NCBI mappings and helper libraries for integration into downstream systems. This enables, for the first time, public programmatic retrieval of current and historical virological taxon names, taxonomic relationships, metadata, and persistent identifiers through stable endpoints. More broadly, this work illustrates a general strategy for transforming structured biological datasets into semantically enriched graph resources exposed through scalable public APIs. These developments enhance interoperability, reduce manual curation, and support FAIR-aligned taxonomic data management in virology and pandemic preparedness.

API↗