Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Identification of trace element-containing proteins in genomic databases.

Development of bioinformatics tools provided researchers with the ability to identify full sets of trace element-containing proteins in organisms for which complete genomic sequences are available. Recently, independent bioinformatics methods were used to identify all, or almost all, genes encoding selenocysteine-containing proteins in human, mouse, and Drosophila genomes, characterizing entire selenoproteomes in these organisms. It also should be possible to search for entire sets of other trace element-associated proteins, such as metal-containing proteins, although methods for their identification are still in development.

Animals↗

The role of informatics in glycobiology research with special emphasis on automatic interpretation of MS spectra.

This paper reviews the current status of bioinformatics applications and databases in glycobiology, which are based on bioinformatics approaches as well as informatics for glycobiology where an explicit encoding of glycan structures is required. The availability of the complete sequence of the human genome has accelerated the systematic identification of so far unidentified glycogenes considerably in many areas of glycobiology using well-established bioinfomatics tools. Although there has been an immense development of new glyco-related data collections as well as informatics tools and several efforts have been started to cross-link and reference the various data deposited in distributed databases, informatics for glycobiology and glycomics is still poorly developed compared to the genomics and proteomics area. The development of algorithms for the automatic interpretation of MS spectra - currently, a severe bottleneck, which hampers the rapid and reliable interpretation of MS data in high-throughput glycomics projects - is reviewed. A comprehensive list of web resources is given. Several lines of progression are discussed. There is an urgent need for the development of decentralised input facilities of experimentally determined glycan structures. Simultaneously, agreements of standards for the structural description of glycans as well as formats for the related data have to be established. The integration of glycomics with genomics/proteomics has to increase.

Computational Biology↗

Web-accessible proteome databases for microbial research.

The analysis of proteomes of biological organisms represents a major challenge of the post-genome era. Classical proteomics combines two-dimensional electrophoresis (2-DE) and mass spectrometry (MS) for the identification of proteins. Novel technologies such as isotope coded affinity tag (ICAT)-liquid chromatography/mass spectrometry (LC/MS) open new insights into protein alterations. The vast amount and diverse types of proteomic data require adequate web-accessible computational and database technologies for storage, integration, dissemination, analysis and visualization. A proteome database system (http://www.mpiib-berlin.mpg.de/2D-PAGE) for microbial research has been constructed which integrates 2-DE/MS, ICAT-LC/MS and functional classification data of proteins with genomic, metabolic and other biological knowledge sources. The two-dimensional polyacrylamide gel electrophoresis database delivers experimental data on microbial proteins including mass spectra for the validation of protein identification. The ICAT-LC/MS database comprises experimental data for protein alterations of mycobacterial strains BCG vs. H37Rv. By formulating complex queries within a functional protein classification database "FUNC_CLASS" for Mycobacterium tuberculosis and Helicobacter pylori the researcher can gather precise information on genes, proteins, protein classes and metabolic pathways. The use of the R language in the database architecture allows high-level data analysis and visualization to be performed "on-the-fly". The database system is centrally administrated, and investigators without specific bioinformatic competence in database construction can submit their data. The database system also serves as a template for a prototype of a European Proteome Database of Pathogenic Bacteria. Currently, the database system includes proteome information for six strains of microorganisms.

Bacterial Proteins↗

The EMBL Nucleotide Sequence Database.

The EMBL Nucleotide Sequence Database (http://www.ebi.ac.uk/embl.html) constitutes Europe's primary nucleotide sequence resource. Main sources for DNA and RNA sequences are direct submissions from individual researchers, genome sequencing projects and patent applications. While automatic procedures allow incorporation of sequence data from large-scale genome sequencing centres and from the European Patent Office (EPO), the preferred submission tool for individual submitters is Webin (WWW). Through all stages, dataflow is monitored by EBI biologists communicating with the sequencing groups. In collaboration with DDBJ and GenBank the database is produced, maintained and distributed at the European Bioinformatics Institute (EBI). Database releases are produced quarterly and are distributed on CD-ROM. Network services allow access to the most up-to-date data collection via Internet and World Wide Web interface. EBI's Sequence Retrieval System (SRS) is a Network Browser for Databanks in Molecular Biology, integrating and linking the main nucleotide and protein databases, plus many specialised databases. For sequence similarity searching a variety of tools (e.g. Blitz, Fasta, Blast etc) are available for external users to compare their own sequences against the most currently available data in the EMBL Nucleotide Sequence Database and SWISS-PROT.

Amino Acid Sequence↗

An integrated biomedical knowledge extraction and analysis platform: using federated search and document clustering technology.

High content screening (HCS) requires time-consuming and often complex iterative information retrieval and assessment approaches to optimally conduct drug discovery programs and biomedical research. Pre- and post-HCS experimentation both require the retrieval of information from public as well as proprietary literature in addition to structured information assets such as compound libraries and projects databases. Unfortunately, this information is typically scattered across a plethora of proprietary bioinformatics tools and databases and public domain sources. Consequently, single search requests must be presented to each information repository, forcing the results to be manually integrated for a meaningful result set. Furthermore, these bioinformatics tools and data repositories are becoming increasingly complex to use; typically they fail to allow for more natural query interfaces. Vivisimo has developed an enterprise software platform to bridge disparate silos of information. The platform automatically categorizes search results into descriptive folders without the use of taxonomies to drive the categorization. A new approach to information retrieval for HCS experimentation is proposed.

Biomedical Research↗

Bioinformatics of proteases in the MEROPS database.

Proteolytic enzymes represent approximately approximately 2% of the total number of proteins present in all types of organisms. Many of these enzymes are of medical importance, and those that are of potential interest as drug targets can be divided into the endogenous enzymes encoded in the human genome, and the exogenous proteases encoded in the genomes of disease-causing organisms. There are also naturally occurring inhibitors of proteases, some of which have pharmaceutical relevance. The MEROPS database provides a rich source of information on proteases and their inhibitors. Storage and retrieval of this information is facilitated by the use of a hierarchical classification system (which was pioneered by the compilers of the database) in which homologous proteases and their inhibitors are divided into clans and families.

Animals↗

Identification and interrogation of highly informative single nucleotide polymorphism sets defined by bacterial multilocus sequence typing databases.

A unified, bioinformatics-driven, single nucleotide polymorphism (SNP)-based approach to microbial genotyping has been developed. Multilocus sequence typing (MLST) databases consist of known variants of standardized housekeeping genes. Normally, seven fragments are defined; a sequence type (ST) consists of the variants of these fragments that are found in a particular isolate. A computer program that can identify highly informative sets of SNPs in entire MLST databases has been constructed. The SNPs either define a particular user-specified ST or provide a high value for Simpson's index of diversity (D), and may thus be generally applicable to that species. SNP sets that are diagnostic for Neisseria meningitidis ST-11 and ST-42, and high-D SNP sets for N. meningitidis and Staphylococcus aureus, were identified and real-time PCR methods to interrogate these SNPs were demonstrated. High-D SNP sets were also identified in other MLST databases. This widely applicable approach allows rapid genetic fingerprinting of infectious agents.

Algorithms↗

GeneX Va: VBC open source microarray database and analysis software.

Developed by the Virginia Bioinformatics Consortium (VBC), GeneX Va is an open source, freeware database and bioinformatics analysis software for archiving and analyzing Affymetrix GeneChip data. It provides an integrated framework for management, documentation, and analysis of microarray experiments and data to support a range of users, from individual research laboratories to institutional microarray facilities. GeneX Va also provides web-based access to a PostgreSQL relational database system with a comprehensive security system. Data can be extracted from the database and delivered to interactive or scriptable statistical analysis protocols. The security system allows each investigator to manage their own array data and analysis output files and also provides custom access privileges for other users, groups, and internal/external collaborators. The analysis interface uses "Analysis Trees," an innovative user interface that allows researchers to interactively create a tree-structured flow chart of analysis routines. The latest GeneX Va software is available from and can be freely downloaded at the Sourceforge web site http://va-genex.sourceforge.net. To allow researchers to access the database and analysis capabilities of the GeneX Va system, microarray data from many VBC GeneChip experiments have been deposited into a public section of the GeneX Va system at the University of Virginia. The VBC GeneX Va sites, which include documentation, are at http://genes.med.virginia.edu/ of the University of Virginia and at http://genex.csbc.vcu.edu/ of the Virginia Commonwealth University.

Computer Security↗

E-MSD: an integrated data resource for bioinformatics.

The Macromolecular Structure Database (MSD) group (http://www.ebi.ac.uk/msd/) continues to enhance the quality and consistency of macromolecular structure data in the worldwide Protein Data Bank (wwPDB) and to work towards the integration of various bioinformatics data resources. One of the major obstacles to the improved integration of structural databases such as MSD and sequence databases like UniProt is the absence of up to date and well-maintained mapping between corresponding entries. We have worked closely with the UniProt group at the EBI to clean up the taxonomy and sequence cross-reference information in the MSD and UniProt databases. This information is vital for the reliable integration of the sequence family databases such as Pfam and Interpro with the structure-oriented databases of SCOP and CATH. This information has been made available to the eFamily group (http://www.efamily.org.uk/) and now forms the basis of the regular interchange of information between the member databases (MSD, UniProt, Pfam, Interpro, SCOP and CATH). This exchange of annotation information has enriched the structural information in the MSD database with annotation from wider sequence-oriented resources. This work was carried out under the 'Structure Integration with Function, Taxonomy and Sequences (SIFTS)' initiative (http://www.ebi.ac.uk/msd-srv/docs/sifts) in the MSD group.

Amino Acid Sequence↗

Farm animal genome databases.

The requirements for bioinformatics resources to support genome research in farm animals is reviewed. The resources developed to meet these needs are described. Resource databases and associated tools have been developed to handle experimental data. Several of these systems serve the needs of multinational collaborations. Genome databases have been established to provide contemporary summaries of the status of genome maps in a range of farm and domestic animals along with experimental details and citations. New resources and tools will be required to address the informatics needs of emerging technologies such as microarrays. However, continued investment is also required to maintain the currency and utility of the current systems, especially the genome databases.

Animals↗

The role of transglutaminase-2 and its substrates in human diseases.

The most characteristic enzymatic function of the class of enzymes known as transglutaminases (TG, EC 2.3.2.13) is the formation of covalent bonds between epsilon-amino groups of primary amines (from lysines or others) and the gamma-carboxamine group of glutamine residues of proteins. In the last years, a growing body of evidence indicate that the most interesting member of the TG family, namely the tissue TG (tTG, also called transglutaminase type 2, TG2), possesses more than one catalytic function. In fact, TG2 is able to catalyze a crosslinking reaction, a deamidation reaction and also shows GTP-binding/hydrolyzing and isopeptidase activities. Therefore, it can act on several classes of substrates, ranging from proteins to peptides, small reactive molecules like mono- and polyamines, and nucleotides. Given the broad spectrum of potentially different activities, elucidating the role of TG2 and its substrates in cellular functions and human diseases is a difficult task. In this study we focus our attention on substrates of TG2 and report a number of interesting considerations about their possible interplay in biological processes and involvement in human diseases, including genetic disorders. A significant improvement in understanding this complex scenario may come from a "multi-interfaced" approach, by exploiting different bioinformatic tools. Starting from a database of known TG2 substrates and using bioinformatic cross-search among other databases, we generated relational tables from which an involvement of TG2 in several genetic disorders can be hypothesized. Developing new bioinformatic tools and strategies to investigate the role of TG2 in molecular mechanisms underlying human diseases will add new light to this fascinating field of research.

Autoimmune Diseases↗

MitBASE: a comprehensive and integrated mitochondrial DNA database.

MitBASE is an integrated and comprehensive database of mitochondrial DNA data which collects all available information from different organisms and from intraspecie variants and mutants. Research institutions from different countries are involved, each in charge of developing, collecting and annotating data for the organisms they are specialised in. The design of the actual structure of the database and its implementation in a user-friendly format are the care of the European Bioinformatics Institute. The database can be accessed on the Web at the following address: http://www.ebi.ac. uk/htbin/Mitbase/mitbase.pl. The impact of this project is intended for both basic and applied research. The study of mitochondrial genetic diseases and mitochondrial DNA intraspecie diversity are key topics in several biotechnological fields. The database has been funded within the EU Biotechnology programme.

Animals↗

Bioinformatic methods for allergenicity assessment using a comprehensive allergen database.

BACKGROUND: A principal aim of the safety assessment of genetically modified crops is to prevent the introduction of known or clinically cross-reactive allergens. Current bioinformatic tools and a database of allergens and gliadins were tested for the ability to identify potential allergens by analyzing 6 Bacillus thuringiensis insecticidal proteins, 3 common non-allergenic food proteins and 50 randomly selected corn (Zea mays) proteins. METHODS: Protein sequences were compared to allergens using the FASTA algorithm and by searching for matches of 6, 7 or 8 contiguous identical amino acids. RESULTS: No significant sequence similarities or matches of 8 contiguous amino acids were found with the B. thuringiensis or food proteins. Surprisingly, 41 of 50 corn proteins matched at least one allergen with 6 contiguous identical amino acids. Only 7 of 50 corn proteins matched an allergen with 8 contiguous identical amino acids. When assessed for overall structural similarity to allergens, these 7 plus 2 additional corn proteins shared >or=35% identity in an overlap of >or=80 amino acids, but only 6 of the 7 were similar across the length of the protein, or shared >50% identity to an allergen. CONCLUSIONS: An evaluation of a protein by the FASTA algorithm is the most predictive of a clinically relevant cross-reactive allergen. An additional search for matches of 8 amino acids may provide an added margin of safety when assessing the potential allergenicity of a protein, but a search with a 6-amino-acid window produces many random, irrelevant matches.

Algorithms↗

E-MSD: an integrated data resource for bioinformatics.

The Macromolecular Structure Database (MSD) group (http://www.ebi.ac.uk/msd/) continues to enhance the quality and consistency of macromolecular structure data in the Protein Data Bank (PDB) and to work towards the integration of various bioinformatics data resources. We have implemented a simple form-based interface that allows users to query the MSD directly. The MSD 'atlas pages' show all of the information in the MSD for a particular PDB entry. The group has designed new search interfaces aimed at specific areas of interest, such as the environment of ligands and the secondary structures of proteins. We have also implemented a novel search interface that begins to integrate separate MSD search services in a single graphical tool. We have worked closely with collaborators to build a new visualization tool that can present both structure and sequence data in a unified interface, and this data viewer is now used throughout the MSD services for the visualization and presentation of search results. Examples showcasing the functionality and power of these tools are available from tutorial webpages (http://www. ebi.ac.uk/msd-srv/docs/roadshow_tutorial/).

Algorithms↗

Mapping XML documents into databases: a Data-Driven Framework for bioinformatic data interchange.

The Data-Driven Framework (DDF) described here addresses two major problems for healthcare Electronic Data Interchange, data formats and software development costs. The use of a standard XML Document Type Definition (DTD) allows robust representation in any application area and leverages industry-standard tools and development directions. The DDF allows reduced software development and maintenance costs since all data-entry and database tools are generated from the DTD. The DTD can change and the tools can be regenerated. The case-study below uses the DDF for reporting cell assays to determine the roles of factors influencing cellular gene expression and regulation.

Computer Communication Networks↗

Advancing glycomics: implementation strategies at the consortium for functional glycomics.

Glycomics-an integrated approach to study structure-function relationships of complex carbohydrates (or glycans)-is an emerging field in this age of post-genomics. Realizing the importance of glycomics, many large scale research initiatives have been established to generate novel resources and technologies to advance glycomics. These initiatives are generating and cataloging diverse data sets necessitating the development of bioinformatic platforms to acquire, integrate, and disseminate these data sets in a meaningful fashion. With the consortium for functional glycomics (CFG) as the model system, this review discusses databases and the bioinformatics platform developed by this consortium to advance glycomics.

Animals↗

Databases, models, and algorithms for functional genomics: a bioinformatics perspective.

A variety of patterns have been observed on the DNA and protein sequences that serve as control points for gene expression and cellular functions. Owing to the vital role of such patterns discovered on biological sequences, they are generally cataloged and maintained within internationally shared databases. Furthermore,the variability in a family of observed patterns is often represented using computational models in order to facilitate their search within an uncharacterized biological sequence. As the biological data is comprised of a mosaic of sequence-levels motifs, it is significant to unravel the synergies of macromolecular coordination utilized in cell-specific differential synthesis of proteins. This article provides an overview of the various pattern representation methodologies and the surveys the pattern databases available for use to the molecular biologists. Our aim is to describe the principles behind the computational modeling and analysis techniques utilized in bioinformatics research, with the objective of providing insight necessary to better understand and effectively utilize the available databases and analysis tools. We also provide a detailed review of DNA sequence level patterns responsible for structural conformations within the Scaffold or Matrix Attachment Regions (S/MARs).

Animals↗