Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Parallelisation of the blast algorithm.

Retrieving homologous DNA and protein sequences from existing databases is a fundamental routine in bioinformatics research. Programs of the NCBI BLAST family are widely used for this purpose. We evaluated paraBLAST, a parallelised version of the NCBI BLAST algorithm, using a Message Passing Interface (MPI) on a multi-node compute cluster. Here, we propose static and dynamic database-partitioning schemes based on the availability of the cluster. We evaluated the application of the algorithm in querying nucleotide sequences against a large-scale sequence database with different numbers of database partitions, and hence, different numbers of CPUs. Since the program's tasks are performed independently of each other, each available CPU can run its own copy of BLAST queries, resulting in reduced interference between processes and leading to a highly scalable solution.

Algorithms↗

SYSTOMONAS--an integrated database for systems biology analysis of Pseudomonas.

To provide an integrated bioinformatics platform for a systems biology approach to the biology of pseudomonads in infection and biotechnology the database SYSTOMONAS (SYSTems biology of pseudOMONAS) was established. Besides our own experimental metabolome, proteome and transcriptome data, various additional predictions of cellular processes, such as gene-regulatory networks were stored. Reconstruction of metabolic networks in SYSTOMONAS was achieved via comparative genomics. Broad data integration is realized using SOAP interfaces for the well established databases BRENDA, KEGG and PRODORIC. Several tools for the analysis of stored data and for the visualization of the corresponding results are provided, enabling a quick understanding of metabolic pathways, genomic arrangements or promoter structures of interest. The focus of SYSTOMONAS is on pseudomonads and in particular Pseudomonas aeruginosa, an opportunistic human pathogen. With this database we would like to encourage the Pseudomonas community to elucidate cellular processes of interest using an integrated systems biology strategy. The database is accessible at http://www.systomonas.de.

Bacterial Proteins↗

Mining thermophile photosynthesis genes: a synthetic operon expressing Chloroflexota species reaction center genes in Rhodobacter sphaeroides.

Photosynthesis is the foundation of the vast majority of life systems, and therefore the most important bioenergetic process on earth, and the greatest diversity in photosynthetic systems are found in microorganisms. However, understanding of the biophysical and biochemical processes that transduce light to chemical energy has derived from the relatively small subset of proteins from microbes that are amenable to cultivation, in contrast to the huge number of microbial DNA sequences encoding proteins that catalyze the initial photochemical reactions that has been deposited in databases, such as from metagenomics. We describe the use of a Rhodobacter sphaeroides laboratory strain for expression of heterologous photosynthesis genes to demonstrate the feasibility of mining this resource, focusing on hot spring Chloroflexota gene sequences. Using a synthetic operon of genes, we produced a photochemically active complex of reaction center proteins in our biological system. We also present bioinformatic analyses of anoxygenic type II reaction center sequences from metagenomic samples collected from hot (42-90° C) springs available through the JGI IMG database, to generate a resource of diverse sequences that potentially are adapted to photosynthesis at such temperatures. These data provide a view into the natural diversity of anoxygenic photosynthesis, through a lens focused on high-temperature environments. The approach we took to express such genes can be applied for potential biotechnology purposes as well as for studies of fundamental catalytic properties of these heretofore inaccessible protein complexes.

Chloroflexota↗

A protocol for maintaining multidatabase referential integrity.

The bioinformatics community is becoming increasingly reliant on the creation of links among biological databases (DBs) as a foundation for DB interoperability. For example, a link might be created from a protein in one DB (such as PIR), to a gene in another DB (such as GDB), by storing the unique identifier (id) of the gene object within an attribute of the protein object. User interfaces can then support navigation from the protein to the gene, and multiDB queries can join the protein with the gene. The unique id of the gene is serving as a foreign key. However, a variety of factors, such as changes in the underlying biology, can cause object ids to become invalid, thus producing invalid links among DBs. Invalid links are a violation of multidatabase referential integrity. We propose a network protocol whereby a database administrator can provide information about changes to the identifiers of objects in their database via Internet, to allow other databases to maintain referential integrity. We request comments from the bioinformatics community for the purpose of building a consensus on the proposed protocol.

Computational Biology↗

Functional inferences from reconstructed evolutionary biology involving rectified databases--an evolutionarily grounded approach to functional genomics.

If bioinformatics tools are constructed to reproduce the natural, evolutionary history of the biosphere, they offer powerful approaches to some of the most difficult tasks in genomics, including the organization and retrieval of sequence data, the updating of massive genomic databases, the detection of database error, the assignment of introns, the prediction of protein conformation from protein sequences, the detection of distant homologs, the assignment of function to open reading frames, the identification of biochemical pathways from genomic data, and the construction of a comprehensive model correlating the history of biomolecules with the history of planet Earth.

Amino Acid Sequence↗

EyeSite: a semi-automated database of protein families in the eye.

The EyeSite is a web-based database of protein families for proteins that function in the eye and their homologous sequences. The resource clusters proteins at different levels of homology in order to facilitate functional annotation of sequences and modelling of proteins from structural homologues. Eye proteins are organized into the tissue types in which they function and are clustered into homologous families using a novel protocol employing the TribeMCL algorithm. Homologous families are further subdivided into sequence clusters for which multiple sequence alignments are generated. Structural annotations from the CATH domain database are provided for nearly 90% of the sequences, and protein family annotations from the Pfam database for approximately 86%. Homology models have also been generated where appropriate. The EyeSite is stored in a relational database and is extensively linked to other online bioinformatics resources to help relate allelic variants, annotations and clinical details to the derived data in the database. The EyeSite is available for online search, sequence information and model retrieval at http://eyesite.cryst.bbk.ac.uk/.

Amino Acid Sequence↗

EMBL Nucleotide Sequence Database in 2006.

The EMBL Nucleotide Sequence Database (http://www.ebi.ac.uk/embl) at the EMBL European Bioinformatics Institute, UK, offers a large and freely accessible collection of nucleotide sequences and accompanying annotation. The database is maintained in collaboration with DDBJ and GenBank. Data are exchanged between the collaborating databases on a daily basis to achieve optimal synchrony. Webin is the preferred tool for individual submissions of nucleotide sequences, including Third Party Annotation, alignments and bulk data. Automated procedures are provided for submissions from large-scale sequencing projects and data from the European Patent Office. In 2006, the volume of data has continued to grow exponentially. Access to the data is provided via SRS, ftp and variety of other methods. Extensive external and internal cross-references enable users to search for related information across other databases and within the database. All available resources can be accessed via the EBI home page at http://www.ebi.ac.uk/. Changes over the past year include changes to the file format, further development of the EMBLCDS dataset and developments to the XML format.

Base Sequence↗

High-throughput protein analysis integrating bioinformatics and experimental assays.

The wealth of transcript information that has been made publicly available in recent years requires the development of high-throughput functional genomics and proteomics approaches for its analysis. Such approaches need suitable data integration procedures and a high level of automation in order to gain maximum benefit from the results generated. We have designed an automatic pipeline to analyse annotated open reading frames (ORFs) stemming from full-length cDNAs produced mainly by the German cDNA Consortium. The ORFs are cloned into expression vectors for use in large-scale assays such as the determination of subcellular protein localization or kinase reaction specificity. Additionally, all identified ORFs undergo exhaustive bioinformatic analysis such as similarity searches, protein domain architecture determination and prediction of physicochemical characteristics and secondary structure, using a wide variety of bioinformatic methods in combination with the most up-to-date public databases (e.g. PRINTS, BLOCKS, INTERPRO, PROSITE SWISSPROT). Data from experimental results and from the bioinformatic analysis are integrated and stored in a relational database (MS SQL-Server), which makes it possible for researchers to find answers to biological questions easily, thereby speeding up the selection of targets for further analysis. The designed pipeline constitutes a new automatic approach to obtaining and administrating relevant biological data from high-throughput investigations of cDNAs in order to systematically identify and characterize novel genes, as well as to comprehensively describe the function of the encoded proteins.

Automation↗

rSNP_Guide: an integrated database-tools system for studying SNPs and site-directed mutations in transcription factor binding sites.

Since the human genome was sequenced in draft, single nucleotide polymorphism (SNP) analysis has become one of the keynote fields of bioinformatics. We have developed an integrated database-tools system, rSNP_Guide (http://wwwmgs.bionet.nsc.ru/mgs/systems/rsnp/), devoted to prediction of transcription factor (TF) binding sites, alterations of which could be associated with disease phenotype. By inputting data on alterations in DNA sequence and in DNA binding pattern of an unknown TF, rSNP_Guide searches for a known TF with alterations in the recognition score calculated on the basis of TF site's sequence and consistent with the input alterations in DNA binding to the unknown TF. Our system has been tested on many relationships between known TF sites and diseases, as well as on site-directed mutagenesis data. Experimental verification of rSNP_Guide system was made on functionally important SNPs in human TDO2and mouse K-ras genes. Additional examples of analysis are reported involving variants in the human gammaA-globin (HBG1), hsp70(HSPA1A), and Factor IX (F9) gene promoters.

Animals↗

Mass spectrometry-based metabolomics.

This review presents an overview of the dynamically developing field of mass spectrometry-based metabolomics. Metabolomics aims at the comprehensive and quantitative analysis of wide arrays of metabolites in biological samples. These numerous analytes have very diverse physico-chemical properties and occur at different abundance levels. Consequently, comprehensive metabolomics investigations are primarily a challenge for analytical chemistry and specifically mass spectrometry has vast potential as a tool for this type of investigation. Metabolomics require special approaches for sample preparation, separation, and mass spectrometric analysis. Current examples of those approaches are described in this review. It primarily focuses on metabolic fingerprinting, a technique that analyzes all detectable analytes in a given sample with subsequent classification of samples and identification of differentially expressed metabolites, which define the sample classes. To perform this complex task, data analysis tools, metabolite libraries, and databases are required. Therefore, recent advances in metabolomics bioinformatics are also discussed.

Databases, Protein↗

The RESID Database of protein structure modifications and the NRL-3D Sequence-Structure Database.

The RESID Database is a comprehensive collection of annotations and structures for protein post-translational modifications including N-terminal, C-terminal and peptide chain cross-link modifications. The RESID Database includes systematic and frequently observed alternate names, Chemical Abstracts Service registry numbers, atomic formulas and weights, enzyme activities, taxonomic range, keywords, literature citations with database cross-references, structural diagrams and molecular models. The NRL-3D Sequence-Structure Database is derived from the three-dimensional structure of proteins deposited with the Research Collaboratory for Structural Bioinformatics Protein Data Bank. The NRL-3D Database includes standardized and frequently observed alternate names, sources, keywords, literature citations, experimental conditions and searchable sequences from model coordinates. These databases are freely accessible through the National Cancer Institute-Frederick Advanced Biomedical Computing Center at these web sites: http://www. ncifcrf.gov/RESID, http://www.ncifcrf.gov/NRL-3D; or at these National Biomedical Research Foundation Protein Information Resource web sites: http://pir.georgetown.edu/pirwww/dbinfo/resid .html, http://pir.georgetown.edu/pirwww/dbinfo/nrl3d .html

Amino Acids↗

Dynamic tables: an architecture for managing evolving, heterogeneous biomedical data in relational database management systems.

Data sparsity and schema evolution issues affecting clinical informatics and bioinformatics communities have led to the adoption of vertical or object-attribute-value-based database schemas to overcome limitations posed when using conventional relational database technology. This paper explores these issues and discusses why biomedical data are difficult to model using conventional relational techniques. The authors propose a solution to these obstacles based on a relational database engine using a sparse, column-store architecture. The authors provide benchmarks comparing the performance of queries and schema-modification operations using three different strategies: (1) the standard conventional relational design; (2) past approaches used by biomedical informatics researchers; and (3) their sparse, column-store architecture. The performance results show that their architecture is a promising technique for storing and processing many types of data that are not handled well by the other two semantic data models.

Computational Biology↗

Detection of novel gene expression in paraffin-embedded tissues by isotopic in situ hybridization in tissue microarrays.

Correlating altered gene expression patterns with particular disease states is a critical step in understanding disease processes and developing treatment strategies. Many thousands of novel gene sequences have recently been annotated in public and private databases and are now available for analysis. Tissue-specific expression patterns of these sequences can be evaluated physically on DNA arrays and other high throughput assays, or virtually by bioinformatics mining of expressed sequence tag (EST) databases. As a secondary screening tool, in situ hybridisation (ISH) not only confirms tissue specificity, but also reveals what is often valuable information about cell-type expression patterns of nov16l sequences. Due to their availability and long-term stability at room temperature, formalin-fixed paraffin-embedded clinical specimens provide an invaluable resource for evaluating expression patterns of novel human genes. We describe a high-throughput approach for identifying and quantifying the expression of novel genes in paraffin-embedded human tissues using isotopic in situ hybridisation and tissue microarrays (TMA).

Blotting, Northern↗

Information management for the study of allergies.

Microarrays and other large-scale screening technologies produce quantities of increasingly complex allergy data. These data link molecular and clinical measurements and observations and provide fertile ground for improving our understanding of the processes involved in allergic reactions. Information technology is employed in gathering, storage, retrieval and analysis of these data. The increasing proportion of allergy data are generated from genomics and proteomics approaches. The major activity focuses on characterization of allergens including IgE reactivity, structural properties, and mapping of IgE and T-cell epitopes. Because of the complexity of allergy data, their utilization requires bioinformatics approaches. Allergen data are stored in the general and specialist databases. At least a dozen of important allergen databases and data repositories have been developed to date. These data are analysed using general and specialist bioinformatics tools. The major applications of bioinformatics include support for allergen characterization, assessment of allergenicity, and identification of allergic cross-reactivity. These applications in turn support the development of vaccines and therapies for allergic disease. In this article we review allergen databases and tools for the analysis of allergens, and discuss the new directions in the field supported by large scale screening involving genomics, proteomics, and bioinformatics support.

Allergens↗

The EMBL Nucleotide Sequence Database.

The EMBL Nucleotide Sequence Database (http://www.ebi.ac.uk/embl), maintained at the European Bioinformatics Institute (EBI) near Cambridge, UK, is a comprehensive collection of nucleotide sequences and annotation from available public sources. The database is part of an international collaboration with DDBJ (Japan) and GenBank (USA). Data are exchanged daily between the collaborating institutes to achieve swift synchrony. Webin is the preferred tool for individual submissions of nucleotide sequences, including Third Party Annotation (TPA) and alignments. Automated procedures are provided for submissions from large-scale sequencing projects and data from the European Patent Office. New and updated data records are distributed daily and the whole EMBL Nucleotide Sequence Database is released four times a year. Access to the sequence data is provided via ftp and several WWW interfaces. With the web-based Sequence Retrieval System (SRS) it is also possible to link nucleotide data to other specialist molecular biology databases maintained at the EBI. Other tools are available for sequence similarity searching (e.g. FASTA and BLAST). Changes over the past year include the removal of the sequence length limit, the launch of the EMBLCDSs dataset, extension of the Sequence Version Archive functionality and the revision of quality rules for TPA data.

Base Sequence↗

The power of an integrated informatic and molecular approach to type 1 diabetes research.

Recent years have witnessed an explosive growth in available biological data. This includes a tremendous quantity of sequence data (e.g., biological structures, genetic and physical maps, pathways) generated by genome and transcriptome projects focused on humans, mice, and a multitude of other species. Diabetes research stands to greatly benefit from this data, which is distributed across public and private databases and the scientific literature. The increasing quantity and complexity of this biological data necessitates use of novel bioinformatics strategies for its efficient retrieval, analysis, and interpretation. Bioinformatic capability is becoming increasingly indispensable for fast and comprehensive analysis of biological data by diabetes researchers. There is great potential for diabetes scientists and clinicians to take advantage of recent bioinformatics and knowledge discovery developments to radically transform and advance this field of research. This paper will review advances in the field of bioinformatics relevant to diabetes research and preview a new specialty diabetes database, Diabetaeta, that we are creating to serve as a central bioinformatic portal for type 1 diabetes research, as well as serving as a public repository for beta cell gene and protein expression data.

Animals↗

The PRINTS protein fingerprint database in its fifth year.

PRINTS is a database of protein family 'fingerprints' offering a diagnostic resource for newly-determined sequences. By contrast with PROSITE, which uses single consensus expressions to characterise particular families, PRINTS exploits groups of motifs to build characteristic signatures. These signatures offer improved diagnostic reliability by virtue of the mutual context provided by motif neighbours. To date, 800 fingerprints have been constructed and stored in PRINTS. The current version, 17.0, encodes approximately 4500 motifs, covering a range of globular and membrane proteins, modular polypeptides, and so on. The database is accessible via the UCL Bioinformatics World Wide Web (WWW) Server at http://www. biochem.ucl.ac.uk/bsm/dbbrowser/ . We have recently enhanced the usefulness of PRINTS by making available new, intuitive search software. This allows both individual query sequence and bulk data submission, permitting easy analysis of single sequences or complete genomes. Preliminary results indicate that use of the PRINTS system is able to assign additional functions not found by other methods, and hence offers a useful adjunct to current genome analysis protocols.

Animals↗

BiologicalNetworks: visualization and analysis tool for systems biology.

Systems level investigation of genomic scale information requires the development of truly integrated databases dealing with heterogeneous data, which can be queried for simple properties of genes or other database objects as well as for complex network level properties, for the analysis and modelling of complex biological processes. Towards that goal, we recently constructed PathSys, a data integration platform for systems biology, which provides dynamic integration over a diverse set of databases [Baitaluk et al. (2006) BMC Bioinformatics 7, 55]. Here we describe a server, BiologicalNetworks, which provides visualization, analysis services and an information management framework over PathSys. The server allows easy retrieval, construction and visualization of complex biological networks, including genome-scale integrated networks of protein-protein, protein-DNA and genetic interactions. Most importantly, BiologicalNetworks addresses the need for systematic presentation and analysis of high-throughput expression data by mapping and analysis of expression profiles of genes or proteins simultaneously on to regulatory, metabolic and cellular networks. BiologicalNetworks Server is available at http://brak.sdsc.edu/pub/BiologicalNetworks.

Computer Graphics↗