Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

The apoptosis database.

The apoptosis database is a public resource for researchers and students interested in the molecular biology of apoptosis. The resource provides functional annotation, literature references, diagrams/images, and alternative nomenclatures on a set of proteins having 'apoptotic domains'. These are the distinctive domains that are often, if not exclusively, found in proteins involved in apoptosis. The initial choice of proteins to be included is defined by apoptosis experts and bioinformatics tools. Users can browse through the web accessible lists of domains, proteins containing these domains and their associated homologs. The database can also be searched by sequence homology using basic local alignment search tool, text word matches of the annotation, and identifiers for specific records. The resource is available at http://www.apoptosis-db.org and is updated on a regular basis.

Animals↗

CHOMPER: a bioinformatic tool for rapid validation of tandem mass spectrometry search results associated with high-throughput proteomic strategies.

Current efforts aimed at developing high-throughput proteomics focus on increasing the speed of protein identification. Although improvements in sample separation, enrichment, automated handling, mass spectrometric analysis, as well as data reduction and database interrogation strategies have done much to increase the quality, quantity and efficiency of data collection, significant bottlenecks still exist. Various separation techniques have been coupled with tandem mass spectrometric (MS/MS) approaches to allow a quicker analysis of complex mixtures of proteins, especially where a high number of unambiguous protein identifications are the exception, rather than the rule. MS/MS is required to provide structural / amino acid sequence information on a peptide and thus allow protein identity to be inferred from individual peptides. Currently these spectra need to be manually validated because: (a) the potential of false positive matches i.e., protein not in database, and (b) observed fragmentation trends may not be incorporated into current MS/MS search algorithms. This validation represents a significant bottleneck associated with high-throughput proteomic strategies. We have developed CHOMPER, a software program which reduces the time required to both visualize and confirm MS/MS search results and generate post-analysis reports and protein summary tables. CHOMPER extracts the identification information from SEQUEST MS/MS search result files, reproduces both the peptide and protein identification summaries, provides a more interactive visualization of the MS/MS spectra and facilitates the direct submission of manually validated identifications to a database.

Algorithms↗

Analysis of gene expression profiles: an application of memetic algorithms to the minimum sum-of-squares clustering problem.

Microarrays have become a key technology in experimental molecular biology since they allow monitoring of gene expression for more than 10,000 genes in parallel producing huge amounts of data. In the exploration of transcriptional regulatory networks, an important task is to cluster gene expression data to identify groups of genes with similar patterns and hence similar function. In this paper, memetic algorithms (MAs)-evolutionary algorithms incorporating local search-are proposed for minimum sum-of-squares clustering (MSSC). In a fitness landscape analysis, it is shown that the MSSC problem has correlation structure exploitable by MAs. The proposed MAs are shown to be superior to multi-start k-means as well as five other clustering algorithms from the bioinformatics literature including hierarchical algorithms and self-organizing maps. Although the fitness values of the different clustering solutions lie close together, it is shown that the solutions differ significantly from each other in terms of cluster memberships which is extremely important for the biological interpretation of the clustering results.

Algorithms↗

EST-PAGE--managing and analyzing EST data.

UNLABELLED: EST-PAGE provides a bioinformatics solution for expressed sequence tags (EST) data entry, database management, GenBank submission, process control and data retrieval from a unified web interface that can be easily customized and adapted by groups working on diverse EST sequencing projects. AVAILABILITY: The system and source code are available upon request from the authors. SUPPLEMENTARY INFORMATION: http://EST-PAGE.binf.gmu.edu

Database Management Systems↗

A combined computational-experimental approach predicts human microRNA targets.

A new paradigm of gene expression regulation has emerged recently with the discovery of microRNAs (miRNAs). Most, if not all, miRNAs are thought to control gene expression, mostly by base pairing with miRNA-recognition elements (MREs) found in their messenger RNA (mRNA) targets. Although a large number of human miRNAs have been reported, many of their mRNA targets remain unknown. Here we used a combined bioinformatics and experimental approach to identify important rules governing miRNA-MRE recognition that allow prediction of human miRNA targets. We describe a computational program, "DIANA-microT", that identifies mRNA targets for animal miRNAs and predicts mRNA targets, bearing single MREs, for human and mouse miRNAs.

Animals↗

Genepi: a blackboard framework for genome annotation.

BACKGROUND: Genome annotation can be viewed as an incremental, cooperative, data-driven, knowledge-based process that involves multiple methods to predict gene locations and structures. This process might have to be executed more than once and might be subjected to several revisions as the biological (new data) or methodological (new methods) knowledge evolves. In this context, although a lot of annotation platforms already exist, there is still a strong need for computer systems which take in charge, not only the primary annotation, but also the update and advance of the associated knowledge. In this paper, we propose to adopt a blackboard architecture for designing such a system RESULTS: We have implemented a blackboard framework (called Genepi) for developing automatic annotation systems. The system is not bound to any specific annotation strategy. Instead, the user will specify a blackboard structure in a configuration file and the system will instantiate and run this particular annotation strategy. The characteristics of this framework are presented and discussed. Specific adaptations to the classical blackboard architecture have been required, such as the description of the activation patterns of the knowledge sources by using an extended set of Allen's temporal relations. Although the system is robust enough to be used on real-size applications, it is of primary use to bioinformatics researchers who want to experiment with blackboard architectures. CONCLUSION: In the context of genome annotation, blackboards have several interesting features related to the way methodological and biological knowledge can be updated. They can readily handle the cooperative (several methods are implied) and opportunistic (the flow of execution depends on the state of our knowledge) aspects of the annotation process.

Algorithms↗

Recent advances in computational genomics.

In the post-genomic era, the new discipline of functional genomics is now facing the challenge of associating a function (as well as estimating its relevance to industrial applications) to about 100,000 microbial, plant or animal genes of known sequence but unknown function. Besides the design of databases, computational methods are increasingly becoming intimately linked with the various experimental approaches. Consequently, bioinformatics is rapidly evolving into independent fields addressing the specific problems of interpreting i) genomic sequences, ii) protein sequences and 3D-structures, as well as iii) transcriptome and macromolecular interaction data. It is thus increasingly difficult for the biologist to choose the computational approaches that perform best in these various areas. This paper attempts to review the most useful developments of the last 2 years.

Computational Biology↗

A graph-theoretic approach to testing associations between disparate sources of functional genomics data.

MOTIVATION: The last few years have seen the advent of high-throughput technologies to analyze various properties of the transcriptome and proteome of several organisms. The congruency of these different data sources, or lack thereof, can shed light on the mechanisms that govern cellular function. A central challenge for bioinformatics research is to develop a unified framework for combining the multiple sources of functional genomics information and testing associations between them, thus obtaining a robust and integrated view of the underlying biology. RESULTS: We present a graph-theoretic approach to test the significance of the association between multiple disparate sources of functional genomics data by proposing two statistical tests, namely edge permutation and node label permutation tests. We demonstrate the use of the proposed tests by finding significant association between a Gene Ontology-derived predictome and data obtained from mRNA expression and phenotypic experiments for Saccharomyces cerevisiae. Moreover, we employ the graph-theoretic framework to recast a surprising discrepancy presented elsewhere between gene expression and knockout phenotype, using expression data from a different set of experiments. AVAILABILITY: An R software package, GraphAT, containing the data and statistical procedures is available from Bioconductor: http://www.bioconductor.org.

Algorithms↗

From masking repeats to identifying functional repeats in the mouse transcriptome.

The back-to-back release of the mouse genome and the functionally annotated RIKEN mouse full-length cDNA collection was an important milestone in mammalian genomics. Yet much of the data remain to be explored in terms of biological effects and mechanisms. For example, interspersed repeats account for 39 per cent of the mouse genome sequence and 11 per cent of representative transcripts. A considerable number of transposable repeat elements are still active and propagating in mouse compared with human. While existing repeat databases and tools assist the classification of repeats or identification of new repeats, there is little bioinformatic support towards exploring the extent and role of repeats in transcriptional variation, modulation of protein function, or gene regulatory events. Since the mouse is used as a model organism to study human genes and their disease associations, this review focuses on information extraction and collation that captures the functional context of repeats in mouse transcripts to facilitate the biological interpretation and extrapolation of findings to the human.

Animals↗

CoMoDis: composite motif discovery in mammalian genomes.

Specificity of mammalian gene regulatory regions is achieved to a large extent through the combinatorial binding of sets of transcription factors to distinct binding sites, discrete combinations of which are often referred to as regulatory modules. Identification and subsequent characterization of gene regulatory modules will be a key step in assembling transcriptional regulatory networks from gene expression profiling data, with the ultimate goal of unravelling the regulatory codes that govern gene expression in various cell types. Here we describe the new bioinformatics tool, Composite Motif Discovery (CoMoDis), which streamlines computational identification of novel regulatory modules starting from a single seed motif. Seed motifs represent binding sites conserved across mammalian species. CoMoDis facilitates novel motif discovery by automating the extraction of DNA sequences flanking seed motifs and streamlining downstream motif discovery using a variety of tools, including several that utilize phylogenetic conservation criteria. CoMoDis is available at http://hscl.cimr.cam.ac.uk/CoMoDis_portal.html.

Animals↗

Talisman--rapid application development for the grid.

In order to make use of the emerging grid and network services offered by various institutes and mandated by many current research projects, some kind of user accessible client is required. In contrast with attempts to build generic workbenches, Talisman is designed to allow a bioinformatics expert to rapidly build custom applications, immediately visible using standard web technology, for users who wish to concentrate on the biology of their problem rather than the informatics aspects. As a component of the MyGrid project, it is intended to allow access to arbitrary resources, including but not limited to relational, object and flat file data sources, analysis programs and grid based storage, tracking and distributed annotation systems.

Computational Biology↗

Supervised identification of allergen-representative peptides for in silico detection of potentially allergenic proteins.

MOTIVATION: Identification of potentially allergenic proteins is needed for the safety assessment of genetically modified foods, certain pharmaceuticals and various other products on the consumer market. Current methods in bioinformatic allergology exploit common features among allergens for the detection of amino acid sequences of potentially allergenic proteins. Features for identification still unexplored include the motifs occurring commonly in allergens, but rarely in ordinary proteins. In this paper, we present an algorithm for the identification of such motifs with the purpose of biocomputational detection of amino acid sequences of potential allergens. RESULTS: Identification of allergen-representative peptides (ARPs) with low or no occurrence in proteins lacking allergenic properties is the essential component of our new method, designated DASARP (Detection based on Automated Selection of Allergen-Representative Peptide). This approach consistently outperforms the criterion based on identical peptide match for predicting allergenicity recommended by ILSI/IFBC and FAO/WHO and shows results comparable to the alignment-based criterion as outlined by FAO/WHO. AVAILABILITY: The detection software and the ARP set needed for the analysis of a query protein reported here are properties of the Swedish National Food Agency and are available upon request. The protein sequence sets used in this work are publicly available on http://www.slv.se/templatesSLV/SLV_Page____9343.asp. Allergenicity assessment for specific protein sequences of interest is also possible via ulfh@slv.se

Algorithms↗

EMDep: a web-based system for the deposition and validation of high-resolution electron microscopy macromolecular structural information.

This paper describes the design and implementation of a Web-based deposition system, EMDep, for macro-molecular volumes determined by electron microscopy and deposited at the European Bioinformatics Institute (EBI) for inclusion in the Electron Microscopy Data Base (EMDB). EMDep is a flexible and portable system (http://www.ebi.ac.uk/msd-srv/emdep/) that allows for the acceptance and validation of data, by an interactive depositor-driven operation. The system takes full advantage of the knowledge and expertise of the experimenters, rather than relying on the database curators, for the complete and accurate description of the structural experiment and its results.

Algorithms↗

cDNA2Genome: a tool for mapping and annotating cDNAs.

BACKGROUND: In the last years several high-throughput cDNA sequencing projects have been funded worldwide with the aim of identifying and characterizing the structure of complete novel human transcripts. However some of these cDNAs are error prone due to frameshifts and stop codon errors caused by low sequence quality, or to cloning of truncated inserts, among other reasons. Therefore, accurate CDS prediction from these sequences first require the identification of potentially problematic cDNAs in order to speed up the posterior annotation process. RESULTS: cDNA2Genome is an application for the automatic high-throughput mapping and characterization of cDNAs. It utilizes current annotation data and the most up to date databases, especially in the case of ESTs and mRNAs in conjunction with a vast number of approaches to gene prediction in order to perform a comprehensive assessment of the cDNA exon-intron structure. The final result of cDNA2Genome is an XML file containing all relevant information obtained in the process. This XML output can easily be used for further analysis such us program pipelines, or the integration of results into databases. The web interface to cDNA2Genome also presents this data in HTML, where the annotation is additionally shown in a graphical form. cDNA2Genome has been implemented under the W3H task framework which allows the combination of bioinformatics tools in tailor-made analysis task flows as well as the sequential or parallel computation of many sequences for large-scale analysis. CONCLUSIONS: cDNA2Genome represents a new versatile and easily extensible approach to the automated mapping and annotation of human cDNAs. The underlying approach allows sequential or parallel computation of sequences for high-throughput analysis of cDNAs.

Chromosome Mapping↗

Use of robust-long serial analysis of gene expression to identify novel fungal and plant genes involved in host-pathogen interactions.

Identification of important transcripts from fungal pathogens and host plants is indispensable for full understanding the molecular events occurring during fungal-plant interactions. Recently, we developed an improved LongSAGE method called robust-long serial analysis of gene expression (RL-SAGE) for deep transcriptome analysis of fungal and plant genomes. Using this method, we made 10 RL-SAGE libraries from two plant species (Oryza sativa and Zea maize) and one fungal pathogen (Magnaporthe grisea). Many of the transcripts identified from these libraries were novel in comparison with their corresponding EST collections. Bioinformatic tools and databases for analyzing the RL-SAGE data were developed. Our results demonstrate that RL-SAGE is an effective approach for large-scale identification of expressed genes in fungal and plant genomes.

Base Sequence↗

Finding exact optimal motifs in matrix representation by partitioning.

MOTIVATION: Finding common patterns, or motifs, in the promoter regions of co-expressed genes is an important problem in bioinformatics. A common representation of the motif is by probability matrix or PSSM (position specific scoring matrix). However, even for a motif of length six or seven, there is no algorithm that can guarantee finding the exact optimal matrix from an infinite number of possible matrices. RESULTS: This paper introduces the first algorithm, called EOMM, for finding the exact optimal matrix-represented motif, or simply optimal motif. Based on branch-and-bound searching by partitioning the solution space recursively, EOMM can find the optimal motif of size up to eight or nine, and a motif of larger size with any desired accuracy on the principle that the smaller the error bound, the longer the running time. Experiments show that for some real and simulated data sets, EOMM finds the motif despite very weak signals when existing software, such as MEME and MITRA-PSSM, fails to do so.

Algorithms↗

Database of p53 gene somatic mutations in human tumors and cell lines: updated compilation and future prospects.

In recent years, there has been an exponential increase in the number of p53 mutations identified in human cancers. The p53 mutation database consists of a list of point mutations in thep53 gene of human tumors and cell lines, compiled from the published literature and made available through electronic media. The database is now maintained at the International Agency for Research on Cancer (IARC) and is updated twice a year. The current version contains records on 5091 published mutations and is expected to surpass the 6000 mark in the January 1997 release. The database is available in various formats through the European Bioinformatics Institute (EBI) ftp server at: ftp://ftp.ebi.ac.uk/pub/databases/p53/ or by request from IARC (p53database@iarc.fr) and will be searchable through the SRS system in the near future. This report provides a description of the criteria for inclusion of data and of the current formats, a summary of the relevance ofp53 mutation analysis to clinical and biological questions, and a brief discussion of the prospects for future developments.

Databases, Factual↗

GeneExpress: a computer system for description, analysis, and recognition of regulatory sequences in eukaryotic genome.

GeneExpress system has been designed to integrate description, analysis, and recognition of eukaryotic regulatory sequences. The system includes 5 basic units: (1) GeneNet contains an object-oriented database for accumulation of data on gene networks and signal transduction pathways and a Java-based viewer that allows an exploration and visualization of the GeneNet information; (2) Transcription Regulation combines the database on transcription regulatory regions of eukaryotic genes (TRRD) and TRRD Viewer; (3) Transcription Factor Binding Site Recognition contains a compilation of transcription factor binding sites (TFBSC) and programs for their analysis and recognition; (4) mRNA Translation is designed for analysis of structural and contextual features of mRNA 5'UTRs and prediction of their translation efficiency; and (5) ACTIVITY is the module for analysis and site activity prediction of a given nucleotide sequence. Integration of the databases in the GeneExpress is based on the Sequence Retrieval System (SRS) created in the European Bioinformatics Institute.

Artificial Intelligence↗