Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

A high productivity/low maintenance approach to high-performance computation for biomedicine: four case studies.

The rapid advances in high-throughput biotechnologies such as DNA microarrays and mass spectrometry have generated vast amounts of data ranging from gene expression to proteomics data. The large size and complexity involved in analyzing such data demand a significant amount of computing power. High-performance computation (HPC) is an attractive and increasingly affordable approach to help meet this challenge. There is a spectrum of techniques that can be used to achieve computational speedup with varying degrees of impact in terms of how drastic a change is required to allow the software to run on an HPC platform. This paper describes a high- productivity/low-maintenance (HP/LM) approach to HPC that is based on establishing a collaborative relationship between the bioinformaticist and HPC expert that respects the former's codes and minimizes the latter's efforts. The goal of this approach is to make it easy for bioinformatics researchers to continue to make iterative refinements to their programs, while still being able to take advantage of HPC. The paper describes our experience applying these HP/LM techniques in four bioinformatics case studies: (1) genome-wide sequence comparison using Blast, (2) identification of biomarkers based on statistical analysis of large mass spectrometry data sets, (3) complex genetic analysis involving ordinal phenotypes, (4) large-scale assessment of the effect of possible errors in analyzing microarray data. The case studies illustrate how the HP/LM approach can be applied to a range of representative bioinformatics applications and how the approach can lead to significant speedup of computationally intensive bioinformatics applications, while making only modest modifications to the programs themselves.

Amino Acid Sequence↗

Biosphere: the interoperation of web services in microarray cluster analysis.

UNLABELLED: The growing use of DNA microarrays in biomedical research has led to the proliferation of analysis tools. These software programs address different aspects of analysis (e.g. normalisation and clustering within and across individual arrays) as well as extended analysis methods (e.g. clustering, annotation and mining of multiple datasets). Therefore, microarray data analysis typically requires the interoperability of multiple software programs involving different analysis types and methods. Such interoperation is often hampered by the heterogeneity inherent in the software tools (which may function by implementing different interfaces and using different programming languages). To address this problem, we employed the simple object access protocol (SOAP)-based web service approach that provides a uniform programmatic interface to these heterogeneous software components. To demonstrate this approach in the microarray context, we created a web server application, Biosphere, which interoperates a number of web services that are geographically widely distributed. These web services include a clustering web service, which is a suite of different clustering algorithms for analysing microarray data; XEMBL, developed at the European Bioinformatics Institute (EBI) for retrieving EMBL Nucleotide Sequence Database sequence data; and three gene annotation web services: GetGO, GetHAPI and GetUMLS. GetGO allows retrieval of Gene Ontology (GO) annotation, and the other two web services retrieve annotation from the biomedical literature that is indexed based on the Medical Subject Headings (MeSH) terms. With these web services, Biosphere allows the users to do the following: (i) cluster gene expression data using seven different algorithms; (ii) visualise the clustering results that are grouped statistically in colour; and (iii) retrieve sequence, annotation and citation data for the genes of interest. AVAILABILITY: Biosphere and its web services described in Web Service Description Language (WSDL) can be accessed at http://rook.cecid.hku.hk:8280/BiosphereServer.

Cluster Analysis↗

The role of informatics in the coordinated management of biological resources collections.

The term 'biological resources' is applied to the living biological material collected, held and catalogued in culture collections: bacterial and fungal cultures; animal, human and plant cells; viruses; and isolated genetic material. A wealth of information on these materials has been accumulated in culture collections, and most of this information is accessible. Digitalisation of data has reached a high level; however, information is still dispersed. Individual and coordinated approaches have been initiated to improve accessibility of biological resource centres, their holdings and related information through the Internet. These approaches cover subjects such as standardisation of data handling and data accessibility, and standardisation and quality control of laboratory procedures. This article reviews some of the most important initiatives implemented so far, as well as the most recent achievements. It also discusses the possible improvements that could be achieved by adopting new communication standards and technologies, such as web services, in view of a deeper and more fruitful integration of biological resources information in the bioinformatics network environment.

Animals↗

[Identification and analysis of a mouse gene homologous to human hepatitis B virus pre-S1 protein-binding protein using the bioinformatics method].

OBJECTIVE: To clone and identify the mouse gene homologous to human hepatitis B virus (HBV) pre-S1 protein-binding protein (PS1BP). METHODS: The human PS1BP cDNA sequence was used as the reference sequence to search homologous mouse cDNA sequence from GenBank established by National Center for Biotechnology (NCBI), National Institute of Health (NIH), for its homologous cDNA sequences of mouse by BLASTn tool. The characteristics of mouse PS1BP protein primary structure were predicted by online software. Finally the genomic DNA structure of mouse PS1BP was deduced and compared. RESULTS: The mouse PS1BP was identified and consisted of 1455 nt, coding a protein of 484 aa. The identity of human and mouse PS1BP protein is 84.92% (411/484). The genomic DNA of mouse PS1BP consisted of 3 exons and 2 introns. CONCLUSION: The identification and characterization of mouse PS1BP cDNA and genomic DNA pave a way for further study of their structures and functions.

Amino Acid Sequence↗

Multiple sequence alignment with the Clustal series of programs.

The Clustal series of programs are widely used in molecular biology for the multiple alignment of both nucleic acid and protein sequences and for preparing phylogenetic trees. The popularity of the programs depends on a number of factors, including not only the accuracy of the results, but also the robustness, portability and user-friendliness of the programs. New features include NEXUS and FASTA format output, printing range numbers and faster tree calculation. Although, Clustal was originally developed to run on a local computer, numerous Web servers have been set up, notably at the EBI (European Bioinformatics Institute) (http://www.ebi.ac.uk/clustalw/).

Algorithms↗

The Ensembl genome database project.

The Ensembl (http://www.ensembl.org/) database project provides a bioinformatics framework to organise biology around the sequences of large genomes. It is a comprehensive source of stable automatic annotation of the human genome sequence, with confirmed gene predictions that have been integrated with external data sources, and is available as either an interactive web site or as flat files. It is also an open source software engineering project to develop a portable system able to handle very large genomes and associated requirements from sequence analysis to data storage and visualisation. The Ensembl site is one of the leading sources of human genome sequence annotation and provided much of the analysis for publication by the international human genome project of the draft genome. The Ensembl system is being installed around the world in both companies and academic sites on machines ranging from supercomputers to laptops.

Computational Biology↗

Recent developments of the chemistry development kit (CDK) - an open-source java library for chemo- and bioinformatics.

The Chemistry Development Kit (CDK) provides methods for common tasks in molecular informatics, including 2D and 3D rendering of chemical structures, I/O routines, SMILES parsing and generation, ring searches, isomorphism checking, structure diagram generation, etc. Implemented in Java, it is used both for server-side computational services, possibly equipped with a web interface, as well as for applications and client-side applets. This article introduces the CDK's new QSAR capabilities and the recently introduced interface to statistical software.

Computational Biology↗

Plant metabolomics: large-scale phytochemistry in the functional genomics era.

Metabolomics or the large-scale phytochemical analysis of plants is reviewed in relation to functional genomics and systems biology. A historical account of the introduction and evolution of metabolite profiling into today's modern comprehensive metabolomics approach is provided. Many of the technologies used in metabolomics, including optical spectroscopy, nuclear magnetic resonance, and mass spectrometry are surveyed. The critical role of bioinformatics and various methods of data visualization are summarized and the future role of metabolomics in plant science assessed.

Computational Biology↗

Computational cluster validation in post-genomic data analysis.

MOTIVATION: The discovery of novel biological knowledge from the ab initio analysis of post-genomic data relies upon the use of unsupervised processing methods, in particular clustering techniques. Much recent research in bioinformatics has therefore been focused on the transfer of clustering methods introduced in other scientific fields and on the development of novel algorithms specifically designed to tackle the challenges posed by post-genomic data. The partitions returned by a clustering algorithm are commonly validated using visual inspection and concordance with prior biological knowledge--whether the clusters actually correspond to the real structure in the data is somewhat less frequently considered. Suitable computational cluster validation techniques are available in the general data-mining literature, but have been given only a fraction of the same attention in bioinformatics. RESULTS: This review paper aims to familiarize the reader with the battery of techniques available for the validation of clustering results, with a particular focus on their application to post-genomic data analysis. Synthetic and real biological datasets are used to demonstrate the benefits, and also some of the perils, of analytical clustervalidation. AVAILABILITY: The software used in the experiments is available at http://dbkweb.ch.umist.ac.uk/handl/clustervalidation/. SUPPLEMENTARY INFORMATION: Enlarged colour plots are provided in the Supplementary Material, which is available at http://dbkweb.ch.umist.ac.uk/handl/clustervalidation/.

Algorithms↗

BioContrasts: extracting and exploiting protein-protein contrastive relations from biomedical literature.

MOTIVATION: Contrasts are useful conceptual vehicles for learning processes and exploratory research of the unknown. For example, contrastive information between proteins can reveal what similarities, divergences and relations there are of the two proteins, leading to invaluable insights for better understanding about the proteins. Such contrastive information are found to be reported in the biomedical literature. However, there have been no reported attempts in current biomedical text mining work that systematically extract and present such useful contrastive information from the literature for exploitation. RESULTS: Our BioContrasts system extracts protein-protein contrastive information from MEDLINE abstracts and presents the information to biologists in a web-application for exploitation. Contrastive information are identified in the text abstracts with contrastive negation patterns such as 'A but not B'. A total of 799 169 pairs of contrastive expressions were successfully extracted from 2.5 million MEDLINE abstracts. Using grounding of contrastive protein names to Swiss-Prot entries, we were able to produce 41 471 pieces of contrasts between Swiss-Prot protein entries. These contrastive pieces of information are then presented via a user-friendly interactive web portal that can be exploited for applications such as the refinement of biological pathways. AVAILABILITY: BioContrasts can be accessed at http://biocontrasts.i2r.a-star.edu.sg. It is also mirrored at http://biocontrasts.biopathway.org. SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online.

Artificial Intelligence↗

SPINE bioinformatics and data-management aspects of high-throughput structural biology.

SPINE (Structural Proteomics In Europe) was established in 2002 as an integrated research project to develop new methods and technologies for high-throughput structural biology. Development areas were broken down into workpackages and this article gives an overview of ongoing activity in the bioinformatics workpackage. Developments cover target selection, target registration, wet and dry laboratory data management and structure annotation as they pertain to high-throughput studies. Some individual projects and developments are discussed in detail, while those that are covered elsewhere in this issue are treated more briefly. In particular, this overview focuses on the infrastructure of the software that allows the experimentalist to move projects through different areas that are crucial to high-throughput studies, leading to the collation of large data sets which are managed and eventually archived and/or deposited.

Computational Biology↗

MuTrack: a genome analysis system for large-scale mutagenesis in the mouse.

BACKGROUND: Modern biological research makes possible the comprehensive study and development of heritable mutations in the mouse model at high-throughput. Using techniques spanning genetics, molecular biology, histology, and behavioral science, researchers may examine, with varying degrees of granularity, numerous phenotypic aspects of mutant mouse strains directly pertinent to human disease states. Success of these and other genome-wide endeavors relies on a well-structured bioinformatics core that brings together investigators from widely dispersed institutions and enables them to seamlessly integrate data, observations and discussions. DESCRIPTION: MuTrack was developed as the bioinformatics core for a large mouse phenotype screening effort. It is a comprehensive collection of on-line computational tools and tracks thousands of mutagenized mice from birth through senescence and death. It identifies the physical location of mice during an intensive phenotype screening process at several locations throughout the state of Tennessee and collects raw and processed experimental data from each domain. MuTrack's statistical package allows researchers to access a real-time analysis of mouse pedigrees for aberrant behavior, and subsequent recirculation and retesting. The end result is the classification of potential and actual heritable mutant mouse strains that become immediately available to outside researchers who have expressed interest in the mutant phenotype. CONCLUSION: MuTrack demonstrates the effectiveness of using bioinformatics techniques in data collection, integration and analysis to identify unique result sets that are beyond the capacity of a solitary laboratory. By employing the research expertise of investigators at several institutions for a broad-ranging study, the TMGC has amplified the effectiveness of any one consortium member. The bioinformatics strategy presented here lends future collaborative efforts a template for a comprehensive approach to large-scale analysis.

Animals↗

Biomedical literature mining: challenges and solutions in the 'omics' era.

It is now obvious that the rate-limiting step in high throughput experimentation is neither data acquisition nor analysis, but rather our ability to interpret data on a genome-wide scale. Indeed, the explosion of data sampling capacity combined with increasing publication rates greatly impairs our ability to find meaning in vast collections of data. In order to support data interpretation, bioinformatic tools are needed to identify critical information contained in large bodies of literature. However, extracting knowledge embedded in free text is an arduous task, compounded in the biomedical field by an inconsistent gene nomenclature, domain-specific language and restricted access to full text articles. This paper presents a selection of currently available biomedical literature mining software. These tools rely on statistic and, more recently, semantic analyses (Natural Language Processing) to automatically extract information from the literature. In addition, a literature mining strategy has been developed to explore patterns of term occurrences in abstracts. This method automatically identifies relevant keywords in collections of abstracts, and uses a pattern discovery algorithm to generate a visual interface for exploring functional associations among genes. Term occurrence heatmaps can also be combined with gene expression profiles to provide valuable functional annotations. Furthermore, as demonstrated with tumor cell line literature profiling results, this approach can be applied to a variety of themes beyond genomic data analysis. Altogether, these examples illustrate how literature analysis can be employed to support knowledge discovery in biomedical research.

Algorithms↗

Potential drug targets in Mycobacterium tuberculosis through metabolic pathway analysis.

The emergence of multidrug resistant varieties of Mycobacterium tuberculosis has led to a search for novel drug targets. We have performed an insilico comparative analysis of metabolic pathways of the host Homo sapiens and the pathogen M. tuberculosis. Enzymes from the biochemical pathways of M. tuberculosis from the KEGG metabolic pathway database were compared with proteins from the host H. sapiens, by performing a BLASTp search against the non-redundant database restricted to the H. sapiens subset. The e-value threshold cutoff was set to 0.005. Enzymes, which do not show similarity to any of the host proteins, below this threshold, were filtered out as potential drug targets. We have identified six pathways unique to the pathogen M. tuberculosis when compared to the host H. sapiens. Potential drug targets from these pathways could be useful for the discovery of broad spectrum drugs. Potential drug targets were also identified from pathways related to lipid metabolism, carbohydrate metabolism, amino acid metabolism, energy metabolism, vitamin and cofactor biosynthetic pathways and nucleotide metabolism. Of the 185 distinct targets identified from these pathways, many are in various stages of progress at the TB Structural Genomics Consortium. However, 67 of our targets are new and can be considered for rational drug design. As a case study, we have built a homology model of one of the potential drug targets MurD ligase using WHAT IF software. The model could be further explored for insilico docking studies with suitable inhibitors. The study was successful in listing out potential drug targets from the M. tuberculosis proteome involved in vital aspects of the pathogen's metabolism, persistence, virulence and cell wall biosynthesis. This systematic evaluation of metabolic pathways of host and pathogen through reliable and conventional bioinformatic methods can be extended to other pathogens of clinical interest.

Amino Acid Sequence↗

TFinder: A Python Web Tool for Predicting Transcription Factor Binding Sites.

Transcription is a key cell process that consists of synthesizing several copies of RNA from a gene DNA sequence. This process is highly regulated and closely linked to the ability of transcription factors to bind specifically to DNA. TFinder is an easy-to-use Python web portal allowing the identification of Individual Motifs (IM) such as Transcription Factor Binding Sites (TFBS). Using the NCBI API, TFinder extracts either promoter or gene terminal regulatory regions, through a simple query of NCBI gene name or ID. It enables simultaneous analysis across five different species for an unlimited number of genes. TFinder searches for Individual Motifs in different formats, including IUPAC codes and JASPAR entries. Moreover, TFinder also allows de novo generations of a Position Weight Matrix (PWM) and the use of already established PWM. Finally, the data are provided in a tabular and a graph format showing the relevance and the P-value of the Individual Motifs found as well as their location relative to the Transcription Start Site (TSS) or the terminal region of the gene. The results are then sent by email to users facilitating the subsequent data analysis and sharing. TFinder is written in Python and freely available on GitHub under the MIT license: https://github.com/Jumitti/TFinder. It can be accessed as a web application implemented in Streamlit at https://tfinder-ipmc.streamlit.app. Resources are available on Streamlit "Resources" tab. TFINDER strength is that it relies on an all-in-one intuitive tool allowing users inexperienced with bioinformatics tools to retrieve gene regulatory regions sequences in multiple species and to search for individual motifs in a huge number of genes.

Transcription Factors↗

Developing an energy landscape for the novel function of a (beta/alpha)8 barrel: ammonia conduction through HisF.

HisH-hisF is a multidomain globular protein complex; hisH is a class I glutamine amidotransferase that hydrolyzes glutamine to form ammonia, and hisF is a (beta/alpha)8 barrel cyclase that completes the ring formation of imidizole glycerol phosphate synthase. Together, hisH and hisF form a glutamine amidotransferase that carries out the fifth step of the histidine biosynthetic pathway. Recently, it has been suggested that the (beta/alpha)8 barrel participates in a novel function: to channel ammonia from the active site of hisH to the active site of hisF. The present study presents a series of molecular dynamic simulations that investigate the channeling function of hisF. This article reconstructs potentials of mean force for the conduction of ammonia through the channel, and the entrance of ammonia through the strictly conserved channel gate, in both a closed and a hypothetical open conformation. The resulting energy landscape within the channel supports the idea that ammonia does indeed pass through the barrel, interacting with conserved hydrophilic residues along the way. The proposed open conformation, which involves an alternate rotamer state of one of the gate residues, presents only an approximately 2.5-kcal energy barrier to ammonia entry. Another alternate open-gate conformation, which may play a role in non-nitrogen-fixing organisms, is deduced through bioinformatics.

Algorithms↗

springScape: visualisation of microarray and contextual bioinformatic data using spring embedding and an 'information landscape'.

The interpretation of microarray and other high-throughput data is highly dependent on the biological context of experiments. However, standard analysis packages are poor at simultaneously presenting both the array and related bioinformatic data. We have addressed this challenge by developing a system springScape based on 'spring embedding' and an 'information landscape' allowing several related data sources to be dynamically combined while highlighting one particular feature. Each data source is represented as a network of nodes connected by weighted edges. The networks are combined and embedded in the 2-D plane by spring embedding such that nodes with a high similarity are drawn close together. Complex relationships can be discovered by varying the weight of each data source and observing the dynamic response of the spring network. By modifying Procrustes analysis, we find that the visualizations have an acceptable degree of reproducibility. The 'information landscape' highlights one particular data source, displaying it as a smooth surface whose height is proportional to both the information being viewed and the density of nodes. The algorithm is demonstrated using several microarray data sets in combination with protein-protein interaction data and GO annotations. Among the features revealed are the spatio-temporal profile of gene expression and the identification of GO terms correlated with gene expression and protein interactions. The power of this combined display lies in its interactive feedback and exploitation of human visual pattern recognition. Overall, springScape shows promise as a tool for the interpretation of microarray data in the context of relevant bioinformatic information.

Algorithms↗

BioWarehouse: a bioinformatics database warehouse toolkit.

BACKGROUND: This article addresses the problem of interoperation of heterogeneous bioinformatics databases. RESULTS: We introduce BioWarehouse, an open source toolkit for constructing bioinformatics database warehouses using the MySQL and Oracle relational database managers. BioWarehouse integrates its component databases into a common representational framework within a single database management system, thus enabling multi-database queries using the Structured Query Language (SQL) but also facilitating a variety of database integration tasks such as comparative analysis and data mining. BioWarehouse currently supports the integration of a pathway-centric set of databases including ENZYME, KEGG, and BioCyc, and in addition the UniProt, GenBank, NCBI Taxonomy, and CMR databases, and the Gene Ontology. Loader tools, written in the C and JAVA languages, parse and load these databases into a relational database schema. The loaders also apply a degree of semantic normalization to their respective source data, decreasing semantic heterogeneity. The schema supports the following bioinformatics datatypes: chemical compounds, biochemical reactions, metabolic pathways, proteins, genes, nucleic acid sequences, features on protein and nucleic-acid sequences, organisms, organism taxonomies, and controlled vocabularies. As an application example, we applied BioWarehouse to determine the fraction of biochemically characterized enzyme activities for which no sequences exist in the public sequence databases. The answer is that no sequence exists for 36% of enzyme activities for which EC numbers have been assigned. These gaps in sequence data significantly limit the accuracy of genome annotation and metabolic pathway prediction, and are a barrier for metabolic engineering. Complex queries of this type provide examples of the value of the data warehousing approach to bioinformatics research. CONCLUSION: BioWarehouse embodies significant progress on the database integration problem for bioinformatics.

Computational Biology↗