Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Enhanced kinship analysis and STR-based DNA typing for human identification in mass fatality incidents: the Swissair flight 111 disaster.

A bioinformatic tool was developed to assist with the victim identification initiative that followed the Swissair Flight 111 disaster. Making use of short tandem repeat (STR) DNA typing data generated with AmpFlSTR Profiler Plus (PP) and AmpFlSTR COfiler(CO) kits, the software systematically compared each available STR genotype with every other genotype. The matching algorithm was based on the search for: (i) direct matches to genotypes derived from personal effects; and (ii) potential kinship associations between victims and next-of-kin, as measured by allele sharing at individual loci. The software greatly assisted parentage analysis by enabling kinship evaluation in situations where complete parentage trios were unavailable and, in some situations, with distantly related relatives. Exclusion of fortuitous kinship associations (FKA) was made possible through the recovery at the disaster site of at least one remains for every sought-after victim, and was incorporated into the kinship software. The data from the 13 combined STR loci produced 6 and 23 times fewer FKAs when compared with PP alone and AmpFlSTR Profiler (PR) alone, respectively. Identification leads or confirmations of identification were obtained for 218 victims for which DNA reference samples (personal effects and kin) had been submitted. Confirmation of an inferred kinship association was sought through frequency and likelihood calculations, as well as corroborative data from other identification modalities. The use of a simple, yet powerful, automated genotype comparison approach and the use of megaplexes with high power of discrimination (PD) values extended considerably the identification capabilities in the case of the Swissair disaster. The DNA typing identification modality proved to be a valuable component of the large arsenal of identification tools deployed in the aftermath of this disaster.

Accidents, Aviation↗

Genome resources and comparative analysis tools for cardiovascular research.

Disorders of the cardiovascular system are often caused by the interaction of genetic and environmental factors that jointly contribute to individual susceptibility. Genomic data and bioinformatics tools generated from genome projects, coupled with functional verification, offer novel approaches to study both rare single-gene and complex multigenic cardiovascular diseases. These approaches include gene mapping using genome variation, especially single-nucleotide polymorphisms and comparative genomics within and between species. This chapter illustrates the major genome resources, associated bioinformatics tools, and their potential application in cardiovascular research.

Cardiovascular Diseases↗

Chemical effects in biological systems (CEBS) object model for toxicology data, SysTox-OM: design and application.

MOTIVATION: The CEBS data repository is being developed to promote a systems biology approach to understand the biological effects of environmental stressors. CEBS will house data from multiple gene expression platforms (transcriptomics), protein expression and protein-protein interaction (proteomics), and changes in low molecular weight metabolite levels (metabolomics) aligned by their detailed toxicological context. The system will accommodate extensive complex querying in a user-friendly manner. CEBS will store toxicological contexts including the study design details, treatment protocols, animal characteristics and conventional toxicological endpoints such as histopathology findings and clinical chemistry measures. All of these data types can be integrated in a seamless fashion to enable data query and analysis in a biologically meaningful manner. RESULTS: An object model, the SysBio-OM (Xirasagar et al., 2004) has been designed to facilitate the integration of microarray gene expression, proteomics and metabolomics data in the CEBS database system. We now report SysTox-OM as an open source systems toxicology model designed to integrate toxicological context into gene expression experiments. The SysTox-OM model is comprehensive and leverages other open source efforts, namely, the Standard for Exchange of Nonclinical Data (http://www.cdisc.org/models/send/v2/index.html) which is a data standard for capturing toxicological information for animal studies and Clinical Data Interchange Standards Consortium (http://www.cdisc.org/models/sdtm/index.html) that serves as a standard for the exchange of clinical data. Such standardization increases the accuracy of data mining, interpretation and exchange. The open source SysTox-OM model, which can be implemented on various software platforms, is presented here. AVAILABILITY: A universal modeling language (UML) depiction of the entire SysTox-OM is available at http://cebs.niehs.nih.gov and the Rational Rose object model package is distributed under an open source license that permits unrestricted academic and commercial use and is available at http://cebs.niehs.nih.gov/cebsdownloads. Currently, the public toxicological data in CEBS can be queried via a web application based on the SysTox-OM at http://cebs.niehs.nih.gov CONTACT: xirasagars@saic.com SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Computational Biology↗

GBPM: GRID-based pharmacophore model: concept and application studies to protein-protein recognition.

MOTIVATION: Automatic procedures to obtain pharmacophore models from experimentally determined macromolecular complexes can help in the drug discovery process, especially when protein-protein recognition plays an important biological role. RESULTS: The GRID-based pharmacophore model (GBPM) is a fully objective method for defining most relevant interaction areas in complexes deriving pharmacophore models from three-dimensional (3D) molecular structure information. It is based on logical and clustering operations with 3D maps computed by the GRID program on structurally known molecular complexes. In this manuscript we describe the concept and discuss application examples regarding protein-protein recognition. In particular two complexes selected in the Protein Data Bank have been tested to evaluate the GBPM capability to identify interaction areas. The results obtained show the capabilities of this new bioinformatic method.

Algorithms↗

Direct identification of proteins from T47D cells and murine brain tissue by matrix-assisted laser desorption/ionization post-source decay/collision-induced dissociation.

The purpose of this study is to determine the feasibility of the direct matrix-assisted laser desorption/ionization (MALDI) identification of proteins in fixed T47D breast cancer cells and murine brain tissues. The ability to identify proteins from cells and tissue may lead to biomarkers that effectively predict the onset of defined disease states, and their dynamic behavior could be an important hint for drug target discoveries. Direct tissue application of trypsin allows protein identification in cells and tissues, while maintaining spatial integrity and intracellular organization. Using a chemical printer, matrix was co-registered on trypsinized human T47D breast cancer cells and cryo-preserved sections of murine brain tissue, followed by MALDI post-source decay (PSD) or MALDI collision-induced dissociation (CID), respectively. Mass-to-charge (m/z) data from the cells and brain tissues were processed using Mascot software interrogation of the National Center for Biotechnology Information (NCBI) database. Histone H2B was identified from cultured T47D human breast cancer cells. Tubulin beta2 was identified from mouse brain cortex following an induced stroke. These results suggest that MALDI PSD/CID, combined with bioinformatics, can be used for the direct identification of proteins from cells and tissues. Refinements in preparation techniques may improve this approach to provide a tool for quantitative proteomics and clinical analysis.

Animals↗

Bayesian inference on biopolymer models.

MOTIVATION: Most existing bioinformatics methods are limited to making point estimates of one variable, e.g. the optimal alignment, with fixed input values for all other variables, e.g. gap penalties and scoring matrices. While the requirement to specify parameters remains one of the more vexing issues in bioinformatics, it is a reflection of a larger issue: the need to broaden the view on statistical inference in bioinformatics. RESULTS: The assignment of probabilities for all possible values of all unknown variables in a problem in the form of a posterior distribution is the goal of Bayesian inference. Here we show how this goal can be achieved for most bioinformatics methods that use dynamic programming. Specifically, a tutorial style description of a Bayesian inference procedure for segmentation of a sequence based on the heterogeneity in its composition is given. In addition, full Bayesian inference algorithms for sequence alignment are described. AVAILABILITY: Software and a set of transparencies for a tutorial describing these ideas are available at http://www.wadsworth.org/res&res/bioinfo/

Bayes Theorem↗

CFinder: locating cliques and overlapping modules in biological networks.

UNLABELLED: Most cellular tasks are performed not by individual proteins, but by groups of functionally associated proteins, often referred to as modules. In a protein association network modules appear as groups of densely interconnected nodes, also called communities or clusters. These modules often overlap with each other and form a network of their own, in which nodes (links) represent the modules (overlaps). We introduce CFinder, a fast program locating and visualizing overlapping, densely interconnected groups of nodes in undirected graphs, and allowing the user to easily navigate between the original graph and the web of these groups. We show that in gene (protein) association networks CFinder can be used to predict the function(s) of a single protein and to discover novel modules. CFinder is also very efficient for locating the cliques of large sparse graphs. AVAILABILITY: CFinder (for Windows, Linux and Macintosh) and its manual can be downloaded from http://angel.elte.hu/clustering. SUPPLEMENTARY INFORMATION: Supplementary data are available on Bioinformatics online.

Biology↗

MinSet: a general approach to derive maximally representative database subsets by using fragment dictionaries and its application to the SCOP database.

MOTIVATION: The size of current protein databases is a challenge for many Bioinformatics applications, both in terms of processing speed and information redundancy. It may be therefore desirable to efficiently reduce the database of interest to a maximally representative subset. RESULTS: The MinSet method employs a combination of a Suffix Tree and a Genetic Algorithm for the generation, selection and assessment of database subsets. The approach is generally applicable to any type of string-encoded data, allowing for a drastic reduction of the database size whilst retaining most of the information contained in the original set. We demonstrate the performance of the method on a database of protein domain structures encoded as strings. We used the SCOP40 domain database by translating protein structures into character strings by means of a structural alphabet and by extracting optimized subsets according to an entropy score that is based on a constant-length fragment dictionary. Therefore, optimized subsets are maximally representative for the distribution and range of local structures. Subsets containing only 10% of the SCOP structure classes show a coverage of >90% for fragments of length 1-4. AVAILABILITY: http://mathbio.nimr.mrc.ac.uk/~jkleinj/MinSet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

Modelling biological processes using workflow and Petri Net models.

MOTIVATION: Biological processes can be considered at many levels of detail, ranging from atomic mechanism to general processes such as cell division, cell adhesion or cell invasion. The experimental study of protein function and gene regulation typically provides information at many levels. The representation of hierarchical process knowledge in biology is therefore a major challenge for bioinformatics. To represent high-level processes in the context of their component functions, we have developed a graphical knowledge model for biological processes that supports methods for qualitative reasoning. RESULTS: We assessed eleven diverse models that were developed in the fields of software engineering, business, and biology, to evaluate their suitability for representing and simulating biological processes. Based on this assessment, we combined the best aspects of two models: Workflow/Petri Net and a biological concept model. The Workflow model can represent nesting and ordering of processes, the structural components that participate in the processes, and the roles that they play. It also maps to Petri Nets, which allow verification of formal properties and qualitative simulation. The biological concept model, TAMBIS, provides a framework for describing biological entities that can be mapped to the workflow model. We tested our model by representing malaria parasites invading host erythrocytes, and composed queries, in five general classes, to discover relationships among processes and structural components. We used reachability analysis to answer queries about the dynamic aspects of the model. AVAILABILITY: The model is available at http://smi.stanford.edu/projects/helix/pubs/process-model/.

Animals↗

Analysis of the golden Syrian hamster anterior pituitary gland proteome by ion trap mass spectrometry.

We utilized mass spectrometry (MS) and bioinformatics to investigate the proteome of the anterior pituitary gland (AP). Subcellular fractions of APs from 2-month-old male Golden Syrian hamsters were prepared for protein denaturation, treatment with trypsin and analyses utilizing micro liquid chromatography MS/MS and the database search software SEQUEST. In the nuclear, non-nuclear 100,000 x g and cytosolic fractions we identified 76, 52 and 52 different proteins, respectively. A total of 145 distinct proteins were detected. We identified growth hormone, prolactin, pro-opiomelanocortin, the alpha-subunit for the glycoprotein hormones, luteinizing hormone-beta and follicle-stimulating hormone-beta. Groups of other identified proteins included hormone processing, secretion granule associated, non-hormonal endoplasmic reticulum associated, calcium binding, protein kinase C associated histone and non-histone chromosomal material, other RNA-binding, splicing factors, heterogeneous nuclear ribonucleoproteins, helicases, lamins, microfilament associated, microtubule associated, adenosine triphosphate and guanosine diphosphate associated, keratins, lysosomal, ribosomal, enzymes in glycolysis and the tricarboxylic and pentose phosphate paths, glutathione associated, transmethylation, catabolic and unknown protein products as well as blood hemoglobins. Proteins previously not reported in the AP, such as fertility protein SP22, were identified. The proteins identified in the present study form a foundation for defining the proteome in normal adult male AP.

Animals↗

Physiological proteomics: cells, organs, biological fluids, and biomarkers.

Proteomic research is accelerating rapidly because of marked advances in protein labeling techniques, mass spectrometry (MS), and bioinformatics. Two-dimensional difference gel electrophoresis (2D-DIGE) is being used effectively in conjunction with liquid chromatography tandem MS (LC-MS/MS) and/or matrix-assisted laser desorption/ionization time-of-flight MS (MALDI-ToF MS) and database search software to quantify relative changes in the levels of proteins in two samples. It is now possible in a single study to identify and quantify large numbers of proteins and their posttranslational modifications in different biological samples. Comparisons can be made between groups of animals in different physiological states or in response to experimental treatment. Differences between normal individuals and those in disease states can form the foundation for elucidation of causative factors of disease and the identification of biomarkers for the diseased state. This symposium includes original research that compares the erythrocyte plasma membrane proteome in the normal and the sickle cell state, evaluates the anterior pituitary gland proteome in the ovariectomized rat in response to estrogen, and assesses proteomic methodology employed to identify potentially useful biomarkers in human cells and fluids for clinical medicine. It is directed not only to investigators working in these fields but also to a diverse group of scientists working in the biological and biomedical fields to stimulate cross-disciplinary awareness, interest, and collaboration.

Anemia, Sickle Cell↗

EXPANDER--an integrative program suite for microarray data analysis.

BACKGROUND: Gene expression microarrays are a prominent experimental tool in functional genomics which has opened the opportunity for gaining global, systems-level understanding of transcriptional networks. Experiments that apply this technology typically generate overwhelming volumes of data, unprecedented in biological research. Therefore the task of mining meaningful biological knowledge out of the raw data is a major challenge in bioinformatics. Of special need are integrative packages that provide biologist users with advanced but yet easy to use, set of algorithms, together covering the whole range of steps in microarray data analysis. RESULTS: Here we present the EXPANDER 2.0 (EXPression ANalyzer and DisplayER) software package. EXPANDER 2.0 is an integrative package for the analysis of gene expression data, designed as a 'one-stop shop' tool that implements various data analysis algorithms ranging from the initial steps of normalization and filtering, through clustering and biclustering, to high-level functional enrichment analysis that points to biological processes that are active in the examined conditions, and to promoter cis-regulatory elements analysis that elucidates transcription factors that control the observed transcriptional response. EXPANDER is available with pre-compiled functional Gene Ontology (GO) and promoter sequence-derived data files for yeast, worm, fly, rat, mouse and human, supporting high-level analysis applied to data obtained from these six organisms. CONCLUSION: EXPANDER integrated capabilities and its built-in support of multiple organisms make it a very powerful tool for analysis of microarray data. The package is freely available for academic users at http://www.cs.tau.ac.il/~rshamir/expander.

Algorithms↗

A multistep bioinformatic approach detects putative regulatory elements in gene promoters.

BACKGROUND: Searching for approximate patterns in large promoter sequences frequently produces an exceedingly high numbers of results. Our aim was to exploit biological knowledge for definition of a sheltered search space and of appropriate search parameters, in order to develop a method for identification of a tractable number of sequence motifs. RESULTS: Novel software (COOP) was developed for extraction of sequence motifs, based on clustering of exact or approximate patterns according to the frequency of their overlapping occurrences. Genomic sequences of 1 Kb upstream of 91 genes differentially expressed and/or encoding proteins with relevant function in adult human retina were analyzed. Methodology and results were tested by analysing 1,000 groups of putatively unrelated sequences, randomly selected among 17,156 human gene promoters. When applied to a sample of human promoters, the method identified 279 putative motifs frequently occurring in retina promoters sequences. Most of them are localized in the proximal portion of promoters, less variable in central region than in lateral regions and similar to known regulatory sequences. COOP software and reference manual are freely available upon request to the Authors. CONCLUSION: The approach described in this paper seems effective for identifying a tractable number of sequence motifs with putative regulatory role.

Algorithms↗

Protein structure prediction methods for drug design.

Along the long path from genomic data to a new drug, the knowledge of three-dimensional protein structure can be of significant help in several places. This paper points out such places, discusses the virtues of protein structure knowledge and reviews bioinformatics methods for gaining such knowledge on the protein structure.

Computational Biology↗

Fast parsers for Entrez Gene.

NCBI completed the transition of its main genome annotation database from Locuslink to Entrez Gene in Spring 2005. However, to this date few parsers exist for the Entrez Gene annotation file. Owing to the widespread use of Locuslink and the popularity of Perl programming language in bioinformatics, a publicly available high performance Entrez Gene parser in Perl is urgently needed. We present four such parsers that were developed using several parsing approaches (Parse::RecDescent, Parse::Yapp, Perl-byacc and Perl 5 regular expressions) and provide the first in-depth comparison of these sophisticated Perl tools. Our fastest parser processes the entire human Entrez Gene annotation file in under 12 min on one Intel Xeon 2.4 GHz CPU and can be of help to the bioinformatics community during and after the transition from Locuslink to Entrez Gene.

Algorithms↗