Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “software tools”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Modelling microbial growth in structured foods: towards a unified approach.

Historically, the ability of foods to support the growth of spoilage organisms and food-borne pathogens has been assessed by inoculating a food with an organism of interest, and following its growth over a period of time. Information gained from such challenge tests, together with knowledge of the organoleptic stability of the product, can then be used to determine an appropriate shelf-life for the food. Whilst this approach may be seen as the "gold-standard" of microbiological assessment of food, it is both time-consuming and costly. A major advance to complement challenge testing was the development of predictive modelling, when it was demonstrated that the growth of a wide range of organisms of interest could be quite accurately modelled as a function of only a few environmental parameters-primarily temperature, pH and water activity (a(w)), with perhaps other factors such as nitrite, organic acids and oxygen. This approach to predictive microbiology is embodied in software tools such as the UK Food MicroModel and the Pathogen Modeling Program from the USA. Whilst modelling of this form yields accurate predictions of the growth of organisms in the majority of foods, there are occasions when there are discrepancies between the model and the observed growth. These discrepancies are most often described as "fail-safe", i.e. the observed growth is slower than predicted by the model. This paper examines the role of food structure in the development of microbial populations and communities, and describes the methodologies we propose to begin to tackle some of these complex and interlinked issues.

Bacteria↗

Data integration and visualization system for enabling conceptual biology.

MOTIVATION: Integration of heterogeneous data in life sciences is a growing and recognized challenge. The problem is not only to enable the study of such data within the context of a biological question but also more fundamentally, how to represent the available knowledge and make it accessible for mining. RESULTS: Our integration approach is based on the premise that relationships between biological entities can be represented as a complex network. The context dependency is achieved by a judicious use of distance measures on these networks. The biological entities and the distances between them are mapped for the purpose of visualization into the lower dimensional space using the Sammon's mapping. The system implementation is based on a multi-tier architecture using a native XML database and a software tool for querying and visualizing complex biological networks. The functionality of our system is demonstrated with two examples: (1) A multiple pathway retrieval, in which, given a pathway name, the system finds all the relationships related to the query by checking available metabolic pathway, transcriptional, signaling, protein-protein interaction and ontology annotation resources and (2) A protein neighborhood search, in which given a protein name, the system finds all its connected entities within a specified depth. These two examples show that our system is able to conceptually traverse different databases to produce testable hypotheses and lead towards answers to complex biological questions.

Computational Biology↗

Meta-modelling: the appropriate solution for a family of applications.

The aim of this paper is to present an appropriate framework able to generate models and to implement them, in the objective of computerizing a family of medico-technical reports. The accelerated rate of technical development makes it necessary to design computerized applications independently of data-processing technology. This apparent paradox is a quite real challenge which needs research and development software environments to support frameworks. In this article, we present a meta-model (i.e. a generic structure - supported by the Méta-Gen software tool) ) which is able to generate various models of medico-technical reports. These models in turn are able to generate various types of instances. This meta-model is a "Meta-medical record", it is constituted of basic concepts : " User Semantic Group " to which are attached a set of " sentence-type ", a set of several corpus of variables with a set of graphs ("navigators"). Five models (echocardiography for hospital "A", echocardiography for hospital "B", gastroscopy, fibercoloscopy A.E.P.). were already generated from this "Meta-Medical Record". A beginning of implementation in echocardiography report is presented here. The advantages are a very thorough personalization of the document for the user, and a greater independence of the design diagram from the technological platform.

Echocardiography↗

Neuroanatomical term generation and comparison between two terminologies.

An approach and software tools are described for identifying and extracting compound terms (CTs), acronyms and their associated contexts from textual material that is associated with neuroanatomical atlases. A set of simple syntactic rules were appended to the output of a commercially available part of speech (POS) tagger (Qtag v 3.01) that extracts CTs and their associated context from the texts of neuroanatomical atlases. This "hybrid" parser. appears to be highly sensitive and recognized 96% of the potentially germane neuroanatomical CTs and acronyms present in the cat and primate thalamic atlases. A comparison of neuroanatomical CTs and acronymsbetween the cat and primate atlas texts was initially performed using exact-term matching. The implementation of string-matching algorithms significantly improved the identification of relevant terms and acronyms between the two domains. The End Gap Free string matcher identified 98% of CTs and the Needleman Wunsch (NW) string matcher matched 36% of acronyms between the two atlases. Combining several simple grammatical and lexical rules with the POS tagger ("hybrid parser") (1) extracted complex neuroanatomical terms and acronyms from selected cat and primate thalamic atlases and (2) and facilitated the semi-automated generation of a highly granular thalamic terminology. The implementation of string-matching algorithms (1) reconciled terminological errors generated by optical character recognition (OCR) software used to generate the neuroanatomical text information and (2) increased the sensitivity of matching neuroanatomical terms and acronyms between the two neuroanatomical domains that were generated by the "hybrid" parser.

Algorithms↗

Measurement of the volume of oral tumors by three-dimensional spiral computed tomography.

OBJECTIVE: To determine the precision and accuracy of in vitro measurements of the volume of oral tumors with three-dimensional (3D) spiral computed tomography (CT) and their precision in vivo. METHODS: Two simulated tumors made of modelling compound mixed with contrast medium were positioned medial to the mandibles of five cadaver heads and examined with subsecond spiral CT. Two observers delineated the simulated tumors twice in axial, coronal and sagittal views and then measured the volume from multiplanar reconstructed images. The software tools automatically displayed the simulated tumors in 3D-reconstructed images with the volumetric measurements. The simulated tumors were removed and their volume measured by water displacement. The volume of 15 oral tumors associated with the mandible were measured in vivo with the same imaging methods and the precision analysed. RESULTS: There were no statistically significant differences between or within observers or between imaging and physical measurements in vitro, nor between inter- and intra-observer measurements in vivo (P > 0.05). CONCLUSION: Volumetric measurements from 3D-reconstructed CT are reliable and accurate in vitro and reliable in vivo. The method is potentially useful for the management of oral neoplasms.

Aged↗

The molecule evoluator. An interactive evolutionary algorithm for the design of drug-like molecules.

We developed a software tool to design drug-like molecules, the "Molecule Evoluator", which we introduce and describe here. An atom-based evolutionary approach was used allowing both several types of mutation and crossover to occur. The novelty, we claim, is the unprecedented interactive evolution, in which the user acts as a fitness function. This brings a human being's creativity, implicit knowledge, and imagination into the design process, next to the more standard chemical rules. Proof-of-concept was demonstrated in a number of ways, both computationally and in the lab. Thus, we synthesized a number of compounds designed with the aid of the Molecule Evoluator. One of these is described here, a new chemical entity with activity on alpha-adrenergic receptors.

Algorithms↗

SPLICE, a computer program for automated extraction of information from GenBank sequence entries.

SPLICE, a software tool for the extraction of sequences from files in GenBank tape format, has been developed. The program can analyze the features table in this format and use any of the information provided to write the corresponding sequences into a standard sequence file format suitable for use with sequence analysis programs. Sequences that are present as several subsequent fragments in a single GenBank file, such as those encoding a peptide, can be spliced together by the program. Further, sequences that are present in more than one Genbank file, such as an exon which spans several different files, can also be spliced into one sequence. SPLICE runs under the MS/DOS and Unix operating systems, can be called as a sub-process by other programs and can process batches of files.

Animals↗

Developing protein documentaries and other multimedia presentations for molecular biology.

Computer-based multimedia technology for distance learning and research has come of age--the price point is acceptable, domain experts using off-the-shelf software can prepare compelling materials, and the material can be efficiently delivered via the Internet to a large audience. While not presenting any new scientific results, this paper outlines experiences with a variety of commercial and free software tools and the associated protocols we have used to prepare protein documentaries and other multimedia presentations relevant to molecular biology. A protein documentary is defined here as a description of the relationship between structure and function in a single protein or in a related family of proteins. A description using text and images which is further enhanced by the use of sound and interactive graphics. Examples of documentaries prepared to describe cAMP dependent protein kinase, the founding structural member of the protein kinase family for which there is now over 40 structures can be found at http://franklin.burnham-inst.org/rcsb. A variety of other prototype multimedia presentations for molecular biology described in this paper can be found at http://fraklin.burnham-inst.org.

Computer-Assisted Instruction↗

Sequence mapping by electronic PCR

The highly specific and sensitive PCR provides the basis for sequence-tagged sites (STSs), unique landmarks that have been used widely in the construction of genetic and physical maps of the human genome. Electronic PCR (e-PCR) refers to the process of recovering these unique sites in DNA sequences by searching for subsequences that closely match the PCR primers and have the correct order, orientation, and spacing that they could plausibly prime the amplification of a PCR product of the correct molecular weight. A software tool was developed to provide an efficient implementation of this search strategy and allow the sort of en masse searching that is required for modern genome analysis. Some sample searches were performed to demonstrate a number of factors that can affect the likelihood of obtaining a match. Analysis of one large sequence database record revealed the presence of several microsatellite and gene-based markers and allowed the exact base-pair distances among them to be calculated. This example provides a demonstration of how e-PCR can be used to integrate the growing body of genomic sequence data with existing maps, reveal relationships among markers that existed previously on different maps, and correlate genetic distances with physical distances.

Base Sequence↗

Comparison of different approaches for comparative genetic analysis using microarray hybridization.

A robust analysis of comparative genomic microarray data is critical for meaningful genomic comparison studies. In this paper, we compare our method (implemented in a new software tool, GENCOM, freely available at http://www.ifr.ac.uk/safety/gencom ) with three commonly used analysis methods: GACK (freely available at http://falkow.stanford.edu ), an empirical cut-off value of twofold difference between the fluorescence intensities after LOWESS normalization or after AVERAGE normalization in which the fluorescence intensity is divided by the average fluorescence intensity of the entire data set. Each method was tested using data sets from real experiments with prior knowledge of conserved and divergent genes. GENCOM and GACK were superior when a high proportion of genes were divergent. GENCOM was the most suitable method for the data set in which the relationship between the fluorescence intensities was not linear. GENCOM has proved robust in an analysis of all the data sets tested.

Algorithms↗

Surveying phylogenetic footprints in large gene clusters: applications to Hox cluster duplications.

Evolutionarily conserved non-coding genomic sequences represent a potentially rich source for the discovery of gene regulatory regions. Since these elements are subject to stabilizing selection they evolve much more slowly than adjacent non-functional DNA. These so-called phylogenetic footprints can be detected by comparison of the sequences surrounding orthologous genes in different species. Therefore the loss of phylogenetic footprints as well as the acquisition of conserved non-coding sequences in some lineages, but not in others, can provide evidence for the evolutionary modification of cis-regulatory elements. We introduce here a statistical model of footprint evolution that allows us to estimate the loss of sequence conservation that can be attributed to gene loss and other structural reasons. This approach to studying the pattern of cis-regulatory element evolution, however, requires the comparison of relatively long sequences from many species. We have therefore developed an efficient software tool for the identification of corresponding footprints in long sequences from multiple species. We apply this novel method to the published sequences of HoxA clusters of shark, human, and the duplicated zebrafish and Takifugu clusters as well as the published HoxB cluster sequences. We find that there is a massive loss of sequence conservation in the intergenic region of the HoxA clusters, consistent with the finding in [Chiu et al., PNAS 99 (2002) 5492]. The loss of conservation after cluster duplication is more extensive than expected from structural reasons. This suggests that binding site turnover and/or adaptive modification may also contribute to the loss of sequence conservation.

Animals↗

An artificial intelligence approach to motif discovery in protein sequences: application to steriod dehydrogenases.

MEME (Multiple Expectation-maximization for Motif Elicitation) is a unique new software tool that uses artificial intelligence techniques to discover motifs shared by a set of protein sequences in a fully automated manner. This paper is the first detailed study of the use of MEME to analyse a large, biologically relevant set of sequences, and to evaluate the sensitivity and accuracy of MEME in identifying structurally important motifs. For this purpose, we chose the short-chain alcohol dehydrogenase superfamily because it is large and phylogenetically diverse, providing a test of how well MEME can work on sequences with low amino acid similarity. Moreover, this dataset contains enzymes of biological importance, and because several enzymes have known X-ray crystallographic structures, we can test the usefulness of MEME for structural analysis. The first six motifs from MEME map onto structurally important alpha-helices and beta-strands on Streptomyces hydrogenans 20beta-hydroxysteroid dehydrogenase. We also describe MAST (Motif Alignment Search Tool), which conveniently uses output from MEME for searching databases such as SWISS-PROT and Genpept. MAST provides statistical measures that permit a rigorous evaluation of the significance of database searches with individual motifs or groups of motifs. A database search of Genpept90 by MAST with the log-odds matrix of the first six motifs obtained from MEME yields a bimodal output, demonstrating the selectivity of MAST. We show for the first time, using primary sequence analysis, that bacterial sugar epimerases are homologs of short-chain dehydrogenases. MEME and MAST will be increasingly useful as genome sequencing provides large datasets of phylogenetically divergent sequences of biomedical interest.

Alcohol Dehydrogenase↗

Virtual bronchoscopy for three--dimensional pulmonary image assessment: state of the art and future needs.

Virtual bronchoscopy is emerging as a useful approach for assessment of three-dimensional (3D) computed tomographic (CT) pulmonary images. A protocol for virtual bronchoscopic assessment of a 3D CT pulmonary image would have two main stages: (a) preprocessing of image data, which involves extracting objects of interest, defining paths through major airways, and preparing the extracted objects for 3D rendering; and (b) interactive image assessment, which involves use of graphics-based software tools such as surface-rendered views, projection images, virtual endoscopic views, tube views, oblique section images, measurement data, global two-dimensional section images, and cross-sectional views. Although a virtual bronchoscope offers a unique opportunity for exploration and quantitation, it cannot replace a real bronchoscope. Limitations of current virtual endoscopy systems include high cost, lack of visual aids beyond simulated endoscopic views, difficulty in performing interactive anatomic exploration, lack of quantitative information, use of surface rendering instead of volume rendering, and need for substantial off-line display computation. Future needs include development of fully integrated user-friendly virtual bronchoscopes, development of optimal CT protocols for generating artifact-free data sets, and improvements in automated preprocessing of 3D CT images.

Bronchoscopy↗

Scoring functions for transcription factor binding site prediction.

BACKGROUND: Transcription factor binding site (TFBS) prediction is a difficult problem, which requires a good scoring function to discriminate between real binding sites and background noise. Many scoring functions have been proposed in the literature, but it is difficult to assess their relative performance, because they are implemented in different software tools using different search methods and different TFBS representations. RESULTS: Here we compare how several scoring functions perform on both real and semi-simulated data sets in a common test environment. We have also developed two new scoring functions and included them in the comparison. The data sets are from the yeast (S. cerevisiae) genome. Our new scoring function LLBG (least likely under the background model) performs best in this study. It achieves the best average rank for the correct motifs. Scoring functions based on positional bias performed quite poorly in this study. CONCLUSION: LLBG may provide an interesting alternative to current scoring functions for TFBS prediction.

Algorithms↗

AnaBench: a Web/CORBA-based workbench for biomolecular sequence analysis.

BACKGROUND: Sequence data analyses such as gene identification, structure modeling or phylogenetic tree inference involve a variety of bioinformatics software tools. Due to the heterogeneity of bioinformatics tools in usage and data requirements, scientists spend much effort on technical issues including data format, storage and management of input and output, and memorization of numerous parameters and multi-step analysis procedures. RESULTS: In this paper, we present the design and implementation of AnaBench, an interactive, Web-based bioinformatics Analysis workBench allowing streamlined data analysis. Our philosophy was to minimize the technical effort not only for the scientist who uses this environment to analyze data, but also for the administrator who manages and maintains the workbench. With new bioinformatics tools published daily, AnaBench permits easy incorporation of additional tools. This flexibility is achieved by employing a three-tier distributed architecture and recent technologies including CORBA middleware, Java, JDBC, and JSP. A CORBA server permits transparent access to a workbench management database, which stores information about the users, their data, as well as the description of all bioinformatics applications that can be launched from the workbench. CONCLUSION: AnaBench is an efficient and intuitive interactive bioinformatics environment, which offers scientists application-driven, data-driven and protocol-driven analysis approaches. The prototype of AnaBench, managed by a team at the Université de Montréal, is accessible on-line at: http://malawimonas.bcm.umontreal.ca:8091/anabench. Please contact the authors for details about setting up a local-network AnaBench site elsewhere.

Computational Biology↗

AUGUSTUS: ab initio prediction of alternative transcripts.

AUGUSTUS is a software tool for gene prediction in eukaryotes based on a Generalized Hidden Markov Model, a probabilistic model of a sequence and its gene structure. Like most existing gene finders, the first version of AUGUSTUS returned one transcript per predicted gene and ignored the phenomenon of alternative splicing. Herein, we present a WWW server for an extended version of AUGUSTUS that is able to predict multiple splice variants. To our knowledge, this is the first ab initio gene finder that can predict multiple transcripts. In addition, we offer a motif searching facility, where user-defined regular expressions can be searched against putative proteins encoded by the predicted genes. The AUGUSTUS web interface and the downloadable open-source stand-alone program are freely available from http://augustus.gobics.de.

Alternative Splicing↗

Automated tracking of gene expression in individual cells and cell compartments.

Many intracellular signal transduction processes involve the reversible translocation from the cytoplasm to the nucleus of transcription factors. The advent of fluorescently tagged protein derivatives has revolutionized cell biology, such that it is now possible to follow the location of such protein molecules in individual cells in real time. However, the quantitative analysis of the location of such proteins in microscopic images is very time consuming. We describe CellTracker, a software tool designed for the automated measurement of the cellular location and intensity of fluorescently tagged proteins. CellTracker runs in the MS Windows environment, is freely available (at http://www.dbkgroup.org/celltracker/), and combines automated cell tracking methods with powerful image-processing algorithms that are optimized for these applications. When tested in an application involving the nuclear transcription factor NF-kappaB, CellTracker is competitive in accuracy with the manual human analysis of such images but is more than 20 times faster, even on a small task where human fatigue is not an issue. This will lead to substantial benefits for time-lapse-based high-content screening.

Animals↗

[Prognosis for life expectancy using the inductive learning method].

The aim of this paper was to find out the possibility of life expectancy achievement (LEA) prognosis in open population by using epidemiological data, to improve it and to determine the differences between regions. We were using inductive learning software tool, ASSISTANT Professional, based on modified Quinlan's inductive learning method. Data from an epidemiological study of oil/fat consuming influence on diabetes incidence in different regions in Croatia, 1887 examinees, have been used. In spite of limited number of attributes that were available, an improvement in prognosis of LEA has been made by selecting the proper attributes and by changing the limiting values of attributes (values that change the meaning of attributes). We reached the absolute accuracy of 75.54%. Even after further pruning of decision tree that lowered this value, Assistant showed 13% better result than those of random selected outcome. Differences between regions were also established that could not been explained with attributes that were used. Expanding the list of attributes and analysis of their influence in particular region can make further improvement in prognosis of LEA.

Croatia↗