Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Software Validation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Finding prokaryotic genes by the 'frame-by-frame' algorithm: targeting gene starts and overlapping genes.

MOTIVATION: Tightly packed prokaryotic genes frequently overlap with each other. This feature, rarely seen in eukaryotic DNA, makes detection of translation initiation sites and, therefore, exact predictions of prokaryotic genes notoriously difficult. Improving the accuracy of precise gene prediction in prokaryotic genomic DNA remains an important open problem. RESULTS: A software program implementing a new algorithm utilizing a uniform Hidden Markov Model for prokaryotic gene prediction was developed. The algorithm analyzes a given DNA sequence in each of six possible global reading frames independently. Twelve complete prokaryotic genomes were analyzed using the new tool. The accuracy of gene finding, predicting locations of protein-coding ORFs, as well as the accuracy of precise gene prediction, and detecting the whole gene including translation initiation codon were assessed by comparison with existing annotation. It was shown that in terms of gene finding, the program performs at least as well as the previously developed tools, such as GeneMark and GLIMMER. In terms of precise gene prediction the new program was shown to be more accurate, by several percentage points, than earlier developed tools, such as GeneMark.hmm, ECOPARSE and ORPHEUS. The results of testing the program indicated the possibility of systematic bias in start codon annotation in several early sequenced prokaryotic genomes. AVAILABILITY: The new gene-finding program can be accessed through the Web site: http:@dixie.biology.gatech.edu/GeneMark/fbf.cgi CONTACT: mark@amber.gatech.edu.

Algorithms↗

Exploiting the past and the future in protein secondary structure prediction.

MOTIVATION: Predicting the secondary structure of a protein (alpha-helix, beta-sheet, coil) is an important step towards elucidating its three-dimensional structure, as well as its function. Presently, the best predictors are based on machine learning approaches, in particular neural network architectures with a fixed, and relatively short, input window of amino acids, centered at the prediction site. Although a fixed small window avoids overfitting problems, it does not permit capturing variable long-rang information. RESULTS: We introduce a family of novel architectures which can learn to make predictions based on variable ranges of dependencies. These architectures extend recurrent neural networks, introducing non-causal bidirectional dynamics to capture both upstream and downstream information. The prediction algorithm is completed by the use of mixtures of estimators that leverage evolutionary information, expressed in terms of multiple alignments, both at the input and output levels. While our system currently achieves an overall performance close to 76% correct prediction--at least comparable to the best existing systems--the main emphasis here is on the development of new algorithmic ideas. AVAILABILITY: The executable program for predicting protein secondary structure is available from the authors free of charge. CONTACT: pfbaldi@ics.uci.edu, gpollast@ics.uci.edu, brunak@cbs.dtu.dk, paolo@dsi.unifi.it.

Algorithms↗

Polymer chromosome models and Monte Carlo simulations of radiation breaking DNA.

MOTIVATION: Chromatin breakage by ionizing radiation is relevant to studies of carcinogenesis, tumor radiotherapy, biodosimetry and molecular biology. This article focuses on computer analysis of chromosome irradiation in mammlian cells. METHODS: Polymer physics and Monte Carlo numerical methods are used to develop a coarse-grained computational approach. Chromatin is modeled as a random walk on a cubic lattice, and the radiation tracks hitting the chromatin are modeled as straight lines hitting lattice sites. Each track can make a cluster of DSBs on a chromosome. RESULTS: The results obtained replace conjectured DNA fragment-size distribution functions in the recently developed RLC formalism by more mechanistically motivated distributions. The discrete lattice algorithm reproduces features of current radiation experiments relevant to chromatin on large scales. It approximates the continuous formalism and experimental data with adequate precision. It was also found that assuming either fixed chromatin with correlations among different clusters of DSBs or moving chromatin with no such correlations gives virtually identical numerical predictions.

Algorithms↗

MaxSub: an automated measure for the assessment of protein structure prediction quality.

MOTIVATION: Evaluating the accuracy of predicted models is critical for assessing structure prediction methods. Because this problem is not trivial, a large number of different assessment measures have been proposed by various authors, and it has already become an active subfield of research (Moult et al. (1997,1999) and CAFASP (Fischer et al. 1999) prediction experiments have demonstrated that it has been difficult to choose one single, 'best' method to be used in the evaluation. Consequently, the CASP3 evaluation was carried out using an extensive set of especially developed numerical measures, coupled with human-expert intervention. As part of our efforts towards a higher level of automation in the structure prediction field, here we investigate the suitability of a fully automated, simple, objective, quantitative and reproducible method that can be used in the automatic assessment of models in the upcoming CAFASP2 experiment. Such a method should (a) produce one single number that measures the quality of a predicted model and (b) perform similarly to human-expert evaluations. RESULTS: MaxSub is a new and independently developed method that further builds and extends some of the evaluation methods introduced at CASP3. MaxSub aims at identifying the largest subset of C(alpha) atoms of a model that superimpose 'well' over the experimental structure, and produces a single normalized score that represents the quality of the model. Because there exists no evaluation method for assessment measures of predicted models, it is not easy to evaluate how good our new measure is. Even though an exact comparison of MaxSub and the CASP3 assessment is not straightforward, here we use a test-bed extracted from the CASP3 fold-recognition models. A rough qualitative comparison of the performance of MaxSub vis-a-vis the human-expert assessment carried out at CASP3 shows that there is a good agreement for the more accurate models and for the better predicting groups. As expected, some differences were observed among the medium to poor models and groups. Overall, the top six predicting groups ranked using the fully automated MaxSub are also the top six groups ranked at CASP3. We conclude that MaxSub is a suitable method for the automatic evaluation of models.

Algorithms↗

BALL--rapid software prototyping in computational molecular biology. Biochemicals Algorithms Library.

MOTIVATION: Rapid software prototyping can significantly reduce development times in the field of computational molecular biology and molecular modeling. Biochemical Algorithms Library (BALL) is an application framework in C++ that has been specifically designed for this purpose. RESULTS: BALL provides an extensive set of data structures as well as classes for molecular mechanics, advanced solvation methods, comparison and analysis of protein structures, file import/export, and visualization. BALL has been carefully designed to be robust, easy to use, and open to extensions. Especially its extensibility which results from an object-oriented and generic programming approach distinguishes it from other software packages. BALL is well suited to serve as a public repository for reliable data structures and algorithms. We show in an example that the implementation of complex methods is greatly simplified when using the data structures and functionality provided by BALL.

Algorithms↗

APDB: a novel measure for benchmarking sequence alignment methods without reference alignments.

MOTIVATION: We describe APDB, a novel measure for evaluating the quality of a protein sequence alignment, given two or more PDB structures. This evaluation does not require a reference alignment or a structure superposition. APDB is designed to efficiently and objectively benchmark multiple sequence alignment methods. RESULTS: Using existing collections of reference multiple sequence alignments and existing alignment methods, we show that APDB gives results that are consistent with those obtained using conventional evaluations. We also show that APDB is suitable for evaluating sequence alignments that are structurally equivalent. We conclude that APDB provides an alternative to more conventional methods used for benchmarking sequence alignment packages.

Algorithms↗

Artificial gene networks for objective comparison of analysis algorithms.

MOTIVATION: Large-scale gene expression profiling generates data sets that are rich in observed features but poor in numbers of observations. The analysis of such data sets is a challenge that has been object of vigorous research. The algorithms in use for this purpose have been poorly documented and rarely compared objectively, posing a problem of uncertainty about the outcomes of the analyses. One way to objectively test such analysis algorithms is to apply them on computational gene network models for which the mechanisms are completely know. RESULTS: We present a system that generates random artificial gene networks according to well-defined topological and kinetic properties. These are used to run in silico experiments simulating real laboratory microarray experiments. Noise with controlled properties is added to the simulation results several times emulating measurement replicates, before expression ratios are calculated. AVAILABILITY: The data sets and kinetic models described here are available from http://www.vbi.vt.edu/~mendes/AGN/as biochemical dynamic models in SBML and Gepasi formats.

Algorithms↗

Evaluation of ontology development tools for bioinformatics.

Ontologies are being used nowadays in many areas, including bioinformatics. To assist users in developing and maintaining ontologies a number of tools have been developed. In this paper we compare four such tools, Protégé-2000, Chimaera, DAG-Edit and OilEd. As test ontologies we have used ontologies from the Gene Ontology Consortium. No system is preferred in all situations, but each system has its own strengths and weaknesses.

Computational Biology↗

GFPE: gene-finding program evaluation.

Gene-finding program evaluation (GFPE) is a set of Java classes for evaluating gene-finding programs. A command-line interface is also provided. Inputs to the program include the sequence data (in FASTA format), annotations of "actual" sequence features, and annotations of "predicted" sequence features. Annotation files are in the General Feature Format promoted by the Sanger center. GFPE calculates a number of metrics of accuracy of predictions at three levels:the coding level, the exon level, and the protein level.

Computational Biology↗

In silico analysis reveals substantial variability in the gene contents of the gamma proteobacteria LexA-regulon.

MOTIVATION: Motif-prediction algorithm capabilities for the analysis of bacterial regulatory networks and the prediction of new regulatory sites can be greatly enhanced by the use of comparative genomics approaches. In this study, we make use of a consensus-building algorithm and comparative genomics to conduct an in-depth analysis of the LexA-regulon of gamma proteobacteria, and we use the inferred results to study the evolution of this regulatory network and to examine the usefulness of the control sequences and gene contents of regulons in phylogenetic analysis. RESULTS: We show, for the first time, the substantial heterogeneity that the LexA-regulon of gamma proteobacteria displays in terms of gene content and we analyze possible branching points in its evolution. We also demonstrate the feasibility of using regulon-related information to derive sound phylogenetic inferences. AVAILABILITY: Complementary analysis data and both the source code and the Windows-executable files of the consensus-building software are available at http://www.cnm.es/~ivan/RCGScanner/

Algorithms↗

Normality of oligonucleotide microarray data and implications for parametric statistical analyses.

MOTIVATION: Experimental limitations have resulted in the popularity of parametric statistical tests as a method for identifying differentially regulated genes in microarray data sets. However, these tests assume that the data follow a normal distribution. To date, the assumption that replicate expression values for any gene are normally distributed, has not been critically addressed for Affymetrix GeneChip data. RESULTS: The normality of the expression values calculated using four different commercial and academic software packages was investigated using a data set consisting of the same target RNA applied to 59 human Affymetrix U95A GeneChips using a combination of statistical tests and visualization techniques. For the majority of probe sets obtained from each analysis suite, the expression data showed a good correlation with normality. The exception was a large number of low-expressed genes in the data set produced using Affymetrix Microarray Suite 5.0, which showed a striking non-normal distribution. In summary, our data provide strong support for the application of parametric tests to GeneChip data sets without the need for data transformation.

Algorithms↗

SaRAD: a Simple and Robust Abbreviation Dictionary.

MOTIVATION: Due to recent interest in the use of textual material to augment traditional experiments it has become necessary to automatically cluster, classify and filter natural language information. RESULTS: The Simple and Robust Abbreviation Dictionary (SaRAD) provides an easy to implement, high performance tool for the construction of a biomedical symbol dictionary. The algorithms, applied to the MEDLINE document set, result in a high quality dictionary and toolset to disambiguate abbreviation symbols automatically.

Abbreviations as Topic↗

Comparative analysis of algorithms for signal quantitation from oligonucleotide microarrays.

MOTIVATION: Recent years' exponential increase in DNA microarrays experiments has motivated the development of many signal quantitation (SQ) algorithms. These algorithms perform various transformations on the actual measurements aimed to enable researchers to compare readings of different genes quantitatively within one experiment and across separate experiments. However, it is relatively unclear whether there is a 'best' algorithm to quantitate microarray data. The ability to compare and assess such algorithms is crucial for any downstream analysis. In this work, we suggest a methodology for comparing different signal quantitation algorithms for gene expression data. Our aim is to enable researchers to compare the effect of different SQ algorithms on the specific dataset they are dealing with. We combine two kinds of tests to assess the effect of an SQ algorithm in terms of signal to noise ratio. To assess noise, we exploit redundancy within the experimental dataset to test the variability of a given SQ algorithm output. For the effect of the SQ on the signal we evaluate the overabundance of differentially expressed genes using various statistical significance tests. RESULTS: We demonstrate our analysis approach with three SQ algorithms for oligonucleotide microarrays. We compare the results of using the dChip software and the RMAExpress software to the ones obtained by using the standard Affymetrix MAS5 on a dataset containing pairs of repeated hybridizations. Our analysis suggests that dChip is more robust and stable than the MAS5 tools for about 60% of the genes while RMAExpress is able to achieve an even greater improvement in terms of signal to noise, for more than 95% of the genes.

Algorithms↗

Simulation tools for biochemical networks: evaluation of performance and usability.

MOTIVATION: Simulation of dynamic biochemical systems is receiving considerable attention due to increasing availability of experimental data of complex cellular functions. Numerous simulation tools have been developed for numerical simulation of the behavior of a system described in mathematical form. However, there exist only a few evaluation studies of these tools. Knowledge of the properties and capabilities of the simulation tools would help bioscientists in building models based on experimental data. RESULTS: We examine selected simulation tools that are intended for the simulation of biochemical systems. We choose four of them for more detailed study and perform time series simulations using a specific pathway describing the concentration of the active form of protein kinase C. We conclude that the simulation results are convergent between the chosen simulation tools. However, the tools differ in their usability, support for data transfer to other programs and support for automatic parameter estimation. From the experimentalists' point of view, all these are properties that need to be emphasized in the future.

Animals↗

Evaluation of iterative alignment algorithms for multiple alignment.

MOTIVATION: Iteration has been used a number of times as an optimization method to produce multiple alignments, either alone or in combination with other methods. Iteration has a great advantage in that it is often very simple both in terms of coding the algorithms and the complexity of the time and memory requirements. In this paper, we systematically test several different iteration strategies by comparing the results on sets of alignment test cases. RESULTS: We tested three schemes where iteration is used to improve an existing alignment. This was found to be remarkably effective and could induce a significant improvement in the accuracy of alignments from most packages. For example the average accuracy of ClustalW was improved by over 6% on the hardest test cases. Iteration was found to be even more powerful when it was directly incorporated into a progressive alignment scheme. Here, iteration was used to improve subalignments at each step of progressive alignment. The beneficial effects of iteration come, in part, from the ability to get round the usual local minimum problem with progressive alignment. This ability can also be used to help reduce the complexity of T-Coffee, without losing accuracy. Alignments can be generated, using T-Coffee, to align subgroups of sequences, which can then be iteratively improved and merged. AVAILABILITY: All of the scripts are freely available on the web at http://www.bioinf.ucd.ie/people/iain/iteration.html CONTACT: iain.wallace@ucd.ie.

Algorithms↗

Statistical evaluation and comparison of a pairwise alignment algorithm that a priori assigns the number of gaps rather than employing gap penalties.

MOTIVATION: Although pairwise sequence alignment is essential in comparative genomic sequence analysis, it has proven difficult to precisely determine the gap penalties for a given pair of sequences. A common practice is to employ default penalty values. However, there are a number of problems associated with using gap penalties. First, alignment results can vary depending on the gap penalties, making it difficult to explore appropriate parameters. Second, the statistical significance of an alignment score is typically based on a theoretical model of non-gapped alignments, which may be misleading. Finally, there is no way to control the number of gaps for a given pair of sequences, even if the number of gaps is known in advance. RESULTS: In this paper, we develop and evaluate the performance of an alignment technique that allows the researcher to assign a priori set of the number of allowable gaps, rather than using gap penalties. We compare this approach with the Smith-Waterman and Needleman-Wunsch techniques on a set of structurally aligned protein sequences. We demonstrate that this approach outperforms the other techniques, especially for short sequences (56-133 residues) with low similarity (<25%). Further, by employing a statistical measure, we show that it can be used to assess the quality of the alignment in relation to the true alignment with the associated optimal number of gaps. AVAILABILITY: The implementation of the described methods SANK_AL is available at http://cbbc.murdoch.edu.au/ CONTACT: matthew@cbbc.murdoch.edu.au.

Algorithms↗

JIGSAW: integration of multiple sources of evidence for gene prediction.

MOTIVATION: Computational gene finding systems play an important role in finding new human genes, although no systems are yet accurate enough to predict all or even most protein-coding regions perfectly. Ab initio programs can be augmented by evidence such as expression data or protein sequence homology, which improves their performance. The amount of such evidence continues to grow, but computational methods continue to have difficulty predicting genes when the evidence is conflicting or incomplete. Genome annotation pipelines collect a variety of types of evidence about gene structure and synthesize the results, which can then be refined further through manual, expert curation of gene models. RESULTS: JIGSAW is a new gene finding system designed to automate the process of predicting gene structure from multiple sources of evidence, with results that often match the performance of human curators. JIGSAW computes the relative weight of different lines of evidence using statistics generated from a training set, and then combines the evidence using dynamic programming. Our results show that JIGSAW's performance is superior to ab initio gene finding methods and to other pipelines such as Ensembl. Even without evidence from alignment to known genes, JIGSAW can substantially improve gene prediction accuracy as compared with existing methods. AVAILABILITY: JIGSAW is available as an open source software package at http://cbcb.umd.edu/software/jigsaw.

Algorithms↗

Comparative analysis of algorithms for identifying amplifications and deletions in array CGH data.

MOTIVATION: Array Comparative Genomic Hybridization (CGH) can reveal chromosomal aberrations in the genomic DNA. These amplifications and deletions at the DNA level are important in the pathogenesis of cancer and other diseases. While a large number of approaches have been proposed for analyzing the large array CGH datasets, the relative merits of these methods in practice are not clear. RESULTS: We compare 11 different algorithms for analyzing array CGH data. These include both segment detection methods and smoothing methods, based on diverse techniques such as mixture models, Hidden Markov Models, maximum likelihood, regression, wavelets and genetic algorithms. We compute the Receiver Operating Characteristic (ROC) curves using simulated data to quantify sensitivity and specificity for various levels of signal-to-noise ratio and different sizes of abnormalities. We also characterize their performance on chromosomal regions of interest in a real dataset obtained from patients with Glioblastoma Multiforme. While comparisons of this type are difficult due to possibly sub-optimal choice of parameters in the methods, they nevertheless reveal general characteristics that are helpful to the biological investigator.

Algorithms↗