Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Software Validation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

New scoring schemes for protein fold recognition based on Voronoi contacts.

MOTIVATION: The genome projects produce a wealth of protein sequences. Theoretical methods to predict possible structures and functions are needed for screening purposes, large-scale comparisons and in-depth analysis to identify worthwhile targets for further experimental research. Sequence-structure alignment is a basic tool for the identification of model folds for protein sequences and the construction of crude structural models. Empirical contact potentials (potentials of mean force) are used to optimize and evaluate such alignments. RESULTS: We propose new scoring schemes based on a contact definition derived from Voronoi decompositions of the three-dimensional coordinates of protein structures. We demonstrate that Voronoi potentials are superior to pure distance-based contact potentials with respect to recognition rate and significance for native folds. Moreover, the scoring scheme has the potential to provide a reasonable balance of detail and ion such that it is also useful for the recognition of distantly related (both homologous and non-homologous) proteins. This is demonstrated here on a set of structural alignments showing much better correspondence of native and model scores for the Voronoi potentials as compared to conventional distance-based potentials. AVAILABILITY: The potentials are made available via the program system ToPLign (URL: http://cartan.gmd.de/ToPLign.html). CONTACT: Ralf.Zimmer,Ralf.Thiele@gmd.de

Algorithms↗

FramePlus: aligning DNA to protein sequences.

MOTIVATION: Automated annotation of Expressed Sequence Tags (ESTs) is becoming increasingly important as EST databases continue to grow rapidly. A common approach to annotation is to align the gene fragments against well-documented databases of protein sequences. The sensitivity of the alignment algorithm is key to the success of such methods. RESULTS: This paper introduces a new algorithm, FramePlus, for DNA-protein sequence alignment. The SCOP database was used to develop a general framework for testing the sensitivity of such alignment algorithms when searching large databases. Using this framework, the performance of FramePlus was found to be somewhat better than other algorithms in the presence of moderate and high rates of frameshift errors, and comparable to Translated Search in the absence of sequencing errors. AVAILABILITY: The source code for FramePlus and the testing datasets are freely available at ftp.compugen.co.il/pub/research. CONTACT: raveh@compugen.co.il.

Algorithms↗

Finding prokaryotic genes by the 'frame-by-frame' algorithm: targeting gene starts and overlapping genes.

MOTIVATION: Tightly packed prokaryotic genes frequently overlap with each other. This feature, rarely seen in eukaryotic DNA, makes detection of translation initiation sites and, therefore, exact predictions of prokaryotic genes notoriously difficult. Improving the accuracy of precise gene prediction in prokaryotic genomic DNA remains an important open problem. RESULTS: A software program implementing a new algorithm utilizing a uniform Hidden Markov Model for prokaryotic gene prediction was developed. The algorithm analyzes a given DNA sequence in each of six possible global reading frames independently. Twelve complete prokaryotic genomes were analyzed using the new tool. The accuracy of gene finding, predicting locations of protein-coding ORFs, as well as the accuracy of precise gene prediction, and detecting the whole gene including translation initiation codon were assessed by comparison with existing annotation. It was shown that in terms of gene finding, the program performs at least as well as the previously developed tools, such as GeneMark and GLIMMER. In terms of precise gene prediction the new program was shown to be more accurate, by several percentage points, than earlier developed tools, such as GeneMark.hmm, ECOPARSE and ORPHEUS. The results of testing the program indicated the possibility of systematic bias in start codon annotation in several early sequenced prokaryotic genomes. AVAILABILITY: The new gene-finding program can be accessed through the Web site: http:@dixie.biology.gatech.edu/GeneMark/fbf.cgi CONTACT: mark@amber.gatech.edu.

Algorithms↗

Exploiting the past and the future in protein secondary structure prediction.

MOTIVATION: Predicting the secondary structure of a protein (alpha-helix, beta-sheet, coil) is an important step towards elucidating its three-dimensional structure, as well as its function. Presently, the best predictors are based on machine learning approaches, in particular neural network architectures with a fixed, and relatively short, input window of amino acids, centered at the prediction site. Although a fixed small window avoids overfitting problems, it does not permit capturing variable long-rang information. RESULTS: We introduce a family of novel architectures which can learn to make predictions based on variable ranges of dependencies. These architectures extend recurrent neural networks, introducing non-causal bidirectional dynamics to capture both upstream and downstream information. The prediction algorithm is completed by the use of mixtures of estimators that leverage evolutionary information, expressed in terms of multiple alignments, both at the input and output levels. While our system currently achieves an overall performance close to 76% correct prediction--at least comparable to the best existing systems--the main emphasis here is on the development of new algorithmic ideas. AVAILABILITY: The executable program for predicting protein secondary structure is available from the authors free of charge. CONTACT: pfbaldi@ics.uci.edu, gpollast@ics.uci.edu, brunak@cbs.dtu.dk, paolo@dsi.unifi.it.

Algorithms↗

Polymer chromosome models and Monte Carlo simulations of radiation breaking DNA.

MOTIVATION: Chromatin breakage by ionizing radiation is relevant to studies of carcinogenesis, tumor radiotherapy, biodosimetry and molecular biology. This article focuses on computer analysis of chromosome irradiation in mammlian cells. METHODS: Polymer physics and Monte Carlo numerical methods are used to develop a coarse-grained computational approach. Chromatin is modeled as a random walk on a cubic lattice, and the radiation tracks hitting the chromatin are modeled as straight lines hitting lattice sites. Each track can make a cluster of DSBs on a chromosome. RESULTS: The results obtained replace conjectured DNA fragment-size distribution functions in the recently developed RLC formalism by more mechanistically motivated distributions. The discrete lattice algorithm reproduces features of current radiation experiments relevant to chromatin on large scales. It approximates the continuous formalism and experimental data with adequate precision. It was also found that assuming either fixed chromatin with correlations among different clusters of DSBs or moving chromatin with no such correlations gives virtually identical numerical predictions.

Algorithms↗

MaxSub: an automated measure for the assessment of protein structure prediction quality.

MOTIVATION: Evaluating the accuracy of predicted models is critical for assessing structure prediction methods. Because this problem is not trivial, a large number of different assessment measures have been proposed by various authors, and it has already become an active subfield of research (Moult et al. (1997,1999) and CAFASP (Fischer et al. 1999) prediction experiments have demonstrated that it has been difficult to choose one single, 'best' method to be used in the evaluation. Consequently, the CASP3 evaluation was carried out using an extensive set of especially developed numerical measures, coupled with human-expert intervention. As part of our efforts towards a higher level of automation in the structure prediction field, here we investigate the suitability of a fully automated, simple, objective, quantitative and reproducible method that can be used in the automatic assessment of models in the upcoming CAFASP2 experiment. Such a method should (a) produce one single number that measures the quality of a predicted model and (b) perform similarly to human-expert evaluations. RESULTS: MaxSub is a new and independently developed method that further builds and extends some of the evaluation methods introduced at CASP3. MaxSub aims at identifying the largest subset of C(alpha) atoms of a model that superimpose 'well' over the experimental structure, and produces a single normalized score that represents the quality of the model. Because there exists no evaluation method for assessment measures of predicted models, it is not easy to evaluate how good our new measure is. Even though an exact comparison of MaxSub and the CASP3 assessment is not straightforward, here we use a test-bed extracted from the CASP3 fold-recognition models. A rough qualitative comparison of the performance of MaxSub vis-a-vis the human-expert assessment carried out at CASP3 shows that there is a good agreement for the more accurate models and for the better predicting groups. As expected, some differences were observed among the medium to poor models and groups. Overall, the top six predicting groups ranked using the fully automated MaxSub are also the top six groups ranked at CASP3. We conclude that MaxSub is a suitable method for the automatic evaluation of models.

Algorithms↗

BALL--rapid software prototyping in computational molecular biology. Biochemicals Algorithms Library.

MOTIVATION: Rapid software prototyping can significantly reduce development times in the field of computational molecular biology and molecular modeling. Biochemical Algorithms Library (BALL) is an application framework in C++ that has been specifically designed for this purpose. RESULTS: BALL provides an extensive set of data structures as well as classes for molecular mechanics, advanced solvation methods, comparison and analysis of protein structures, file import/export, and visualization. BALL has been carefully designed to be robust, easy to use, and open to extensions. Especially its extensibility which results from an object-oriented and generic programming approach distinguishes it from other software packages. BALL is well suited to serve as a public repository for reliable data structures and algorithms. We show in an example that the implementation of complex methods is greatly simplified when using the data structures and functionality provided by BALL.

Algorithms↗

APDB: a novel measure for benchmarking sequence alignment methods without reference alignments.

MOTIVATION: We describe APDB, a novel measure for evaluating the quality of a protein sequence alignment, given two or more PDB structures. This evaluation does not require a reference alignment or a structure superposition. APDB is designed to efficiently and objectively benchmark multiple sequence alignment methods. RESULTS: Using existing collections of reference multiple sequence alignments and existing alignment methods, we show that APDB gives results that are consistent with those obtained using conventional evaluations. We also show that APDB is suitable for evaluating sequence alignments that are structurally equivalent. We conclude that APDB provides an alternative to more conventional methods used for benchmarking sequence alignment packages.

Algorithms↗

Evaluation of ontology development tools for bioinformatics.

Ontologies are being used nowadays in many areas, including bioinformatics. To assist users in developing and maintaining ontologies a number of tools have been developed. In this paper we compare four such tools, Protégé-2000, Chimaera, DAG-Edit and OilEd. As test ontologies we have used ontologies from the Gene Ontology Consortium. No system is preferred in all situations, but each system has its own strengths and weaknesses.

Computational Biology↗

A prospective controlled trial of computerized decision support for lipid management in primary care.

OBJECTIVES: This study aimed to assess the uptake and effect in primary care of a computerized decision support system (DSS) for the management of hyperlipidaemia. METHOD: A prospective controlled trial was conducted in 25 practices covering a population of 150,000 in the city of Birmingham. The Primed system, a specialist developed, rule based DSS for general practice, was introduced prospectively after a 3-month baseline data collection. The main outcome measures were nine months' data on prescribing of lipid lowering agents; use of laboratory tests; and referrals to secondary care for the investigation of hyperlipidaemia. RESULTS: System use was lower than expected. A shift was observed towards requests for appropriate follow-up of previously abnormal lipid results and a greater emphasis on full lipid profiles, in line with the DSS guidelines. Referrals showed a 55% decrease on those expected (NS). The prescribing evaluation revealed a large variation between practices, but no significant alteration following system use. Views of users favoured decision support as a concept, but criticised technical problems with the system. CONCLUSIONS: Greater integration of DSS software and practice based data handling systems is needed. The mode of data capture, and hence both the content and form of knowledge representation, in DSS must take greater account of the primary care consultation process if such systems are to be of use to practitioners.

Attitude to Computers↗

Towards improvement of the accuracy and completeness of medication registration with the use of an electronic medical record (EMR).

BACKGROUND: Approximately 80% of GPs use a GP information system (GIS) and an electronic medical record (EMR) in their daily practice. To reap the full benefits of an EMR for patient care, post-graduate education and research, the data input must be well structured and accurately coded. OBJECTIVES: The quality and user-friendliness of the software positively influence the completeness and reliability of the data recorded in the GIS. To assess this in actual practice, this study examined whether or not an increase occurred in the accuracy and completeness of indication-related medication registration after the GIS's software package was upgraded. METHOD: GPs recorded data for the Registration Network Groningen (RNG) concerning four medication groups: insulin, trimethoprim, the contraceptive pill and beta-blocking agents. The completeness and accuracy of the registered data were assessed both before and after the change to the new software package. The completeness is evaluated on the basis of the indications missing for the prescribed medications. To assess accuracy, a check was made to determine whether the indications corresponded to those deemed relevant for that particular medication according to National Pharmaceutical Guidelines. RESULTS: The percentage of missing indications decreased notably, especially in the chronically prescribed medication groups. For insulin, the percentage decreased from 40.5 to 3% and for the contraceptive pill from 34.5 to 1%. For trimethoprim, the percentage decreased from 10 to 1%, and for beta-blocking agents from 22 to 1.5%. Of the indications present, the percentage of relevant indications showed a slight increase, with the largest increase observed for the contraceptive pill where the percentage rose from 86 to 96%. CONCLUSIONS: The completeness of recorded indications improved considerably after the change of software. This is due mostly to the efforts of the GPs, their practice assistants and the support of the RNG organization involved in the conversion procedure. Accuracy improved slightly, especially due to the software modifications which ensured that non-existent codes could not be entered. To summarize, with increased user-friendliness of the software, combined with the training of motivated GPs, the quality of recorded data improved.

Adrenergic Antagonists↗

An adaptive, object oriented strategy for base calling in DNA sequence analysis.

An algorithm has been developed for the determination of nucleotide sequence from data produced in fluorescence-based automated DNA sequencing instruments employing the four-color strategy. This algorithm takes advantage of object oriented programming techniques for modularity and extensibility. The algorithm is adaptive in that data sets from a wide variety of instruments and sequencing conditions can be used with good results. Confidence values are provided on the base calls as an estimate of accuracy. The algorithm iteratively employs confidence determinations from several different modules, each of which examines a different feature of the data for accurate peak identification. Modules within this system can be added or removed for increased performance or for application to a different task. In comparisons with commercial software, the algorithm performed well.

Algorithms↗

Reproducibility and accuracy of angle measurements obtained under static conditions with the Motion Analysis video system.

The development of computerized and semi-automated motion analysis systems has made the study of human motion more widely available in research and clinical settings. Although many of these systems are currently used by physical therapists, the accuracy and reproducibility of some of these systems in estimating joint angles have not been reported. In this study, the accuracy and reproducibility of angle measurements obtained by use of the Motion Analysis video system were evaluated under static conditions using a standard goniometer. Reflective markers placed on a goniometer were recorded by two video cameras at 17 angles, from 20 to 180 degrees, in 10-degree increments. Recordings of the goniometer were made at three locations within the field of view of the cameras. The intraclass correlation coefficient for each location tested was .99. Average within-trial variability was less than 0.4 degree at all locations. A linear regression of the system-calculated angles and reference angles for all locations had slopes near unity (ie, 1) and intercepts that were not statistically different from zero. A preliminary evaluation of the system under dynamic conditions revealed that distances were slightly underestimated, regardless of where the movement occurred within the calibration cube.

Algorithms↗

Evaluating computerized health information systems: hardware, software and human ware: experiences from the Northern Province, South Africa.

Despite enormous investment world-wide in computerized health information systems their overall benefits and costs have rarely been fully assessed. A major new initiative in South Africa provides the opportunity to evaluate the introduction of information technology from a global perspective and assess its impact on public health. The Northern Province is implementing a comprehensive integrated hospital information system (HIS) in all of its 42 hospitals. These include two mental health institutions, eight regional hospitals (two acting as a tertiary complex with teaching responsibilities) and 32 district hospitals. The overall goal of the HIS is to improve the efficiency and effectiveness of health (and welfare) services through the creation and use of information, for clinical, administrative and monitoring purposes. This multi-site implementation is being undertaken as a single project at a cost of R130 million (which represents 2.5 per cent of the health and welfare budget on an annual basis). The implementation process commenced on 1 September 1998 with the introduction of the system into Mankweng Hospital as the pilot site and is to be completed in the year 2001. An evaluation programme has been designed to maximize the likelihood of success of the implementation phase (formative evaluation) as well as providing an overall assessment of its benefits and costs (summative evaluation). The evaluation was designed as a form of health technology assessment; the system will have to prove its worth (in terms of cost-effectiveness) relative to other interventions. This is more extensive than the traditional form of technical assessment of hardware and software functionality, and moves into assessing the day-to-day utility of the system, the clinical and managerial environment in which it is situated (humanware), and ultimately its effects on the quality of patient care and public health. In keeping with new South African legislation the evaluation process sought to involve as many stakeholders as possible at the same time as creating a methodologically rigorous study that lived within realistic resource limits. The design chosen for the summative assessment was a randomized controlled trial (RCT) in which 24 district hospitals will receive the HIS either early or late. This is the first attempt to carry out an RCT evaluation of a multi-site implementation of an HIS in the world. Within this design the evaluation will utilize a range of qualitative and quantitative techniques over varying time scales, each addressing specific aims of the evaluation programme. In addition, it will attempt to provide an overview of the general impact on people and organizations of introducing high-technology solutions into a relatively unprepared environment. The study should help to stimulate an evaluation culture in the health and welfare services in the Northern Province as well as building the capacity to undertake such evaluations in the future.

Computers↗

Role of contact and noncontact mapping in the curative ablation of tachyarrhythmias.

Assessment of the timing of electrical activation recorded by multiple electrodes positioned in various locations within the heart has been the conventional method for mapping cardiac arrhythmias. This technique requires fluoroscopy for catheter manipulation, which in addition to being harmful (ionizing radiation), is inadequate for visualizing the complex three-dimensional cardiac anatomy and lacks reproducibility regarding localization of sites of interest. Because of these limitations, several new mapping systems that can function in a complimentary role to the conventional mapping technique, or can be used independently, have been developed. These new mapping strategies have unique advantages. They overcome the limitations of fluoroscopy by creating accurate three-dimensional intracardiac maps. The ability to localize and accurately display intracardiac catheter positioning and ablation lesion sites facilitate increasingly complex catheter ablation procedures.

Body Surface Potential Mapping↗

Influence of a computer database and problem exercises on students' knowledge of bacteriology.

This study compared the performances of students at the University of North Carolina at Chapel Hill School of Medicine who had access to sets of problem exercises and a computer database to support their learning of bacteriology with the performances of students at the University of Iowa College of Medicine who did not have such access. The study also examined the extent of a student's database use as a predictor of posttest performance. The students studied were randomly selected groups of 32-44 first-year students per year at each school; the study was conducted in three academic years (1988-1990) with some modifications in the intervention as the host environment evolved. The criterion measure was a posttest created from the same pool of problems used to generate the problem sets. The students at the intervention school scored significantly higher on the posttest in two of the three years, and overall. Also in two of the three years and overall, there was a significant relationship between the extent of a student's database use and his or her posttest score. Although the observed effects may have been due to other factors in this quasi-experimental design, the authors conclude that the use of problem sets and a computer database had a positive influence on the students' learning.

Bacteriology↗

A reexamination of the NRMP matching algorithm. National Resident Matching Program.

Most graduating medical students in the United States find their first professional appointments through the National Resident Matching Program (NRMP). This service receives rank-order lists of preferences from students and from hospitals, and then generates final assignments of students to hospitals through the use of a specific computerized matching algorithm. The author uses recent findings from the mathematics and economics literatures to demonstrate three difficulties with the NRMP's matching algorithm and the official descriptions thereof. First, the algorithm favors hospitals over students, a feature known to the NRMP since at least 1976, but, in the author's opinion, not made clear in NRMP literature for students. Second, the author argues that the NRMP's justification that its algorithm mimics orderly, noncentralized admission processes is not correct. Institutions operating under non-centralized procedures must typically make more initial offers than there are positions, in the realization that some fraction of their offers will be declined. This arrangement enlarges the choices available to many applicants, and thereby benefits them, whereas the NRMP's algorithm unrealistically assumes that no institution would ever send out any extra offers. Third, the NRMP's algorithm contains incentives for students to misrepresent their true preferences when constructing their rank-order lists. This feature is a substantial disadvantage of the current algorithm and is incorrectly described in literature distributed to students and in published articles from the NRMP.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms↗

The NRMP matching algorithm revisited: theory versus practice. National Resident Matching Program.

The authors examine the algorithm used by the National Resident Matching Program (NRMP) in its centralized matching of applicants to U.S. residency programs ("the Match"). Their goal is to evaluate the current NRMP matching algorithm to determine whether it still fulfills its intended purpose adequately and whether changes could be made that would improve the Match. They describe the basic NRMP algorithm and many of the variations of the matching process ("match variations") incorporated over the last 20 years to meet participants' requirements. An overview of the current state of the theory of preference matching is presented, including descriptions of the characteristics of stable matches in general, program-optimal and applicant-optimal matchings, and strategies for formulating preference lists. The characteristics of the current NRMP algorithm are then compared with the theoretical findings. Research conducted long after the original NRMP algorithm was devised has shown that an algorithm that produces stable matches is the best approach for matching applicants to positions. In the absence of requirements to satisfy match variations, the NRMP's deferred-acceptance algorithm produces a program-optimal stable match. When match variations, such as those handled by the NRMP, must be introduced, it is possible that no stable matching exists, and the resulting matching produced by the NRMP algorithm may not be program-optimal. The question of program-optimal versus applicant-optimal matchings is discussed. Theoretical and empirical evidence currently available suggest that differences between these two kinds of matchings are likely to be small. However, further tests and research are needed to assess the real differences in the results produced by different stable matching algorithms that produce program-optimal or applicant-optimal stable matches.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms↗