Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Software Validation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,585 records · Page 88Linked to original sources

A vectorized Monte Carlo code for radiotherapy treatment planning dose calculation.

How to speed up Monte Carlo (MC) simulation in dose calculation without losing its intrinsic accuracy is one of the key issues of making a clinical MC dose engine. In this study we intensively investigated a special parallel computation technique, the vectorization technique, to boost simulation efficiency on a personal computer (PC) without extra hardware investment. A MC code, dose planning method (DPM), was extensively modified into a vectorized code, V-DPM, using the streaming single-instruction-multiple-data extension (SSE) parallel computation model. Comparative simulations were conducted for typical simulation cases in both DPM and V-DPM codes. We found that in every case the V-DPM code runs 1.5 times faster than the DPM code with variance of 0.6%.

Algorithms↗

The effect of optimization on surface dose in intensity modulated radiotherapy (IMRT).

Although IMRT has been shown clinically to increase skin doses for some patients, it has also been shown that intensity modulated delivery does not, of itself, increase skin doses. The reason for this apparent difference is that inverse planning can result in solutions that give high fluence to tangential beam segments near the skin surface, in an attempt to counter the build-up region. In cases where the clinical target volume (CTV) stops short of the skin surface, but the planning target volume (PTV) does not, there is no clinical reason to treat the skin. The CTV-PTV margin exists purely to ensure that fields are large enough to allow for geometrical uncertainties. With an objective function based on the doses to the PTV, it is possible for a plan that gives excess fluence to the skin to have a lower objective function, and hence to be preferred in an optimization. We describe a technique of plan evaluation, based on analysis of a plan by recalculating several plans in which the isocentre has been offset by a distance equal to the CTV-PTV margin. We demonstrate that changes to a plan that reduce a PTV-based objective can give a worse dose distribution to the CTV when systematic and random set-up errors are accounted for, and increase skin dose. Several possible strategies for avoiding this problem are discussed, including the use of the skin as an organ at risk, modification of the PTV to avoid the skin, and the use of 'pretend bolus' applied in planning but not in treatment. The latter gave the best results. The possibility of using the evaluation method itself, as the basis of an objective function for optimization, is discussed.

Algorithms↗

Impedance cardiography revisited.

UNLABELLED: Previously reported comparisons between cardiac output (CO) results in patients with cardiac conditions measured by thoracic impedance cardiography (TIC) versus thermodilution (TD) reveal upper and lower limits of agreement with two standard deviations (2SD) of approximately +/-2.2 l min(-1), a 44% disparity between the two technologies. We show here that if the electrodes are placed on one wrist and on a contralateral ankle instead of on the chest, a configuration designated as regional impedance cardiography (RIC), the 2SD limit of agreement between RIC and TD is +/-1.0 l min(-1), approximately 20% disparity between the two methods. To compare the performances of the TIC and RIC algorithms, the raw data of peripheral impedance changes yielded by RIC in 43 cardiac patients were used here for software processing and calculating the CO with the TIC algorithm. The 2SD between the TIC and TD was +/-1.7 l min(-1), and after annexing the correcting factors of the RIC formula to the TIC formula, the disparity between TIC and TD further declined to +/-1.25 l min(-1). CONCLUSIONS: (1) in cardiac conditions, the RIC technology is twice as accurate as TIC; (2) the advantage of RIC is the use of peripheral rather than thoracic impedance signals, supported by correcting factors.

Algorithms↗

The degenerate primer design problem.

A PCR primer sequence is called degenerate if some of its positions have several possible bases. The degeneracy of the primer is the number of unique sequence combinations it contains. We study the problem of designing a pair of primers with prescribed degeneracy that match a maximum number of given input sequences. Such problems occur when studying a family of genes that is known only in part, or is known in a related species. We prove that various simplified versions of the problem are hard, show the polynomiality of some restricted cases, and develop approximation algorithms for one variant. Based on these algorithms, we implemented a program called HYDEN for designing highly-degenerate primers for a set of genomic sequences. We report on the success of the program in an experimental scheme for identifying all human olfactory receptor (OR) genes. In that project, HYDEN was used to design primers with degeneracies up to 10(10) that amplified with high specificity many novel genes of that family, tripling the number of OR genes known at the time.

Algorithms↗

Finding composite regulatory patterns in DNA sequences.

Pattern discovery in unaligned DNA sequences is a fundamental problem in computational biology with important applications in finding regulatory signals. Current approaches to pattern discovery focus on monad patterns that correspond to relatively short contiguous strings. However, many of the actual regulatory signals are composite patterns that are groups of monad patterns that occur near each other. A difficulty in discovering composite patterns is that one or both of the component monad patterns in the group may be 'too weak'. Since the traditional monad-based motif finding algorithms usually output one (or a few) high scoring patterns, they often fail to find composite regulatory signals consisting of weak monad parts. In this paper, we present a MITRA (MIsmatch TRee Algorithm) approach for discovering composite signals. We demonstrate that MITRA performs well for both monad and composite patterns by presenting experiments over biological and synthetic data.

Algorithms↗

Variance stabilization applied to microarray data calibration and to the quantification of differential expression.

We introduce a statistical model for microarray gene expression data that comprises data calibration, the quantification of differential expression, and the quantification of measurement error. In particular, we derive a transformation h for intensity measurements, and a difference statistic Deltah whose variance is approximately constant along the whole intensity range. This forms a basis for statistical inference from microarray data, and provides a rational data pre-processing strategy for multivariate analyses. For the transformation h, the parametric form h(x)=arsinh(a+bx) is derived from a model of the variance-versus-mean dependence for microarray intensity data, using the method of variance stabilizing transformations. For large intensities, h coincides with the logarithmic transformation, and Deltah with the log-ratio. The parameters of h together with those of the calibration between experiments are estimated with a robust variant of maximum-likelihood estimation. We demonstrate our approach on data sets from different experimental platforms, including two-colour cDNA arrays and a series of Affymetrix oligonucleotide arrays.

Algorithms↗

PDP: protein domain parser.

UNLABELLED: We have developed a program for automatic identification of domains in protein three-dimensional structures. Performance of the program was assessed by three different benchmarks: (i) by comparison with the expert-curated SCOP database of structural domains; (ii) by comparison with a collection of manual domain assignments; and (iii) by comparison with a set of 55 proteins, frequently used as a benchmark for automatic domain assignment. In all these benchmarks PDP identified domains correctly in more than 80% of proteins. AVAILABILITY: http://123d.ncifcrf.gov/.

Algorithms↗

AFLPinSilico, simulating AFLP fingerprints.

SUMMARY: A drawback of the Amplified Fragment Length Polymorphism (AFLP) fingerprinting method is the difficulty to correlate the different fragments with their DNA sequence. The AFLPinSilico application presented here simulates AFLP experiments run on either cDNA or genomic sequences, producing virtual fingerprints that allow high throughput identification of AFLP fragments. The program also enables biologists to manage experiments through simulations done beforehand, thereby reducing the number of experiments that have to be run. AFLPinSilico is available through the www or as a stand-alone version, through a command line executable (available upon request, for any platform running PERL).

Algorithms↗

Deriving phylogenetic trees from the similarity analysis of metabolic pathways.

MOTIVATION: Comparative analysis of metabolic pathways in different genomes can give insights into the understanding of evolutionary and organizational relationships among species. This type of analysis allows one to measure the evolution of complete processes (with different functional roles) rather than the individual elements of a conventional analysis. We present a new technique for the phylogenetic analysis of metabolic pathways based on the topology of the underlying graphs. A distance measure between graphs is defined using the similarity between nodes of the graphs and the structural relationship between them. This distance measure is applied to the enzyme-enzyme relational graphs derived from metabolic pathways. Using this approach, pathways and group of pathways of different organisms are compared to each other and the resulting distance matrix is used to obtain a phylogenetic tree. RESULTS: We apply the method to the Citric Acid Cycle and the Glycolysis pathways of different groups of organisms, as well as to the Carbohydrate metabolic networks. Phylogenetic trees obtained from the experiments were close to existing phylogenies and revealed interesting relationships among organisms.

Algorithms↗

Computation method to identify differential allelic gene expression and novel imprinted genes.

MOTIVATION: Genomic imprinting plays an important role in both normal development and diseases. Abnormal imprinting is strongly associated with several human diseases including cancers. Most of the imprinted genes were discovered in the neighborhood of the known imprinted genes. This approach is difficult to extend to analyze the whole genome. We have decided to take a computational approach to systematically search the whole genome for the presence of mono-allelic expressed genes and imprinted genes in human genome. RESULTS: A computational method was developed to identify novel imprinted or mono-allelic genes. Individuals represented in human cDNA libraries were genotyped using Bayesian statistics, and differential expression of polymorphic alleles was identified. A significant reduction in the number of libraries that expressed both alleles, measured by Z-statistics, is a strong indicator for an imprinted or a mono-allelic gene. AVAILABILITY: The data sets are available at http://leelab.nci.nih.gov/leelab/jsp/IGDM/IGDM.html

Algorithms↗

Q-Gene: processing quantitative real-time RT-PCR data.

SUMMARY: Q-Gene is an application for the processing of quantitative real-time RT-PCR data. It offers the user the possibility to freely choose between two principally different procedures to calculate normalized gene expressions as either means of Normalized Expressions or Mean Normalized Expressions. In this contribution it will be shown that the calculation of Mean Normalized Expressions has to be used for processing simplex PCR data, while multiplex PCR data should preferably be processed by calculating Normalized Expressions. The two procedures, which are currently in widespread use and regarded as more or less equivalent alternatives, should therefore specifically be applied according to the quantification procedure used. AVAILABILITY: Web access to this program is provided at http://www.biotechniques.com/softlib/qgene.html

Algorithms↗

SATCHMO: sequence alignment and tree construction using hidden Markov models.

MOTIVATION: Aligning multiple proteins based on sequence information alone is challenging if sequence identity is low or there is a significant degree of structural divergence. We present a novel algorithm (SATCHMO) that is designed to address this challenge. SATCHMO simultaneously constructs a tree and a set of multiple sequence alignments, one for each internal node of the tree. The alignment at a given node contains all sequences within its sub-tree, and predicts which positions in those sequences are alignable and which are not. Aligned regions therefore typically get shorter on a path from a leaf to the root as sequences diverge in structure. Current methods either regard all positions as alignable (e.g. ClustalW), or align only those positions believed to be homologous across all sequences (e.g. profile HMM methods); by contrast SATCHMO makes different predictions of alignable regions in different subgroups. SATCHMO generates profile hidden Markov models at each node; these are used to determine branching order, to align sequences and to predict structurally alignable regions. RESULTS: In experiments on the BAliBASE benchmark alignment database, SATCHMO is shown to perform comparably to ClustalW and the UCSC SAM HMM software. Results using SATCHMO to identify protein domains are demonstrated on potassium channels, with implications for the mechanism by which tumor necrosis factor alpha affects potassium current. AVAILABILITY: The software is available for download from http://www.drive5.com/lobster/index.htm

Algorithms↗

ESTprep: preprocessing cDNA sequence reads.

MOTIVATION: High accuracy of data always governs the large-scale gene discovery projects. The data should not only be trustworthy but should be correctly annotated for various features it contains. Sequence errors are inherent in single-pass sequences such as ESTs obtained from automated sequencing. These errors further complicate the automated identification of EST-related sequencing. A tool is required to prepare the data prior to advanced annotation processing and submission to public databases. RESULTS: This paper describes ESTprep, a program designed to preprocess expressed sequence tag (EST) sequences. It identifies the location of features present in ESTs and allows the sequence to pass only if it meets various quality criteria. Use of ESTprep has resulted in substantial improvement in accurate EST feature identification and fidelity of results submitted to GenBank. AVAILABILITY: The program is freely available for download from http://genome.uiowa.edu/pubsoft/software.html

Algorithms↗

A hidden Markov model for progressive multiple alignment.

MOTIVATION: Progressive algorithms are widely used heuristics for the production of alignments among multiple nucleic-acid or protein sequences. Probabilistic approaches providing measures of global and/or local reliability of individual solutions would constitute valuable developments. RESULTS: We present here a new method for multiple sequence alignment that combines an HMM approach, a progressive alignment algorithm, and a probabilistic evolution model describing the character substitution process. Our method works by iterating pairwise alignments according to a guide tree and defining each ancestral sequence from the pairwise alignment of its child nodes, thus, progressively constructing a multiple alignment. Our method allows for the computation of each column minimum posterior probability and we show that this value correlates with the correctness of the result, hence, providing an efficient mean by which unreliably aligned columns can be filtered out from a multiple alignment.

Algorithms↗

A comprehensive set of protein complexes in yeast: mining large scale protein-protein interaction screens.

MOTIVATION: The analysis of protein-protein interactions allows for detailed exploration of the cellular machinery. The biochemical purification of protein complexes followed by identification of components by mass spectrometry is currently the method, which delivers the most reliable information--albeit that the data sets are still difficult to interpret. Consolidating individual experiments into protein complexes, especially for high-throughput screens, is complicated by many contaminants, the occurrence of proteins in otherwise dissimilar purifications due to functional re-use and technical limitations in the detection. A non-redundant collection of protein complexes from experimental data would be useful for biological interpretation, but manual assembly is tedious and often inconsistent. RESULTS: Here, we introduce a measure to define similarity within collections of purifications and generate a set of minimally redundant, comprehensive complexes using unsupervised clustering. AVAILABILITY: Programs and results are freely available from http://www.bork.embl-heidelberg.de/Docu/purclust/

Algorithms↗

Protein beta-turn prediction using nearest-neighbor method.

MOTIVATION: With the emerging success of protein secondary structure prediction through the applications of various statistical and machine learning techniques, similar techniques have been applied to protein beta-turn prediction. In this study, we perform protein beta-turn prediction using a k-nearest neighbor method, which is combined with a filter that uses predicted protein secondary structure information. Traditional beta-turn prediction from k-nearest neighbor method is modified to account for the unbalanced ratio of the natural occurrence of beta-turns and non-beta-turns. RESULTS: Our prediction scheme is tested on a set of 426 non-homologous protein sequences. The prediction scheme consists of two stages: k-nearest neighbor method stage and filtering stage. Variations of the k-nearest neighbor method were used to take property of beta-turns into consideration. Our filtering method uses beta-turn/non-beta-turn estimates from the k-nearest neighbor method stage and predicted protein secondary structure information from PSI-PRED in order to get new beta-turn/non-beta-turn estimate. Our result is compared with the previously best known beta-turn prediction method on the dataset of 426 non-homologous protein sequences and is shown to give slightly superior performance at significantly lower computational complexity. AVAILABILITY: Contact the author for information on the source code of the programs used.

Algorithms↗

PHIRE, a deterministic approach to reveal regulatory elements in bacteriophage genomes.

MOTIVATION: In silico genome analysis of bacteriophage genomes focuses mainly on gene discovery and functional assignment. The search for regulatory elements contained within these genome sequences is often based on prior knowledge of other genomic elements or on learning algorithms of experimentally determined data, potentially leading to a biased prediction output. The PHage In silico Regulatory Elements (PHIRE) program is a standalone program in Visual Basic. It performs an algorithmic string-based search on bacteriophage genome sequences to uncover and extract subsequence alignments hinting at regulatory elements contained within these genomes, in a deterministic manner without any prior experimental or predictive knowledge. RESULTS: The PHIRE program was tested on known phage genomes with experimentally verified regulatory elements. PHIRE was able to extract phage regulatory sequences correctly for bacteriophages T7, T3, YeO3-12 and lambda, based solely on the genome sequence. For 11 bacteriophages, new predictions of conserved phage-specific putative regulatory elements were made, further corroborating this approach. AVAILABILITY: http://www.agr.kuleuven.ac.be/logt/PHIRE.htm. Freely available for academic use. Commercial users should contact the corresponding author.

Bacteriophages↗

STAM: simple transmembrane alignment method.

MOTIVATION: The database of transmembrane protein (TMP) structures is still very small. At the same time, more and more TMP sequences are being determined. Molecular modeling is an interim answer that may bridge the gap between the two databases. The first step in homology modeling is to achieve a good alignment between the target sequences and the template structure. However, since most algorithms to obtain the alignments were constructed with data derived from globular proteins, they perform poorly when applied to TMPs. In our application, we automate the alignment procedure and design it specifically for TMP. We first identify segments likely to form transmembrane alpha-helices. We then apply different sets of criteria for transmembrane and non-transmembrane segments. For example, the penalty for insertion/deletions in the transmembrane segments is much higher than that of a penalty in the loop region. Different substitution matrices are used since the frequencies of occurrence of the various amino acids differ for transmembrane segments and water-soluble domains. RESULTS: This program leads to better models since it does not treat the protein as a single entity with the same properties, but accounts for the different physical properties of the various segments. STAM is the first multisequence alignment program that is directly targeted at transmembrane proteins. AVAILABILITY: Source code and installation package are available on request from the authors. Web access is currently implemented.

Algorithms↗