Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Software Validation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Highly complex proofs and implications of such proofs.

Conventional wisdom says the ideal proof should be short, simple, and elegant. However there are now examples of very long, complicated proofs, and as mathematics continues to mature, more examples are likely to appear. Such proofs raise various issues. For example it is impossible to write out a very long and complicated argument without error, so is such a 'proof' really a proof? What conditions make complex proofs necessary, possible, and of interest? Is the mathematics involved in dealing with information rich problems qualitatively different from more traditional mathematics?

Algorithms↗

The justification of mathematical statements.

The uncompromising ethos of pure mathematics in the early post-war period was that any theorem should be provided with a proof which the reader could and should check. Two things have made this no longer realistic: (i) the appearance of increasingly long and complicated proofs and (ii) the involvement of computers. This paper discusses what compromises the mathematical community needs to make as a result.

Algorithms↗

Functionality of system components: conservation of protein function in protein feature space.

Many protein features useful for prediction of protein function can be predicted from sequence, including posttranslational modifications, subcellular localization, and physical/chemical properties. We show here that such protein features are more conserved among orthologs than paralogs, indicating they are crucial for protein function and thus subject to selective pressure. This means that a function prediction method based on sequence-derived features may be able to discriminate between proteins with different function even when they have highly similar structure. Also, such a method is likely to perform well on organisms other than the one on which it was trained. We evaluate the performance of such a method, ProtFun, which relies on protein features as its sole input, and show that the method gives similar performance for most eukaryotes and performs much better than anticipated on archaea and bacteria. From this analysis, we conclude that for the posttranslational modifications studied, both the cellular use and the sequence motifs are conserved within Eukarya.

Amino Acid Sequence↗

Hierarchical scaffolding with Bambus.

The output of a genome assembler generally comprises a collection of contiguous DNA sequences (contigs) whose relative placement along the genome is not defined. A procedure called scaffolding is commonly used to order and orient these contigs using paired read information. This ordering of contigs is an essential step when finishing and analyzing the data from a whole-genome shotgun project. Most recent assemblers include a scaffolding module; however, users have little control over the scaffolding algorithm or the information produced. We thus developed a general-purpose scaffolder, called Bambus, which affords users significant flexibility in controlling the scaffolding parameters. Bambus was used recently to scaffold the low-coverage draft dog genome data. Most significantly, Bambus enables the use of linking data other than that inferred from mate-pair information. For example, the sequence of a completed genome can be used to guide the scaffolding of a related organism. We present several applications of Bambus: support for finishing, comparative genomics, analysis of the haplotype structure of genomes, and scaffolding of a mammalian genome at low coverage. Bambus is available as an open-source package from our Web site.

Algorithms↗

First pass annotation of promoters on human chromosome 22.

The publication of the first almost complete sequence of a human chromosome (chromosome 22) is a major milestone in human genomics. Together with the sequence, an excellent annotation of genes was published which certainly will serve as an information resource for numerous future projects. We noted that the annotation did not cover regulatory regions; in particular, no promoter annotation has been provided. Here we present an analysis of the complete published chromosome 22 sequence for promoters. A recent breakthrough in specific in silico prediction of promoter regions enabled us to attempt large-scale prediction of promoter regions on chromosome 22. Scanning of sequence databases revealed only 20 experimentally verified promoters, of which 10 were correctly predicted by our approach. Nearly 40% of our 465 predicted promoter regions are supported by the currently available gene annotation. Promoter finding also provides a biologically meaningful method for "chromosomal scaffolding", by which long genomic sequences can be divided into segments starting with a gene. As one example, the combination of promoter region prediction with exon/intron structure predictions greatly enhances the specificity of de novo gene finding. The present study demonstrates that it is possible to identify promoters in silico on the chromosomal level with sufficient reliability for experimental planning and indicates that a wealth of information about regulatory regions can be extracted from current large-scale (megabase) sequencing projects. Results are available on-line at http://genomatix.gsf.de/chr22/.

Algorithms↗

Gene structure prediction and alternative splicing analysis using genomically aligned ESTs.

With the availability of a nearly complete sequence of the human genome, aligning expressed sequence tags (EST) to the genomic sequence has become a practical and powerful strategy for gene prediction. Elucidating gene structure is a complex problem requiring the identification of splice junctions, gene boundaries, and alternative splicing variants. We have developed a software tool, Transcript Assembly Program (TAP), to delineate gene structures using genomically aligned EST sequences. TAP assembles the joint gene structure of the entire genomic region from individual splice junction pairs, using a novel algorithm that uses the EST-encoded connectivity and redundancy information to sort out the complex alternative splicing patterns. A method called polyadenylation site scan (PASS) has been developed to detect poly-A sites in the genome. TAP uses these predictions to identify gene boundaries by segmenting the joint gene structure at polyadenylated terminal exons. Reconstructing 1007 known transcripts, TAP scored a sensitivity (Sn) of 60% and a specificity (Sp) of 92% at the exon level. The gene boundary identification process was found to be accurate 78% of the time. also reports alternative splicing patterns in EST alignments. An analysis of alternative splicing in 1124 genic regions suggested that more than half of human genes undergo alternative splicing. Surprisingly, we saw an absolute majority of the detected alternative splicing events affect the coding region. Furthermore, the evolutionary conservation of alternative splicing between human and mouse was analyzed using an EST-based approach. (See http://stl.wustl.edu/~zkan/TAP/)

Alternative Splicing↗

Basecalling with LifeTrace.

A pivotal step in electrophoresis sequencing is the conversion of the raw, continuous chromatogram data into the actual sequence of discrete nucleotides, a process referred to as basecalling. We describe a novel algorithm for basecalling implemented in the program LifeTrace. Like Phred, currently the most widely used basecalling software program, LifeTrace takes processed trace data as input. It was designed to be tolerant to variable peak spacing by means of an improved peak-detection algorithm that emphasizes local chromatogram information over global properties. LifeTrace is shown to generate high-quality basecalls and reliable quality scores. It proved particularly effective when applied to MegaBACE capillary sequencing machines. In a benchmark test of 8372 dye-primer MegaBACE chromatograms, LifeTrace generated 17% fewer substitution errors, 16% fewer insertion/deletion errors, and 2.4% more aligned bases to the finished sequence than did Phred. For two sets totaling 6624 dye-terminator chromatograms, the performance improvement was 15% fewer substitution errors, 10% fewer insertion/deletion errors, and 2.1% more aligned bases. The processing time required by LifeTrace is comparable to that of Phred. The predicted quality scores were in line with observed quality scores, permitting direct use for quality clipping and in silico single nucleotide polymorphism (SNP) detection. Furthermore, we introduce a new type of quality score associated with every basecall: the gap-quality. It estimates the probability of a deletion error between the current and the following basecall. This additional quality score improves detection of single basepair deletions when used for locating potential basecalling errors during the alignment. We also describe a new protocol for benchmarking that we believe better discerns basecaller performance differences than methods previously published.

Algorithms↗

Identification of candidate disease genes by EST alignments, synteny, and expression and verification of Ensembl genes on rat chromosome 1q43-54.

We aligned Incyte ESTs and publicly available sequences to the rat genome and analyzed rat chromosome 1q43-54, a region in which several quantitative trait loci (QTLs) have been identified, including renal disease, diabetes, hypertension, body weight, and encephalomyelitis. Within this region, which contains 255 Ensembl gene predictions, the aligned sequences clustered into 568 Incyte genes and gene fragments. Of the Incyte genes, 261 (46%) overlapped 184 (72%) of the Ensembl gene predictions, whereas 307 were unique to Incyte. The rat-to-human syntenic map displays rearrangement of this region on rat chr. 1 onto human chromosomes 9 and 10. The mapping of corresponding human disease phenotypes to either one of these chromosomes has allowed us to focus in on genes associated with disease phenotypes. As an example, we have used the syntenic information for the rat Rf-1 disease region and the orthologous human ESRD disease region to reduce the size of the original rat QTL to only 11.5 Mb. Using the syntenic information in combination with expression data from ESTs and microarrays, we have selected a set of 66 candidate disease genes for Rf-1. The combination of the results from these different analyses represents a powerful approach for narrowing the number of genes that could play a role in the development of complex diseases.

Animals↗

Single nucleotide polymorphisms associated with rat expressed sequences.

Single nucleotide polymorphisms (SNPs) are the most common source of genetic variation in populations and are thus most likely to account for the majority of phenotypic and behavioral differences between individuals or strains. Although the rat is extensively studied for the latter, data on naturally occurring polymorphisms are mostly lacking. We have used publicly available sequences consisting of whole-genome shotgun (WGS), expressed sequence tag (EST), and mRNA data as a source for the in silico identification of SNPs in gene-coding regions and have identified a large collection of 33,305 high-quality candidate SNPs. Experimental verification of 471 candidate SNPs using a limited set of rat isolates revealed a confirmation rate of approximately 50%. Although the majority of SNPs were identified between Sprague-Dawley (EST data) and Brown Norway (WGS data) strains, we found that 66% of the verified variations are common among different rat strains. All SNPs were extensively annotated, including chromosomal and genetic map information, and nonsynonymous SNPs were analyzed by SIFT and PolyPhen prediction programs for their potential deleterious effect on protein function. Interestingly, we retrieved three SNPs from the database that result in the introduction of a premature stop codon and that could be confirmed experimentally. Two of these "in silico-identified knockouts" reside in interesting QTL regions. Data are publicly available via a Web interface (http://cascad.niob.knaw.nl), allowing simple and advanced search queries.

Animals↗

Regulog analysis: detection of conserved regulatory networks across bacteria: application to Staphylococcus aureus.

A transcriptional regulatory network encompasses sets of genes (regulons) whose expression states are directly altered in response to an activating signal, mediated by trans-acting regulatory proteins and cis-acting regulatory sequences. Enumeration of these network components is an essential step toward the creation of a framework for systems-based analysis of biological processes. Profile-based methods for the detection of cis-regulatory elements are often applied to predict regulon members, but they suffer from poor specificity. In this report we describe Regulogger, a novel computational method that uses comparative genomics to eliminate spurious members of predicted gene regulons. Regulogger produces regulogs, sets of coregulated genes for which the regulatory sequence has been conserved across multiple organisms. The quantitative method assigns a confidence score to each predicted regulog member on the basis of the degree of conservation of protein sequence and regulatory mechanisms. When applied to a reference collection of regulons from Escherichia coli, Regulogger increased the specificity of predictions up to 25-fold over methods that use cis-element detection in isolation. The enhanced specificity was observed across a wide range of biologically meaningful parameter combinations, indicating a robust and broad utility for the method. The power of computational pattern discovery methods coupled with Regulogger to unravel transcriptional networks was demonstrated in an analysis of the genome of Staphylococcus aureus. A total of 125 regulogs were found in this organism, including both well-defined functional groups and a subset with unknown functions.

Bacillus subtilis↗

Mauve: multiple alignment of conserved genomic sequence with rearrangements.

As genomes evolve, they undergo large-scale evolutionary processes that present a challenge to sequence comparison not posed by short sequences. Recombination causes frequent genome rearrangements, horizontal transfer introduces new sequences into bacterial chromosomes, and deletions remove segments of the genome. Consequently, each genome is a mosaic of unique lineage-specific segments, regions shared with a subset of other genomes and segments conserved among all the genomes under consideration. Furthermore, the linear order of these segments may be shuffled among genomes. We present methods for identification and alignment of conserved genomic DNA in the presence of rearrangements and horizontal transfer. Our methods have been implemented in a software package called Mauve. Mauve has been applied to align nine enterobacterial genomes and to determine global rearrangement structure in three mammalian genomes. We have evaluated the quality of Mauve alignments and drawn comparison to other methods through extensive simulations of genome evolution.

Chromosomes, Bacterial↗

ProbCons: Probabilistic consistency-based multiple sequence alignment.

To study gene evolution across a wide range of organisms, biologists need accurate tools for multiple sequence alignment of protein families. Obtaining accurate alignments, however, is a difficult computational problem because of not only the high computational cost but also the lack of proper objective functions for measuring alignment quality. In this paper, we introduce probabilistic consistency, a novel scoring function for multiple sequence comparisons. We present ProbCons, a practical tool for progressive protein multiple sequence alignment based on probabilistic consistency, and evaluate its performance on several standard alignment benchmark data sets. On the BAliBASE, SABmark, and PREFAB benchmark alignment databases, ProbCons achieves statistically significant improvement over other leading methods while maintaining practical speed. ProbCons is publicly available as a Web resource.

Algorithms↗

An optimal controller for an electric ventricular-assist device: theory, implementation, and testing.

This paper addresses the development and testing of an optimal position feedback controller for the Penn State electric ventricular-assist device (EVAD). The control law is designed to minimize the expected value of the EVAD's power consumption for a targeted patient population. The closed-loop control law is implemented on an Intel 8096 microprocessor and in vitro test runs show that this controller improves the EVAD's efficiency by 15-21%, when compared with the performance of the currently used feedforward control scheme.

Computers↗

Optimal discrimination and classification of neuronal action potential waveforms from multiunit, multichannel recordings using software-based linear filters.

We describe advanced protocols for the discrimination and classification of neuronal spike waveforms within multichannel electrophysiological recordings. The programs are capable of detecting and classifying the spikes from multiple, simultaneously active neurons, even in situations where there is a high degree of spike waveform superposition on the recording channels. The protocols are based on the derivation of an optimal linear filter for each individual neuron. Each filter is tuned to selectively respond to the spike waveform generated by the corresponding neuron, and to attenuate noise and the spike waveforms from all other neurons. The protocol is essentially an extension of earlier work [1], [13], [18]. However, the protocols extend the power and utility of the original implementations in two significant respects. First, a general single-pass automatic template estimation algorithm was derived and implemented. Second, the filters were implemented within a software environment providing a greatly enhanced functional organization and user interface. The utility of the analysis approach was demonstrated on samples of multiunit electrophysiological recordings from the cricket abdominal nerve cord.

Action Potentials↗

A new algorithm for autoreconstruction of catheters in computed tomography-based brachytherapy treatment planning.

This paper describes innovative software for automatic reconstruction, which we term autoreconstruction, of plastic and metallic brachytherapy catheters using computed tomography (CT) data. No such automatic facility has previously existed in any treatment planning software. The patient data consists of a set of post-implantation CT images with the catheters in situ in their final positions. This new software solves those difficulties which arise when the catheters are intersecting or when loop techniques are used. With the software algorithms, catheter reconstruction time is significantly reduced and accuracy is also improved when compared with that achieved using the classical manual method of CT-slice-by-CT-slice reconstruction.

Algorithms↗

An efficient method for quantifying the multichannel EEG spatial-temporal complexity.

The complexity index (delta) quantifies the intrinsic dimensionality of the global complexity of a point set, and was shown to be able to characterize electroencephalogram spatial-temporal features. The complexity index is conceptually comprehensible and easily implemented, yet, it is time consuming. In this paper, we present an efficient computational method based on the projection of the high-dimensional state-space points onto a one-dimensional axis. The computational time decreases by at least 50%, without affecting the measure accuracy.

Algorithms↗

Volume delineation by fusion of fuzzy sets obtained from multiplanar tomographic images.

Techniques of three-dimensional (3-D) volume delineation from tomographic medical imaging are usually based on 2-D contour definition. For a given structure, several different contours can be obtained depending on the segmentation method used or the user's choice. The goal of this work is to develop a new method that reduces the inaccuracies generally observed. A minimum volume that is certain to be included in the volume concerned (membership degree mu = 1), and a maximum volume outside which no part of the volume is expected to be found (membership degree mu = 0), are defined semi-automatically. The intermediate fuzziness region (0 < mu < 1) is processed using the theory of possibility. The resulting fuzzy volume is obtained after data fusion from multiplanar slices. The influence of the contrast-to-noise ratio was tested on simulated images. The influence of slice thickness as well as the accuracy of the method were studied on phantoms. The absolute volume error was less than 2% for phantom volumes of 2-8 cm3, whereas the values obtained with conventional methods were much larger than the actual volumes. Clinical experiments were conducted, and the fuzzy logic method gave a volume lower than that obtained with the conventional method. Our fuzzy logic method allows volumes to be determined with better accuracy and reproducibility.

Artificial Intelligence↗