Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “software tools”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

An efficient algorithm for optimizing whole genome alignment with noise.

MOTIVATION: This paper is concerned with algorithms for aligning two whole genomes so as to identify regions that possibly contain conserved genes. Motivated by existing heuristic-based software tools, we initiate the study of an optimization problem that attempts to uncover conserved genes with a global concern. Another interesting feature in our formulation is the tolerance of noise, which also complicates the optimization problem. A brute-force approach takes time exponential in the noise level. RESULTS: We show how an insight into the optimization structure can lead to a drastic improvement in the time and space requirement [precisely, to O(k2n2) and O(k2n), respectively, where n is the size of the input and k is the noise level]. The reduced space requirement allows us to implement the new algorithm, called MaxMinCluster, on a PC. It is exciting to see that when tested with different real data sets, MaxMinCluster consistently uncovers a high percentage of conserved genes that have been published by GenBank. Its performance is indeed favorably compared to MUMmer (perhaps the most popular software tool for uncovering conserved genes in a whole-genome scale). AVAILABILITY: The source code is available from the website http://www.csis.hku.hk/~colly/maxmincluster/ detailed proof of the propositions can also be found there.

Algorithms↗

Exploiting EST databases for the development and characterization of gene-derived SSR-markers in barley (Hordeum vulgare L.).

A software tool was developed for the identification of simple sequence repeats (SSRs) in a barley ( Hordeum vulgare L.) EST (expressed sequence tag) database comprising 24,595 sequences. In total, 1,856 SSR-containing sequences were identified. Trimeric SSR repeat motifs appeared to be the most abundant type. A subset of 311 primer pairs flanking SSR loci have been used for screening polymorphisms among six barley cultivars, being parents of three mapping populations. As a result, 76 EST-derived SSR-markers were integrated into a barley genetic consensus map. A correlation between polymorphism and the number of repeats was observed for SSRs built of dimeric up to tetrameric units. 3'-ESTs yielded a higher portion of polymorphic SSRs (64%) than 5'-ESTs did. The estimated PIC (polymorphic information content) value was 0.45 +/- 0.03. Approximately 80% of the SSR-markers amplified DNA fragments in Hordeum bulbosum, followed by rye, wheat (both about 60%) and rice (40%). A subset of 38 EST-derived SSR-markers comprising 114 alleles were used to investigate genetic diversity among 54 barley cultivars. In accordance with a previous, RFLP-based, study, spring and winter cultivars, as well as two- and six-rowed barleys, formed separate clades upon PCoA analysis. The results show that: (1) with the software tool developed, EST databases can be efficiently exploited for the development of cDNA-SSRs, (2) EST-derived SSRs are significantly less polymorphic than those derived from genomic regions, (3) a considerable portion of the developed SSRs can be transferred to related species, and (4) compared to RFLP-markers, cDNA-SSRs yield similar patterns of genetic diversity.

DNA, Plant↗

Advances in functional and structural MR image analysis and implementation as FSL.

The techniques available for the interrogation and analysis of neuroimaging data have a large influence in determining the flexibility, sensitivity, and scope of neuroimaging experiments. The development of such methodologies has allowed investigators to address scientific questions that could not previously be answered and, as such, has become an important research area in its own right. In this paper, we present a review of the research carried out by the Analysis Group at the Oxford Centre for Functional MRI of the Brain (FMRIB). This research has focussed on the development of new methodologies for the analysis of both structural and functional magnetic resonance imaging data. The majority of the research laid out in this paper has been implemented as freely available software tools within FMRIB's Software Library (FSL).

Bayes Theorem↗

Research software for radiotherapy gel dosimetry.

Gel dosimetry using magnetic resonance imaging is a technique which allows measurement of three-dimensional absorbed dose distributions in radiation therapy. This paper presents details of a software tool written specifically to provide facilities to perform image processing required in research and development of gel dosimetry. Collections of magnetic resonance images can be converted into either longitudinal or transverse nuclear magnetic resonance relaxation images. The conversions are accomplished by means of a pixel-by-pixel non-linear least squares fitting algorithm. Adjustments can be made to the number of parameters used in the fitting algorithm. Fundamental image manipulation tools such as window width/level display adjustment, zooming, profile and region of interest tools are provided. The software has been developed using MATLAB (The MathWorks Inc., Natick, MA) running on Windows 95. User interaction is via a windows graphical user interface (GUI). Data such as statistics from regions of interest can be exported to other windows applications for further processing. Flexibility is incorporated in the GUI design by taking advantage of the developmental aspects of the MATLAB environment. Although originally designed for gel dosimetry, the software can be used in any application of MRI which requires production and manipulation of relaxation time images.

Gels↗

Reviewing and managing syndromic surveillance SaTScan datasets using an open source data visualization tool.

SaTScan is a popular, free software tool used to identify disease clusters early in the course of an outbreak. Using geographic and time-based surveillance data, SaTScan can generate large datasets that are difficult for humans to interpret. Tracing disease clusters through space and time using text tables is a challenging cognitive task. To simplify this process, we developed a Java-based open-source tool to transform SaTScan analytic datasets into easily navigable data visualizations.

Cluster Analysis↗

GenomeComp: a visualization tool for microbial genome comparison.

We have developed a software tool, GenomeComp, for summarizing, parsing and visualizing the genome sequences comparison results derived from voluminous BLAST textual output. With GenomeComp, the variation between genomes can be easily highlighted, such as repeat regions, insertions, deletions and rearrangements of genomic segments. This software provides a new visualizing tool for microbe comparative genomics.

Computational Biology↗

Sequence assembly with CAFTOOLS.

Large-scale genomic sequencing requires a software infrastructure to support and integrate applications that are not directly compatible. We describe a suite of software tools built around the Common Assembly Format (CAF), a comprehensive representation of a sequence assembly as a text file. These tools form the backbone of sequencing informatics at the Sanger Centre and the Genome Sequencing Center. The CAF format is intentionally flexible, and our Perl and C libraries, which parse and manipulate it, provide powerful tools for creating new applications as well as wrappers to incorporate other software. The tools are available free by anonymous FTP from ftp://ftp.sanger.ac.uk/pub/badger/.

Algorithms↗

A clinical tool for nursing.

The background of this software tool stretches back nearly 20 years to original systems techniques developed within an area of computer science called artificial intelligence. Research reflects efforts to capture the reasoning processes of healthcare experts.

Clinical Protocols↗

vcfsim: flexible simulation of all-sites VCFs with missing data.

BACKGROUND |: VCFs are the most widely used data format for encoding genetic variation. By design, standard VCFs do not include data from sites where all individuals are homozygous for the reference allele ("invariant sites") and thus do not differentiate these from sites where data are completely missing. However, missing data are a key feature of biological datasets across all domains of genomics, and many recent studies have shown that missing data can introduce a variety of statistical biases in the estimation of key population genetic parameters. A solution to this limitation is to include invariant sites in a standard VCF, creating an "all-sites VCF", exposing missing and invariant sites explicitly. One hurdle to the wider adoption of all-sites VCFs is a reliable parameterized simulation framework for generating biologically realistic all-sites VCFs. RESULTS |: Here, we introduce an open-source command line tool, vcfsim, that interfaces with the popular coalescent simulation platform msprime and provides convenience functions for simulating all-sites VCFs with variable levels of ploidy and missing data. We show that the post-processed VCFs generated using vcfsim align precisely with population genetic expectations (i.e. are statistically identical to raw msprime output), accurately introduce missing data, and permit the simulation of data with varying ploidy levels, including the simulation of intraindividual ploidy variation (e.g. heterogametic sex chromosomes) and population structures. CONCLUSIONS |: Our results vcfsim is a useful and easy-to-use tool for the benchmarking of new software tools, performing population genetic inference, training of machine learning models, and the exploration of the effects of missing data in genomics data sets.

Benchmarking↗

The Habitability Mini-Laboratory: testing the tools of space habitat architecture.

Living in the closed, confined environment of a space station for a long period and under microgravity conditions, crew members can encounter problems of a physiological and also a psychological nature. The architecture of their living quarters can greatly influence their well-being and their efficiency. Simulation of a proposed architecture and of human movement within that architecture is the most effective way to evaluate the design. Two simulation tools, Computer-Aided Design (CAD) software tools and a mock-up on a smaller scale, were used to 'construct' several proposed architectures. Those models were then evaluated to determine the investigation methods that should be used in future architectural development projects.

Computer Simulation↗

MetaCyc: a multiorganism database of metabolic pathways and enzymes.

MetaCyc is a database of metabolic pathways and enzymes located at http://MetaCyc.org/. Its goal is to serve as a metabolic encyclopedia, containing a collection of non-redundant pathways central to small molecule metabolism, which have been reported in the experimental literature. Most of the pathways in MetaCyc occur in microorganisms and plants, although animal pathways are also represented. MetaCyc contains metabolic pathways, enzymatic reactions, enzymes, chemical compounds, genes and review-level comments. Enzyme information includes substrate specificity, kinetic properties, activators, inhibitors, cofactor requirements and links to sequence and structure databases. Data are curated from the primary literature by curators with expertise in biochemistry and molecular biology. MetaCyc serves as a readily accessible comprehensive resource on microbial and plant pathways for genome analysis, basic research, education, metabolic engineering and systems biology. Querying, visualization and curation of the database is supported by SRI's Pathway Tools software. The PathoLogic component of Pathway Tools is used in conjunction with MetaCyc to predict the metabolic network of an organism from its annotated genome. SRI and the European Bioinformatics Institute employed this tool to create pathway/genome databases (PGDBs) for 165 organisms, available at the BioCyc.org website. These PGDBs also include predicted operons and pathway hole fillers.

Animals↗

Evaluating eukaryotic secreted protein prediction.

BACKGROUND: Improvements in protein sequence annotation and an increase in the number of annotated protein databases has fueled development of an increasing number of software tools to predict secreted proteins. Six software programs capable of high throughput and employing a wide range of prediction methods, SignalP 3.0, SignalP 2.0, TargetP 1.01, PrediSi, Phobius, and ProtComp 6.0, are evaluated. RESULTS: Prediction accuracies were evaluated using 372 unbiased, eukaryotic, SwissProt protein sequences. TargetP, SignalP 3.0 maximum S-score and SignalP 3.0 D-score were the most accurate single scores (90-91% accurate). The combination of a positive TargetP prediction, SignalP 2.0 maximum Y-score, and SignalP 3.0 maximum S-score increased accuracy by six percent. CONCLUSION: Single predictive scores could be highly accurate, but almost all accuracies were slightly less than those reported by program authors. Predictive accuracy could be substantially improved by combining scores from multiple methods into a single composite prediction.

Animals↗

Online annotation tool for digital mammography.

RATIONALE AND OBJECTIVES: To develop a software tool for radiologists to annotate mammograms online. MATERIALS AND METHODS: The tool is web-based with a Java development environment and composed of a Digital Image and Communication in Medicine (DICOM) image viewer, a clinical information collector, and a database. The use of the tool with a sample case is demonstrated in this article. RESULTS: An online tool for the annotation of digital mammograms for teaching and research purposes has been developed. CONCLUSION: This annotation tool can be used to provide an annotated-case library on mammography education for residents and fellows.

Education, Medical, Continuing↗

The Hunt Lab Guide to De Novo Peptide Sequence Analysis by Tandem Mass Spectrometry.

Donald Hunt has made seminal contributions to the fields of proteomics, immunology, epigenetics, and glycobiology. The foundation of every important work to come out of the Hunt Laboratory is de novo peptide sequencing. For decades, he taught hundreds of students, postdocs, engineers, and scientists to directly interpret mass spectral data. To honor his legacy and ensure that the art of de novo sequencing is not lost, we have adapted his teaching materials into "The Hunt Lab Guide to De Novo Peptide Sequence Analysis by Tandem Mass Spectrometry". In addition to the de novo sequencing tutorials, we present two freely available software tools that facilitate manual interpretation of mass spectra and validation of search results. The first, "Hunt Lab Peptide Fragment Calculator", calculates precursor and fragment mass-to-charge ratios for any peptide. The second program, "Predator Protein Fragment Calculator", was inspired in part by the fragment calculator developed in the Hunt Lab. Its capabilities are enhanced to facilitate interpretation of mass spectral data derived from intact proteins. We hope that the combination of these educational tools will continue to benefit students and researchers by empowering them to interpret data on their own.

Tandem Mass Spectrometry↗

A flexible data analysis tool for chemical genetic screens.

High-throughput assays generate immense quantities of data that require sophisticated data analysis tools. We have created a freely available software tool, SLIMS (Small Laboratory Information Management System), for chemical genetics which facilitates the collection and analysis of large-scale chemical screening data. Compound structures, physical locations, and raw data can be loaded into SLIMS. Raw data from high-throughput assays are normalized using flexible analysis protocols, and systematic spatial errors are automatically identified and corrected. Various computational analyses are performed on tested compounds, and dilution-series data are processed using standard or user-defined algorithms. Finally, published literature associated with active compounds is automatically retrieved from Medline and processed to yield potential mechanisms of actions. SLIMS provides a framework for analyzing high-throughput assay data both as a laboratory information management system and as a platform for experimental analysis.

Cyclic AMP Response Element-Binding Protein↗

Complexity: an internet resource for analysis of DNA sequence complexity.

The search for DNA regions with low complexity is one of the pivotal tasks of modern structural analysis of complete genomes. The low complexity may be preconditioned by strong inequality in nucleotide content (biased composition), by tandem or dispersed repeats or by palindrome-hairpin structures, as well as by a combination of all these factors. Several numerical measures of textual complexity, including combinatorial and linguistic ones, together with complexity estimation using a modified Lempel-Ziv algorithm, have been implemented in a software tool called 'Complexity' (http://wwwmgs.bionet.nsc.ru/mgs/programs/low_complexity/). The software enables a user to search for low-complexity regions in long sequences, e.g. complete bacterial genomes or eukaryotic chromosomes. In addition, it estimates the complexity of groups of aligned sequences.

Algorithms↗

An interactive beam-weight optimization tool for three-dimensional radiotherapy treatment planning.

A computer software tool has been developed to aid the treatment planner in selecting beam weights for three-dimensional radiotherapy treatment planning. The program consists of a feasibility search algorithm embedded in an interactive, user-friendly driving program. The feasibility search algorithm is based on the iterative relaxation algorithm of Cimmino [La Ricerca Scientifica, Vol. I, pp. 326-333 (1938)] as applied to the radiotherapy inverse problem by Altschuler et al. [Med. Phys. 13, 590 (1986)]. Relative importances of structures based upon clinical considerations can be incorporated into the algorithm. In order to speed convergence, the relaxation parameter is made to vary, with its value based upon a measure of deviation from feasibility. The interactive driving program is designed so that the treatment planner can make reasonable judgments regarding the acceptability of a plan in the event that the dose constraints yield no feasible solution. An example of the use of this program applied to a problem in three-dimensional radiotherapy treatment planning is illustrated.

Algorithms↗

The RNA Ontology Consortium: an open invitation to the RNA community.

The aim of the RNA Ontology Consortium (ROC) is to create an integrated conceptual framework-an RNA Ontology (RO)-with a common, dynamic, controlled, and structured vocabulary to describe and characterize RNA sequences, secondary structures, three-dimensional structures, and dynamics pertaining to RNA function. The RO should produce tools for clear communication about RNA structure and function for multiple uses, including the integration of RNA electronic resources into the Semantic Web. These tools should allow the accurate description in computer-interpretable form of the coupling between RNA architecture, function, and evolution. The purposes for creating the RO are, therefore, (1) to integrate sequence and structural databases; (2) to allow different computational tools to interoperate; (3) to create powerful software tools that bring advanced computational methods to the bench scientist; and (4) to facilitate precise searches for all relevant information pertaining to RNA. For example, one initial objective of the ROC is to define, identify, and classify RNA structural motifs described in the literature or appearing in databases and to agree on a computer-interpretable definition for each of these motifs. To achieve these aims, the ROC will foster communication and promote collaboration among RNA scientists by coordinating frequent face-to-face workshops to discuss, debate, and resolve difficult conceptual issues. These meeting opportunities will create new directions at various levels of RNA research. The ROC will work closely with the PDB/NDB structural databases and the Gene, Sequence, and Open Biomedical Ontology Consortia to integrate the RO with existing biological ontologies to extend existing content while maintaining interoperability.

Databases, Genetic↗