Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “software tools”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Nonuniform hybridization: a potential source of error in oligonucleotide-chip experiments with low amounts of starting material.

Low amounts of starting material are a significant limitation of gene-expression profiling of microprepared pathologic specimens. Linear RNA amplification has become the method of choice to overcome this problem. Thus, transcriptomal analyses by oligonucleotide-chips or cDNA microarrays are now feasible with labeled complementary RNA generated from total RNA samples in the lower nanogram range. However, in case of oligonucleotide-chips, it has been underestimated so far that individual complementary RNA molecules are shorter in length than and display a 3' bias in comparison to the sequence stretch represented by oligonucleotides on the chip. This can lead to incorrect interpretation of raw data. We have analyzed this problem testing ex vivo-microprepared endothelial cells with Affymetrix GeneChips U133A. Only a small subset of housekeeping genes showed adequate uniform hybridization. We developed a software tool for objective evaluation of oligonucleotide-chips based on automated analysis of as well as normalization to this subset of housekeeping genes. We analyzed the gene expression profile of microprepared lymphatic vascular endothelial cells. We show that optimized normalization prevented exclusion of angiopoietin-2, a lymphatic endothelial marker, from the lymphovascular transcriptome.

Endothelial Cells↗

Population kinetics and conditional assessment of the optimal dosage regimen using the P-PHARM software package.

The adjustment of individual dosage regimen is an adaptive control process based upon an individual response to a pharmacokinetic model. To attain this objective, it is very helpful to know the characteristics of the population to which the subject belongs, in terms of mean parameters and interindividual variability. Usually the available information consists of incomplete and sparse data. For this reason it is essential to employ a computational methodology based on non-linear mixed-effect procedures in order to obtain a population parameter estimate. A Bayesian methodology can then be applied from the population parameters to the specific data for the individual requiring a dosage adjustment (such data includes drug concentration(s) of the active drug, demographic data, etc). The result of the Bayesian calculation supplies the required individual pharmacokinetic parameters. An optimal dosage regimen can be defined on the basis of therapeutical criteria (concentration ranges) as well as practical constraints such as: the size of available unitary drug dosages, feasible drug intake times, penalties associated with expected concentrations falling outside the therapeutic concentration ranges. In this paper we present the methodology and results obtained using the P-Pharm software tool. P-Pharm implements a non-linear mixed-effect population parameter estimation algorithm based on the EM algorithm. This method allows the inclusion of explicit variables into the calculations, it implements an individual Bayesian parameter estimation procedure and also an algorithm for the conditional assessment of the optimal dosage regimen given a list of practical constraints.

Algorithms↗

A tool for sharing annotated research data: the "Category 0" UMLS (Unified Medical Language System) vocabularies.

BACKGROUND: Large biomedical data sets have become increasingly important resources for medical researchers. Modern biomedical data sets are annotated with standard terms to describe the data and to support data linking between databases. The largest curated listing of biomedical terms is the the National Library of Medicine's Unified Medical Language System (UMLS). The UMLS contains more than 2 million biomedical terms collected from nearly 100 medical vocabularies. Many of the vocabularies contained in the UMLS carry restrictions on their use, making it impossible to share or distribute UMLS-annotated research data. However, a subset of the UMLS vocabularies, designated Category 0 by UMLS, can be used to annotate and share data sets without violating the UMLS License Agreement. METHODS: The UMLS Category 0 vocabularies can be extracted from the parent UMLS metathesaurus using a Perl script supplied with this article. There are 43 Category 0 vocabularies that can be used freely for research purposes without violating the UMLS License Agreement. Among the Category 0 vocabularies are: MESH (Medical Subject Headings), NCBI (National Center for Bioinformatics) Taxonomy and ICD-9-CM (International Classification of Diseases-9-Clinical Modifiers). RESULTS: The extraction file containing all Category 0 terms and concepts is 72,581,138 bytes in length and contains 1,029,161 terms. The UMLS Metathesaurus MRCON file (January, 2003) is 151,048,493 bytes in length and contains 2,146,899 terms. Therefore the Category 0 vocabularies, in aggregate, are about half the size of the UMLS metathesaurus.A large publicly available listing of 567,921 different medical phrases were automatically coded using the full UMLS metatathesaurus and the Category 0 vocabularies. There were 545,321 phrases with one or more matches against UMLS terms while 468,785 phrases had one or more matches against the Category 0 terms. This indicates that when the two vocabularies are evaluated by their fitness to find at least one term for a medical phrase, the Category 0 vocabularies performed 86% as well as the complete UMLS metathesaurus. CONCLUSION: The Category 0 vocabularies of UMLS constitute a large nomenclature that can be used by biomedical researchers to annotate biomedical data. These annotated data sets can be distributed for research purposes without violating the UMLS License Agreement. These vocabularies may be of particular importance for sharing heterogeneous data from diverse biomedical data sets. The software tools to extract the Category 0 vocabularies are freely available Perl scripts entered into the public domain and distributed with this article.

Algorithms↗

Multimedia technologies in education.

In general multimedia is the combination of visual and audio representations. These representations could include elements of texts, graphic arts, sound, animation, and video. However, multimedia is restricted in such systems where information is digitalized and is processed by a computer. Interactive multimedia and hypermedia consist of multimedia applications that the user has more active role. Education is perhaps the most useful destination for multimedia and the place where multimedia has the most effective applications, as it enriches the learning process. Multimedia both in nursing education and in medical informatics education has several applications as well. A multimedia project can be developed even as a "stand alone" application (on CD-ROM), or on World Wide Web pages on Internet. However several technical constraints exist for developing multimedia applications on Internet. For developing multimedia projects we need hardware and software, talent and skill. The software requirements for multimedia development consist of one or more authoring systems and various editing applications for text, images, sounds and video. In this chapter different software tools for creating multimedia applications are presented. In the last part of this chapter, two examples of multimedia educational training programs are discussed. Both are "stand alone" applications (CD-ROMs). The first, examines several aspects of the electronic patient record by using videos, audio descriptions, lectures and glossary, while the second one presents several topics regarding epidemiology and epidemiological research by using graphics, sound and animation.

CD-ROM↗

A hitchhiker's guide to expressed sequence tag (EST) analysis.

Expressed sequence tag (EST) sequencing projects are underway for numerous organisms, generating millions of short, single-pass nucleotide sequence reads, accumulating in EST databases. Extensive computational strategies have been developed to organize and analyse both small- and large-scale EST data for gene discovery, transcript and single nucleotide polymorphism analysis as well as functional annotation of putative gene products. We provide an overview of the significance of ESTs in the genomic era, their properties and the applications of ESTs. Methods adopted for each step of EST analysis by various research groups have been compared. Challenges that lie ahead in organizing and analysing the ever increasing EST data have also been identified. The most appropriate software tools for EST pre-processing, clustering and assembly, database matching and functional annotation have been compiled (available online from http://biolinfo.org/EST). We propose a road map for EST analysis to accelerate the effective analyses of EST data sets. An investigation of EST analysis platforms reveals that they all terminate prior to downstream functional annotation including gene ontologies, motif/pattern analysis and pathway mapping.

Animals↗

High throughput proteome screening for biomarker detection.

Mass spectrometry-based quantitative proteomics has become an important component of biological and clinical research. Current methods, while highly developed and powerful, are falling short of their goal of routinely analyzing whole proteomes mainly because the wealth of proteomic information accumulated from prior studies is not used for the planning or interpretation of present experiments. The consequence of this situation is that in every proteomic experiment the proteome is rediscovered. In this report we describe an approach for quantitative proteomics that builds on the extensive prior knowledge of proteomes and a platform for the implementation of the method. The method is based on the selection and chemical synthesis of isotopically labeled reference peptides that uniquely identify a particular protein and the addition of a panel of such peptides to the sample mixture consisting of tryptic peptides from the proteome in question. The platform consists of a peptide separation module for the generation of ordered peptide arrays from the combined peptide sample on the sample plate of a MALDI mass spectrometer, a high throughput MALDI-TOF/TOF mass spectrometer, and a suite of software tools for the selective analysis of the targeted peptides and the interpretation of the results. Applying the method to the analysis of the human blood serum proteome we demonstrate the feasibility of using mass spectrometry-based proteomics as a high throughput screening technology for the detection and quantification of targeted proteins in a complex system.

Automation↗

Autonomous system for Web-based microarray image analysis.

Software-based feature extraction from DNA microarray images still requires human intervention on various levels. Manual adjustment of grid and metagrid parameters, precise alignment of superimposed grid templates and gene spots, or simply identification of large-scale artifacts have to be performed beforehand to reliably analyze DNA signals and correctly quantify their expression values. Ideally, a Web-based system with input solely confined to a single microarray image and a data table as output containing measurements for all gene spots would directly transform raw image data into abstracted gene expression tables. Sophisticated algorithms with advanced procedures for iterative correction function can overcome imminent challenges in image processing. Herein is introduced an integrated software system with a Java-based interface on the client side that allows for decentralized access and furthermore enables the scientist to instantly employ the most updated software version at any given time. This software tool is extended from PixClust as used in Extractiff incorporated with Java Web Start deployment technology. Ultimately, this setup is destined for high-throughput pipelines in genome-wide medical diagnostics labs or microarray core facilities aimed at providing fully automated service to its users.

Algorithms↗

Genomic pathways database and biological data management.

In this paper, we discuss the properties of biological data and challenges it poses for data management, and argue that, in order to meet the data management requirements for 'digital biology', careful integration of the existing technologies and the development of new data management techniques for biological data are needed. Based on this premise, we present PathCase: Case Pathways Database System. PathCase is an integrated set of software tools for modelling, storing, analysing, visualizing and querying biological pathways data at different levels of genetic, molecular, biochemical and organismal detail. The novel features of the system include: (i) genomic information integrated with other biological data and presented starting from pathways; (ii) design for biologists who are possibly unfamiliar with genomics, but whose research is essential for annotating gene and genome sequences with biological functions; (iii) database design, implementation and graphical tools which enable users to visualize pathways data in multiple abstraction levels and to pose exploratory queries; (iv) a wide range of different types of queries including, 'path' and 'neighbourhood queries' and graphical visualization of query outputs; and (v) an implementation that allows for web (XML)-based dissemination of query outputs (i.e. pathways data in BIOPAX format) to researchers in the community, giving them control on the use of pathways data.

Computational Biology↗

Comparison of PCR-based methods for typing Escherichia coli.

OBJECTIVE: To establish a library typing system appropriate for studying cross-transmission of Escherichia coli. METHODS: Eighteen epidemiologically unrelated isolates were genotyped by means of pulsed-field gel electrophoresis (PFGE), random amplified polymorphic DNA (RAPD), repetitive (rep) PCR, and fluorescent amplified fragment length polymorphism (AFLP). Fingerprints were analyzed either by Pearson correlation or, in the case of AFLP, by Dice coefficients employing the novel 'uncertain band' software tool from GelCompar II. During a nine-month period, 112 isolates taken from 93 patients hospitalized in five intensive care units were analyzed by use of the two most discriminative PCR typing methods. RESULTS: Genotyping by RAPD and rep-PCR revealed insufficient discrimination. Among 18 epidemiologically unrelated strains with 17 different PFGE patterns, IS3 rep-PCR and AFLP distinguished 10 and 18 types, respectively. Comparison of the different methods for analysis of AFLP fingerprints showed that the Dice coefficients, which ignore 'uncertain bands', offered the best concordance with visual interpretation. Consecutive isolates originating from the same patient differed in less than three fragments. CONCLUSIONS: AFLP analysis showed the highest discriminative capacity for PCR typing of E. coli isolates. Analysis of fingerprints employing the Dice coefficients proved the most efficient method for an automated software-based retrieval of visually indistinguishable genotypes in an AFLP fingerprint database.

Bacterial Typing Techniques↗

[Information system of the Federal Health Monitoring System. An online database offering a wide range of health information].

The information system of the Federal Health Monitoring System (http://www.gbe-bund.de) offers as an online database a wide range of coordinated health information in words and figures. High-performance software tools allow the information to be researched, analyzed flexibly, and downloaded onto the user's own PC for subsequent processing. At present, the core of the information system represents more than 650 million data from more than 100 data sources, combined in sensible indicators. Clearly arranged diagrams, easy-to-understand texts, and precise definitions complete the offer. Documentations contain additional information on more than 200 data sources, features of surveys, methodological questions of surveys, and contact persons.

Databases as Topic↗

A hybrid micro-macroevolutionary approach to gene tree reconstruction.

Gene family evolution is determined by microevolutionary processes (e.g., point mutations) and macroevolutionary processes (e.g., gene duplication and loss), yet macroevolutionary considerations are rarely incorporated into gene phylogeny reconstruction methods. We present a dynamic program to find the most parsimonious gene family tree with respect to a macroevolutionary optimization criterion, the weighted sum of the number of gene duplications and losses. The existence of a polynomial delay algorithm for duplication/loss phylogeny reconstruction stands in contrast to most formulations of phylogeny reconstruction, which are NP-complete. We next extend this result to obtain a two-phase method for gene tree reconstruction that takes both micro- and macroevolution into account. In the first phase, a gene tree is constructed from sequence data, using any of the previously known algorithms for gene phylogeny construction. In the second phase, the tree is refined by rearranging regions of the tree that do not have strong support in the sequence data to minimize the duplication/lost cost. Components of the tree with strong support are left intact. This hybrid approach incorporates both micro- and macroevolutionary considerations, yet its computational requirements are modest in practice because the two-phase approach constrains the search space. Our hybrid algorithm can also be used to resolve nonbinary nodes in a multifurcating gene tree. We have implemented these algorithms in a software tool, NOTUNG 2.0, that can be used as a unified framework for gene tree reconstruction or as an exploratory analysis tool that can be applied post hoc to any rooted tree with bootstrap values. The NOTUNG 2.0 graphical user interface can be used to visualize alternate duplication/loss histories, root trees according to duplication and loss parsimony, manipulate and annotate gene trees, and estimate gene duplication times. It also offers a command line option that enables high-throughput analysis of a large number of trees.

ATP-Binding Cassette Transporters↗

Oligonucleotide fingerprint identification for microarray-based pathogen diagnostic assays.

MOTIVATION: Advances in DNA microarray technology and computational methods have unlocked new opportunities to identify 'DNA fingerprints', i.e. oligonucleotide sequences that uniquely identify a specific genome. We present an integrated approach for the computational identification of DNA fingerprints for design of microarray-based pathogen diagnostic assays. We provide a quantifiable definition of a DNA fingerprint stated both from a computational as well as an experimental point of view, and the analytical proof that all in silico fingerprints satisfying the stated definition are found using our approach. RESULTS: The presented computational approach is implemented in an integrated high-performance computing (HPC) software tool for oligonucleotide fingerprint identification termed TOFI. We employed TOFI to identify in silico DNA fingerprints for several bacteria and plasmid sequences, which were then experimentally evaluated as potential probes for microarray-based diagnostic assays. Results and analysis of approximately 150 in silico DNA fingerprints for Yersinia pestis and 250 fingerprints for Francisella tularensis are presented. AVAILABILITY: The implemented algorithm is available upon request.

Algorithms↗

MASIA: recognition of common patterns and properties in multiple aligned protein sequences.

SUMMARY: MASIA is a software tool for pattern recognition in multiple aligned protein sequences. MASIA converts a sequence to a properties matrix that can be scanned in both vertical and horizontal steps. Consistent patterns are recognized based on the statistical significance of their occurrence. Preset macros can be altered on-line to seek any combination of amino acid properties or sequence characteristics. MASIA output can be used directly by our programs to predict the 3D structure of proteins. AVAILABILITY: Access MASIA at http://www.scsb.utmb.edu/masia/ma sia.html.

Sequence Alignment↗

A combined approach for locating box H/ACA snoRNAs in the human genome.

A novel combined method for locating box H/ACA small nucleolar RNAs (snoRNAs) is described, together with a software tool. The method adopts both a probabilistic hidden Markov model (HMM) and a minimum free energy (MFE) rule, and filters possible candidate box H/ACA snoRNAs obtained from genomic DNA sequences. With our novel method 12 known box H/ACA snoRNAs, and one strong candidate were identified in 30 nucleolar protein genomic sequences.

Algorithms↗

A novel method of automated skull registration for forensic facial approximation.

Modern forensic facial reconstruction techniques are based on an understanding of skeletal variation and tissue depths. These techniques rely upon a skilled practitioner interpreting limited data. To (i) increase the amount of data available and (ii) lessen the subjective interpretation, we use medical imaging and statistical techniques. We introduce a software tool, reality enhancement/facial approximation by computational estimation (RE/FACE) for computer-based forensic facial reconstruction. The tool applies innovative computer-based techniques to a database of human head computed tomography (CT) scans in order to derive a statistical approximation of the soft tissue structure of a questioned skull. A core component of this tool is an algorithm for removing the variation in facial structure due to skeletal variation. This method uses models derived from the CT scans and does not require manual measurement or placement of landmarks. It does not require tissue-depth tables, can be tailored to specific racial categories by adding CT scans, and removes much of the subjectivity of manual reconstructions.

Algorithms↗

Improvement-focused information technology for the clinical office practice: a patient registry for disease management.

A patient registry for disease management identifies and tracks key information concerning identified patients enrolled in a care management program. The software tool makes the proactive tracking and outreach of population-based management feasible. It is highly focused on the limited patient demographic and clinical information needed for this purpose. Automated medical records are beyond the reach of most physician practices, and many of those in use lack the features necessary for population-based management. We believe the patient registry will meet both the needs for disease-management focused support and the budget of the typical physician practice.

Chronic Disease↗

Analysis of internal loops within the RNA secondary structure in almost quadratic time.

MOTIVATION: Evaluating all possible internal loops is one of the key steps in predicting the optimal secondary structure of an RNA molecule. The best algorithm available runs in time O(L(3)), L is the length of the RNA. RESULTS: We propose a new algorithm for evaluating internal loops, its run-time is O(M(*)log(2)L), M < L(2) is a number of possible nucleotide pairings. We created a software tool Afold which predicts the optimal secondary structure of RNA molecules of lengths up to 28 000 nt, using a computer with 2 Gb RAM. We also propose algorithms constructing sets of conditionally optimal multi-branch loop free (MLF) structures, e.g. the set that for every possible pairing (x, y) contains an optimal MLF structure in which nucleotides x and y form a pair. All the algorithms have run-time O(M(*)log(2)L).

Algorithms↗

3D-VIEWER: an atlas-based system for individual and statistical investigations of the human brain.

3D-VIEWER is a new software tool for neurosurgical planning and population studies. It is based on digitized three-dimensional brain atlases derived from standard stereotactic atlases that can be adapted to an individual's brain and shown as a series of displayed images. If the patient's brain has been imaged in different modalities, the standardized anatomical information can be adapted to the individual images, which will bring the images into registration. The 3D-VIEWER can be used as a tool for combining multimodal information from the same patient. In addition, several tools are available that allow oblique views of anatomical structures or the view along the intended trajectory during a neurosurgical intervention. Furthermore, using the atlas transformation matrices, anatomical information can be determined when comparing an individual's brain to the anatomy of the atlas brain. Thus, standardized anatomical information from the atlas can be introduced into individual images. This standardization is used to perform individual-group and group-by-group comparisons between patients and normal controls in anatomical studies.

Brain Mapping↗