Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “software tools”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Evaluation of a target contouring protocol for 3D conformal radiotherapy in non-small cell lung cancer.

BACKGROUND: A protocol for the contouring of target volumes in lung cancer was implemented. Subsequently, a study was performed in order to determine the intra and inter-clinician variations in contoured volumes. MATERIALS AND METHODS: Six radiation oncologists (RO) contoured the gross tumour volume (GTV) and/or clinical target volume (CTV), and planning target volume (PTV) for three patients with non-small cell lung cancer (NSCLC), on two separate occasions. These were, respectively, a well-circumscribed T1N0M0 lesion, an irregularly shaped T2N0M0 lesion, and a T2N2M0 tumour. Detailed diagnostic radiology reports were provided and contours were entered into a 3D planning system. The target volumes were calculated and beams-eye view (BEV) plots were generated to visualise differences in contouring. A software tool was used to expand the GTV and CTV in three dimensions for an automatically derived PTV. RESULTS: Significant inter-RO variations in contoured target volumes were observed for all patients, and these were greater than intra-RO differences. The ratio of the largest to smallest contoured volume ranged from 1.6 for the GTV in the T1N0 lesion, to 2.0 for the PTV in the T2N2 lesion. The BEV plots revealed significant inter-RO variations in contouring the mediastinal CTV. The PTV's derived using a 3D margin programme were larger than manually contoured PTV's. These variations did not correlate with the experience of ROs. CONCLUSIONS: Despite the use of an institutional contouring protocol, significant interclinician variations persist in contouring target volumes in NSCLC. Additional measures to decrease such variations should be incorporated into clinical trials.

Carcinoma, Non-Small-Cell Lung↗

SeqState: primer design and sequence statistics for phylogenetic DNA datasets.

Choosing and designing primers based on available DNA sequence data and statistical contrasting of domains or structural features is a common routine among molecular biologists. Currently available, free software tools were found to lack desirable features related to these tasks. This was the motivation for developing a new program, SeqState. SeqState locates regions that remain to be sequenced in phylogenetic DNA datasets, evaluates user-provided primers and selects primers best suited to fill gaps in the sequences. If the primers provided by the user are unsuitable, new primers are designed. Primers can be loaded from a primer database, be supplied as part of the alignment or be entered manually. The position of internal primers is automatically localised in the loaded data file. Primers can be edited, and changes and new primers can be saved to the database. Primer sheets allow the user to view internal dimers, complements to a second primer, mismatches to all loaded sequences, and other primer characteristics. Calculation of various sequence statistics can be requested for the whole dataset or parts thereof (character sets), with standard errors estimated by bootstrapping. Insertion-deletion events can be evaluated statistically and encoded for subsequent phylogenetic analysis according to several published coding principles.

Algorithms↗

[Topographical anatomy of the female pelvis in ultrasound].

AIM: Achieving a high quality gynaecological ultrasound examination requires thorough knowledge of topographic anatomy. To date, there are no guidelines for a standardised course of the examination. The goal of the study was to define exact planes by means of cross-sectional anatomy and then to standardise the gynaecological ultrasound examination with the transabdominal, introital and transvaginal technique. METHOD: We developed a software tool based on IDL (Interactive Data Language) for the female data set of the Visible Human Project which generates free determinable planes in the volume. The organs of the female pelvis were divided into landmark- and target structures according to the ultrasonic visibility and the variability of the position, shape and structure. From this, a course for the gynaecological ultrasound examination was created and verified on 65 patients each with an inconspicuous ultrasound finding. In addition, the average duration of the examination was determined. RESULTS: The landmark structures could be demonstrated in all patients. Five planes were defined for each technique, and the course of the whole examination with 15 exact planes was described. The average duration of the examination was 4.5 minutes. CONCLUSION: As of now, the digitally reconstructed anatomical illustrations have achieved the best image resolution and quality regardless of the position of the plane in the examination volume. The standardised course of the gynaecological ultrasound examination can serve as a basis for the improvement of training quality and the evaluation of a general gynaecological ultrasound screening.

Female↗

Exchange of Veterans Affairs medical data using national and local networks.

Remote data exchange is extremely useful to a number of medical applications. It requires an infrastructure including systems, network and software tools. With such an infrastructure, existing local applications can be extended to serve national needs. There are many approaches to providing remote data exchange. Selection of an approach for an application requires balancing of various factors, including the need for rapid interactive access to data and ad hoc queries, the adequacy of access to predefined data sets, the need for an integrated view of the data, the ability to provide adequate security protection, the amount of data required, and the time frame in which data is required. The applications described here demonstrate new ways that the VA is reaping benefits from its infrastructure and its compatible integrated hospital information systems located at its facilities. The needs that have been met are also needs of private hospitals. However, in many cases the infrastructure to allow data exchange is not present. The VA's experiences may serve to establish the benefits that can be obtained by all hospitals.

CD-ROM↗

PRESTA: associating promoter sequences with information on gene expression.

BACKGROUND: Large sets of well-characterized promoter sequences are required to facilitate the understanding of promoter architecture. The major sequence databases are a prospective source of upstream regulatory regions, but suffer from inaccurate annotation. The software tool PRESTA (PRomoter EST Association) presented in this study is designed for efficient recovery of characterized and partially verified promoters from GenBank and EMBL libraries. RESULTS: The PRESTA algorithm examines the putative GenBank/EMBL promoters and automatically removes most of the poorly annotated entries. The remaining records are connected to expressed sequence tags (ESTs) through a high-stringency BLAST search. The frequency and source of recovered ESTs provide an estimate of the activity and expression pattern of the promoter, and the ESTs' 5' ends assist in transcription start-site verification. The PRESTA database provides easy access to non-redundant upstream regulatory regions recently extracted by the PRESTA algorithm. The current size of this resource is 552 human and 241 mouse promoters. Surprisingly, no overlap between the PRESTA database and the Eukaryotic Promoter Database (EPD) was detected by sequence comparison. CONCLUSIONS: The PRESTA algorithm demonstrates the principle of promoter verification by mapping EST 5' ends. The publicly available PRESTA database collects hundreds of characterized and partially verified promoter sequences and is complementary to other promoter databases.

Algorithms↗

A bank of protein family patterns for rapid identification of possible functions of amino acid sequences.

A method and software tool to develop patterns of protein families has been designed. These patterns are intended for the identification of local similarities in arbitrary amino acid sequences with proteins of the SWISS-PROT bank. The method is based on the physical, chemical and structural properties of amino acids. It assembles a 'best set' of elements (a pattern) for a given group of aligned related proteins. These elements provide discrimination between proteins of a family and representatives of other families or random sequences. The method combines the advantages of BLOCKS (automatic generation of multiple elements for protein groups), PROSITE (simplicity of element presentation) and matrices/profiles (different distinctions between amino acids for different positions of aligned sequences). Using our method, a data bank of protein family patterns, PROF_PAT, is produced. This data bank is based on the 27,752 amino acid sequences of SWISS-PROT bank release 24. The characteristics of patterns of 743 related protein groups are described. The results of comparisons of PROF_PAT patterns with the proteins of the SWISS-PROT bank are discussed.

Algorithms↗

Gene structure prediction and alternative splicing analysis using genomically aligned ESTs.

With the availability of a nearly complete sequence of the human genome, aligning expressed sequence tags (EST) to the genomic sequence has become a practical and powerful strategy for gene prediction. Elucidating gene structure is a complex problem requiring the identification of splice junctions, gene boundaries, and alternative splicing variants. We have developed a software tool, Transcript Assembly Program (TAP), to delineate gene structures using genomically aligned EST sequences. TAP assembles the joint gene structure of the entire genomic region from individual splice junction pairs, using a novel algorithm that uses the EST-encoded connectivity and redundancy information to sort out the complex alternative splicing patterns. A method called polyadenylation site scan (PASS) has been developed to detect poly-A sites in the genome. TAP uses these predictions to identify gene boundaries by segmenting the joint gene structure at polyadenylated terminal exons. Reconstructing 1007 known transcripts, TAP scored a sensitivity (Sn) of 60% and a specificity (Sp) of 92% at the exon level. The gene boundary identification process was found to be accurate 78% of the time. also reports alternative splicing patterns in EST alignments. An analysis of alternative splicing in 1124 genic regions suggested that more than half of human genes undergo alternative splicing. Surprisingly, we saw an absolute majority of the detected alternative splicing events affect the coding region. Furthermore, the evolutionary conservation of alternative splicing between human and mouse was analyzed using an EST-based approach. (See http://stl.wustl.edu/~zkan/TAP/)

Alternative Splicing↗

A computerized induction analysis of possible co-variations among different elements in human tooth enamel.

In recent decades software tools in the area of artificial intelligence have rapidly developed for use in personal computers. Interactive rule induction utilizing mathematical algorithms has become a powerful tool in data analysis and in making rules and patterns explicit. Data from a Secondary Ion Mass Spectrometry (SIMS) elemental analysis of human dental enamel were used to elucidate co-variations between certain elements. A co-variation analysis was performed employing a computerized induction analysis program, as well as a neural network program. Both analyses, confirming each other, revealed co-variations between certain elements in dental enamel in addition to exclusion of data of no importance for chosen outcomes. The results are presented in hierarchic diagrams, in which the importance for every specific element is given by its position and level in the diagram (decision tree). From the results it became evident that elements such as chlorine and sodium expressed a high co-variation level. Similarly fluorine and potassium co-varied, as well as magnesium and the trace element strontium. It was demonstrated that data from an elemental analysis could be processed by an induction analysis to reveal co-variations between certain elements in tooth enamel. The biological significance of these data is not fully understood, and further analyses in the field are needed.

Algorithms↗

Appropriate medical data categorization for data mining classification techniques.

Some data mining (DM) methods, or software tools, require normalized data, others rely on categorized data, and some can accommodate multiple data scales. Each DM technique has a specific background theory; therefore, different results are expected when applying multiple methods. The purpose of this study is to find the data format appropriate for each DM classification technique for wider applications, and efficiently to obtain trustworthy results. Considering the nature of medical data, categorical variables are sometimes useful for making decisions and can make it easier to extrapolate knowledge. In this study, three mathematical data categorization methods (Fusinter, minimum description length principle [MDLPC] and Chi-merge) were applied to accommodate five data mining classification techniques (statistics discriminant analysis, supervised classification with Neural Networks, Decision trees, Genetic supervised clustering and Bayesian classification [probability neural networks; PNN]) using a heart disease database with four types of data (continuous data, binary data, nominal data, and ordinal data). Compared with original or normalized data, data categorized by the MDLPC categorization method was found to perform better in most of the DM classification techniques used in this study. Categorical data is good for most DM classification techniques (e.g. classification of disease and non-disease groups) and is relatively easy to use for extracting medical knowledge.

Bayes Theorem↗

Composite Module Analyst: identification of transcription factor binding site combinations using genetic algorithm.

Composite Module Analyst (CMA) is a novel software tool aiming to identify promoter-enhancer models based on the composition of transcription factor (TF) binding sites and their pairs. CMA is closely interconnected with the TRANSFAC database. In particular, CMA uses the positional weight matrix (PWM) library collected in TRANSFAC and therefore provides the possibility to search for a large variety of different TF binding sites. We model the structure of the long gene regulatory regions by a Boolean function that joins several local modules, each consisting of co-localized TF binding sites. Having as an input a set of co-regulated genes, CMA builds the promoter model and optimizes the parameters of the model automatically by applying a genetic-regression algorithm. We use a multicomponent fitness function of the algorithm which includes several statistical criteria in a weighted linear function. We show examples of successful application of CMA to a microarray data on transcription profiling of TNF-alpha stimulated primary human endothelial cells. The CMA web server is freely accessible at http://www.gene-regulation.com/pub/programs/cma/CMA.html. An advanced version of CMA is also a part of the commercial system ExPlaintrade mark (www.biobase.de) designed for causal analysis of gene expression data.

Algorithms↗

BioMoby extensions to the Taverna workflow management and enactment software.

BACKGROUND: As biology becomes an increasingly computational science, it is critical that we develop software tools that support not only bioinformaticians, but also bench biologists in their exploration of the vast and complex data-sets that continue to build from international genomic, proteomic, and systems-biology projects. The BioMoby interoperability system was created with the goal of facilitating the movement of data from one Web-based resource to another to fulfill the requirements of non-expert bioinformaticians. In parallel with the development of BioMoby, the European myGrid project was designing Taverna, a bioinformatics workflow design and enactment tool. Here we describe the marriage of these two projects in the form of a Taverna plug-in that provides access to many of BioMoby's features through the Taverna interface. RESULTS: The exposed BioMoby functionality aids in the design of "sensible" BioMoby workflows, aids in pipelining BioMoby and non-BioMoby-based resources, and ensures that end-users need only a minimal understanding of both BioMoby, and the Taverna interface itself. Users are guided through the construction of syntactically and semantically correct workflows through plug-in calls to the Moby Central registry. Moby Central provides a menu of only those BioMoby services capable of operating on the data-type(s) that exist at any given position in the workflow. Moreover, the plug-in automatically and correctly connects a selected service into the workflow such that users are not required to understand the nature of the inputs or outputs for any service, leaving them to focus on the biological meaning of the workflow they are constructing, rather than the technical details of how the services will interoperate. CONCLUSION: With the availability of the BioMoby plug-in to Taverna, we believe that BioMoby-based Web Services are now significantly more useful and accessible to bench scientists than are more traditional Web Services.

Biology↗

WindowMasker: window-based masker for sequenced genomes.

MOTIVATION: Matches to repetitive sequences are usually undesirable in the output of DNA database searches. Repetitive sequences need not be matched to a query, if they can be masked in the database. RepeatMasker/Maskeraid (RM), currently the most widely used software for DNA sequence masking, is slow and requires a library of repetitive template sequences, such as a manually curated RepBase library, that may not exist for newly sequenced genomes. RESULTS: We have developed a software tool called WindowMasker (WM) that identifies and masks highly repetitive DNA sequences in a genome, using only the sequence of the genome itself. WM is orders of magnitude faster than RM because WM uses a few linear-time scans of the genome sequence, rather than local alignment methods that compare each library sequence with each piece of the genome. We validate WM by comparing BLAST outputs from large sets of queries applied to two versions of the same genome, one masked by WM, and the other masked by RM. Even for genomes such as the human genome, where a good RepBase library is available, searching the database as masked with WM yields more matches that are apparently non-repetitive and fewer matches to repetitive sequences. We show that these results hold for transcribed regions as well. WM also performs well on genomes for which much of the sequence was in draft form at the time of the analysis. AVAILABILITY: WM is included in the NCBI C++ toolkit. The source code for the entire toolkit is available at ftp://ftp.ncbi.nih.gov/toolbox/ncbi_tools++/CURRENT/. Once the toolkit source is unpacked, the instructions for building WindowMasker application in the UNIX environment can be found in file src/app/winmasker/README.build. SUPPLEMENTARY INFORMATION: Supplementary data are available at ftp://ftp.ncbi.nlm.nih.gov/pub/agarwala/windowmasker/windowmasker_suppl.pdf

Algorithms↗

ArrayQuest: a web resource for the analysis of DNA microarray data.

BACKGROUND: Numerous microarray analysis programs have been created through the efforts of Open Source software development projects. Providing browser-based interfaces that allow these programs to be executed over the Internet enhances the applicability and utility of these analytic software tools. RESULTS: Here we present ArrayQuest, a web-based DNA microarray analysis process controller. Key features of ArrayQuest are that (1) it is capable of executing numerous analysis programs such as those written in R, BioPerl and C++; (2) new analysis programs can be added to ArrayQuest Methods Library at the request of users or developers; (3) input DNA microarray data can be selected from public databases (i.e., the Medical University of South Carolina (MUSC) DNA Microarray Database or Gene Expression Omnibus (GEO)) or it can be uploaded to the ArrayQuest center-point web server into a password-protected area; and (4) analysis jobs are distributed across computers configured in a backend cluster. To demonstrate the utility of ArrayQuest we have populated the methods library with methods for analysis of Affymetrix DNA microarray data. CONCLUSION: ArrayQuest enables browser-based implementation of DNA microarray data analysis programs that can be executed on a Linux-based platform. Importantly, ArrayQuest is a platform that will facilitate the distribution and implementation of new analysis algorithms and is therefore of use to both developers of analysis applications as well as users. ArrayQuest is freely available for use at http://proteogenomics.musc.edu/arrayquest.html.

Algorithms↗

Measurements of characteristics of time pattern in dose delivery in step-and-shoot IMRT.

BACKGROUND AND PURPOSE: Although intensity-modulated radiotherapy (IMRT) has already shown its clinical benefit, there are some issues which are not yet fully understood. Among these is the question whether the protracted dose delivery due to the lowered dose rate has any radiobiological consequences. To investigate this question, an exact characterization of dose rate profiles in typical clinical plans is needed. Furthermore, such a characterization may lead to an increased knowledge how to improve IMRT technically. MATERIAL AND METHODS: A new IMRT phantom which allows precise measurement of up to nine points of interest simultaneously with pin-point ionization chambers was developed. To examine dose rates, a new software tool (GRAYHOUND) was developed which can measure doses in short time intervals of up to 0.5 s. 250 points in four clinical IMRT plans were examined. A set of parameters was defined to describe the dose rate profiles including the effective fraction time (eft, which is the percentage of the fraction time in which any dose is delivered to a specific point), and a quotient of the percentage of dose delivered in high dose pulses (> 0.01 Gy/s) divided by the percentage of fraction time needed to deliver this dose (d(HD)/t(HD)). RESULTS: These quotients are excellent markers for the inhomogeneity of dose rate delivery in IMRT. In both parameters a wide variance in points of the same plan and between different plans was found. For example, eft ranged between 11.6% and 37.3% in high dose points and the time in which high dose rates are delivered to a single high dose point ranged between 3.6% and 10.1% of total fraction time. CONCLUSIONS: These data show a great inhomogeneity of dose rates not only between different plans but also between different points in the same plan. Biological investigations are needed to quantify the relevance of these inhomogeneities. The parameters which are introduced in this work may be suitable to compare different optimization algorithms in IMRT.

Algorithms↗

[Migration analysis of cemented Müller polyethylene acetabular cups versus cement-free Zweymüller screw-attached acetabular cups].

OBJECTIVE: Is the cementless Zweymüller hip cup superior to the cemented Müller cup? METHOD: This article presents a radiographic analysis of 25 cemented Müller acetabular cups versus 22 cementless Zweymüller cups using the Einbildröntgenanalyse (EBRA), a software tool for radiographic measurement of acetabular cup migration. In addition, we determined the effects of the cup anteversion and inclination, the polyethylene wear, the lateral bone coverage of the acetabular cup, the position of the center of rotation, and individual factors on the incidence of cup migration. RESULTS: The incidence of cup migration was 64% in the cementless group and 48% in the cemented group after a mean follow-up of 6 years. The average migration rate was 0.33 mm/a for cementless Zweymüller cups and 0.38 mm/a for cemented Müller cups. Cup anteversion and inclination showed no effect on the incidence of cup migration. The combination metal-polyethylene (0.17 mm per year) demonstrated a significantly higher wear rate in comparison to the ceramic-polyethylene combination (0.11 mm per year). Incompletely lateral covered cups demonstrated a significantly higher incidence of cup migration. Cranial or medial deviations of the center of rotation up to 5 mm are tolerable, in contrast to caudal or lateral deviations that lead to a significantly higher incidence of cup migration. CONCLUSION: The superiority of the cementless Zweymüller cup was not observed. We recommend a complete lateral bone coverage of the hip cup. Cranial and medial deviations of the center of rotation up to 5 mm are tolerable. In the present study the polyethylene wear of the ceramic-polyethylene combination was significantly less as compared with the metal-polyethylene combination.

Acetabulum↗

An OGSA-based integration of life-scientific resources for drug discovery.

OBJECTIVES: The rapid progress of life-scientific research has the potential to dramatically change the paradigm of drug discovery. Efficient utilization of life-scientific resources, i.e., databases and analytic software tools, poses a challenging issue with regard to the reduction of time and cost in the drug discovery process. In this paper, a variety of heterogeneous Web-based life-scientific resources are integrated toward the improvement of drug discovery performance. METHODS: For the integration of heterogeneous life-scientific resources, a database federation technique based on three-layer architecture has been utilized. With the federation technique, life-scientific resources are integrated step by step through database layers, database integration layers and analysis layers to encapsulate complexity and heterogeneity. In this study, we have taken advantage of the latest Grid technology based on OGSA (Open Grid Services Architecture) for the implementation of our approach. RESULTS: The actual case of life-scientific resources for drug discovery demonstrates that our prototype system developed with the proposed technique works well for the identification process of candidate compounds to a target protein. In other words, the prototype system allows a researcher to retrieve candidate compounds with less effort than before. CONCLUSIONS: The usefulness of the prototypic system represents the ability of our approach to integrate heterogeneous life-scientific resources, which have the potential to dramatically improve efficiency in drug discovery, resulting in the shortening of drug development. On the other hand, the system requires further consideration from the aspect of practical use. Dynamic aggregation of the resources is one example of such a consideration.

Biological Science Disciplines↗

FlexE: efficient molecular docking considering protein structure variations.

Side-chain or even backbone adjustments upon docking of different ligands to the same protein structure, a phenomenon known as induced fit, are frequently observed. Sometimes point mutations within the active site influence the ligand binding of proteins. Furthermore, for homology derived protein structures there are often ambiguities in side-chain placement and uncertainties in loop modeling which may be critical for docking applications. Nevertheless, only very few molecular docking approaches have taken into account such variations in protein structures. We present the new software tool FlexE which addresses the problem of protein structure variations during docking calculations. FlexE can dock flexible ligands into an ensemble of protein structures which represents the flexibility, point mutations, or alternative models of a protein. The FlexE approach is based on a united protein description generated from the superimposed structures of the ensemble. For varying parts of the protein, discrete alternative conformations are explicitly taken into account, which can be combinatorially joined to create new valid protein structures.FlexE was evaluated using ten protein structure ensembles containing 105 crystal structures from the PDB and one modeled structure with 60 ligands in total. For 50 ligands (83 %) FlexE finds a placement with an RMSD to the crystal structure below 2.0 A. In all cases our results are of similar quality to the best solution obtained by sequentially docking the ligands into all protein structures (cross docking). In most cases the computing time is significantly lower than the accumulated run times for the single structures. FlexE takes about five and a half minutes on average for placing one ligand into the united protein description on a common workstation. The example of the aldose reductase demonstrates the necessity of considering protein structure variations for docking calculations. We docked three potent inhibitors into four protein structures with substantial conformational changes within the active site. Using only one rigid protein structure for screening would have missed potential inhibitors whereas all inhibitors can be docked taking all protein structures into account.

Aldehyde Reductase↗

High-throughput mass spectrometric discovery of protein post-translational modifications.

The availability of genome sequences, affordable mass spectrometers and high-resolution two-dimensional gels has made possible the identification of hundreds of proteins from many organisms by peptide mass fingerprinting. However, little attention has been paid to how information generated by these means can be utilised for detailed protein characterisation. Here we present an approach for the systematic characterisation of proteins using mass spectrometry and a software tool FindMod. This tool, available on the internet at http://www.expasy.ch/sprot/findmod.html , examines peptide mass fingerprinting data for mass differences between empirical and theoretical peptides. Where mass differences correspond to a post-translational modification, intelligent rules are applied to predict the amino acids in the peptide, if any, that might carry the modification. FindMod rules were constructed by examining 5153 incidences of post-translational modifications documented in the SWISS-PROT database, and for the 22 post-translational modifications currently considered (acetylation, amidation, biotinylation, C-mannosylation, deamidation, flavinylation, farnesylation, formylation, geranyl-geranylation, gamma-carboxyglutamic acids, hydroxylation, lipoylation, methylation, myristoylation, N -acyl diglyceride (tripalmitate), O-GlcNAc, palmitoylation, phosphorylation, pyridoxal phosphate, phospho-pantetheine, pyrrolidone carboxylic acid, sulphation) a total of 29 different rules were made. These consider which amino acids can carry a modification, whether the modification occurs on N-terminal, C-terminal or internal amino acids, and the type of organisms on which the modification can be found. We illustrate the utility of the approach with proteins from 2-D gels of Escherichia coli and sheep wool, where post-translational modifications predicted by FindMod were confirmed by MALDI post-source decay peptide fragmentation. As the approach is amenable to automation, it presents a potentially large-scale means of protein characterisation in proteome projects.

Acetylation↗