Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

The European Bioinformatics Institute (EBI) databases.

The European Bioinformatics Institute (EBI) maintains and distributes the EMBL Nucleotide Sequence database, Europe's primary nucleotide sequence data resource. The EBI also maintains and distributes the SWISS-PROT Protein Sequence database, in collaboration with Amos Bairoch of the University of Geneva. Over fifty additional specialist molecular biology databases, as well as software and documentation of interest to molecular biologists are available. The EBI network services include database searching and sequence similarity searching facilities.

Amino Acid Sequence↗

Abstract shapes of RNA.

The function of a non-protein-coding RNA is often determined by its structure. Since experimental determination of RNA structure is time-consuming and expensive, its computational prediction is of great interest, and efficient solutions based on thermodynamic parameters are known. Frequently, however, the predicted minimum free energy structures are not the native ones, leading to the necessity of generating suboptimal solutions. While this can be accomplished by a number of programs, the user is often confronted with large outputs of similar structures, although he or she is interested in structures with more fundamental differences, or, in other words, with different abstract shapes. Here, we formalize the concept of abstract shapes and introduce their efficient computation. Each shape of an RNA molecule comprises a class of similar structures and has a representative structure of minimal free energy within the class. Shape analysis is implemented in the program RNAshapes. We applied RNAshapes to the prediction of optimal and suboptimal abstract shapes of several RNAs. For a given energy range, the number of shapes is considerably smaller than the number of structures, and in all cases, the native structures were among the top shape representatives. This demonstrates that the researcher can quickly focus on the structures of interest, without processing up to thousands of near-optimal solutions. We complement this study with a large-scale analysis of the growth behaviour of structure and shape spaces. RNAshapes is available for download and as an online version on the Bielefeld Bioinformatics Server.

5' Untranslated Regions↗

Whole-genome automated assembly pipeline for Chlamydia trachomatis strains from reference, in vitro and clinical samples using the integrated CtGAP pipeline.

Whole genome sequencing (WGS) is pivotal for the molecular characterization of Chlamydia trachomatis (Ct)-the leading bacterial cause of sexually transmitted infections and infectious blindness worldwide. Ct WGS can inform epidemiologic, public health and outbreak investigations of these human-restricted pathogens. However, challenges persist in generating high-quality genomes for downstream analyses given its obligate intracellular nature and difficulty with in vitro propagation. No single tool exists for the entirety of Ct genome assembly, necessitating the adaptation of multiple programs with varying success. Compounding this issue is the absence of reliable Ct reference strain genomes. We, therefore, developed CtGAP-Chlamydia trachomatisGenome Assembly Pipeline-as an integrated 'one-stop-shop' pipeline for assembly and characterization of Ct genome sequencing data from various sources including isolates, in vitro samples, clinical swabs and urine. CtGAP, written in Snakemake, enables read quality statistics output, adapter and quality trimming, host read removal, de novo and reference-guided assembly, contig scaffolding, selective ompA, multi-locus-sequence and plasmid typing, phylogenetic tree construction, and recombinant genome identification. Twenty Ct reference genomes were also generated. Successfully validated on a diverse collection of 363 samples containing Ct, CtGAP represents a novel pipeline requiring minimal bioinformatics expertise with easy adaptation for use with other bacterial species.

Chlamydia trachomatis↗

HotSwap for bioinformatics: a STRAP tutorial.

BACKGROUND: Bioinformatics applications are now routinely used to analyze large amounts of data. Application development often requires many cycles of optimization, compiling, and testing. Repeatedly loading large datasets can significantly slow down the development process. We have incorporated HotSwap functionality into the protein workbench STRAP, allowing developers to create plugins using the Java HotSwap technique. RESULTS: Users can load multiple protein sequences or structures into the main STRAP user interface, and simultaneously develop plugins using an editor of their choice such as Emacs. Saving changes to the Java file causes STRAP to recompile the plugin and automatically update its user interface without requiring recompilation of STRAP or reloading of protein data. This article presents a tutorial on how to develop HotSwap plugins. STRAP is available at http://strapjava.de and http://www.charite.de/bioinf/strap. CONCLUSION: HotSwap is a useful and time-saving technique for bioinformatics developers. HotSwap can be used to efficiently develop bioinformatics applications that require loading large amounts of data into memory.

Algorithms↗

How well do we understand the clusters found in microarray data?

We wished to quantify the state-of-the-art of our understanding of clusters in microarray data. To do this we systematically compared the clusters produced on sets of microarray data using a representative set of clustering algorithms (hierarchical, k-means, and a modified version of QT_CLUST) with the annotation schemes MIPS, GeneOntology and GenProtEC. We assumed that if a cluster reflected known biology its members would share related ontological annotations. This assumption is the basis of "guilt-by-association" and is commonly used to assign the putative function of proteins. To statistically measure the relationship between cluster and annotation we developed a new predictive discriminatory measure. We found that the clusters found in microarray data do not in general agree with functional annotation classes. Although many statistically significant relationships can be found, the majority of clusters are not related to known biology (as described in annotation ontologies). This implies that use of guilt-by-association is not supported by annotation ontologies. Depending on the estimate of the amount of noise in the data, our results suggest that bioinformatics has only codified a small proportion of the biological knowledge required to understand microarray data.

Algorithms↗

Application of bioinformatics in search for cleavable peptides of SARS-CoV M(pro) and chemical modification of octapeptides.

According to the "distorted key" theory as elaborated in a review article years ago (Chou, K.C.: Analytical Biochemistry, 1996, 233, 1-14), the knowledge of the cleavable peptides by SARS-CoV M(pro) (severe acute respiratory syndrome coronavirus main proteinase) can provide very useful insights on developing drugs against SARS. In view of this, the softwares, ZCURVE_CoV 1.0 and ZCURVE_CoV 2.0 (http://tubic.tju.edu.cn/sars/), developed recently for SARS-Coronavirus are used to analyze the 36 complete SARS-Coronavirus RNA sequences in the gene bank NCBI (http://www.ncbi.nlm.nih.gov/) from different sources for protein coding genes, and to search for the cleavage sites of SARS-CoV M(pro) in polyproteins pp1a and pp1ab. A total of 396 cleavage points are found in the 36 SARS-Coronavirus and 11 cleavable octapeptides abstracted from the 396 cleavage sites. The statistical distributions of amino acids for the cleavable octapeptides at the subsites R4, R3, R2, R1, R1', R2', R3' and R4' are calculated. The cleavage-specific positions are on R2, R1 and R1', and the positions R3 and R4 are featured by some certain specificity for SARS-CoV M(pro). The structural characters of amino acid residues around the cleavage-specific positions are discussed. Two most promising octapeptides, i.e., NH(2)-ATLQ downward arrowAIAS-COOH and NH(2)-ATLQ downward arrowAENV-COOH, are selected to be the candidates for chemical modification, converting into the inhibitors of SARS-CoV M(pro). A possible strategy to convert a cleavable octapeptide by SARS enzyme into a drug candidate against SARS is elucidated.

Amino Acid Sequence↗

Extending SEE for large-scale phenotyping of mouse open-field behavior.

SEE (Software for the Exploration of Exploration) is a visualization and analysis tool designed for the study of open-field behavior in rodents. In this paper, I present new extensions of SEE that were designed to facilitate its use for mouse behavioral phenotyping and, especially, for the problems of discrimination of genotypes and the replication of results across laboratories and experimental conditions. These extensions were specifically designed to promote a new approach in behavioral phenotyping, reminiscent of the approach that has been successfully employed in bioinformatics during recent years. The path coordinates of all animals from many experiments are stored in a database. SEE can be used to query, visualize, and analyze any desirable subsection of this database and to design new measures (endpoints) with increasingly better discriminative power and replicability. The use of the new extensions is demonstrated here in the analysis of results from several experiments and laboratories, with an emphasis on this approach.

Animals↗

A new set of bioinformatics tools for genome projects.

A new tool called System for Automated Bacterial Integrated Annotation--SABIA (SABIA being a very well-known bird in Brazil) was developed for the assembly and annotation of bacterial genomes. This system performs automatic tasks of assembly analysis, ORFs identification/analysis, and extragenic region analyses. Genome assembly and contig automatic annotation data are also available in the same working environment. The system integrates several public domains and newly developed software programs capable of dealing with several types of databases, and it is portable to other operational systems. These programs interact with most of the well-known biological database/softwares, such as Glimmer, Genemark, the BLAST family programs, InterPro, COG, Kegg, PSORT, GO, tRNAScan and RBSFinder, and can also be used to identify metabolic pathways.

Brazil↗

Investigating semantic similarity measures across the Gene Ontology: the relationship between sequence and annotation.

MOTIVATION: Many bioinformatics data resources not only hold data in the form of sequences, but also as annotation. In the majority of cases, annotation is written as scientific natural language: this is suitable for humans, but not particularly useful for machine processing. Ontologies offer a mechanism by which knowledge can be represented in a form capable of such processing. In this paper we investigate the use of ontological annotation to measure the similarities in knowledge content or 'semantic similarity' between entries in a data resource. These allow a bioinformatician to perform a similarity measure over annotation in an analogous manner to those performed over sequences. A measure of semantic similarity for the knowledge component of bioinformatics resources should afford a biologist a new tool in their repertoire of analyses. RESULTS: We present the results from experiments that investigate the validity of using semantic similarity by comparison with sequence similarity. We show a simple extension that enables a semantic search of the knowledge held within sequence databases. AVAILABILITY: Software available from http://www.russet.org.uk.

Artificial Intelligence↗

Proteomics and bioinformatics approaches for identification of serum biomarkers to detect breast cancer.

BACKGROUND: Surface-enhanced laser desorption/ionization (SELDI) is an affinity-based mass spectrometric method in which proteins of interest are selectively adsorbed to a chemically modified surface on a biochip, whereas impurities are removed by washing with buffer. This technology allows sensitive and high-throughput protein profiling of complex biological specimens. METHODS: We screened for potential tumor biomarkers in 169 serum samples, including samples from a cancer group of 103 breast cancer patients at different clinical stages [stage 0 (n = 4), stage I (n = 38), stage II (n = 37), and stage III (n = 24)], from a control group of 41 healthy women, and from 25 patients with benign breast diseases. Diluted serum samples were applied to immobilized metal affinity capture Ciphergen ProteinChip Arrays previously activated with Ni2+. Proteins bound to the chelated metal were analyzed on a ProteinChip Reader Model PBS II. Complex protein profiles of different diagnostic groups were compared and analyzed using the ProPeak software package. RESULTS: A panel of three biomarkers was selected based on their collective contribution to the optimal separation between stage 0-I breast cancer patients and noncancer controls. The same separation was observed using independent test data from stage II-III breast cancer patients. Bootstrap cross-validation demonstrated that a sensitivity of 93% for all cancer patients and a specificity of 91% for all controls were achieved by a composite index derived by multivariate logistic regression using the three selected biomarkers. CONCLUSIONS: Proteomics approaches such as SELDI mass spectrometry, in conjunction with bioinformatics tools, could greatly facilitate the discovery of new and better biomarkers. The high sensitivity and specificity achieved by the combined use of the selected biomarkers show great potential for the early detection of breast cancer.

Adult↗

Segmentation and intensity estimation of microarray images using a gamma-t mixture model.

MOTIVATION: We present a new approach to the analysis of images for complementary DNA microarray experiments. The image segmentation and intensity estimation are performed simultaneously by adopting a two-component mixture model. One component of this mixture corresponds to the distribution of the background intensity, while the other corresponds to the distribution of the foreground intensity. The intensity measurement is a bivariate vector consisting of red and green intensities. The background intensity component is modeled by the bivariate gamma distribution, whose marginal densities for the red and green intensities are independent three-parameter gamma distributions with different parameters. The foreground intensity component is taken to be the bivariate t distribution, with the constraint that the mean of the foreground is greater than that of the background for each of the two colors. The degrees of freedom of this t distribution are inferred from the data but they could be specified in advance to reduce the computation time. Also, the covariance matrix is not restricted to being diagonal and so it allows for nonzero correlation between R and G foreground intensities. This gamma-t mixture model is fitted by maximum likelihood via the EM algorithm. A final step is executed whereby nonparametric (kernel) smoothing is undertaken of the posterior probabilities of component membership. The main advantages of this approach are: (1) it enjoys the well-known strengths of a mixture model, namely flexibility and adaptability to the data; (2) it considers the segmentation and intensity simultaneously and not separately as in commonly used existing software, and it also works with the red and green intensities in a bivariate framework as opposed to their separate estimation via univariate methods; (3) the use of the three-parameter gamma distribution for the background red and green intensities provides a much better fit than the normal (log normal) or t distributions; (4) the use of the bivariate t distribution for the foreground intensity provides a model that is less sensitive to extreme observations; (5) as a consequence of the aforementioned properties, it allows segmentation to be undertaken for a wide range of spot shapes, including doughnut, sickle shape and artifacts. RESULTS: We apply our method for gridding, segmentation and estimation to cDNA microarray real images and artificial data. Our method provides better segmentation results in spot shapes as well as intensity estimation than Spot and spotSegmentation R language softwares. It detected blank spots as well as bright artifact for the real data, and estimated spot intensities with high-accuracy for the synthetic data. AVAILABILITY: The algorithms were implemented in Matlab. The Matlab codes implementing both the gridding and segmentation/estimation are available upon request. SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online.

Algorithms↗

RELIC--a bioinformatics server for combinatorial peptide analysis and identification of protein-ligand interaction sites.

Phage display technology provides a versatile tool for exploring the interactions between proteins, peptides and small molecule ligands. Quantitative analysis of peptide population sequence diversity and bias patterns has the power to significantly enhance the impact of these methods [1, 2]. We have developed a suite of computational tools for the analysis of peptide populations and made them accessible by integrating fifteen software programs for the analysis of combinatorial peptide sequences into the REceptor LIgand Contacts (RELIC) relational database and web-server. These programs have been developed for the analysis of statistical properties of peptide populations; identification of weak consensus sequences within these populations; and the comparison of these peptide sequences to those of naturally occurring proteins. RELIC is particularly suited to the analysis of peptide populations affinity selected with a small molecule ligand such as a drug or metabolite. Within this functional context, the ability to identify potential small molecule binding proteins using combinatorial peptide screening will accelerate as more ligands are screened and more genome sequences become available. The broader impact of this work is the addition of a novel means of analyzing peptide populations to the phage display community.

Algorithms↗

Efficient exact p-value computation for small sample, sparse, and surprising categorical data.

A major obstacle in applying various hypothesis testing procedures to datasets in bioinformatics is the computation of ensuing p-values. In this paper, we define a generic branch-and-bound approach to efficient exact p-value computation and enumerate the required conditions for successful application. Explicit procedures are developed for the entire Cressie-Read family of statistics, which includes the widely used Pearson and likelihood ratio statistics in a one-way frequency table goodness-of-fit test. This new formulation constitutes a first practical exact improvement over the exhaustive enumeration performed by existing statistical software. The general techniques we develop to exploit the convexity of many statistics are also shown to carry over to contingency table tests, suggesting that they are readily extendible to other tests and test statistics of interest. Our empirical results demonstrate a speed-up of orders of magnitude over the exhaustive computation, significantly extending the practical range for performing exact tests. We also show that the relative speed-up gain increases as the null hypothesis becomes sparser, that computation precision increases with increase in speed-up, and that computation time is very moderately affected by the magnitude of the computed p-value. These qualities make our algorithm especially appealing in the regimes of small samples, sparse null distributions, and rare events, compared to the alternative asymptotic approximations and Monte Carlo samplers. We discuss several established bioinformatics applications, where small sample size, small expected counts in one or more categories (sparseness), and very small p-values do occur. Our computational framework could be applied in these, and similar cases, to improve performance.

Computational Biology↗

A qualitative study of the implementation of a bioinformatics tool in a biological research laboratory.

OBJECTIVE: To explore how the implementation of a comprehensive new bioinformatics analysis system would affect workflow, collaboration and information management in a small genetic research lab. DESIGN: This was a longitudinal qualitative study of seven individuals involved in genomic and proteomic research. The study data were gathered using the illuminative/responsive approach of immersion in the environment. Additional qualitative data were gathered using informal semi-structured interviews, participant observation in lab meetings, and direct observation of lab researchers engaged in specific tasks. MEASUREMENTS: Interview, observation and field note data were coded and analyzed based on three analysis perspectives. A subset of the data was independently evaluated by an external researcher to enhance the trustworthiness of results. RESULTS: Three reoccurring themes were observed in the study. (1) Satisfaction and acceptance of software tools tended to be role and goal specific. (2) The system was seen primarily as a measurement system rather than a "total laboratory analysis system". (3) Lab meetings deemphasized the system, preferring more traditional data analysis techniques. These themes support the observations that the system was not used to its full potential in the lab. CONCLUSION: Themes identified in this study suggest that sophisticated genetic researchers face similar problems of technology implementation as do professionals in other fields. We recommend that leadership support and on-going training and evolution of academic curricula can improve chances of bioinformatics analysis systems becoming used more effectively.

Computational Biology↗

Depositing electron microscopy maps.

A meeting was held at the European Bioinformatics Institute (EBI) in Hinxton, United Kingdom to discuss recent progress in the development of EMD, a database for maps determined by electron microscopy that is now integrated with MSD, the macromolecular structure database at EBI. This meeting of representatives of many of the major image processing groups in electron microscopy also discussed possible software developments that would ease the documentation and deposition of such datasets. The meeting concluded with a strong endorsement of map deposition in electron microscopy and its linkage with the family of archival databases in biomedical research.

Computational Biology↗

Multiple Alignment of protein structures and sequences for VMD.

Multiple Alignment is a new interface for performing and analyzing multiple protein structure alignments. It enables viewing levels of sequence and structure similarity on the aligned structures and performing a variety of evolutionary and bioinformatic tasks, including the construction of structure-based phylogenetic trees and minimal basis sets of structures that best represent the topology of the phylogenetic tree. It is implemented as a plugin for VMD (Visual Molecular Dynamics), which is distributed by the NIH Resource for Macromolecular Modeling and Bioinformatics at the University of Illinois.

Amino Acid Sequence↗

Indelign: a probabilistic framework for annotation of insertions and deletions in a multiple alignment.

MOTIVATION: A quantitative study of molecular evolutionary events such as substitutions, insertions and deletions from closely related genomes requires (1) an accurate multiple sequence alignment program and (2) a method to annotate the insertions and deletions that explain the 'gaps' in the alignment. Although the former requirement has been extensively addressed, the latter problem has received little attention, especially in a comprehensive probabilistic framework. RESULTS: Here, we present Indelign, a program that uses a probabilistic evolutionary model to compute the most likely scenario of insertions and deletions consistent with an input multiple alignment. It is also capable of modifying the given alignment so as to obtain a better agreement with the evolutionary model. We find close to optimal performance and substantial improvement over alternative methods, in tests of Indelign on synthetic data. We use Indelign to analyze regulatory sequences in Drosophila, and find an excess of insertions over deletions, which is different from what has been reported for neutral sequences. AVAILABILITY: The Indelign program may be downloaded from the website http://veda.cs.uiuc.edu/indelign/ SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online.

Algorithms↗

Identification of 17 Pseudomonas aeruginosa sRNAs and prediction of sRNA-encoding genes in 10 diverse pathogens using the bioinformatic tool sRNAPredict2.

sRNAs are small, non-coding RNA species that control numerous cellular processes. Although it is widely accepted that sRNAs are encoded by most if not all bacteria, genome-wide annotations for sRNA-encoding genes have been conducted in only a few of the nearly 300 bacterial species sequenced to date. To facilitate the efficient annotation of bacterial genomes for sRNA-encoding genes, we developed a program, sRNAPredict2, that identifies putative sRNAs by searching for co-localization of genetic features commonly associated with sRNA-encoding genes. Using sRNAPredict2, we conducted genome-wide annotations for putative sRNA-encoding genes in the intergenic regions of 11 diverse pathogens. In total, 2759 previously unannotated candidate sRNA loci were predicted. There was considerable range in the number of sRNAs predicted in the different pathogens analyzed, raising the possibility that there are species-specific differences in the reliance on sRNA-mediated regulation. Of 34 previously unannotated sRNAs predicted in the opportunistic pathogen Pseudomonas aeruginosa, 31 were experimentally tested and 17 were found to encode sRNA transcripts. Our findings suggest that numerous genes have been missed in the current annotations of bacterial genomes and that, by using improved bioinformatic approaches and tools, much remains to be discovered in 'intergenic' sequences.

Computational Biology↗