Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

SS-Wrapper: a package of wrapper applications for similarity searches on Linux clusters.

BACKGROUND: Large-scale sequence comparison is a powerful tool for biological inference in modern molecular biology. Comparing new sequences to those in annotated databases is a useful source of functional and structural information about these sequences. Using software such as the basic local alignment search tool (BLAST) or HMMPFAM to identify statistically significant matches between newly sequenced segments of genetic material and those in databases is an important task for most molecular biologists. Searching algorithms are intrinsically slow and data-intensive, especially in light of the rapid growth of biological sequence databases due to the emergence of high throughput DNA sequencing techniques. Thus, traditional bioinformatics tools are impractical on PCs and even on dedicated UNIX servers. To take advantage of larger databases and more reliable methods, high performance computation becomes necessary. RESULTS: We describe the implementation of SS-Wrapper (Similarity Search Wrapper), a package of wrapper applications that can parallelize similarity search applications on a Linux cluster. Our wrapper utilizes a query segmentation-search (QS-search) approach to parallelize sequence database search applications. It takes into consideration load balancing between each node on the cluster to maximize resource usage. QS-search is designed to wrap many different search tools, such as BLAST and HMMPFAM using the same interface. This implementation does not alter the original program, so newly obtained programs and program updates should be accommodated easily. Benchmark experiments using QS-search to optimize BLAST and HMMPFAM showed that QS-search accelerated the performance of these programs almost linearly in proportion to the number of CPUs used. We have also implemented a wrapper that utilizes a database segmentation approach (DS-BLAST) that provides a complementary solution for BLAST searches when the database is too large to fit into the memory of a single node. CONCLUSIONS: Used together, QS-search and DS-BLAST provide a flexible solution to adapt sequential similarity searching applications in high performance computing environments. Their ease of use and their ability to wrap a variety of database search programs provide an analytical architecture to assist both the seasoned bioinformaticist and the wet-bench biologist.

Algorithms↗

The use of microarrays to study the anaerobic response in Arabidopsis.

BACKGROUND AND AIMS: The use of microarrays to characterize the transcript profile of Arabidopsis under various experimental conditions is rapidly expanding. This technique provides a huge amount of expression data, requiring bioinformatics tools to allow the proposal of working hypotheses. The aim of this study was to test the usefulness of this approach to examine the anaerobic response of Arabidopsis by evaluating the reliability of microarray data sets and by interrogation of microarray databases for the expression data of a set of anoxia-inducible genes. METHODS: User-driven software tools that display large gene expression datasets onto diagrams of metabolic pathways were used. The Genevestigator software was used to explore the expression of anoxia-inducible genes throughout the life cycle of Arabidopsis as well as relative to plant organs. T-DNA tagged mutants for selected genes identified from our microarray analysis were searched in the Arabidopsis thaliana Insertion Database, looking for insertional mutants from the Salk collection. KEY RESULTS: The results indicate that microarray data can provide the basis for new hypotheses in the field of plant responses to anaerobiosis and also provide knowledge for a targeted screening of Arabidopsis mutants. CONCLUSIONS: Research on plant responses to anaerobiosis can enormously benefit from the microarray technology.

Anaerobiosis↗

The European Bioinformatics Institute (EBI) databases.

This paper describes the databases and services of the European Bioinformatics Institute (EBI). In collaboration with DDBJ and GenBank/NCBI, the EBI maintains and distributes the EMBL Nucleotide Sequence Database, Europe's primary nucleotide sequence data resource. The EBI also maintains and distributes the SWISS-PROT Protein Sequence Database, in collaboration with Amos Bairoch of the University of Geneva. Over thirty additional specialist molecular biology databases, as well as software and documentation of interest to molecular biologists, are also available. The EBI network services include database searching, entry retrieval, and sequence similarity searching facilities.

Amino Acid Sequence↗

Introduction to bioinformatics.

This article introduces the field of bioinformatics and describes bioinformatic approaches and their application to the study of protein allergens. The predominant bioinformatics tools and resources are listed and discussed.

Allergens↗

Computational methods for gene annotation: the Arabidopsis genome.

Since the structure of the DNA molecule was identified half a century ago, the complete genome sequence has been determined for 37 prokaryotes and several eukaryotes. With the exponential growth of genetic information, bioinformatics has attempted to predict gene locations and functions in cyberspace prior to experimental confirmation at the bench.

Arabidopsis↗

Storing biological sequence databases in relational form.

SUMMARY: We have created a set of applications using Perl and Java in combination with XML technology to install biological sequence databases into an Oracle RDBMS. An easy-to-use interface using Java has been created for database query and other tools developed to integrate with our in-house bioinformatics applications. AVAILIBILITY: The database schema, DTD file, and source codes are available from the authors via email. CONTACT: guochun_ xie@merck. com

Amino Acid Sequence↗

Proteomic approaches to studying drug targets and resistance in Plasmodium.

Ever increasing drug resistance by Plasmodium falciparum, the most virulent of human malaria parasites, is creating new challenges in malaria chemotherapy. The entire genome sequences of P. falciparum and the rodent malaria parasite, P. yoelii yoelii are now available. Extensive genome sequence data from other Plasmodium species including another important human malaria parasite, P. vivax are also available. Powerful research techniques coupled to genomic resources are needed to help identify new drug and vaccine targets against malaria. Applied to Plasmodium, proteomics combines high-resolution protein or peptide separation with mass spectrometry and computer software to rapidly identify large numbers of proteins expressed from various stages of parasite development. Proteomic methods can be applied to study sub-cellular localization, cell function, organelle composition, changes in protein expression patterns in response to drug exposure, drug-protein binding and validation of data from genomic annotation and transcript expression studies. Recent high-throughput proteomic approaches have provided a wealth of protein expression data on P. falciparum, while smaller-scale studies examining specific drug-related hypotheses are also appearing. Of particular interest is the study of mechanisms of action and resistance of drugs such as the quinolines, whose targets currently may not be predictable from genomic data. Coupling the Plasmodium sequence data with bioinformatics, proteomics and RNA transcript expression profiling opens unprecedented opportunities for exploring new malaria control strategies. This review will focus on pharmacological research in malaria and other intracellular parasites using proteomic techniques, emphasizing resources and strategies available for Plasmodium.

Animals↗

Representation and processing of complex DNA spatial architecture and its annotated genomic content.

This paper presents a new general approach for the spatial representation and visualization of DNA molecule and its annotated information. This approach is based on a biological 3D model that predicts the complex spatial trajectory of huge naked DNA. With such modeling, a global vision of the sequence is possible, which is different and complementary to other representations as textual, linguistics or syntactic ones. The DNA is well known as a three-dimensional structure. Whereas, the spatial information plays a great part during its evolution and its interaction with the other biological elements This work will motivate investigations in order to launch new bioinformatics studies for the analysis of the spatial architecture of the genome. Besides, in order to obtain a friendly interactive visualization, a powerful graphic modeling is proposed including DNA complex trajectory management and its annotated-based content structuring. The paper describes spatial architecture modeling, with consideration of both biological and computational constraints. This work is implemented through a powerful graphic software tool, named ADN-Viewer. Several examples of visualization are shown for various organisms and biological elements.

DNA↗

[Changes in apoptosis-related genes expression profile in human breast carcinoma cell line Bcap-37 induced by flavonoids from seed residues of Hippophae Rhamnoides L].

BACKGROUND & OBJECTIVE: Hippophae rhamnoides L. possesses functions of antioxidation and radioprotection. This study was designed to investigate changes in apoptosis-related genes expression profile in human breast carcinoma cell line Bcap-37 induced by flavonoids from seed residues of Hippophae rhamnoides L. (FHR) with cDNA microarray, and to explore possible mechanism of signal transduction on apoptosis. METHODS: Total RNA was extracted from Bcap-37 cells before and after treatment of FHR. Two cDNA probes, labeled by Cy3-dUTP or Cy5-dUTP fluorescent dyes, were synthesized via reverse transcription, and hybridized with a microarray contained 13 824 human 14K cDNA. Differential gene expression profiles of FHR group and control group were analyzed by Genespring software. RESULTS: After treatment of FHR, 305 genes were up-regulated, and 361 were down-regulated; 32 apoptosis-related genes were differentially expressed, and accounted for 0.23% of the total genes in cDNA microarray. Of the 32 apoptosis-related genes, 25 were up-regulated (average Ratio: 3.071), and 7 were down-regulated (average Ratio: 0.418). Bioinformatic analyses showed that the 32 genes, including CTNNB1, TSSC3, IGFBP4, IGFBP6, GADD34, TNFRSF10B, Caspase-9, and PCNA, related with apoptosis of Bcap-37 cells when treated with FHR. CONCLUSION: Apoptosis of Bcap-37 cells induced by FHR relates with various genes through co-regulating of intracellular and extracellular signal transduction pathways.

Apoptosis↗

Arby: automatic protein structure prediction using profile-profile alignment and confidence measures.

MOTIVATION: Arby is a new server for protein structure prediction that combines several homology-based methods for predicting the three-dimensional structure of a protein, given its sequence. The methods used include a threading approach, which makes use of structural information, and a profile-profile alignment approach that incorporates secondary structure predictions. The combination of the different methods with the help of empirically derived confidence measures affords reliable template selection. RESULTS: According to the recent CAFASP3 experiment, the server is one of the most sensitive methods for predicting the structure of single domain proteins. The quality of template selection is assessed using a fold-recognition experiment. AVAILABILITY: The Arby server is available through the portal of the Helmholtz Network for Bioinformatics at http://www.hnbioinfo.de under the protein structure category.

Algorithms↗

FlyRNAi: the Drosophila RNAi screening center database.

RNA interference (RNAi) has become a powerful tool for genetic screening in Drosophila. At the Drosophila RNAi Screening Center (DRSC), we are using a library of over 21,000 double-stranded RNAs targeting known and predicted genes in Drosophila. This library is available for the use of visiting scientists wishing to perform full-genome RNAi screens. The data generated from these screens are collected in the DRSC database (http://flyRNAi.org/cgi-bin/RNAi_screens.pl) in a flexible format for the convenience of the scientist and for archiving data. The long-term goal of this database is to provide annotations for as many of the uncharacterized genes in Drosophila as possible. Data from published screens are available to the public through a highly configurable interface that allows detailed examination of the data and provides access to a number of other databases and bioinformatics tools.

Animals↗

A compilation of molecular biology web servers: 2006 update on the Bioinformatics Links Directory.

The Bioinformatics Links Directory is a public online resource that lists the servers published in this and all previously published Nucleic Acids Research Web Server issues together with other useful tools, databases and resources for bioinformatics and molecular biology research. This rich directory of tools and websites can be browsed and searched with all listed links freely accessible to the public. The 2006 update includes the 149 websites highlighted in the July 2006 issue of Nucleic Acids Research and brings the total number of servers listed in the Bioinformatics Links Directory to over 1000 links. To aid navigation through this growing resource, all link entries contain a brief synopsis, a citation list and are classified by function in descriptive biological categories. The most up-to-date version of this actively maintained listing of bioinformatics resources is available at the Bioinformatics Links Directory website, http://bioinformatics.ubc.ca/resources/links_directory/. A complete list of all links listed in this Nucleic Acids Research 2006 Web Server issue can be accessed online at http://bioinformatics.ubc.ca/resources/links_directory/narweb2006/. The 2006 update of the Bioinformatics Links Directory, which includes the Web Server list and summaries, is also available online at the Nucleic Acids Research website, http://nar.oupjournals.org/.

Computational Biology↗

Identification of novel highly expressed genes in pancreatic ductal adenocarcinomas through a bioinformatics analysis of expressed sequence tags.

In most microarray experiments, a significant fraction of the differentially expressed mRNAs identified correspond to expressed sequence tags (ESTs) and are generally discarded from further analyses. We used careful bioinformatics analyses to characterize those ESTs that were found to be highly overexpressed in a series of pancreatic adenocarcinomas. cDNA was prepared from 60 non-neoplastic samples (normal pancreas [n = 20], normal colon [n = 10], or normal duodenal mucosal [n = 30]) and from 64 pancreatic cancers (resected cancers [n = 50] or cancer cell lines [n = 14]) and hybridized to the complete Affymetrix Human Genome U133 GeneChip(R) set (arrays U133A and B) for simultaneous analysis of 45,000 fragments corresponding to 33,000 known genes and 6,000 ESTs. The GeneExpress(R) software system Fold Change Analysis Tool was used and 60 ESTs were identified that were expressed at levels at least 3-fold greater in the pancreatic cancers as compared to normal tissues. Searches against the human genomic sequence and comparative genomic analysis of human and mouse genomes was carried out using basic local alignment search tools (BLAST), BLASTN, and BLASTX, for identifying protein coding genes corresponding to the ESTs. Subsequently, in order to pick the most relevant candidate genes for a more detailed analysis, we looked for domains/motifs in the open reading frames using SMART and Pfam programs. We were able to definitively map 43 of the 60 ESTs to known or novel genes, and 15 of the ESTs could be localized in close proximity to a gene in the human genome although we were unable to establish that the EST was indeed derived from those genes. The differential expression of a subset of genes was confirmed at the protein level by immunohistochemical labeling of tissue microarrays (inhibin beta A [INHBA] and CD29) and/or at the transcript level by RT-PCR (INHBA, AKAP12, ELK3, FOXQ1, EIF5A2, and EFNA5). We conclude that bioinformatics tools can be used to characterize differentially overexpressed ESTs, and that some of these ESTs may represent diagnostically and therapeutically useful targets that might be missed using data solely from currently annotated databases.

Adenocarcinoma↗

Poxvirus orthologous clusters: toward defining the minimum essential poxvirus genome.

Increasingly complex bioinformatic analysis is necessitated by the plethora of sequence information currently available. A total of 21 poxvirus genomes have now been completely sequenced and annotated, and many more genomes will be available in the next few years. First, we describe the creation of a database of continuously corrected and updated genome sequences and an easy-to-use and extremely powerful suite of software tools for the analysis of genomes, genes, and proteins. These tools are available free to all researchers and, in most cases, alleviate the need for using multiple Internet sites for analysis. Further, we describe the use of these programs to identify conserved families of genes (poxvirus orthologous clusters) and have named the software suite POCs, which is available at www.poxvirus.org. Using POCs, we have identified a set of 49 absolutely conserved gene families-those which are conserved between the highly diverged families of insect-infecting entomopoxviruses and vertebrate-infecting chordopoxviruses. An additional set of 41 gene families conserved in chordopoxviruses was also identified. Thus, 90 genes are completely conserved in chordopoxviruses and comprise the minimum essential genome, and these will make excellent drug, antibody, vaccine, and detection targets. Finally, we describe the use of these tools to identify necessary annotation and sequencing updates in poxvirus genomes. For example, using POCs, we identified 19 genes that were widely conserved in poxviruses but missing from the vaccinia virus strain Tian Tan 1998 GenBank file. We have reannotated and resequenced fragments of this genome and verified that these genes are conserved in Tian Tan. The results for poxvirus genes and genomes are discussed in light of evolutionary processes.

Amino Acid Sequence↗

Large-scale proteomic analysis of membrane proteins.

Proteomic analysis of membrane proteins is a promising approach for the identification of novel drug targets and/or disease biomarkers. Despite notable technological developments, obstacles related to extraction and solublization of membrane proteins are encountered. A critical discussion of the different preparative methods of membrane proteins is offered in relation to downstream proteomic applications, mainly gel-based analyses and mass spectrometry. Frequently, unknown proteins are identified by high-throughput profiling of membrane proteins. In search for novel membrane proteins, analysis of protein sequences using computational tools is performed to predict the presence of transmembrane domains. This review also presents these bioinformatic tools with the human proteome as a case study. Along with technological innovations, advancements in the areas of sample preparation and computational prediction of membrane proteins will lead to exciting discoveries.

Mass Spectrometry↗

PIMWalker: visualising protein interaction networks using the HUPO PSI molecular interaction format.

UNLABELLED: This article reports on PIMWalker, a free and interactive tool for visualising protein interaction networks. PIMWalker handles the unified molecular interaction (MI) format defined by members of the Proteomics Standards Initiative (the PSI MI format), and it is thus directly and easily usable by bench biologists. PIMWalker also comes with a documented, open-source Javatrade mark application programming interface allowing the bioinformatic programmer to easily extend the functions. AVAILABILITY: PIMWalker is available under a free license from http://pim.hybrigenics.com/pimwalker.

Computer Graphics↗

Drawing phylogenetic trees in LATEX and Microsoft Word.

UNLABELLED: newicktree is a PSTricks-based LATEX package which enables phylogenetic trees described in the Newick format to be drawn directly into LATEX documents. mswordtree is a macro for producing phylogenetic trees using the drawing elements available in Microsoft Word. AVAILABILITY: Both programs are available free from the John Innes Centre's Bioinformatics Research Group website at http://jic-bioinfo.bbsrc.ac.uk/bioinformatics-research/software/index.html. SUPPLEMENTARY INFORMATION: A full user-guide for newicktree and installation and usage instructions for mswordtree and available at http://jic-bioinfo.bbsrc.ac.uk/bioinformatics-research/software/index.html

Computer Graphics↗

Consensus shapes: an alternative to the Sankoff algorithm for RNA consensus structure prediction.

MOTIVATION: The well-known Sankoff algorithm for simultaneous RNA sequence alignment and folding is currently considered an ideal, but computationally over-expensive method. Available tools implement this algorithm under various pragmatic restrictions. They are still expensive to use, and it is difficult to judge if the moderate quality of results is because of the underlying model or to its imperfect implementation. RESULTS: We propose to redefine the consensus structure prediction problem in a way that does not imply a multiple sequence alignment step. For a family of RNA sequences, our method explicitly and independently enumerates the near-optimal abstract shape space, and predicts as the consensus an abstract shape common to all sequences. For each sequence, it delivers the thermodynamically best structure which has this common shape. Since the shape space is much smaller than the structure space, and identification of common shapes can be done in linear time (in the number of shapes considered), the method is essentially linear in the number of sequences. Our evaluation shows that the new method compares favorably with available alternatives. AVAILABILITY: The new method has been implemented in the program RNAcast and is available on the Bielefeld Bioinformatics Server. CONTACT: jreeder@TechFak.Uni-Bielefeld.DE, robert@TechFak.Uni-Bielefeld.DE SUPPLEMENTARY INFORMATION: Available at http://bibiserv.techfak.uni-bielefeld.de/rnacast/supplementary.html

Algorithms↗