Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

SSEP: Secondary structural elements of proteins.

SSEP is a comprehensive resource for accessing information related to the secondary structural elements present in the 25 and 90% non-redundant protein chains. The database contains 1771 protein chains from 1670 protein structures and 6182 protein chains from 5425 protein structures in 25 and 90% non-redundant protein chains, respectively. The current version provides information about the alpha-helical segments and beta-strand fragments of varying lengths. In addition, it also contains the information about 3(10)-helix, beta- and nu-turns and hairpin loops. The free graphics program RASMOL has been interfaced with the search engine to visualize the three-dimensional structures of the user queried secondary structural fragment. The database is updated regularly and is available through Bioinformatics web server at http://cluster.physics.iisc.ernet.in/ssep/ or http://144.16.71.148/ssep/.

Databases, Protein↗

Structural principles governing domain motions in proteins.

With the use of a recently developed method, twenty-four proteins for which two or more X-ray conformers are known have been analyzed to reveal structural principles that govern domain motions in proteins. In all 24 cases, the domain motion is a rotation about a physical axis created through local interactions both covalent and noncovalent. In many cases, two or more mechanical hinges separated in space create a stable hinge axis for precise control of the domain closure. The terminal regions of alpha-helices and beta-sheets have been found to act as mechanical hinges in a significant number of cases. In some cases, the two terminal regions of neighboring strands of a single beta-sheet can create a hinge axis, as can the two termini of a single alpha-helix. These two structures have been termed the "double-hinged beta-sheet" and "double-hinged alpha-helix," respectively. A flexible loop that attaches one domain to another and through which the effective hinge axis passes is another construct that is used to create a hinge. Noncovalent interactions between segments remote along the polypeptide chain can also form hinges. In addition alpha-helices that preserve their hydrogen bonding structure when bent have been found to behave as mechanical hinges. It is suggested that these alpha-helices act as a store of elastic energy that drives the closing of domains for rapid capture of the substrate. If the repertoire of possible interdomain structures is as limited as this study suggests, the dynamic behavior of proteins could soon be predicted using bioinformatics techniques. Proteins 1999;36:425-435.

Animals↗

Beyond Mfold: recent advances in RNA bioinformatics.

Computational analysis of RNA secondary structure is a classical field of biosequence analysis, which has recently gained momentum due to the manyfold regulatory functions of RNA that have become apparent. We present five recent computational approaches that address the problems of synoptic folding space analysis, pseudoknot prediction, structure alignment, comparative structure prediction, and miRNA target prediction. All these programs are in current use and are available via the Bielefeld Bioinformatics Server at .

Animals↗

FatiGO: a web tool for finding significant associations of Gene Ontology terms with groups of genes.

We present a simple but powerful procedure to extract Gene Ontology (GO) terms that are significantly over- or under-represented in sets of genes within the context of a genome-scale experiment (DNA microarray, proteomics, etc.). Said procedure has been implemented as a web application, FatiGO, allowing for easy and interactive querying. FatiGO, which takes the multiple-testing nature of statistical contrast into account, currently includes GO associations for diverse organisms (human, mouse, fly, worm and yeast) and the TrEMBL/Swissprot GOAnnotations@EBI correspondences from the European Bioinformatics Institute.

Algorithms↗

Structural bioinformatics of DNA: a web-based tool for the analysis of molecular dynamics results and structure prediction.

UNLABELLED: We report here the release of a web-based tool (MDDNA) to study and model the fine structural details of DNA on the basis of data extracted from a set of molecular dynamics (MD) trajectories of DNA sequences involving all the unique tetranucleotides. The dynamic web interface can be employed to analyze the first neighbor sequence context effects on the 10 unique dinucleotide steps of DNA. Functionality is included to build all atom models of any user-defined sequence based on the MD results. The backend of this interface is a relational database storing the conformational details of DNA obtained in 39 different MD simulation trajectories comprising all the 136 unique tetranucleotide steps. Examples of the use of this data to predict DNA structures are included. AVAILABILITY: http://humphry.chem.wesleyan.edu:8080/MDDNA. SUPPLEMENTARY INFORMATION: Supplementary data including color figures are available at Bioinformatics online.

Base Sequence↗

HTself: self-self based statistical test for low replication microarray studies.

Different statistical methods have been used to classify a gene as differentially expressed in microarray experiments. They usually require a number of experimental observations to be adequately applied. However, many microarray experiments are constrained to low replication designs for different reasons, from financial restrictions to scarcely available RNA samples. Although performed in a high-throughput framework, there are few experimental replicas for each gene to allow the use of traditional or state-of-art statistical methods. In this work, we present a web-based bioinformatics tool that deals with real-life problems concerning low replication experiments. It uses an empirically derived criterion to classify a gene as differentially expressed by combining two widely accepted ideas in microarray analysis: self-self experiments to derive intensity-dependent cutoffs and non-parametric estimation techniques. To help laboratories without a bioinformatics infrastructure, we implemented the tool in a user-friendly website (http://blasto.iq.usp.br/~rvencio/HTself).

Algorithms↗

Protein structure prediction servers at University College London.

A number of state-of-the-art protein structure prediction servers have been developed by researchers working in the Bioinformatics Unit at University College London. The popular PSIPRED server allows users to perform secondary structure prediction, transmembrane topology prediction and protein fold recognition. More recent servers include DISOPRED for the prediction of protein dynamic disorder and DomPred for domain boundary prediction. These servers are available from our software home page at http://bioinf.cs.ucl.ac.uk/software.html.

Computational Biology↗

3dSS: 3D structural superposition.

3dSS is a web-based interactive computing server, primarily designed to aid researchers, to superpose two or several 3D protein structures. In addition, the server can be effectively used to find the invariant and common water molecules present in the superposed homologous protein structures. The molecular visualization tool RASMOL is interfaced with the server to visualize the superposed 3D structures with the water molecules (invariant or common) in the client machine. Furthermore, an option is provided to save the superposed 3D atomic coordinates in the client machine. To perform the above, users need to enter Protein Data Bank (PDB)-id(s) or upload the atomic coordinates in PDB format. This server uses a locally maintained PDB anonymous FTP server that is being updated weekly. This program can be accessed through our Bioinformatics web server at the URL http://cluster.physics.iisc.ernet.in/3dss/ or http://10.188.1.15/3dss/.

Computer Graphics↗

EGassembler: online bioinformatics service for large-scale processing, clustering and assembling ESTs and genomic DNA fragments.

Expressed sequence tag (EST) sequencing has proven to be an economically feasible alternative for gene discovery in species lacking a draft genome sequence. Ongoing large-scale EST sequencing projects feel the need for bioinformatics tools to facilitate uniform EST handling. This brings about a renewed importance for a universal tool for processing and functional annotation of large sets of ESTs. EGassembler (http://egassembler.hgc.jp/) is a web server, which provides an automated as well as a user-customized analysis tool for cleaning, repeat masking, vector trimming, organelle masking, clustering and assembling of ESTs and genomic fragments. The web server is publicly available and provides the community a unique all-in-one online application web service for large-scale ESTs and genomic DNA clustering and assembling. Running on a Sun Fire 15K supercomputer, a significantly large volume of data can be processed in a short period of time. The results can be used to functionally annotate genes, to facilitate splice alignment analysis, to link the transcripts to genetic and physical maps, design microarray chips, to perform transcriptome analysis and to map to KEGG metabolic pathways. The service provides an excellent bioinformatics tool to research groups in wet-lab as well as an all-in-one-tool for sequence handling to bioinformatics researchers.

Computational Biology↗

OPAAS: a web server for optimal, permuted, and other alternative alignments of protein structures.

The large number of experimentally determined protein 3D structures is a rich resource for studying protein function and evolution, and protein structure comparison (PSC) is a key method for such studies. When comparing two protein structures, almost all currently available PSC servers report a single and sequential (i.e. topological) alignment, whereas the existence of good alternative alignments, including those involving permutations (i.e. non-sequential or non-topological alignments), is well known. We have recently developed a novel PSC method that can detect alternative alignments of statistical significance (alignment similarity P-value <10(-5)), including structural permutations at all levels of complexity. OPAAS, the server of this PSC method freely accessible at our website (http://opaas.ibms.sinica.edu.tw), provides an easy-to-read hierarchical layout of output to display detailed information on all of the significant alternative alignments detected. Because these alternative alignments can offer a more complete picture on the structural, evolutionary and functional relationship between two proteins, OPAAS can be used in structural bioinformatics research to gain additional insight that is not readily provided by existing PSC servers.

Internet↗

Diagnosis of Ovarian Cancer Using Decision Tree Classification of Mass Spectral Data.

Recent reports from our laboratory and others support the SELDI ProteinChip technology as a potential clinical diagnostic tool when combined with $n$ -dimensional analyses algorithms. The objective of this study was to determine if the commercially available classification algorithm biomarker patterns software (BPS), which is based on a classification and regression tree (CART), would be effective in discriminating ovarian cancer from benign diseases and healthy controls. Serum protein mass spectrum profiles from 139 patients with either ovarian cancer, benign pelvic diseases, or healthy women were analyzed using the BPS software. A decision tree, using five protein peaks resulted in an accuracy of 81.5% in the cross-validation analysis and 80%in a blinded set of samples in differentiating the ovarian cancer from the control groups. The potential, advantages, and drawbacks of the BPS system as a bioinformatic tool for the analysis of the SELDI high-dimensional proteomic data are discussed.

Journal Article↗

ToothPrint, a proteomic database for dental tissues.

Increasing demand exists to disseminate and integrate proteomic data as proteome analysis assumes a commanding role in the postgenome era. Databases on the World Wide Web are an effective means to share information obtained from two-dimensional gels and allied proteomic approaches. Here we report the establishment of ToothPrint, a proteomic database for dental tissues accessed at http://toothprint.otago.ac.nz. Using developing rat enamel as a prototype, ToothPrint provides a variety of functionally relevant data (ligand binding, subcellular localisation, developmental regulation) in addition to protein identification maps. Features designed to enhance usability of the website and simplify its computing requirements are also outlined. Customized for mineralizing tissues, ToothPrint should prove to be an effective bioinformatic resource for investigations of dental biology.

Animals↗

Exploiting large scale computing to construct high resolution linkage disequilibrium maps of the human genome.

UNLABELLED: Linkage disequilibrium (LD) maps increase power and precision in association mapping, define optimal marker spacing and identify recombination hot-spots and regions influenced by natural selection. Phase II of HapMap provides approximately 2.8-fold more single nucleotide polymorphisms (SNPs) than phase I for constructing higher resolution maps. LDMAP-cluster, is a parallel program for rapid map construction in a Linux environment used here to construct genome-wide LD maps with >8.2 million SNPs from the phase II data. AVAILABILITY: The LD maps, LDMAP-cluster and documentation are available from: http://www.som.soton.ac.uk/research/geneticsdiv/epidemiology/LDMAP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

Secured distributed service to manage biological data on EGEE grid.

Biological data are most times published and then become public ones. They, then, do not need to be isolated or encrypted. But, in some cases, these data stemed from patients or are analyzed with, for instance, pharmaceutical or agronomics goals. Also in simple ways , these data, before to become public, have to be kept confidential while researchers haven't been able to publish their work or to register them. So they are a lot of cases where the integrity and the confidentiality of biological data have to be protected against unauthorized accesses. But, as these private data are also large datasets, they need high-throughput computing and huge data storage to processed, such as ones produced by complete genome projects. These requirements are enhanced in the context of a Grid such EGEE, where the computing and storage resources are distributed across a large-scale platform. We have developed a secured distributed service to manage biological data on grid: the EncFile encrypted files management system. We have deployed it on the production platform of the EGEE grid project. Thus we provided grid users with a user-friendly component that doesn't require any user privileges. And we have integrated into a bioinformatics grid portal associated to encrypted representative biological resources: world-famous databases and programs.

Computational Biology↗

Domain analysis of fatty acid synthase protein (NP_217040) from Mycobacterium tuberculosis H37Rv--a bioinformatics study.

Different domains of fatty acid synthase (FAS) protein of Mycobacterium tuberculosis H37Rv, involved in mycolic acid synthesis were analyzed using various bioinformatics tools. Based on different database searches (CDD and Pfam), FAS protein of Mycobacterium tuberculosis was grouped into eight domains, five of which showed close similarity with pdb templates (1MLA, 1IQ6A, 2BMOA, and 1J3NA). Based on the PSI blast analysis, 3D structures of only five domains were predicted using MODELLER software, and loop modeling was done for only those regions that were predicted as loops by predict protein server. Compared to the original structure, the loop modeled structure showed a lower DOPE score value for FAS protein. The X-ray determined templates that were used for predicting the 3D structure suggest that, FAS protein has "Malonyl-coenzyme A-Hydratase-Nitrobenzene dioxygenase-3-oxoacyl-(acp) synthase" activity. Accuracy of the prediction of 3D structure of different domains of FAS protein was further validated by Ramachandran plot and PROCHECK (G-value).

Amino Acid Sequence↗

Ensembl 2004.

The Ensembl (http://www.ensembl.org/) database project provides a bioinformatics framework to organize biology around the sequences of large genomes. It is a comprehensive and integrated source of annotation of large genome sequences, available via interactive website, web services or flat files. As well as being one of the leading sources of genome annotation, Ensembl is an open source software engineering project to develop a portable system able to handle very large genomes and associated requirements. The facilities of the system range from sequence analysis to data storage and visualization and installations exist around the world both in companies and at academic sites. With a total of nine genome sequences available from Ensembl and more genomes to follow, recent developments have focused mainly on closer integration between genomes and external data.

Animals↗

VANTED: a system for advanced data analysis and visualization in the context of biological networks.

BACKGROUND: Recent advances with high-throughput methods in life-science research have increased the need for automatized data analysis and visual exploration techniques. Sophisticated bioinformatics tools are essential to deduct biologically meaningful interpretations from the large amount of experimental data, and help to understand biological processes. RESULTS: We present VANTED, a tool for the visualization and analysis of networks with related experimental data. Data from large-scale biochemical experiments is uploaded into the software via a Microsoft Excel-based form. Then it can be mapped on a network that is either drawn with the tool itself, downloaded from the KEGG Pathway database, or imported using standard network exchange formats. Transcript, enzyme, and metabolite data can be presented in the context of their underlying networks, e. g. metabolic pathways or classification hierarchies. Visualization and navigation methods support the visual exploration of the data-enriched networks. Statistical methods allow analysis and comparison of multiple data sets such as different developmental stages or genetically different lines. Correlation networks can be automatically generated from the data and substances can be clustered according to similar behavior over time. As examples, metabolite profiling and enzyme activity data sets have been visualized in different metabolic maps, correlation networks have been generated and similar time patterns detected. Some relationships between different metabolites were discovered which are in close accordance with the literature. CONCLUSION: VANTED greatly helps researchers in the analysis and interpretation of biochemical data, and thus is a useful tool for modern biological research. VANTED as a Java Web Start Application including a user guide and example data sets is available free of charge at http://vanted.ipk-gatersleben.de.

Algorithms↗

RBR: library-less repeat detection for ESTs.

MOTIVATION: Repeat sequences in ESTs are a source of problems, in particular for clustering. ESTs are therefore commonly masked against a library of known repeats. High quality repeat libraries are available for the widely studied organisms, but for most other organisms the lack of such libraries is likely to compromise the quality of EST analysis. RESULTS: We present a fast, flexible and library-less method for masking repeats in EST sequences, based on match statistics within the EST collection. The method is not linked to a particular clustering algorithm. Extensive testing on datasets using different clustering methods and a genomic mapping as reference shows that this method gives results that are better than or as good as those obtained using RepeatMasker with a repeat library. AVAILABILITY: The implementation of RBR is available under the terms of the GPL from http://www.ii.uib.no/~ketil/bioinformatics CONTACT: ketil.malde@bccs.uib.no SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗