The Bioinformatics Open Access option.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
In the few short years since its discovery, RNA interference (RNAi) has revolutionized the functional analysis of genomes: both technical and conceptual approaches to the investigation of gene function are being transformed as a result of this new technology. Genome-scale RNAi analyses have already been performed in the model organisms Caenorhabditis elegans (in vivo) and Drosophila melanogaster (in cell lines), ushering in a new era of RNAi-based approaches to probing the inner workings of the cell. The transformation of complex phenotypic data into mineable 'digitized' formats is fostering the emergence of a new area of bioinformatics related to the phenome.
Vast amounts of life sciences data are scattered around the world in the form of a variety of heterogeneous data sources. The need to be able to co-relate relevant information is fundamental to increase the overall knowledge and understanding of a specific subject. Bioinformaticians aspire to find ways to integrate biological data sources for this purpose and system integration is a very important research topic. The purpose of this paper is to provide an overview of important integration issues that should be considered when designing a bioinformatics integration system. The currently prevailing approach for integration is presented with examples of bioinformatics information systems together with their main characteristics. Here, we introduce agent technology and we argue why it provides an appropriate solution for designing bioinformatics integration systems.
MOTIVATION: High-throughput technologies now allow the acquisition of biological data, such as comprehensive biochemical time-courses at unprecedented rates. These temporal profiles carry topological and kinetic information regarding the biochemical network from which they were drawn. Retrieving this information will require systematic application of both experimental and computational methods. RESULTS: S-systems are non-linear mathematical approximative models based on the power-law formalism. They provide a general framework for the simulation of integrated biological systems exhibiting complex dynamics, such as genetic circuits, signal transduction and metabolic networks. We describe how the heuristic optimization technique simulated annealing (SA) can be effectively used for estimating the parameters of S-systems from time-course biochemical data. We demonstrate our methods using three artificial networks designed to simulate different network topologies and behavior. We then end with an application to a real biochemical network by creating a working model for the cadBA system in Escherichia coli. AVAILABILITY: The source code written in C++ is available at http://www.engg.upd.edu.ph/~naval/bioinformcode.html. All the necessary programs including the required compiler are described in a document archived with the source code. SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online.
Using the Distributed Annotation System (DAS) we have created a protein annotation resource available at our web page: http://www.cbs.dtu.dk, as a part of the BioSapiens Network of Excellence EU FP6 project. The DAS protocol allows us to gather layers of annotation data for a given sequence and thereby gain an overview of the sequence's features. A user-friendly graphical client has also been developed (http://www.cbs.dtu.dk/cgi-bin/das), which demonstrates the possibility of integrating DAS annotation data from multiple sources into a simple graphical view. The client displays protein feature annotations from the Center for Biological Sequence Analysis as well as from the BioSapiens reference UniProt server (http://www.ebi.ac.uk/das-srv/uniprot/das) at the European Bioinformatics Institute. Other DAS data sources for protein annotation will be added as they become available.
There are four different types of N-terminal amino acid sequences (F-I-0, F-I, F-II, F-III) in the multicomponents of earthworm fibrinolytic enzymes (EFE). In GenBank 21 nucleic acid sequences of EFE have been reported. Among them, most of the N-terminal amino acid sequences belong to the F-III type,few belong to the F-II type. Only one is similar to the F-I type, but none to F-I-0. In this research we hoped to obtain the gene encoding component F-I-0 of EFE by the bioinformatics tools. Based on the N-terminal amino acid sequence VVGGSDTTIGQYPHQL of the F-I-0 type from Lumbricus rubellus, a nucleic acid sequence was obtained by in silico cloning from dbEST of Lumbricidae using the software DNAMAN. A new gene of EFE from Eisenia foetida was successfully obtained by RT-PCR using specific primers designed according to this sequence. The new gene named EfP-0 was cloned in pMAL-c2x and expressed as the fusion protein MBP-EfP-0 in the supernatant of lysate. The fusion protein MBP-EfP-0 purified by affinity chromatography had hydrolytic activity on casein plate. Sequencing result shows, EfP-0 has 678bp and encodes a protein of 225 amino acids. The protein is a serine protease belonging to trypsin family. It has similar amino acid composition to F-I-0. BLAST in GenBank shows that the similarity is lower than 40% between EJP-0 gene and other EFE genes. By this we conclude that EfP-0 gene of EFE is a novel gene and it is the first time to be reported, its accession number for Genbank is DQ836917.
The functional characterization of available genomic sequences is the major task of the research in the post-genome era. This complex task requires an integrative approach of high-throughput systems with in vitro and in vivo models in order to have a reliable evaluation of the biological function. The oligonucleotide antisense technology is one of the most promising approaches for the investigation of gene function; the crucial point of antisense experiments is the identification of optimal target sites for hybridisation. In this paper we have applied a bioinformatic tool for the recognition of optimal antisense targets. In order to evaluate the effect of mutational events on target selection we have tested the program on a sample of human beta-hemoglobin variants. The proposed algorithm software will be integrated in a web based tool at the site: http://www.nettab.org/agewa.
OBJECTIVE: To design oligonucleotide probes for microarray detection of dengue virus. METHODS: By analyzing the cDNAs of dengue viruses of 4 different serotypes with BLAST program, a group of specific sequences for the candidate probes was acquired. Oligo6.0 software was applied to analyze the candidates to select the probes with high specificity, identical length and similar melting temperature (Tm). RESULT: Altogether 48 oligonucleotide probes were designed, and deposited on oligonucleotide chips as the microarray for dengue virus detection. CONCLUSION: BLAST program and Oligo6.0 software are simple and effective means for designing the oligonucleotide probes.
MOTIVATION: RNA secondary structure analysis often requires searching for potential helices in large sequence data. RESULTS: We present a utility program GUUGle that efficiently locates potential helical regions under RNA base pairing rules, which include Watson-Crick as well as G-U pairs. It accepts a positive and a negative set of sequences, and determines all exact matches under RNA rules between positive and negative sequences that exceed a specified length. The GUUGle algorithm can also be adapted to use a precomputed suffix array of the positive sequence set. We show how this program can be effectively used as a filter preceding a more computationally expensive task such as miRNA target prediction. AVAILABILITY: GUUGle is available via the Bielefeld Bioinformatics Server at http://bibiserv.techfak.uni-bielefeld.de/guugle
In this study a systematic attempt has been made to integrate various approaches in order to predict allergenic proteins with high accuracy. The dataset used for testing and training consists of 578 allergens and 700 non-allergens obtained from A. K. Bjorklund, D. Soeria-Atmadja, A. Zorzet, U. Hammerling and M. G. Gustafsson (2005) Bioinformatics, 21, 39-50. First, we developed methods based on support vector machine using amino acid and dipeptide composition and achieved an accuracy of 85.02 and 84.00%, respectively. Second, a motif-based method has been developed using MEME/MAST software that achieved sensitivity of 93.94 with 33.34% specificity. Third, a database of known IgE epitopes was searched and this predicted allergenic proteins with 17.47% sensitivity at specificity of 98.14%. Fourth, we predicted allergenic proteins by performing BLAST search against allergen representative peptides. Finally hybrid approaches have been developed, which combine two or more than two approaches. The performance of all these algorithms has been evaluated on an independent dataset of 323 allergens and on 101 725 non-allergens obtained from Swiss-Prot. A web server AlgPred has been developed for the predicting allergenic proteins and for mapping IgE epitopes on allergenic proteins (http://www.imtech.res.in/raghava/algpred/). AlgPred is available at www.imtech.res.in/raghava/algpred/.
Optimal spaced seeds were developed as a method to increase sensitivity of local alignment programs similar to BLASTN. Such seeds have been used before in the program PatternHunter, and have given improved sensitivity and running time relative to BLASTN in genome-genome comparison. We study the problem of computing optimal spaced seeds for detecting homologous coding regions in unannotated genomic sequences. By using well-chosen seeds, we are able to improve the sensitivity of coding sequence alignment over that of TBLASTX, while keeping runtime comparable to BLASTN. We identify good seeds by first giving effective hidden Markov models of conservation in alignments of homologous coding regions. We give an efficient algorithm to compute the optimal spaced seed when conservation patterns are generated by these models. Our results offer the hope of improved gene finding due to fewer missed exons in DNA/DNA comparison, and more effective homology search in general, and may have applications outside of bioinformatics.
Database interoperation is becoming a bottleneck for the research community in biology. In this paper, we first discuss the question of interoperability and give a brief overview of CORBA. Then, an example is explained in some detail: a simple but realistic data bank of STSs is implemented. The Object Request Broker is the media for communication between an object server (the data bank) and a client (possibly a genome center). Since CORBA enables easy development of networked applications, we meant this paper to provide an incentive for the bioinformatics community to develop distributed objects.
UNLABELLED: The aim of the MCQTL software package is to perform QTL mapping in multi-cross designs. It allows the analysis of the usual populations derived from inbred lines and can link the families by assuming that the QTL locations are the same in all of them. Moreover, a diallel modelling of the QTL genotypic effects is allowed in multiple related families. The implemented model is a linear regression model. A composite interval mapping and an iterative QTL mapping are implemented to deal with multiple QTL models. Marker cofactor selections by forward or backward stepwise methods are implemented as well as computation of threshold test value by permutation. AVAILABILITY: The program is available on request after signing a licence agreement; free of charge for academic and non-profit organizations at http://www.genoplante.org (Bioinformatics products).
Explore the source record for details and available documents.
OBJECTIVE: The biological significance of the chemokine ligand C-X-C motif chemokine ligand 13 (CXCL13) may play a significant role in the pathogenesis of uterine corpus endometrial carcinoma (UCEC). This study aims to identify and verify CXCL13 with predictive value for prognosis in UCEC. METHODS: CXCL13 mRNA expression differences were analyzed using R software in three independent datasets: one each from The Cancer Genome Atlas (TCGA) and two from the Gene Expression Omnibus (GEO), namely GSE17025 and GSE106191. The correlation between CXCL13 expression and prognosis was evaluated by Kaplan-Meier analysis. Univariate and multivariate Cox analyses were utilized to construct a prognostic nomogram. Tumor Immune Estimation Resource (TIMER) and the Tumor and Immune System Interaction Database (TISIDB) were employed to assess the relationship between CXCL13 and tumor immune infiltration. Coexpressed genes with CXCL13 were identified by the Spearman correlation analysis. A CXCL13 protein-protein interaction (PPI) network was constructed with the STRING website tool and hub genes were screened out. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genome (KEGG) analyses were performed with the "clusterProfiler" R package. Gene set enrichment analysis (GSEA) was used to identify underlying biological mechanisms. A drug-gene interaction network was constructed in the Comparative Toxicogenomics Database (CTD). RESULTS: High CXCL13 mRNA expression were validated in UCEC in the above three independent datasets. High CXCL13 expression was associated with favorable prognosis in UCEC. A nomogram for predicting the 1-, 3-, and 5-year survival probability in UCEC was construct based on CXCL13 expression and other clinical parameters. The use of Spearman correlation indicated certain correlation between CXCL13 and immune cells and immune checkpoint (ICP) genes. Seven hub genes were upregulated in UCEC, namely CXCL9, IFNG, CXCL10, CXCL11, GBP5, CCL18, and GZMB. The expression and prognostic relevance of CXCL9, IFNG, GBP5, and GZMB were in accordance with CXCL13. The main biological processes enriched were cytokine-cytokine receptor interaction and chemokine signaling pathway. CONCLUSIONS: The above comprehensive analyses suggest that CXCL13 may serve as a potential prognostic biomarker for UCEC, specifically for early-stage UCEC.
A large part of mammalian proteomes is represented by hypothetical proteins (HP), i.e. proteins predicted from nucleic acid sequences only and protein sequences with unknown function. Databases are far from being complete and errors are expected. The legion of HP is awaiting experiments to show their existence at the protein level and subsequent bioinformatic handling in order to assign proteins a tentative function is mandatory. Two-dimensional gel-electrophoresis with subsequent mass spectrometrical identification of protein spots is an appropriate tool to search for HP in the high-throughput mode. Spots are identified by MS or by MS/MS measurements (MALDI-TOF, MALDI-TOF-TOF) and subsequent software as e.g. Mascot or ProFound. In many cases proteins can thus be unambiguously identified and characterised; if this is not the case, de novo sequencing or Q-TOF analysis is warranted. If the protein is not identified, the sequence is being sent to databases for BLAST searches to determine identities/similarities or homologies to known proteins. If no significant identity to known structures is observed, the protein sequence is examined for the presence of functional domains (databases PROSITE, PRINTS, InterPro, ProDom, Pfam and SMART), subjected to searches for motifs (ELM) and finally protein-protein interaction databases (InterWeaver, STRING) are consulted or predictions from conformations are performed. We here provide information about hypothetical proteins in terms of protein chemical analysis, independent of antibody availability and specificity and bioinformatic handling to contribute to the extension/completion of protein databases and include original work on HP in the brain to illustrate the processes of HP identification and functional assignment.
Maize (Zea mays) protoporphyrinogen IX oxidase (PPO: EC 1.3.3.4) possesses a chloroplast transit peptide (CTP) that delivers the enzyme into the chloroplast. The cleavage site yielding the mature protein was predicted by using the ChloroP software and by comparing conserved regions of the available plant PPO sequences. In parallel, the processed NH(2)-terminus of native PPO was identified experimentally by microsequencing the immunoprecipitated plant PPO from maize etioplasts. The cleavage sites identified using the bioinformatic approaches did not match the experimental result. The three sequences have been cloned and expressed in bacteria and their kinetics were compared in order to understand if the generated proteins had biochemically relevant differences. Recombinant PPO corresponding to the native PPO accumulated at higher level and was more active than the two homologues. A cysteine present in the CTP seems to be able to modify the redox state of the enzyme and to be responsible for the alteration of the kinetic features. In contrast, the sensitivity to different herbicides was unaffected by modifications at the NH(2)-terminus, suggesting that the mode of action is non-competitive and that the NH(2)-terminus is involved in the recognition of the natural substrate.