Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

A scale of functional divergence for yeast duplicated genes revealed from analysis of the protein-protein interaction network.

BACKGROUND: Studying the evolution of the function of duplicated genes usually implies an estimation of the extent of functional conservation/divergence between duplicates from comparison of actual sequences. This only reveals the possible molecular function of genes without taking into account their cellular function(s). We took into consideration this latter dimension of gene function to approach the functional evolution of duplicated genes by analyzing the protein-protein interaction network in which their products are involved. For this, we derived a functional classification of the proteins using PRODISTIN, a bioinformatics method allowing comparison of protein function. Our work focused on the duplicated yeast genes, remnants of an ancient whole-genome duplication. RESULTS: Starting from 4,143 interactions, we analyzed 41 duplicated protein pairs with the PRODISTIN method. We showed that duplicated pairs behaved differently in the classification with respect to their interactors. The different observed behaviors allowed us to propose a functional scale of conservation/divergence for the duplicated genes, based on interaction data. By comparing our results to the functional information carried by GO annotations and sequence comparisons, we showed that the interaction network analysis reveals functional subtleties, which are not discernible by other means. Finally, we interpreted our results in terms of evolutionary scenarios. CONCLUSIONS: Our analysis might provide a new way to analyse the functional evolution of duplicated genes and constitutes the first attempt of protein function evolutionary comparisons based on protein-protein interactions.

Computational Biology↗

APID: Agile Protein Interaction DataAnalyzer.

Agile Protein Interaction DataAnalyzer (APID) is an interactive bioinformatics web tool developed to integrate and analyze in a unified and comparative platform main currently known information about protein-protein interactions demonstrated by specific small-scale or large-scale experimental methods. At present, the application includes information coming from five main source databases enclosing an unified sever to explore >35 000 different proteins and 111 000 different proven interactions. The web includes search tools to query and browse upon the data, allowing selection of the interaction pairs based in calculated parameters that weight and qualify the reliability of each given protein interaction. Such parameters are for the 'proteins': connectivity, cluster coefficient, Gene Ontology (GO) functional environment, GO environment enrichment; and for the 'interactions': number of methods, GO overlapping, iPfam domain-domain interaction. APID also includes a graphic interactive tool to visualize selected sub-networks and to navigate on them or along the whole interaction network. The application is available open access at http://bioinfow.dep.usal.es/apid/.

Computational Biology↗

Analysis of the expressed genome of the lone star tick, Amblyomma americanum (Acari: Ixodidae) using an expressed sequence tag approach.

An expressed sequence tag (EST) approach was used to study the genome of two developmental stages of the lone star tick, Amblyomma americanum. cDNA libraries were constructed from the larval and adult stages of A. americanum. In total, 1942 ESTs were sequenced (1462 adult ESTs and 480 larval ESTs) and analyzed using bioinformatic programs. Contig assembly using the CAPII program revealed 11% and 15% redundancy of sequences in the larval and adult ESTs, respectively. Of the 1942 ESTs, 1738 sequences were considered quality sequences and of these, 771 or approximately 44.4% of the sequences were putatively identified based on amino acid identity using the protein Basic Local Alignment Search Tool (BLAST) algorithm. Putatively identified sequences were classified according to their predicted gene function. In total, 967 sequences, or 55.6% of the quality sequences, had limited or no protein similarity to previously identified gene products. Sequences lacking protein homology were analyzed using an automated sequence annotation system for predicted protein characteristics such as open reading frames, signal peptides, protein motifs, and transmembrane regions. In this paper we describe the sequencing of the largest number of ESTs obtained from an arachnid species to date and the subsequent detailed analysis of these sequences.

Algorithms↗

MyHits: a new interactive resource for protein annotation and domain identification.

The MyHits web server (http://myhits.isb-sib.ch) is a new integrated service dedicated to the annotation of protein sequences and to the analysis of their domains and signatures. Guest users can use the system anonymously, with full access to (i) standard bioinformatics programs (e.g. PSI-BLAST, ClustalW, T-Coffee, Jalview); (ii) a large number of protein sequence databases, including standard (Swiss-Prot, TrEMBL) and locally developed databases (splice variants); (iii) databases of protein motifs (Prosite, Interpro); (iv) a precomputed list of matches ('hits') between the sequence and motif databases. All databases are updated on a weekly basis and the hit list is kept up to date incrementally. The MyHits server also includes a new collection of tools to generate graphical representations of pairwise and multiple sequence alignments including their annotated features. Free registration enables users to upload their own sequences and motifs to private databases. These are then made available through the same web interface and the same set of analytical tools. Registered users can manage their own sequences and annotations using only web tools and freeze their data in their private database for publication purposes.

Computer Graphics↗

OLS4: a new Ontology Lookup Service for a growing interdisciplinary knowledge ecosystem.

SUMMARY: The Ontology Lookup Service (OLS) is an open source search engine for ontologies which is used extensively in the bioinformatics and chemistry communities to annotate biological and biomedical data with ontology terms. Recently, there has been a significant increase in the size and complexity of ontologies due to new scales of biological knowledge, such as spatial transcriptomics, new ontology development methodologies, and curation on an increased scale. Existing Web-based tools for ontology browsing such as BioPortal and OntoBee do not support the full range of definitions used by today's ontologies. In order to support the community going forward, we have developed OLS4, implementing the complete OWL2 specification, internationalization support for multiple languages, and a new user interface with UX enhancements such as links out to external databases. OLS4 has replaced OLS3 in production at EMBL-EBI and has a backward compatible API supporting users of OLS3 to transition. AVAILABILITY AND IMPLEMENTATION: The source code of OLS is available at https://github.com/EBISPOT/ols4 and DOI 10.5281/zenodo.14960290 with Apache 2.0 License. A freely available implementation is accessible at https://www.ebi.ac.uk/ols4.

Biological Ontologies↗

LCR-modules: a collection of workflows for cancer genome analysis.

MOTIVATION: The surge of genomic data from advanced sequencing technologies is outpacing current analytical pipelines. We introduce LCR-modules, an open-source suite of bioinformatics tools designed for flexible and automated cancer genome data analysis. LCR-modules enables reproducible analysis of diverse cancer genomics data at scale. The suite comprises 49 Snakemake-based workflows organized into three levels, facilitating tasks from low-level quality control to complex cohort-level analyses. LCR-modules supports various sequencing types and integrates pipelines such as mutation calling, expression quantification, and cohort-level aggregation, ensuring flexibility and reproducibility. LCR-modules represents a significant advancement in genomic data analysis, reducing barriers in reproducibility and scalability and has already been applied to a combination of exomes and genomes from over 10 800 samples. AVAILABILITY: No new data were generated in support of this research. The source code for the LCR-modules is openly available at https://github.com/LCR-BCCRC/lcr-modules.

Software↗

The group B streptococcal sialic acid O-acetyltransferase is encoded by neuD, a conserved component of bacterial sialic acid biosynthetic gene clusters.

Nearly two dozen microbial pathogens have surface polysaccharides or lipo-oligosaccharides that contain sialic acid (Sia), and several Sia-dependent virulence mechanisms are known to enhance bacterial survival or result in host tissue injury. Some pathogens are also known to O-acetylate their Sias, although the role of this modification in pathogenesis remains unclear. We report that neuD, a gene located within the Group B Streptococcus (GBS) Sia biosynthetic gene cluster, encodes a Sia O-acetyltransferase that is itself required for capsular polysaccharide (CPS) sialylation. Homology modeling and site-directed mutagenesis identified Lys-123 as a critical residue for Sia O-acetyltransferase activity. Moreover, a single nucleotide polymorphism in neuD can determine whether GBS displays a "high" or "low" Sia O-acetylation phenotype. Complementation analysis revealed that Escherichia coli K1 NeuD also functions as a Sia O-acetyltransferase in GBS. In fact, NeuD homologs are commonly found within Sia biosynthetic gene clusters. A bioinformatic approach identified 18 bacterial species with a Sia biosynthetic gene cluster that included neuD. Included in this list are the sialylated human pathogens Legionella pneumophila, Vibrio parahemeolyticus, Pseudomonas aeruginosa, and Campylobacter jejuni, as well as an additional 12 bacterial species never before analyzed for Sia expression. Phylogenetic analysis shows that NeuD homologs of sialylated pathogens share a common evolutionary lineage distinct from the poly-Sia O-acetyltransferase of E. coli K1. These studies define a molecular genetic approach for the selective elimination of GBS Sia O-acetylation without concurrent loss of sialylation, a key to further studies addressing the role(s) of this modification in bacterial virulence.

Acetyltransferases↗

AllerTool: a web server for predicting allergenicity and allergic cross-reactivity in proteins.

UNLABELLED: Assessment of potential allergenicity and patterns of cross-reactivity is necessary whenever novel proteins are introduced into human food chain. Current bioinformatic methods in allergology focus mainly on the prediction of allergenic proteins, with no information on cross-reactivity patterns among known allergens. In this study, we present AllerTool, a web server with essential tools for the assessment of predicted as well as published cross-reactivity patterns of allergens. The analysis tools include graphical representation of allergen cross-reactivity information; a local sequence comparison tool that displays information of known cross-reactive allergens; a sequence similarity search tool for assessment of cross-reactivity in accordance to FAO/WHO Codex alimentarius guidelines; and a method based on support vector machine (SVM). A 10-fold cross-validation results showed that the area under the receiver operating curve (A(ROC)) of SVM models is 0.90 with 86.00% sensitivity (SE) at specificity (SP) of 86.00%. AVAILABILITY: AllerTool is freely available at http://research.i2r.a-star.edu.sg/AllerTool/.

Algorithms↗

An assessment of three dinucleotide parameters to predict DNA curvature by quantitative comparison with experimental data.

Curved DNA fragments are often found near functionally important sites such as promoters and origins of replication, and hence sequence-dependent DNA curvature prediction is of great utility in genomics and bioinformatics. In light of this, an assessment of three different dinucleotide step parameters (based on gel retardation as well as crystal structure data) is carried out. These parameters (BMHT, LB and CS) are evaluated quantitatively for their ability to predict correctly the experimental results of a large set of nucleic acid sequences containing A-tracts as well as GC-rich motifs. This set contained around 40 synthetic as well as natural sequences whose solution properties have been well characterized experimentally. All three models could account reasonably well for curvature in the various DNA sequences. The CS model, where dinucleotide parameters are calculated from crystal structure data, consistently shows slightly better correlation with experimental data. Our simple analysis also indicates that presently available trinucleotide parameters fail to predict curvature in some of the well-characterized sequences. The study shows that the dinucleotide parameters with some further refinement can be used to predict sequence-dependent curvature correctly in genomic sequences.

Base Sequence↗

STRUCTURELAB: a heterogeneous bioinformatics system for RNA structure analysis.

STRUCTURELAB is a computational system that has been developed to permit the use of a broad array of approaches for the analysis of the structure of RNA. The goal of the development is to provide a large set of tools that can be well integrated with experimental biology to aid in the process of the determination of the underlying structure of RNA sequences. The approach taken views the structure determination problem as one of dealing with a database of many computationally generated structures and provides the capability to analyze this data set from different perspectives. Many algorithms are integrated into one system that also utilizes a heterogeneous computing approach permitting the use of several computer architectures to help solve the posed problems. These different computational platforms make it relatively easy to incorporate currently existing programs as well as newly developed algorithms and to best match these algorithms to the appropriate hardware. The system has been written in Common Lisp running on SUN or SGI Unix workstations, and it utilizes a network of participating machines defined in reconfigurable tables. A window-based interface makes this heterogeneous environment as transparent to the user as possible.

Algorithms↗

Analysing the ability to retain sidechain hydrogen-bonds in mutant proteins.

MOTIVATION: Hydrogen bonds are one of the most important inter-atomic interactions in biology. Previous experimental, theoretical and bioinformatics analyses have shown that the hydrogen bonding potential of amino acids is generally satisfied and that buried unsatisfied hydrogen-bond-capable residues are destabilizing. When studying mutant proteins, or introducing mutations to residues involved in hydrogen bonding, one needs to know whether a hydrogen bond can be maintained. Our aim, therefore, was to develop a rapid method to evaluate whether a sidechain can form a hydrogen-bond. RESULTS: A novel knowledge-based approach was developed in which the conformations accessible to the residues involved are taken into account. Residues involved in hydrogen bonds in a set of high resolution crystal structures were analyzed and this analysis is then applied to a given protein. The program was applied to assess mutations in the tumour-suppressor protein, p53. This raised the number of distinct mutations identified as disrupting sidechain-sidechain hydrogen bonding from 181 in our previous analysis to 202 in this analysis.

Algorithms↗

Response of REV3 promoter to N-methyl-N'-nitro-N-nitrosoguanidine.

Previously, we have shown that low concentration of N-methyl-N'-nitro-N-nitrosoguanidine (MNNG) led to the upregulation of REV3 gene at transcriptional level in cultured human amnion FL cells. In this study, using bioinformatic analysis the putative binding sites for different transcription factors were found to exist in REV3 gene promoter region. A 2570-bp fragment of the 5' flanking region of REV3 gene was amplified by PCR from PAC clone RP3-415N12 and inserted into the pGL3-Basic reporter vector. Dual-luciferase reporter assay demonstrated that the reconstructed plasmid did respond to MNNG exposure in transfected FL cells. Several variants of the reporter plasmids with different deletions of the REV3 promoter region were also constructed and their promoter strength was analyzed. It was found that the MNNG response element might locate at the REV3 gene promoter region -404 to -102 between two Sma1 sites. The shortest responsive fragment containing the putative binding sites for transcription factors CREBP, AP-2, NF-kappaB, and SP1 was also identified.

Base Sequence↗

ONCOMINE: a cancer microarray database and integrated data-mining platform.

DNA microarray technology has led to an explosion of oncogenomic analyses, generating a wealth of data and uncovering the complex gene expression patterns of cancer. Unfortunately, due to the lack of a unifying bioinformatic resource, the majority of these data sit stagnant and disjointed following publication, massively underutilized by the cancer research community. Here, we present ONCOMINE, a cancer microarray database and web-based data-mining platform aimed at facilitating discovery from genome-wide expression analyses. To date, ONCOMINE contains 65 gene expression datasets comprising nearly 48 million gene expression measurements form over 4700 microarray experiments. Differential expression analyses comparing most major types of cancer with respective normal tissues as well as a variety of cancer subtypes and clinical-based and pathology-based analyses are available for exploration. Data can be queried and visualized for a selected gene across all analyses or for multiple genes in a selected analysis. Furthermore, gene sets can be limited to clinically important annotations including secreted, kinase, membrane, and known gene-drug target pairs to facilitate the discovery of novel biomarkers and therapeutic targets.

Databases, Genetic↗

ExPASy: The proteomics server for in-depth protein knowledge and analysis.

The ExPASy (the Expert Protein Analysis System) World Wide Web server (http://www.expasy.org), is provided as a service to the life science community by a multidisciplinary team at the Swiss Institute of Bioinformatics (SIB). It provides access to a variety of databases and analytical tools dedicated to proteins and proteomics. ExPASy databases include SWISS-PROT and TrEMBL, SWISS-2DPAGE, PROSITE, ENZYME and the SWISS-MODEL repository. Analysis tools are available for specific tasks relevant to proteomics, similarity searches, pattern and profile searches, post-translational modification prediction, topology prediction, primary, secondary and tertiary structure analysis and sequence alignment. These databases and tools are tightly interlinked: a special emphasis is placed on integration of database entries with related resources developed at the SIB and elsewhere, and the proteomics tools have been designed to read the annotations in SWISS-PROT in order to enhance their predictions. ExPASy started to operate in 1993, as the first WWW server in the field of life sciences. In addition to the main site in Switzerland, seven mirror sites in different continents currently serve the user community.

Databases, Protein↗

The MPI Bioinformatics Toolkit for protein sequence analysis.

The MPI Bioinformatics Toolkit is an interactive web service which offers access to a great variety of public and in-house bioinformatics tools. They are grouped into different sections that support sequence searches, multiple alignment, secondary and tertiary structure prediction and classification. Several public tools are offered in customized versions that extend their functionality. For example, PSI-BLAST can be run against regularly updated standard databases, customized user databases or selectable sets of genomes. Another tool, Quick2D, integrates the results of various secondary structure, transmembrane and disorder prediction programs into one view. The Toolkit provides a friendly and intuitive user interface with an online help facility. As a key feature, various tools are interconnected so that the results of one tool can be forwarded to other tools. One could run PSI-BLAST, parse out a multiple alignment of selected hits and send the results to a cluster analysis tool. The Toolkit framework and the tools developed in-house will be packaged and freely available under the GNU Lesser General Public Licence (LGPL). The Toolkit can be accessed at http://toolkit.tuebingen.mpg.de.

Computational Biology↗

InterProScan: protein domains identifier.

InterProScan [E. M. Zdobnov and R. Apweiler (2001) Bioinformatics, 17, 847-848] is a tool that combines different protein signature recognition methods from the InterPro [N. J. Mulder, R. Apweiler, T. K. Attwood, A. Bairoch, A. Bateman, D. Binns, P. Bradley, P. Bork, P. Bucher, L. Cerutti et al. (2005) Nucleic Acids Res., 33, D201-D205] consortium member databases into one resource. At the time of writing there are 10 distinct publicly available databases in the application. Protein as well as DNA sequences can be analysed. A web-based version is accessible for academic and commercial organizations from the EBI (http://www.ebi.ac.uk/InterProScan/). In addition, a standalone Perl version and a SOAP Web Service [J. Snell, D. Tidwell and P. Kulchenko (2001) Programming Web Services with SOAP, 1st edn. O'Reilly Publishers, Sebastopol, CA, http://www.w3.org/TR/soap/] are also available to the users. Various output formats are supported and include text tables, XML documents, as well as various graphs to help interpret the results.

Databases, Protein↗

A contact energy function considering residue hydrophobic environment and its application in protein fold recognition.

The three-dimensional (3D) structure prediction of proteins is an important task in bioinformatics. Finding energy functions that can better represent residue-residue and residue-solvent interactions is a crucial way to improve the prediction accuracy. The widely used contact energy functions mostly only consider the contact frequency between different types of residues; however, we find that the contact frequency also relates to the residue hydrophobic environment. Accordingly, we present an improved contact energy function to integrate the two factors, which can reflect the influence of hydrophobic interaction on the stabilization of protein 3D structure more effectively. Furthermore, a fold recognition (threading) approach based on this energy function is developed. The testing results obtained with 20 randomly selected proteins demonstrate that, compared with common contact energy functions, the proposed energy function can improve the accuracy of the fold template prediction from 20% to 50%, and can also improve the accuracy of the sequence-template alignment from 35% to 65%.

Amino Acid Sequence↗

A deterministic finite automaton for faster protein hit detection in BLAST.

BLAST is the most popular bioinformatics tool and is used to run millions of queries each day. However, evaluating such queries is slow, taking typically minutes on modern workstations. Therefore, continuing evolution of BLAST--by improving its algorithms and optimizations--is essential to improve search times in the face of exponentially increasing collection sizes. We present an optimization to the first stage of the BLAST algorithm specifically designed for protein search. It produces the same results as NCBI-BLAST but in around 59% of the time on Intel-based platforms; we also present results for other popular architectures. Overall, this is a saving of around 15% of the total typical BLAST search time. Our approach uses a deterministic finite automaton (DFA), inspired by the original scheme used in the 1990 BLAST algorithm. The techniques are optimized for modern hardware, making careful use of cache-conscious approaches to improve speed. Our optimized DFA approach has been integrated into a new version of BLAST that is freely available for download at http://www.fsa-blast.org/.

Algorithms↗