Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences.

MOTIVATION: In 2001 and 2002, we published two papers (Bioinformatics, 17, 282-283, Bioinformatics, 18, 77-82) describing an ultrafast protein sequence clustering program called cd-hit. This program can efficiently cluster a huge protein database with millions of sequences. However, the applications of the underlying algorithm are not limited to only protein sequences clustering, here we present several new programs using the same algorithm including cd-hit-2d, cd-hit-est and cd-hit-est-2d. Cd-hit-2d compares two protein datasets and reports similar matches between them; cd-hit-est clusters a DNA/RNA sequence database and cd-hit-est-2d compares two nucleotide datasets. All these programs can handle huge datasets with millions of sequences and can be hundreds of times faster than methods based on the popular sequence comparison and database search tools, such as BLAST.

Algorithms↗

State-of-the-art in membrane protein prediction.

Membrane proteins are crucial for many biological functions and have become attractive targets for pharmacological agents. About 10%-30% of all proteins contain membrane-spanning helices. Despite recent successes, high-resolution structures for membrane proteins remain exceptional. The gap between known sequences and known structures calls for finding solutions through bioinformatics. While many methods predict membrane helices, very few predict membrane strands. The good news is that most methods for helical membrane proteins are available and are more often right than wrong. The best current prediction methods appear to correctly predict all membrane helices for about 50%-70% of all proteins, and to falsely predict membrane helices for about 10% of all globular proteins. The bad news is that developers have seriously overestimated the accuracy of their methods. In particular, while simple hydrophobicity scales identify many membrane helices, they frequently and incorrectly predict membrane helices in globular proteins. Additionally, all methods tend to confuse signal peptides with membrane helices. Nonetheless, wet-lab biologists can reach into an impressive toolbox for membrane protein predictions. However, the computational biologists will have to improve their methods considerably before they reach the levels of accuracy they claim.

Computational Biology↗

Toward prediction of class II mouse major histocompatibility complex peptide binding affinity: in silico bioinformatic evaluation using partial least squares, a robust multivariate statistical technique.

The accurate identification of T-cell epitopes remains a principal goal of bioinformatics within immunology. As the immunogenicity of peptide epitopes is dependent on their binding to major histocompatibility complex (MHC) molecules, the prediction of binding affinity is a prerequisite to the reliable prediction of epitopes. The iterative self-consistent (ISC) partial-least-squares (PLS)-based additive method is a recently developed bioinformatic approach for predicting class II peptide-MHC binding affinity. The ISC-PLS method overcomes many of the conceptual difficulties inherent in the prediction of class II peptide-MHC affinity, such as the binding of a mixed population of peptide lengths due to the open-ended class II binding site. The method has applications in both the accurate prediction of class II epitopes and the manipulation of affinity for heteroclitic and competitor peptides. The method is applied here to six class II mouse alleles (I-Ab, I-Ad, I-Ak, I-As, I-Ed, and I-Ek) and included peptides up to 25 amino acids in length. A series of regression equations highlighting the quantitative contributions of individual amino acids at each peptide position was established. The initial model for each allele exhibited only moderate predictivity. Once the set of selected peptide subsequences had converged, the final models exhibited a satisfactory predictive power. Convergence was reached between the 4th and 17th iterations, and the leave-one-out cross-validation statistical terms--q2, SEP, and NC--ranged between 0.732 and 0.925, 0.418 and 0.816, and 1 and 6, respectively. The non-cross-validated statistical terms r2 and SEE ranged between 0.98 and 0.995 and 0.089 and 0.180, respectively. The peptides used in this study are available from the AntiJen database (http://www.jenner.ac.uk/AntiJen). The PLS method is available commercially in the SYBYL molecular modeling software package. The resulting models, which can be used for accurate T-cell epitope prediction, will be made freely available online (http://www.jenner.ac.uk/MHCPred).

Animals↗

Super pairwise alignment (SPA): an efficient approach to global alignment for homologous sequences.

Sequence analysis is the basis of bioinformatics, while sequence alignment is a fundamental task for sequence analysis. The widely used alignment algorithm, Dynamic Programming, though generating optimal alignment, takes too much time due to its high computation complexity O(N(2)). In order to reduce computation complexity without sacrificing too much accuracy, we have developed a new approach to align two homologous sequences. The new approach presented here, adopting our novel algorithm which combines the methods of probabilistic and combinatorial analysis, reduces the computation complexity to as low as O(N). The computation speed by our program is at least 15 times faster than traditional pairwise alignment algorithms without a loss of much accuracy. We hence named the algorithm Super Pairwise Alignment (SPA). The pairwise alignment execution program based on SPA and the detailed results of the aligned sequences discussed in this article are available upon request.

Algorithms↗

The web server of IBM's Bioinformatics and Pattern Discovery group: 2004 update.

In this report, we provide an update on the services and content which are available on the web server of IBM's Bioinformatics and Pattern Discovery group. The server, which is operational around the clock, provides access to a large number of methods that have been developed and published by the group's members. There is an increasing number of problems that these tools can help tackle; these problems range from the discovery of patterns in streams of events and the computation of multiple sequence alignments, to the discovery of genes in nucleic acid sequences, the identification--directly from sequence--of structural deviations from alpha-helicity and the annotation of amino acid sequences for antimicrobial activity. Additionally, annotations for more than 130 archaeal, bacterial, eukaryotic and viral genomes are now available on-line and can be searched interactively. The tools and code bundles continue to be accessible from http://cbcsrv.watson.ibm.com/Tspd.html whereas the genomics annotations are available at http://cbcsrv.watson.ibm.com/Annotations/.

Anti-Infective Agents↗

A systematic approach to infer biological relevance and biases of gene network structures.

The development of high-throughput technologies has generated the need for bioinformatics approaches to assess the biological relevance of gene networks. Although several tools have been proposed for analysing the enrichment of functional categories in a set of genes, none of them is suitable for evaluating the biological relevance of the gene network. We propose a procedure and develop a web-based resource (BIOREL) to estimate the functional bias (biological relevance) of any given genetic network by integrating different sources of biological information. The weights of the edges in the network may be either binary or continuous. These essential features make our web tool unique among many similar services. BIOREL provides standardized estimations of the network biases extracted from independent data. By the analyses of real data we demonstrate that the potential application of BIOREL ranges from various benchmarking purposes to systematic analysis of the network biology.

Computational Biology↗

Dynamic tables: an architecture for managing evolving, heterogeneous biomedical data in relational database management systems.

Data sparsity and schema evolution issues affecting clinical informatics and bioinformatics communities have led to the adoption of vertical or object-attribute-value-based database schemas to overcome limitations posed when using conventional relational database technology. This paper explores these issues and discusses why biomedical data are difficult to model using conventional relational techniques. The authors propose a solution to these obstacles based on a relational database engine using a sparse, column-store architecture. The authors provide benchmarks comparing the performance of queries and schema-modification operations using three different strategies: (1) the standard conventional relational design; (2) past approaches used by biomedical informatics researchers; and (3) their sparse, column-store architecture. The performance results show that their architecture is a promising technique for storing and processing many types of data that are not handled well by the other two semantic data models.

Computational Biology↗

A diagnostic test for prostate cancer from gene expression profiling data.

PURPOSE: Multiple recent studies show excellent classification accuracy using bioinformatics tools applied to expression profiling data on various tumors. However, the clinical applicability of these techniques remains unfulfilled because of difficulty in translating complex multigene mathematical algorithms into reproducible, platform independent tests. We recently developed a broadly applicable platform independent method based on simple ratios of gene expression to diagnose and predict outcome in cancer. In the current study we applied this technique to the diagnosis of prostate cancer. MATERIALS AND METHODS: We developed a ratio based predictive model using a training set of 32 samples with previously published gene profiling data. We then tested and refined the model using additional independent samples with previously published microarray data from another source (that is the test set of 34 samples). Finally, the optimal ratio based test was examined with quantitative reverse transcriptase-polymerase chain reaction for data acquisition in a third cohort of samples consisting of 10 frozen normal and 10 tumor prostate tissues. RESULTS: A 3-ratio test using 4 genes was 90% accurate (18 of 20 samples) for distinguishing normal prostate and prostate cancer samples obtained at surgery (Fisher's exact test p = 0.0007). This test did not result in any false-negative findings. CONCLUSIONS: We describe and validate a new gene ratio based test for the diagnosis of prostate cancer, which was developed from the analysis of extensive gene profiling data for the diagnosis of prostate cancer. This test can be easily adapted to the clinical arena without the need for complex computer software or hardware. We anticipate that the gene ratio based diagnosis of prostate cancer using fine needle aspirations could serve as a useful adjunct to standard histopathological techniques.

Gene Expression Profiling↗

Representations of molecular pathways: an evaluation of SBML, PSI MI and BioPAX.

MOTIVATION: Analysis and simulation of pathway data is of high importance in bioinformatics. Standards for representation of information about pathways are necessary for integration and analysis of data from various sources. Recently, a number of representation formats for pathway data, SBML, PSI MI and BioPAX, have been proposed. RESULTS: In this paper we compare these formats and evaluate them with respect to their underlying models, information content and possibilities for easy creation of tools. The evaluation shows that the main structure of the formats is similar. However, SBML is tuned towards simulation models of molecular pathways while PSI MI is more suitable for representing details about particular interactions and experiments. BioPAX is the most general and expressive of the formats. These differences are apparent in allowed information and the structure for representation of interactions. We discuss the impact of these differences both with respect to information content in existing databases and computational properties for import and analysis of data.

Computational Biology↗

SHARP2: protein-protein interaction predictions using patch analysis.

UNLABELLED: SHARP2 is a flexible web-based bioinformatics tool for predicting potential protein-protein interaction sites on protein structures. It implements a predictive algorithm that calculates multiple parameters for overlapping patches of residues on the surface of a protein. Six parameters are calculated: solvation potential, hydrophobicity, accessible surface area, residue interface propensity, planarity and protrusion (SHARP2). Parameter scores for each patch are combined, and the patch with the highest combined score is predicted as a potential interaction site. SHARP2 enables users to upload 3D protein structure files in PDB format, to obtain information on potential interaction sites as downloadable HTML tables and to view the location of the sites on the 3D structure using Jmol. The server allows for the input of multiple structures and multiple combinations of parameters. Therefore predictions can be made for complete datasets, as well as individual structures. AVAILABILITY: http://www.bioinformatics.sussex.ac.uk/SHARP2.

Algorithms↗

Identification of protein pattern in kidney cancer using ProteinChip arrays and bioinformatics.

Tumor biology of renal cell carcinoma (RCC) is not very well understood, although many studies on molecular and cellular biology have been performed. It is accepted now that cancer research has to be performed also with proteomic tools, because proteins are the real actors in the genesis and progression of cancer. Therefore, we used a ProteinChip System(R) (SELDI) which is able to detect minute amounts of protein and moreover to analyze a complex protein pattern. We analyzed 37 cases of clear cell RCC as a training set including corresponding normal tissue. From all samples protein lysates were made and spotted directly on different chip surfaces (SAX2, WCX). After a washing procedure the arrays were analyzed in the ProteinChip Reader. All profiles were subjected to a bioinformatical analysis including normalization, clustering, rule extraction and rating. Defined rules (markers) were evaluated using a test set of 24 samples (13 tumor tissues and 11 normal kidney tissues). The generated rule base for the SAX2 surface showed a sensitivity of 100% and a specificity of 97.3%. For the WCX arrays the optimal rule base showed worse results. A combined rule base for SAX2 and WCX did not result in a higher sensitivity or specificity. Using the optimal rule base for the SAX2 chip in the test set, sensitivity and specificity reached 76.9% and 100%, respectively. The ProteinChip System represents a key technology for the rapid detection of cancer specific proteomic patterns. It is possible to identify clear cell renal cancer with high sensitivity and specificity from minimal amounts of cells.

Cell Line, Tumor↗

Bioinformatics: use in bacterial vaccine discovery.

Bioinformatics has now become a common laboratory name for groups studying genomic sequences. It is composed of many different, yet interrelated scientific fields such as genomics, proteomics, and transcriptional profiling. The availability of complete genomic sequences, especially prokaryotic organisms, allows one to rapidly identify, analyze, and clone genes of interest. For bacterial vaccine discovery, one can "mine" the genomic sequence for potential surface targets using various algorithms, characterize these gene targets, and produce primers for cloning, all before one enters the wet laboratory. This review will focus on various genomic mining tools/algorithms available for predicting open reading frames and their associated annotation (if known), physical and functional characterization, and cellular localization. Finally, examples are given of how all of this is being used for the identification of potential bacterial vaccine candidates.

Animals↗

AI-HOPE: an AI-driven conversational agent for enhanced clinical and genomic data integration in precision medicine research.

MOTIVATION: The growing complexity of clinical cancer research has fueled a surge in demand for automated bioinformatics tools capable of integrating clinical and genomic data to accelerate discovery efforts. RESULTS: We present the Artificial Intelligence Agent for High-Optimization and Precision Medicine (AI-HOPE), an AI-driven system that enables domain experts to conduct integrative data analyses through natural language interactions. Powered by Large Language Models, AI-HOPE interprets user instructions, converts them into executable code, and autonomously analyzes locally stored data. It supports flexible association studies, subset comparisons, clinical prevalence assessments and survival analyses. In addition, AI-HOPE enables global variable scans to identify features significantly associated with a user-defined outcome, making a powerful and intuitive tool for advancing precision medicine research. Importantly, its closed-system design prevents clinical data leakage. To demonstrate its utility, AI-HOPE was applied to The Cancer Genome Atlas data to address two clinical questions. First, it identified significant enrichment of TP53 mutations in late-stage colorectal cancer compared to early-stage cases. Second, it uncovered a strong association between KRAS mutations and poorer progression-free survival in FOLFOX-treated patients. These findings align with established literature and demonstrate AI-HOPE's ability to generate meaningful insights independently, without prior assumptions. By removing programming barriers and simplifying complex analyses, AI-HOPE bridges the gap between data complexity and research needs. With its scalable and adaptable framework, AI-HOPE has the potential to support diverse biomedical research fields, driving innovation and efficiency in translational studies. AVAILABILITY AND IMPLEMENTATION: The AI-HOPE software and demonstration data is available at https://github.com/Velazquez-Villarreal-Lab/AI-HOPE.

Precision Medicine↗

An attempt to define allergen-specific molecular surface features: a bioinformatic approach.

Allergens are proteins that elicit T helper lymphocyte type 2 (Th2) responses culminating in IgE antibody production and allergic disease. However, we have no answer to the fundamental question of why certain proteins are allergens, while others are not. We hypothesized that analysis of the surface of diverse allergens may reveal common structural features which might enable them to be recognized as Th2-inducing antigens by cells of the innate immune system. We have therefore used the ConSurf server to search for allergen-specific motifs. This has enabled us to identify residue conservation patterns in the homologues of Ara t 8 (plant profilin), Act c 1 (actinidin), Bet v 1 (plant pathogenesis-related protein) and Ves v 5 (venom allergen). The results demonstrate the presence of allergen-specific patches consisting of an unusually high proportion of surface-exposed hydrophobic residues. The patches that have been identified may represent molecular patterns recognizable by cells of the innate immune system.

Algorithms↗

A motif-based profile scanning approach for genome-wide prediction of signaling pathways.

The rapid increase in genomic information requires new techniques to infer protein function and predict protein-protein interactions. Bioinformatics identifies modular signaling domains within protein sequences with a high degree of accuracy. In contrast, little success has been achieved in predicting short linear sequence motifs within proteins targeted by these domains to form complex signaling networks. Here we describe a peptide library-based searching algorithm, accessible over the World Wide Web, that identifies sequence motifs likely to bind to specific protein domains such as 14-3-3, SH2, and SH3 domains, or likely to be phosphorylated by specific protein kinases such as Src and AKT. Predictions from database searches for proteins containing motifs matching two different domains in a common signaling pathway provides a much higher success rate. This technology facilitates prediction of cell signaling networks within proteomes, and could aid in the identification of drug targets for the treatment of human diseases.

Algorithms↗

Evolutionary mapping of the SHV beta-lactamase and evidence for two separate IS26-dependent blaSHV mobilization events from the Klebsiella pneumoniae chromosome.

OBJECTIVES: To determine the most likely evolutionary pathway that has led to the development of extended-spectrum SHV derivatives, and to the mobilization of blaSHV. MATERIALS AND METHODS: Evolutionary mapping used a shortest-path analysis of aligned blaSHV variants, and other basic bioinformatic approaches, such as CLUSTAL W and Blast were employed. RESULTS: Two main branches of the blaSHV evolutionary tree were located; both are derived from variant blaSHV-1 alleles. Identical mutations, responsible for extended-spectrum SHV substrate profiles, have been selected independently in each branch. There is evidence for a pool of non-mobile blaSHV framework sequences. Analysis of the genome sequence of Klebsiella pneumoniae confirms the chromosomal origin of blaSHV, whose mobilization has occurred at least twice, once for each of the main evolutionary branches. Both these mobilization events have been catalysed by IS26. Evolution of blaSHV to give common extended-spectrum variants is most likely to have occurred following mobilization. CONCLUSIONS: These data shed new light on the evolution and mobilization of blaSHV, and these observations may be useful in predicting what might happen in future, both for blaSHV, and for other beta-lactamase genes.

Alleles↗

Analysis of molecular recognition features (MoRFs).

Several proteomic studies in the last decade revealed that many proteins are either completely disordered or possess long structurally flexible regions. Many such regions were shown to be of functional importance, often allowing a protein to interact with a large number of diverse partners. Parallel to these findings, during the last five years structural bioinformatics has produced an explosion of results regarding protein-protein interactions and their importance for cell signaling. We studied the occurrence of relatively short (10-70 residues), loosely structured protein regions within longer, largely disordered sequences that were characterized as bound to larger proteins. We call these regions molecular recognition features (MoRFs, also known as molecular recognition elements, MoREs). Interestingly, upon binding to their partner(s), MoRFs undergo disorder-to-order transitions. Thus, in our interpretation, MoRFs represent a class of disordered region that exhibits molecular recognition and binding functions. This work extends previous research showing the importance of flexibility and disorder for molecular recognition. We describe the development of a database of MoRFs derived from the RCSB Protein Data Bank and present preliminary results of bioinformatics analyses of these sequences. Based on the structure adopted upon binding, at least three basic types of MoRFs are found: alpha-MoRFs, beta-MoRFs, and iota-MoRFs, which form alpha-helices, beta-strands, and irregular secondary structure when bound, respectively. Our data suggest that functionally significant residual structure can exist in MoRF regions prior to the actual binding event. The contribution of intrinsic protein disorder to the nature and function of MoRFs has also been addressed. The results of this study will advance the understanding of protein-protein interactions and help towards the future development of useful protein-protein binding site predictors.

Algorithms↗

AEGIS: an annotation extraction and genomic integration resource.

MOTIVATION: Genome annotation files (GFF3/GTF) are the standard for storing genomic feature data, yet their flexibility often results in formatting inconsistencies that create bottlenecks for downstream bioinformatics analyses. A robust, unified framework is required to parse, standardise, and validate these files to ensure interoperability and facilitate complex comparative genomic tasks. RESULTS: We present AEGIS (Annotation Extraction and Genomic Integration Suite), a comprehensive toolkit designed to parse, correct, and standardise genome annotations. Beyond quality control, AEGIS provides advanced modules for flexible feature extraction (e.g., coding sequences, promoters) and comparative genomic analysis. Uniquely, it integrates multiple lines of evidence, including sequence homology, synteny, and coordinate-based lift-overs, to assess gene model correspondence and infer orthology. We demonstrate the utility of AEGIS by quantifying complex structural changes between Arabidopsis annotation versions and identifying high-confidence orthologues across diverse plant genomes. AVAILABILITY: AEGIS is implemented in Python. Source code and documentation are freely available under the GPL-3 license at https://github.com/Tomsbiolab/aegis and as a Docker container at https://hub.docker.com/r/tomsbiolab/aegis. The package is also available on PyPI (pip install aegis-bio).

Software↗