Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Predicting rRNA-, RNA-, and DNA-binding proteins from primary structure with support vector machines.

In the post-genome era, the prediction of protein function is one of the most demanding tasks in the study of bioinformatics. Machine learning methods, such as the support vector machines (SVMs), greatly help to improve the classification of protein function. In this work, we integrated SVMs, protein sequence amino acid composition, and associated physicochemical properties into the study of nucleic-acid-binding proteins prediction. We developed the binary classifications for rRNA-, RNA-, DNA-binding proteins that play an important role in the control of many cell processes. Each SVM predicts whether a protein belongs to rRNA-, RNA-, or DNA-binding protein class. Self-consistency and jackknife tests were performed on the protein data sets in which the sequences identity was < 25%. Test results show that the accuracies of rRNA-, RNA-, DNA-binding SVMs predictions are approximately 84%, approximately 78%, approximately 72%, respectively. The predictions were also performed on the ambiguous and negative data set. The results demonstrate that the predicted scores of proteins in the ambiguous data set by RNA- and DNA-binding SVM models were distributed around zero, while most proteins in the negative data set were predicted as negative scores by all three SVMs. The score distributions agree well with the prior knowledge of those proteins and show the effectiveness of sequence associated physicochemical properties in the protein function prediction. The software is available from the author upon request.

Amino Acid Sequence↗

Optimizing sparse and skew hashing: faster k-mer dictionaries.

MOTIVATION: Representing a set of k-mers-strings of length k-in small space under fast lookup queries is a fundamental requirement for several applications in Bioinformatics. A data structure based on sparse and skew hashing (SSHash) was recently proposed for this purpose (Pibiri 2022): it combines good space effectiveness with fast lookup and streaming queries. It is also order-preserving, i.e. consecutive k-mers (sharing a prefix-suffix overlap of length k-1) are assigned consecutive hash codes which helps compressing satellite data typically associated with k-mers, like abundances and color sets in colored De Bruijn graphs. RESULTS: We study the problem of accelerating queries under the sparse and skew hashing indexing paradigm, without compromising its space effectiveness. We propose a refined data structure with less complex lookups and fewer cache misses. We give a simpler and faster algorithm for streaming lookup queries. The refined architecture translates to substantial performance gains, outperforming the original version of SSHash in both index construction speed and query efficiency. Compared to indexes with similar capabilities and based on the Burrows-Wheeler transform, like SBWT and FMSI, SSHash is significantly faster to build and query. SSHash is competitive in space with the fast (and default) modality of SBWT when both k-mer strands are indexed. While larger than FMSI, it is also more than one order of magnitude faster to query. AVAILABILITY AND IMPLEMENTATION: The SSHash software is available at https://github.com/jermp/sshash, and also distributed via Bioconda. A benchmark of data structures for k-mer sets is available at https://github.com/jermp/kmer_sets_benchmark. The datasets used in this article are described and available at https://zenodo.org/records/17582116.

Algorithms↗

A tool for sharing annotated research data: the "Category 0" UMLS (Unified Medical Language System) vocabularies.

BACKGROUND: Large biomedical data sets have become increasingly important resources for medical researchers. Modern biomedical data sets are annotated with standard terms to describe the data and to support data linking between databases. The largest curated listing of biomedical terms is the the National Library of Medicine's Unified Medical Language System (UMLS). The UMLS contains more than 2 million biomedical terms collected from nearly 100 medical vocabularies. Many of the vocabularies contained in the UMLS carry restrictions on their use, making it impossible to share or distribute UMLS-annotated research data. However, a subset of the UMLS vocabularies, designated Category 0 by UMLS, can be used to annotate and share data sets without violating the UMLS License Agreement. METHODS: The UMLS Category 0 vocabularies can be extracted from the parent UMLS metathesaurus using a Perl script supplied with this article. There are 43 Category 0 vocabularies that can be used freely for research purposes without violating the UMLS License Agreement. Among the Category 0 vocabularies are: MESH (Medical Subject Headings), NCBI (National Center for Bioinformatics) Taxonomy and ICD-9-CM (International Classification of Diseases-9-Clinical Modifiers). RESULTS: The extraction file containing all Category 0 terms and concepts is 72,581,138 bytes in length and contains 1,029,161 terms. The UMLS Metathesaurus MRCON file (January, 2003) is 151,048,493 bytes in length and contains 2,146,899 terms. Therefore the Category 0 vocabularies, in aggregate, are about half the size of the UMLS metathesaurus.A large publicly available listing of 567,921 different medical phrases were automatically coded using the full UMLS metatathesaurus and the Category 0 vocabularies. There were 545,321 phrases with one or more matches against UMLS terms while 468,785 phrases had one or more matches against the Category 0 terms. This indicates that when the two vocabularies are evaluated by their fitness to find at least one term for a medical phrase, the Category 0 vocabularies performed 86% as well as the complete UMLS metathesaurus. CONCLUSION: The Category 0 vocabularies of UMLS constitute a large nomenclature that can be used by biomedical researchers to annotate biomedical data. These annotated data sets can be distributed for research purposes without violating the UMLS License Agreement. These vocabularies may be of particular importance for sharing heterogeneous data from diverse biomedical data sets. The software tools to extract the Category 0 vocabularies are freely available Perl scripts entered into the public domain and distributed with this article.

Algorithms↗

Toward an atomistic model for predicting transcription-factor binding sites.

Identifying the specific DNA-binding sites of transcription-factor proteins is essential to understanding the regulation of gene expression in the cell. Bioinformatics approaches are fast compared to experiments, but require prior knowledge of multiple binding sites for each protein. Here, we present an atomistic force-field method to predict binding sites based only on the X-ray structure of a related bound complex. Specific flexible contacts between the protein and DNA are modeled by a library of amino acid side-chain rotamers. Using the example of the mouse transcription factor, Zif268, a well-studied zinc-finger protein, we show that the protein sequence alone, without the detailed experimental structure, gives a strong bias toward the consensus binding site.

Algorithms↗

Membrane topology of the yeast endoplasmic reticulum-localized ubiquitin ligase Doa10 and comparison with its human ortholog TEB4 (MARCH-VI).

Quality control machinery in the endoplasmic reticulum (ER) helps ensure that only properly folded and assembled proteins accumulate in the ER or continue along the secretory pathway. Aberrant proteins are retrotranslocated to the cytosol and degraded by the proteasome, a process called ER-associated degradation. Doa10, a transmembrane protein of the ER/nuclear envelope, is one of the primary ubiquitin ligases (E3s) participating in ER-associated degradation in Saccharomyces cerevisiae. Here we report the membrane organization of the 1319-residue Doa10 polypeptide. The topology was determined by fusing a dual-topology reporter after 16 different Doa10 fragments. Our results indicate that Doa10 contains 14 transmembrane helices (TMs). Based on protease digestion of yeast microsomes, both the N-terminal RING-CH domain and the C terminus face the cytosol. Notably, the experimentally derived topology was not predicted correctly by any of the generally available TM prediction algorithms. Bioinformatic analysis and in silico mutagenesis guided the topological studies through problematic regions. The conserved TD domain in Doa10 includes three TMs. These TMs might function in cofactor binding or substrate recognition, or they might be part of a retrotranslocation channel. The Derlins were previously proposed to provide such channels, but we show that the two yeast Derlins are not required for degradation of Doa10 membrane substrates, as was found before for the Sec61 translocon. Finally, we provide evidence that the likely human Doa10 ortholog, TEB4 (MARCH-VI), adopts a topology similar to that of Doa10.

Algorithms↗

PLAID: ultrafast single-sample gene set enrichment scoring.

SUMMARY: In recent years, computational methods have emerged that calculate enrichment of gene signatures within individual samples. These signatures offer critical insights into the coordinated activity of functionally related genes, proteins or metabolites, enabling the identification of unique molecular profiles in individual cells and patients. This strategy is pivotal for patient stratification and advancement of personalized medicine. However, the rise of large-scale datasets, including single-cell profiles and population biobanks, has exposed significant computational inefficiencies in existing methods. Current methods often demand excessive runtime and memory resources, becoming impractical for large datasets. Overcoming these limitations is a focus of current efforts by bioinformatics teams in academia and the pharmaceutical industry, as essential to support basic and clinical biomedical research. To address this critical need, we developed PLAID (Pathway Level Average Intensity Detection), an ultrafast and memory optimized single sample gene set enrichment algorithm that utilizes sparse matrix computation. PLAID delivers highly accurate gene set scoring and surpasses the performance of current methods in single-cell and bulk transcriptomics, and proteomics data. PLAID uniquely integrates the most widely used gene set scoring algorithms, enabling researchers to apply multiple methods for cross-validation with outstanding runtime efficiency and minimal memory requirement. AVAILABILITY AND IMPLEMENTATION: PLAID is implemented in the R language for statistical computing. PLAID source code and installation instructions are available with no restrictions at https://github.com/bigomics/plaid.

Algorithms↗

High performance workflow implementation for protein surface characterization using grid technology.

BACKGROUND: This study concerns the development of a high performance workflow that, using grid technology, correlates different kinds of Bioinformatics data, starting from the base pairs of the nucleotide sequence to the exposed residues of the protein surface. The implementation of this workflow is based on the Italian Grid.it project infrastructure, that is a network of several computational resources and storage facilities distributed at different grid sites. METHODS: Workflows are very common in Bioinformatics because they allow to process large quantities of data by delegating the management of resources to the information streaming. Grid technology optimizes the computational load during the different workflow steps, dividing the more expensive tasks into a set of small jobs. RESULTS: Grid technology allows efficient database management, a crucial problem for obtaining good results in Bioinformatics applications. The proposed workflow is implemented to integrate huge amounts of data and the results themselves must be stored into a relational database, which results as the added value to the global knowledge. CONCLUSION: A web interface has been developed to make this technology accessible to grid users. Once the workflow has started, by means of the simplified interface, it is possible to follow all the different steps throughout the data processing. Eventually, when the workflow has been terminated, the different features of the protein, like the amino acids exposed on the protein surface, can be compared with the data present in the output database.

Automation↗

CEAS: cis-regulatory element annotation system.

The recent availability of high-density human genome tiling arrays enables biologists to conduct ChIP-chip experiments to locate the in vivo-binding sites of transcription factors in the human genome and explore the regulatory mechanisms. Once genomic regions enriched by transcription factor ChIP-chip are located, genome-scale downstream analyses are crucial but difficult for biologists without strong bioinformatics support. We designed and implemented the first web server to streamline the ChIP-chip downstream analyses. Given genome-scale ChIP regions, the cis-regulatory element annotation system (CEAS) retrieves repeat-masked genomic sequences, calculates GC content, plots evolutionary conservation, maps nearby genes and identifies enriched transcription factor-binding motifs. Biologists can utilize CEAS to retrieve useful information for ChIP-chip validation, assemble important knowledge to include in their publication and generate novel hypotheses (e.g. transcription factor cooperative partner) for further study. CEAS helps the adoption of ChIP-chip in mammalian systems and provides insights towards a more comprehensive understanding of transcriptional regulatory mechanisms. The URL of the server is http://ceas.cbi.pku.edu.cn.

Binding Sites↗

PoPS: a computational tool for modeling and predicting protease specificity.

Proteases play a fundamental role in the control of intra- and extracellular processes by binding and cleaving specific amino acid sequences. Identifying these targets is extremely challenging. Current computational attempts to predict cleavage sites are limited, representing these amino acid sequences as patterns or frequency matrices. Here we present PoPS, a publicly accessible bioinformatics tool (http://pops.csse.monash.edu.au/) which provides a novel method for building computational models of protease specificity that, while still being based on these amino acid sequences, can be built from any experimental data or expert knowledge available to the user. PoPS specificity models can be used to predict and rank likely cleavages within a single substrate, and within entire proteomes. Other factors, such as the secondary or tertiary structure of the substrate, can be used to screen unlikely sites. Furthermore, the tool also provides facilities to infer, compare and test models, and to store them in a publicly accessible database.

Amino Acid Sequence↗

Toward computer-based cleavage site prediction of cysteine endopeptidases.

Identification of relevant substrates is essential for elucidation of in vivo functions of peptidases. The recent availability of the complete genome sequences of many eukaryotic organisms holds the promise of identifying specific peptidase substrates by systematic proteome analyses in combination with computer-based screening of genome databases. Currently available proteomics and bioinformatics tools are not sufficient for reliable endopeptidase substrate predictions. To address these shortcomings the bioinformatics tool 'PEPS' (Prediction of Endopeptidase Substrates) has been developed and is presented here. PEPS uses individual rule-based endopeptidase cleavage site scoring matrices (CSSM). The efficiency of PEPS in predicting putative caspase 3, cathepsin B and cathepsin L cleavage sites is demonstrated in comparison to established algorithms. Mortalin, a member of the heat shock protein family HSP70, was identified by PEPS as a putative cathepsin L substrate. Comparative proteome analyses of cathepsin L-deficient and wild-type mouse fibroblasts showed that mortalin is enriched in the absence of cathepsin L. These results indicate that CSSM/PEPS can correctly predict relevant peptidase substrates.

Animals↗

A rapid bioinformatic method identifies novel genes with direct clinical relevance to colon cancer.

Identifying genes whose differential expression affect the survival of patients after primary tumor surgery is a major aim of clinical cancer research. To address this issue we combined rapid bioinformatic search algorithms with quantitative RT-PCR in a panel of clearly defined cases of colorectal carcinomas with detailed patient histories. Search algorithms were written that identified Expressed Sequence Tags (ESTs) from the Unigene EST collection of putative open reading frames (ORFs). Expression ratios of healthy to cancerous tissue of each Unigene ORF were calculated. The first 35 candidates arising from bioinformatic searches were examined for mRNA expression in a panel of 20 well documented cases of colon cancer. Four of these 35 genes showed significant correlations with histopathological parameters. Therefore, their expression was further analysed by quantitative RT-PCR in a larger patient cohort. Kaplan-Meier/log rank statistical tests of up to 49 patients in three of the four genes demonstrated significant association of gene expression with poor survival. All four genes demonstrated a strong association with metastatic tumor progression. Expression of the genes was localized to epithelial cells by in-situ hybridization.

Algorithms↗

The HUPO PSI's molecular interaction format--a community standard for the representation of protein interaction data.

A major goal of proteomics is the complete description of the protein interaction network underlying cell physiology. A large number of small scale and, more recently, large-scale experiments have contributed to expanding our understanding of the nature of the interaction network. However, the necessary data integration across experiments is currently hampered by the fragmentation of publicly available protein interaction data, which exists in different formats in databases, on authors' websites or sometimes only in print publications. Here, we propose a community standard data model for the representation and exchange of protein interaction data. This data model has been jointly developed by members of the Proteomics Standards Initiative (PSI), a work group of the Human Proteome Organization (HUPO), and is supported by major protein interaction data providers, in particular the Biomolecular Interaction Network Database (BIND), Cellzome (Heidelberg, Germany), the Database of Interacting Proteins (DIP), Dana Farber Cancer Institute (Boston, MA, USA), the Human Protein Reference Database (HPRD), Hybrigenics (Paris, France), the European Bioinformatics Institute's (EMBL-EBI, Hinxton, UK) IntAct, the Molecular Interactions (MINT, Rome, Italy) database, the Protein-Protein Interaction Database (PPID, Edinburgh, UK) and the Search Tool for the Retrieval of Interacting Genes/Proteins (STRING, EMBL, Heidelberg, Germany).

Database Management Systems↗

MedKit: a helper toolkit for automatic mining of MEDLINE/PubMed citations.

UNLABELLED: MEDLINE/PubMed is one of the most important information sources for bioinformatics text mining. However, there remain limitations in working with MEDLINE/PubMed citations. For example, PubMed imposes an upper limit of 10,000 for downloading PMID list or citations; and MEDLINE files are too large for most off-the-shelf XML parsers. We developed a Java package, MedKit, to work-around the limitations, as well as provide other useful functionalities, e.g. random sampling. Its four modules (querier, sampler, fetcher and parser) can work independently, or be pipelined in various combinations. It can be used as a stand-alone GUI application, or integrated into other text-mining systems. Text mining researchers and others may download and use the toolkit free for non-commercial purposes. AVAILABILITY: http://metnetdb.gdcb.iastate.edu/medkit CONTACT: berleant@iastate.edu.

Abstracting and Indexing↗

Discovering patterns to extract protein-protein interactions from the literature: Part II.

MOTIVATION: An enormous number of protein-protein interaction relationships are buried in millions of research articles published over the years, and the number is growing. Rediscovering them automatically is a challenging bioinformatics task. Solutions to this problem also reach far beyond bioinformatics. RESULTS: We study a new approach that involves automatically discovering English expression patterns, optimizing them and using them to extract protein-protein interactions. In a sister paper, we described how to generate English expression patterns related to protein-protein interactions, and this approach alone has already achieved precision and recall rates significantly higher than those of other automatic systems. This paper continues to present our theory, focusing on how to improve the patterns. A minimum description length (MDL)-based pattern-optimization algorithm is designed to reduce and merge patterns. This has significantly increased generalization power, and hence the recall and precision rates, as confirmed by our experiments. AVAILABILITY: http://spies.cs.tsinghua.edu.cn.

Abstracting and Indexing↗

ESTAnnotator: A tool for high throughput EST annotation.

In high throughput sequence analysis, it is often necessary to combine the results of contemporary bioinformatics tools, because no individual tool alone computes all the requested information. ESTAnnotator is a tool for the high throughput annotation of expressed sequence tags (ESTs) by automatically running a collection of bioinformatics applications. In the first step, a quality check is performed and repeats, vector parts and low quality sequences are masked. Then successive steps of database searching and EST clustering are performed. Already known transcripts present within mRNA and genomic DNA reference databases are identified. Subsequently, tools for the clustering of anonymous ESTs, and for further database searches at the protein level, are applied. Finally, the outputs of each individual tool are gathered and the relevant results presented in a descriptive summary. ESTAnnotator was already successfully applied for the systematic identification and characterisation of novel human genes involved in cartilage/bone formation, growth, differentiation and homeostasis. ESTAnnotator is available at http://genome.dkfz-heidelberg.de, contact: genome@dkfz.de.

Cartilage↗

sRNAPredict: an integrative computational approach to identify sRNAs in bacterial genomes.

Small non-coding bacterial RNAs (sRNAs) play important regulatory roles in a variety of cellular processes. Nearly all known sRNAs have been identified in Escherichia coli and most of these are not conserved in the majority of other bacterial species. Many of the E.coli sRNAs were initially predicted through bioinformatic approaches based on their common features, namely that they are encoded between annotated open reading frames and are flanked by predictable transcription signals. Because promoter consensus sequences are undetermined for most species, the successful use of bioinformatics to identify sRNAs in bacteria other than E.coli has been limited. We have created a program, sRNAPredict, which uses coordinate-based algorithms to integrate the respective positions of individual predictive features of sRNAs and rapidly identify putative intergenic sRNAs. Relying only on sequence conservation and predicted Rho-independent terminators, sRNAPredict was used to search for sRNAs in Vibrio cholerae. This search identified 9 of the 10 known or putative V.cholerae sRNAs and 32 candidates for novel sRNAs. Small transcripts for 6 out of 9 candidate sRNAs were observed by Northern analysis. Our findings suggest that sRNAPredict can be used to efficiently identify novel sRNAs even in bacteria for which promoter consensus sequences are not available.

Algorithms↗

Prediction of nuclear hormone receptor response elements.

The nuclear receptor (NR) class of transcription factors controls critical regulatory events in key developmental processes, homeostasis maintenance, and medically important diseases and conditions. Identification of the members of a regulon controlled by a NR could provide an accelerated understanding of development and disease. New bioinformatics methods for the analysis of regulatory sequences are required to address the complex properties associated with known regulatory elements targeted by the receptors because the standard methods for binding site prediction fail to reflect the diverse target site configurations. We have constructed a flexible Hidden Markov Model framework capable of predicting NHR binding sites. The model allows for variable spacing and orientation of half-sites. In a genome-scale analysis enabled by the model, we show that NRs in Fugu rubripes have a significant cross-regulatory potential. The model is implemented in a web interface, freely available for academic researchers, available at http://mordor.cgb.ki.se/NHR-scan.

Algorithms↗

ALPAR: automated learning pipeline for antimicrobial resistance.

SUMMARY: The field of machine learning in antimicrobial resistance (AMR) research has experienced rapid growth, fueled by advancements in high-throughput genome sequencing and the growing capacity of computational resources. However, the complexity and lack of standardized data preparation and bioinformatic analyses present significant challenges, especially for newcomers to the domain. In response to these challenges, we introduce ALPAR (Automated Learning Pipeline for Antimicrobial Resistance), a comprehensive AMR data analysis tool covering the entire process from processing of raw genomic data to training machine learning models to interpretation of results. Our method relies on a reproducible pipeline that integrates widely used bioinformatics tools, presenting a simplified, automatic workflow specifically tailored for single-reference AMR analysis. Accepting genomic data in the form of FASTA files as input, ALPAR facilitates the generation of machine learning-ready data tables and both the training of machine learning and the execution of genome-wide association studies (GWAS) experiments. Additionally, our tool offers supplementary functionalities such as phylogeny-based analysis of the distribution of mutations, enhancing its utility for researchers. The tool has also proven its performance in competitive benchmarks, winning the 2024 CAMDA Anti-Microbial Resistance Prediction Challenge and placing third in the 2025 edition. AVAILABILITY AND IMPLEMENTATION: ALPAR is open-source and freely accessible via GitHub (https://github.com/kalininalab/ALPAR). The pipeline is fully reproducible and can be easily installed as a Conda package (https://anaconda.org/kalininalab/ALPAR).

Machine Learning↗