Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

An XML standard for the dissemination of annotated 2D gel electrophoresis data complemented with mass spectrometry results.

BACKGROUND: Many proteomics initiatives require a seamless bioinformatics integration of a range of analytical steps between sample collection and systems modeling immediately assessable to the participants involved in the process. Proteomics profiling by 2D gel electrophoresis to the putative identification of differentially expressed proteins by comparison of mass spectrometry results with reference databases, includes many components of sample processing, not just analysis and interpretation, are regularly revisited and updated. In order for such updates and dissemination of data, a suitable data structure is needed. However, there are no such data structures currently available for the storing of data for multiple gels generated through a single proteomic experiments in a single XML file. This paper proposes a data structure based on XML standards to fill the void that exists between data generated by proteomics experiments and storing of data. RESULTS: In order to address the resulting procedural fluidity we have adopted and implemented a data model centered on the concept of annotated gel (AG) as the format for delivery and management of 2D Gel electrophoresis results. An eXtensible Markup Language (XML) schema is proposed to manage, analyze and disseminate annotated 2D Gel electrophoresis results. The structure of AG objects is formally represented using XML, resulting in the definition of the AGML syntax presented here. CONCLUSION: The proposed schema accommodates data on the electrophoresis results as well as the mass-spectrometry analysis of selected gel spots. A web-based software library is being developed to handle data storage, analysis and graphic representation. Computational tools described will be made available at http://bioinformatics.musc.edu/agml. Our development of AGML provides a simple data structure for storing 2D gel electrophoresis data.

Computational Biology↗

Two-stage multi-class support vector machines to protein secondary structure prediction.

Bioinformatics techniques to protein secondary structure (PSS) prediction are mostly single-stage approaches in the sense that they predict secondary structures of proteins by taking into account only the contextual information in amino acid sequences. In this paper, we propose two-stage Multi-class Support Vector Machine (MSVM) approach where a MSVM predictor is introduced to the output of the first stage MSVM to capture the sequential relationship among secondary structure elements for the prediction. By using position specific scoring matrices, generated by PSI-BLAST, the two-stage MSVM approach achieves Q3 accuracies of 78.0% and 76.3% on the RS126 dataset of 126 nonhomologous globular proteins and the CB396 dataset of 396 nonhomologous proteins, respectively, which are better than the highest scores published on both datasets to date.

Computational Biology↗

Data mining of sequences and 3D structures of allergenic proteins.

MOTIVATION: Many sequences, and in some cases structures, of proteins that induce an allergic response in atopic individuals have been determined in recent years. This data indicates that allergens, regardless of source, fall into discreet protein families. Similarities in the sequence may explain clinically observed cross-reactivities between different biological triggers. However, previously available allergy databases group allergens according to their biological sources, or observed clinical cross-reactivities, without providing data about the proteins. A computer-aided data mining system is needed to compare the sequential and structural details of known allergens. This information will aid in predicting allergenic cross-responses and eventually in determining possible common characteristics of IgE recognition. RESULTS: The new web-based Structural Database of Allergenic Proteins (SDAP) permits the user to quickly compare the sequence and structure of allergenic proteins. Data from literature sources and previously existing lists of allergens are combined in a MySQL interactive database with a wide selection of bioinformatics applications. SDAP can be used to rapidly determine the relationship between allergens and to screen novel proteins for the presence of IgE or T-cell epitopes they may share with known allergens. Further, our novel similarity search method, based on five dimensional descriptors of amino acid properties, can be used to scan the SDAP entries with a peptide sequence. For example, when a known IgE binding epitope from shrimp tropomyosin was used as a query, the method rapidly identified a similar sequence in known shellfish and insect allergens. This prediction of cross-reactivity between allergens is consistent with clinical observations. AVAILABILITY: SDAP is available on the web at http://fermi.utmb.edu/SDAP/index.html

Allergens↗

Integrative bioinformatics for functional genome annotation: trawling for G protein-coupled receptors.

G protein-coupled receptors (GPCR) are amongst the best studied and most functionally diverse types of cell-surface protein. The importance of GPCRs as mediates or cell function and organismal developmental underlies their involvement in key physiological roles and their prominence as targets for pharmacological therapeutics. In this review, we highlight the requirement for integrated protocols which underline the different perspectives offered by different sequence analysis methods. BLAST and FastA offer broad brush strokes. Motif-based search methods add the fine detail. Structural modelling offers another perspective which allows us to elucidate the physicochemical properties that underlie ligand binding. Together, these different views provide a more informative and a more detailed picture of GPCR structure and function. Many GPCRs remain orphan receptors with no identified ligand, yet as computer-driven functional genomics starts to elaborate their functions, a new understanding of their roles in cell and developmental biology will follow.

Algorithms↗

Do we want our data raw? Including binary mass spectrometry data in public proteomics data repositories.

With the human Plasma Proteome Project (PPP) pilot phase completed, the largest and most ambitious proteomics experiment to date has reached its first milestone. The correspondingly impressive amount of data that came from this pilot project emphasized the need for a centralized dissemination mechanism and led to the development of a detailed, PPP specific data gathering infrastructure at the University of Michigan, Ann Arbor as well as the protein identifications database project at the European Bioinformatics Institute as a general proteomics data repository. One issue that crept up while discussing which data to store for the PPP concerns whether the raw, binary data coming from the mass spectrometers should be stored, or rather the more compact and already significantly processed peak lists. As this debate is not restricted to the PPP but relates to the proteomics community in general, we will attempt to detail the relative merits and caveats associated with centralized storage and dissemination of raw data and/or peak lists, building on the extensive experience gained during the PPP pilot phase. Finally, some suggestions are made for both immediate and future storage of MS data in public repositories.

Computational Biology↗

FuNTB: a functional network clustering tool for the analysis of genome-wide genetic variants in Mycobacterium tuberculosis.

MOTIVATION: Tuberculosis (TB), caused by Mycobacterium tuberculosis (Mtb), still claims around 1.25 million lives each year. The growing threat of drug resistance-often driven by single‑nucleotide polymorphisms (SNPs) in Mtb genomes underscores the need for high‑quality genomic data and powerful bioinformatics tools. We present FuNTB, a python‑based pipeline that detects non‑synonymous SNPs in Mtb and builds functional network clusters to reveal genotype-phenotype relationships. RESULTS: FuNTB profiles non‑synonymous SNPs at the gene level across user‑defined phenotypes, pinpointing both shared and unique mutations. It ingests annotated Variant Call Format (VCF) files or MTBseq outputs and merges them with clinical metadata to produce network‑XML files compatible with Cytoscape and Gephi. When applied to the CRyPTIC Mtb collection, FuNTB rapidly recovered established resistance genes and surfaced novel candidates, validating its utility for mapping genotype-phenotype associations. AVAILABILITY AND IMPLEMENTATION: FuNTB is implemented in Python 3.8+ and is freely available under the MIT license at https://doi.org/10.5281/zenodo.15399917.

Mycobacterium tuberculosis↗

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software↗

SNPeffect v2.0: a new step in investigating the molecular phenotypic effects of human non-synonymous SNPs.

UNLABELLED: Single nucleotide polymorphisms (SNPs) constitute the most fundamental type of genetic variation in human populations. About 75 000 of these reported variations cause an amino acid change in the translated protein. An important goal in genomic research is to understand how this variability affects protein function, and whether or not particular SNPs are associated to disease susceptibility. Accordingly, the SNPeffect database uses sequence- and structure-based bioinformatics tools to predict the effect of non-synonymous SNPs on the molecular phenotype of proteins. SNPeffect analyses the effect of SNPs on three categories of functional properties: (1) structural and thermodynamic properties affecting protein dynamics and stability (2) the integrity of functional and binding sites and (3) changes in posttranslational processing and cellular localization of proteins. The search interface of the database can be used to search specifically for polymorphisms that are predicted to cause a change in one of these properties. Now based on the Ensembl human databases, the SNPeffect database has been remodeled to better fit an automatically updatable structure. The current edition holds the molecular phenotype of 74 567 nsSNPs in 23 426 proteins. AVAILABILITY: SNPeffect can be accessed through http://snpeffect.vib.be.

Algorithms↗

Biopipe: a flexible framework for protocol-based bioinformatics analysis.

We identify several challenges facing bioinformatics analysis today. Firstly, to fulfill the promise of comparative studies, bioinformatics analysis will need to accommodate different sources of data residing in a federation of databases that, in turn, come in different formats and modes of accessibility. Secondly, the tsunami of data to be handled will require robust systems that enable bioinformatics analysis to be carried out in a parallel fashion. Thirdly, the ever-evolving state of bioinformatics presents new algorithms and paradigms in conducting analysis. This means that any bioinformatics framework must be flexible and generic enough to accommodate such changes. In addition, we identify the need for introducing an explicit protocol-based approach to bioinformatics analysis that will lend rigorousness to the analysis. This makes it easier for experimentation and replication of results by external parties. Biopipe is designed in an effort to meet these goals. It aims to allow researchers to focus on protocol design. At the same time, it is designed to work over a compute farm and thus provides high-throughput performance. A common exchange format that encapsulates the entire protocol in terms of the analysis modules, parameters, and data versions has been developed to provide a powerful way in which to distribute and reproduce results. This will enable researchers to discuss and interpret the data better as the once implicit assumptions are now explicitly defined within the Biopipe framework.

Amino Acid Sequence↗

Prelude and Fugue, predicting local protein structure, early folding regions and structural weaknesses.

UNLABELLED: Prelude&Fugue are bioinformatics tools aiming at predicting the local 3D structure of a protein from its amino acid sequence in terms of seven backbone torsion angle domains, using database-derived potentials. Prelude(&Fugue) computes all lowest free energy conformations of a protein or protein region, ranked by increasing energy, and possibly satisfying some interresidue distance constraints specified by the user. (Prelude&)Fugue detects sequence regions whose predicted structure is significantly preferred relative to other conformations in the absence of tertiary interactions. These programs can be used for predicting secondary structure, tertiary structure of short peptides, flickering early folding sequences and peptides that adopt a preferred conformation in solution. They can also be used for detecting structural weaknesses, i.e. sequence regions that are not optimal with respect to the tertiary fold. AVAILABILITY: http://babylone.ulb.ac.be/Prelude_and_Fugue.

Algorithms↗

Phasis: a software tool for register-resolved discovery of plant phased small RNA loci.

Plant PHAS locus discovery remains challenging because phasiRNA-producing loci must be distinguished from other sRNA-producing regions with high abundance or apparent periodicity. This problem is especially acute for reproductive 24-PHAS loci, which occur within genomes that also produce abundant 24-nt siRNAs from nonPHAS regions. We present Phasis, an open-source Python software tool for plant PHAS-locus discovery from small RNA sequencing data. Phasis combines statistical evidence for phased accumulation with locus-level features and a Register-Resolved Locus Interpretation Layer that evaluates whether candidate loci show coherent phased architecture. Across diverse plant datasets, Phasis recovered validated or annotated 21- and 24-PHAS loci with a strong balance between call-level precision and reference-locus recall, and generally outperformed PhaseTank and ShortStack in matched benchmark analyses. The register-resolved interpretation layer reduced unsupported calls by separating coherent phased loci from ambiguous sRNA-producing regions. In maize dcl5 mutant libraries, Phasis showed strong depletion of 24-PHAS recovery, supporting DCL5-dependent recovery of reproductive 24-PHAS signal. Together, these results support Phasis as a biologically interpretable tool for large-scale discovery of plant DCL-dependent phasiRNA loci.

bioinformatics↗

CpGene: a web application for epigenetic signature identification from DNA methylation arrays.

MOTIVATION: DNA methylation (DNAme) is the best studied epigenetic mechanism that plays pivotal role in tissue differentiation and epigenetic disruption has been correlated to diverse disease types (e.g. cancer, metabolic disorders). While various DNAme array platforms have been discovered, data analysis remains a challenging task which often requires in-depth bioinformatic expertise. Here, we developed a user-friendly web-based application for data analysis and visualization that accommodates users ranging from early-career basic/translational researchers to experienced bioinformaticians. RESULTS: CpGene is a web application for analyzing DNA methylation array data. It supports Illumina 450K, EPIC, and EPICv2 methylation array platforms and processes .idat files with integrated preprocessing, normalization, and quality control. Biomarker discovery is available through either classic differential methylation point analysis or machine learning-based feature selection as well as gene enrichment analysis. Results are summarized with clear visualizations, to aid interpretation. By combining these functions in a unified interface, CpGene streamlines methylation analysis and helps identify CpG sites and genes with biological and clinical relevance. AVAILABILITY AND IMPLEMENTATION: CpGene is openly accessible as a web service through http://cpgene.duckdns.org:8001/ and it's source code is available on https://github.com/kostaslazaros/cpgenene.

DNA Methylation↗

Gene networks as a tool to understand transcriptional regulation.

Gene regulatory networks, or simply gene networks (GNs), have shown to be a promising approach that the bioinformatics community has been developing for studying regulatory mechanisms in biological systems. GNs are built from the genome-wide high-throughput gene expression data that are often available from DNA microarray experiments. Conceptually, GNs are (un)directed graphs, where the nodes correspond to the genes and a link between a pair of genes denotes a regulatory interaction that occurs at transcriptional level. In the present study, we had two objectives: 1) to develop a framework for GN reconstruction based on a Bayesian network model that captures direct interactions between genes through nonparametric regression with B-splines, and 2) to demonstrate the potential of GNs in the analysis of expression data of a real biological system, the yeast pheromone response pathway. Our framework also included a number of search schemes to learn the network. We present an intuitive notion of GN theory as well as the detailed mathematical foundations of the model. A comprehensive analysis of the consistency of the model when tested with biological data was done through the analysis of the GNs inferred for the yeast pheromone pathway. Our results agree fairly well with what was expected based on the literature, and we developed some hypotheses about this system. Using this analysis, we intended to provide a guide on how GNs can be effectively used to study transcriptional regulation. We also discussed the limitations of GNs and the future direction of network analysis for genomic data. The software is available upon request.

Bayes Theorem↗

Identification and integrative analysis of 28 novel genes specifically expressed and developmentally regulated in murine spermatogenic cells.

Mammalian spermatogenesis is a highly ordered process that occurs in mitotic, meiotic, and postmeiotic phases. The unique mechanisms responsible for this tightly regulated developmental process suggest the presence of an intrinsic genetic program composed of spermatogenic cell-specific genes. In this study, we analyzed the mouse round spermatid UniGene library currently containing 2124 gene-oriented transcript clusters, predicting that 467 of them are testis-specific genes, and systematically identified 28 novel genes with evident testis-specific expression by in silico and in vitro approaches. We analyzed these genes by Northern blot hybridization and cDNA cloning, demonstrating the presence of additional transcript sequences in five genes and multiple transcript isoforms in six genes. Genomic analysis revealed lack of human orthologues for 10 genes, implying a relationship between these genes and male reproduction unique to mouse. We found that all of the novel genes are expressed in developmentally regulated and stage-specific patterns, suggesting that they are primary regulators of male germ cell development. Using computational bioinformatics tools, we found that 20 gene products are potentially involved in various processes during spermatogenesis or fertilization. Taken together, we predict that over 20% of the genes from the round spermatid library are testis-specific, have discovered the 28 authentic, novel genes with probable spermatogenic cell-specific expression by the integrative approach, and provide new and thorough information about the novel genes by various in vitro and in silico analyses. Thus, the study establishes on a comprehensive scale a new basis for studies to uncover molecular mechanisms underlying the reproductive process.

Animals↗

The National Microbial Pathogen Database Resource (NMPDR): a genomics platform based on subsystem annotation.

The National Microbial Pathogen Data Resource (NMPDR) (http://www.nmpdr.org) is a National Institute of Allergy and Infections Disease (NIAID)-funded Bioinformatics Resource Center that supports research in selected Category B pathogens. NMPDR contains the complete genomes of approximately 50 strains of pathogenic bacteria that are the focus of our curators, as well as >400 other genomes that provide a broad context for comparative analysis across the three phylogenetic Domains. NMPDR integrates complete, public genomes with expertly curated biological subsystems to provide the most consistent genome annotations. Subsystems are sets of functional roles related by a biologically meaningful organizing principle, which are built over large collections of genomes; they provide researchers with consistent functional assignments in a biologically structured context. Investigators can browse subsystems and reactions to develop accurate reconstructions of the metabolic networks of any sequenced organism. NMPDR provides a comprehensive bioinformatics platform, with tools and viewers for genome analysis. Results of precomputed gene clustering analyses can be retrieved in tabular or graphic format with one-click tools. NMPDR tools include Signature Genes, which finds the set of genes in common or that differentiates two groups of organisms. Essentiality data collated from genome-wide studies have been curated. Drug target identification and high-throughput, in silico, compound screening are in development.

Bacteria↗

A task framework for the web interface W2H.

SUMMARY: The W3H task framework allows the execution of compound jobs utilizing the description of work and data flows in a heterogeneous bioinformatics environment using meta-data information. By means of these descriptions, the task system can schedule the necessary execution of applications available in the environment, depending on rules specified in the meta-data. By integrating this task framework into the publicly available web interface W2H, similarly based on meta-data, web access and data management are immediately available for each task description. Authors of task descriptions can base their work on the underlying classes and objects to be able to describe dependency rules between previously independent applications. The result of a compound task is given as XML data that is translated according to XSLT data into web pages or plain text to report the result of the task to the user. AVAILABILITY: Within the HUSAR environment at DKFZ http://genome.dkfz-heidelberg.de/

Database Management Systems↗

Versatile and declarative dynamic programming using pair algebras.

BACKGROUND: Dynamic programming is a widely used programming technique in bioinformatics. In sharp contrast to the simplicity of textbook examples, implementing a dynamic programming algorithm for a novel and non-trivial application is a tedious and error prone task. The algebraic dynamic programming approach seeks to alleviate this situation by clearly separating the dynamic programming recurrences and scoring schemes. RESULTS: Based on this programming style, we introduce a generic product operation of scoring schemes. This leads to a remarkable variety of applications, allowing us to achieve optimizations under multiple objective functions, alternative solutions and backtracing, holistic search space analysis, ambiguity checking, and more, without additional programming effort. We demonstrate the method on several applications for RNA secondary structure prediction. CONCLUSION: The product operation as introduced here adds a significant amount of flexibility to dynamic programming. It provides a versatile testbed for the development of new algorithmic ideas, which can immediately be put to practice.

Algorithms↗

PheGe, the platform for exploring genotype-phenotype relations on cellular and organism level.

One major challenge of bioinformatics is to extract biological information into a form that gives access to analyses and predictive models and that sheds new light on cellular and organism function. In order to approach automated network analysis on organism level the relational platform PheGe was generated. PheGe enables a) presentation of cell-specific regulatory and metabolic pathways, b) sorting and coordination of the various molecules, genes and reactions to their particular signaling systems, c) visualization of signaling par distance, d) organization of downstream events on a multi-cellular level, e) recording and evaluation of pathological relevant data, f) coordination of the aberrant genes and gene products into the various regulatory pathways balancing phenotypic patterns g) modeling of cellular differentiation and finally h) tracing of network components that balance differentiation programs.

Algorithms↗