Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

The SuMo server: 3D search for protein functional sites.

UNLABELLED: We provide the scientific community with a web server which gives access to SuMo, a bioinformatic system for finding similarities in arbitrary 3D structures or substructures of proteins. SuMo is based on a unique representation of macromolecules using selected triplets of chemical groups having their own geometry and symmetry, regardless of the restrictive notions of main chain and lateral chains of amino acids. The heuristic for extracting similar sites was used to drive two major large-scale approaches. First, searching for ligand binding sites onto a query structure has been made possible by comparing the structure against each of the ligand binding sites found in the Protein Data Bank (PDB). Second, the reciprocal process, i.e. searching for a given 3D site of interest among the structures of the PDB is also possible and helps detect cross-reacting targets in drug design projects. AVAILABILITY: The web server is freely accessible to academia through http://sumo-pbil.ibcp.fr and full support is available from MEDIT (http://www.medit.fr). CONTACT: mjambon@burnham.org.

Amino Acid Sequence↗

An algorithm for mapping positively selected members of quasispecies-type viruses.

BACKGROUND: Many RNA viruses do not have a single, representative genome but instead form a set of related variants that has been called a quasispecies. The sequence variability of such viruses presents a significant bioinformatics challenge. In order for the sequence information to be understood, the complete mutational spectrum needs to be distilled to a biologically relevant and analyzable representation. RESULTS: Here, we develop a "selection mapping" algorithm--QUASI--that identifies the positively selected variants of viral proteins. The key to the selection mapping algorithm is the identification of particular replacement mutations that are overabundant relative to silent mutations at each codon (e.g., threonine at hemagglutinin position 262). Selection mapping identifies such replacement mutations as positively selected. Conversely, selection mapping recognizes negatively selected variants as mutational "noise" (e.g., serine at hemagglutinin position 262). CONCLUSION: Selection mapping is a fundamental improvement over earlier methods (e.g., dN/dS) that identify positive selection at codons but do not identify which amino acids at these codons confer selective advantage. Using QUASI's selection maps, we characterize the selected mutational landscapes of influenza A H3 hemagglutinin, HIV-1 reverse transcriptase, and HIV-1 gp120.

Algorithms↗

Integrating partonomic hierarchies in anatomy ontologies.

BACKGROUND: Anatomy ontologies play an increasingly important role in developing integrated bioinformatics applications. One of the primary relationships between anatomical tissues represented in such ontologies is part-of. As there are a number of ways to divide up the anatomical structure of an organism, each may be represented by more than one valid partonomic (part-of) hierarchy. This raises the issue of how to represent and integrate multiple such hierarchies. RESULTS: In this paper we describe a solution that is based on our work on an anatomy ontology for mouse embryo development, which is part of the Edinburgh Mouse Atlas Project (EMAP). The paper describes the basic conceptual aspects of our approach and discusses strengths and limitations of the proposed solution. A prototype was implemented in Prolog for evaluation purposes. CONCLUSION: With the proposed name set approach, rather than having to standardise hierarchies, it is sufficient to agree on a suitable set of basic tissue terms and their meaning in order to facilitate the integration of multiple partonomic hierarchies.

Anatomy↗

High-throughput confocal microscopy for beta-arrestin-green fluorescent protein translocation G protein-coupled receptor assays using the Evotec Opera.

Ligand-activated G protein-coupled receptors (GPCRs) are known to regulate a myriad of homeostatic functions. Inappropriate signaling is associated with several pathophysiological states. GPCRs belong to a approximately 800 member superfamily of seven transmembrane-spanning receptor proteins that respond to a diversity of ligands. As such, they present themselves as potential points of therapeutic intervention. Furthermore, orphan GPCRs, which are GPCRs without a known cognate ligand, offer new opportunities as drug development targets. This chapter describes a systems-based biological approach, one that combines in silico bioinformatics, genomics, high-throughput screening, and high-content cell-based confocal microscopy strategies to (1) identify a relevant subset of protein family targets, (2) within the therapeutic area of energy metabolism/obesity, (3) and to identify small molecule leads as tractable combinatorial and medicinal chemistry starting points. Our choice of screening platform was the Transfluor beta-arrestin-green fluorescent protein translocation assay in which full-length human orphan GPCRs were stably expressed in a U-2 OS cell background. These cells lend themselves to high-speed confocal imaging techniques using the Evotec Technologies Opera automated microscope system. The basic assay system can be implemented in any laboratory using a fluorescent probe, a stably expressed GPCR of interest, automation-assisted plate and liquid-handling techniques, an optimized image analysis algorithm, and a high-speed confocal microscope with sophisticated data analysis tools.

Algorithms↗

DIALIGN P: fast pair-wise and multiple sequence alignment using parallel processors.

BACKGROUND: Parallel computing is frequently used to speed up computationally expensive tasks in Bioinformatics. RESULTS: Herein, a parallel version of the multi-alignment program DIALIGN is introduced. We propose two ways of dividing the program into independent sub-routines that can be run on different processors: (a) pair-wise sequence alignments that are used as a first step to multiple alignment account for most of the CPU time in DIALIGN. Since alignments of different sequence pairs are completely independent of each other, they can be distributed to multiple processors without any effect on the resulting output alignments. (b) For alignments of large genomic sequences, we use a heuristics by splitting up sequences into sub-sequences based on a previously introduced anchored alignment procedure. For our test sequences, this combined approach reduces the program running time of DIALIGN by up to 97%. CONCLUSIONS: By distributing sub-routines to multiple processors, the running time of DIALIGN can be crucially improved. With these improvements, it is possible to apply the program in large-scale genomics and proteomics projects that were previously beyond its scope.

Computational Biology↗

Experimental study on the role and biomarker potential of CX3CR1 in osteoarthritis.

BACKGROUND: Osteoarthritis (OA) is a chronic joint disorder marked by progressive degeneration of articular cartilage and the formation of secondary osteophytes. Despite extensive research, the underlying molecular mechanisms remain poorly understood. This study aimed to identify OA-associated genes and elucidate the molecular pathways implicated, with the goal of discovering reliable diagnostic biomarkers. METHODS: The microarray dataset was retrieved from the Gene Expression Omnibus (GEO) and analyzed using R software to identify the signature gene, CX3CR1. Differentially expressed genes (DEGs) correlated with CX3CR1 were subsequently subjected to Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), and immune infiltration analyses. A ceRNA regulatory network was also constructed. Vali-dation of CX3CR1 expression was conducted through qRT-PCR, Western blotting, and immunohistochemistry. RESULTS: CX3CR1 emerged as a candidate gene significantly associated with OA, exhibiting regulatory roles primarily in lipid metabolism-related and extra-cellular matrix-related biological processes and signaling cascades. The infiltration levels of immune cells, particularly activated mast cells, appeared to modulate OA progression. Both in vitro and in vivo experiments demonstrated elevated CX3CR1 expression in OA tissues relative to controls, with a robust positive correlation observed between CX3CR1 and MMP13 levels. CONCLUSION: CX3CR1 represents a potential biomarker for OA diagnosis and therapeutic targeting, exerting its effects by modulating lipid metabolism, extracellular matrix dynamics, and immune cell infiltration.

CX3C Chemokine Receptor 1↗

A new progressive-iterative algorithm for multiple structure alignment.

MOTIVATION: Multiple structure alignments are becoming important tools in many aspects of structural bioinformatics. The current explosion in the number of available protein structures demands multiple structural alignment algorithms with an adequate balance of accuracy and speed, for large scale applications in structural genomics, protein structure prediction and protein classification. RESULTS: A new multiple structural alignment program, MAMMOTH-mult, is described. It is demonstrated that the alignments obtained with the new method are an improvement over previous manual or automatic alignments available in several widely used databases at all structural levels. Detailed analysis of the structural alignments for a few representative cases indicates that MAMMOTH-mult delivers biologically meaningful trees and conservation at the sequence and structural levels of functional motifs in the alignments. An important improvement over previous methods is the reduction in computational cost. Typical alignments take only a median time of 5 CPU seconds in a single R12000 processor. MAMMOTH-mult is particularly useful for large scale applications. AVAILABILITY: http://ub.cbm.uam.es/mammoth/mult.

Algorithms↗

Visualization and interpretation of protein networks in Mycobacterium tuberculosis based on hierarchical clustering of genome-wide functional linkage maps.

Genome-wide functional linkages among proteins in cellular complexes and metabolic pathways can be inferred from high throughput experimentation, such as DNA microarrays, or from bioinformatic analyses. Here we describe a method for the visualization and interpretation of genome-wide functional linkages inferred by the Rosetta Stone, Phylogenetic Profile, Operon and Conserved Gene Neighbor computational methods. This method involves the construction of a genome-wide functional linkage map, where each significant functional linkage between a pair of proteins is displayed on a two-dimensional scatter-plot, organized according to the order of genes along the chromosome. Subsequent hierarchical clustering of the map reveals clusters of genes with similar functional linkage profiles and facilitates the inference of protein function and the discovery of functionally linked gene clusters throughout the genome. We illustrate this method by applying it to the genome of the pathogenic bacterium Mycobacterium tuberculosis, assigning cellular functions to previously uncharacterized proteins involved in cell wall biosynthesis, signal transduction, chaperone activity, energy metabolism and polysaccharide biosynthesis.

Bacterial Proteins↗

BOD: a customizable bioinformatics on demand system accommodating multiple steps and parallel tasks.

The integration of bioinformatics resources worldwide is one of the major concerns of the biological community. We herein established the BOD (Bioinformatics on demand) system to use Grid computing technology to set up a virtual workbench via a web-based platform, to assist researchers performing customized comprehensive bioinformatics work. Users will be able to submit entire search queries and computation requests, e.g. from DNA assembly to gene prediction and finally protein folding, from their own office using the BOD end-user web interface. The BOD web portal parses the user's job requests into steps, each of which may contain multiple tasks in parallel. The BOD task scheduler takes an entire task, or splits it into multiple subtasks, and dispatches the task or subtasks proportionally to computation node(s) associated with the BOD portal server. A node may further split and distribute an assigned task to its sub-nodes using a similar strategy. In the end, the BOD portal server receives and collates all results and returns them to the user. BOD uses a pipeline model to describe the user's submitted data and stores the job requests/status/results in a relational database. In addition, an XML criterion is established to capture task computation program details.

Computational Biology↗

TreeDomViewer: a tool for the visualization of phylogeny and protein domain structure.

Phylogenetic analysis and examination of protein domains allow accurate genome annotation and are invaluable to study proteins and protein complex evolution. However, two sequences can be homologous without sharing statistically significant amino acid or nucleotide identity, presenting a challenging bioinformatics problem. We present TreeDomViewer, a visualization tool available as a web-based interface that combines phylogenetic tree description, multiple sequence alignment and InterProScan data of sequences and generates a phylogenetic tree projecting the corresponding protein domain information onto the multiple sequence alignment. Thereby it makes use of existing domain prediction tools such as InterProScan. TreeDomViewer adopts an evolutionary perspective on how domain structure of two or more sequences can be aligned and compared, to subsequently infer the function of an unknown homolog. This provides insight into the function assignment of, in terms of amino acid substitution, very divergent but yet closely related family members. Our tool produces an interactive scalar vector graphics image that provides orthological relationship and domain content of proteins of interest at one glance. In addition, PDF, JPEG or PNG formatted output is also provided. These features make TreeDomViewer a valuable addition to the annotation pipeline of unknown genes or gene products. TreeDomViewer is available at http://www.bioinformatics.nl/tools/treedom/.

Computer Graphics↗

A method of precise mRNA/DNA homology-based gene structure prediction.

BACKGROUND: Accurate and automatic gene finding and structural prediction is a common problem in bioinformatics, and applications need to be capable of handling non-canonical splice sites, micro-exons and partial gene structure predictions that span across several genomic clones. RESULTS: We present a mRNA/DNA homology based gene structure prediction tool, GIGOgene. We use a new affine gap penalty splice-enhanced global alignment algorithm running in linear memory for a high quality annotation of splice sites. Our tool includes a novel algorithm to assemble partial gene structure predictions using interval graphs. GIGOgene exhibited a sensitivity of 99.08% and a specificity of 99.98% on the Genie learning set, and demonstrated a higher quality of gene structural prediction when compared to Sim4, est2genome, Spidey, Galahad and BLAT, including when genes contained micro-exons and non-canonical splice sites. GIGOgene showed an acceptable loss of prediction quality when confronted with a noisy Genie learning set simulating ESTs. CONCLUSION: GIGOgene shows a higher quality of gene structure prediction for mRNA/DNA spliced alignment when compared to other available tools.

Algorithms↗

Visualization using NIPTviewer support the clinical interpretation of noninvasive prenatal testing results.

BACKGROUND: Noninvasive prenatal testing (NIPT) is increasingly used to screen for fetal chromosomal aneuploidy by analyzing cell-free DNA (cfDNA) in peripheral maternal blood. The method provides an opportunity for early detection of large genetic abnormalities without an increased risk of miscarriage due to invasive procedures. Commercial applications for use at clinical laboratories often take advantage of DNA sequencing technologies and include the bioinformatic workup of the sequence data. The interpretation of the test results and the clinical report writing, however, remains the responsibility of the diagnostic laboratory. In order to facilitate this step, we developed NIPTviewer, a web-based application to visualize and guide the interpretation of NIPT data results. RESULTS: NIPTviewer has a database functionality to store the NIPT results and a web interface for user interaction and visualization. The application has been implemented as part of a novel analysis pipeline for NIPT in a diagnostic laboratory at Uppsala University Hospital. The validation data set included 84 previously analyzed plasma samples with known results regarding chromosomes 13, 18, 21, X and Y. They were sequenced in six different experiments, uploaded to NIPTviewer and assigned to a clinical laboratory geneticist for interpretation. The results of all previously analyzed samples were replicated. CONCLUSION: NIPTviewer facilitates NIPT results interpretation and has been implemented as part of a NIPT analysis routine that was accredited by the national accreditation body for Sweden (Swedac).

Humans↗

Linear indices of the 'macromolecular graph's nucleotides adjacency matrix' as a promising approach for bioinformatics studies. Part 1: prediction of paromomycin's affinity constant with HIV-1 psi-RNA packaging region.

The design of novel anti-HIV compounds has now become a crucial area for scientists around the world. In this paper a new set of macromolecular descriptors (that are calculated from the macromolecular graph's nucleotide adjacency matrix) of relevance to nucleic acid QSAR/QSPR studies, nucleic acids' linear indices. A study of the interaction of the antibiotic Paromomycin with the packaging region of the HIV-1 psi-RNA has been performed as example of this approach. A multiple linear regression model predicted the local binding affinity constants [Log K (10(-4) M(-1))] between a specific nucleotide and the aforementioned antibiotic. The linear model explains more than 87% of the variance of the experimental Log K (R = 0.93 and s = 0.102 x 10(-4) M(-1)) and leave-one-out press statistics evidenced its predictive ability (q2 = 0.82 and s(cv) = 0.108 x 10(-4) M(-1)). The comparison with other approaches (macromolecular quadratic indices, Markovian Negentropies and 'stochastic' spectral moments) reveals a good behavior of our method.

Anti-Bacterial Agents↗

Role of protein kinases in neuropeptide gene regulation by PACAP in chromaffin cells: a pharmacological and bioinformatic analysis.

Pituitary adenylate cyclase-activating polypeptide (PACAP) is an adrenomedullary cotransmitter that along with acetylcholine is responsible for driving catecholamine and neuropeptide biosynthesis and secretion from chromaffin cells in response to stimulation of the splanchnic nerve. Two neuropeptides whose biosynthesis is regulated by PACAP include enkephalin and vasoactive intestinal polypeptide (VIP). Occupancy of PAC1 PACAP receptors on chromaffin cells can result in elevation of cyclic AMP, inositol phosphates, and intracellular calcium. The proenkephalin A and VIP genes are transcriptionally responsive to signals generated within all three pathways, and potentially by combinatorial activation of these pathways as well. The characteristics of PACAP regulation of enkephalin and VIP biosynthesis were examined pharmacologically for evidence of involvement of several serine/threonine protein kinases activated by cAMP, IP3, and/or calcium, including calmodulin kinase II, protein kinase A, and protein kinase C. Evidence is presented for the differential involvement of these protein kinases in regulation of enkephalin and VIP biosynthesis in chromaffin cells, and for a prominent role of the mixed-function (tyrosine and serine/threonine) MAP kinase family in mediating transcriptional activation of neuropeptide genes by PACAP.

Animals↗

In silico analysis of Burkholderia pseudomallei genome sequence for potential drug targets.

Recent advances in DNA sequencing technology have enabled elucidation of whole genome information from a plethora of organisms. In parallel with this technology, various bioinformatics tools have driven the comparative analysis of the genome sequences between species and within isolates. While drawing meaningful conclusions from a large amount of raw material, computer-aided identification of suitable targets for further experimental analysis and characterization, has also led to the prediction of non-human homologous essential genes in bacteria as promising candidates for novel drug discovery. Here, we present a comparative genomic analysis to identify essential genes in Burkholderia pseudomallei. Our in silico prediction has identified 312 essential genes which could also be potential drug candidates. These genes encode essential proteins to support the survival of B. pseudomallei including outer-inner membrane and surface structures, regulators, proteins involved in pathogenenicity, adaptation, chaperones as well as degradation of small and macromolecules, energy metabolism, information transfer, central/intermediate/miscellaneous metabolism pathways and some conserved hypothetical proteins of unknown function. Therefore, our in silico approach has enabled rapid screening and identification of potential drug targets for further characterization in the laboratory.

Burkholderia pseudomallei↗

The role of pattern databases in sequence analysis.

In the wake of the numerous now-fruitful genome projects, we are entering an era rich in biological data. The field of bioinformatics is poised to exploit this information in increasingly powerful ways, but the abundance and growing complexity both of the data and of the tools and resources required to analyse them are threatening to overwhelm us. Databases and their search tools are now an essential part of the research environment. However, the rate of sequence generation and the haphazard proliferation of databases have made it difficult to keep pace with developments. In an age of information overload, researchers want rapid, easy-to-use, reliable tools for functional characterisation of newly determined sequences. But what are those tools? How do we access them? Which should we use? This review focuses on a particular type of database that is increasingly used in the task of routine sequence analysis--the so-called pattern database. The paper aims to provide an overview of the current status of pattern databases in common use, outlining the methods behind them and giving pointers on their diagnostic strengths and weaknesses.

Amino Acid Motifs↗

Predicting coarse-grained representations of biogeochemical cycles from metabarcoding data.

MOTIVATION: Taxonomic analysis of environmental microbial communities is now routinely performed thanks to advances in DNA sequencing. Determining the role of these communities in global biogeochemical cycles requires the identification of their metabolic functions, such as hydrogen oxidation, sulfur reduction, and carbon fixation. These functions can be directly inferred from metagenomics data, but in many environmental applications metabarcoding is still the method of choice. The reconstruction of metabolic functions from metabarcoding data and their integration into coarse-grained representations of biogeochemical cycles remains a difficult bioinformatics problem today. RESULTS: We developed a pipeline, called Tabigecy, which exploits taxonomic affiliations to predict metabolic functions constituting biogeochemical cycles. In a first step, Tabigecy uses the tool EsMeCaTa to predict consensus proteomes from input affiliations. To optimize this process, we generated a precomputed database containing information about 2404 taxa from UniProt. The consensus proteomes are searched using bigecyhmm, a newly developed Python package relying on Hidden Markov Models to identify key enzymes involved in metabolic function of biogeochemical cycles. The metabolic functions are then projected on coarse-grained representation of the cycles. We applied Tabigecy to two salt cavern datasets and validated its predictions with microbial activity and hydrochemistry measurements performed on the samples. The results highlight the utility of the approach to investigate the impact of microbial communities on biogeochemical processes. AVAILABILITY AND IMPLEMENTATION: The Tabigecy pipeline is available at https://github.com/ArnaudBelcour/tabigecy. The Python package bigecyhmm and the precomputed EsMeCaTa database are also separately available at https://github.com/ArnaudBelcour/bigecyhmm and https://doi.org/10.5281/zenodo.13354073, respectively.

Metagenomics↗

Tracing regulatory element networks using epigenetic traits to identify key transcription factors: TENET R/Bioconductor package.

SUMMARY: There is a lack of publicly available bioinformatic tools that can be widely used by researchers to identify transcription factors (TFs) that regulate cell type-specific regulatory elements (REs). To address this, we developed the Tracing regulatory Element Networks using Epigenetic Traits (TENET) R/Bioconductor package. By collecting hundreds of histone mark and open chromatin datasets from a variety of cell lines, primary cells, and tissues, and comparing these features along with matched DNA methylation and gene expression data, TENET identifies TFs and REs linked to a specific cell type. Moreover, we developed methods to interrogate findings using motifs, clinical information, and other genomic and chromatin conformation capture datasets, and applied them to pan-cancer data, highlighting TFs and REs associated with ten different cancer types. TENET enables researchers to better characterize the 3D epigenomes of cell types of interest for future clinical applications. AVAILABILITY AND IMPLEMENTATION: TENET is available at http://bioconductor.org/packages/TENET. Curated functional genomic datasets utilized by TENET are available at http://bioconductor.org/packages/TENET.AnnotationHub. Example datasets are available at http://bioconductor.org/packages/TENET.ExperimentHub.

Transcription Factors↗