Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

A distributed environment for physical map construction.

MOTIVATION: With the main focus of the Human Genome Project shifting to sequencing, bioinformatics support for constructing large-scale genomic maps of other organisms is still required. We attempt to provide for this with our work, aimed at the delivery of robust and user-friendly contig-building software on the WWW. RESULTS: We present a prototype distributed analytical environment for molecular biologists working in the area of genomic mapping. It consists of the WWW server for constructing contigs from users' data with a hypertext output connected to Java-based map visualization software. AVAILABILITY: Freely available on http://www.mpimg-berlin-dahlem.mpg. de/ approximately andy/server/ CONTACT: andy@rag3.rz-berlin.mpg.de

Algorithms↗

Graphically-enabled integration of bioinformatics tools allowing parallel execution.

Rapid analysis of large amounts of genomic data is of great biological as well as medical interest. This type of analysis will greatly benefit from the ability to rapidly assemble a set of related analysis programs and to exploit the power of parallel computing. TurboGenomics, which is a software package currently in its alpha-testing phase, allows integration of heterogeneous software components to be done graphically. In addition, the tool is capable of making the integrated components run in parallel. To demonstrate these abilities, we use the tool to develop a Web-based application that allows integrated access to a set of large-scale sequence data analysis programs used by a transposon-insertion based yeast genome project. We also contrast the differences in building such an application with and without using the TurboGenomics software.

Computational Biology↗

Ubiquitous distributed objects with CORBA.

Database interoperation is becoming a bottleneck for the research community in biology. In this paper, we first discuss the question of interoperability and give a brief overview of CORBA. Then, an example is explained in some detail: a simple but realistic data bank of STSs is implemented. The Object Request Broker is the media for communication between an object server (the data bank) and a client (possibly a genome center). Since CORBA enables easy development of networked applications, we meant this paper to provide an incentive for the bioinformatics community to develop distributed objects.

Base Sequence↗

Expression and prognosis of CXCL13 in uterine corpus endometrial carcinoma based on bioinformatics analysis.

OBJECTIVE: The biological significance of the chemokine ligand C-X-C motif chemokine ligand 13 (CXCL13) may play a significant role in the pathogenesis of uterine corpus endometrial carcinoma (UCEC). This study aims to identify and verify CXCL13 with predictive value for prognosis in UCEC. METHODS: CXCL13 mRNA expression differences were analyzed using R software in three independent datasets: one each from The Cancer Genome Atlas (TCGA) and two from the Gene Expression Omnibus (GEO), namely GSE17025 and GSE106191. The correlation between CXCL13 expression and prognosis was evaluated by Kaplan-Meier analysis. Univariate and multivariate Cox analyses were utilized to construct a prognostic nomogram. Tumor Immune Estimation Resource (TIMER) and the Tumor and Immune System Interaction Database (TISIDB) were employed to assess the relationship between CXCL13 and tumor immune infiltration. Coexpressed genes with CXCL13 were identified by the Spearman correlation analysis. A CXCL13 protein-protein interaction (PPI) network was constructed with the STRING website tool and hub genes were screened out. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genome (KEGG) analyses were performed with the "clusterProfiler" R package. Gene set enrichment analysis (GSEA) was used to identify underlying biological mechanisms. A drug-gene interaction network was constructed in the Comparative Toxicogenomics Database (CTD). RESULTS: High CXCL13 mRNA expression were validated in UCEC in the above three independent datasets. High CXCL13 expression was associated with favorable prognosis in UCEC. A nomogram for predicting the 1-, 3-, and 5-year survival probability in UCEC was construct based on CXCL13 expression and other clinical parameters. The use of Spearman correlation indicated certain correlation between CXCL13 and immune cells and immune checkpoint (ICP) genes. Seven hub genes were upregulated in UCEC, namely CXCL9, IFNG, CXCL10, CXCL11, GBP5, CCL18, and GZMB. The expression and prognostic relevance of CXCL9, IFNG, GBP5, and GZMB were in accordance with CXCL13. The main biological processes enriched were cytokine-cytokine receptor interaction and chemokine signaling pathway. CONCLUSIONS: The above comprehensive analyses suggest that CXCL13 may serve as a potential prognostic biomarker for UCEC, specifically for early-stage UCEC.

CXCL13↗

Bayesian inference on biopolymer models.

MOTIVATION: Most existing bioinformatics methods are limited to making point estimates of one variable, e.g. the optimal alignment, with fixed input values for all other variables, e.g. gap penalties and scoring matrices. While the requirement to specify parameters remains one of the more vexing issues in bioinformatics, it is a reflection of a larger issue: the need to broaden the view on statistical inference in bioinformatics. RESULTS: The assignment of probabilities for all possible values of all unknown variables in a problem in the form of a posterior distribution is the goal of Bayesian inference. Here we show how this goal can be achieved for most bioinformatics methods that use dynamic programming. Specifically, a tutorial style description of a Bayesian inference procedure for segmentation of a sequence based on the heterogeneity in its composition is given. In addition, full Bayesian inference algorithms for sequence alignment are described. AVAILABILITY: Software and a set of transparencies for a tutorial describing these ideas are available at http://www.wadsworth.org/res&res/bioinfo/

Bayes Theorem↗

Protein structure prediction methods for drug design.

Along the long path from genomic data to a new drug, the knowledge of three-dimensional protein structure can be of significant help in several places. This paper points out such places, discusses the virtues of protein structure knowledge and reviews bioinformatics methods for gaining such knowledge on the protein structure.

Computational Biology↗

CoMR: an integrative scoring pipeline for comprehensive mitochondrial proteome reconstruction across eukaryotes.

Mitochondrial proteome reconstruction from eukaryotic sequence data typically relies on prediction of mitochondrial targeting signals (MTSs). However, MTS predictors are primarily trained on model organisms and may perform poorly in phylogenetically divergent lineages or in organisms with atypical or reduced targeting sequences. Accurate reconstruction therefore requires integration of complementary sources of evidence beyond targeting prediction alone. We developed Comprehensive Mitochondrial Reconstructor (CoMR), an integrative workflow that combines targeting prediction, curated homology searches, large-scale similarity searches, and automated phylogenetic analysis within a unified scoring framework. Benchmarking on the model yeast Saccharomyces cerevisiae yielded strong discriminatory performance [receiver operating characteristic (ROC)-area under the curve (AUC) = 0.92], exceeding standalone prediction with TargetP2, a predictor of N-terminal targeting peptides (ROC-AUC = 0.72). In the divergent anaerobic protist Paratrimastix pyriformis, CoMR maintained robust performance (ROC-AUC = 0.86) validated with an experimental proteome despite extreme class imbalance, achieving a precision-recall AUC of 0.183 (~78-fold enrichment over random expectation and ~10-fold improvement over TargetP2). Ablation analyses demonstrate that predictive performance is robust to individual evidence-layer removal, while overlap analyses showed that homology-based searches recovered candidates missed by targeting predictors, particularly in P. pyriformis. Overall, CoMR improves mitochondrial proteome reconstruction over targeting prediction alone and provides a reproducible workflow for predicting mitochondrial and mitochondrion-related organelle protein repertoires across eukaryotes to aid investigations of organelle evolution and proteome reduction.

Proteome↗

AutoPVPrimer: A comprehensive AI-Enhanced pipeline for efficient plant virus primer design and assessment.

Plant viruses pose a significant threat to global agriculture and require efficient tools for their timely detection. We present AutoPVPrimer, an innovative pipeline that integrates artificial intelligence (AI) and machine learning to accelerate the development of plant virus primers. The pipeline uses Biopython to automatically retrieve different genomic sequences from the NCBI database to increase the robustness of the subsequent primer design. The design_primers_with_tuning module uses a random forest classifier that optimizes parameters and provides flexibility for different experimental conditions. Quality control measures, including the evaluation of poly-X content and melting temperature, increase primer reliability. Unique to AutoPVPrimer is the visualize_primer_dimer module, which supports the visual evaluation of primer dimers-a feature missing in other tools. Primer specificity is validated via primer BLAST, which contributes to the overall efficiency of the pipeline. AutoPVPrimer has been successfully applied to the tomato mosaic virus, proving its adaptability and efficiency. The modular design allows customization by the user and extends the applicability to different plant viruses and experimental scenarios. The pipeline represents a significant advance in primer design and provides researchers with an effective tool to accelerate molecular biology experiments. Future developments aim to extend compatibility and incorporate user feedback to consolidate AutoPVPrimer as an innovative contribution to the bioinformatics toolbox and a promising resource for the advancement of plant virology research.

DNA Primers↗

Bioinformatics -- a patenting view.

The use of bioinformatics in the biological sciences has brought about a change in the way that biological inventions can be protected by patent laws. Using approaches developed in the fields of computer science and business, patent applicants now seek to protect certain aspects of their inventions, which include software, methods of doing business and uses of information as well as more traditional biotechnological products and processes. These approaches are useful in resolving some of the difficulties now faced in prosecuting patent applications directed to biological inventions that are claimed in more conventional terms.

Algorithms↗

Rose: generating sequence families.

MOTIVATION: We present a new probabilistic model of the evolution of RNA-, DNA-, or protein-like sequences and a software tool, Rose, that implements this model. Guided by an evolutionary tree, a family of related sequences is created from a common ancestor sequence by insertion, deletion and substitution of characters. During this artificial evolutionary process, the 'true' history is logged and the 'correct' multiple sequence alignment is created simultaneously. The model also allows for varying rates of mutation within the sequences, making it possible to establish so-called sequence motifs. RESULTS: The data created by Rose are suitable for the evaluation of methods in multiple sequence alignment computation and the prediction of phylogenetic relationships. It can also be useful when teaching courses in or developing models of sequence evolution and in the study of evolutionary processes. AVAILABILITY: Rose is available on the Bielefeld Bioinformatics WebServer under the following URL: http://bibiserv.TechFak.Uni-Bielefeld.DE/rose/ The source code is available upon request. CONTACT: folker@TechFak.Uni-Bielefeld.DE

Algorithms↗

Structural principles governing domain motions in proteins.

With the use of a recently developed method, twenty-four proteins for which two or more X-ray conformers are known have been analyzed to reveal structural principles that govern domain motions in proteins. In all 24 cases, the domain motion is a rotation about a physical axis created through local interactions both covalent and noncovalent. In many cases, two or more mechanical hinges separated in space create a stable hinge axis for precise control of the domain closure. The terminal regions of alpha-helices and beta-sheets have been found to act as mechanical hinges in a significant number of cases. In some cases, the two terminal regions of neighboring strands of a single beta-sheet can create a hinge axis, as can the two termini of a single alpha-helix. These two structures have been termed the "double-hinged beta-sheet" and "double-hinged alpha-helix," respectively. A flexible loop that attaches one domain to another and through which the effective hinge axis passes is another construct that is used to create a hinge. Noncovalent interactions between segments remote along the polypeptide chain can also form hinges. In addition alpha-helices that preserve their hydrogen bonding structure when bent have been found to behave as mechanical hinges. It is suggested that these alpha-helices act as a store of elastic energy that drives the closing of domains for rapid capture of the substrate. If the repertoire of possible interdomain structures is as limited as this study suggests, the dynamic behavior of proteins could soon be predicted using bioinformatics techniques. Proteins 1999;36:425-435.

Animals↗

The European Bioinformatics Institute (EBI) databases.

The European Bioinformatics Institute (EBI) maintains and distributes the EMBL Nucleotide Sequence database, Europe's primary nucleotide sequence data resource. The EBI also maintains and distributes the SWISS-PROT Protein Sequence database, in collaboration with Amos Bairoch of the University of Geneva. Over fifty additional specialist molecular biology databases, as well as software and documentation of interest to molecular biologists are available. The EBI network services include database searching and sequence similarity searching facilities.

Amino Acid Sequence↗

Whole-genome automated assembly pipeline for Chlamydia trachomatis strains from reference, in vitro and clinical samples using the integrated CtGAP pipeline.

Whole genome sequencing (WGS) is pivotal for the molecular characterization of Chlamydia trachomatis (Ct)-the leading bacterial cause of sexually transmitted infections and infectious blindness worldwide. Ct WGS can inform epidemiologic, public health and outbreak investigations of these human-restricted pathogens. However, challenges persist in generating high-quality genomes for downstream analyses given its obligate intracellular nature and difficulty with in vitro propagation. No single tool exists for the entirety of Ct genome assembly, necessitating the adaptation of multiple programs with varying success. Compounding this issue is the absence of reliable Ct reference strain genomes. We, therefore, developed CtGAP-Chlamydia trachomatisGenome Assembly Pipeline-as an integrated 'one-stop-shop' pipeline for assembly and characterization of Ct genome sequencing data from various sources including isolates, in vitro samples, clinical swabs and urine. CtGAP, written in Snakemake, enables read quality statistics output, adapter and quality trimming, host read removal, de novo and reference-guided assembly, contig scaffolding, selective ompA, multi-locus-sequence and plasmid typing, phylogenetic tree construction, and recombinant genome identification. Twenty Ct reference genomes were also generated. Successfully validated on a diverse collection of 363 samples containing Ct, CtGAP represents a novel pipeline requiring minimal bioinformatics expertise with easy adaptation for use with other bacterial species.

Chlamydia trachomatis↗

TFinder: A Python Web Tool for Predicting Transcription Factor Binding Sites.

Transcription is a key cell process that consists of synthesizing several copies of RNA from a gene DNA sequence. This process is highly regulated and closely linked to the ability of transcription factors to bind specifically to DNA. TFinder is an easy-to-use Python web portal allowing the identification of Individual Motifs (IM) such as Transcription Factor Binding Sites (TFBS). Using the NCBI API, TFinder extracts either promoter or gene terminal regulatory regions, through a simple query of NCBI gene name or ID. It enables simultaneous analysis across five different species for an unlimited number of genes. TFinder searches for Individual Motifs in different formats, including IUPAC codes and JASPAR entries. Moreover, TFinder also allows de novo generations of a Position Weight Matrix (PWM) and the use of already established PWM. Finally, the data are provided in a tabular and a graph format showing the relevance and the P-value of the Individual Motifs found as well as their location relative to the Transcription Start Site (TSS) or the terminal region of the gene. The results are then sent by email to users facilitating the subsequent data analysis and sharing. TFinder is written in Python and freely available on GitHub under the MIT license: https://github.com/Jumitti/TFinder. It can be accessed as a web application implemented in Streamlit at https://tfinder-ipmc.streamlit.app. Resources are available on Streamlit "Resources" tab. TFINDER strength is that it relies on an all-in-one intuitive tool allowing users inexperienced with bioinformatics tools to retrieve gene regulatory regions sequences in multiple species and to search for individual motifs in a huge number of genes.

Transcription Factors↗

The European Bioinformatics Institute (EBI) databases.

This paper describes the databases and services of the European Bioinformatics Institute (EBI). In collaboration with DDBJ and GenBank/NCBI, the EBI maintains and distributes the EMBL Nucleotide Sequence Database, Europe's primary nucleotide sequence data resource. The EBI also maintains and distributes the SWISS-PROT Protein Sequence Database, in collaboration with Amos Bairoch of the University of Geneva. Over thirty additional specialist molecular biology databases, as well as software and documentation of interest to molecular biologists, are also available. The EBI network services include database searching, entry retrieval, and sequence similarity searching facilities.

Amino Acid Sequence↗

Computational methods for gene annotation: the Arabidopsis genome.

Since the structure of the DNA molecule was identified half a century ago, the complete genome sequence has been determined for 37 prokaryotes and several eukaryotes. With the exponential growth of genetic information, bioinformatics has attempted to predict gene locations and functions in cyberspace prior to experimental confirmation at the bench.

Arabidopsis↗

Storing biological sequence databases in relational form.

SUMMARY: We have created a set of applications using Perl and Java in combination with XML technology to install biological sequence databases into an Oracle RDBMS. An easy-to-use interface using Java has been created for database query and other tools developed to integrate with our in-house bioinformatics applications. AVAILIBILITY: The database schema, DTD file, and source codes are available from the authors via email. CONTACT: guochun_ xie@merck. com

Amino Acid Sequence↗

PhyloNaP: a user-friendly database of phylogeny for natural product-producing enzymes.

SUMMARY: Phylogenetic analysis is widely used to predict enzyme function, yet building annotated and reusable trees is labor-intensive and requires extensive knowledge about the specific enzymes. Existing resources rarely cover biosynthetic enzymes and lack the context needed for meaningful analysis. We present PhyloNaP, the first large-scale resource dedicated to phylogenies of biosynthetic enzymes. PhyloNaP provides ∼51 000 annotated and interactive trees enriched with chemical, functional, and taxonomic information. Users can classify their own sequences via phylogenetic placement, enabling functional inference in an evolutionary context. A contribution portal allows the community to submit curated trees. By combining scale, breadth of annotation, and interactive functionality, PhyloNaP fills a major gap in bioinformatics resources for enzyme discovery and annotation, with immediate applications to secondary metabolism and beyond. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at https://phylonap.cs.uni-tuebingen.de.

Phylogeny↗