Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Identification and characterization of DISP3 gene in silico.

Lipid-modified soluble proteins Hedgehog (SHH, DHH and IHH) and WNT (WNT1, WNT2, WNT2B, WNT3, etc.) share distantly related mechanisms for ligand modification as well as for signaling through seven-transmembrane protein with Frizzled domain. Hedgehog and WNT signaling pathways network together during embryogenesis and carcinogenesis. Dispatched 1 (DISP1) and Dispatched 2 (DISP2) are human homologs for Drosophila Dispatched implicated in the release of lipid-anchored Hedgehog from producing cells. Here, we identified and characterized Dispatched 3 (DISP3) gene by using bioinformatics. DISP3 complete coding sequence was determined by assembling BU170953 EST and KIAA1337 uncharacterized cDNA. DISP3 gene at human chromosome 1p36.22 was linked to D1S2667 microsatellite maker and TERE1 gene, whose locus is associated with prostate cancer, bladder cancer, and liver cancer. DISP3 mRNA was expressed in human embryonic stem (ES) cells, brain, testis, lung carcinoid, neuroblastoma, retinoblastoma and brain tumor. DISPH1 domain with five transmembrane regions (codon 452-637 of DISP3) and DISPH2 domain with four transmembrane regions (codon 1116-1319 of DISP3) were identified as novel domains conserved between DISP3 (1392 aa) and DISP1. The region around DISPH1 and DISPH2 domains of DISP3 protein was the Patched homologous region conserved among Patched family members and DISP family members. Because DISP3 and DISP1 are multi-span transmembrane proteins with the Patched homologous region, DISP3 is predicted to be implicated in the release of lipid-anchored secreted proteins. This is the first report on identification and characterization of the DISP3 gene.

Amino Acid Sequence↗

A programming course in bioinformatics for computer and information science students.

We have created a course entitled "Representations and Algorithms for Computational Molecular Biology" with three specific goals in mind. First, we want to provide a technical introduction for computer science and medical information science students to the challenges of computing with molecular biology data, particularly the advantages of having easy access to real-world data sets. Second, we want to equip the students with the skills required of productive research assistants in molecular biology computing research projects. Finally, we want to provide a showcase for local investigators to describe their work in the context of a course that provide adequate background information. In order to achieve these goals, we have created a programming course, in which three major projects and six smaller assignments are assigned during the quarter. We stress fundamental representations and algorithms during the first part of the course in lectures given by the core faculty, and then have more focused lectures in which faculty research interests are highlighted. The course stressed issues of structural molecular biology, in order to better motivate the critical issues in sequence analysis. The culmination of the course was a challenge to the students to use a version of protein threading to predict which members of a set of unknown sequences were globins. The course was well received, and has been made a core requirement in the Medical Information Sciences program.

Algorithms↗

MAO: a Multiple Alignment Ontology for nucleic acid and protein sequences.

The application of high-throughput techniques such as genomics, proteomics or transcriptomics means that vast amounts of heterogeneous data are now available in the public databases. Bioinformatics is responding to the challenge with new integrated management systems for data collection, validation and analysis. Multiple alignments of genomic and protein sequences provide an ideal environment for the integration of this mass of information. In the context of the sequence family, structural and functional data can be evaluated and propagated from known to unknown sequences. However, effective integration is being hindered by syntactic and semantic differences between the different data resources and the alignment techniques employed. One solution to this problem is the development of an ontology that systematically defines the terms used in a specific domain. Ontologies are used to share data from different resources, to automatically analyse information and to represent domain knowledge for non-experts. Here, we present MAO, a new ontology for multiple alignments of nucleic and protein sequences. MAO is designed to improve interoperation and data sharing between different alignment protocols for the construction of a high quality, reliable multiple alignment in order to facilitate knowledge extraction and the presentation of the most pertinent information to the biologist.

Databases, Genetic↗

GOTree Machine (GOTM): a web-based platform for interpreting sets of interesting genes using Gene Ontology hierarchies.

BACKGROUND: Microarray and other high-throughput technologies are producing large sets of interesting genes that are difficult to analyze directly. Bioinformatics tools are needed to interpret the functional information in the gene sets. RESULTS: We have created a web-based tool for data analysis and data visualization for sets of genes called GOTree Machine (GOTM). This tool was originally intended to analyze sets of co-regulated genes identified from microarray analysis but is adaptable for use with other gene sets from other high-throughput analyses. GOTree Machine generates a GOTree, a tree-like structure to navigate the Gene Ontology Directed Acyclic Graph for input gene sets. This system provides user friendly data navigation and visualization. Statistical analysis helps users to identify the most important Gene Ontology categories for the input gene sets and suggests biological areas that warrant further study. GOTree Machine is available online at http://genereg.ornl.gov/gotm/. CONCLUSION: GOTree Machine has a broad application in functional genomic, proteomic and other high-throughput methods that generate large sets of interesting genes; its primary purpose is to help users sort for interesting patterns in gene sets.

Animals↗

Discovery of Time-Delayed Gene Regulatory Networks based on temporal gene expression profiling.

BACKGROUND: It is one of the ultimate goals for modern biological research to fully elucidate the intricate interplays and the regulations of the molecular determinants that propel and characterize the progression of versatile life phenomena, to name a few, cell cycling, developmental biology, aging, and the progressive and recurrent pathogenesis of complex diseases. The vast amount of large-scale and genome-wide time-resolved data is becoming increasing available, which provides the golden opportunity to unravel the challenging reverse-engineering problem of time-delayed gene regulatory networks. RESULTS: In particular, this methodological paper aims to reconstruct regulatory networks from temporal gene expression data by using delayed correlations between genes, i.e., pairwise overlaps of expression levels shifted in time relative each other. We have thus developed a novel model-free computational toolbox termed TdGRN (Time-delayed Gene Regulatory Network) to address the underlying regulations of genes that can span any unit(s) of time intervals. This bioinformatics toolbox has provided a unified approach to uncovering time trends of gene regulations through decision analysis of the newly designed time-delayed gene expression matrix. We have applied the proposed method to yeast cell cycling and human HeLa cell cycling and have discovered most of the underlying time-delayed regulations that are supported by multiple lines of experimental evidence and that are remarkably consistent with the current knowledge on phase characteristics for the cell cyclings. CONCLUSION: We established a usable and powerful model-free approach to dissecting high-order dynamic trends of gene-gene interactions. We have carefully validated the proposed algorithm by applying it to two publicly available cell cycling datasets. In addition to uncovering the time trends of gene regulations for cell cycling, this unified approach can also be used to study the complex gene regulations related to the development, aging and progressive pathogenesis of a complex disease where potential dependences between different experiment units might occurs.

Algorithms↗

SIMAP--the similarity matrix of proteins.

MOTIVATION: Sequence similarity searches are of great importance in bioinformatics. Exhaustive searches for homologous proteins in databases are computationally expensive and can be replaced by a database of pre-calculated homologies in many cases. Retrieving similarities from an incrementally updated database instead of repeatedly recalculating them should provide homologs much faster and frees computational resources for other purposes. RESULTS: We have implemented SIMAP-a database containing the similarity space formed by almost all amino acid sequences from public databases and completely sequenced genomes. The database is capable of handling very large datasets and allows incremental updates. We have implemented a powerful backbone for similarity computation, which is based on FASTA heuristics. By providing WWW interfaces as well as web services, we make our data accessible to the worldwide community. We have also adapted procedures to detect putative orthologs as example applications. AVAILABILITY: The SIMAP portal page providing links to SIMAP services is publicly available: http://mips.gsf.de/services/analysis/simap/. The web services can be accessed under http://mips.gsf.de/proj/hobitws/services/RPCSimapService?wsdl and http://mips.gsf.de/proj/hobitws/services/DocSimapService?wsdl

Algorithms↗

PHACCS, an online tool for estimating the structure and diversity of uncultured viral communities using metagenomic information.

BACKGROUND: Phages, viruses that infect prokaryotes, are the most abundant microbes in the world. A major limitation to studying these viruses is the difficulty of cultivating the appropriate prokaryotic hosts. One way around this limitation is to directly clone and sequence shotgun libraries of uncultured viral communities (i.e., metagenomic analyses). PHACCS http://phage.sdsu.edu/phaccs, Phage Communities from Contig Spectrum, is an online bioinformatic tool to assess the biodiversity of uncultured viral communities. PHACCS uses the contig spectrum from shotgun DNA sequence assemblies to mathematically model the structure of viral communities and make predictions about diversity. RESULTS: PHACCS builds models of possible community structure using a modified Lander-Waterman algorithm to predict the underlying contig spectrum. PHACCS finds the most appropriate structure model by optimizing the model parameters until the predicted contig spectrum is as close as possible to the experimental one. This model is the basis for making estimates of uncultured viral community richness, evenness, diversity index and abundance of the most abundant genotype. CONCLUSION: PHACCS analysis of four different environmental phage communities suggests that the power law is an important rank-abundance form to describe uncultured viral community structure. The estimates support the fact that the four phage communities were extremely diverse and that phage community biodiversity and structure may be correlated with that of their hosts.

Algorithms↗

Telemedical database of Hodgkin's disease.

The creation of a complex telemedical system oriented towards childhood Hodgkin's disease has been undertaken at the Department of Bioinformatics and Telemedicine of the Jagiellonian University Medical College in cooperation with the Department of Oncology and Pediatric Hematology of the Polish-American Institute of Pediatrics, JU MC. Data collecting, data processing and data transmission is aimed to aid and/or supervise surgical and drug treatment. The Tele-Database of Childhood Hodgkin's Disease (TDCHD) is not a simple Internet database project. A few hundred data items are presented in each patient's record, covering the complete medical treatment period. Efficient management and proper data protection are necessary for a medical database. Therefore, the interface for entering data has been divided into several parts. Each part is subjected to separate editing and transfer. A double-layer debugging system has been applied in the program: the first pass occurs on the client side (programmed in JavaScript and XML), the second - on the server side (programmed in PHP). Strict authorization is requested for all participants. Clinical data collected according to management standards and information governance (data quality, security and confidentiality) is organized in a way that facilitates practical and scientific use.

Child↗

Ontological visualization of protein-protein interactions.

BACKGROUND: Cellular processes require the interaction of many proteins across several cellular compartments. Determining the collective network of such interactions is an important aspect of understanding the role and regulation of individual proteins. The Gene Ontology (GO) is used by model organism databases and other bioinformatics resources to provide functional annotation of proteins. The annotation process provides a mechanism to document the binding of one protein with another. We have constructed protein interaction networks for mouse proteins utilizing the information encoded in the GO annotations. The work reported here presents a methodology for integrating and visualizing information on protein-protein interactions. RESULTS: GO annotation at Mouse Genome Informatics (MGI) captures 1318 curated, documented interactions. These include 129 binary interactions and 125 interaction involving three or more gene products. Three networks involve over 30 partners, the largest involving 109 proteins. Several tools are available at MGI to visualize and analyze these data. CONCLUSIONS: Curators at the MGI database annotate protein-protein interaction data from experimental reports from the literature. Integration of these data with the other types of data curated at MGI places protein binding data into the larger context of mouse biology and facilitates the generation of new biological hypotheses based on physical interactions among gene products.

Animals↗

In silico study of breast cancer associated gene 3 using LION Target Engine and other tools.

Sequence analysis of individual targets is an important step in annotation and validation. As a test case, we investigated human breast cancer associated gene 3 (BCA3) with LION Target Engine and with other bioinformatics tools. LION Target Engine confirmed that the BCA3 gene is located on 11p15.4 and that the two most likely splice variants (lacking exon 3 and exons 3 and 5, respectively) exist. Based on our manual curation of sequence data, it is proposed that an additional variant (missing only exon 5) published in a public sequence repository, is a prediction artifact. A significant number of new orthologs were also identified, and these were the basis for a high-quality protein secondary structure prediction. Moreover, our research confirmed several distinct functional domains as described in earlier reports. Sequence conservation from multiple sequence alignments, splice variant identification, secondary structure predictions, and predicted phosphorylation sites suggest that the removal of interaction sites through alternative splicing might play a modulatory role in BCA3. This in silico approach shows the depth and relevance of an analysis that can be accomplished by including a variety of publicly available tools with an integrated and customizable life science informatics platform.

Adaptor Proteins, Signal Transducing↗

ChromBERT-tools: a versatile toolkit for context-specific regulatory representations of transcription regulators across different cell types.

SUMMARY: Representations that encode the genome-wide regulatory behavior of transcription regulators provide a foundation for flexible transcription modeling and in silico regulatory analysis. Existing regulator representations are commonly derived from gene co-expression, motif annotations, or static protein features, which capture useful but limited aspects of regulator identity but do not directly model how regulators participate in region-specific regulatory programs across the genome. ChromBERT addresses this gap by learning context-aware regulatory representations from large-scale ChIP-seq data. However, routine bioinformatics applications require lightweight, accessible, and modular tools for generating, adapting, and interpreting these representations in user-defined biological contexts. Here, we present ChromBERT-tools, a user-oriented toolkit built upon ChromBERT that converts its regulatory representation framework into practical workflows for customizable analysis across cellular contexts. ChromBERT-tools provides command-line interfaces and Python APIs organized into three functional layers: representation generation, predictive modeling, and regulatory interpretation. The representation generation layer produces representations of genomic regions and transcription regulators. The predictive modeling layer fine-tunes ChromBERT for genome-wide regulatory activity prediction through classification or regression tasks, with optimized implementation to reduce running time and computational resource requirements. The regulatory interpretation layer supports inference of the context-specific roles of cis-regulatory elements and transcription regulators. These modules can be used independently or integrated into end-to-end workflows, enabling flexible analyses across diverse datasets. ChromBERT-tools lowers the barrier to applying context-specific regulatory representations in routine genomic analyses. AVAILABILITY AND IMPLEMENTATION: ChromBERT-tools is freely available at https://github.com/TongjiZhanglab/ChromBERT-tools, with documentation at https://chrombert-tools.readthedocs.io/en/latest/. A frozen archival snapshot is available on Zenodo under DOI: 10.5281/zenodo.20094206.

Software↗

NAST: a multiple sequence alignment server for comparative analysis of 16S rRNA genes.

Microbiologists conducting surveys of bacterial and archaeal diversity often require comparative alignments of thousands of 16S rRNA genes collected from a sample. The computational resources and bioinformatics expertise required to construct such an alignment has inhibited high-throughput analysis. It was hypothesized that an online tool could be developed to efficiently align thousands of 16S rRNA genes via the NAST (Nearest Alignment Space Termination) algorithm for creating multiple sequence alignments (MSA). The tool was implemented with a web-interface at http://greengenes.lbl.gov/NAST. Each user-submitted sequence is compared with Greengenes' 'Core Set', comprising approximately 10,000 aligned non-chimeric sequences representative of the currently recognized diversity among bacteria and archaea. User sequences are oriented and paired with their closest match in the Core Set to serve as a template for inserting gap characters. Non-16S data (sequence from vector or surrounding genomic regions) are conveniently removed in the returned alignment. From the resulting MSA, distance matrices can be calculated for diversity estimates and organisms can be classified by taxonomy. The ability to align and categorize large sequence sets using a simple interface has enabled researchers with various experience levels to obtain bacterial and archaeal community profiles.

Algorithms↗

Expanding and improving analyses of nucleotide recoding RNA-seq experiments with the EZbakR suite.

Nucleotide recoding RNA sequencing methods (NR-seq; TimeLapse-seq, SLAM-seq, TUC-seq, etc.) are powerful approaches for assaying transcript population dynamics. In addition, these methods have been extended to probe a host of regulated steps in the RNA life cycle. Current bioinformatic tools significantly constrain analyses of NR-seq data. To address this limitation, we developed EZbakR (https://github.com/isaacvock/EZbakR), an R package to facilitate a more comprehensive set of NR-seq analyses, and fastq2EZbakR (https://github.com/isaacvock/fastq2EZbakR), a Snakemake pipeline for flexible preprocessing of NR-seq datasets, collectively referred to as the EZbakR suite. Together, these tools generalize many aspects of the NR-seq analysis workflow. The fastq2EZbakR pipeline can assign reads to a diverse set of genomic features (e.g., genes, exons, splice junctions), and EZbakR can perform analyses on any combination of these features. EZbakR extends standard NR-seq mutational modeling to support multi-label analyses (e.g., s4U and s6G dual labeling), and implements an improved hierarchical model to better account for transcript-to-transcript variance in metabolic label incorporation. EZbakR also generalizes dynamical systems modeling of NR-seq data to support analyses of premature mRNA processing and flow between subcellular compartments. Finally, EZbakR implements flexible and well-powered comparative analyses of all estimated parameters via design matrix-specified generalized linear modeling. The EZbakR suite will thus allow researchers to make full, effective use of NR-seq data.

Software↗

Five hierarchical levels of sequence-structure correlation in proteins.

This article reviews recent work towards modelling protein folding pathways using a bioinformatics approach. Statistical models have been developed for sequence-structure correlations in proteins at five levels of structural complexity: (i) short motifs; (ii) extended motifs; (iii) nonlocal pairs of motifs; (iv) 3-dimensional arrangements of multiple motifs; and (v) global structural homology. We review statistical models, including sequence profiles, hidden Markov models (HMMs) and interaction potentials, for the first four levels of structural detail. The I-sites (folding Initiation sites) Library models short local structure motifs. Each succeeding level has a statistical model, as follows: HMMSTR (HMM for STRucture) is an HMM for extended motifs; HMMSTR-CM (Contact Maps) is a model for pairwise interactions between motifs; and SCALI-HMM (HMMs for Structural Core ALIgnments) is a set of HMMs for the spatial arrangements of motifs. The parallels between the statistical models and theoretical models for folding pathways are discussed in this article; however, global sequence models are not discussed because they have been extensively reviewed elsewhere. The data used and algorithms presented in this article are available at http://www.bioinfo.rpi.edu/~bystrc/ (click on "servers" or "downloads") or by request to bystrc@rpi.edu .

Algorithms↗

Shortest triplet clustering: reconstructing large phylogenies using representative sets.

BACKGROUND: Understanding the evolutionary relationships among species based on their genetic information is one of the primary objectives in phylogenetic analysis. Reconstructing phylogenies for large data sets is still a challenging task in Bioinformatics. RESULTS: We propose a new distance-based clustering method, the shortest triplet clustering algorithm (STC), to reconstruct phylogenies. The main idea is the introduction of a natural definition of so-called k-representative sets. Based on k-representative sets, shortest triplets are reconstructed and serve as building blocks for the STC algorithm to agglomerate sequences for tree reconstruction in O(n2) time for n sequences. Simulations show that STC gives better topological accuracy than other tested methods that also build a first starting tree. STC appears as a very good method to start the tree reconstruction. However, all tested methods give similar results if balanced nearest neighbor interchange (BNNI) is applied as a post-processing step. BNNI leads to an improvement in all instances. The program is available at http://www.bi.uni-duesseldorf.de/software/stc/. CONCLUSION: The results demonstrate that the new approach efficiently reconstructs phylogenies for large data sets. We found that BNNI boosts the topological accuracy of all methods including STC, therefore, one should use BNNI as a post-processing step to get better topological accuracy.

Algorithms↗

AQuA Tools: clear and reliable BEDPE operations for 3D genomics.

MOTIVATION: The genome interacts with itself within the volume of the cell nucleus to process information. These interactions mediate signal integration, gene regulation, and cell identity. The identification of new therapeutic targets from non-coding disease-associated variants relies critically on correctly assigning variants to genes through 3D interactions. Experimental techniques in 3D genomics, such as HiC and HiChIP, allow the mapping of interactions through sequencing. Bioinformatics for 3D genomics contends primarily with contact matrices that contain interaction frequencies for all possible element pairs, and BEDPE files that store element pairs that interact. Whereas the tools available for processing linear genomic data are mature, operating on contact matrices and BEDPE files remains cumbersome, opaque, and error-prone, as researchers have had to shoehorn tools originally designed for linear data. A genome arithmetic designed from the ground up for 3D genomics does not yet exist. RESULTS: We present AQuA Tools, a suite of shell- and R-based command-line tools that provide a set of core operations on contact matrices and BEDPE files motivated by key questions in population genetics, cancer research, and precision medicine. We have designed our core operations to be clear, reliable, intuitive and versatile. Core operations can be chained together along with standard UNIX commands. Our goal is to make AQuA Tools easy for the novice to learn and the go-to choice for power users. We hope our tools will motivate more researchers to use 3D genomic data in their projects. AVAILABILITY AND IMPLEMENTATION: We provide and maintain AQuA Tools at https://github.com/axiotl/aqua-tools.

Genomics↗

Target and biomarker exploration portal for drug discovery.

MOTIVATION: The discovery of novel drug targets and precision biomarkers remains a major challenge in drug development, with traditional differential expression analysis often overlooking key regulatory proteins. Here, we present a novel, web-based bioinformatics tool, the Target and Biomarker Exploration Portal (TBEP), designed to accelerate the drug discovery process by integrating large-scale biomedical data with network analysis techniques. RESULTS: TBEP harnesses machine-learning approaches to mine and combine multimodal datasets, including human genetics, functional genomics, and protein-protein interaction networks, to decode causal disease mechanisms and uncover novel therapeutic targets and precision biomarkers for specific phenotypes. A unique feature of the tool is its ability to process large-scale data in real-time, facilitated by an efficient cloud-based architecture. Additionally, the tool incorporates an integrated large language model (LLM), which assists researchers in exploring and interpreting complex biological relationships within the generated networks and multi-omics data using natural language (English). By offering an intuitive, interactive interface, the LLM enhances the exploration of biological insights, making it easier for scientists to derive actionable conclusions. This powerful integration of network analysis, multi-omics data, and LLM provides a robust framework for accelerating the identification of novel drug targets. AVAILABILITY AND IMPLEMENTATION: The tool is publicly available at https://tbep.missouri.edu. The source code, documentation and installation instructions are available at GitHub repository: https://github.com/mizzoudbl/tbep.

Drug Discovery↗

Dynamic covariation between gene expression and proteome characteristics.

BACKGROUND: Cells react to changing intra- and extracellular signals by dynamically modulating complex biochemical networks. Cellular responses to extracellular signals lead to changes in gene and protein expression. Since the majority of genes encode proteins, we investigated possible correlations between protein parameters and gene expression patterns to identify proteome-wide characteristics indicative of trends common to expressed proteins. RESULTS: Numerous bioinformatics methods were used to filter and merge information regarding gene and protein annotations. A new statistical time point-oriented analysis was developed for the study of dynamic correlations in large time series data. The method was applied to investigate microarray datasets for different cell types, organisms and processes, including human B and T cell stimulation, Drosophila melanogaster life span, and Saccharomyces cerevisiae cell cycle. CONCLUSION: We show that the properties of proteins synthesized correlate dynamically with the gene expression profile, indicating that not only is the actual identity and function of expressed proteins important for cellular responses but that several physicochemical and other protein properties correlate with gene expression as well. Gene expression correlates strongly with amino acid composition, composition- and sequence-derived variables, functional, structural, localization and gene ontology parameters. Thus, our results suggest that a dynamic relationship exists between proteome properties and gene expression in many biological systems, and therefore this relationship is fundamental to understanding cellular mechanisms in health and disease.

Animals↗