Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

The modular nature of genetic diseases.

Evidence from many sources suggests that similar phenotypes are begotten by functionally related genes. This is most obvious in the case of genetically heterogeneous diseases such as Fanconi anemia, Bardet-Biedl or Usher syndrome, where the various genes work together in a single biological module. Such modules can be a multiprotein complex, a pathway, or a single cellular or subcellular organelle. This observation suggests a number of hypotheses about the human phenome that are now beginning to be explored. First, there is now good evidence from bioinformatic analyses that human genetic diseases can be clustered on the basis of their phenotypic similarities and that such a clustering represents true biological relationships of the genes involved. Second, one may use such phenotypic similarity to predict and then test for the contribution of apparently unrelated genes to the same functional module. This concept is now being systematically tested for several diseases. Most recently, a systematic yeast two-hybrid screen of all known genes for inherited ataxias indicated that they all form part of a single extended protein-protein interaction network. Third, one can use bioinformatics to make predictions about new genes for diseases that form part of the same phenotype cluster. This is done by starting from the known disease genes and then searching for genes that share one or more functional attributes such as gene expression pattern, coevolution, or gene ontology. Ultimately, one may expect that a modular view of disease genes should help the rapid identification of additional disease genes for multifactorial diseases once the first few contributing genes (or environmental factors) have been reliably identified.

Computational Biology↗

Expressed sequence tags: analysis and annotation.

Expressed sequence tags (ESTs) present a special set of problems for bioinformatic analysis. They are partial and error-prone, and large datasets can have significant internal redundancy. To facilitate analysis of small EST datasets from in-house projects, we present an integrated "pipeline" of tools that take EST data from sequence trace to database submission. These tools also can be used to provide clustering of ESTs into putative genes and to annotate these genes with preliminary sequence similarity searches. The systems are written to use the public-domain LINUX environment and other openly available analytical tools.

Computational Biology↗

BioInfo3D: a suite of tools for structural bioinformatics.

Here, we describe BioInfo3D, a suite of freely available web services for protein structural analysis. The FlexProt method performs flexible structural alignment of protein molecules. FlexProt simultaneously detects the hinge regions and aligns the rigid subparts of the molecules. It does not require an a priori knowledge of the flexible hinge regions. MultiProt and MASS perform simultaneous comparison of multiple protein structures. PatchDock performs prediction of protein-protein and protein-small molecule interactions. The input to all services is either protein PDB codes or protein structures uploaded to the server. All the services are available at http://bioinfo3d.cs.tau.ac.il.

Algorithms↗

SNOW: standard nomenclature wizard to help searching for (bio) chemical standardized names.

UNLABELLED: When developing bioinformatical tools dealing with enzymatic activity, metabolism or enzymatic networks, the problem of the lack of a clear nomenclature for biochemical compounds often arises. This problem leads us to develop a small web-based tool (SNOW, Standard NOmenclature Wizard) which may help to find recommended and trivial names or the correct closest spelling for a query compound name, if it exists. AVAILABILITY: Web-based interface available at http://ibb.uab.es/snow/ SUPPLEMENTARY INFORMATION: http://ibb.uab.es/snow/snow_moreinfo.html

Algorithms↗

PAT: a protein analysis toolkit for integrated biocomputing on the web.

PAT, for Protein Analysis Toolkit, is an integrated biocomputing server. The main goal of its design was to facilitate the combination of different processing tools for complex protein analyses and to simplify the automation of repetitive tasks. The PAT server provides a standardized web interface to a wide range of protein analysis tools. It is designed as a streamlined analysis environment that implements many features which strongly simplify studies dealing with protein sequences and structures and improve productivity. PAT is able to read and write data in many bioinformatics formats and to create any desired pipeline by seamlessly sending the output of a tool to the input of another tool. PAT can retrieve protein entries from identifier-based queries by using pre-computed database indexes. Users can easily formulate complex queries combining different analysis tools with few mouse clicks, or via a dedicated macro language, and a web session manager provides direct access to any temporary file generated during the user session. PAT is freely accessible on the Internet at http://pat.cbs.cnrs.fr.

Computational Biology↗

SeqHound: biological sequence and structure database as a platform for bioinformatics research.

BACKGROUND: SeqHound has been developed as an integrated biological sequence, taxonomy, annotation and 3-D structure database system. It provides a high-performance server platform for bioinformatics research in a locally-hosted environment. RESULTS: SeqHound is based on the National Center for Biotechnology Information data model and programming tools. It offers daily updated contents of all Entrez sequence databases in addition to 3-D structural data and information about sequence redundancies, sequence neighbours, taxonomy, complete genomes, functional annotation including Gene Ontology terms and literature links to PubMed. SeqHound is accessible via a web server through a Perl, C or C++ remote API or an optimized local API. It provides functionality necessary to retrieve specialized subsets of sequences, structures and structural domains. Sequences may be retrieved in FASTA, GenBank, ASN.1 and XML formats. Structures are available in ASN.1, XML and PDB formats. Emphasis has been placed on complete genomes, taxonomy, domain and functional annotation as well as 3-D structural functionality in the API, while fielded text indexing functionality remains under development. SeqHound also offers a streamlined WWW interface for simple web-user queries. CONCLUSIONS: The system has proven useful in several published bioinformatics projects such as the BIND database and offers a cost-effective infrastructure for research. SeqHound will continue to develop and be provided as a service of the Blueprint Initiative at the Samuel Lunenfeld Research Institute. The source code and examples are available under the terms of the GNU public license at the Sourceforge site http://sourceforge.net/projects/slritools/ in the SLRI Toolkit.

Amino Acid Sequence↗

InSilicoSpectro: an open-source proteomics library.

We present a new proteomics open-source project, InSilicoSpectro, aimed at implementing recurrent computations that are necessary for proteomics data analysis. Illustrative examples are mass list file format conversions, protein sequence digestion, theoretical peptide and fragment mass computations, graphical display, matching with experimental data, isoelectric point estimation, and peptide retention time prediction. The project library is written in Perl, a widely used scripting language in bioinformatics, and it offers a unique framework of integrated objects to implement complex proteomics data analyses. For instance, only a few lines of code are required to digest a protein with fixed and variable modifications, label peptides with 18O, compute the fragmentation spectra and display their match with experimental spectra. We believe that InSilicoSpectro will be of great help to bioinformaticians, without detailed knowledge of proteomics specifics, and to mass spectrometrists with computer programming interest as well.

Amino Acid Sequence↗

CINEMA-MX: a modular multiple alignment editor.

UNLABELLED: Analyzing and visualizing multiple sequence alignments is a common task in many areas of molecular biology and bioinformatics. Many tools exist for this purpose, but are not easily customizable for specific in-house uses. Here we report the development of an editor, CINEMA-MX, that addresses these issues. CINEMA-MX is highly modular and configurable, and we present examples to illustrate its extensibility. AVAILABILITY: The program and full source code, which are available from http://www.bioinf.man.ac.uk/dbbrowser/cinema-mx, are being released under a combination of the LGPL and GPL, for Unix or Windows platforms.

Computer Graphics↗

VirDetector: a bioinformatic pipeline for virus surveillance using nanopore sequencing.

SUMMARY: Virus surveillance programmes are designed to counter the growing threat of viral outbreaks to human health. Nanopore sequencing, in particular, has proven to be suitable for this purpose, as it is readily available and provides rapid results. However, as special bioinformatic programs are required to extract the relevant information from the sequencing data, applications are needed that allow users without extensive bioinformatics knowledge to carry out the relevant analysis steps. We present VirDetector, a bioinformatic pipeline for virus surveillance using nanopore sequencing. The pipeline automatically installs all required programs and databases and allows all its steps to be executed with a single console command. After preprocessing the samples, including the possibility for basecalling, the pipeline classifies each sample taxonomically and reconstructs the viral consensus genomes, which are then used in phylogenetic analyses. This streamlined workflow provides a user-friendly and efficient solution for monitoring viral pathogens. AVAILABILITY AND IMPLEMENTATION: VirDetector is freely available at https://github.com/NLKaiser/VirDetector and https://zenodo.org/records/14637302 (10.5281/zenodo.14637302).

Nanopore Sequencing↗

TOUCAN 2: the all-inclusive open source workbench for regulatory sequence analysis.

We present the second and improved release of the TOUCAN workbench for cis-regulatory sequence analysis. TOUCAN implements and integrates fast state-of-the-art methods and strategies in gene regulation bioinformatics, including algorithms for comparative genomics and for the detection of cis-regulatory modules. This second release of TOUCAN has become open source and thereby carries the potential to evolve rapidly. The main goal of TOUCAN is to allow a user to come to testable hypotheses regarding the regulation of a gene or of a set of co-regulated genes. TOUCAN can be launched from this location: http://www.esat.kuleuven.ac.be/~saerts/software/toucan.php.

Algorithms↗

404 not found: the stability and persistence of URLs published in MEDLINE.

MOTIVATION: The advent of the World Wide Web has enabled unprecedented supplementation of traditional journal publications, allowing access to resources, such as video, sound, software, databases, datasets too large to publish, and even supplementary information and discussion. However, unlike traditional publications, continued availability of these online resources is not guaranteed. An automated survey was conducted to quantify the growth in Uniform Resource Locators (URLs) published to date in MEDLINE abstracts, their current availability and distribution by journal. RESULTS: Of 1630 unique URLs identified, formatting and/or spelling errors were detected within 201 (12%) of them as published. After corrections were made, a survey revealed that approximately 63% of these URLs were consistently available, and another 19% were available intermittently. The rate of failure was far worse for anonymous login to FTP sites, with only 12 of 33 sites (36%) responding. This survey also shows that journals vary disproportionately in the number of web citations published, suggesting policy implementation among a few could have a profound impact overall. Out of the 306 journals with a URL published in an abstract, Bioinformatics published the most (12% of total). AVAILABILITY: URL database and program available by request.

Abstracting and Indexing↗

Extending MapMan: application to legume genome arrays.

MOTIVATION: Based on a gene classification into hierarchical categories ('BINs'), MapMan was originally developed to display Arabidopsis thaliana gene expression in a functional context. We have created a bioinformatics system to extend MapMan to any organism by using a new BIN structure based on the KEGG database. Gene sequences are assigned to this ontology by homology relationships in four reference databases: KEGG, COG, Swiss-Prot and Gene Ontology. We applied this system to tailor MapMan to the GeneChips of two model legumes, Glycine max and Medicago truncatula. We also developed a module to identify the most relevant pathways involved. AVAILABILITY: All mapping files, pathway pictures and the analysis method are available at http://bioinfoserver.rsbs.anu.edu.au/

Algorithms↗

A simple spreadsheet-based, MIAME-supportive format for microarray data: MAGE-TAB.

BACKGROUND: Sharing of microarray data within the research community has been greatly facilitated by the development of the disclosure and communication standards MIAME and MAGE-ML by the MGED Society. However, the complexity of the MAGE-ML format has made its use impractical for laboratories lacking dedicated bioinformatics support. RESULTS: We propose a simple tab-delimited, spreadsheet-based format, MAGE-TAB, which will become a part of the MAGE microarray data standard and can be used for annotating and communicating microarray data in a MIAME compliant fashion. CONCLUSION: MAGE-TAB will enable laboratories without bioinformatics experience or support to manage, exchange and submit well-annotated microarray data in a standard format using a spreadsheet. The MAGE-TAB format is self-contained, and does not require an understanding of MAGE-ML or XML.

Computational Biology↗

Comparative map and trait viewer (CMTV): an integrated bioinformatic tool to construct consensus maps and compare QTL and functional genomics data across genomes and experiments.

In the past few decades, a wealth of genomic data has been produced in a wide variety of species using a diverse array of functional and molecular marker approaches. In order to unlock the full potential of the information contained in these independent experiments, researchers need efficient and intuitive means to identify common genomic regions and genes involved in the expression of target phenotypic traits across diverse conditions. To address this need, we have developed a Comparative Map and Trait Viewer (CMTV) tool that can be used to construct dynamic aggregations of a variety of types of genomic datasets. By algorithmically determining correspondences between sets of objects on multiple genomic maps, the CMTV can display syntenic regions across taxa, combine maps from separate experiments into a consensus map, or project data from different maps into a common coordinate framework using dynamic coordinate translations between source and target maps. We present a case study that illustrates the utility of the tool for managing large and varied datasets by integrating data collected by CIMMYT in maize drought tolerance research with data from public sources. This example will focus on one of the visualization features for Quantitative Trait Locus (QTL) data, using likelihood ratio (LR) files produced by generic QTL analysis software and displaying the data in a unique visual manner across different combinations of traits, environments and crosses. Once a genomic region of interest has been identified, the CMTV can search and display additional QTLs meeting a particular threshold for that region, or other functional data such as sets of differentially expressed genes located in the region; it thus provides an easily used means for organizing and manipulating data sets that have been dynamically integrated under the focus of the researcher's specific hypothesis.

Adaptation, Physiological↗

Health On the Net automated database of health and medical information.

With the number of World Wide Web sites growing every day, the problem is not just to find information, but to locate the right piece of information. Current World Wide Web search engines have not resolved this problem as they most often return a long list of documents. The search result is then unusable because of the large number of answers from different domains and topics. Only complex queries may, in a given situation, produce a limited number of potentially relevant documents. To make searches more efficient and usable by common users, we now need intelligent and specialised search engines on the Net [1,2]. Health On the Net Foundation and the Molecular Imaging and Bioinformatics Laboratory at Geneva University Hospital have developed Multi-Agent Retrieval Vagabond on Information Networks (MARVIN), a robot that searches sites and documents specifically related to a given specialised field. One such robot has already been implemented and used for the medical and the 2D electrophoresis domains. Health On the Net Foundation has implemented the corresponding search engines, MedHunt (http://www.hon.ch/cgi-bin/find) for the medical field and 2DHunt (http://www.hon.ch/cgi-bin/2DHunt/find) for the 2D electrophoresis field.

Computer Communication Networks↗

Globally distributed object identification for biological knowledgebases.

The World-Wide Web provides a globally distributed communication framework that is essential for almost all scientific collaboration, including bioinformatics. However, several limits and inadequacies have become apparent, one of which is the inability to programmatically identify locally named objects that may be widely distributed over the network. This shortcoming limits our ability to integrate multiple knowledgebases, each of which gives partial information of a shared domain, as is commonly seen in bioinformatics. The Life Science Identifier (LSID) and LSID Resolution System (LSRS) provide simple and elegant solutions to this problem, based on the extension of existing internet technologies. LSID and LSRS are consistent with next-generation semantic web and semantic grid approaches. This article describes the syntax, operations, infrastructure compatibility considerations, use cases and potential future applications of LSID and LSRS. We see the adoption of these methods as important steps toward simpler, more elegant and more reliable integration of the world's biological knowledgebases, and as facilitating stronger global collaboration in biology.

Animals↗

PULPO: pipeline of understanding large-scale patterns of oncogenomic signatures.

SUMMARY: PULPO v1.0 is a novel; fully automated pipeline designed for the preprocess and extraction of mutational signatures from raw Optical Genome Mapping (OGM) data. Built using Snakemake and executed within an isolated, Conda-managed environment, PULPO transforms complex cytogenetic alterations, captured at ultra-high resolution, into Catalogue of somatic mutations in cancer mutational signatures (COSMIC). This innovative approach not only enables researchers to work directly from raw OGM inputs but also streamlines the traditionally complex process of signature extraction, making advanced oncogenomic analyses accessible to users with varying levels of bioinformatics expertise. By facilitating the integration of comprehensive structural variants (SVs) and copy number variants (CNVs) data with established signature catalogues, PULPO paves the way for improved diagnostic accuracy and personalized therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The pipeline is open source and freely available under the MIT License at https://github.com/OncologyHNJ/PULPO-v.1.0 and DOI in Zenodo: https://zenodo.org/records/17749097.

Software↗

An ant colony optimisation algorithm for the 2D and 3D hydrophobic polar protein folding problem.

BACKGROUND: The protein folding problem is a fundamental problems in computational molecular biology and biochemical physics. Various optimisation methods have been applied to formulations of the ab-initio folding problem that are based on reduced models of protein structure, including Monte Carlo methods, Evolutionary Algorithms, Tabu Search and hybrid approaches. In our work, we have introduced an ant colony optimisation (ACO) algorithm to address the non-deterministic polynomial-time hard (NP-hard) combinatorial problem of predicting a protein's conformation from its amino acid sequence under a widely studied, conceptually simple model - the 2-dimensional (2D) and 3-dimensional (3D) hydrophobic-polar (HP) model. RESULTS: We present an improvement of our previous ACO algorithm for the 2D HP model and its extension to the 3D HP model. We show that this new algorithm, dubbed ACO-HPPFP-3, performs better than previous state-of-the-art algorithms on sequences whose native conformations do not contain structural nuclei (parts of the native fold that predominantly consist of local interactions) at the ends, but rather in the middle of the sequence, and that it generally finds a more diverse set of native conformations. CONCLUSIONS: The application of ACO to this bioinformatics problem compares favourably with specialised, state-of-the-art methods for the 2D and 3D HP protein folding problem; our empirical results indicate that our rather simple ACO algorithm scales worse with sequence length but usually finds a more diverse ensemble of native states. Therefore the development of ACO algorithms for more complex and realistic models of protein structure holds significant promise.

Algorithms↗