Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

MAPPP: MHC class I antigenic peptide processing prediction.

MAPPP is a bioinformatics tool for the prediction of potential antigenic epitopes presented on the cell surface by major histocompatibility complex class I (MHC I) molecules to CD8 positive T lymphocytes. It combines existing predictions for proteasomal cleavage with peptide anchoring to MHC I molecules.

Algorithms↗

Theme discovery from gene lists for identification and viewing of multiple functional groups.

BACKGROUND: High throughput methods of the genome era produce vast amounts of data in the form of gene lists. These lists are large and difficult to interpret without advanced computational or bioinformatic tools. Most existing methods analyse a gene list as a single entity although it is comprised of multiple gene groups associated with separate biological functions. Therefore it is imperative to define and visualize gene groups with unique functionality within gene lists. RESULTS: In order to analyse the functional heterogeneity within a gene list, we have developed a method that clusters genes to groups with homogenous functionalities. The method uses Non-negative Matrix Factorization (NMF) to create several clustering results with varying numbers of clusters. The obtained clustering results are combined into a simple graphical presentation showing the functional groups over-represented in the analyzed gene list. We demonstrate its performance on two data sets and show results that improve upon existing methods. The comparison also shows that our method creates a more simplified view that aids in discovery of biological themes within the list and discards less informative classes from the results. CONCLUSION: The presented method and associated software are useful for the identification and interpretation of biological functions associated with gene lists and are especially useful for the analysis of large lists.

Cluster Analysis↗

nf-core/pacvar: a pipeline for analyzing long-read PacBio whole genome and repeat expansion sequencing data.

MOTIVATION: Pacific Biosciences (PacBio) single-molecule, long-read sequencing enables whole genome annotation and the characterization of 20 complex repetitive repeat regions, especially relevant to neurodegenerative diseases, through their PureTarget panel. Long-read whole-genome sequencing (WGS) also allows for the detection of structural variants that would be difficult to detect with traditional short-read sequencing. However, the raw unaligned Binary Alignment Map data need to be processed before analysis. There is a need for an intuitive comprehensive bioinformatic pipeline that can analyze these data. RESULTS: We present nf-core/pacvar, a comprehensive pipeline for analyzing both PacBio single-molecule PureTarget and WGS data that demultiplexes and parallelizes pre-processing, variant calling and repeat characterization. nf-core/pacvar is compatible with little configuration and has few dependencies. This pipeline enables rapid end-to-end, parallel processing of PacBio single-molecule whole genome and targeted repeat expansion sequencing. AVAILABILITY AND IMPLEMENTATION: nf-core/pacvar is available on nf-core website (https://nf-co.re/pacvar/) and on github (https://github.com/nf-core/pacvar) under MIT License (DOI: 10.5281/zenodo.14813048).

Software↗

SOAP-based services provided by the European Bioinformatics Institute.

SOAP (Simple Object Access Protocol) (http://www.w3.org/TR/soap) based Web Services technology (http://www.w3.org/ws) has gained much attention as an open standard enabling interoperability among applications across heterogeneous architectures and different networks. The European Bioinformatics Institute (EBI) is using this technology to provide robust data retrieval and data analysis mechanisms to the scientific community and to enhance utilization of the biological resources it already provides [N. Harte, V. Silventoinen, E. Quevillon, S. Robinson, K. Kallio, X. Fustero, P. Patel, P. Jokinen and R. Lopez (2004) Nucleic Acids Res., 32, 3-9]. These services are available free to all users from http://www.ebi.ac.uk/Tools/webservices.

Biotechnology↗

Genetic information resources: a new field for medical librarians.

The Human Genome Project is generating unprecedented quantities of new genetic information. The new discipline of bioinformatics has created many new molecular biology databanks to store the results of the Human Genome Project. This data is expected to be the information source for biomedical science in the 21st century. As molecular biology research moves out of the laboratory and becomes molecular medicine, a growing number of people need access to genetic information. Medical students, healthcare practitioners and patients need help in finding appropriate information. Medical librarians should know how to search key genetic information resources and should have a basic understanding of the type of information contained in each.

Databases, Bibliographic↗

System-based proteomic analysis of the interferon response in human liver cells.

BACKGROUND: Interferons (IFNs) play a critical role in the host antiviral defense and are an essential component of current therapies against hepatitis C virus (HCV), a major cause of liver disease worldwide. To examine liver-specific responses to IFN and begin to elucidate the mechanisms of IFN inhibition of virus replication, we performed a global quantitative proteomic analysis in a human hepatoma cell line (Huh7) in the presence and absence of IFN treatment using the isotope-coded affinity tag (ICAT) method and tandem mass spectrometry (MS/MS). RESULTS: In three subcellular fractions from the Huh7 cells treated with IFN (400 IU/ml, 16 h) or mock-treated, we identified more than 1,364 proteins at a threshold that corresponds to less than 5% false-positive error rate. Among these, 54 were induced by IFN and 24 were repressed by more than two-fold, respectively. These IFN-regulated proteins represented multiple cellular functions including antiviral defense, immune response, cell metabolism, signal transduction, cell growth and cellular organization. To analyze this proteomics dataset, we utilized several systems-biology data-mining tools, including Gene Ontology via the GoMiner program and the Cytoscape bioinformatics platform. CONCLUSIONS: Integration of the quantitative proteomics with global protein interaction data using the Cytoscape platform led to the identification of several novel and liver-specific key regulatory components of the IFN response, which may be important in regulating the interplay between HCV, interferon and the host response to virus infection.

Cell Line, Tumor↗

An easy-to-use pipeline to analyze amplicon-based Next Generation Sequencing results of human mitochondrial DNA from degraded samples.

Genome and transcriptome examinations have become more common due to Next-Generation Sequencing (NGS), which significantly increases throughput and depth coverage while reducing costs and time. Mitochondrial DNA (mtDNA) is often the marker of choice in degraded samples from archaeological and forensic contexts, as its higher number of copies can improve the success of the experiment. Among other sequencing strategies, amplicon-based NGS techniques are currently being used to obtain enough data to be analyzed. There are some pipelines designed for the analysis of ancient mtDNA samples and others for the analysis of amplicon data. However, these pipelines pose a challenge for non-expert users and cannot often address both ancient and forensic DNA particularities and amplicon-based sequencing simultaneously. To overcome these challenges, a user-friendly bioinformatic tool was developed to analyze the non-coding region of human mtDNA from degraded samples recovered in archaeological and forensic contexts. The tool can be easily modified to fit the specifications of other amplicon-based NGS experiments. A comparative analysis between two tools, MarkDuplicates from Picard and dedup parameter from fastp, both designed for duplicate removal was conducted. Additionally, various thresholds of PMDtools, a specialized tool designed for extracting reads affected by post-mortem damage, were used. Finally, the depth coverage of each amplicon was correlated with its level of damage. The results obtained indicated that, for removing duplicates, dedup is a better tool since retains more non-repeated reads, that are removed by MarkDuplicates. On the other hand, a PMDS = 1 in PMDtools was the threshold that allowed better differentiation between present-day and ancient samples, in terms of damage, without losing too many reads in the process. These two bioinformatic tools were added to a pipeline designed to obtain both haplotype and haplogroup of mtDNA. Furthermore, the pipeline presented in the present study generates information about the quality and possible contamination of the sample. This pipeline is designed to automatize mtDNA analysis, however, particularly for ancient samples, some manual analyses may be required to fully validate results since the amplicons that used to be more easily recovered were the ones that had fewer reads with damage, indicating that special care must be taken for poor recovered samples.

DNA, Mitochondrial↗

GERMINATE. a generic database for integrating genotypic and phenotypic information for plant genetic resource collections.

The extensive germplasm resource collections that are now available for major crop plants and their wild relatives will increasingly provide valuable biological and bioinformatics resources for plant physiologists and geneticists to dissect the molecular basis of key traits and to develop highly adapted plant material to sustain future breeding programs. A key to the efficient deployment of these resources is the development of information systems that will enable the collection and storage of biological information for these plant lines to be integrated with the molecular information that is now becoming available through the use of high-throughput genomics and post-genomics technologies. The GERMINATE database has been designed to hold a diverse variety of data types, ranging from molecular to phenotypic, and to allow querying between such data for any plant species. Data are stored in GERMINATE in a technology-independent manner, such that new technologies can be accommodated in the database as they emerge, without modification of the underlying schema. Users can access data in GERMINATE databases either via a lightweight Perl-CGI Web interface or by the more complex Genomic Diversity and Phenotype Connection software. GERMINATE is released under the GNU General Public License and is available at http://germinate.scri.sari.ac.uk/germinate/.

Computational Biology↗

Nanopore cheminformatics.

A cheminformatics method is described for classification, and biophysical examination, of individual molecules. A novel molecular detector is used--one based on current blockade measurements through a nanometer-scale ion channel (alpha-hemolysin). Classification results are described for blockades caused by DNA molecules in the alpha-hemolysin nanopore detector, with signal analysis and pattern recognition performed using a combination of methods from bioinformatics and machine learning. Due to the size of the alpha-hemolysin protein channel, the blockade events report on one DNA molecule at a time, which enables a variety of reproducible, single-molecule biophysical experiments. To capture the full sensitivity of the nanopore detector's blockade signal, Hidden Markov Models (HMMs) were used with Expectation/Maximization for denoising and for associating a feature vector with the ionic current blockade of each captured DNA molecule. Support Vector Machines (SVMs) that employ novel kernel designs were then used as discriminators. With SVM training performed off-line, and economical HMM processing on-line, blockade classification was possible during capture. HMMs were also used in conjunction with a time-domain finite state automaton (off-line) for feature discovery and kinetics analysis. Analysis of the DNA data indicates a variety of binding (DNA-protein), fraying, and conformational shifts that are consistent with data obtained from thermodynamic analyses (melting curves), X-ray crystallography, and NMR studies. The software tools are designed for analysis of generic blockades in ionic channels, including those in other biological pore-forming toxins, other biological channels in general, and semiconductor-based channels.

Base Sequence↗

ClusterControl: a web interface for distributing and monitoring bioinformatics applications on a Linux cluster.

UNLABELLED: ClusterControl is a web interface to simplify distributing and monitoring bioinformatics applications on Linux cluster systems. We have developed a modular concept that enables integration of command line oriented program into the application framework of ClusterControl. The systems facilitate integration of different applications accessed through one interface and executed on a distributed cluster system. The package is based on freely available technologies like Apache as web server, PHP as server-side scripting language and OpenPBS as queuing system and is available free of charge for academic and non-profit institutions. AVAILABILITY: http://genome.tugraz.at/Software/ClusterControl

Computational Biology↗

Using GOstats to test gene lists for GO term association.

MOTIVATION: Functional analyses based on the association of Gene Ontology (GO) terms to genes in a selected gene list are useful bioinformatic tools and the GOstats package has been widely used to perform such computations. In this paper we report significant improvements and extensions such as support for conditional testing. RESULTS: We discuss the capabilities of GOstats, a Bioconductor package written in R, that allows users to test GO terms for over or under-representation using either a classical hypergeometric test or a conditional hypergeometric that uses the relationships among GO terms to decorrelate the results. AVAILABILITY: GOstats is available as an R package from the Bioconductor project: http://bioconductor.org

Algorithms↗

The Mouse Genome Database (MGD): updates and enhancements.

The Mouse Genome Database (MGD) integrates genetic and genomic data for the mouse in order to facilitate the use of the mouse as a model system for understanding human biology and disease processes. A core component of the MGD effort is the acquisition and integration of genomic, genetic, functional and phenotypic information about mouse genes and gene products. MGD works within the broader bioinformatics community to define referential and semantic standards to facilitate data exchange between resources including the incorporation of information from the biomedical literature. MGD is also a platform for computational assessment of integrated biological data with the goal of identifying candidate genes associated with complex phenotypes. MGD is web accessible at http://www.informatics.jax.org. Recent improvements in MGD described here include the incorporation of an interactive genome browser, the enhancement of phenotype resources and the further development of functional annotation resources.

Animals↗

Identification and characterization of Bombyx mori eIF5A gene through bioinformatics approaches.

As the genome of B. mori is available in GenBank and the EST database of B. mori is expanding, identification of novel genes of B. mori was conceivable by data-mining techniques and bioinformatics tools. In this study, we used the in silico cloning method to identify eukaryotic initiation factor 5A (eIF5A) gene in B. mori. With the hypusine formation, eIF5A is involved in the regulation of cell proliferation and apoptosis. Using the computer program MEGA3, we conducted a search for homologs of eIF5A among many eukaryotic species and confirmed that the eIF5A was conserved in all organisms investigated. This gene has been registered in GenBank under the accession number DQ104412.

Amino Acid Sequence↗

PAINT: a promoter analysis and interaction network generation tool for gene regulatory network identification.

We have developed a bioinformatics tool named PAINT that automates the promoter analysis of a given set of genes for the presence of transcription factor binding sites. Based on coincidence of regulatory sites, this tool produces an interaction matrix that represents a candidate transcriptional regulatory network. This tool currently consists of (1) a database of promoter sequences of known or predicted genes in the Ensembl annotated mouse genome database, (2) various modules that can retrieve and process the promoter sequences for binding sites of known transcription factors, and (3) modules for visualization and analysis of the resulting set of candidate network connections. This information provides a substantially pruned list of genes and transcription factors that can be examined in detail in further experimental studies on gene regulation. Also, the candidate network can be incorporated into network identification methods in the form of constraints on feasible structures in order to render the algorithms tractable for large-scale systems. The tool can also produce output in various formats suitable for use in external visualization and analysis software. In this manuscript, PAINT is demonstrated in two case studies involving analysis of differentially regulated genes chosen from two microarray data sets. The first set is from a neuroblastoma N1E-115 cell differentiation experiment, and the second set is from neuroblastoma N1E-115 cells at different time intervals following exposure to neuropeptide angiotensin II. PAINT is available for use as an agent in BioSPICE simulation and analysis framework (www.biospice.org), and can also be accessed via a WWW interface at www.dbi.tju.edu/dbi/tools/paint/.

Angiotensin II↗

Taverna: a tool for the composition and enactment of bioinformatics workflows.

MOTIVATION: In silico experiments in bioinformatics involve the co-ordinated use of computational tools and information repositories. A growing number of these resources are being made available with programmatic access in the form of Web services. Bioinformatics scientists will need to orchestrate these Web services in workflows as part of their analyses. RESULTS: The Taverna project has developed a tool for the composition and enactment of bioinformatics workflows for the life sciences community. The tool includes a workbench application which provides a graphical user interface for the composition of workflows. These workflows are written in a new language called the simple conceptual unified flow language (Scufl), where by each step within a workflow represents one atomic task. Two examples are used to illustrate the ease by which in silico experiments can be represented as Scufl workflows using the workbench application.

Computational Biology↗

BPROMPT: A consensus server for membrane protein prediction.

Protein structure prediction is a cornerstone of bioinformatics research. Membrane proteins require their own prediction methods due to their intrinsically different composition. A variety of tools exist for topology prediction of membrane proteins, many of them available on the Internet. The server described in this paper, BPROMPT (Bayesian PRediction Of Membrane Protein Topology), uses a Bayesian Belief Network to combine the results of other prediction methods, providing a more accurate consensus prediction. Topology predictions with accuracies of 70% for prokaryotes and 53% for eukaryotes were achieved. BPROMPT can be accessed at http://www.jenner.ac.uk/BPROMPT.

Bayes Theorem↗

CryptoDB: a Cryptosporidium bioinformatics resource update.

The database, CryptoDB (http://CryptoDB.org), is a community bioinformatics resource for the AIDS-related apicomplexan-parasite, Cryptosporidium. CryptoDB integrates whole genome sequence and annotation with expressed sequence tag and genome survey sequence data and provides supplemental bioinformatics analyses and data-mining tools. A simple, yet comprehensive web interface is available for mining and visualizing the data. CryptoDB is allied with the databases PlasmoDB and ToxoDB via ApiDB, an NIH/NIAID-fundedBioinformatics Resource Center. Recent updates to CryptoDB include the deposition of annotated genome sequences for Cryptosporidium parvum and Cryptosporidium hominis, migration to a relational database (GUS), a new query and visualization interface and the introduction of Web services.

Animals↗

BiologicalNetworks: visualization and analysis tool for systems biology.

Systems level investigation of genomic scale information requires the development of truly integrated databases dealing with heterogeneous data, which can be queried for simple properties of genes or other database objects as well as for complex network level properties, for the analysis and modelling of complex biological processes. Towards that goal, we recently constructed PathSys, a data integration platform for systems biology, which provides dynamic integration over a diverse set of databases [Baitaluk et al. (2006) BMC Bioinformatics 7, 55]. Here we describe a server, BiologicalNetworks, which provides visualization, analysis services and an information management framework over PathSys. The server allows easy retrieval, construction and visualization of complex biological networks, including genome-scale integrated networks of protein-protein, protein-DNA and genetic interactions. Most importantly, BiologicalNetworks addresses the need for systematic presentation and analysis of high-throughput expression data by mapping and analysis of expression profiles of genes or proteins simultaneously on to regulatory, metabolic and cellular networks. BiologicalNetworks Server is available at http://brak.sdsc.edu/pub/BiologicalNetworks.

Computer Graphics↗