Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Systematically investigating and identifying bacteriocins in the human gut microbiome.

Human gut microbiota produces unmodified bacteriocins, natural antimicrobial peptides that protect against pathogens and regulate host physiology. However, current bioinformatic tools limit the comprehensive investigation of bacteriocins' biosynthesis, obstructing research into their biological functions. Here, we introduce IIBacFinder, a superior analysis pipeline for identifying unmodified class II bacteriocins. Through large-scale bioinformatic analysis and experimental validation, we demonstrate their widespread distribution across the bacterial kingdom, with most being habitat specific. Analyzing over 280,000 bacterial genomes, we reveal the diverse potential of human gut bacteria to produce these bacteriocins. Guided by meta-omics analysis, we synthesized 26 hypothetical bacteriocins from gut commensal species, with 16 showing antibacterial activities. Further ex vivo tests show minimal impact of narrow-spectrum bacteriocins on human fecal microbiota. Our study highlights the huge biosynthetic potential of unmodified bacteriocins in the human gut, paving the way for understanding their biological functions and health implications.

Humans↗

Energy balance for analysis of complex metabolic networks.

Predicting behavior of large-scale biochemical networks represents one of the greatest challenges of bioinformatics and computational biology. Computational tools for predicting fluxes in biochemical networks are applied in the fields of integrated and systems biology, bioinformatics, and genomics, and to aid in drug discovery and identification of potential drug targets. Approaches, such as flux balance analysis (FBA), that account for the known stoichiometry of the reaction network while avoiding implementation of detailed reaction kinetics are promising tools for the analysis of large complex networks. Here we introduce energy balance analysis (EBA)--the theory and methodology for enforcing the laws of thermodynamics in such simulations--making the results more physically realistic and revealing greater insight into the regulatory and control mechanisms operating in complex large-scale systems. We show that EBA eliminates thermodynamically infeasible results associated with FBA.

Biophysics↗

Light-weight integration of molecular biological databases.

MOTIVATION: Due to the increasing number of molecular biological databases and the exponential growth of their contents, database integration is an important topic of research in bioinformatics. Existing approaches in this area have in common that considerable efforts are needed to provide integrated access to heterogeneous data sources. RESULTS: This article describes the LIMBO architecture as a light-weight approach to molecular biological database integration. By building systems upon this architecture, the efforts needed for database integration can be significantly lowered. AVAILABILITY: As an illustration of the principle usefulness of the underlying ideas, a prototypical implementation based upon the LIMBO architecture is described. This implementation is exclusively based on freely available open source components like the PostgreSQL database management system and the BioRuby project. Additional files and modified components are available upon request from the author.

Computational Biology↗

Predicting reliable regions in protein sequence alignments.

MOTIVATION: Protein sequence alignments have a myriad of applications in bioinformatics, including secondary and tertiary structure prediction, homology modeling, and phylogeny. Unfortunately, all alignment methods make mistakes, and mistakes in alignments often yield mistakes in their application. Thus, a method to identify and remove suspect alignment positions could benefit many areas in protein sequence analysis. RESULTS: We tested four predictors of alignment position reliability, including near-optimal alignment information, column score, and secondary structural information. We validated each predictor against a large library of alignments, removing positions predicted as unreliable. Near-optimal alignment information was the best predictor, removing 70% of the substantially-misaligned positions and 58% of the over-aligned positions, while retaining 86% of those aligned accurately.

Algorithms↗

The Bioperl toolkit: Perl modules for the life sciences.

The Bioperl project is an international open-source collaboration of biologists, bioinformaticians, and computer scientists that has evolved over the past 7 yr into the most comprehensive library of Perl modules available for managing and manipulating life-science information. Bioperl provides an easy-to-use, stable, and consistent programming interface for bioinformatics application programmers. The Bioperl modules have been successfully and repeatedly used to reduce otherwise complex tasks to only a few lines of code. The Bioperl object model has been proven to be flexible enough to support enterprise-level applications such as EnsEMBL, while maintaining an easy learning curve for novice Perl programmers. Bioperl is capable of executing analyses and processing results from programs such as BLAST, ClustalW, or the EMBOSS suite. Interoperation with modules written in Python and Java is supported through the evolving BioCORBA bridge. Bioperl provides access to data stores such as GenBank and SwissProt via a flexible series of sequence input/output modules, and to the emerging common sequence data storage format of the Open Bioinformatics Database Access project. This study describes the overall architecture of the toolkit, the problem domains that it addresses, and gives specific examples of how the toolkit can be used to solve common life-sciences problems. We conclude with a discussion of how the open-source nature of the project has contributed to the development effort.

Algorithms↗

BioMOBY successfully integrates distributed heterogeneous bioinformatics Web Services. The PlaNet exemplar case.

The burden of non-interoperability between on-line genomic resources is increasingly the rate-limiting step in large-scale genomic analysis. BioMOBY is a biological Web Service interoperability initiative that began as a retreat of representatives from the model organism database community in September, 2001. Its long-term goal is to provide a simple, extensible platform through which the myriad of on-line biological databases and analytical tools can offer their information and analytical services in a fully automated and interoperable way. Of the two branches of the larger BioMOBY project, the Web Services branch (MOBY-S) has now been deployed over several dozen data sources worldwide, revealing some significant observations about the nature of the integrative biology problem; in particular, that Web Service interoperability in the domain of bioinformatics is, unexpectedly, largely a syntactic rather than a semantic problem. That is to say, interoperability between bioinformatics Web Services can be largely achieved simply by specifying the data structures being passed between the services (syntax) even without rich specification of what those data structures mean (semantics). Thus, one barrier of the integrative problem has been overcome with a surprisingly simple solution. Here, we present a non-technical overview of the critical components that give rise to the interoperable behaviors seen in MOBY-S and discuss an exemplar case, the PlaNet consortium, where MOBY-S has been deployed to integrate the on-line plant genome databases and analytical services provided by a European consortium of databases and data service providers.

Computational Biology↗

Bioinformatic insights from metagenomics through visualization.

Cutting-edge biological and bioinformatics research seeks a systems perspective through the analysis of multiple types of high-throughput and other experimental data for the same sample. Systems-level analysis requires the integration and fusion of such data, typically through advanced statistics and mathematics. Visualization is a complementary computational approach that supports integration and analysis of complex data or its derivatives. We present a bioinformatics visualization prototype, Juxter, which depicts categorical information derived from or assigned to these diverse data for the purpose of comparing patterns across categorizations. The visualization allows users to easily discern correlated and anomalous patterns in the data. These patterns, which might not be detected automatically by algorithms, may reveal valuable information leading to insight and discovery. We describe the visualization and interaction capabilities and demonstrate its utility in a new field, metagenomics, which combines molecular biology and genetics to identify and characterize genetic material from multi-species microbial samples.

Algorithms↗

AmpSeqR: an R package for amplicon deep sequencing data analysis.

Amplicon sequencing (AmpSeq) is a methodology that targets specific genomic regions of interest for polymerase chain reaction (PCR) amplification so that they can be sequenced to a high depth of coverage. Amplicons are typically chosen to be highly polymorphic, usually with several highly informative, high frequency single nucleotide polymorphisms (SNPs) segregating in an amplicon of 100-200 base pair (bp). This allows high sensitivity detection and quantification of the frequency of each sequence within each sample making it suitable for applications such as low frequency somatic mosaicism detection or minor clone detection in mixed samples. AmpSeq is being increasingly applied to both biological and medical studies, in applications such as cancer, infectious diseases and brain mosaicism studies. Current bioinformatics pipelines for AmpSeq data processing lack downstream analysis, have difficulty distinguishing between true sequences and PCR sequencing errors and artifacts, and often require bioinformatic expertise. We present a new R package: AmpSeqR, designed for the processing of deep short-read amplicon sequencing data, with a focus on infectious diseases. The pipeline integrates several existing R packages combining them with newly developed functions to perform optimal filtering of reads to remove noise and improve the accuracy of the detected sequences data, permitting detection of very low frequency clones in mixed samples. The package provides useful functions including data pre-processing, amplicon sequence variants (ASVs) estimation, data post-processing, data visualization, and automatically generates a comprehensive Rmarkdown report that contains all essential results facilitating easy inclusion into reports and publications. AmpSeqR is publicly available at https://github.com/bahlolab/AmpSeqR.

High-Throughput Nucleotide Sequencing↗

A novel method for protein secondary structure prediction using dual-layer SVM and profiles.

A high-performance method was developed for protein secondary structure prediction based on the dual-layer support vector machine (SVM) and position-specific scoring matrices (PSSMs). SVM is a new machine learning technology that has been successfully applied in solving problems in the field of bioinformatics. The SVM's performance is usually better than that of traditional machine learning approaches. The performance was further improved by combining PSSM profiles with the SVM analysis. The PSSMs were generated from PSI-BLAST profiles, which contain important evolution information. The final prediction results were generated from the second SVM layer output. On the CB513 data set, the three-state overall per-residue accuracy, Q3, reached 75.2%, while segment overlap (SOV) accuracy increased to 80.0%. On the CB396 data set, the Q3 of our method reached 74.0% and the SOV reached 78.1%. A web server utilizing the method has been constructed and is available at http://www.bioinfo.tsinghua.edu.cn/pmsvm.

Computational Biology↗

Assigning new GO annotations to protein data bank sequences by combining structure and sequence homology.

Accompanying the discovery of an increasing number of proteins, there is the need to provide functional annotation that is both highly accurate and consistent. The Gene Ontology (GO) provides consistent annotation in a computer readable and usable form; hence, GO annotation (GOA) has been assigned to a large number of protein sequences based on direct experimental evidence and through inference determined by sequence homology. Here we show that this annotation can be extended and corrected for cases where protein structures are available. Specifically, using the Combinatorial Extension (CE) algorithm for structure comparison, we extend the protein annotation currently provided by GOA at the European Bioinformatics Institute (EBI) to further describe the contents of the Protein Data Bank (PDB). Specific cases of biologically interesting annotations derived by this method are given. Given that the relationship between sequence, structure, and function is complicated, we explore the impact of this relationship on assigning GOA. The effect of superfolds (folds with many functions) is considered and, by comparison to the Structural Classification of Proteins (SCOP), the individual effects of family, superfamily, and fold.

Algorithms↗

PROFILER: a tool for automatic searching of internally maintained databases.

A new application program, the BIOINFORMATICS PROFILER, is described which simplifies the analysis of new genome sequence information appearing on a daily basis. Control tasks are defined through an intuitive graphical user interface and are executed at user defined nightly intervals. Electronic mail is sent to indicate that search results have been found. All task output is presented to users in the form of hypertext (HTML), allowing easy browsing. Currently supported tasks include BLAST and FastA sequence searching, keyword based searching of network news articles and WAIS databases, examination of GenBank sequence entries using regular expressions and boolean operations and protein sequence motif searching.

Amino Acid Sequence↗

Target gene identification from expression array data by promoter analysis.

DNA microchips and expression arrays yield enormous amounts of data linking cDNA sequences to gene expression patterns. This now allows the characterization of gene expression in normal and diseased tissues as well as the response of tissues to the application of therapeutic reagents. Software currently exists to analyze DNA array/chip data with respect to corresponding mRNA sequences, which facilitates the precise determination of when and where certain groups of genes are expressed. The information concerning transcriptional regulatory networks responsible for the observed expression patterns is not contained within the cDNA sequences used to generate the arrays, but resides often within the promoter sequences of the individual genes (and/or enhancers). The complete sequence of the human genome will provide the molecular basis for the identification of such regulatory regions. Promoter sequences for specific cDNAs can be obtained reliably from genomic sequences simply by exon mapping. Promoter prediction tools can also be used to locate promoters directly in the genomic sequence in many cases in which cDNAs are 5'-incomplete. Once sufficient numbers of promoter sequences have been obtained, the comparative promoter analysis of the co-regulated genes and groups of genes can be applied in order to generate models describing the higher order levels of the transcription factor binding site organization within these promoter regions. As evident from several examples, this approach can identify promoter modules responsible for the common regulation of promoters solely by the application of bioinformatics methods. Such modules represent the molecular mechanisms through which regulatory networks influence gene expression. Another advantage of this approach is that it also provides a powerful alternative for elucidating functional features of genes with no detectable sequence similarity, by linking them to other genes on the basis of their common promoter structures.

Computational Biology↗

GPAC: benchmarking the sensitivity of genome informatics analysis to genome annotation completeness.

In view of the recent explosion in genome sequence data, and the 200 or more complete genome sequences currently available, the importance of genome-scale bioinformatics analysis is increasing rapidly. However, computational genome informatics analyses often lack a statistical assessment of their sensitivity to the completeness of the functional annotation. Therefore, a pre-analysis method to automatically validate the sensitivity of computational genome analyses with regard to genome annotation completeness is useful for this purpose. In this report we developed the Gene Prediction Accuracy Classification (GPAC) test, which provides statistical evidence of sensitivity by repeating the same analysis for five different gene groups (classified according to annotation accuracy level), and for randomly sampled gene groups, with the same number of genes as each of the five classified groups. Variability in these results is then assessed, and if the results vary significantly with different data subsets, the analysis is considered "sensitive" to annotation completeness, and careful selection of data is advised prior to the actual in silico analysis. The GPAC test has been applied to the analyses of Sakai et al., 2001, and Ohno et al., 2001, and it revealed that the analysis of Ohno et al. was more sensitive to annotation completeness. It showed that GPAC could be employed to ascertain the sensitivity of an analysis. The GPAC bendhmarking software is freely available in the latest G-language Genome Analysis Environment package, at http://www.g-language.org/.

Benchmarking↗

Whole genome shotgun sequencing guided by bioinformatics pipelines--an optimized approach for an established technique.

While the sequencing of bacterial genomes has become a routine procedure at major sequencing centers, there are still a number of genome projects at small- or medium-size facilities. For these facilities a maximum of control over sequencing, assembling and finishing is essential. At the same time, facilities have to be able to co-operate at minimum costs for the overall project. We have established a pipeline for the distributed sequencing of Alcanivorax borkumensis SK2, Azoarcus sp. BH72, Clavibacter michiganensis subsp. michiganensis NCPPB382, Sorangium cellulosum So ce56 and Xanthomonas campestris pv. vesicatoria 85-10. Our pipeline relies on standard tools (e.g. PHRED/PHRAP, CAP3 and Consed/Autofinish) wherever possible, supplementing them with new tools (BioMake and BACCardI) to achieve the aims described above.

Algorithms↗

Agents in bioinformatics, computational and systems biology.

The adoption of agent technologies and multi-agent systems constitutes an emerging area in bioinformatics. In this article, we report on the activity of the Working Group on Agents in Bioinformatics (BIOAGENTS) founded during the first AgentLink III Technical Forum meeting on the 2nd of July, 2004, in Rome. The meeting provided an opportunity for seeding collaborations between the agent and bioinformatics communities to develop a different (agent-based) approach of computational frameworks both for data analysis and management in bioinformatics and for systems modelling and simulation in computational and systems biology. The collaborations gave rise to applications and integrated tools that we summarize and discuss in context of the state of the art in this area. We investigate on future challenges and argue that the field should still be explored from many perspectives ranging from bio-conceptual languages for agent-based simulation, to the definition of bio-ontology-based declarative languages to be used by information agents, and to the adoption of agents for computational grids.

Artificial Intelligence↗

Bioinformatics approach to predicting HIV drug resistance.

The emergence of drug resistance remains one of the most challenging issues in the treatment of HIV-1 infection. The extreme replication dynamics of HIV facilitates its escape from the selective pressure exerted by the human immune system and by the applied combination drug therapy. This article reviews computational methods whose combined use can support the design of optimal antiretroviral therapies based on viral genotypic and phenotypic data. Genotypic assays are based on the analysis of mutations associated with reduced drug susceptibility, but are difficult to interpret due to the numerous mutations and mutational patterns that confer drug resistance. Phenotypic resistance or susceptibility can be experimentally evaluated by measuring the inhibition of the viral replication in cell culture assays. However, this procedure is expensive and time consuming.

Anti-HIV Agents↗

Array of hope for gene technology.

A Washington-based bioinformatics company is developing sophisticated DNA microarrays that should help researchers measure and analyze gene expression faster, more economically, and with greater precision than ever before possible. The FlexJet system, as the microarray product is known, uses inkjet technology to propel microscopic strands of DNA nucleotides onto slides, "printing" arrays of DNA molecules in a process not unlike the manner in which a printer deposits ink onto paper, forming distinct patterns of characters and images. Microarray technology may revolutionize the field of toxicogenomics by helping scientists target new drugs, discover gene function, determine biologic pathways, and better understand diseases such as cancer, cystic fibrosis, and cardiovascular disease at the molecular level.

Computers↗

Mistaken identifiers: gene name errors can be introduced inadvertently when using Excel in bioinformatics.

BACKGROUND: When processing microarray data sets, we recently noticed that some gene names were being changed inadvertently to non-gene names. RESULTS: A little detective work traced the problem to default date format conversions and floating-point format conversions in the very useful Excel program package. The date conversions affect at least 30 gene names; the floating-point conversions affect at least 2,000 if Riken identifiers are included. These conversions are irreversible; the original gene names cannot be recovered. CONCLUSIONS: Users of Excel for analyses involving gene names should be aware of this problem, which can cause genes, including medically important ones, to be lost from view and which has contaminated even carefully curated public databases. We provide work-arounds and scripts for circumventing the problem.

Animals↗