Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Gene Class expression: analysis tool of Gene Ontology terms with gene expression data.

Serial analysis of gene expression (SAGE) technology produces large sets of interesting genes that are difficult to analyze directly. Bioinformatics tools are needed to interpret the functional information in these gene sets. We present an interactive web-based tool, called Gene Class, which allows functional annotation of SAGE data using the Gene Ontology (GO) database. This tool performs searches in the GO database for each SAGE tag, making associations in the selected GO category for a level selected in the hierarchy. This system provides user-friendly data navigation and visualization for mapping SAGE data onto the gene ontology structure. This tool also provides graphical visualization of the percentage of SAGE tags in each GO category, along with confidence intervals and hypothesis testing.

Animals↗

Bioinformatics approaches and resources for single nucleotide polymorphism functional analysis.

Since the initial sequencing of the human genome, many projects are underway to understand the effects of genetic variation between individuals. Predicting and understanding the downstream effects of genetic variation using computational methods are becoming increasingly important for single nucleotide polymorphism (SNP) selection in genetics studies and understanding the molecular basis of disease. According to the NIH, there are now more than four million validated SNPs in the human genome. The volume of known genetic variations lends itself well to an informatics approach. Bioinformaticians have become very good at functional inference methods derived from functional and structural genomics. This review will present a broad overview of the tools and resources available to collect and understand functional variation from the perspective of structure, expression, evolution and phenotype. Additionally, public resources available for SNP identification and characterisation are summarised.

Algorithms↗

Approaches to the automatic discovery of patterns in biosequences.

This paper surveys approaches to the discovery of patterns in biosequences and places these approaches within a formal framework that systematises the types of patterns and the discovery algorithms. Patterns with expressive power in the class of regular languages are considered, and a classification of pattern languages in this class is developed, covering the patterns that are the most frequently used in molecular bioinformatics. A formulation is given of the problem of the automatic discovery of such patterns from a set of sequences, and an analysis is presented of the ways in which an assessment can be made of the significance of the discovered patterns. It is shown that the problem is related to problems studied in the field of machine learning. The major part of this paper comprises a review of a number of existing methods developed to solve the problem and how these relate to each other, focusing on the algorithms underlying the approaches. A comparison is given of the algorithms, and examples are given of patterns that have been discovered using the different methods.

Algorithms↗

'Harvester': a fast meta search engine of human protein resources.

SUMMARY: We have developed a Web-based tool named 'Harvester' that bulk-collects bioinformatic data on human proteins from various databases and prediction servers. The information on every single protein is assembled on a single HTML page as a combination of database screen-shots and plain text. A full text meta search engine, similar to Google trade mark, allows screening of the whole genome proteome for current protein functions and predictions in a few seconds. With Harvester it is now possible to compare and check the quality of different database entries and prediction algorithms on a single page. A feedback forum allows users to comment on Harvester and to report database inconsistencies. AVAILABILITY: The service is freely available to the academic community at http://harvester.embl.de.

Database Management Systems↗

The NeuARt II system: a viewing tool for neuroanatomical data based on published neuroanatomical atlases.

BACKGROUND: Anatomical studies of neural circuitry describing the basic wiring diagram of the brain produce intrinsically spatial, highly complex data of great value to the neuroscience community. Published neuroanatomical atlases provide a spatial framework for these studies. We have built an informatics framework based on these atlases for the representation of neuroanatomical knowledge. This framework not only captures current methods of anatomical data acquisition and analysis, it allows these studies to be collated, compared and synthesized within a single system. RESULTS: We have developed an atlas-viewing application ('NeuARt II') in the Java language with unique functional properties. These include the ability to use copyrighted atlases as templates within which users may view, save and retrieve data-maps and annotate them with volumetric delineations. NeuARt II also permits users to view multiple levels on multiple atlases at once. Each data-map in this system is simply a stack of vector images with one image per atlas level, so any set of accurate drawings made onto a supported atlas (in vector graphics format) could be uploaded into NeuARt II. Presently the database is populated with a corpus of high-quality neuroanatomical data from the laboratory of Dr Larry Swanson (consisting 64 highly-detailed maps of PHAL tract-tracing experiments, made up of 1039 separate drawings that were published in 27 primary research publications over 17 years). Herein we take selective examples from these data to demonstrate the features of NeuArt II. Our informatics tool permits users to browse, query and compare these maps. The NeuARt II tool operates within a bioinformatics knowledge management platform (called 'NeuroScholar') either as a standalone or a plug-in application. CONCLUSION: Anatomical localization is fundamental to neuroscientific work and atlases provide an easily-understood framework that is widely used by neuroanatomists and non-neuroanatomists alike. NeuARt II, the neuroinformatics tool presented here, provides an accurate and powerful way of representing neuroanatomical data in the context of commonly-used brain atlases for visualization, comparison and analysis. Furthermore, it provides a framework that supports the delivery and manipulation of mapped data either as a standalone system or as a component in a larger knowledge management system.

Anatomy, Artistic↗

Docking to single-domain and multiple-domain proteins: old and new challenges.

The diverse selection of targets in the CAPRI experiments provides grounds for determining the limits of our rigid-body docking program MolFit, and for extending it. We find that the sensitivity of MolFit is high, enabling it to produce reasonably accurate docking solutions when the structures undergo moderate local conformation changes upon complex formation or when the docked molecules are modeled. Yet the ranks of these solutions are sometimes too low to meet the requirements of CAPRI assessment. This indicates that the selectivity of MolFit, which was optimized for docking of unbound X-ray structures, and which relies on the availability of external data from biochemical and bioinformatic sources, needs readjustment in order to meet the challenges presented by NMR or modeled structures. A different challenge is presented by large global conformation changes such as movements of domains. We show that such changes can be accommodated within the rigid-body approximation by employing rigid multibody multistage docking procedures. We also address the difficulty of ranking results from 2-body and multibody docking scans in cases in which there are no external data favoring one option over the other.

Algorithms↗

A high level interface to SCOP and ASTRAL implemented in python.

BACKGROUND: Benchmarking algorithms in structural bioinformatics often involves the construction of datasets of proteins with given sequence and structural properties. The SCOP database is a manually curated structural classification which groups together proteins on the basis of structural similarity. The ASTRAL compendium provides non redundant subsets of SCOP domains on the basis of sequence similarity such that no two domains in a given subset share more than a defined degree of sequence similarity. Taken together these two resources provide a 'ground truth' for assessing structural bioinformatics algorithms. We present a small and easy to use API written in python to enable construction of datasets from these resources. RESULTS: We have designed a set of python modules to provide an abstraction of the SCOP and ASTRAL databases. The modules are designed to work as part of the Biopython distribution. Python users can now manipulate and use the SCOP hierarchy from within python programs, and use ASTRAL to return sequences of domains in SCOP, as well as clustered representations of SCOP from ASTRAL. CONCLUSION: The modules make the analysis and generation of datasets for use in structural genomics easier and more principled.

Database Management Systems↗

Detecting horizontal gene transfer with T-REX and RHOM programs.

As the Human Genome Project and other genome projects experience remarkable success and a flood of biological data is produced by means of high-throughout sequencing techniques, detection of horizontal gene transfer (HGT) becomes a promising field in bioinformatics. This review describes two freeware programs: T-REX for MS Windows and RHOM for Linux. T-REX is a graphical user interface program that offers functions to reconstruct the HGT network among the donor and receptor hosts from the gene and species distance matrices. RHOM is a set of command-line driven programs used to detect HGT in genomes. While T-REX impresses with a user-friendly interface and drawing of the reticulation network, the strength of RHOM is an extensive statistical framework of genome and the graphical display of the estimated sequence position probabilities for the candidate horizontally transferred genes.

Algorithms↗

PDB file parser and structure class implemented in Python.

UNLABELLED: The biopython project provides a set of bioinformatics tools implemented in Python. Recently, biopython was extended with a set of modules that deal with macromolecular structure. Biopython now contains a parser for PDB files that makes the atomic information available in an easy-to-use but powerful data structure. The parser and data structure deal with features that are often left out or handled inadequately by other packages, e.g. atom and residue disorder (if point mutants are present in the crystal), anisotropic B factors, multiple models and insertion codes. In addition, the parser performs some sanity checking to detect obvious errors. AVAILABILITY: The Biopython distribution (including source code and documentation) is freely available (under the Biopython license) from http://www.biopython.org

Computer Simulation↗

BioAfrica's HIV-1 proteomics resource: combining protein data with bioinformatics tools.

Most Internet online resources for investigating HIV biology contain either bioinformatics tools, protein information or sequence data. The objective of this study was to develop a comprehensive online proteomics resource that integrates bioinformatics with the latest information on HIV-1 protein structure, gene expression, post-transcriptional/post-translational modification, functional activity, and protein-macromolecule interactions. The BioAfrica HIV-1 Proteomics Resource http://bioafrica.mrc.ac.za/proteomics/index.html is a website that contains detailed information about the HIV-1 proteome and protease cleavage sites, as well as data-mining tools that can be used to manipulate and query protein sequence data, a BLAST tool for initiating structural analyses of HIV-1 proteins, and a proteomics tools directory. The Proteome section contains extensive data on each of 19 HIV-1 proteins, including their functional properties, a sample analysis of HIV-1HXB2, structural models and links to other online resources. The HIV-1 Protease Cleavage Sites section provides information on the position, subtype variation and genetic evolution of Gag, Gag-Pol and Nef cleavage sites. The HIV-1 Protein Data-mining Tool includes a set of 27 group M (subtypes A through K) reference sequences that can be used to assess the influence of genetic variation on immunological and functional domains of the protein. The BLAST Structure Tool identifies proteins with similar, experimentally determined topologies, and the Tools Directory provides a categorized list of websites and relevant software programs. This combined database and software repository is designed to facilitate the capture, retrieval and analysis of HIV-1 protein data, and to convert it into clinically useful information relating to the pathogenesis, transmission and therapeutic response of different HIV-1 variants. The HIV-1 Proteomics Resource is readily accessible through the BioAfrica website at: http://bioafrica.mrc.ac.za/proteomics/index.html.

Africa↗

A novel secreted protein toxin from the insect pathogenic bacterium Xenorhabdus nematophila.

The bacterium Xenorhabdus nematophila is an insect pathogen that produces several proteins that enable it to kill insects. Screening of a cosmid library constructed from X. nematophila strain A24 identified a gene that encoded a novel protein that was toxic to insects. The 42-kDa protein encoded by the toxin gene was expressed and purified from a recombinant system, and was shown to kill the larvae of insects such as Galleria mellonella and Helicoverpa armigera when injected at doses of around 30-40 ng/g larvae. Sequencing and bioinformatic analysis suggested that the toxin was a novel protein, and that it was likely to be part of a genomic island involved in pathogenicity. When the native bacteria were grown under laboratory conditions, a soluble form of the 42-kDa toxin was secreted only by bacteria in the phase II state. Preliminary histological analysis of larvae injected with recombinant protein suggested that the toxin primarily acted on the midgut of the insect. Finally, some of the common strategies used by the bacterial pathogens of insects, animals, and plants are discussed.

Amino Acid Sequence↗

Implicit motif distribution based hybrid computational kernel for sequence classification.

MOTIVATION: We designed a general computational kernel for classification problems that require specific motif extraction and search from sequences. Instead of searching for explicit motifs, our approach finds the distribution of implicit motifs and uses as a feature for classification. Implicit motif distribution approach may be used as modus operandi for bioinformatics problems that require specific motif extraction and search, which is otherwise computationally prohibitive. RESULTS: A system named P2SL that infer protein subcellular targeting was developed through this computational kernel. Targeting-signal was modeled by the distribution of subsequence occurrences (implicit motifs) using self-organizing maps. The boundaries among the classes were then determined with a set of support vector machines. P2SL hybrid computational system achieved approximately 81% of prediction accuracy rate over ER targeted, cytosolic, mitochondrial and nuclear protein localization classes. P2SL additionally offers the distribution potential of proteins among localization classes, which is particularly important for proteins, shuttle between nucleus and cytosol. AVAILABILITY: http://staff.vbi.vt.edu/volkan/p2sl and http://www.i-cancer.fen.bilkent.edu.tr/p2sl CONTACT: rengul@bilkent.edu.tr.

Algorithms↗

The genome revolution in vaccine research.

The conventional approach to vaccine development is based on dissection of the pathogen using biochemical, immunological and microbiological methods. Although successful in several cases, this approach has failed to provide a solution to prevent several major bacterial infections. The availability of complete genome sequences in combination with novel advanced technologies, such as bioinformatics, microarrays and proteomics, have revolutionized the approach to vaccine development and provided a new impulse to microbial research. The genomic revolution allows the design of vaccines starting from the prediction of all antigens in silico, independently of their abundance and without the need to grow the pathogen in vitro. This new genome-based approach, which we have named "Reverse Vaccinology", has been successfully applied for Neisseria meningitidis serogroup B for which conventional strategies have failed to provide an efficacious vaccine. The concept of "Reverse Vaccinology" can be easily applied to all the pathogens for which vaccines are not yet available and can be extended to parasites and viruses.

Bacterial Proteins↗

Tolerating some redundancy significantly speeds up clustering of large protein databases.

MOTIVATION: Sequence clustering replaces groups of similar sequences in a database with single representatives. Clustering large protein databases like the NCBI Non-Redundant database (NR) using even the best currently available clustering algorithms is very time-consuming and only practical at relatively high sequence identity thresholds. Our previous program, CD-HI, clustered NR at 90% identity in approximately 1 h and at 75% identity in approximately 1 day on a 1 GHz Linux PC (Li et al., Bioinformatics, 17, 282, 2001); however even faster clustering speed is needed because the size of protein databases are rapidly growing and many applications desire a lower attainable thresholds. RESULTS: For our previous algorithm (CD-HI), we have employed short-word filters to speed up the clustering. In this paper, we show that tolerating some redundancy makes for more efficient use of these short-word filters and increases the program's speed 100 times. Our new program implements this technique and clusters NR at 70% identity within 2 h, and at 50% identity in approximately 5 days. Although some redundancy is present after clustering, our new program's results only differ from our previous program's by less than 0.4%.

Algorithms↗

qcCHIP: an R package to identify clonal hematopoiesis variants using cohort-specific data characteristics.

SUMMARY: Clonal hematopoiesis (CH) is a molecular biomarker associated with various adverse outcomes in both healthy individuals and those with underlying conditions, including cancer. Detecting CH usually involves genomic sequencing of individual blood samples followed by robust bioinformatics data filtering. We report an R package, qcCHIP, a bioinformatics pipeline that implements permutation-based parameter optimization to guide quality control filtering and cohort-specific CH identification. We benchmark qcCHIP under various data settings, including different sequencing depths, ranges of cohort sizes, with and without normal-tumor paired samples, and across different cancer types. We show that qcCHIP allows users to customize analysis needs to generate CH calls based on cohort-specific data characteristics. AVAILABILITY AND IMPLEMENTATION: qcCHIP R package is freely accessible at GitHub https://github.com/tenglab/qcCHIP and DOI: 10.5281/zenodo.16421861.

Humans↗

PoPS: a computational tool for modeling and predicting protease specificity.

Proteases play a fundamental role in the control of intra- and extra-cellular processes by binding and cleaving specific amino acid sequences. Identifying these targets is extremely challenging. Current computational attempts to predict cleavage sites are limited, representing these amino acid sequences as patterns or frequency matrices. Here we present PoPS, a publicly accessible bioinformatics tool (http://pops.csse.monash.edu.au/) that provides a novel method for building computational models of protease specificity, which while still being based on these amino acid sequences, can be built from any experimental data or expert knowledge available to the user. PoPS specificity models can be used to predict and rank likely cleavages within a single substrate, and within entire proteomes. Other factors, such as the secondary or tertiary structure of the substrate, can be used to screen unlikely sites. Furthermore, the tool also provides facilities to infer, compare and test models, and to store them in a publicly accessible database.

Algorithms↗

The EMBL Nucleotide Sequence Database.

The EMBL Nucleotide Sequence Database (http://www.ebi.ac.uk/embl.html) constitutes Europe's primary nucleotide sequence resource. Main sources for DNA and RNA sequences are direct submissions from individual researchers, genome sequencing projects and patent applications. While automatic procedures allow incorporation of sequence data from large-scale genome sequencing centres and from the European Patent Office (EPO), the preferred submission tool for individual submitters is Webin (WWW). Through all stages, dataflow is monitored by EBI biologists communicating with the sequencing groups. In collaboration with DDBJ and GenBank the database is produced, maintained and distributed at the European Bioinformatics Institute (EBI). Database releases are produced quarterly and are distributed on CD-ROM. Network services allow access to the most up-to-date data collection via Internet and World Wide Web interface. EBI's Sequence Retrieval System (SRS) is a Network Browser for Databanks in Molecular Biology, integrating and linking the main nucleotide and protein databases, plus many specialised databases. For sequence similarity searching a variety of tools (e.g. Blitz, Fasta, Blast etc) are available for external users to compare their own sequences against the most currently available data in the EMBL Nucleotide Sequence Database and SWISS-PROT.

Amino Acid Sequence↗

ELM server: A new resource for investigating short functional sites in modular eukaryotic proteins.

Multidomain proteins predominate in eukaryotic proteomes. Individual functions assigned to different sequence segments combine to create a complex function for the whole protein. While on-line resources are available for revealing globular domains in sequences, there has hitherto been no comprehensive collection of small functional sites/motifs comparable to the globular domain resources, yet these are as important for the function of multidomain proteins. Short linear peptide motifs are used for cell compartment targeting, protein-protein interaction, regulation by phosphorylation, acetylation, glycosylation and a host of other post-translational modifications. ELM, the Eukaryotic Linear Motif server at http://elm.eu.org/, is a new bioinformatics resource for investigating candidate short non-globular functional motifs in eukaryotic proteins, aiming to fill the void in bioinformatics tools. Sequence comparisons with short motifs are difficult to evaluate because the usual significance assessments are inappropriate. Therefore the server is implemented with several logical filters to eliminate false positives. Current filters are for cell compartment, globular domain clash and taxonomic range. In favourable cases, the filters can reduce the number of retained matches by an order of magnitude or more.

Amino Acid Motifs↗