Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Construction of a human glycogene library and comprehensive functional analysis.

Eighteen years have passed after the first mammalian glycosyltransferase was cloned. At the beginning of April, 2001, 110 genes for human glycosyltransferases, including modifying enzymes for carbohydrate chains such as sulfotransferases, had been cloned and analyzed. We started the Glycogene Project (GG project) in April 2001, a comprehensive study on human glycogenes with the aid of bioinformatic technology. The term glycogene includes the genes for glycosyltransferases, sulfotransferases adding sulfate to carbohydrates and sugar-nucleotide transporters, etc. Firstly, as many novel genes, which are the candidates for glycogenes, as possible were searched using bioinformatic technology in databases. They were then cloned and expressed in various expression systems to detect the activity for carbohydrate synthesis. Their substrate specificity was determined using various acceptors.

Animals↗

PACK: Profile Analysis using Clustering and Kurtosis to find molecular classifiers in cancer.

MOTIVATION: Elucidating the molecular taxonomy of cancers and finding biological and clinical markers from microarray experiments is problematic due to the large number of variables being measured. Feature selection methods that can identify relevant classifiers or that can remove likely false positives prior to supervised analysis are therefore desirable. RESULTS: We present a novel feature selection procedure based on a mixture model and a non-gaussianity measure of a gene's expression profile. The method can be used to find genes that define either small outlier subgroups or major subdivisions, depending on the sign of kurtosis. The method can also be used as a filtering step, prior to supervised analysis, in order to reduce the false discovery rate. We validate our methodology using six independent datasets by rediscovering major classifiers in ER negative and ER positive breast cancer and in prostate cancer. Furthermore, our method finds two novel subtypes within the basal subgroup of ER negative breast tumours, associated with apoptotic and immune response functions respectively, and with statistically different clinical outcome. AVAILABILITY: An R-function pack that implements the methods used here has been added to vabayelMix, available from (www.cran.r-project.org). CONTACT: aet21@cam.ac.uk SUPPLEMENTARY INFORMATION: Supplementary information is available at Bioinformatics online.

Algorithms↗

BioMoby extensions to the Taverna workflow management and enactment software.

BACKGROUND: As biology becomes an increasingly computational science, it is critical that we develop software tools that support not only bioinformaticians, but also bench biologists in their exploration of the vast and complex data-sets that continue to build from international genomic, proteomic, and systems-biology projects. The BioMoby interoperability system was created with the goal of facilitating the movement of data from one Web-based resource to another to fulfill the requirements of non-expert bioinformaticians. In parallel with the development of BioMoby, the European myGrid project was designing Taverna, a bioinformatics workflow design and enactment tool. Here we describe the marriage of these two projects in the form of a Taverna plug-in that provides access to many of BioMoby's features through the Taverna interface. RESULTS: The exposed BioMoby functionality aids in the design of "sensible" BioMoby workflows, aids in pipelining BioMoby and non-BioMoby-based resources, and ensures that end-users need only a minimal understanding of both BioMoby, and the Taverna interface itself. Users are guided through the construction of syntactically and semantically correct workflows through plug-in calls to the Moby Central registry. Moby Central provides a menu of only those BioMoby services capable of operating on the data-type(s) that exist at any given position in the workflow. Moreover, the plug-in automatically and correctly connects a selected service into the workflow such that users are not required to understand the nature of the inputs or outputs for any service, leaving them to focus on the biological meaning of the workflow they are constructing, rather than the technical details of how the services will interoperate. CONCLUSION: With the availability of the BioMoby plug-in to Taverna, we believe that BioMoby-based Web Services are now significantly more useful and accessible to bench scientists than are more traditional Web Services.

Biology↗

Ensembl 2002: accommodating comparative genomics.

The Ensembl (http://www.ensembl.org/) database project provides a bioinformatics framework to organise biology around the sequences of large genomes. It is a comprehensive source of stable automatic annotation of human, mouse and other genome sequences, available as either an interactive web site or as flat files. Ensembl also integrates manually annotated gene structures from external sources where available. As well as being one of the leading sources of genome annotation, Ensembl is an open source software engineering project to develop a portable system able to handle very large genomes and associated requirements. These range from sequence analysis to data storage and visualisation and installations exist around the world in both companies and at academic sites. With both human and mouse genome sequences available and more vertebrate sequences to follow, many of the recent developments in Ensembl have focusing on developing automatic comparative genome analysis and visualisation.

Animals↗

A database of unique protein sequence identifiers for proteome studies.

In proteome studies, identification of proteins requires searching protein sequence databases. The public protein sequence databases (e.g., NCBInr, UniProt) each contain millions of entries, and private databases add thousands more. Although much of the sequence information in these databases is redundant, each database uses distinct identifiers for the identical protein sequence and often contains unique annotation information. Users of one database obtain a database-specific sequence identifier that is often difficult to reconcile with the identifiers from a different database. When multiple databases are used for searches or the databases being searched are updated frequently, interpreting the protein identifications and associated annotations can be problematic. We have developed a database of unique protein sequence identifiers called Sequence Globally Unique Identifiers (SEGUID) derived from primary protein sequences. These identifiers serve as a common link between multiple sequence databases and are resilient to annotation changes in either public or private databases throughout the lifetime of a given protein sequence. The SEGUID Database can be downloaded (http://bioinformatics.anl.gov/SEGUID/) or easily generated at any site with access to primary protein sequence databases. Since SEGUIDs are stable, predictions based on the primary sequence information (e.g., pI, Mr) can be calculated just once; we have generated approximately 500 different calculations for more than 2.5 million sequences. SEGUIDs are used to integrate MS and 2-DE data with bioinformatics information and provide the opportunity to search multiple protein sequence databases, thereby providing a higher probability of finding the most valid protein identifications.

Amino Acid Sequence↗

PhyloNaP: a user-friendly database of phylogeny for natural product-producing enzymes.

SUMMARY: Phylogenetic analysis is widely used to predict enzyme function, yet building annotated and reusable trees is labor-intensive and requires extensive knowledge about the specific enzymes. Existing resources rarely cover biosynthetic enzymes and lack the context needed for meaningful analysis. We present PhyloNaP, the first large-scale resource dedicated to phylogenies of biosynthetic enzymes. PhyloNaP provides ∼51 000 annotated and interactive trees enriched with chemical, functional, and taxonomic information. Users can classify their own sequences via phylogenetic placement, enabling functional inference in an evolutionary context. A contribution portal allows the community to submit curated trees. By combining scale, breadth of annotation, and interactive functionality, PhyloNaP fills a major gap in bioinformatics resources for enzyme discovery and annotation, with immediate applications to secondary metabolism and beyond. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at https://phylonap.cs.uni-tuebingen.de.

Phylogeny↗

Genometrics as an essential tool for the assembly of whole genome sequences: the example of the chromosome of Bifidobacterium longum NCC2705.

BACKGROUND: Analysis of the first reported complete genome sequence of Bifidobacterium longum NCC2705, an actinobacterium colonizing the gastrointestinal tract, uncovered its proteomic relatedness to Streptomyces coelicolor and Mycobacterium tuberculosis. However, a rapid scrutiny by genometric methods revealed a genome organization totally different from all so far sequenced high-GC Gram-positive chromosomes. RESULTS: Generally, the cumulative GC- and ORF orientation skew curves of prokaryotic genomes consist of two linear segments of opposite slope: the minimum and the maximum of the curves correspond to the origin and the terminus of chromosome replication, respectively. However, analyses of the B. longum NCC2705 chromosome yielded six, instead of two, linear segments, while its dnaA locus, usually associated with the origin of replication, was not located at the minimum of the curves. Furthermore, the coorientation of gene transcription with replication was very low. Comparison with closely related actinobacteria strongly suggested that the chromosome of B. longum was misassembled, and the identification of two pairs of relatively long homologous DNA sequences offers the possibility for an alternative genome assembly proposed here below. By genometric criteria, this configuration displays all of the characters common to bacteria, in particular to related high-GC Gram-positives. In addition, it is compatible with the partially sequenced genome of DJO10A B. longum strain. Recently, a corrected sequence of B. longum NCC2705, with a configuration similar to the one proposed here below, has been deposited in GenBank, confirming our predictions. CONCLUSION: Genometric analyses, in conjunction with standard bioinformatic tools and knowledge of bacterial chromosome architecture, represent fast and straightforward methods for the evaluation of chromosome assembly.

Bifidobacterium↗

Interpretation of mass spectrometry data for high-throughput proteomics.

Recent developments in proteomics have revealed a bottleneck in bioinformatics: high-quality interpretation of acquired MS data. The ability to generate thousands of MS spectra per day, and the demand for this, makes manual methods inadequate for analysis and underlines the need to transfer the advanced capabilities of an expert human user into sophisticated MS interpretation algorithms. The identification rate in current high-throughput proteomics studies is not only a matter of instrumentation. We present software for high-throughput PMF identification, which enables robust and confident protein identification at higher rates. This has been achieved by automated calibration, peak rejection, and use of a meta search approach which employs various PMF search engines. The automatic calibration consists of a dynamic, spectral information-dependent algorithm, which combines various known calibration methods and iteratively establishes an optimised calibration. The peak rejection algorithm filters signals that are unrelated to the analysed protein by use of automatically generated and dataset-dependent exclusion lists. In the "meta search" several known PMF search engines are triggered and their results are merged by use of a meta score. The significance of the meta score was assessed by simulation of PMF identification with 10,000 artificial spectra resembling a data situation close to the measured dataset. By means of this simulation the meta score is linked to expectation values as a statistical measure. The presented software is part of the proteome database ProteinScape which links the information derived from MS data to other relevant proteomics data. We demonstrate the performance of the presented system with MS data from 1891 PMF spectra. As a result of automatic calibration and peak rejection the identification rate increased from 6% to 44%.

Algorithms↗

Comprehensive circRNA expression profile and hub genes screening during human liver development.

BACKGROUND: Understanding the expression of non-coding RNA in the liver during embryonic development provides important insights into liver diseases. Therefore, we investigated circular RNA (circRNA) roles in human liver development, an unexplored research domain. METHODS: Using high-throughput sequencing and bioinformatics, we analysed foetal liver samples across developmental stages (7-20 weeks post-conception). Differentially expressed (DE) genes were identified and subjected to enrichment analysis using Gene Ontology (GO), Kyoto Encyclopaedia of Genes and Genomes (KEGG), and Disease Ontology (DO). Modular analysis was performed using the Search Tool for Retrieval of Interacting Genes (STRING), followed by construction of a protein-protein interaction (PPI) network using Cytoscape software. The key genes were screened using Molecular Complex Detection (MCODE). The mRNA levels of hub genes were validated using quantitative reverse transcription polymerase chain reaction (qRT-PCR). RESULTS: There were 645 DE circRNAs and 5,145 DE mRNAs between human livers at the three growth stages (HB, EH, and LH). It was found that the activity of circRNAs was boosted remarkably in the hepatoblastic stage. Enrichment analysis found they mainly involved in nervous system regulation of liver function, embryonic organ development and digestive system development. In addition, DE circRNAs were primarily involved in the PI3K-AKT, MAPK and calcium pathways, potentially contributing to adult liver diseases. Notably, only hsa_circ_001471 and novel_circ_017382 were simultaneously identified at all stages and were persistently downregulated. A co-expression regulatory network involving these circRNAs was established. Three hub genes (LGR5, FOXL1 and RSPO3) were identified from the PPI network of 167 genes and may play key roles in human liver development. The RT-qPCR validation results were in agreement with the sequencing data. CONCLUSIONS: Our findings provide the first insights into the roles and regulatory networks of circRNAs in human liver development, laying the groundwork for further investigations of molecular and signalling networks.

Humans↗

Exploring Williams-Beuren syndrome using myGrid.

MOTIVATION: In silico experiments necessitate the virtual organization of people, data, tools and machines. The scientific process also necessitates an awareness of the experience base, both of personal data as well as the wider context of work. The management of all these data and the co-ordination of resources to manage such virtual organizations and the data surrounding them needs significant computational infra-structure support. RESULTS: In this paper, we show that (my)Grid, middleware for the Semantic Grid, enables biologists to perform and manage in silico experiments, then explore and exploit the results of their experiments. We demonstrate (my)Grid in the context of a series of bioinformatics experiments focused on a 1.5 Mb region on chromosome 7 which is deleted in Williams-Beuren syndrome (WBS). Due to the highly repetitive nature of sequence flanking/in the WBS critical region (WBSCR), sequencing of the region is incomplete leaving documented gaps in the released sequence. (my)Grid was used in a series of experiments to find newly sequenced human genomic DNA clones that extended into these 'gap' regions in order to produce a complete and accurate map of the WBSCR. Once placed in this region, these DNA sequences were analysed with a battery of prediction tools in order to locate putative genes and regulatory elements possibly implicated in the disorder. Finally, any genes discovered were submitted to a range of standard bioinformatics tools for their characterization. We report how (my)Grid has been used to create workflows for these in silico experiments, run those workflows regularly and notify the biologist when new DNA and genes are discovered. The (my)Grid services collect and co-ordinate data inputs and outputs for the experiment, as well as much provenance information about the performance of experiments on WBS. AVAILABILITY: The (my)Grid software is available via http://www.mygrid.org.uk

Algorithms↗

Adapters, shims, and glue--service interoperability for in silico experiments.

MOTIVATION: Computationally, in silico experiments in biology are workflows describing the collaboration of people, data and methods. The Grid and Web services are proposed to be the next generation infrastructure supporting the deployment of bioinformatics workflows. But the growing number of autonomous and heterogeneous services pose challenges to the used middleware w.r.t. composition, i.e. discovery and interoperability of services required within in silico experiments. In the IRIS project, we handle the problem of service interoperability by a semi-automatic procedure for identifying and placing customizable adapters into workflows built by service composition. RESULTS: We show the effectiveness and robustness of the software-aided composition procedure by a case study in the field of life science. In this study we combine different database services with different analysis services with the objective of discovering required adapters. Our experiments show that we can identify relevant adapters with high precision and recall.

Computational Biology↗

Text-mining and information-retrieval services for molecular biology.

Text-mining in molecular biology -- defined as the automatic extraction of information about genes, proteins and their functional relationships from text documents -- has emerged as a hybrid discipline on the edges of the fields of information science, bioinformatics and computational linguistics. A range of text-mining applications have been developed recently that will improve access to knowledge for biologists and database annotators.

Computational Biology↗

Bioinformatics support for high-throughput proteomics.

In the "post-genome" era, mass spectrometry (MS) has become an important method for the analysis of proteome data. The rapid advancement of this technique in combination with other methods used in proteomics results in an increasing number of high-throughput projects. This leads to an increasing amount of data that needs to be archived and analyzed. To cope with the need for automated data conversion, storage, and analysis in the field of proteomics, the open source system ProDB was developed. The system handles data conversion from different mass spectrometer software, automates data analysis, and allows the annotation of MS spectra (e.g. assign gene names, store data on protein modifications). The system is based on an extensible relational database to store the mass spectra together with the experimental setup. It also provides a graphical user interface (GUI) for managing the experimental steps which led to the MS data. Furthermore, it allows the integration of genome and proteome data. Data from an ongoing experiment was used to compare manual and automated analysis. First tests showed that the automation resulted in a significant saving of time. Furthermore, the quality and interpretability of the results was improved in all cases.

Algorithms↗

Assessing the impact of alternative splicing on domain interactions in the human proteome.

We have constructed a database of alternatively spliced protein forms (ASP), consisting of 13,384 protein isoform sequences of 4422 human genes (www.bioinformatics.ucla.edu/ASP). We identified fifty protein domain types that were selectively removed by alternative splicing at much higher frequencies than average (p-value < 0.01). These include many well-known protein-interaction domains (e.g., KRAB; ankyrin repeats; Kelch) including some that have been previously shown to be regulated functionally by alternative splicing (e.g., collagen domain). We present a number of novel examples (Kruppel transcription factors; Pbx2; Enc1) from the ASP database, illustrating how this pattern of alternative splicing changes the structure of a biological pathway, by redirecting protein interaction networks at key switch points. Our bioinformatics analysis indicates that a major impact of alternative splicing is removal of protein-protein interaction domains that mediate key linkages in protein interaction networks. ASP expands the available dataset of human alternatively spliced protein forms from 1989 human genes (SwissProt release 42) to 5413 (nonredundant set, ASP + SwissProt), a nearly 3-fold increase. ASP will enhance the existing pool of protein sequences that are searched by mass spectroscopy software during the identification of peptide fragments.

Alternative Splicing↗

MOLE: a data management application based on a protein production data model.

MOLE (mining, organizing, and logging experiments) has been developed to meet the growing data management and target tracking needs of molecular biologists and protein crystallographers. The prototype reported here will become a Laboratory Information Management System (LIMS) to help protein scientists manage the large amounts of laboratory data being generated due to the acceleration in proteome research and will furthermore facilitate collaborations between groups based at different sites. To achieve this, MOLE is based on the data model for protein production devised at the European Bioinformatics Institute (Pajon A, et al., Proteins in press).

Algorithms↗

GLAD: a system for developing and deploying large-scale bioinformatics grid.

MOTIVATION: Grid computing is used to solve large-scale bioinformatics problems with gigabytes database by distributing the computation across multiple platforms. Until now in developing bioinformatics grid applications, it is extremely tedious to design and implement the component algorithms and parallelization techniques for different classes of problems, and to access remotely located sequence database files of varying formats across the grid. In this study, we propose a grid programming toolkit, GLAD (Grid Life sciences Applications Developer), which facilitates the development and deployment of bioinformatics applications on a grid. RESULTS: GLAD has been developed using ALiCE (Adaptive scaLable Internet-based Computing Engine), a Java-based grid middleware, which exploits the task-based parallelism. Two bioinformatics benchmark applications, such as distributed sequence comparison and distributed progressive multiple sequence alignment, have been developed using GLAD.

Computational Biology↗

SeqVISTA: a graphical tool for sequence feature visualization and comparison.

BACKGROUND: Many readers will sympathize with the following story. You are viewing a gene sequence in Entrez, and you want to find whether it contains a particular sequence motif. You reach for the browser's "find in page" button, but those darn spaces every 10 bp get in the way. And what if the motif is on the opposite strand? Subsequently, your favorite sequence analysis software informs you that there is an interesting feature at position 13982-14013. By painstakingly counting the 10 bp blocks, you are able to examine the sequence at this location. But now you want to see what other features have been annotated close by, and this information is buried several screenfuls higher up the web page. RESULTS: SeqVISTA presents a holistic, graphical view of features annotated on nucleotide or protein sequences. This interactive tool highlights the residues in the sequence that correspond to features chosen by the user, and allows easy searching for sequence motifs or extraction of particular subsequences. SeqVISTA is able to display results from diverse sequence analysis tools in an integrated fashion, and aims to provide much-needed unity to the bioinformatics resources scattered around the Internet. Our viewer may be launched on a GenBank record by a single click of a button installed in the web browser. CONCLUSION: SeqVISTA allows insights to be gained by viewing the totality of sequence annotations and predictions, which may be more revealing than the sum of their parts. SeqVISTA runs on any operating system with a Java 1.4 virtual machine. It is freely available to academic users at http://zlab.bu.edu/SeqVISTA.

Amino Acid Sequence↗