Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Detection of hypopharyngeal squamous cell carcinoma using serum proteomics.

CONCLUSIONS: The combination of surface-enhanced laser desorption/ionization (SELDI) with bioinformatics tools could help find serum proteome biomarkers and establish a predictive model for early detection of hypopharyngeal squamous cell carcinoma (HSCC). OBJECTIVES: Proteomic profiling of serum using surface-enhanced laser desorption/ionization time-of-flight mass spectrometry (SELDI-TOF-MS) is an emerging technique to identify new biomarkers in biological fluids and to establish clinically useful diagnostic computational models. We used it to find new potential biomarkers and to establish a predictive model for early detection of HSCC. MATERIALS AND METHODS: One hundred serum samples including 48 from HSCC patients and 52 from normal controls which were divided into a training set and a blind testing set were treated on WCX2 and IMAC3 protein chip, and serum protein or peptide patterns were detected by SELDI-TOF-MS. The data of spectra were analyzed by Biomarker Wizard software to screen serum proteome biomarkers of HSCC. A decision tree classification algorithm and blind validation were determined by Biomarker Pattern Software (BPS). RESULTS: Ranging from 2 to 30 kDa, 45 potential biomarkers could differentiate HSCC patients from normal controls (p < 0.05). Among them four candidate protein peaks with m/z values of 7796, 4216, 5927, and 5361Da were selected to establish a predictive model by BPS with sensitivity of 94% and specificity of 89%. A sensitivity of 92% and specificity of 82% were validated in the blind testing set.

Adult↗

FOUNTAIN: a JAVA open-source package to assist large sequencing projects.

BACKGROUND: Better automation, lower cost per reaction and a heightened interest in comparative genomics has led to a dramatic increase in DNA sequencing activities. Although the large sequencing projects of specialized centers are supported by in-house bioinformatics groups, many smaller laboratories face difficulties managing the appropriate processing and storage of their sequencing output. The challenges include documentation of clones, templates and sequencing reactions, and the storage, annotation and analysis of the large number of generated sequences. RESULTS: We describe here a new program, named FOUNTAIN, for the management of large sequencing projects http://genetics.hpi.uni-hamburg.de/FOUNTAIN.html. FOUNTAIN uses the JAVA computer language and data storage in a relational database. Starting with a collection of sequencing objects (clones), the program generates and stores information related to the different stages of the sequencing project using a web browser interface for user input. The generated sequences are subsequently imported and annotated based on BLAST searches against the public databases. In addition, simple algorithms to cluster sequences and determine putative polymorphic positions are implemented. CONCLUSIONS: A simple, but flexible and scalable software package is presented to facilitate data generation and storage for large sequencing projects. Open source and largely platform and database independent, we wish FOUNTAIN to be improved and extended in a community effort.

Algorithms↗

YAdumper: extracting and translating large information volumes from relational databases to structured flat files.

Downloading the information stored in relational databases into XML and other flat formats is a common task in bioinformatics. This periodical dumping of information requires considerable CPU time, disk and memory resources. YAdumper has been developed as a purpose-specific tool to deal with the integral structured information download of relational databases. YAdumper is a Java application that organizes database extraction following an XML template based on an external Document Type Declaration. Compared with other non-native alternatives, YAdumper substantially reduces memory requirements and considerably improves writing performance.

Algorithms↗

Expression Profiler: next generation--an online platform for analysis of microarray data.

Expression Profiler (EP, http://www.ebi.ac.uk/expressionprofiler) is a web-based platform for microarray gene expression and other functional genomics-related data analysis. The new architecture, Expression Profiler: next generation (EP:NG), modularizes the original design and allows individual analysis-task-related components to be developed by different groups and yet still seamlessly to work together and share the same user interface look and feel. Data analysis components for gene expression data preprocessing, missing value imputation, filtering, clustering methods, visualization, significant gene finding, between group analysis and other statistical components are available from the EBI (European Bioinformatics Institute) web site. The web-based design of Expression Profiler supports data sharing and collaborative analysis in a secure environment. Developed tools are integrated with the microarray gene expression database ArrayExpress and form the exploratory analytical front-end to those data. EP:NG is an open-source project, encouraging broad distribution and further extensions from the scientific community.

Gene Expression Profiling↗

Accelerated probabilistic inference of RNA structure evolution.

BACKGROUND: Pairwise stochastic context-free grammars (Pair SCFGs) are powerful tools for evolutionary analysis of RNA, including simultaneous RNA sequence alignment and secondary structure prediction, but the associated algorithms are intensive in both CPU and memory usage. The same problem is faced by other RNA alignment-and-folding algorithms based on Sankoff's 1985 algorithm. It is therefore desirable to constrain such algorithms, by pre-processing the sequences and using this first pass to limit the range of structures and/or alignments that can be considered. RESULTS: We demonstrate how flexible classes of constraint can be imposed, greatly reducing the computational costs while maintaining a high quality of structural homology prediction. Any score-attributed context-free grammar (e.g. energy-based scoring schemes, or conditionally normalized Pair SCFGs) is amenable to this treatment. It is now possible to combine independent structural and alignment constraints of unprecedented general flexibility in Pair SCFG alignment algorithms. We outline several applications to the bioinformatics of RNA sequence and structure, including Waterman-Eggert N-best alignments and progressive multiple alignment. We evaluate the performance of the algorithm on test examples from the RFAM database. CONCLUSION: A program, Stemloc, that implements these algorithms for efficient RNA sequence alignment and structure prediction is available under the GNU General Public License.

Algorithms↗

Bioinformatics issues for automating the annotation of genomic sequences.

The rapid explosion in the amount of biological data being generated worldwide is surpassing efforts to manage analysis of the data. As part of an ongoing project to automate and manage bioinformatics analysis, the authors have designed and implemented a simple automated annotation system, which is described in this paper. The system is applied to existing GenBank/DDBJ/EMBL entries and compared with existing annotations to illustrate not only potential errors but also that they are generally not up-to-date, as a result of new versions of analysis tools and updates of genomic repositories. We highlight the important Bioinformatics issues of storage and management of information to ensure data and results are kept up-to-date in light of new information becoming available. Surprisingly, from just four database entries, a significant number of new features were found. We describe the results as well as identify important issues that need to be addressed in order to automate the re-analysis/re-annotation of genomic sequences within a reasonable timeframe.

Computational Biology↗

Increased power of microarray analysis by use of an algorithm based on a multivariate procedure.

MOTIVATION: The power of microarray analyses to detect differential gene expression strongly depends on the statistical and bioinformatical approaches used for data analysis. Moreover, the simultaneous testing of tens of thousands of genes for differential expression raises the 'multiple testing problem', increasing the probability of obtaining false positive test results. To achieve more reliable results, it is, therefore, necessary to apply adjustment procedures to restrict the family-wise type I error rate (FWE) or the false discovery rate. However, for the biologist the statistical power of such procedures often remains abstract, unless validated by an alternative experimental approach. RESULTS: In the present study, we discuss a multiplicity adjustment procedure applied to classical univariate as well as to recently proposed multivariate gene-expression scores. All procedures strictly control the FWE. We demonstrate that the use of multivariate scores leads to a more efficient identification of differentially expressed genes than the widely used MAS5 approach provided by the Affymetrix software tools (Affymetrix Microarray Suite 5 or GeneChip Operating Software). The practical importance of this finding is successfully validated using real time quantitative PCR and data from spike-in experiments. AVAILABILITY: The R-code of the statistical routines can be obtained from the corresponding author. CONTACT: Schuster@imise.uni-leipzig.de

Algorithms↗

Statistical bioinformatic methods in microbial genome analysis.

It is probable that, increasingly, genome investigations are going to be based on statistical formalization. This review summarizes the state of art and potentiality of using statistics in microbial genome analysis. First, I focus on recent advances in functional genomics, such as finding genes and operons, identifying gene conversion events, detecting DNA replication origins and analysing regulatory sites. Then I describe how to use phylogenetic methods in genome analysis and methods for genome-wide scanning for positively selected amino acids. I conclude with speculations on the future course of genome statistical modeling.

Computational Biology↗

Biological systems modeling and analysis: a biomolecular technique of the twenty-first century.

It is proposed that computational systems biology should be considered a biomolecular technique of the twenty-first century, because it complements experimental biology and bioinformatics in unique ways that will eventually lead to insights and a depth of understanding not achievable without systems approaches. This article begins with a summary of traditional and novel modeling techniques. In the second part, it proposes concept map modeling as a useful link between experimental biology and biological systems modeling and analysis. Concept map modeling requires the collaboration between biologist and modeler. The biologist designs a regulated connectivity diagram of processes comprising a biological system and also provides semi-quantitative information on stimuli and measured or expected responses of the system. The modeler converts this information through methods of forward and inverse modeling into a mathematical construct that can be used for simulations and to generate and test new hypotheses. The biologist and the modeler collaboratively interpret the results and devise improved concept maps. The third part of the article describes software, BST-Box, supporting the various modeling activities.

Computational Biology↗

Screening of genes for proteins interacting with the PS1TP5 protein of hepatitis B virus: probing a human leukocyte cDNA library using the yeast two-hybrid system.

BACKGROUND: The hepatitis B virus (HBV) genome includes S, C, P and X regions. The S region is divided into four subregions of pre-pre-S, pre-S1, pre-S2 and S. PS1TP5 (human gene 5 transactivated by pre-S1 protein of HBV) is a novel target gene transactivated by the pre-S1 protein that has been screened with a suppression subtractive hybridization technique in our laboratory (GenBank accession: AY427953). In order to investigate the biological function of the PS1TP5 protein, we performed a yeast two-hybrid system 3 to screen proteins from a human leukocyte cDNA library interacting with the PS1TP5 protein. METHODS: The reverse transcription polymerase chain reaction (RT-PCR) was performed to amplify the gene of PS1TP5 from the mRNA of HepG2 cells and the gene was then cloned into the pGEM-T vector. After being sequenced and analyzed with Vector NTI 9.1 and NCBI BLAST software, the target gene of PS1TP5 was cut from the pGEM-T vector and cloned into a yeast expression plasmid pGBKT7, then "bait" plasmid pGBKT7-PS1TP5 was transformed into the yeast strain AH109. The yeast protein was isolated and analyzed with sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) and Western blotting hybridization. After expression of the pGBKT7-PS1TP5 fusion protein in the AH109 yeast strain was accomplished, a yeast two-hybrid screening was performed by mating AH109 with Y187 containing a leukocyte cDNA library plasmid. The mated yeast was plated on quadruple dropout medium and assayed for alpha-gal activity. The interaction between the PS1TP5 protein and the proteins obtained from positive colonies was further confirmed by repeating the yeast two-hybrid screen. After extracting and sequencing of plasmids from blue colonies we carried out a bioinformatic analysis. RESULTS: Forty true positive colonies were selected and sequenced, full length sequences were obtained and we searched for homologous DNA sequences from GenBank. Among the 40 positive colonies, 23 coding genes with known functions were obtained, including Homo sapien leukocyte adhesion protein p150, 95, interleukin 2 receptor gamma chain, PALM2-AKAP2 protein (PALM2-AKAP2), eukaryotic translation initiation factor 4A, beta-2-microglobin, solute carrier family 9 (sodium/hydrogen exchanger), calreticulin, asialoglycoprotein receptor 1 (ASGR1), MHC class II lymphocyte antigen, cytochrome c oxidase subunit 1, lymphocyte antigen 86 (LY86) and lymphocyte cytosolic protein 1. One novel gene with unknown function was found and named as PS1TP5BP1. After being electronically spliced, it was deposited in GenBank (accession number: DQ471327). CONCLUSIONS: Genes of proteins interacting with PS1TP5 were successfully screened from leukocyte cDNA library. These results suggested that PS1TP5 was closely correlated with immunoregulation, carbohydrate metabolism, signal transduction, the formation of hepatic fibrosis and initiation and development of tumors and also brought some new clues for further studying the biological functions of the pre-S1 protein.

Amino Acid Sequence↗

FeatureExtract--extraction of sequence annotation made easy.

Work on a large number of biological problems benefits tremendously from having an easy way to access the annotation of DNA sequence features, such as intron/exon structure, the contents of promoter regions and the location of other genes in upsteam and downstream regions. For example, taking the placement of introns within a gene into account can help in a phylogenetic analysis of homologous genes. Designing experiments for investigating UTR regions using PCR or DNA microarrays require knowledge of known elements in UTR regions and the positions and strandness of other genes nearby on the chromosome. A wealth of such information is already known and documented in databases such as GenBank and the NCBI Human Genome builds. However, it usually requires significant bioinformatics skills and intimate knowledge of the data format to access this information. Presented here is a highly flexible and easy-to-use tool for extracting feature annotation from GenBank entries. The tool is also useful for extracting datasets corresponding to a particular feature (e.g. promoters). Most importantly, the output data format is highly consistent, easy to handle for the user and easy to parse computationally. The FeatureExtract web server is freely available for both academic and commercial use at http://www.cbs.dtu.dk/services/FeatureExtract/.

Chromosomes↗

A web services choreography scenario for interoperating bioinformatics applications.

BACKGROUND: Very often genome-wide data analysis requires the interoperation of multiple databases and analytic tools. A large number of genome databases and bioinformatics applications are available through the web, but it is difficult to automate interoperation because: 1) the platforms on which the applications run are heterogeneous, 2) their web interface is not machine-friendly, 3) they use a non-standard format for data input and output, 4) they do not exploit standards to define application interface and message exchange, and 5) existing protocols for remote messaging are often not firewall-friendly. To overcome these issues, web services have emerged as a standard XML-based model for message exchange between heterogeneous applications. Web services engines have been developed to manage the configuration and execution of a web services workflow. RESULTS: To demonstrate the benefit of using web services over traditional web interfaces, we compare the two implementations of HAPI, a gene expression analysis utility developed by the University of California San Diego (UCSD) that allows visual characterization of groups or clusters of genes based on the biomedical literature. This utility takes a set of microarray spot IDs as input and outputs a hierarchy of MeSH Keywords that correlates to the input and is grouped by Medical Subject Heading (MeSH) category. While the HTML output is easy for humans to visualize, it is difficult for computer applications to interpret semantically. To facilitate the capability of machine processing, we have created a workflow of three web services that replicates the HAPI functionality. These web services use document-style messages, which means that messages are encoded in an XML-based format. We compared three approaches to the implementation of an XML-based workflow: a hard coded Java application, Collaxa BPEL Server and Taverna Workbench. The Java program functions as a web services engine and interoperates with these web services using a web services choreography language (BPEL4WS). CONCLUSION: While it is relatively straightforward to implement and publish web services, the use of web services choreography engines is still in its infancy. However, industry-wide support and push for web services standards is quickly increasing the chance of success in using web services to unify heterogeneous bioinformatics applications. Due to the immaturity of currently available web services engines, it is still most practical to implement a simple, ad-hoc XML-based workflow by hard coding the workflow as a Java application. For advanced web service users the Collaxa BPEL engine facilitates a configuration and management environment that can fully handle XML-based workflow.

Computational Biology↗

Convergent functional genomics: a Bayesian candidate gene identification approach for complex disorders.

Identifying genes involved in complex neuropsychiatric disorders through classic human genetic approaches has proven difficult. To overcome that barrier, we have developed a translational approach called Convergent Functional Genomics (CFG), which cross-matches animal model microarray gene expression data with human genetic linkage data as well as human postmortem brain data and biological role data, as a Bayesian way of cross-validating findings and reducing uncertainty. Our approach produces a short list of high probability candidate genes out of the hundreds of genes changed in microarray datasets and the hundreds of genes present in a linkage peak chromosomal area. These genes can then be prioritized, pursued, and validated in an individual fashion using: (1) human candidate gene association studies and (2) cell culture and mouse transgenic models. Further bioinformatics analysis of groups of genes identified through CFG leads to insights into pathways and mechanisms that may be involved in the pathophysiology of the illness studied. This simple but powerful approach is likely generalizable to other complex, non-neuropsychiatric disorders, for which good animal models, as well as good human genetic linkage datasets and human target tissue gene expression datasets exist.

Animals↗

An integrated genetic data environment (GDE)-based LINUX interface for analysis of HIV-1 and other microbial sequences.

MOTIVATION: Sequence databases encode a wealth of information needed to develop improved vaccination and treatment strategies for the control of HIV and other important pathogens. To facilitate effective utilization of these datasets, we developed a user-friendly GDE-based LINUX interface that reduces input/output file formatting. DESIGN AND RESULTS: GDE was adapted to the Linux operating system, bioinformatics tools were integrated with microbe-specific databases, and up-to-date GDE menus were developed for several clinically important viral, bacterial and parasitic genomes. Each microbial interface was designed for local access and contains Genbank, BLAST-formatted and phylogenetic databases. AVAILABILITY: GDE-Linux is available for research purposes by direct application to the corresponding author. Application-specific menus and support files can be downloaded from (http://www.bioafrica.net).

Database Management Systems↗

Wrapping up BLAST and other applications for use on Unix clusters.

UNLABELLED: We have developed two programs that speed up common bioinformatic applications by spreading them across a UNIX cluster.(1) BLAST.pm, a new module for the 'MOLLUSC' package. (2) WRAPID, a simple tool for parallelizing large numbers of small instances of programs such as BLAST, FASTA and CLUSTALW. AVAILABILITY: The packages were developed in Perl on a 20-node Linux cluster and are provided together with a configuration script and documentation. They can be freely downloaded from http://wolfe.gen.tcd.ie/wrapper.

Computer Communication Networks↗

Bioinformatics-driven, rational engineering of protein thermostability.

A longstanding goal in protein engineering is to identify specific sequence changes that endow proteins with desired functional properties. As opposed to traditional rational and random protein engineering techniques, we have employed a bioinformatic approach to identify specific sequence changes that influence key functional properties of a protein within a defined superfamily. Specifically, we have used the Bayesian sequence-based algorithms PROBE and Classifier to identify a strand-turn-strand motif that contributes to thermophilicity among members of the serine protease subtilase superfamily. By replacing a 16 amino acid sequence in the mesophilic subtilisin E (from Bacillus subtilis) with a bioinformatics-generated thermophilic model sequence, the melting temperature of subtilisin E was increased by 13 degrees C. While wild-type subtilisin E was inactive at 90 degrees C, the mutant retained a substantial fraction of its function, with ca. one-third of the activity that it has at 45 degrees C.

Algorithms↗

DiscoverySpace: an interactive data analysis application.

DiscoverySpace is a graphical application for bioinformatics data analysis. Users can seamlessly traverse references between biological databases and draw together annotations in an intuitive tabular interface. Datasets can be compared using a suite of novel tools to aid in the identification of significant patterns. DiscoverySpace is of broad utility and its particular strength is in the analysis of serial analysis of gene expression (SAGE) data. The application is freely available online.

Animals↗

Ontology for immunogenetics: the IMGT-ONTOLOGY.

MOTIVATION: IMGT, the international ImMunoGeneTics database (http:@imgt.cines.fr:8104), created by M.-P. Lefranc, is an integrated database specializing in antigen receptors (immunoglobulins and T-cell receptors) and major histocompatibility complex (MHC) of all vertebrate species. IMGT accurate immunogenetics data are based on the standardization of the biological knowledge provided by the 'ImMunoGeneTics' IMGT-ONTOLOGY. The IMGT-ONTOLOGY describes the classification and specification of terms needed for immunogenetics and bioinformatics. IMGT-ONTOLOGY covers four main concepts: 'IDENTIFICATION', 'DESCRIPTION', 'CLASSIFICATION' and 'OBTENTION'. These concepts allow an extensive and standardized description and characterization of immunoglobulin and T-cell receptor data. The controlled vocabulary and the annotation rules are indispensable to ensure accuracy, consistency and coherence in IMGT. IMGT-ONTOLOGY allows scientists and clinicians to use, for the first time, identical terms with the same meaning in immunogenetics. It provides a semantic repository that will improve interoperability between specialist and generalist databases.

Animals↗