Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

e2g: an interactive web-based server for efficiently mapping large EST and cDNA sets to genomic sequences.

e2g is a web-based server which efficiently maps large expressed sequence tag (EST) and cDNA datasets to genomic DNA. It significantly extends the volume of data that can be mapped in reasonable time, and makes this improved efficiency available as a web service. Our server hosts large collections of EST sequences (e.g. 4.1 million mouse ESTs of 1.87 Gb) in precomputed indexed data structures for efficient sequence comparison. The user can upload a genomic DNA sequence of interest and rapidly compare this to the complete collection of ESTs on the server. This delivers a mapping of the ESTs on the genomic DNA. The e2g web interface provides a graphical overview of the mapping. Alignments of the mapped EST regions with parts of the genomic sequence are visualized. Zooming functions allow the user to interactively explore the results. Mapped sequences can be downloaded for further analysis. e2g is available on the Bielefeld University Bioinformatics Server at http://bibiserv.techfak.uni-bielefeld.de/e2g/.

Base Sequence↗

REMUS: a tool for identification of unique peptide segments as epitopes.

We provide a 'R(E)MUS' (reinforced merging techniques for unique peptide segments) web server for identification of the locations and compositions of unique peptide segments from a set of protein family sequences. Different levels of uniqueness are determined according to substitutional relationship in the amino acids, frequency of appearance and biological properties such as priority for serving as candidates for epitopes where antibodies recognize. R(E)MUS also provides interactive visualization of 3D structures for allocation and comparison of the identified unique peptide segments. Accuracy of the algorithm was found to be 70% in terms of mapping a unique peptide segment as an epitope. The R(E)MUS web server is available at http://biotools.cs.ntou.edu.tw/REMUS and the PC version software can be freely downloaded either at http://bioinfo.life.nthu.edu.tw/REMUS or http://spider.cs.ntou.edu.tw/BioTools/REMUS. User guide and working examples for PC version are available at http://spider.cs.ntou.edu.tw/BioTools/REMUS-DOCS.html, and details of the proposed algorithm can be referred to the documents as described previously [H. T. Chang, T. W. Pai, T. C. Fan, B. H. Su, P. C. Wu, C. Y. Tang, C. T. Chang, S. H. Liu and M. D. T. Chang (2006) BMC Bioinformatics, 7, 38 and T. W. Pai, B. H. Su, P. C. Wu, M. D. T. Chang, H. T. Chang, T. C. Fan and S. H. Liu (2006) J. Bioinform. Comput. Biol., 4, 75-92].

Algorithms↗

CoMR: an integrative scoring pipeline for comprehensive mitochondrial proteome reconstruction across eukaryotes.

Mitochondrial proteome reconstruction from eukaryotic sequence data typically relies on prediction of mitochondrial targeting signals (MTSs). However, MTS predictors are primarily trained on model organisms and may perform poorly in phylogenetically divergent lineages or in organisms with atypical or reduced targeting sequences. Accurate reconstruction therefore requires integration of complementary sources of evidence beyond targeting prediction alone. We developed Comprehensive Mitochondrial Reconstructor (CoMR), an integrative workflow that combines targeting prediction, curated homology searches, large-scale similarity searches, and automated phylogenetic analysis within a unified scoring framework. Benchmarking on the model yeast Saccharomyces cerevisiae yielded strong discriminatory performance [receiver operating characteristic (ROC)-area under the curve (AUC) = 0.92], exceeding standalone prediction with TargetP2, a predictor of N-terminal targeting peptides (ROC-AUC = 0.72). In the divergent anaerobic protist Paratrimastix pyriformis, CoMR maintained robust performance (ROC-AUC = 0.86) validated with an experimental proteome despite extreme class imbalance, achieving a precision-recall AUC of 0.183 (~78-fold enrichment over random expectation and ~10-fold improvement over TargetP2). Ablation analyses demonstrate that predictive performance is robust to individual evidence-layer removal, while overlap analyses showed that homology-based searches recovered candidates missed by targeting predictors, particularly in P. pyriformis. Overall, CoMR improves mitochondrial proteome reconstruction over targeting prediction alone and provides a reproducible workflow for predicting mitochondrial and mitochondrion-related organelle protein repertoires across eukaryotes to aid investigations of organelle evolution and proteome reduction.

Proteome↗

Serum proteomic features for detection of endometrial cancer.

To find new potential biomarkers for detection of endometrial cancer (EC), 70 serum samples including 40 from EC patients and 30 from normal healthy females were detected by surface-enhanced laser desorption-ionization time-of-flight mass spectrometry (SELDI-TOF-MS) using WCX2 (weak cation exchange) protein chip. Mass spectra were then assessed with three powerful data-mining tools: a tree classifier, Biomarker Wizard software, and Biomarker Patterns System. The diagnostic pattern combined with 13 potential biomarkers could differentiate EC patients from healthy persons, with a specificity of 100%, sensitivity of 92.5%, and total coincidence of 95.7%. The combination of surface-enhanced laser desorption-ionization with bioinformatics tools could help find new biomarkers and establish with high sensitivity and specificity for the detection of EC.

Adult↗

[Analysis of p53 mutational spectra of esophageal squamous cell carcinomas from Linzhou, comparison with esophageal and other cancers from other areas].

OBJECTIVE: To study the etiological clues involved (in esophageal cancer in Linzhou, Henan, a high-incidence area for esophageal cancer) by analyzing p53 mutational spectrum from esophageal precancerous and cancerous lesions. METHODS: Using bolt bioinformatic and Monte Carlo methods to analyze p53 mutation spectra from "The IARC Database of Somatic p53 Mutations in Human Tumors and Cell Lines", "p53 Database at Institute Curic" and to establish a local database based on these data using the FileMark Pro 3.0 software to allow fast and off-line analysis on a PC from the authors' laboratory. RESULTS: We found that esophageal squamous cell carcinomas from Linzhou had a lower prevalence of base substitutions associated with strand bias than those from other areas (32.8% vs 39.8%). However, a higher prevalence of G:C-->A:T transitions at CpG site (29.6% vs16.4%) was found. Esophageal squamous cell carcinomas from Linzhou displayed a distinctive profile of mutation hotspots, including codons 273 (covers 11.3% of all missense mutations), 175 (9.7%), 158 (9.7%), 159 (6.5%) and 282 (6.5%), all of which were at the CpG site. Statistical analysis showed that the p53 mutation profiles between esophageal squamous cell carcinomas from Linzhou and those from other areas were different (P = 0.02). The p53 mutation profiles of esophageal squamous cell carcinomas from Linzhou and from other areas were also different from cancer of the head and neck. CONCLUSION: Data showed that the p53 mutational spectrum of esophageal squamous cell carcinomas from Linzhou baring the characteristics of those caused both by endogenous and exogenous mutagenic agents, suggesting the potential involvement of chronic inflammation, unique dietary habits and carcinogen exposure in the pathogenesis of esophageal squamous cell carcinoma in Linzhou.

Carcinogens, Environmental↗

Biological data becomes computer literate: new advances in bioinformatics.

Bioinformatics is an art and science concerned with the use of computing in biological research areas such as genomics, transcriptomics, proteomics, genetics, and evolution. This review paints a broad picture of bioinformatics, drawing examples from genomic sequencing and microarray analysis. I highlight the role of bioinformatics at multiple points along the path from high-tech data generation to biological discovery.

Computational Biology↗

OBIYagns: a grid-based biochemical simulator with a parameter estimator.

UNLABELLED: OBIYagns (yet another gene network simulator) is a biochemical system simulator that comprises a multiple-user Web-based graphical interface, an ordinary differential equation solver and a parameter estimators distributed over an open bioinformatics grid (OBIGrid). This grid-based biochemical simulation system can achieve high performance and provide a secure simulation environment for estimating kinetic parameters in an acceptable time period. OBIYagns can be applied to larger system biology-oriented simulation projects. AVAILABILITY: OBIYagns example models, methods and user guide are available at https://access.obigrid.org/yagns/ SUPPLEMENTARY INFORMATION: Please refer to Bioinformatics online.

Algorithms↗

GAME: detecting cis-regulatory elements using a genetic algorithm.

MOTIVATION: Identification of a transcription factor binding sites is an important aspect of the analysis of genetic regulation. Many programs have been developed for the de novo discovery of a binding motif (collection of binding sites). Recently, a scoring function formulation was derived that allows for the comparison of discovered motifs from different programs [S.T. Jensen, X.S. Liu, Q. Zhou and J.S. Liu (2004) Stat. Sci., 19, 188-204.] A simple program, BioOptimizer, was proposed in [S.T. Jensen and J.S. Liu (2004) Bioinformatics, 20, 1557-1564.] that improved discovered motifs by optimizing a scoring function. However, BioOptimizer is a very simple algorithm that can only make local improvements upon an already discovered motif and so BioOptimizer can only be used in conjunction with other motif-finding software. RESULTS: We introduce software, GAME, which utilizes a genetic algorithm to find optimal motifs in DNA sequences. GAME evolves motifs with high fitness from a population of randomly generated starting motifs, which eliminate the reliance on additional motif-finding programs. In addition to using standard genetic operations, GAME also incorporates two additional operators that are specific to the motif discovery problem. We demonstrate the superior performance of GAME compared with MEME, BioProspector and BioOptimizer in simulation studies as well as several real data applications where we use an extended version of the GAME algorithm that allows the motif width to be unknown.

Algorithms↗

SIRW: A web server for the Simple Indexing and Retrieval System that combines sequence motif searches with keyword searches.

SIRW (http://sirw.embl.de/) is a World Wide Web interface to the Simple Indexing and Retrieval System (SIR) that is capable of parsing and indexing various flat file databases. In addition it provides a framework for doing sequence analysis (e.g. motif pattern searches) for selected biological sequences through keyword search. SIRW is an ideal tool for the bioinformatics community for searching as well as analyzing biological sequences of interest.

Abstracting and Indexing↗

The Bioinformatics Links Directory: a compilation of molecular biology web servers.

The Bioinformatics Links Directory is an online community resource that contains a directory of freely available tools, databases, and resources for bioinformatics and molecular biology research. The listing of the servers published in this and previous issues of Nucleic Acids Research together with other useful tools and websites represents a rich repository of resources that are openly provided to the research community using internet technologies. The 166 servers highlighted in the 2005 90002 are included in the more than 700 links to useful online resources that are currently contained within the descriptive biological categories of the Bioinformatics Links Directory. This curated listing of bioinformatics resources is available online at the Bioinformatics Links Directory web site, http://bioinformatics.ubc.ca/resources/links_directory/. A complete listing of the 2005 Nucleic Acids Research 90002 servers is available online at the Nucleic Acids web site, http://nar.oupjournals.org/, and on the Bioinformatics Links Directory web site, http://bioinformatics.ubc.ca/resources/links_directory/narweb2005/.

Computational Biology↗

Mass spectrometric identification of proteins in complex post-genomic projects. Soluble proteins of the metabolically versatile, denitrifying 'Aromatoleum' sp. strain EbN1.

The rapidly developing proteomics technologies help to advance the global understanding of physiological and cellular processes. The lifestyle of a study organism determines the type and complexity of a given proteomic project. The complexity of this study is characterized by a broad collection of pathway-specific subproteomes, reflecting the metabolic versatility as well as the regulatory potential of the aromatic-degrading, denitrifying bacterium 'Aromatoleum' sp. strain EbN1. Differences in protein profiles were determined using a gel-based approach. Protein identification was based on a progressive application of MALDI-TOF-MS, MALDI-TOF-MS/MS and LC-ESI-MS/MS. This progression was result-driven and automated by software control. The identification rate was increased by the assembly of a project-specific list of background signals that was used for internal calibration of the MS spectra, and by the combination of two search engines using a dedicated MetaScoring algorithm. In total, intelligent bioinformatics could increase the identification yield from 53 to 70% of the analyzed 5,050 gel spots; a total of 556 different proteins were identified. MS identification was highly reproducible: most proteins were identified more than twice from parallel 2DE gels with an average sequence coverage of >50% and rather restrictive score thresholds (Mascot >or=95, ProFound >or=2.2, MetaScore >or=97). The MS technologies and bioinformatics tools that were implemented and integrated to handle this complex proteomic project are presented. In addition, we describe the basic principles and current developments of the applied technologies and provide an overview over the current state of microbial proteome research.

Amino Acid Sequence↗

GibbsST: a Gibbs sampling method for motif discovery with enhanced resistance to local optima.

BACKGROUND: Computational discovery of transcription factor binding sites (TFBS) is a challenging but important problem of bioinformatics. In this study, improvement of a Gibbs sampling based technique for TFBS discovery is attempted through an approach that is widely known, but which has never been investigated before: reduction of the effect of local optima. RESULTS: To alleviate the vulnerability of Gibbs sampling to local optima trapping, we propose to combine a thermodynamic method, called simulated tempering, with Gibbs sampling. The resultant algorithm, GibbsST, is then validated using synthetic data and actual promoter sequences extracted from Saccharomyces cerevisiae. It is noteworthy that the marked improvement of the efficiency presented in this paper is attributable solely to the improvement of the search method. CONCLUSION: Simulated tempering is a powerful solution for local optima problems found in pattern discovery. Extended application of simulated tempering for various bioinformatic problems is promising as a robust solution against local optima problems.

Algorithms↗

AutoPVPrimer: A comprehensive AI-Enhanced pipeline for efficient plant virus primer design and assessment.

Plant viruses pose a significant threat to global agriculture and require efficient tools for their timely detection. We present AutoPVPrimer, an innovative pipeline that integrates artificial intelligence (AI) and machine learning to accelerate the development of plant virus primers. The pipeline uses Biopython to automatically retrieve different genomic sequences from the NCBI database to increase the robustness of the subsequent primer design. The design_primers_with_tuning module uses a random forest classifier that optimizes parameters and provides flexibility for different experimental conditions. Quality control measures, including the evaluation of poly-X content and melting temperature, increase primer reliability. Unique to AutoPVPrimer is the visualize_primer_dimer module, which supports the visual evaluation of primer dimers-a feature missing in other tools. Primer specificity is validated via primer BLAST, which contributes to the overall efficiency of the pipeline. AutoPVPrimer has been successfully applied to the tomato mosaic virus, proving its adaptability and efficiency. The modular design allows customization by the user and extends the applicability to different plant viruses and experimental scenarios. The pipeline represents a significant advance in primer design and provides researchers with an effective tool to accelerate molecular biology experiments. Future developments aim to extend compatibility and incorporate user feedback to consolidate AutoPVPrimer as an innovative contribution to the bioinformatics toolbox and a promising resource for the advancement of plant virology research.

DNA Primers↗

Bioinformatics -- a patenting view.

The use of bioinformatics in the biological sciences has brought about a change in the way that biological inventions can be protected by patent laws. Using approaches developed in the fields of computer science and business, patent applicants now seek to protect certain aspects of their inventions, which include software, methods of doing business and uses of information as well as more traditional biotechnological products and processes. These approaches are useful in resolving some of the difficulties now faced in prosecuting patent applications directed to biological inventions that are claimed in more conventional terms.

Algorithms↗

SNP-PHAGE--High throughput SNP discovery pipeline.

BACKGROUND: Single nucleotide polymorphisms (SNPs) as defined here are single base sequence changes or short insertion/deletions between or within individuals of a given species. As a result of their abundance and the availability of high throughput analysis technologies SNP markers have begun to replace other traditional markers such as restriction fragment length polymorphisms (RFLPs), amplified fragment length polymorphisms (AFLPs) and simple sequence repeats (SSRs or microsatellite) markers for fine mapping and association studies in several species. For SNP discovery from chromatogram data, several bioinformatics programs have to be combined to generate an analysis pipeline. Results have to be stored in a relational database to facilitate interrogation through queries or to generate data for further analyses such as determination of linkage disequilibrium and identification of common haplotypes. Although these tasks are routinely performed by several groups, an integrated open source SNP discovery pipeline that can be easily adapted by new groups interested in SNP marker development is currently unavailable. RESULTS: We developed SNP-PHAGE (SNP discovery Pipeline with additional features for identification of common haplotypes within a sequence tagged site (Haplotype Analysis) and GenBank (-dbSNP) submissions. This tool was applied for analyzing sequence traces from diverse soybean genotypes to discover over 10,000 SNPs. This package was developed on UNIX/Linux platform, written in Perl and uses a MySQL database. Scripts to generate a user-friendly web interface are also provided with common queries for preliminary data analysis. A machine learning tool developed by this group for increasing the efficiency of SNP discovery is integrated as a part of this package as an optional feature. The SNP-PHAGE package is being made available open source at http://bfgl.anri.barc.usda.gov/ML/snp-phage/. CONCLUSION: SNP-PHAGE provides a bioinformatics solution for high throughput SNP discovery, identification of common haplotypes within an amplicon, and GenBank (dbSNP) submissions. SNP selection and visualization are aided through a user-friendly web interface. This tool is useful for analyzing sequence tagged sites (STSs) of genomic sequences, and this software can serve as a starting point for groups interested in developing SNP markers.

Base Sequence↗

[Proteomic analysis of prostate cancer using surface enhanced laser desorption/ionization mass spectrometry].

OBJECTIVE: To identify the serum biomarkers of prostate cancer by using protein chip and bioinformatics. METHODS: Eighty three prostate cancer (PCA) patients and ninety five healthy people from mass screen in Changchun were detected by surface-enhanced laser desorption/ionization mass spectrometry (SELDI-MS). The data of spectra were analyzed by bioinformatics tools-Biomarker Wizard and Biomarker Pattern. RESULTS: Compared with the spectra of healthy people, there were 18 potential markers detected in the spectra of the PCA patients, the protein expression was high in 4 of which and low in the 10 of which. The softwares Biomarkerwizard and Biomarker Pattern automatically, under given conditions, selected 8 biomarker proteins to be used to establish a five layer decision tree differentiate to diagnose PCA and differentiate PCA from healthy people with a specificity of 92.632% and a sensitivity of 96.386%. CONCLUSION: New serum biomarkers of PCA have been identified, and this SELDI mass spectrometry coupled with decision tree classification algorithm will provide a highly accurate and innovative approach for the early diagnosis of PCA.

Aged↗

Rose: generating sequence families.

MOTIVATION: We present a new probabilistic model of the evolution of RNA-, DNA-, or protein-like sequences and a software tool, Rose, that implements this model. Guided by an evolutionary tree, a family of related sequences is created from a common ancestor sequence by insertion, deletion and substitution of characters. During this artificial evolutionary process, the 'true' history is logged and the 'correct' multiple sequence alignment is created simultaneously. The model also allows for varying rates of mutation within the sequences, making it possible to establish so-called sequence motifs. RESULTS: The data created by Rose are suitable for the evaluation of methods in multiple sequence alignment computation and the prediction of phylogenetic relationships. It can also be useful when teaching courses in or developing models of sequence evolution and in the study of evolutionary processes. AVAILABILITY: Rose is available on the Bielefeld Bioinformatics WebServer under the following URL: http://bibiserv.TechFak.Uni-Bielefeld.DE/rose/ The source code is available upon request. CONTACT: folker@TechFak.Uni-Bielefeld.DE

Algorithms↗

Reconstruction of metabolic networks from genome data and analysis of their global structure for various organisms.

MOTIVATION: Information from fully sequenced genomes makes it possible to reconstruct strain-specific global metabolic network for structural and functional studies. These networks are often very large and complex. To properly understand and analyze the global properties of metabolic networks, methods for rationally representing and quantitatively analyzing their structure are needed. RESULTS: In this work, the metabolic networks of 80 fully sequenced organisms are in silico reconstructed from genome data and an extensively revised bioreaction database. The networks are represented as directed graphs and analyzed by using the 'breadth first searching algorithm to identify the shortest pathway (path length) between any pair of the metabolites. The average path length of the networks are then calculated and compared for all the organisms. Different from previous studies the connections through current metabolites and cofactors are deleted to make the path length analysis physiologically more meaningful. The distribution of the connection degree of these networks is shown to follow the power law, indicating that the overall structure of all the metabolic networks has the characteristics of a small world network. However, clear differences exist in the network structure of the three domains of organisms. Eukaryotes and archaea have a longer average path length than bacteria. AVAILABILITY: The reaction database in excel format and the programs in VBA (Visual Basic for Applications) are available upon request. SUPPLEMENTARY MATERIAL: Bioinformatics Online.

Archaea↗