Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

e2g: an interactive web-based server for efficiently mapping large EST and cDNA sets to genomic sequences.

e2g is a web-based server which efficiently maps large expressed sequence tag (EST) and cDNA datasets to genomic DNA. It significantly extends the volume of data that can be mapped in reasonable time, and makes this improved efficiency available as a web service. Our server hosts large collections of EST sequences (e.g. 4.1 million mouse ESTs of 1.87 Gb) in precomputed indexed data structures for efficient sequence comparison. The user can upload a genomic DNA sequence of interest and rapidly compare this to the complete collection of ESTs on the server. This delivers a mapping of the ESTs on the genomic DNA. The e2g web interface provides a graphical overview of the mapping. Alignments of the mapped EST regions with parts of the genomic sequence are visualized. Zooming functions allow the user to interactively explore the results. Mapped sequences can be downloaded for further analysis. e2g is available on the Bielefeld University Bioinformatics Server at http://bibiserv.techfak.uni-bielefeld.de/e2g/.

Base Sequence↗

CoMR: an integrative scoring pipeline for comprehensive mitochondrial proteome reconstruction across eukaryotes.

Mitochondrial proteome reconstruction from eukaryotic sequence data typically relies on prediction of mitochondrial targeting signals (MTSs). However, MTS predictors are primarily trained on model organisms and may perform poorly in phylogenetically divergent lineages or in organisms with atypical or reduced targeting sequences. Accurate reconstruction therefore requires integration of complementary sources of evidence beyond targeting prediction alone. We developed Comprehensive Mitochondrial Reconstructor (CoMR), an integrative workflow that combines targeting prediction, curated homology searches, large-scale similarity searches, and automated phylogenetic analysis within a unified scoring framework. Benchmarking on the model yeast Saccharomyces cerevisiae yielded strong discriminatory performance [receiver operating characteristic (ROC)-area under the curve (AUC) = 0.92], exceeding standalone prediction with TargetP2, a predictor of N-terminal targeting peptides (ROC-AUC = 0.72). In the divergent anaerobic protist Paratrimastix pyriformis, CoMR maintained robust performance (ROC-AUC = 0.86) validated with an experimental proteome despite extreme class imbalance, achieving a precision-recall AUC of 0.183 (~78-fold enrichment over random expectation and ~10-fold improvement over TargetP2). Ablation analyses demonstrate that predictive performance is robust to individual evidence-layer removal, while overlap analyses showed that homology-based searches recovered candidates missed by targeting predictors, particularly in P. pyriformis. Overall, CoMR improves mitochondrial proteome reconstruction over targeting prediction alone and provides a reproducible workflow for predicting mitochondrial and mitochondrion-related organelle protein repertoires across eukaryotes to aid investigations of organelle evolution and proteome reduction.

Proteome↗

[Analysis of p53 mutational spectra of esophageal squamous cell carcinomas from Linzhou, comparison with esophageal and other cancers from other areas].

OBJECTIVE: To study the etiological clues involved (in esophageal cancer in Linzhou, Henan, a high-incidence area for esophageal cancer) by analyzing p53 mutational spectrum from esophageal precancerous and cancerous lesions. METHODS: Using bolt bioinformatic and Monte Carlo methods to analyze p53 mutation spectra from "The IARC Database of Somatic p53 Mutations in Human Tumors and Cell Lines", "p53 Database at Institute Curic" and to establish a local database based on these data using the FileMark Pro 3.0 software to allow fast and off-line analysis on a PC from the authors' laboratory. RESULTS: We found that esophageal squamous cell carcinomas from Linzhou had a lower prevalence of base substitutions associated with strand bias than those from other areas (32.8% vs 39.8%). However, a higher prevalence of G:C-->A:T transitions at CpG site (29.6% vs16.4%) was found. Esophageal squamous cell carcinomas from Linzhou displayed a distinctive profile of mutation hotspots, including codons 273 (covers 11.3% of all missense mutations), 175 (9.7%), 158 (9.7%), 159 (6.5%) and 282 (6.5%), all of which were at the CpG site. Statistical analysis showed that the p53 mutation profiles between esophageal squamous cell carcinomas from Linzhou and those from other areas were different (P = 0.02). The p53 mutation profiles of esophageal squamous cell carcinomas from Linzhou and from other areas were also different from cancer of the head and neck. CONCLUSION: Data showed that the p53 mutational spectrum of esophageal squamous cell carcinomas from Linzhou baring the characteristics of those caused both by endogenous and exogenous mutagenic agents, suggesting the potential involvement of chronic inflammation, unique dietary habits and carcinogen exposure in the pathogenesis of esophageal squamous cell carcinoma in Linzhou.

Carcinogens, Environmental↗

Biological data becomes computer literate: new advances in bioinformatics.

Bioinformatics is an art and science concerned with the use of computing in biological research areas such as genomics, transcriptomics, proteomics, genetics, and evolution. This review paints a broad picture of bioinformatics, drawing examples from genomic sequencing and microarray analysis. I highlight the role of bioinformatics at multiple points along the path from high-tech data generation to biological discovery.

Computational Biology↗

OBIYagns: a grid-based biochemical simulator with a parameter estimator.

UNLABELLED: OBIYagns (yet another gene network simulator) is a biochemical system simulator that comprises a multiple-user Web-based graphical interface, an ordinary differential equation solver and a parameter estimators distributed over an open bioinformatics grid (OBIGrid). This grid-based biochemical simulation system can achieve high performance and provide a secure simulation environment for estimating kinetic parameters in an acceptable time period. OBIYagns can be applied to larger system biology-oriented simulation projects. AVAILABILITY: OBIYagns example models, methods and user guide are available at https://access.obigrid.org/yagns/ SUPPLEMENTARY INFORMATION: Please refer to Bioinformatics online.

Algorithms↗

SIRW: A web server for the Simple Indexing and Retrieval System that combines sequence motif searches with keyword searches.

SIRW (http://sirw.embl.de/) is a World Wide Web interface to the Simple Indexing and Retrieval System (SIR) that is capable of parsing and indexing various flat file databases. In addition it provides a framework for doing sequence analysis (e.g. motif pattern searches) for selected biological sequences through keyword search. SIRW is an ideal tool for the bioinformatics community for searching as well as analyzing biological sequences of interest.

Abstracting and Indexing↗

The Bioinformatics Links Directory: a compilation of molecular biology web servers.

The Bioinformatics Links Directory is an online community resource that contains a directory of freely available tools, databases, and resources for bioinformatics and molecular biology research. The listing of the servers published in this and previous issues of Nucleic Acids Research together with other useful tools and websites represents a rich repository of resources that are openly provided to the research community using internet technologies. The 166 servers highlighted in the 2005 90002 are included in the more than 700 links to useful online resources that are currently contained within the descriptive biological categories of the Bioinformatics Links Directory. This curated listing of bioinformatics resources is available online at the Bioinformatics Links Directory web site, http://bioinformatics.ubc.ca/resources/links_directory/. A complete listing of the 2005 Nucleic Acids Research 90002 servers is available online at the Nucleic Acids web site, http://nar.oupjournals.org/, and on the Bioinformatics Links Directory web site, http://bioinformatics.ubc.ca/resources/links_directory/narweb2005/.

Computational Biology↗

AutoPVPrimer: A comprehensive AI-Enhanced pipeline for efficient plant virus primer design and assessment.

Plant viruses pose a significant threat to global agriculture and require efficient tools for their timely detection. We present AutoPVPrimer, an innovative pipeline that integrates artificial intelligence (AI) and machine learning to accelerate the development of plant virus primers. The pipeline uses Biopython to automatically retrieve different genomic sequences from the NCBI database to increase the robustness of the subsequent primer design. The design_primers_with_tuning module uses a random forest classifier that optimizes parameters and provides flexibility for different experimental conditions. Quality control measures, including the evaluation of poly-X content and melting temperature, increase primer reliability. Unique to AutoPVPrimer is the visualize_primer_dimer module, which supports the visual evaluation of primer dimers-a feature missing in other tools. Primer specificity is validated via primer BLAST, which contributes to the overall efficiency of the pipeline. AutoPVPrimer has been successfully applied to the tomato mosaic virus, proving its adaptability and efficiency. The modular design allows customization by the user and extends the applicability to different plant viruses and experimental scenarios. The pipeline represents a significant advance in primer design and provides researchers with an effective tool to accelerate molecular biology experiments. Future developments aim to extend compatibility and incorporate user feedback to consolidate AutoPVPrimer as an innovative contribution to the bioinformatics toolbox and a promising resource for the advancement of plant virology research.

DNA Primers↗

Bioinformatics -- a patenting view.

The use of bioinformatics in the biological sciences has brought about a change in the way that biological inventions can be protected by patent laws. Using approaches developed in the fields of computer science and business, patent applicants now seek to protect certain aspects of their inventions, which include software, methods of doing business and uses of information as well as more traditional biotechnological products and processes. These approaches are useful in resolving some of the difficulties now faced in prosecuting patent applications directed to biological inventions that are claimed in more conventional terms.

Algorithms↗

[Proteomic analysis of prostate cancer using surface enhanced laser desorption/ionization mass spectrometry].

OBJECTIVE: To identify the serum biomarkers of prostate cancer by using protein chip and bioinformatics. METHODS: Eighty three prostate cancer (PCA) patients and ninety five healthy people from mass screen in Changchun were detected by surface-enhanced laser desorption/ionization mass spectrometry (SELDI-MS). The data of spectra were analyzed by bioinformatics tools-Biomarker Wizard and Biomarker Pattern. RESULTS: Compared with the spectra of healthy people, there were 18 potential markers detected in the spectra of the PCA patients, the protein expression was high in 4 of which and low in the 10 of which. The softwares Biomarkerwizard and Biomarker Pattern automatically, under given conditions, selected 8 biomarker proteins to be used to establish a five layer decision tree differentiate to diagnose PCA and differentiate PCA from healthy people with a specificity of 92.632% and a sensitivity of 96.386%. CONCLUSION: New serum biomarkers of PCA have been identified, and this SELDI mass spectrometry coupled with decision tree classification algorithm will provide a highly accurate and innovative approach for the early diagnosis of PCA.

Aged↗

Rose: generating sequence families.

MOTIVATION: We present a new probabilistic model of the evolution of RNA-, DNA-, or protein-like sequences and a software tool, Rose, that implements this model. Guided by an evolutionary tree, a family of related sequences is created from a common ancestor sequence by insertion, deletion and substitution of characters. During this artificial evolutionary process, the 'true' history is logged and the 'correct' multiple sequence alignment is created simultaneously. The model also allows for varying rates of mutation within the sequences, making it possible to establish so-called sequence motifs. RESULTS: The data created by Rose are suitable for the evaluation of methods in multiple sequence alignment computation and the prediction of phylogenetic relationships. It can also be useful when teaching courses in or developing models of sequence evolution and in the study of evolutionary processes. AVAILABILITY: Rose is available on the Bielefeld Bioinformatics WebServer under the following URL: http://bibiserv.TechFak.Uni-Bielefeld.DE/rose/ The source code is available upon request. CONTACT: folker@TechFak.Uni-Bielefeld.DE

Algorithms↗

Reconstruction of metabolic networks from genome data and analysis of their global structure for various organisms.

MOTIVATION: Information from fully sequenced genomes makes it possible to reconstruct strain-specific global metabolic network for structural and functional studies. These networks are often very large and complex. To properly understand and analyze the global properties of metabolic networks, methods for rationally representing and quantitatively analyzing their structure are needed. RESULTS: In this work, the metabolic networks of 80 fully sequenced organisms are in silico reconstructed from genome data and an extensively revised bioreaction database. The networks are represented as directed graphs and analyzed by using the 'breadth first searching algorithm to identify the shortest pathway (path length) between any pair of the metabolites. The average path length of the networks are then calculated and compared for all the organisms. Different from previous studies the connections through current metabolites and cofactors are deleted to make the path length analysis physiologically more meaningful. The distribution of the connection degree of these networks is shown to follow the power law, indicating that the overall structure of all the metabolic networks has the characteristics of a small world network. However, clear differences exist in the network structure of the three domains of organisms. Eukaryotes and archaea have a longer average path length than bacteria. AVAILABILITY: The reaction database in excel format and the programs in VBA (Visual Basic for Applications) are available upon request. SUPPLEMENTARY MATERIAL: Bioinformatics Online.

Archaea↗

SSEP: Secondary structural elements of proteins.

SSEP is a comprehensive resource for accessing information related to the secondary structural elements present in the 25 and 90% non-redundant protein chains. The database contains 1771 protein chains from 1670 protein structures and 6182 protein chains from 5425 protein structures in 25 and 90% non-redundant protein chains, respectively. The current version provides information about the alpha-helical segments and beta-strand fragments of varying lengths. In addition, it also contains the information about 3(10)-helix, beta- and nu-turns and hairpin loops. The free graphics program RASMOL has been interfaced with the search engine to visualize the three-dimensional structures of the user queried secondary structural fragment. The database is updated regularly and is available through Bioinformatics web server at http://cluster.physics.iisc.ernet.in/ssep/ or http://144.16.71.148/ssep/.

Databases, Protein↗

Structural principles governing domain motions in proteins.

With the use of a recently developed method, twenty-four proteins for which two or more X-ray conformers are known have been analyzed to reveal structural principles that govern domain motions in proteins. In all 24 cases, the domain motion is a rotation about a physical axis created through local interactions both covalent and noncovalent. In many cases, two or more mechanical hinges separated in space create a stable hinge axis for precise control of the domain closure. The terminal regions of alpha-helices and beta-sheets have been found to act as mechanical hinges in a significant number of cases. In some cases, the two terminal regions of neighboring strands of a single beta-sheet can create a hinge axis, as can the two termini of a single alpha-helix. These two structures have been termed the "double-hinged beta-sheet" and "double-hinged alpha-helix," respectively. A flexible loop that attaches one domain to another and through which the effective hinge axis passes is another construct that is used to create a hinge. Noncovalent interactions between segments remote along the polypeptide chain can also form hinges. In addition alpha-helices that preserve their hydrogen bonding structure when bent have been found to behave as mechanical hinges. It is suggested that these alpha-helices act as a store of elastic energy that drives the closing of domains for rapid capture of the substrate. If the repertoire of possible interdomain structures is as limited as this study suggests, the dynamic behavior of proteins could soon be predicted using bioinformatics techniques. Proteins 1999;36:425-435.

Animals↗

FatiGO: a web tool for finding significant associations of Gene Ontology terms with groups of genes.

We present a simple but powerful procedure to extract Gene Ontology (GO) terms that are significantly over- or under-represented in sets of genes within the context of a genome-scale experiment (DNA microarray, proteomics, etc.). Said procedure has been implemented as a web application, FatiGO, allowing for easy and interactive querying. FatiGO, which takes the multiple-testing nature of statistical contrast into account, currently includes GO associations for diverse organisms (human, mouse, fly, worm and yeast) and the TrEMBL/Swissprot GOAnnotations@EBI correspondences from the European Bioinformatics Institute.

Algorithms↗

Structural bioinformatics of DNA: a web-based tool for the analysis of molecular dynamics results and structure prediction.

UNLABELLED: We report here the release of a web-based tool (MDDNA) to study and model the fine structural details of DNA on the basis of data extracted from a set of molecular dynamics (MD) trajectories of DNA sequences involving all the unique tetranucleotides. The dynamic web interface can be employed to analyze the first neighbor sequence context effects on the 10 unique dinucleotide steps of DNA. Functionality is included to build all atom models of any user-defined sequence based on the MD results. The backend of this interface is a relational database storing the conformational details of DNA obtained in 39 different MD simulation trajectories comprising all the 136 unique tetranucleotide steps. Examples of the use of this data to predict DNA structures are included. AVAILABILITY: http://humphry.chem.wesleyan.edu:8080/MDDNA. SUPPLEMENTARY INFORMATION: Supplementary data including color figures are available at Bioinformatics online.

Base Sequence↗

HTself: self-self based statistical test for low replication microarray studies.

Different statistical methods have been used to classify a gene as differentially expressed in microarray experiments. They usually require a number of experimental observations to be adequately applied. However, many microarray experiments are constrained to low replication designs for different reasons, from financial restrictions to scarcely available RNA samples. Although performed in a high-throughput framework, there are few experimental replicas for each gene to allow the use of traditional or state-of-art statistical methods. In this work, we present a web-based bioinformatics tool that deals with real-life problems concerning low replication experiments. It uses an empirically derived criterion to classify a gene as differentially expressed by combining two widely accepted ideas in microarray analysis: self-self experiments to derive intensity-dependent cutoffs and non-parametric estimation techniques. To help laboratories without a bioinformatics infrastructure, we implemented the tool in a user-friendly website (http://blasto.iq.usp.br/~rvencio/HTself).

Algorithms↗

Protein structure prediction servers at University College London.

A number of state-of-the-art protein structure prediction servers have been developed by researchers working in the Bioinformatics Unit at University College London. The popular PSIPRED server allows users to perform secondary structure prediction, transmembrane topology prediction and protein fold recognition. More recent servers include DISOPRED for the prediction of protein dynamic disorder and DomPred for domain boundary prediction. These servers are available from our software home page at http://bioinf.cs.ucl.ac.uk/software.html.

Computational Biology↗