Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Design and implementation of an introductory course for computer applications in molecular genetics. A case study.

Formal training in computational biology was initiated at Wayne State University in 1990 to meet the needs of the faculty. This was still at a time when the molecular databases and analysis tools could be housed in what is now equivalent to a modern but dated desktop computer. In 1995 the course was expanded to include graduate students to provide these senior students with a foundation in computational biology. This course has armed our students with a requisite set of basic skills that are necessary for a successful career in molecular genetics. It is now an integral component of the graduate program of the Center for Molecular Medicine and Genetics and our experiences in course delivery have been detailed (BioInformatics Methods and Protocols, S. Misener and S. A. Krawetz, eds., Humana Press, Totowa, NJ, 2000.). The course was expanded to a campus-wide unlimited enrollment program for the summer of 2000 to address the needs of our student body. In this review we present our experience with delivering a multidisciplinary campus-wide computational biology course to a new and widely diverse student body.

Computational Biology↗

Identification of mouse retinal genes differentially regulated by dim and bright cyclic light rearing.

Bright cyclic light rearing protects BALB/c mice from light-induced photoreceptor apoptosis compared to dim cyclic light rearing. We used a microarray approach to search for putative neuroprotection genes that were up- or down-regulated under these environmental conditions. Retinal protection by bright cyclic rearing was determined by quantitative histology and DNA fragmentation analysis. Total RNA was isolated from 5-week-old mice raised in bright (400 lux) or dim (5 lux) cyclic light and prepared for analysis on microarrays produced using a 70-mer oligonucleotide library that represented 16,463 mouse genes. Genes of interest were identified using statistically robust bioinformatics analysis methods that were developed in-house. Changes in some genes were confirmed with quantitative real time PCR. We found that 952 genes were up- or down-regulated by bright cyclic light rearing compared to dim cyclic light rearing. One hundred and eighty-four of them, having >/=2-fold differences, were grouped into 13 categories, and selected for further consideration. Eleven up-regulated and two down-regulated genes were confirmed by semi-quantitative PCR. Five neuroprotection-associated genes were up-regulated by bright cyclic light rearing as confirmed by real-time PCR. The human orthologue chromosomal location of 22 differentially expressed genes map to known retinal degeneration loci. Using PathwayAssist software, we modeled the pathway networks of up- and down-regulated genes that are functionally related to the retina. We identified retinal genes that are differentially regulated by environmental light history. Those that directly affect cell processes such as survival, apoptosis, and transcription are likely play a pivotal role in the regulation of retinal neuroprotection against light-induced photoreceptor apoptosis.

Adaptation, Ocular↗

BTW: a web server for Boltzmann time warping of gene expression time series.

UNLABELLED: Dynamic time warping (DTW) is a well-known quadratic time algorithm to determine the smallest distance and optimal alignment between two numerical sequences, possibly of different length. Originally developed for speech recognition, this method has been used in data mining, medicine and bioinformatics. For gene expression time series data, time warping distance is arguably a more flexible tool to determine genes having similar temporal expression, hence possibly related biological function, than either Euclidean distance or correlation coefficient--especially since time warping accommodates sequences of different length. The BTW web server allows a user to upload two tab-separated text files A,B of gene expression data, each possibly having a different number of time intervals of different durations. BTW then computes time warping distance between each gene of A with each gene of B, using a recently developed symmetric algorithm which additionally computes the Boltzmann partition function and outputs Boltzmann pair probabilities. The Boltzmann pair probabilities, not available with any other existent software, suggest possible biological significance of certain positions in an optimal time warping alignment. AVAILABILITY: http://bioinformatics.bc.edu/clotelab/BTW/.

Algorithms↗

Web-ware bioinformatical analysis and structure modelling of N-terminus of human multisynthetase complex auxiliary component protein p43.

Human multisynthetase complex auxiliary component, protein p43 is an endothelial monocyte-activating polypeptide II precursor. In this study, comprehensive sequence analysis of N-terminus has been performed to identify structural domains, motifs, sites of post-translation modification and other functionally important parameters. The spatial structure model of full-chain protein p43 is obtained.

Amino Acid Motifs↗

Evolution of web services in bioinformatics.

Bioinformaticians have developed large collections of tools to make sense of the rapidly growing pool of molecular biological data. Biological systems tend to be complex and in order to understand them, it is often necessary to link many data sets and use more than one tool. Therefore, bioinformaticians have experimented with several strategies to try to integrate data sets and tools. Owing to the lack of standards for data sets and the interfaces of the tools this is not a trivial task. Over the past few years building services with web-based interfaces has become a popular way of sharing the data and tools that have resulted from many bioinformatics projects. This paper discusses the interoperability problem and how web services are being used to try to solve it, resulting in the evolution of tools with web interfaces from HTML/web form-based tools not suited for automatic workflow generation to a dynamic network of XML-based web services that can easily be used to create pipelines.

Computational Biology↗

Evaluation of ontology development tools for bioinformatics.

Ontologies are being used nowadays in many areas, including bioinformatics. To assist users in developing and maintaining ontologies a number of tools have been developed. In this paper we compare four such tools, Protégé-2000, Chimaera, DAG-Edit and OilEd. As test ontologies we have used ontologies from the Gene Ontology Consortium. No system is preferred in all situations, but each system has its own strengths and weaknesses.

Computational Biology↗

Combined analysis of expression data and transcription factor binding sites in the yeast genome.

BACKGROUND: The analysis of gene expression using DNA microarrays provides genome wide profiles of the genes controlled by the presence or absence of a specific transcription factor. However, the question arises of whether a change in the level of transcription of a specific gene is caused by the transcription factor acting directly at the promoter of the gene or through regulation of other transcription factors working at the promoter. RESULTS: To address this problem we have devised a computational method that combines microarray expression and site preference data. We have tested this approach by identifying functional targets of the a1-alpha2 complex, which represses haploid-specific genes in the yeast Saccharomyces cerevisiae. Our analysis identified many known or suspected haploid-specific genes that are direct targets of the a1-alpha2 complex, as well as a number of previously uncharacterized targets. We were also able to identify a number of haploid-specific genes which do not appear to be direct targets of the a1-alpha2 complex, as well as a1-alpha2 target sites that do not repress transcription of nearby genes. Our method has a much lower false positive rate when compared to some of the conventional bioinformatic approaches. CONCLUSIONS: These findings show advantages of combining these two forms of data to investigate the mechanism of co-regulation of specific sets of genes.

Algorithms↗

Cluster analysis of an extensive human breast cancer cell line protein expression map database.

In the current study, the protein expression maps (PEMs) of 26 breast cancer cell lines and three cell lines derived from normal breast or benign disease tissue were visualised by high resolution two-dimensional gel electrophoresis. Analysis of this data was performed with ChiClust and ChiMap, two analytical bioinformatics tools that are described here. These tools are designed to facilitate recognition of specific patterns shared by two or more (a series) PEMs. Both tools use PEMs that were matched by an image analysis program and locally written programs to create a match table that is saved in an object relational database. The ChiClust tool uses clustering and subclustering methods to extract statistically significant protein expression patterns from a large series of PEMs. The ChiMap tool calculates a differential value (either as percentage change or a fold change) and represents these graphically. All such differentials or just those identified using ChiClust can be submitted to ChiMap. These methods are not dependent on any particular commercial image analysis program, and the whole software package gives an integrated procedure for the comparison and analysis of a series of PEMs. The ChiClust tool was used here to order the breast cell lines into groups according to biological characteristics including morphology in vitro and tumour forming ability in vivo. ChiMap was then used to highlight eight major protein feature-changes detected between breast cancer cell lines that either do or do not proliferate in nude mice. Mass spectrometry was used to identify the proteins. The possible role of these proteins in cancer is discussed.

Algorithms↗

Design and implementation of a CORBA-based genome mapping system prototype.

MOTIVATION: CORBA (Common Object Request Broker Architecture), as an open standard, is considered to be a good solution for the development and deployment of applications in distributed heterogeneous environments. This technology can be applied in the bioinformatics area to enhance utilization, management and interoperation between biological resources. RESULTS: This paper investigates issues in developing CORBA applications for genome mapping information systems in the Internet environment with emphasis on database connectivity and graphical user interfaces. The design and implementation of a CORBA prototype for an animal genome mapping database are described. AVAILABILITY: The prototype demonstration is available via: http://www.ri.bbsrc.ac.uk/ark_corba/. CONTACT: jian.hu@bbsrc.ac.uk

Chromosome Mapping↗

ParPEST: a pipeline for EST data analysis based on parallel computing.

BACKGROUND: Expressed Sequence Tags (ESTs) are short and error-prone DNA sequences generated from the 5' and 3' ends of randomly selected cDNA clones. They provide an important resource for comparative and functional genomic studies and, moreover, represent a reliable information for the annotation of genomic sequences. Because of the advances in biotechnologies, ESTs are daily determined in the form of large datasets. Therefore, suitable and efficient bioinformatic approaches are necessary to organize data related information content for further investigations. RESULTS: We implemented ParPEST (Parallel Processing of ESTs), a pipeline based on parallel computing for EST analysis. The results are organized in a suitable data warehouse to provide a starting point to mine expressed sequence datasets. The collected information is useful for investigations on data quality and on data information content, enriched also by a preliminary functional annotation. CONCLUSION: The pipeline presented here has been developed to perform an exhaustive and reliable analysis on EST data and to provide a curated set of information based on a relational database. Moreover, it is designed to reduce execution time of the specific steps required for a complete analysis using distributed processes and parallelized software. It is conceived to run on low requiring hardware components, to fulfill increasing demand, typical of the data used, and scalability at affordable costs.

Algorithms↗

Prediction of MHC class II-binding peptides using an evolutionary algorithm and artificial neural network.

MOTIVATION: Prediction methods for identifying binding peptides could minimize the number of peptides required to be synthesized and assayed, and thereby facilitate the identification of potential T-cell epitopes. We developed a bioinformatic method for the prediction of peptide binding to MHC class II molecules. RESULTS: Experimental binding data and expert knowledge of anchor positions and binding motifs were combined with an evolutionary algorithm (EA) and an artificial neural network (ANN): binding data extraction --> peptide alignment --> ANN training and classification . This method, termed PERUN, was implemented for the prediction of peptides that bind to HLA-DR4(B1*0401). The respective positive predictive values of PERUN predictions of high-, moderate-, low- and zero-affinity binders were assessed as 0.8, 0.7, 0.5 and 0.8 by cross-validation, and 1.0, 0.8, 0.3 and 0.7 by experimental binding. This illustrates the synergy between experimentation and computer modeling, and its application to the identification of potential immunotherapeutic peptides. AVAILABILITY: Software and data are available from the authors upon request. CONTACT: vladimir@wehi.edu. au

Algorithms↗

Ontology-based knowledge representation for bioinformatics.

Much of biology works by applying prior knowledge ('what is known') to an unknown entity, rather than the application of a set of axioms that will elicit knowledge. In addition, the complex biological data stored in bioinformatics databases often require the addition of knowledge to specify and constrain the values held in that database. One way of capturing knowledge within bioinformatics applications and databases is the use of ontologies. An ontology is the concrete form of a conceptualisation of a community's knowledge of a domain. This paper aims to introduce the reader to the use of ontologies within bioinformatics. A description of the type of knowledge held in an ontology will be given.The paper will be illustrated throughout with examples taken from bioinformatics and molecular biology, and a survey of current biological ontologies will be presented. From this it will be seen that the use to which the ontology is put largely determines the content of the ontology. Finally, the paper will describe the process of building an ontology, introducing the reader to the techniques and methods currently in use and the open research questions in ontology development.

Artificial Intelligence↗

KISS for STRAP: user extensions for a protein alignment editor.

SUMMARY: The Structural Alignment Program STRAP is a comfortable comprehensive editor and analyzing tool for protein alignments. A wide range of functions related to protein sequences and protein structures are accessible with an intuitive graphical interface. Recent features include mapping of mutations and polymorphisms onto structures and production of high quality figures for publication. Here we address the general problem of multi-purpose program packages to keep up with the rapid development of bioinformatical methods and the demand for specific program functions. STRAP was remade implementing a novel design which aims at Keeping Interfaces in STRAP Simple (KISS). KISS renders STRAP extendable to bio-scientists as well as to bio-informaticians. Scientists with basic computer skills are capable of implementing statistical methods or embedding existing bioinformatical tools in STRAP themselves. For bio-informaticians STRAP may serve as an environment for rapid prototyping and testing of complex algorithms such as automatic alignment algorithms or phylogenetic methods. Further, STRAP can be applied as an interactive web applet to present data related to a particular protein family and as a teaching tool. REQUIREMENTS: JAVA-1.4 or higher. AVAILABILITY: http://www.charite.de/bioinf/strap/

Computer Graphics↗

CEP: a conformational epitope prediction server.

CEP server (http://bioinfo.ernet.in/cep.htm) provides a web interface to the conformational epitope prediction algorithm developed in-house. The algorithm, apart from predicting conformational epitopes, also predicts antigenic determinants and sequential epitopes. The epitopes are predicted using 3D structure data of protein antigens, which can be visualized graphically. The algorithm employs structure-based Bioinformatics approach and solvent accessibility of amino acids in an explicit manner. Accuracy of the algorithm was found to be 75% when evaluated using X-ray crystal structures of Ag-Ab complexes available in the PDB. This is the first and the only method available for the prediction of conformational epitopes, which is an attempt to map probable antibody-binding sites of protein antigens.

Algorithms↗

The MicrobesOnline Web site for comparative genomics.

At present, hundreds of microbial genomes have been sequenced, and hundreds more are currently in the pipeline. The Virtual Institute for Microbial Stress and Survival has developed a publicly available suite of Web-based comparative genomic tools (http://www.microbesonline.org) designed to facilitate multispecies comparison among prokaryotes. Highlights of the MicrobesOnline Web site include operon and regulon predictions, a multispecies genome browser, a multispecies Gene Ontology browser, a comparative KEGG metabolic pathway viewer, a Bioinformatics Workbench for in-depth sequence analysis, and Gene Carts that allow users to save genes of interest for further study while they browse. In addition, we provide an interface for genome annotation, which like all of the tools reported here, is freely available to the scientific community.

Animals↗

Sequence handling by sequence analysis toolbox v1.0.

The fact that mass spectrometry have become a high-throughput method calls for bioinformatic tools for automated sequence handling and prediction. For efficient use of bioinformatic tools, it is important that these tools are integrated or interfaced with each other. The purpose of sequence analysis toolbox v1.0 was to have a general purpose sequence analyzing tool that can import sequences obtained by high-throughput sequencing methods. The program includes algorithms for calculation or prediction of isoelectric point, hydropathicity index, transmembrane segments, and glycosylphosphatidyl inositol-anchored proteins.

Computational Biology↗

ReGAIN: a bioinformatics platform for assessing probabilistic co-occurrence between resistance genes in bacterial pathogens.

MOTIVATION: Multidrug-resistant bacterial pathogens continue to rise globally, yet scalable methods are needed to infer how resistance determinants co-occur across pathogen populations and to quantify conditional dependencies underlying co-occurrence and shared genetic context. RESULTS: We present ReGAIN (Resistance Gene Association and Inference Network), an open-source platform that applies Bayesian network structure learning to infer probabilistic, conditional dependency relationships among antibiotic resistance, heavy metal tolerance, stress response, and virulence determinants in bacteria. In contrast to pairwise co-occurrence analyses, ReGAIN reports conditional probabilities, relative risks, and absolute risk differences with confidence intervals to prioritize candidate relationships for downstream prioritization. Applied across ESKAPEE pathogens, ReGAIN recapitulated established resistance gene relationships and identified additional candidate patterns consistent with co-selection and shared genetic context. Together, these results support scalable, reproducible population-wide analysis of resistance networks for surveillance, comparative genomics and epidemiology. AVAILABILITY: ReGAIN analyses are performed using Python v3.11.5 and R v4.4.1 and is available as open-source software through Bioconda at {https://anaconda.org/bioconda/regain-cli}. Source code and documentation can be found at {https://github.com/ERBringHorvath/regain_CLI}. All genomes used in this publication were downloaded from the National Center for Biotechnology Information database. Large supplementary tables and results data from the ESKAPEE pathogen example network analyses can be downloaded from https://figshare.com/articles/dataset/ReGAIN_command_line_software_and_supplemental_figures_/28959431.

Computational Biology↗

Serum protein fingerprinting coupled with artificial neural network distinguishes glioma from healthy population or brain benign tumor.

To screen and evaluate protein biomarkers for the detection of gliomas (Astrocytoma grade I-IV) from healthy individuals and gliomas from brain benign tumors by using surface enhanced laser desorption/ionization time of flight mass spectrometry (SELDI-TOF-MS) coupled with an artificial neural network (ANN) algorithm. SELDI-TOF-MS protein fingerprinting of serum from 105 brain tumor patients and healthy individuals, included 28 patients with glioma (Astrocytoma I-IV), 37 patients with brain benign tumor, and 40 age-matched healthy individuals. Two thirds of the total samples of every compared pair as training set were used to set up discriminating patterns, and one third of total samples of every compared pair as test set were used to cross-validate; simultaneously, discriminate-cluster analysis derived SPSS 10.0 software was used to compare Astrocytoma grade I-II with grade III-IV ones. An accuracy of 95.7%, sensitivity of 88.9%, specificity of 100%, positive predictive value of 90% and negative predictive value of 100% were obtained in a blinded test set comparing gliomas patients with healthy individuals; an accuracy of 86.4%, sensitivity of 88.9%, specificity of 84.6%, positive predictive value of 90% and negative predictive value of 85.7% were obtained when patient's gliomas was compared with benign brain tumor. Total accuracy of 85.7%, accuracy of grade I-II Astrocytoma was 86.7%, accuracy of III-IV Astrocytoma was 84.6% were obtained when grade I-II Astrocytoma was compared with grade III-IV ones (discriminant analysis). SELDI-TOF-MS combined with bioinformatics tools, could greatly facilitate the discovery of better biomarkers. The high sensitivity and specificity achieved by the use of selected biomarkers showed great potential application for the discrimination of gliomas patients from healthy individuals and gliomas from brain benign tumors.

Adult↗