Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Exploring protein domain structure.

The protein databank contains coordinates of over 10,000 protein structures, which constitute more than 25,000 structural domains in total. The investigation of protein structural, functional and evolutionary relationships is fundamental to many important fields in bioinformatics research, and will be crucial in determining the function of the human and other genomes. This review describes the SCOP and CATH databases of protein structure classification, which define, classify and annotate each domain in the protein databank. The hierarchical structure, use and annotation of the databases are explained. Other tools for exploring protein structure relationships are also described.

Computational Biology↗

Characterization of renal allograft rejection by urinary proteomic analysis.

OBJECTIVE: To develop a diagnostic method with no morbidity or mortality for the detection of acute renal transplant rejection. SUMMARY BACKGROUND DATA: Rejection constitutes the major impediment to the success of transplantation. Currently available methods, including clinical presentation and biochemical organ function parameters, often fail to detect rejection until late stages of progression. Renal biopsies have associated morbidity and mortality and provide only a limited sample of the organ. METHODS: Thirty-four urine samples were collected from 32 renal transplant patients at various stages posttransplantation. Samples were collected from 17 transplant recipients with acute rejection and 15 patients with no rejection. Samples from patients less than 4 days posttransplant were omitted from data analysis due to the presence of excessive inflammatory response proteins. Rejection status was confirmed by kidney biopsy. Specimens were analyzed in triplicate using SELDI mass spectrometry. The obtained spectra were subjected to bioinformatic analysis using ProPeak as well as CART (Classification and Regression Tree) algorithms to identify rejection biomarker candidates. These candidates were identified by their molecular weight and ranked by their ability to distinguish between nonrejection and rejection based on receiver operating characteristic (ROC) analysis. The candidates with the highest area under the ROC curve (AUC) exhibited the best diagnostic performance. RESULTS: The best candidate biomarkers demonstrated highly successful diagnostic performance: 6.5 kd (AUC = 0.839, P <.0001), 6.7 kd (AUC = 0.839, P <.0001), 6.6 kd (AUC = 0.807, P <.0001), 7.1 kd (AUC = 0.807, P <.0001), and 13.4 kd (AUC = 0.804, P <.0001). A separate analysis using the CART algorithm in the Ciphergen Biomarker Pattern Software correctly classified 91% of the 34 specimens in the training set, giving a sensitivity of 83% and specificity of 100% using two separate biomarker candidates at 10.0 kd and 3.4 kd. CONCLUSIONS: Biomarker candidates exist in urine that have the ability to distinguish between renal transplant patients with no rejection and those with acute rejection. These biomarker candidates are the basis for development of a noninvasive method of diagnosing acute rejection without the morbidity and mortality associated with needle biopsy. The combination of biomarkers into a panel for diagnosis leads to the possibility of enhanced diagnostic performance.

Acute Disease↗

Non-small cell lung cancer and tumor-educated platelets: screening of biomarkers and construction of a prognostic model.

BACKGROUND: Lung cancer is a leading cause of cancer-related mortality worldwide, emphasizing the urgent need for effective early detection strategies. Traditional Chinese medicine (TCM) provides a unique perspective on tumor pathogenesis, focusing on concepts such as "long-term stasis leading to accumulation". Tumor-educated platelets (TEPs) offer potential as biomarkers due to their ability to reflect cancer heterogeneity and facilitate less invasive diagnostic approaches. This study aims to identify TEP-related prognostic biomarkers for non-small cell lung cancer (NSCLC) and to construct and validate a multigene prognostic model by integrating platelet transcriptomic data with tumor tissue datasets. METHODS: We performed comprehensive analysis of gene expression datasets obtained from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) repositories to characterize transcriptomic differences among lung cancer specimens, normal tissue samples, and TEPs. Using R software, we identified Differentially expressed genes (DEGs) and subsequently applied a multi-stage analytical pipeline to TEP-associated DEGs, incorporating univariate Cox proportional hazards regression, least absolute shrinkage and selection operator (LASSO) regression, multivariate Cox regression, and stepwise regression modeling to pinpoint genes with prognostic significance. These prognostically relevant genes served as the foundation for developing a risk stratification model. We computed individual risk scores across both training and validation cohorts, enabling patient stratification into high- and low-risk categories. Model robustness was assessed through internal cross-validation and external validation procedures, while predictive performance was quantified using risk calibration metrics and receiver operating characteristic (ROC) curve analysis. RESULTS: Through systematic bioinformatics screening, we identified a four-gene prognostic signature comprising NELL2, C4orf48, PRAM1, and KLHL35, which served as the foundation for developing our risk stratification algorithm. Rigorous internal cross-validation and external cohort validation substantiated the moderate predictive performance of this signature. Comprehensive clinicopathological correlation analysis revealed that elevated risk indices, advanced pathological staging (stage III-IV), increased primary tumor dimensions, regional lymph node metastasis, and distant organ dissemination each demonstrated statistically significant associations with diminished overall survival (OS) outcomes in lung cancer patients. The clinical nomogram exhibited acceptable calibration, with calibration plots showing reasonable concordance between predicted and observed survival probabilities across all time points. Discriminative capacity assessment via time-dependent ROC analysis yielded area under the curve (AUC) values consistently surpassing 0.6, confirming moderate prognostic discrimination. Furthermore, decision curve analysis (DCA) demonstrated that our integrated multi-gene model conferred potential net clinical benefit compared to individual prognostic variables across the full spectrum of clinically relevant threshold probabilities (0-1 range), thereby establishing its potential utility for risk-informed clinical decision-making. CONCLUSIONS: This study identified NELL2, C4orf48, PRAM1, and KLHL35 as candidate TEP-related prognostic biomarkers for non-small cell lung cancer (NSCLC). The developed prognostic model shows preliminary potential for patient stratification, but its clinical application, particularly as a platelet-based liquid biopsy tool, requires further validation in independent TEP-based cohorts.

Tumor-educated platelets (TEPs)↗

The World-Wide Web: an interface between research and teaching in bioinformatics.

The rapid expansion occurring in World-Wide Web activity is beginning to make the concepts of 'global hypermedia' and 'universal document readership realistic objectives of the new revolution in information technology. One consequence of this increase in usage is that educators and students are becoming more aware of the diversity of the knowledge base which can be accessed via the Internet. Although computerised databases and information services have long played a key role in bioinformatics these same resources can also be used to provide core materials for teaching and learning. The large datasets and archives that have been compiled for biomedical research can be enhanced with the addition of a variety of multimedia elements (images, digital videos, animation etc.). The use of this digitally stored information in structured and self-directed learning environments is likely to increase as activity across World-Wide Web increases.

Information Systems↗

Design and implementation of an introductory course for computer applications in molecular genetics. A case study.

Formal training in computational biology was initiated at Wayne State University in 1990 to meet the needs of the faculty. This was still at a time when the molecular databases and analysis tools could be housed in what is now equivalent to a modern but dated desktop computer. In 1995 the course was expanded to include graduate students to provide these senior students with a foundation in computational biology. This course has armed our students with a requisite set of basic skills that are necessary for a successful career in molecular genetics. It is now an integral component of the graduate program of the Center for Molecular Medicine and Genetics and our experiences in course delivery have been detailed (BioInformatics Methods and Protocols, S. Misener and S. A. Krawetz, eds., Humana Press, Totowa, NJ, 2000.). The course was expanded to a campus-wide unlimited enrollment program for the summer of 2000 to address the needs of our student body. In this review we present our experience with delivering a multidisciplinary campus-wide computational biology course to a new and widely diverse student body.

Computational Biology↗

Evaluation of ontology development tools for bioinformatics.

Ontologies are being used nowadays in many areas, including bioinformatics. To assist users in developing and maintaining ontologies a number of tools have been developed. In this paper we compare four such tools, Protégé-2000, Chimaera, DAG-Edit and OilEd. As test ontologies we have used ontologies from the Gene Ontology Consortium. No system is preferred in all situations, but each system has its own strengths and weaknesses.

Computational Biology↗

Cluster analysis of an extensive human breast cancer cell line protein expression map database.

In the current study, the protein expression maps (PEMs) of 26 breast cancer cell lines and three cell lines derived from normal breast or benign disease tissue were visualised by high resolution two-dimensional gel electrophoresis. Analysis of this data was performed with ChiClust and ChiMap, two analytical bioinformatics tools that are described here. These tools are designed to facilitate recognition of specific patterns shared by two or more (a series) PEMs. Both tools use PEMs that were matched by an image analysis program and locally written programs to create a match table that is saved in an object relational database. The ChiClust tool uses clustering and subclustering methods to extract statistically significant protein expression patterns from a large series of PEMs. The ChiMap tool calculates a differential value (either as percentage change or a fold change) and represents these graphically. All such differentials or just those identified using ChiClust can be submitted to ChiMap. These methods are not dependent on any particular commercial image analysis program, and the whole software package gives an integrated procedure for the comparison and analysis of a series of PEMs. The ChiClust tool was used here to order the breast cell lines into groups according to biological characteristics including morphology in vitro and tumour forming ability in vivo. ChiMap was then used to highlight eight major protein feature-changes detected between breast cancer cell lines that either do or do not proliferate in nude mice. Mass spectrometry was used to identify the proteins. The possible role of these proteins in cancer is discussed.

Algorithms↗

Design and implementation of a CORBA-based genome mapping system prototype.

MOTIVATION: CORBA (Common Object Request Broker Architecture), as an open standard, is considered to be a good solution for the development and deployment of applications in distributed heterogeneous environments. This technology can be applied in the bioinformatics area to enhance utilization, management and interoperation between biological resources. RESULTS: This paper investigates issues in developing CORBA applications for genome mapping information systems in the Internet environment with emphasis on database connectivity and graphical user interfaces. The design and implementation of a CORBA prototype for an animal genome mapping database are described. AVAILABILITY: The prototype demonstration is available via: http://www.ri.bbsrc.ac.uk/ark_corba/. CONTACT: jian.hu@bbsrc.ac.uk

Chromosome Mapping↗

Prediction of MHC class II-binding peptides using an evolutionary algorithm and artificial neural network.

MOTIVATION: Prediction methods for identifying binding peptides could minimize the number of peptides required to be synthesized and assayed, and thereby facilitate the identification of potential T-cell epitopes. We developed a bioinformatic method for the prediction of peptide binding to MHC class II molecules. RESULTS: Experimental binding data and expert knowledge of anchor positions and binding motifs were combined with an evolutionary algorithm (EA) and an artificial neural network (ANN): binding data extraction --> peptide alignment --> ANN training and classification . This method, termed PERUN, was implemented for the prediction of peptides that bind to HLA-DR4(B1*0401). The respective positive predictive values of PERUN predictions of high-, moderate-, low- and zero-affinity binders were assessed as 0.8, 0.7, 0.5 and 0.8 by cross-validation, and 1.0, 0.8, 0.3 and 0.7 by experimental binding. This illustrates the synergy between experimentation and computer modeling, and its application to the identification of potential immunotherapeutic peptides. AVAILABILITY: Software and data are available from the authors upon request. CONTACT: vladimir@wehi.edu. au

Algorithms↗

Ontology-based knowledge representation for bioinformatics.

Much of biology works by applying prior knowledge ('what is known') to an unknown entity, rather than the application of a set of axioms that will elicit knowledge. In addition, the complex biological data stored in bioinformatics databases often require the addition of knowledge to specify and constrain the values held in that database. One way of capturing knowledge within bioinformatics applications and databases is the use of ontologies. An ontology is the concrete form of a conceptualisation of a community's knowledge of a domain. This paper aims to introduce the reader to the use of ontologies within bioinformatics. A description of the type of knowledge held in an ontology will be given.The paper will be illustrated throughout with examples taken from bioinformatics and molecular biology, and a survey of current biological ontologies will be presented. From this it will be seen that the use to which the ontology is put largely determines the content of the ontology. Finally, the paper will describe the process of building an ontology, introducing the reader to the techniques and methods currently in use and the open research questions in ontology development.

Artificial Intelligence↗

ReGAIN: a bioinformatics platform for assessing probabilistic co-occurrence between resistance genes in bacterial pathogens.

MOTIVATION: Multidrug-resistant bacterial pathogens continue to rise globally, yet scalable methods are needed to infer how resistance determinants co-occur across pathogen populations and to quantify conditional dependencies underlying co-occurrence and shared genetic context. RESULTS: We present ReGAIN (Resistance Gene Association and Inference Network), an open-source platform that applies Bayesian network structure learning to infer probabilistic, conditional dependency relationships among antibiotic resistance, heavy metal tolerance, stress response, and virulence determinants in bacteria. In contrast to pairwise co-occurrence analyses, ReGAIN reports conditional probabilities, relative risks, and absolute risk differences with confidence intervals to prioritize candidate relationships for downstream prioritization. Applied across ESKAPEE pathogens, ReGAIN recapitulated established resistance gene relationships and identified additional candidate patterns consistent with co-selection and shared genetic context. Together, these results support scalable, reproducible population-wide analysis of resistance networks for surveillance, comparative genomics and epidemiology. AVAILABILITY: ReGAIN analyses are performed using Python v3.11.5 and R v4.4.1 and is available as open-source software through Bioconda at {https://anaconda.org/bioconda/regain-cli}. Source code and documentation can be found at {https://github.com/ERBringHorvath/regain_CLI}. All genomes used in this publication were downloaded from the National Center for Biotechnology Information database. Large supplementary tables and results data from the ESKAPEE pathogen example network analyses can be downloaded from https://figshare.com/articles/dataset/ReGAIN_command_line_software_and_supplemental_figures_/28959431.

Computational Biology↗

The apoptosis database.

The apoptosis database is a public resource for researchers and students interested in the molecular biology of apoptosis. The resource provides functional annotation, literature references, diagrams/images, and alternative nomenclatures on a set of proteins having 'apoptotic domains'. These are the distinctive domains that are often, if not exclusively, found in proteins involved in apoptosis. The initial choice of proteins to be included is defined by apoptosis experts and bioinformatics tools. Users can browse through the web accessible lists of domains, proteins containing these domains and their associated homologs. The database can also be searched by sequence homology using basic local alignment search tool, text word matches of the annotation, and identifiers for specific records. The resource is available at http://www.apoptosis-db.org and is updated on a regular basis.

Animals↗

CHOMPER: a bioinformatic tool for rapid validation of tandem mass spectrometry search results associated with high-throughput proteomic strategies.

Current efforts aimed at developing high-throughput proteomics focus on increasing the speed of protein identification. Although improvements in sample separation, enrichment, automated handling, mass spectrometric analysis, as well as data reduction and database interrogation strategies have done much to increase the quality, quantity and efficiency of data collection, significant bottlenecks still exist. Various separation techniques have been coupled with tandem mass spectrometric (MS/MS) approaches to allow a quicker analysis of complex mixtures of proteins, especially where a high number of unambiguous protein identifications are the exception, rather than the rule. MS/MS is required to provide structural / amino acid sequence information on a peptide and thus allow protein identity to be inferred from individual peptides. Currently these spectra need to be manually validated because: (a) the potential of false positive matches i.e., protein not in database, and (b) observed fragmentation trends may not be incorporated into current MS/MS search algorithms. This validation represents a significant bottleneck associated with high-throughput proteomic strategies. We have developed CHOMPER, a software program which reduces the time required to both visualize and confirm MS/MS search results and generate post-analysis reports and protein summary tables. CHOMPER extracts the identification information from SEQUEST MS/MS search result files, reproduces both the peptide and protein identification summaries, provides a more interactive visualization of the MS/MS spectra and facilitates the direct submission of manually validated identifications to a database.

Algorithms↗

Recent advances in computational genomics.

In the post-genomic era, the new discipline of functional genomics is now facing the challenge of associating a function (as well as estimating its relevance to industrial applications) to about 100,000 microbial, plant or animal genes of known sequence but unknown function. Besides the design of databases, computational methods are increasingly becoming intimately linked with the various experimental approaches. Consequently, bioinformatics is rapidly evolving into independent fields addressing the specific problems of interpreting i) genomic sequences, ii) protein sequences and 3D-structures, as well as iii) transcriptome and macromolecular interaction data. It is thus increasingly difficult for the biologist to choose the computational approaches that perform best in these various areas. This paper attempts to review the most useful developments of the last 2 years.

Computational Biology↗

Talisman--rapid application development for the grid.

In order to make use of the emerging grid and network services offered by various institutes and mandated by many current research projects, some kind of user accessible client is required. In contrast with attempts to build generic workbenches, Talisman is designed to allow a bioinformatics expert to rapidly build custom applications, immediately visible using standard web technology, for users who wish to concentrate on the biology of their problem rather than the informatics aspects. As a component of the MyGrid project, it is intended to allow access to arbitrary resources, including but not limited to relational, object and flat file data sources, analysis programs and grid based storage, tracking and distributed annotation systems.

Computational Biology↗

cDNA2Genome: a tool for mapping and annotating cDNAs.

BACKGROUND: In the last years several high-throughput cDNA sequencing projects have been funded worldwide with the aim of identifying and characterizing the structure of complete novel human transcripts. However some of these cDNAs are error prone due to frameshifts and stop codon errors caused by low sequence quality, or to cloning of truncated inserts, among other reasons. Therefore, accurate CDS prediction from these sequences first require the identification of potentially problematic cDNAs in order to speed up the posterior annotation process. RESULTS: cDNA2Genome is an application for the automatic high-throughput mapping and characterization of cDNAs. It utilizes current annotation data and the most up to date databases, especially in the case of ESTs and mRNAs in conjunction with a vast number of approaches to gene prediction in order to perform a comprehensive assessment of the cDNA exon-intron structure. The final result of cDNA2Genome is an XML file containing all relevant information obtained in the process. This XML output can easily be used for further analysis such us program pipelines, or the integration of results into databases. The web interface to cDNA2Genome also presents this data in HTML, where the annotation is additionally shown in a graphical form. cDNA2Genome has been implemented under the W3H task framework which allows the combination of bioinformatics tools in tailor-made analysis task flows as well as the sequential or parallel computation of many sequences for large-scale analysis. CONCLUSIONS: cDNA2Genome represents a new versatile and easily extensible approach to the automated mapping and annotation of human cDNAs. The underlying approach allows sequential or parallel computation of sequences for high-throughput analysis of cDNAs.

Chromosome Mapping↗

Database of p53 gene somatic mutations in human tumors and cell lines: updated compilation and future prospects.

In recent years, there has been an exponential increase in the number of p53 mutations identified in human cancers. The p53 mutation database consists of a list of point mutations in thep53 gene of human tumors and cell lines, compiled from the published literature and made available through electronic media. The database is now maintained at the International Agency for Research on Cancer (IARC) and is updated twice a year. The current version contains records on 5091 published mutations and is expected to surpass the 6000 mark in the January 1997 release. The database is available in various formats through the European Bioinformatics Institute (EBI) ftp server at: ftp://ftp.ebi.ac.uk/pub/databases/p53/ or by request from IARC (p53database@iarc.fr) and will be searchable through the SRS system in the near future. This report provides a description of the criteria for inclusion of data and of the current formats, a summary of the relevance ofp53 mutation analysis to clinical and biological questions, and a brief discussion of the prospects for future developments.

Databases, Factual↗

GeneExpress: a computer system for description, analysis, and recognition of regulatory sequences in eukaryotic genome.

GeneExpress system has been designed to integrate description, analysis, and recognition of eukaryotic regulatory sequences. The system includes 5 basic units: (1) GeneNet contains an object-oriented database for accumulation of data on gene networks and signal transduction pathways and a Java-based viewer that allows an exploration and visualization of the GeneNet information; (2) Transcription Regulation combines the database on transcription regulatory regions of eukaryotic genes (TRRD) and TRRD Viewer; (3) Transcription Factor Binding Site Recognition contains a compilation of transcription factor binding sites (TFBSC) and programs for their analysis and recognition; (4) mRNA Translation is designed for analysis of structural and contextual features of mRNA 5'UTRs and prediction of their translation efficiency; and (5) ACTIVITY is the module for analysis and site activity prediction of a given nucleotide sequence. Integration of the databases in the GeneExpress is based on the Sequence Retrieval System (SRS) created in the European Bioinformatics Institute.

Artificial Intelligence↗