Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

BIBI, a bioinformatics bacterial identification tool.

BIBI was designed to automate DNA sequence analysis for bacterial identification in the clinical field. BIBI relies on the use of BLAST and CLUSTAL W programs applied to different subsets of sequences extracted from GenBank. These sequences are filtered and stored in a new database, which is adapted to bacterial identification.

Bacteria↗

A computer system to perform structure comparison using TOPS representations of protein structure.

We describe the design and implementation of a fast topology-based method for protein structure comparison. The approach uses the TOPS topological representation of protein structure, aligning two structures using a common discovered pattern and generating measure of distance derived from an insert score. Heavy use is made of a constraint-based pattern-matching algorithm for TOPS diagrams that we have designed and described elsewhere (Bioinformatics 15(4) (1999) 317). The comparison system is maintained at the European Bioinformatics Institute and is available over the Web at tops.ebi.ac.uk/tops. Users submit a structure description in Protein Data Bank (PDB) format and can compare it with structures in the entire PDB or a representative subset of protein domains, receiving the results by email.

Algorithms↗

Cambridge Healthtech Institute's Third Annual Conference on Lab-on-a-Chip and Microarrays. 22-24 January 2001, Zurich, Switzerland.

Cambridge Healthtech Institute's Third Annual Conference on Lab-on-a-Chip and Microarray technology covered the latest advances in this technology and applications in life sciences. Highlights of the meetings are reported briefly with emphasis on applications in genomics, drug discovery and molecular diagnostics. There was an emphasis on microfluidics because of the wide applications in laboratory and drug discovery. The lab-on-a-chip provides the facilities of a complete laboratory in a hand-held miniature device. Several microarray systems have been used for hybridisation and detection techniques. Oligonucleotide scanning arrays provide a versatile tool for the analysis of nucleic acid interactions and provide a platform for improving the array-based methods for investigation of antisense therapeutics. A method for analysing combinatorial DNA arrays using oligonucleotide-modified gold nanoparticle probes and a conventional scanner has considerable potential in molecular diagnostics. Various applications of microarray technology for high-throughput screening in drug discovery and single nucleotide polymorphisms (SNP) analysis were discussed. Protein chips have important applications in proteomics. With the considerable amount of data generated by the different technologies using microarrays, it is obvious that the reading of the information and its interpretation and management through the use of bioinformatics is essential. Various techniques for data analysis were presented. Biochip and microarray technology has an essential role to play in the evolving trends in healthcare, which integrate diagnosis with prevention/treatment and emphasise personalised medicines.

Computational Biology↗

Tutorial section: domains and motifs - proteins in bite-sized chunks.

Possibly the ultimate goal of bioinformatics is to be able to predict protein tertiary structure and chemical functionality from the initial amino acid sequence. Despite the best efforts of many researchers over the past two decades, a reliable modelling method has yet to be found and the folding problem continues to be a hurdle for scientists.

Amino Acid Motifs↗

Deciphering regulatory patterns of inflammatory gene expression from interleukin-1-stimulated human endothelial cells.

OBJECTIVE: Endothelial cells comprise a key component of the inflammatory response. We set out to obtain a comprehensive overview of the immediate-early to early gene expression program of interleukin-1 (IL-1)-stimulated endothelial cells and to identify novel transcription factors and regulatory elements. METHODS AND RESULTS: Human umbilical vein endothelial cells (HUVECs) were stimulated with IL-1 for 0, 0.5, 1, 2.5, and 6 hours and analyzed using Affymetrix U133 microarrays. A total of 137 genes were found to be regulated >4-fold, including 18 transcription factors. The expression of selected genes was confirmed by real-time polymerase chain reaction. Cluster analysis was performed in order to group genes according to their expression profiles. To identify novel transcription factor-binding sites, the corresponding promoters were extracted from databases and analyzed for regulatory elements that were over-represented in specific clusters. Several potentially novel DNA binding sites were identified, and one was shown to specifically bind an IL-1-inducible protein from HUVEC. CONCLUSIONS: These results demonstrate that in the early phase after stimulation, IL-1 evokes a complex gene expression program that includes positive but also negative (feedback) regulators of diverse endothelial cell functions. Furthermore, the identification of a new promoter regulatory element demonstrates the feasibility of the bioinformatics-driven approach to discover novel regulatory mechanisms.

Binding Sites↗

Development under extreme conditions: forensic bioinformatics in the wake of the World Trade Center disaster.

The terrorist attacks of September 11, 2001 resulted in death and devastation in three locations, and extraordinary efforts have been exerted to identify the remains of all victims. As mass fatalities go, this one has been unusual at a policy level because the goal has been not merely to identify remains for every decedent, but to identify every bit of remains found so that even small pieces of tissue can be returned to families for burial. While the human impact at the Pentagon and Shanksville, PA was horrific, the World Trade Center site presented a particularly complex challenge for forensic DNA matching and data handling. A complete and definitive list of all those killed is still elusive, and human remains were crushed and co-mingled by the falling towers. Software tools had never been considered for a problem of this scale and scope. New data handling systems had to be created under extreme software development conditions characterized by incomplete requirements specifications, chaotically changing priorities, truly impossible deadlines and rapidly rolling production releases. Partly because of the company's experience with mtDNA tools built for the Armed Forces DNA Identification Lab starting in 1997, the New York City Office of Chief Medical Examiner [OCME] contacted Gene Codes Corporation in late September as existing data-handling tools began to fail. We began work on the project in mid-October, 2001. Our approach to the problem included: Extreme Programming [XP] methodology for functional software development, On-site time and motion analysis at the OCME for user interface design, Evidentiary references between STR, SNP and mtDNA analysis results, and Separate data Quality Control [QC] and software Quality Assurance [QA] initiatives. A substantial software suite was developed called M-FISys, an acronym for Mass-Fatality Identification System.

Computational Biology↗

PyEvoMotion: a Python tool for population-based time-course analysis of genome evolution.

SUMMARY: We present PyEvoMotion, an open-source Python tool for inferring molecular clock models with time-dependent Gaussian noise from high-throughput genomic datasets. PyEvoMotion features a command-line interface and a modular architecture, allowing seamless integration into larger bioinformatic pipelines. The tool supports customizable filtering, temporal discretization definition, and mutation classification, making it adaptable to diverse research needs. While traditional phylogenetic methods may encounter computational challenges with large datasets, PyEvoMotion can process thousands to millions of sequences to compute statistical parameters associated with a stochastic differential equation model, thereby weighting the genetic variation within the population. Using viral genomic data, we demonstrate its capability to infer evolutionary rates and detect non-Brownian evolutionary motions with subdiffusive behavior. PyEvoMotion shows potential to provide overlooked insights into genome evolution in different contexts. AVAILABILITY AND IMPLEMENTATION: The open source software is available on GitHub at https://github.com/luksgrin/PyEvoMotion and on SourceForge at https://sourceforge.net/projects/pyevomotion.

Software↗

abCRISPR: deep learning-based design of abasic gRNA sequences for specific CRISPR-Cas9 genome editing.

SUMMARY: CRISPR-Cas9 has become a widely used tool for genome editing. However, its off-target cleavage caused by partial sequence matches with guide RNAs (gRNAs) remains a critical limitation. Recently, abasic gRNAs (ØXØ) have been developed to enhance target specificity, but their effects vary depending on the positional sequence context. Here, we present abCRISPR, a deep neural network (DNN) framework for the rational design of ØXØ sequences with minimized off-target activity. abCRISPR leverages informative few-shot training with paired datasets of abasic and unmodified gRNAs, using high-quality random mismatch target libraries, exhaustively sequenced for mismatched off-target substrates (n = 97583) in in vitro CRISPR-Cas9 cleavage experiments. Predicted off-target activities for both abasic and unmodified gRNAs showed strong correlation with experimental data (r ≥ 0.95, 10-fold cross-validation). Notably, these comprehensive training sets provide robust ground-truth negatives, enabling accurate and sensitive prediction of off-targets. For unmodified gRNAs, abCRISPR (AUC = 0.98) was validated to outperform existing deep learning-based methods (AUC = 0.45-0.68). When applied to the human genome, abCRISPR generated ØXØ sequences, covering 58 875 004 potent CRISPR-targetable sites with improved target specificity. Together, this work provides a comprehensive bioinformatics resource for safe and precise CRISPR-Cas9 genome editing. AVAILABILITY AND IMPLEMENTATION: The source code for abCRISPR and training data are available at https://doi.org/10.5281/zenodo.20398246. abCRISPR results for the human genome are available at http://clip.korea.ac.kr/abCRISPR/.

Deep Learning↗

CCRR: a user-friendly platform for analyzing complex chromosomal rearrangements in tumors.

SUMMARY: Complex chromosomal rearrangements in tumors involve intricate genomic alterations that significantly affect gene function and contribute to cancer development. Identifying these events is crucial for cancer research but is often challenging due to the complexity and limitations of existing tools. We developed the Complex Chromosomal Rearrangements Resolver (CCRR), a comprehensive, reproducible, and user-friendly platform for analyzing complex rearrangements in tumors. CCRR integrates multiple SV and CNV detection tools within a Docker container environment, simplifying installation and configuration. It can be easily deployed, automating the execution and merging of results, providing high-confidence consensus SV and CNV calls, allowing researchers to efficiently analyze complex chromosomal rearrangements in tumors without extensive bioinformatics expertise. CCRR also includes a web server for one-click analysis and customized visualization. AVAILABILITY AND IMPLEMENTATION: The CCRR platform is freely available at https://www.ccrr.life. Source code and executables can be accessed at https://github.com/laslk/CCRR. An archived version is available at Zenodo: https://doi.org/10.5281/zenodo.15386513.

Software↗

GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequences.

MOTIVATION: Short tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. RESULTS: Here, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653.

Microsatellite Repeats↗

A simple and fast secondary structure prediction method using hidden neural networks.

MOTIVATION: In this paper, we present a secondary structure prediction method YASPIN that unlike the current state-of-the-art methods utilizes a single neural network for predicting the secondary structure elements in a 7-state local structure scheme and then optimizes the output using a hidden Markov model, which results in providing more information for the prediction. RESULTS: YASPIN was compared with the current top-performing secondary structure prediction methods, such as PHDpsi, PROFsec, SSPro2, JNET and PSIPRED. The overall prediction accuracy on the independent EVA5 sequence set is comparable with that of the top performers, according to the Q3, SOV and Matthew's correlations accuracy measures. YASPIN shows the highest accuracy in terms of Q3 and SOV scores for strand prediction. AVAILABILITY: YASPIN is available on-line at the Centre for Integrative Bioinformatics website (http://ibivu.cs.vu.nl/programs/yaspinwww/) at the Vrije University in Amsterdam and will soon be mirrored on the Mathematical Biology website (http://www.mathbio.nimr.mrc.ac.uk) at the NIMR in London. CONTACT: kxlin@nimr.mrc.ac.uk

Algorithms↗

Proposed classification of cells in the Foundational Model of Anatomy.

A logical and principled representation of cell types and their component parts could serve as a framework for correlating the various ontologies that are emerging in bioinformatics with a focus on cells and subcellular biological entities. In order to address this need we have extended the Foundational Model of Anatomy (FMA)1,2 from macroscopic to cellular and subcellular anatomical entities. The poster will provide a live demonstration of this implementation.

Anatomy↗

Deriving folds of macromolecular complexes through electron cryomicroscopy and bioinformatics approaches.

Intermediate-resolution (7-9A) structures of large macromolecular complexes can be obtained by electron cryomicroscopy. This structural information, combined with bioinformatics data for the individual protein components or domains, can lead to a fold model for the entire complex. Such approaches have been demonstrated with the 6.8 A structure of the rice dwarf virus to derive models for the major capsid shell proteins.

Amino Acid Sequence↗

Bioinformatic approaches to assigning protein function from novel sequence data.

The current pace of functional genomic initiatives and genome sequencing projects has provided researchers with a bewildering array of sequence and biological data to analyze. The disease system-driven approach to identifying key genes frequently identifies nucleotide and protein sequences for which the gene and protein function are not known in sufficient detail to allow informed follow-up. Using a range of bioinformatic tools and sequence-based clues, most of unassigned sequences can now be annotated. This chapter takes as an example an unannotated expressed sequence tag, describing how to identify its related gene, and how to annotate the encoded protein using sequence, profile, and structure-based annotation methodologies.

Computational Biology↗

Integrating Application Programs for Bioinformatics Using a Web Browser.

We have constructed a general framework for integrating application programs with control through a local Web browser. This method is based on a simple inter-process message function from an external process to application programs. Commands to a target program are prepared in a script file, which is parsed by a message dispatcher program. When it is used as a helper application to a Web browser, these messages will be sent from the browser by clicking a hyper-link in a Web document. Our framework also supports pluggable extension-modules for application programs by means of dynamic linking. A prototype system is implemented on our molecular structure-viewer program, MOSBY. It successfully featured a function to load an extension-module required for the docking study of molecular fragments from a Web page. Our simple framework facilitates the concise configuration of Web softwares without complicated knowledge on network computation and security issues. It is also applicable for a wide range of network computations processing private data using a Web browser.

Journal Article↗

Using CAVE technology for functional genomics studies.

We have established the first Java 3D-enabled CAVE (CAVE automated virtual environment). The Java application programming interface allows the complete separation of the program development from the program execution, opening new application domains for the CAVE technology. Programs can be developed on any Java-enabled computer platform, including Windows, Macintosh, and Linux workstations, and executed in the CAVE without modification. The introduction of Java, one of the major programming environments for bioinformatics, into the CAVE environment allows the rapid development applications for genome research, especially for the analysis of the spatial and temporal data that are being produced by functional genomics experiments. The CAVE technology will play a major role in the modeling of biological systems that is necessary to understand how these systems are organized and how they function.

Automation↗

A new approach to sequence comparison: normalized sequence alignment.

The Smith-Waterman algorithm for local sequence alignment is one of the most important techniques in computational molecular biology. This ingenious dynamic programming approach was designed to reveal the highly conserved fragments by discarding poorly conserved initial and terminal segments. However, the existing notion of local similarity has a serious flaw: it does not discard poorly conserved intermediate segments. The Smith-Waterman algorithm finds the local alignment with maximal score but it is unable to find local alignment with maximum degree of similarity (e.g. maximal percent of matches). Moreover, there is still no efficient algorithm that answers the following natural question: do two sequences share a (sufficiently long) fragment with more than 70% of similarity? As a result, the local alignment sometimes produces a mosaic of well-conserved fragments artificially connected by poorly-conserved or even unrelated fragments. This may lead to problems in comparison of long genomic sequences and comparative gene prediction as recently pointed out by Zhang et al. (Bioinformatics, 15, 1012-1019, 1999). In this paper we propose a new sequence comparison algorithm (normalized local alignment ) that reports the regions with maximum degree of similarity. The algorithm is based on fractional programming and its running time is O(n2log n). In practice, normalized local alignment is only 3-5 times slower than the standard Smith-Waterman algorithm.

Algorithms↗

Finding potential ligands for PDZ domains by tailfit, a JAVA program.

OBJECTIVE: To deduce all potential ligands undiscovered experimentally by searching all the proteins containing same C-termini, which can bind a certain PDZ domain. METHODS: We developed a JAVA program for searching short exact sequence matches at C-terminus. According to the known C-termini, which PDZ domains recognized experimentally, Swissprot database has been searched by this program for all potential ligands. RESULTS: Some PDZ domains may have more potential ligand proteins, which are undiscovered yet experimentally. These bioinformatic results also provide clues for studying functions of hypothetical proteins and PDZ domains' protein interactions in many different organisms. CONCLUSION: The results may provide useful clues for discovering potential functions of hypothetical proteins and new functions of known proteins.

Amino Acid Sequence↗