Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Identification of autophagy-related genes as potential biomarkers correlated with immune infiltration in bipolar disorder: a bioinformatics analysis.

BACKGROUND: Bipolar disorder (BPD) is a kind of manic and depressive phase alternate episodes of serious mental illness, and it is correlated with well-documented cortical brain abnormalities. Emerging evidence supports that autophagy dysfunction in neuronal system contributes to pathophysiological changes in neurological disease. However, the role of autophagy in bipolar disorder has rarely been elucidated. This study aimed to identify the autophagy-related gene as a potential biomarker Correlated to immune infiltration in BPD. METHODS: The microarray dataset GSE23848 and autophagy-related genes (ARGs) were downloaded. Differentially expressed genes (DEGs) between normal and BPD samples were screened using the R software. Machine learning algorithms were performed to screen the significant candidate biomarker from autophagy-related differentially expressed genes (ARDEGs). The correlation between the screened ARDEGs and infiltrating immune cells was explored through correlation analysis. RESULTS: In this study, the autophagy pathway was abundantly enriched and activated in BPD, as indicated by Pathway enrichment analysis. We identified 16 ARDEGs in BPD compared to the normal group. A signature of 4 ARDEGs (ERN1, ATG3, CTSB, and EIF2AK3) was screened. ROC analysis showed that the above genes have good diagnostic performance. In addition, immune correlation analysis considered that the above four genes significantly correlated with immune cells in BPD. CONCLUSIONS: Autophagy - immune cell axis mediates pathophysiological changes in BPD. Four important ARDEGs are prospective to be potential biomarkers associated with immune infiltration in BPD and helpful for the prediction or diagnosis of BPD.

Bipolar Disorder↗

Semiparametric efficient estimation of small genetic effects in large-scale population cohorts.

Population genetics seeks to quantify DNA variant associations with traits or diseases, as well as interactions among variants and with environmental factors. Computing millions of estimates in large cohorts in which small effect sizes and tight confidence intervals are expected, necessitates minimizing model-misspecification bias to increase power and control false discoveries. We present TarGene, a unified statistical workflow for the semi-parametric efficient and double robust estimation of genetic effects including $ k $-point interactions among categorical variables in the presence of confounding and weak population dependence. $ k $-point interactions, or Average Interaction Effects (AIEs), are a direct generalization of the usual average treatment effect (ATE). We estimate genetic effects with cross-validated and/or weighted versions of Targeted Minimum Loss-based Estimators (TMLE) and One-Step Estimators (OSE). The effect of dependence among data units on variance estimates is corrected by using sieve plateau variance estimators based on genetic relatedness across the units. We present extensive realistic simulations to demonstrate power, coverage, and control of type I error. Our motivating application is the targeted estimation of genetic effects on trait, including two-point and higher-order gene-gene and gene-environment interactions, in large-scale genomic databases such as UK Biobank and All of Us. All cross-validated and/or weighted TMLE and OSE for the AIE $ k $-point interaction, as well as ATEs, conditional ATEs and functions thereof, are implemented in the general purpose Julia package TMLE.jl. For high-throughput applications in population genomics, we provide the open-source Nextflow pipeline and software TarGene which integrates seamlessly with modern high-performance and cloud computing platforms.

Humans↗

TAMBIS: transparent access to multiple bioinformatics information sources.

UNLABELLED: TAMBIS (Transparent Access to Multiple Bioinformatics Information Sources) is an application that allows biologists to ask rich and complex questions over a range of bioinformatics resources. It is based on a model of the knowledge of the concepts and their relationships in molecular biology and bioinformatics. AVAILABILITY: TAMBIS is available as an applet from http://img.cs.man.ac.uk/tambis SUPPLEMENTARY: A full manual, tutorial and videos can be found at http://img.cs.man.ac.uk/tambis. CONTACT: tambis@cs.man.ac.uk

Computational Biology↗

Tricross : using dot-plots in sequence-id space to detect uncataloged intergenic features.

MOTIVATION: The process of determining the functional sequence content of an organism is confounded by several factors. Large protein coding sequences are relatively easy to find by statistical methods. Smaller proteins however may escape detection due to their size falling below some arbitrary researcher-defined minimum cutoff, or the inability to precisely define a promoter, or translational start (Delcher et al., Nucleic Acids Res., 27, 4636-4641, 1999). Promoter and regulatory sequences themselves are difficult to define due to a significant amount of allowable sequence variation, as well as a probable lack of any completely accurate whole-organismal gene catalogs to date. Finally, certain genes coding functional RNAs may have insufficient structural or sequence constraints to be detectable by normal sequence structure/pattern searching methods (Eddy and Rivas, Bioinformatics, 16, 583-605, 2000). In those cases where there are multiple closely related organisms that have been sequenced, there is additional information that may be used in the investigation of sequence content-that being the possible conserved nature of functional sequences between the organisms. We present a method for the utilization of this conserved information to detect genes and other potentially functional sequences that may be missed by standard ORF-calling, RNA finding, and pattern matching software. The tricross programs produce a multi-way cross comparison of three sets of sequences, determine which are conserved in all three sets, and produce a graphical (Virtual Reality Modelling Language-VRML; (ISO/IEC 14772-1: 1997, VDC), 1997) representation as well as alignments of all sequence triples found. The software can also be applied to a pair of sequence sets, though the noise in the results increases. RESULTS: Tricross has been used to examine the intergenic-sequence content of the three archaeal Pyrococcus genomes to determine the most highly related sequences remaining between the annotated protein and RNA coding sequences. Set to relatively stringent similarity requirements for the search, tricross found 101 intergenic sequences conserved among the three organisms. Interestingly, 29 of these appear to contain members of a family of small RNA molecules (Kiss-Laszlo et al., EMBO J., 17, 797-807, 1998) only recently discovered in the Archaea (Armbruster, OSU, Diss., 1988; Omer et al., Science, 288, 517-522, 2000; Gaspin et al., J. Mol. Biol., 297, 895-906, 2000). While some of the remaining 72 appear to be individual highly conserved promoter sequences, others have no currently known biological significance. Although originally developed to facilitate the examination of intergenic sequences, none of the tricross logic is inherently specific to intergenic sequences. The software can also be applied to gene sequences, and has been used to produce inter-genomic gene order dot-plots for Haemophilus influenzae (Fleischmann et al., Science, 269, 496-512, 1995) versus H.ducreyi (unpublished data), and Neisseria meningiditis Z2491 (serogroup A) (Parkhill et al., Nature, 404, 502-506, 2000) versus Neisseria meningiditis Z58 (serogroup B) (Tettelin et al., Science, 287, 1809-1815, 2000) versus Neisseria gonorrhoeae (Lewis et al., http://micro-gen.ouhsc.edu/, 2000). AVAILABILITY: The tricross software package is available from http://www.biosci.ohio-state.edu/~ray/bioinformatics/tricross.html. CONTACT: ray@biosci.ohio-state.edu; daniels.7@osu.edu; munsonr@pediatrics.ohio-state.edu SUPPLEMENTARY INFORMATION: Additional data from the cross-genomic comparisons examined in the discussion section are linked from http://www.biosci.ohio-state.edu/~ray/bioinformatics/tricross.html.

Base Sequence↗

UK CropNet: a collection of databases and bioinformatics resources for crop plant genomics.

The UK Crop Plant Bioinformatics Network (UK CropNet) was established in 1996 in order to harness the extensive work in genome mapping in crop plants in the UK. Since this date we have published five databases from our central UK CropNet WWW site (http://synteny.nott.ac.uk/) with a further three to follow shortly. Our resource facilitates the identification and manipulation of agronomically important genes by laying a foundation for comparative analysis among crop plants and model species. In addition, we have developed a number of software tools that facilitate the visualisation and analysis of our data. Many of our tools are made freely available for use with both crop plant data and with data from other species.

Crops, Agricultural↗

Elucidating the Mechanism of Xiaoqinglong Decoction in Chronic Urticaria Treatment: An Integrated Approach of Network Pharmacology, Bioinformatics Analysis, Molecular Docking, and Molecular Dynamics Simulations.

INTRODUCTION: Xiaoqinglong Decoction (XQLD) is a traditional Chinese medicinal formula commonly used to treat chronic urticaria (CU). However, its underlying therapeutic mechanisms remain incompletely characterized. This study employed an integrated approach combining network pharmacology, bioinformatics, molecular docking, and molecular dynamics simulations to identify the active components, potential targets, and related signaling pathways involved in XQLD's therapeutic action against CU, thereby providing a mechanistic foundation for its clinical application. METHODS: The active components of XQLD and their corresponding targets were identified using the Traditional Chinese Medicine Systems Pharmacology (TCMSP) database. CU-related targets were retrieved from the OMIM and GeneCards databases. Subsequently, core components and targets were determined via protein-protein interaction (PPI) network analysis and component-target-pathway network construction. Topological analyses were performed using Cytoscape software to prioritize core nodes within these networks. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were conducted via the DAVID database to identify enriched biological processes and signaling pathways. Molecular docking was performed to evaluate binding interactions between key components and core targets, while molecular dynamics (MD) simulations were employed to assess the stability of the component-target complexes with the lowest binding energy. Finally, CU-related targets of XQLD were validated using datasets from the Gene Expression Omnibus (GEO) database. RESULTS: A total of 135 active components and 249 potential targets of XQLD were identified, alongside 1,711 CU-related targets. Core components, such as quercetin, kaempferol, beta-sitosterol, naringenin, stigmasterol, and luteolin, exhibited high degree values in the constructed networks. The core targets identified included AKT1, TNF, IL6, TP53, PTGS2, CASP3, BCL2, ESR1, PPARG, and MAPK3. GO and KEGG pathway enrichment analyses revealed the PI3K-Akt signaling pathway as a central regulatory mechanism. Molecular docking studies demonstrated strong binding affinities between active components and core targets, with the stigmasterol-AKT1 complex exhibiting the lowest binding energy (-11.4 kcal/mol) and high stability in MD simulations. Validation using GEO datasets identified 12 core genes shared between CU-related targets and XQLD-associated targets, including PTGS2 and IL6, which were also prioritized as core targets in the network pharmacology analyses. DISCUSSION: This study comprehensively integrates multidisciplinary approaches to clarify the potential molecular mechanisms of XQLD in treating CU, highlighting its multitarget and multipathway synergistic effects. Molecular docking and dynamics simulations confirm the stable interaction between stigmasterol and the core target AKT1. Additionally, GEO dataset analysis verifies the pathogenic relevance of targets such as PTGS2 and IL6, significantly enhancing the credibility of our findings. These results provide a modern scientific basis for the traditional therapeutic effects of XQLD on CU and have important implications for developing multitarget treatments for this condition. However, this study mainly relies on database mining and computational simulations. Further in vitro and in vivo experimental validations are needed to confirm the predicted component-target-pathway interactions. CONCLUSION: This study identifies the active components, potential targets, and pathways through which XQLD exerts therapeutic effects on CU. These findings provide a theoretical foundation for further mechanistic studies and support their clinical application in the treatment of CU.

Molecular Docking Simulation↗

A novel splice-altering TNC variant (c.5247A > T, p.Gly1749Gly) in an Chinese family with autosomal dominant non-syndromic hearing loss.

BACKGROUND: This study aims to analyze the pathogenic gene in a Chinese family with non-syndromic hearing loss and identify a novel mutation site in the TNC gene. METHODS: A five-generation Chinese family from Anhui Province, presenting with autosomal dominant non-syndromic hearing loss, was recruited for this study. By analyzing the family history, conducting clinical examinations, and performing genetic analysis, we have thoroughly investigated potential pathogenic factors in this family. The peripheral blood samples were obtained from 20 family members, and the pathogenic genes were identified through whole exome sequencing. Subsequently, the mutation of gene locus was confirmed using Sanger sequencing. The conservation of TNC mutation sites was assessed using Clustal Omega software. We utilized functional prediction software including dbscSNV_AdaBoost, dbscSNV_RandomForest, NNSplice, NetGene2, and Mutation Taster to accurately predict the pathogenicity of these mutations. Furthermore, exon deletions were validated through RT-PCR analysis. RESULTS: The family exhibited autosomal dominant, progressive, post-lingual, non-syndromic hearing loss. A novel synonymous variant (c.5247A > T, p.Gly1749Gly) in TNC was identified in affected members. This variant is situated at the exon-intron junction boundary towards the end of exon 18. Notably, glycine residue at position 1749 is highly conserved across various species. Bioinformatics analysis indicates that this synonymous mutation leads to the disruption of the 5' end donor splicing site in the 18th intron of the TNC gene. Meanwhile, verification experiments have demonstrated that this synonymous mutation disrupts the splicing process of exon 18, leading to complete exon 18 skipping and direct splicing between exons 17 and 19. CONCLUSION: This novel splice-altering variant (c.5247A > T, p.Gly1749Gly) in exon 18 of the TNC gene disrupts normal gene splicing and causes hearing loss among HBD families.

Adult↗

bioTk:componentry for genome informatics graphical user interfaces.

bioTk is a collection of graphical "widgets" and utilities that support application programming in the domain of bioinformatics. It is intended to establish a framework that encourages the development of communicating window-based applications and flexible, non-modal user interaction. The current release of bioTk has domain-specific widgets for chromosome ideogram displays, genome maps, and scrolling sequence windows.

Base Sequence↗

Exploring shared biomarkers and their mechanisms in thyroid cancer and systemic lupus erythematosus via bioinformatics analysis.

BACKGROUND: Systemic lupus erythematosus (SLE), an autoimmune disorder, is linked to a heightened risk of multiple malignancies, including thyroid cancer. Thyroid cancer is the most prevalent malignancy of the endocrine system, and its autoimmune-related pathological features render it an optimal subject for investigating the mechanisms of their comorbidity. The molecular mechanisms underlying this comorbidity are still ambiguous. The accurate diagnosis and treatment of thyroid cancer urgently necessitate innovative molecular targets that extend beyond conventional pathological characteristics. This study seeks to employ integrated bioinformatics approaches to elucidate potential shared molecular mechanisms and immunological features between thyroid cancer and systemic lupus erythematosus (SLE), aiming to enhance understanding of their comorbidity and identify novel intervention targets. METHODS: This study initially acquired gene expression data for TC and SLE from the GEO database and subsequently screened and identified differentially expressed genes (DEGs) shared by both diseases. Subsequently, we conducted Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), and Reactome functional enrichment analyses on these 46 shared differentially expressed genes (DEGs) and further assessed the activation status of pertinent pathways using Gene Set Enrichment Analysis (GSEA). Subsequently, we employed CIBERSORTx to examine immune infiltration patterns and developed protein-protein interaction networks utilising the STRING database. We identified hub genes utilising the MCODE and cytoHubba plugins and visualised the findings with Cytoscape software. We additionally assessed the diagnostic efficacy of these core hub genes in an independent dataset utilising ROC curves and investigated their prognostic relevance in thyroid cancer through Kaplan-Meier survival analysis and multivariate Cox proportional hazards regression. Ultimately, we employed the Network Analyst platform to forecast transcription factor-gene and miRNA-gene regulatory networks and identified potential targeted therapeutic compounds utilising the DSigDB database. RESULTS: This study identified 46 differentially expressed genes (DEGs) commonly linked to thyroid cancer and systemic lupus erythematosus (SLE), which were significantly enriched in signalling pathways associated with immune-inflammatory activation, type I interferon responses, and complement pathway activation. Moreover, GSEA findings validated that immune-inflammatory and autoimmune-related pathways are markedly activated in both conditions. Twelve hub genes were discerned through protein-protein interaction networks. Analysis of immune infiltration indicated that thyroid cancer and systemic lupus erythematosus exhibit a shared characteristic of innate immune dysregulation, marked by the infiltration of myeloid cells (neutrophils, M0/M2 macrophages). Receiver operating characteristic (ROC) curve analysis identified six significant core hub genes with substantial diagnostic value: C1QB, LCN2, C1QC, LTF, VSIG4, and C3AR1. Univariate survival analysis indicated that elevated expression of C1QC and C3AR1 significantly enhances overall survival in thyroid cancer patients; however, multivariate COX regression analysis revealed that their independent prognostic significance necessitates further validation. This study predicted the interaction networks of transcription factors and miRNAs regulating key genes, with LCN2 demonstrating the highest connectivity to miRNAs, and identified candidate therapeutic compounds linked to it. CONCLUSION: This study employed bioinformatics analysis to identify critical shared hub genes and molecular pathways connecting thyroid cancer and systemic lupus erythematosus, offering novel insights into their shared pathogenesis and the advancement of targeted biomarkers and therapeutic strategies.

Bioinformatics analysis↗

The untranslated regions of eukaryotic mRNAs: structure, function, evolution and bioinformatic tools for their analysis.

The crucial role of the non-coding portion of genomes is now widely acknowledged. In particular, mRNA untranslated regions are involved in many post-transcriptional regulatory pathways that control mRNA localisation, stability and translation efficiency. A review is given of the most recent research works on the functional characterisation of eukaryotic mRNA untranslated regions. In order to make possible a systematic and detailed sequence analysis of mRNA untranslated regions (UTRs), a non-redundant database of metazoan mRNA untranslated sequences annotated for the occurrence of specific functional elements, UTRdb, was devised. These elements, whose consensus structure has been devised on the basis of experimental assays and of comparative analyses, have been collected in the UTRsite database. A suitable pattern-matching software has been devised to search UTRsite patterns in user-submitted sequences, also assessing their statistical significance. Structural, compositional and evolutionary features of untranslated sequences of metazoan mRNAs have been investigated showing peculiar intra- and interspecific patterns.

Animals↗

The Merck Gene Index browser: an extensible data integration system for gene finding, gene characterization and EST data mining.

MOTIVATION: To make effective use of the vast amounts of expressed sequence tag (EST) sequence data generated by the Merck-sponsored EST project and other similar efforts, sequences must be organized into gene classes, and scientists must be able to 'mine' the gene class data in the context of related genomic data. RESULTS: This paper presents the Merck Gene Index browser, an easily extensible, World Wide Web-based system for mining the Merck Gene Index (MGI) and related genomic data. The MGI is a non-redundant set of clones and sequences, each representing a distinct gene, constructed from all high-quality 3' EST sequences generated by the Merck-sponsored EST project. The MGI browser integrates data from a variety of sources and storage formats, both local and remote, using an eclectic integration strategy, including a federation of relational databases, a local data warehouse and simple hypertext links. Data currently integrated include: LENS cDNA clone and EST data, dbEST protein and non-EST nucleic acid similarity data, WashU sequence chromatograms. Entrez sequence and Medline entries, and UniGene gene clusters. Flatfile sequence data are accessed using the Bioapps server, an internally developed client-server system that supports generic sequence analysis applications. Browser data are retrieved and formatted by means of the Bioinformatics Data Integration Toolkit (B-DIT), a new suite of Perl routines.

Abstracting and Indexing↗

A computer system to perform structure comparison using TOPS representations of protein structure.

We describe the design and implementation of a fast topology-based method for protein structure comparison. The approach uses the TOPS topological representation of protein structure, aligning two structures using a common discovered pattern and generating measure of distance derived from an insert score. Heavy use is made of a constraint-based pattern-matching algorithm for TOPS diagrams that we have designed and described elsewhere (Bioinformatics 15(4) (1999) 317). The comparison system is maintained at the European Bioinformatics Institute and is available over the Web at tops.ebi.ac.uk/tops. Users submit a structure description in Protein Data Bank (PDB) format and can compare it with structures in the entire PDB or a representative subset of protein domains, receiving the results by email.

Algorithms↗

Cambridge Healthtech Institute's Third Annual Conference on Lab-on-a-Chip and Microarrays. 22-24 January 2001, Zurich, Switzerland.

Cambridge Healthtech Institute's Third Annual Conference on Lab-on-a-Chip and Microarray technology covered the latest advances in this technology and applications in life sciences. Highlights of the meetings are reported briefly with emphasis on applications in genomics, drug discovery and molecular diagnostics. There was an emphasis on microfluidics because of the wide applications in laboratory and drug discovery. The lab-on-a-chip provides the facilities of a complete laboratory in a hand-held miniature device. Several microarray systems have been used for hybridisation and detection techniques. Oligonucleotide scanning arrays provide a versatile tool for the analysis of nucleic acid interactions and provide a platform for improving the array-based methods for investigation of antisense therapeutics. A method for analysing combinatorial DNA arrays using oligonucleotide-modified gold nanoparticle probes and a conventional scanner has considerable potential in molecular diagnostics. Various applications of microarray technology for high-throughput screening in drug discovery and single nucleotide polymorphisms (SNP) analysis were discussed. Protein chips have important applications in proteomics. With the considerable amount of data generated by the different technologies using microarrays, it is obvious that the reading of the information and its interpretation and management through the use of bioinformatics is essential. Various techniques for data analysis were presented. Biochip and microarray technology has an essential role to play in the evolving trends in healthcare, which integrate diagnosis with prevention/treatment and emphasise personalised medicines.

Computational Biology↗

PyEvoMotion: a Python tool for population-based time-course analysis of genome evolution.

SUMMARY: We present PyEvoMotion, an open-source Python tool for inferring molecular clock models with time-dependent Gaussian noise from high-throughput genomic datasets. PyEvoMotion features a command-line interface and a modular architecture, allowing seamless integration into larger bioinformatic pipelines. The tool supports customizable filtering, temporal discretization definition, and mutation classification, making it adaptable to diverse research needs. While traditional phylogenetic methods may encounter computational challenges with large datasets, PyEvoMotion can process thousands to millions of sequences to compute statistical parameters associated with a stochastic differential equation model, thereby weighting the genetic variation within the population. Using viral genomic data, we demonstrate its capability to infer evolutionary rates and detect non-Brownian evolutionary motions with subdiffusive behavior. PyEvoMotion shows potential to provide overlooked insights into genome evolution in different contexts. AVAILABILITY AND IMPLEMENTATION: The open source software is available on GitHub at https://github.com/luksgrin/PyEvoMotion and on SourceForge at https://sourceforge.net/projects/pyevomotion.

Software↗

abCRISPR: deep learning-based design of abasic gRNA sequences for specific CRISPR-Cas9 genome editing.

SUMMARY: CRISPR-Cas9 has become a widely used tool for genome editing. However, its off-target cleavage caused by partial sequence matches with guide RNAs (gRNAs) remains a critical limitation. Recently, abasic gRNAs (ØXØ) have been developed to enhance target specificity, but their effects vary depending on the positional sequence context. Here, we present abCRISPR, a deep neural network (DNN) framework for the rational design of ØXØ sequences with minimized off-target activity. abCRISPR leverages informative few-shot training with paired datasets of abasic and unmodified gRNAs, using high-quality random mismatch target libraries, exhaustively sequenced for mismatched off-target substrates (n = 97583) in in vitro CRISPR-Cas9 cleavage experiments. Predicted off-target activities for both abasic and unmodified gRNAs showed strong correlation with experimental data (r ≥ 0.95, 10-fold cross-validation). Notably, these comprehensive training sets provide robust ground-truth negatives, enabling accurate and sensitive prediction of off-targets. For unmodified gRNAs, abCRISPR (AUC = 0.98) was validated to outperform existing deep learning-based methods (AUC = 0.45-0.68). When applied to the human genome, abCRISPR generated ØXØ sequences, covering 58 875 004 potent CRISPR-targetable sites with improved target specificity. Together, this work provides a comprehensive bioinformatics resource for safe and precise CRISPR-Cas9 genome editing. AVAILABILITY AND IMPLEMENTATION: The source code for abCRISPR and training data are available at https://doi.org/10.5281/zenodo.20398246. abCRISPR results for the human genome are available at http://clip.korea.ac.kr/abCRISPR/.

Deep Learning↗

CCRR: a user-friendly platform for analyzing complex chromosomal rearrangements in tumors.

SUMMARY: Complex chromosomal rearrangements in tumors involve intricate genomic alterations that significantly affect gene function and contribute to cancer development. Identifying these events is crucial for cancer research but is often challenging due to the complexity and limitations of existing tools. We developed the Complex Chromosomal Rearrangements Resolver (CCRR), a comprehensive, reproducible, and user-friendly platform for analyzing complex rearrangements in tumors. CCRR integrates multiple SV and CNV detection tools within a Docker container environment, simplifying installation and configuration. It can be easily deployed, automating the execution and merging of results, providing high-confidence consensus SV and CNV calls, allowing researchers to efficiently analyze complex chromosomal rearrangements in tumors without extensive bioinformatics expertise. CCRR also includes a web server for one-click analysis and customized visualization. AVAILABILITY AND IMPLEMENTATION: The CCRR platform is freely available at https://www.ccrr.life. Source code and executables can be accessed at https://github.com/laslk/CCRR. An archived version is available at Zenodo: https://doi.org/10.5281/zenodo.15386513.

Software↗

GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequences.

MOTIVATION: Short tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. RESULTS: Here, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653.

Microsatellite Repeats↗

Deriving folds of macromolecular complexes through electron cryomicroscopy and bioinformatics approaches.

Intermediate-resolution (7-9A) structures of large macromolecular complexes can be obtained by electron cryomicroscopy. This structural information, combined with bioinformatics data for the individual protein components or domains, can lead to a fold model for the entire complex. Such approaches have been demonstrated with the 6.8 A structure of the rice dwarf virus to derive models for the major capsid shell proteins.

Amino Acid Sequence↗