Search PubMedSearch

SEARCH · Search PubMed

Results for “hybrid framework”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Design of an innovative framework based hybrid catalyst for simultaneous and sensitive monitoring of food additive and preservative of vanillin and nitrite in direct samples.

As vanillin (VAN) and nitrite (NIT) contamination in the food chain poses substantial threats to environmental and public health, rapid and portable detection is essential. The present study presents the first electrochemical sensor report based on a hybrid composite of Ni-TPA-MOF and MoS2/Co3O4. The oxidation of VAN and NIT exhibited sharp peaks and less over-potential on Ni-TPA-MOF/MoS2/Co3O4/GCE than on control electrode surfaces. On modified composite electrode surfaces, pH and scan rate were investigated for VAN and NIT. Further, the oxidation current exhibited high linearity at VAN and NIT concentrations of 5 nM-1000 μM and 3 nM-1250 μM, with detection limits of 0.102 nM and 0.073 nM (S/N = 3). We also applied anti-interfering ability (five/ten-fold excess of co-interfering compounds) and practical tests to various food-based real samples, with high recoveries of 98.85-102.41%. This study highlights the catalytic properties of Ni-TPA-MOF/MoS2/Co3O4 and demonstrates the sensor as a promising tool for food safety.

Benzaldehydes

A decentralized future for the open-science databases.

The continuous and reliable open access to curated biological data repositories is indispensable for accelerating rigorous scientific inquiry and fostering reproducible research outcomes. However, the current paradigm, which relies heavily on centralized infrastructure for the storage and distribution of foundational biomedical datasets, inherently introduces significant vulnerabilities. This centralized model is susceptible to single points of failure, including cyberattacks, technical malfunctions, natural disasters, and even political or funding uncertainties. Such disruptions can lead to widespread data unavailability, data loss, integrity compromises, and substantial delays in critical research, ultimately impeding scientific progress. The downstream effect of such interruptions can be the widespread paralysis of diverse research activities, including computational, clinical, molecular, and climate studies. This scenario vividly illustrates the inherent dangers of consolidating essential scientific resources within a single geopolitical or institutional locus. As data generation is accelerating and the global landscape continues to fluctuate, the sustainability of centralized models must be critically re-evaluated. A shift toward federated and decentralized architectures may offer a robust and forward-looking approach to enhancing the resilience of scientific data infrastructures by reducing exposure to governance instability, infrastructural fragility, and funding volatility, while also promoting equity and global accessibility. Inspired by established models such as ELIXIR's federated infrastructure and the policy and funding frameworks developed by CODATA and the Global Biodata Coalition (GBC), emerging Decentralized Science (DeSci) initiatives can contribute to building more resilient, fair, and incentive-aligned data ecosystems. The future of open science depends on integrating these complementary approaches to establish a globally distributed, economically sustainable, and institutionally robust infrastructure that safeguards scientific data as a public good, further ensuring continued accessibility, interoperability, and preservation for generations to come. Here, we examine the structural limitations of centralized repositories, evaluate federated and decentralized models, and propose a hybrid framework for resilient, fair, and sustainable scientific data stewardship.

data accessibility

Q RadFusion: Hybrid Quantum Classical Radiogenomic Framework for Breast Cancer Diagnosis.

BACKGROUND AND PURPOSE: Breast cancer remains the most common cancer in women worldwide, with early and accurate diagnosis critical for patient survival. Radiogenomics integrates imaging phenotypes with genomic profiles, offering a pathway to precision diagnostics. However, existing classical machine learning models often struggle with the high dimensionality and heterogeneity of multimodal data, leading to issues in calibration and reproducibility. This study presents Q RadFusion, a hybrid quantum-classical framework designed to enhance breast cancer diagnosis by fusing mammography and genomics data. METHODS: Q RadFusion was implemented on two publicly available datasets: CBIS-DDSM (2,600 curated mammography cases, TCIA) and TCGA-BRCA (1,000 genomic profiles, GDC). Imaging preprocessing included bias-field correction, segmentation, and harmonization, while genomic data underwent normalization and imputation. Feature selection was performed using the Quantum Approximate Optimization Algorithm (QAOA), and features were mapped into a quantum Hilbert space using Variational Quantum Circuits (VQC). For multimodal fusion, ResNet encoded mammography features, and a Transformer encoded genomic features. Patient-level and site-held-out splits were used for evaluation. RESULTS: Q RadFusion achieved an AUC of 0.96 and accuracy of 94%, outperforming baselines including CNN-LSTM, ResNet + XGBoost, and multimodal Transformers. Ablation studies confirmed the contribution of quantum components, with optimal performance observed at circuit depth, qubits, and QAOA layers. The model also demonstrated improved calibration and ~ 80% fewer parameters compared to deep fusion networks. CONCLUSION: Q RadFusion demonstrates that hybrid quantum-classical radiogenomic integration can deliver accurate, reproducible, and clinically meaningful diagnostic support for breast cancer, with strong potential for future clinical translation.

Breast Cancer

A reinforcement learning-enhanced fuzzy multi-objective equilibrium optimization framework for multiple sequence alignment.

Multiple sequence alignment (MSA) is a fundamental task in bioinformatics, underpinning comparative genomics, structural analysis, and evolutionary inference. However, MSA remains a challenging multi-objective optimization problem due to the need to simultaneously maximize alignment accuracy, preserve conserved regions, and control gap proliferation, particularly in large and heterogeneous sequence collections. In this work, we propose MOFSACEO-MSA, a novel hybrid optimization framework for multiple sequence alignment that integrates a fuzzy multi-objective evaluation scheme with the Equilibrium Optimizer (EO) and a Soft Actor-Critic (SAC)-based adaptive control mechanism. The proposed framework formulates MSA as a dynamic multi-objective optimization problem, in which alignment quality is assessed using complementary residue-level and column-level criteria, including Sum-of-Pairs score, column conservation, entropy, and gap statistics. Fuzzy membership functions are employed to harmonize competing objectives into a unified optimization landscape, while EO provides robust global exploration. To further enhance adaptability, SAC dynamically regulates key EO parameters during the search process, enabling an effective balance between exploration and exploitation across datasets of varying size and heterogeneity. Extensive experiments werew conducted on diverse biological sequence datasets, with a primary focus on RNA benchmarks, including structured families from Rfam, large-scale repositories from RNAcentral and GenBank, and organism-specific tRNA datasets from GtRNAdb. Comparative evaluations against classical alignment tools (ClustalW, MAFFT, MUSCLE, PRANK, KAlign, and T-Coffee), metaheuristic methods (SAGA, Sequoya and EAFSA), and a reinforcement learning-based approach (RLALIGN) demonstrate that MOFSACEO-MSA consistently achieves competitive or superior Sum-of-Pairs scores while significantly reducing gap proportions and maintaining compact alignment lengths. Notably, the proposed framework exhibits improved robustness on large and highly heterogeneous datasets, where existing methods often suffer from excessive gap insertion or unstable convergence. Overall, MOFSACEO-MSA provides a flexible and extensible optimization paradigm that effectively bridges evolutionary search and reinforcement learning for high-quality multiple sequence alignment, with demonstrated effectiveness on challenging RNA alignment tasks.

Sequence Alignment

A complete and near-perfect rhesus macaque reference genome: lessons from subtelomeric repeats and sequencing bias.

A truly complete, telomere-to-telomere (T2T), and error-free reference genome remains a foundational resource-and long-standing goal-for unbiased comparative and functional genomics. While recent T2T assemblies of humans and other primates have made substantial progress, most still contain thousands of base-level errors, particularly within highly repetitive regions. Here, we present T2T-MMU8v2.0, a near-perfect T2T assembly of the rhesus macaque (Macaca mulatta), representing the highest base-level accuracy reported in a primate genome to date. By employing an optimized ONT-only assembly strategy, we identify subtelomeric satellite-rich regions as the principal bottleneck to improving assembly quality, owing to technological biases in long-read platforms and limitations in current hybrid assembly frameworks. We discover 268 previously unannotated repeat families and resolve ~8 Mbp of SATR satellite arrays, with over 99-fold enrichment in historically misassembled subtelomeric regions. These satellites form four distinct genomic architectures, each with unique SATR satellite composition, segmental duplication organization, and epigenetic signatures, distinct from the subtelomeric architectures observed in hominid genomes. Notably, in contrast to the largely gene-poor subtelomeric regions in African hominids, the SATR architectures in macaques harbor 58 actively transcribed genes, supported by open chromatin and expression data, suggesting gene innovation within these repetitive regions. Functionally, T2T-MMU8v2.0 improves read mappability and accuracy across sequencing platforms, and results in a 19% improvement of transcription start site enrichment scores and 5,821 additional chromatin accessibility peaks on average, thereby enhancing variant detection, regulatory annotation, and transcriptomic resolution in population genetics or single-nucleus studies. Together, this work establishes a new benchmark for genomics, offers a roadmap for resolving complex repetitive regions, and reveals previously unrecognized features of subtelomeric genome structure and evolution.

Journal Article

Gene expression is stable despite widespread cis and trans regulatory divergence in Saccharomyces yeasts.

Regulatory evolution can alter phenotypes, but cis- and trans-regulatory mechanisms may also diverge extensively while total transcript abundance remains stable. Comparisons of parental expression with allele-specific expression in F1 hybrids provide a framework for separating cis- and trans-regulatory effects because both parental alleles are measured in a shared trans-regulatory environment. Here, we analyzed RNA sequencing data from Saccharomyces cerevisiae, Saccharomyces paradoxus, and their F1 hybrid. Among the 4,164 genes with sufficient allele-specific support for strict classification, 2,134 (51.2%) showed detectable cis and/or trans regulatory divergence. However, hybrid expression remained largely conserved, with 81.5% of genes not significantly different from either parent. Compensatory cis-trans divergence predominated over reinforcing divergence; cross-replicate estimation reduced the apparent magnitude of this excess, but opposite-sign effects remained predominant in all 20 non-overlapping replicate comparisons. To connect gene expression to genome sequence, we analyzed the strongly cis-diverged locus LYS2 and found species differences in promoter architecture, including an S. cerevisiae-specific AT-rich insertion, altered spacing among candidate regulatory features, and a promoter-proximal TATA-like element unique to S. cerevisiae. Sequence-based nucleosome prediction suggests that these differences create a broader promoter-proximal nucleosome-depleted region in S. cerevisiae than in S. paradoxus. We also quantified allele-resolved intron retention and found that allele-resolved intron retention was broadly conserved, with only rare locus-specific hybrid-associated shifts. Together, these results show that regulatory divergence is widespread but often buffered in the hybrid, whereas intron-retention divergence is comparatively limited.

Saccharomyces

Gene expression is stable despite widespread cis and trans regulatory divergence in Saccharomyces yeasts.

Regulatory evolution can alter phenotypes, but cis- and trans-regulatory mechanisms may also diverge extensively while total transcript abundance remains stable. Comparisons of parental expression with allele-specific expression in F1 hybrids provide a framework for separating cis- and trans-regulatory effects because both parental alleles are measured in a shared trans-regulatory environment. Here, we analyzed RNA sequencing data from Saccharomyces cerevisiae, Saccharomyces paradoxus, and their F1 hybrid. Regulatory divergence was widespread, with 61.3% of tested orthologs showing significant divergence in at least one cis or trans component. However, hybrid expression remained largely conserved, with 81.6% of genes not significantly different from either parent. Compensatory cis-trans divergence predominated over reinforcing divergence, consistent with widespread buffering of transcript abundance. To connect genome-wide patterns to mechanism, we analyzed the strongly cis-diverged locus LYS2 and found species differences in promoter architecture, including an S. cerevisiae-specific AT-rich insertion, altered spacing among candidate regulatory features, and a promoter-proximal TATA-like element unique to S. cerevisiae. Sequence-based nucleosome prediction suggests that these differences create a broader promoter-proximal nucleosome-depleted region in S. cerevisiae than in S. paradoxus. We also quantified allele-resolved intron retention and found that splicing was broadly conserved, with only rare locus-specific hybrid-associated shifts. Together, these results show that regulatory divergence is widespread but often buffered in the hybrid, whereas post-transcriptional divergence is comparatively limited.

Gene expression

Whole-Genome Deep Learning Predicts Chemotherapy Response in Colorectal Cancer.

Chemotherapy response in colorectal cancer (CRC) exhibits significant heterogeneity, with current clinical predictors failing to capture complex genomic determinants of resistance. We developed a hybrid deep learning framework integrating convolutional neural networks (CNNs) and bidirectional long short-term memory (BiLSTM) networks to analyze whole-genome somatic mutations, evolutionary conservation, chromatin accessibility, and 3D genome architecture in 2,546 TCGA patients. An attention mechanism identified predictive genomic regions. The model achieved an AUC of 0.92 (95% CI: 0.89-0.94) in cross-validation and 0.88 (95% CI: 0.85-0.91) in independent validation, outperforming clinical models (&#x394;AUC = +0.18, p < 0.001). Key predictors included non-coding variants in TP53, KRAS, and PIK3CA regulatory regions. Triple-positive patients (mutations in all 3 regions) had significantly worse progression-free survival (HR = 4.7, p < 0.001). Our framework enables accurate chemotherapy response prediction and reveals novel non-coding resistance mechanisms, advancing precision oncology in CRC.

Humans

Beyond parental lines: multi-omics analyses reveal epigenetic and transcriptional mechanisms underlying heterosis in Oryza sativa &#xd7; Oryza rufipogon hybrids.

Heterosis, or hybrid vigor, refers to the superior phenotypes of a hybrid compared with their parents and is widely exploited in agriculture. Interspecific hybrids within the Oryza genus demonstrate significant potential for the systematic improvement of rice varieties. Nevertheless, the mechanistic basis underlying heterosis in interspecific Oryza hybrids remains poorly understood. Here, we systematically performed phenotypic characterization, whole-genome bisulfite sequencing, RNA sequencing, and small RNA profiling using Oryza sativa L. ssp. japonica cv. Nipponbare (NIP), Oryza rufipogon Griff. acc. CWR, and their resulting F1 hybrid (named as NC). NIP and CWR showed distinct phenotypic and molecular differences. The interspecific hybrid, NC, exhibited significant yield heterosis. In the hybrid, most epigenetic and transcriptional features displayed additive inheritance patterns relative to parental lines. Analysis revealed that domestication-selected genes maintained relatively low DNA methylation coupled with high expression levels in both hybrid and parental lines. Additionally, we identified that non-additive miRNAs were potentially involved in regulating fertility, cell growth, and cell division processes in the hybrid. A significant negative correlation was observed between DNA methylation level and gene expression. Functional enrichment analysis revealed that hybrid-MPV DEGs were significantly associated with flowering time regulation, carbohydrate metabolism, photosynthesis, protein phosphorylation, seed development, and defense responses. Through weighted gene co-expression network analysis, we identified 102 functional gene modules, six of which were significantly associated with yield-related heterosis. Collectively, our results provide a multi-omics framework for understanding interspecific hybridization between elite cultivars and wild rice relatives, highlighting CWR as an untapped genetic reservoir for rice improvement.

Oryza

Recent advances in electrode materials for electrochemical detection of zearalenone.

Zearalenone (ZEN) is an estrogenic mycotoxin commonly found in cereals, animal feed, and processed foods, making it an important concern for food safety and public health. Conventional chromatographic and immunological methods can detect ZEN; however, they often require expensive instruments, lengthy sample preparation, and skilled personnel, which restrict their use for rapid and on-site testing. Electrochemical sensors have attracted enormous interest of the scientific community because of their high sensitivity, rapid response, low cost, miniaturization potential, and compatibility with portable systems. The analytical performance of the electrochemical sensors is strongly influenced by electrode materials, morphology, conductivity, porosity, surface functionality, and the efficiency of bioreceptor immobilization. Despite several reviews on mycotoxin detection, a systematic assessment connecting electrode-material design, modification strategies, sensing mechanisms, and electroanalytical performance specifically for ZEN sensing remain limited. This review critically evaluates recent advances in metal oxides, carbon-based materials, metal-organic- and covalent organic frameworks, MXenes, polymers, and hybrid composites for electrochemical ZEN detection. Particular attention has been given to their roles in electron transfer, analyte enrichment, selectivity, and real-sample analysis. The review also compares the major limitations of current sensing systems, including complex fabrication, matrix interference, insufficient long-term stability, poor inter-electrode reproducibility, and limited scalability. Finally, future directions for developing robust, cost-effective, portable, and commercially viable ZEN sensors are discussed.

Journal Article

seq2ribo: structure-aware integration of machine learning and simulation to predict ribosome location profiles from RNA sequences.

MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context. AVAILABILITY: seq2ribo is available at https://github.com/Kingsford-Group/seq2ribo.

Machine Learning

seq2ribo: Structure-aware integration of machine learning and simulation to predict ribosome location profiles from RNA sequences.

MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context.

Journal Article

A high-quality draft genome assembly of Johnsongrass illuminates relationships between polyploidization, crop-wild hybridization, and reproductive biology.

Johnsongrass [Sorghum halepense (L.) Pers.] is an allopolyploid, rhizomatous, perennial grass species and one of the most troublesome weeds in global agriculture. We assembled the first Johnsongrass genome to clarify poorly understood genetic factors influencing variable rates of crop-wild hybridization with cultivated sorghum [S. bicolor (L.) Moench]. The draft genome assembly has a total size of 3.26 Gb and BUSCO completeness of 95.3%. We also report the first evolutionary analysis of INHIBITION OF ALIEN POLLEN (IAP), the only known cross-(in)compatibility locus in the genus. Our results reveal an evolutionary history of genome instability, including the loss of distinct parental subgenomes, and suggest that Nebraska accession 'J-37,' the genome donor, is a segmental allotetraploid that may function as a diploid or aneuploid during meiosis. Genome instability could explain observations of variable ploidies in Johnsongrass and facilitate ongoing hybridization with sorghum where gamete ploidies and IAP alleles match. Given this information, we provide a suggested research framework for studying evolution and gene expression in the Sorghum genus where crop-wild hybridization occurs and for predicting the potential for hybridization between specific crossing partners. Collectively, this work will bolster efforts to study and manage reproductive biology in other crop-wild polyploid complexes.

Sorghum

A Molecularly Anchored Spatial Transcriptomic Framework for Precise CA1-Subiculum Parcellation and Region-Resolved Analysis in Alzheimer's Disease.

BACKGROUND: The precise molecular delineation of the interface between the Subiculum (Sub) and cornu ammonis 1 (CA1) is a challenge in hippocampal research, as conventional cytoarchitectural boundaries are often ambiguous and limit reproducible regional annotation. Here, we developed a molecularly anchored spatial transcriptomic framework to define CA1-Sub regional identities using high-definition spatial transcriptomics (Stereo-seq) and single-nucleus RNA sequencing (snRNA-seq) references. FINDINGS: Using a human hippocampal Stereo-seq dataset from 12 donors, we established a data-driven parcellation framework that defines reproducible molecular features distinguishing CA1 and Sub while capturing the transition between these regions. FN1 was identified as a Sub-enriched marker in a subset of EX_Sub and, together with ETV1 and additional regional markers, enabled molecular assignment of CA1 and Sub identities across datasets. The Sub association of FN1 and ETV1 was further supported by human 10X Genomics spatial transcriptomics, mouse in situ hybridization data, and a mouse spatial transcriptomic dataset. Applying this framework to Alzheimer's disease (AD) tissues revealed region-specific transcriptional alterations across CA1 and Sub, including enrichment of mitochondrial energy metabolism-related transcripts in the Sub, suggesting exploratory transcriptional associations of altered metabolic function. CONCLUSIONS: This study provides a molecularly anchored framework for human CA1-Sub parcellation that complements conventional annotation. By defining regional molecular states while preserving the biological continuum across CA1-Sub interface, this approach enables more consistent regional analysis of human hippocampus tissue across donors, datasets, and disease conditions.

Journal Article

MetaflowX: a scalable and resource-efficient workflow for multi-strategy metagenomic analysis.

Microbiomes play crucial roles in diverse ecosystems, spanning environmental, agricultural, and human health domains. However, in-depth metagenomic data analysis presents significant technical and resource challenges, particularly at scale. Existing computational pipelines are typically limited to either reference-based or reference-free approaches and exhibit inefficiencies in process large datasets. Here, we introduce MetaflowX (https://github.com/01life/MetaflowX), an open-resource workflow integrating both analytical paradigms for enhanced metagenomic investigations. This modular framework encompasses short-read quality control, rapid microbial profiling, hybrid contig assembly and binning, high-quality metagenome-assembled genome (MAG) identification, as well as bin refinement and reassembly. Benchmarking tests showed that MetaflowX completed full metagenomic analyses up to 14-fold faster and with 38% less disk usage than existing workflows. It also recovered the highest number of high-quality and taxonomically diverse MAGs. A dedicated reassembly module further improved MAG quality, increasing completeness by 5.6% and reducing contamination by 53% on average. Functional annotation modules enable detection of key features, including virulence and antibiotic resistance genes. Designed for extensibility, MetaflowX provides an efficient solution addressing current and emerging demands in large-scale metagenomic research.

Metagenomics

N6-methyladenine identification using deep learning and discriminative feature integration.

N6-methyladenine (6&#xa0;mA) is a pivotal DNA modification that plays a crucial role in epigenetic regulation, gene expression, and various biological processes. With advancements in sequencing technologies and computational biology, there is an increasing focus on developing accurate methods for 6&#xa0;mA site identification to enhance early detection and understand its biological significance. Despite the rapid progress of machine learning in bioinformatics, accurately detecting 6&#xa0;mA sites remains a challenge due to the limited generalizability and efficiency of existing approaches. In this study, we present Deep-N6mA, a novel Deep Neural Network (DNN) model incorporating optimal hybrid features for precise 6&#xa0;mA site identification. The proposed framework captures complex patterns from DNA sequences through a comprehensive feature extraction process, leveraging k-mer, Dinucleotide-based Cross Covariance (DCC), Trinucleotide-based Auto Covariance (TAC), Pseudo Single Nucleotide Composition (PseSNC), Pseudo Dinucleotide Composition (PseDNC), and Pseudo Trinucleotide Composition (PseTNC). To optimize computational efficiency and eliminate irrelevant or noisy features, an unsupervised Principal Component Analysis (PCA) algorithm is employed, ensuring the selection of the most informative features. A multilayer DNN serves as the classification algorithm to identify N6-methyladenine sites accurately. The robustness and generalizability of Deep-N6mA were rigorously validated using fivefold cross-validation on two benchmark datasets. Experimental results reveal that Deep-N6mA achieves an average accuracy of 97.70% on the F. vesca dataset and 95.75% on the R. chinensis dataset, outperforming existing methods by 4.12% and 4.55%, respectively. These findings underscore the effectiveness of Deep-N6mA as a reliable tool for early 6&#xa0;mA site detection, contributing to epigenetic research and advancing the field of computational biology.

Deep Learning

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348&#xa0;handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding

Improving insurance deduction identification: a hybrid artificial intelligence model using machine learning and expert systems.

PURPOSE: Financial challenges in healthcare systems worldwide, especially in low- and middle-income countries like Iran, have increased hospitals' reliance on insurance reimbursements. Unrecognized insurance deductions often cause severe financial shortages, making efficient deduction management crucial. This study aimed to design a hybrid intelligent system for identifying and predicting insurance deductions by combining machine learning and expert system frameworks. DESIGN/METHODOLOGY/APPROACH: A mixed-methods design was applied in four stages. First, a scoping review identified the causes and patterns of insurance deductions. Second, interviews with 15 insurance experts produced a validated checklist and a dataset from inpatient billing records. Third, using the CRISP-DM methodology, machine learning algorithms were developed and tested in SPSS Modeler alongside a fuzzy expert system developed in MATLAB. Finally, the model was validated using the holdout method. FINDINGS: Four categories of deduction drivers were identified: service provision, registration errors, document submission issues, and revenue conversion processes. The CHAID decision tree outperformed other algorithms with a 99% precision rate and the lowest Mean Absolute Error (9.43). A brief assessment of potential overfitting was conducted to ensure that the CHAID model's high accuracy was interpreted cautiously and supported by the validation results. The fuzzy expert system with validated rules was adaptable for deduction classification, especially for cases unsuitable for quantitative modeling. ORIGINALITY/VALUE: The hybrid model improves detection and prevention of deductions, offering actionable insights for hospital administrators, insurers, and policymakers. Its implementation can enhance hospital information systems, streamline claims processing, and optimize revenue management amid financial constraints.

Machine Learning