Search PubMedSearch

Biomedical subjects

Ranjit Prasad Bahadur

Publications and source records attributed to Ranjit Prasad Bahadur.

2 recordsLinked to original sources

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348 handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding

An atlas of non-redundant sequences and structures of transcription factor assemblies across domains of life.

Transcription factors (TFs) regulate gene expression by controlling the recruitment of transcriptional machinery to regulatory regions of the genome. Nearly 10% of the human genome encodes TFs, making them one of the largest protein families. Despite their central roles in gene regulation, TFs are historically considered challenging therapeutic targets due to their complex interactions with DNA, RNA and associated proteins. Although recent progress in studying TFs both at molecular and structural level excels our understanding on their function, yet a universal rule decoding their recognition process remains elusive. Here, we present a curated non-redundant dataset of TFs with 3570 sequences and 377 structures. We further characterize "unique interfaces" by quantifying interface identity across interacting chains in TF assemblies. Surprisingly, our data shows that the "unique interfaces" have optimal size ranging from 2000 Å2 to 4000 Å2 irrespective of their quaternary assembly. To understand the functional diversity, we integrate sequence motifs, structural domains, subcellular localization and functional enrichment of TFs. We have also catalogued association of TFs with various human diseases. Our dataset provides a comprehensive platform to perform large scale analysis of TF-assemblies and aid in computational methods for their prediction across domains of life.

Gene regulation