Search PubMedSearch

SEARCH · Search PubMed

Results for “Base Sequence”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

572 recordsLinked to original sources

Evaluation of one-step amplicon-based targeted enrichment for SARS-CoV-2 whole-genome sequencing using the Midnight amplicon scheme.

Genomic surveillance proved invaluable during the COVID-19 pandemic for tracking SARS-CoV-2 variants and guiding outbreak responses, underscoring the ongoing need to reduce whole-genome sequencing (WGS) costs and improve workflow efficiency to ensure accessibility in resource limited settings. Here, we evaluated a one-step reverse transcription polymerase chain reaction (RT-PCR) approach using the Midnight V2 primer scheme for targeted amplification of the SARS-CoV-2 genome, assessed its compatibility with Illumina sequencing, and compared its performance to a well-established two-step method. Initially, we determined optimal RT-PCR reaction conditions using the Midnight V2 primer panel for the one-step RT-PCR kit and scaled reaction volumes for both RT-PCR and library preparation. Clinical specimens (n = 53) that had undergone routine WGS for surveillance purposes using the established two-step RT-PCR method were compared using the one-step RT-PCR assay. For samples with genome completeness greater than 70%, both methods gave comparable results with similar sequence coverage and 100% concordance for lineage assignment. Further investigation revealed a higher percentage of reads aligning to the SARS-CoV-2 genome with a greater depth of coverage using the one-step method compared to the two-step method. Finally, analysis of scaled one-step and library reaction volumes revealed significant cost savings for samples undergoing WGS. Overall, the results presented here verify the accuracy and reproducibility of one-step targeted amplification and offer an efficient and cost-effective workflow for routine SARS-CoV-2 genomic surveillance.

Humans

Nanopore-based epigenomic profiling reveals the absence of widespread CpG methylation in the African swine fever virus genome.

DNA methylation is a critical epigenetic mechanism implicated in regulating replication and transcription in DNA viruses. However, the epigenetic landscape of African swine fever virus (ASFV), a large double-stranded DNA virus infecting pigs, remains controversial. Here, we systematically profiled the DNA methylome of the first ASFV strain isolated in Hong Kong (HK_NT_202103) using Oxford Nanopore Technologies (ONT) R10.4.1 sequencing. We employed a paired design: native whole-genome sequencing (WGS) against a methylation-free whole-genome amplification (WGA) control. Using conservative thresholds, we found no evidence of 5-methylcytosine (5mC), especially typical CpG methylation, across the viral genome. Importantly, clear CpG methylation signals were successfully detected in the host genome from WGS data, confirming the functionality of the workflow to detect 5mC at CG sites. While widespread 5mC seems absent, a small number of putative N6-methyladenine (6mA) loci were identified. A specific 6mA candidate exhibited raw ionic current disruptions and gene-level intersection with another ASFV isolate (CAS19-01/2019), although it lacked single-base consensus across different methylation callers or between the two isolates. Although our biological findings are restricted to a single isolate under specific experimental conditions, this study introduces a novel, highly rigorous ONT framework for viral epigenomics research. Furthermore, the absence of ASFV CpG methylation indicates that host CpG-depletion remains a viable strategy for viral metagenomic enrichment. Ultimately, our work offers a critical methodological baseline for ASFV surveillance and highlights the necessity of targeted experimental validation for rare viral modifications.

African Swine Fever Virus

A genome-wide coverage-based pipeline for the identification of host-derived candidate DNA biomarkers from cell-free blood.

We have created a new data-analysis pipeline for the discovery of host-specific candidate DNA biomarkers derived from sequencing data of cell-free blood. Unlike approaches that rely on specific molecular or genetic signatures, our method leverages the coverage distribution of cell-free DNA sequences mapped to a reference genome, applying statistical analyses to identify informative short genomic regions for biomarker discovery. The pipeline is applicable to diverse diseases and can be used to analyze cell-free DNA sequences from plasma or serum to identify candidate biomarkers that are characteristic of disease states in mammals. Core functionalities were developed in Java and integrated with open-source software tools for the preprocessing of raw sequencing data, complemented by Python scripts for the machine-learning analysis and statistical validation. The pipeline is designed for HPC use and users can access the pipeline through a Galaxy workflow, which offers a user-friendly web interface for input selection prior to execution and analysis progress monitoring. Performance tests, carried out using duplicate sets of COVID-19 samples and controls, showed linear scalability of execution time with an increasing dataset size, as well as a substantial reduction in execution time through parallelized computation, whereby each HPC node is used to process the data of one chromosome. Further statistical tests confirmed the quality of the pipeline's results by showing that the set of identified candidate biomarkers remained stable across varying dataset sizes.

Biomarkers

Performance comparison of rapid and native barcoding methods for Oxford Nanopore sequencing of Poliovirus Viral Protein 1 (VP1) amplicons.

Accurate and timely sequencing of poliovirus is critical for global eradication efforts, particularly for molecular epidemiology based on the typing region of the genome, viral protein 1 (VP1). While Oxford Nanopore Technologies (ONT) sequencing has expanded capabilities for poliovirus surveillance, the relative performance of different ONT library preparation methods, including ligation-based (Native Barcoding) and transposase-based (Rapid Barcoding) approaches, has not been systematically evaluated. In this study, we compared rapid barcoding and native barcoding workflows for sequencing VP1 amplicons from 17 type 2 poliovirus-positive samples, each processed in triplicate. Native barcoding generated significantly more sequencing output, producing approximately 2.3-fold greater total read yield than rapid barcoding, and demonstrated higher run-to-run reproducibility (R2 = 0.979-0.998 vs. 0.847-0.929, respectively; p&#x202f;<&#x202f;0.001). In addition, native barcoding generated 80% of the total yield achieved by rapid barcoding within approximately 7&#x202f;h, whereas rapid barcoding required approximately 40&#x202f;h to reach the same output. Despite these differences, both methods produced identical VP1 consensus sequences across all samples, with comparable read quality (median per-base Q-scores of approximately Q17-Q18). Rapid barcoding provided substantial practical advantages, reducing hands-on library preparation time (55 vs. 200&#x202f;min) and per-sample cost ($12.82 vs. $16.54), while simplifying workflow and reducing technical complexity. These findings indicate that sequencing yield may not be a determinant of downstream analytical outcomes for poliovirus VP1 ONT sequencing. Rapid barcoding therefore represents a cost-effective and efficient approach for routine poliovirus surveillance, whereas native barcoding remains advantageous in applications requiring rapid data generation or maximal sequencing depth.

Poliovirus

LitCTL1: A novel C-type lectin involved in the mucosal and cellular immunity of the common periwinkle Littorinalittorea.

C-type lectins (CTLs) are vital pattern-recognition receptors (PRRs) that mediate innate immune responses in mollusks, yet their characterization in Caenogastropoda, the largest gastropod group, remains limited. This study characterizes LitCTL1, a novel secreted single-domain C-type lectin from the common periwinkle, Littorina littorea. The 199-amino acid polypeptide contains a conserved carbohydrate recognition domain with canonical QPD and WND motifs and is predicted to form a homodimer. Uniquely, LitCTL1 was localized in both circulating hemocytes and mucus-secreting epithelial cells of the foot, mantle, and hypobranchial gland - the first report of such dual localization for a molluscan lectin, linking systemic and mucosal defense. Expression analysis revealed that LitCTL1 is constitutively expressed in hemocytes. Functional assays with recombinant LitCTL1 demonstrated its role as a potent opsonin with hemagglutinating activity, significantly enhancing hemocyte spreading and the phagocytosis of zymosan. Genomic analysis reveals that LitCTL1 belongs to a rapidly diversifying, genus-specific expansion distinct from conserved perlucin-like lineages. These results identify LitCTL1 as a key effector molecule in both systemic and mucosal innate immunity, likely reflecting an evolutionary adaptation to the microbial challenges of the intertidal environment.

Animals

Uce-based phylogeny and classification of Megachilini.

The generic-level classification of the bee tribe Megachilini (Megachilidae) has remained controversial due to poor phylogenetic resolution at the base of the group, particularly among the brood parasitic genera and the numerous dauber ("Chalicodoma s. l.") lineages. We present a phylogenomic analysis of Megachilini based on ultraconserved elements (UCEs), sampling 52 ingroup taxa with emphasis on the dauber lineages. We also present a combined UCE&#xa0;+&#xa0;six-gene analysis to improve taxon coverage, resulting in a dataset with 127 ingroup taxa. Maximum likelihood, coalescent, and Bayesian analyses of multiple UCE matrices recover largely congruent topologies with substantially improved support relative to previous studies. Our results strongly support the monophyly of Megachilini, the early divergence of Noteriades and Gronoceras, and a single origin of brood parasitism. All remaining non-parasitic Megachilini form a moderately supported clade sister to the brood parasitic lineage. The leafcutter bees are monophyletic and nested within dauber lineages. Several major dauber clades are consistently recovered, including an exclusively Australian clade corresponding to the Hackeriapis group of subgenera, while several recognized subgenera are paraphyletic. The lineage known as Morphella, previously placed in synonymy with the subgenus Callomegachile, was not closely related to that subgenus and is here treated as a valid subgenus. Divergence-time analyses place the crown age of Megachilini in the late Eocene to early Oligocene, with major extant lineages diversifying during the Miocene. Limited morphological diagnosability of several clades indicates that splitting non-parasitic lineages into numerous genera would result in an impractical classification that would widen the gap between taxonomists and non-specialists and exacerbate the taxonomic impediment in bees. We therefore advocate retaining a single genus Megachile for non-parasitic Megachilini (excluding Noteriades and Gronoceras), as the classification best supported by phylogenomic evidence and most robust to future taxon sampling.

Animals

Whole genome sequencing of unusual Hepatitis C virus subtypes and drug resistance analysis during direct-acting antiviral therapy in India.

INTRODUCTION AND OBJECTIVES: Pangenotypic direct-acting antivirals (DAA) are effective against highly prevalent Hepatitis C virus (HCV) subtypes, but have been clinically validated almost exclusively in high-income countries. Unusual HCV subtypes may carry natural polymorphisms, potentially impacting DAA susceptibility. We conducted full-genome characterization and resistance analysis of unusual HCV subtypes in patients receiving DAA treatment. PATIENTS AND METHODS: In this prospective hospital-based study, eligible patients were screened for anti-HCV antibodies and active infection was confirmed by diagnostic 5'NCR-based HCV RNA detection. Genotyping was performed by core region sequencing, and viral load quantified by real-time PCR. For whole genome sequencing, multiplex primers were designed using alignments of global reference sequences. Sequencing was carried out using the Oxford Nanopore Technology platform. Phylogenetic analysis used multiple sequence alignment and the HCV-GLUE resource for resistance-associated substitution (RAS) analysis. RESULTS: Predominant genotype was genotype 3 in 64.3% (n = 45); genotype 6 in 21.4% (n = 15); and genotype 1 in 14.2% (n = 10). Unusual HCV subtype 6xa was detected in two patients and showed no NS5A resistance mutations. One genotype 3b patient relapsed at 24 weeks post-DAA treatment completion and carried NS5A resistance-associated substitutions 30 K and 31 M both at baseline and at relapse, conferring high-level resistance to NS5A inhibitors. CONCLUSION: This is the first report from India of whole genome sequencing of HCV subtype 6xa. The identification of NS5A resistance mutations in the 3b relapse case underscores challenges for global HCV elimination strategies.

Humans

Routine methods misidentify Serratia spp.: Limitations of MALDI-TOF MS revealed by whole-genome sequencing.

Accurate species-level identification within the genus Serratia remains challenging due to extensive phenotypic overlap and high genomic relatedness among closely related and recently described taxa. This study presents an evaluation of routine and genome-based identification approaches applied to clinical Serratia isolates, integrating phenotypic assays, MALDI-TOF MS (Bruker Daltonics), 16S rRNA gene sequencing, and Whole-Genome Sequencing (WGS). A total of 103 isolates collected from a teaching hospital were analyzed. WGS was performed on a subset of isolates. Conventional biochemical methods classified all isolates as Serratia marcescens, whereas MALDI-TOF MS identified 60.1% as S. marcescens, 11.6% as S. ureilytica, and 28.1% just at the genus level. Peak analysis from MALDI-TOF MS revealed specific peaks associated with S. marcescens and S. ureilytica, but limited discriminatory power. WGS of six isolates initially identified as S. ureilytica by MALDI-TOF MS revealed reclassification as Serratia sarumanii (n = 5) and Serratia montpellierensis (n = 1), supported by Average Nucleotide Identity (ANI), Average Amino Acid Identity (AAI), and Digital DNA-DNA Hybridization (dDDH) thresholds. In contrast, 16S rRNA analysis showed limited species-level resolution. Phylogenomic and SNP-based analyses confirmed these classifications with strong support. Overall, this study underscores the critical role of high-resolution genomic approaches for precise species identification and highlights the need for continuous expansion and curation of MALDI-TOF MS reference databases to support reliable clinical diagnostics and epidemiological surveillance of emerging Serratia species.

Spectrometry, Mass, Matrix-Assisted Laser Desorpti

Comparative genomic epidemiology of food- and patient-derived diarrheagenic Escherichia coli from sentinel surveillance in Southeast China.

Diarrheagenic Escherichia coli (DEC) remains an important foodborne pathogen, yet long-term comparative genomic surveillance data jointly characterizing food-derived and patient-derived isolates remain limited. This surveillance-based comparative study integrated antimicrobial susceptibility testing and whole-genome sequencing to characterize diarrheagenic Escherichia coli isolates recovered from food and patient sources in Lishui, Southeast China, during 2018-2025, with emphasis on occurrence, resistance profiles, genomic backgrounds, and plasmid replicon-associated features. Antimicrobial susceptibility testing was performed for 258 selected isolates, and whole-genome sequencing was conducted for a curated analytical subset of 204 isolates. The sequenced subset was used for diversity-oriented comparative genomic analysis rather than for unbiased prevalence estimation of the entire DEC collection. EAEC predominated in both sources, although food-associated occurrence was heterogeneous across categories, with the highest recovery rate observed in raw meat. Patient-derived isolates showed a broader overall resistance burden, whereas food-derived isolates retained substantial resistance to tetracycline, chloramphenicol, and florfenicol. Phylogenetic analysis showed partial overlap in genomic backgrounds between food-derived and patient-derived isolates, while representative resistance determinants displayed both broadly distributed and lineage-enriched patterns. Replicon-based plasmid profiling identified 42 plasmid types, including 12 detected in both sources, with IncF-related replicons predominating among these shared profiles. Several food-derived isolates carried multiple plasmid replicon types that were also observed in patient-derived isolates. Overall, food-derived and patient-derived DEC showed partial overlap in genomic backgrounds, resistance determinants, and replicon-defined plasmid profiles within this surveillance setting, while retaining source-associated heterogeneity. These findings should be interpreted as surveillance-based comparative evidence rather than as evidence of direct source attribution or transmission.

Humans

Molecular Landscape and Advanced Diagnostic Technologies for BRAF Mutations in Cancer: From Quantitative PCR and ddPCR to CRISPR-Based Platforms.

BRAF mutations are key oncogenic alterations across multiple malignancies, including melanoma, thyroid carcinoma, colorectal cancer, non-small cell lung cancer, glioma, and hairy cell leukemia. The most prevalent variant, BRAF-V600E, induces constitutive activation of the MAPK signaling pathway, promoting tumor progression and influencing therapeutic responsiveness. Accurate detection of BRAF alterations is therefore essential for molecular classification, prognostic assessment, treatment selection, and resistance surveillance. This review summarizes the molecular heterogeneity of BRAF mutations and critically evaluates current diagnostic methodologies. Conventional approaches such as allele-specific PCR and Sanger sequencing are compared with advanced quantitative platforms, including high-resolution melting analysis, droplet digital PCR, and next-generation sequencing, with emphasis on analytical sensitivity, mutation coverage, and clinical applicability. Emerging technologies such as CRISPR-based assays, rolling circle amplification systems, and nanoparticle-based biosensors and point-of-care diagnostic platforms are also discussed for their potential to enhance ultra-sensitive detection, particularly in liquid biopsy settings. These emerging tools are highlighted for their potential to enable ultra-sensitive, rapid, and decentralized mutation detection, particularly in liquid biopsy settings. Key challenges, including intratumoral heterogeneity, low allele-frequency variants, FFPE-associated artifacts, and clonal evolution under therapeutic pressure, are examined within a translational framework. In addition, we examine critical barriers to clinical implementation, including standardization, cost, and global accessibility of molecular diagnostics, and outline potential solutions through scalable technologies and decentralized testing strategies. We propose that optimal BRAF testing requires a mutation subclass-informed and clinically integrated strategy combining comprehensive baseline profiling with longitudinal molecular monitoring. Future diagnostic paradigms will likely integrate multi-omics data and artificial intelligence (AI)-assisted interpretation to refine precision oncology implementation. Looking forward, we propose that optimal BRAF testing will require integration of multi-omics profiling with AI-assisted interpretation, enabling automated variant classification, real-time clinical decision support, and improved prediction of therapeutic response and resistance.

Humans

Targeted sequencing reveals a distinct genetic alteration landscape in oral multiple primary squamous cell carcinomas.

OBJECTIVE: Oral multiple primary cancers (MPCs) are associated with poor clinical outcomes, yet their genomic characteristics remain insufficiently understood. DESIGN: Fifty-four formalin-fixed paraffin-embedded (FFPE) tumor samples from 30 patients with oral MPCs were analyzed using high-depth targeted sequencing of a customized 14-gene panel derived from prior whole-exome sequencing data. Detected alterations were analyzed after removal of synonymous mutations. RESULTS: Non-silent genomic alterations were identified in 59.3% (32/54) of samples, involving 19 patients. A total of 70 variant loci across 13 genes were detected. AKAP13 was the most frequently mutated gene at both the sample (22.2%, 12/54), with recurrent mutations observed across multiple patients. In contrast, TP53 mutations occurred at a substantially lower frequency (11.1%, 6/54). Marked inter- and intra-patient mutational heterogeneity was observed. CONCLUSIONS: FFPE-based targeted sequencing enabled an initial characterization of genomic alterations in oral MPCs. Recurrent alterations in AKAP13, GLI2, JMJD1C, and DNAH8, together with the relatively low frequency of TP53 alterations, identify candidate genomic features for further investigation and provide a basis for future studies of the molecular basis of oral MPCs.

Humans

Transcriptomic and RNAi analyses reveal chloride channel 3-associated osmoregulation in Litopenaeus vannamei under low-salinity stress.

Chloride channels and transporters are important for cellular volume regulation and salinity adaptation in euryhaline crustaceans, yet the intestinal transcriptional relationship between plasma-membrane and intracellular chloride pathways remains unclear in Litopenaeus vannamei. In this study, RNA interference of anoctamin 1 (ANO1) was combined with intestinal transcriptome sequencing under the production-relevant low-salinity condition of salinity 3. ANO1 silencing produced a focused transcriptional response, with 16 differentially expressed genes (DEGs) identified (11 upregulated and 5 downregulated). Functional enrichment indicated that these genes were associated with transporter activity, cytoskeletal organization, extracellular matrix-receptor interaction, membrane lipid metabolism, and vesicular processes. Notably, a transcript encoding chloride channel protein 3 (CLC-3) was significantly upregulated following ANO1 knockdown, suggesting a potential transcriptional relationship between ANO1 and CLC-3 in chloride homeostasis. Based on this finding, CLC-3 was selected for full-length cDNA cloning, sequence characterization, salinity-gradient expression analysis, and RNAi-based functional assessment. The cloned CLC-3 cDNA was 2883&#xa0;bp in length and encoded an 850 amino acid protein containing a conserved voltage-gated chloride channel (Voltage-CLC) domain and two cystathionine &#x3b2;-synthase domains. Phylogenetic analysis placed LvCLC-3 within the intracellular CLC-c clade, and tissue distribution analysis showed the highest CLC-3 expression in the intestine. Intestinal CLC-3 expression responded nonlinearly to salinity variation, peaking at salinity 20. Under salinity 3, CLC-3 knockdown reduced ANO1, Na+/K+-ATPase alpha subunit, and Na+-K+-2Cl- cotransporter transcript levels, whereas glutamate-gated chloride channel expression increased. Mild hepatopancreatic structural alterations were also observed after CLC-3 knockdown. These findings suggest that CLC-3 is a salinity-responsive intracellular chloride-transporter candidate associated with intestinal ion-transport-related transcriptional responses after ANO1 suppression in L. vannamei, although the underlying physiological mechanism requires further validation.

Animals

Single-cell RNA sequencing provides further insights into the immunostimulatory action of freeze-dried Lactiplantibacillus plantarum on Penaeus vannamei shrimp.

Immunostimulation through dietary interventions opened new avenues in developing disease control and prevention tools for shrimp aquaculture. We have previously shown that feeding with freeze-dried Lactiplantibacillus plantarum (LAB) increased disease resistance of Penaeus vannamei against both Vibrio parahaemolyticus and white spot syndrome virus (WSSV) based on bulk RNA sequencing of shrimp gills. This tissue participates in ion transport and serves as a first line of defense against environmental stressors and pathogenic infections. However, characterization of their cell composition and functions remains limited. Here, we implemented a single-cell RNA sequencing approach to further gather insights into how feeding with freeze-dried LAB modulates host immunity which may not be evident with bulk RNA sequencing approach. A total of five clusters with unique transcriptional signatures were identified, corresponding to pillar cells, septal cells, and sessile hemocytes. Pseudo-bulk analyses at global- and cluster-levels showed differential expression of genes related to host immunity and metabolism. We further revealed how overall transcriptomic changes are not exclusively caused by gene expression changes but may also be driven by cell population dynamics. This study highlighted how single-cell RNA sequencing approach may shed light on the mechanisms of action of immunostimulants which may be masked in bulk transcriptome analyses.

Animals

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (&#x3a8;) represents one of the most abundant and conserved RNA modifications. &#x3a8; provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of &#x3a8; sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel &#x3a8; site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA &#x3a8;-site prediction. The &#x3a8; modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA &#x3a8;-site prediction. Meta-PseU offers a new framework for robust &#x3a8;-site identification by using long sequences.

Pseudouridine

Whole-transcriptome RNA sequencing and ceRNA network analyses provide novel insights into the antibacterial immune response of Hippocampus abdominalis against Vibrio harveyi.

Long non-coding RNAs (lncRNAs) stand as newly-arisen molecular types that exert regulatory effects, able to operate as competitive endogenous RNAs (ceRNAs) to engage microRNAs (miRNAs) in interaction, resulting in the recovery of target mRNA expression and activity. Increasing evidences indicate that the ceRNA network affects various biological processes in mammals, including development, cellular differentiation, metabolism, immune response, and disease pathogenesis. In teleost fish, the lncRNA-miRNA-mRNA regulatory networks have been reported occasionally. However, up to now, the roles of lncRNAs in the big-belly seahorse (Hippocampus abdominalis) remains unclear. In this study, we reported for the first time, via whole-transcriptome RNA sequencing, the lncRNA mediated ceRNA regulatory network in Vibrio harveyi-infected H. abdominalis. A total of 4197 differentially expressed mRNAs (DE-mRNAs), 1317 DE-lncRNAs, and 183 DE-miRNAs were identified. Furthermore, the crosstalk between miRNAs and lncRNAs as well as between miRNAs and mRNAs was inferred based on the negative correlations between miRNAs and their target lncRNAs/mRNAs. A core immune associated lncRNA-miRNA-mRNA putative regulatory network was thus constructed, comprising 211 lncRNA-miRNA and 224 mRNA-miRNA pairs. In conclusion, our findings provide an integrative overview of the ceRNA regulatory networks on the underlying immune responses to V. harveyi infection in the big-belly seahorse, and offer a solid theoretical foundation for the comparative immunological research of teleost fish.

Animals

A contextual activity score (CAS) for inferring ADAR-associated transcriptional activity across RNA-seq, single-cell, and spatial transcriptomics.

BACKGROUND AND OBJECTIVE: Adenosine-to-inosine RNA editing, catalyzed by Adenosine Deaminases Acting on RNA (ADARs), is a widespread modification involved in neural function, immune regulation, and cancer. The Alu Editing Index (AEI) is the standard metric to estimate ADAR activity but requires raw sequencing reads and is poorly suited for single-cell and spatial transcriptomic data. This study aimed to develop an alternative framework for inferring ADAR-associated transcriptional activity from gene expression data across diverse transcriptomic technologies. METHODS: We developed the Contextual Activity Score (CAS), a framework based on transcriptional signatures from ADAR perturbation experiments. Context-specific signatures were generated for human neurons, mouse neurons, and cancer models to infer ADAR1 and ADAR2 activity. CAS was computed from normalized gene expression matrices using regulon-based enrichment analysis. Performance was evaluated by comparing with the Alu Editing Index across bulk RNA sequencing datasets, simulated sequencing depths, and library preparation protocols. RESULTS: CAS showed strong concordance with the Alu Editing Index across multiple datasets, while remaining robust to reduced sequencing depth and different library protocols. Unlike the Alu Editing Index, CAS can be applied to single-cell and spatial transcriptomic data and enables the independent assessment of ADAR2 activity. In cancer and neuronal contexts, CAS captured biologically meaningful variations in ADAR-associated transcriptional activity at sample, cell-type, and spatial levels. CONCLUSION: CAS provides a scalable approach applicable across multiple RNA-seq protocols for estimating ADAR-associated transcriptional activity using gene expression data. This method, implemented in an open-source R package for broad adoption, expands the ability to study ADAR-associated transcriptional activity across transcriptomic modalities where direct editing quantification is challenging, such as single-cell and spatial transcriptomics.

Adenosine Deaminase

A conserved distal-tail helical extension defines a tailspike attachment architecture in Gram-negative siphophages.

Rapid growth of bacteriophage genome collections has outpaced functional annotation of tail-tip proteins, limiting comparative analysis of host-recognition structures. Starting from a shared distal-tail gene organization in the Salmonella phages 9NA and Jersey, I developed a morphogenetic bioinformatic framework integrating gene synteny, sequence comparison, profile hidden Markov model (HMM) screening, structural evidence, structure-aware searching, and AlphaFold modeling. Comparison with the experimentally characterized lambda and Sf11 tail assemblies identified a predominantly alpha-helical C-terminal extension of the distal-tail (DT) protein associated with tailspike attachment, termed the distal-tail helical extension (DT-helix). Screening 541,986 proteins from 5167 complete NCBI RefSeq tailed-phage genomes, followed by evidence-based evaluation of sequence, genomic context, and structural architecture, identified 165 curated DT-helical-extension-associated phages. Their DT proteins segregated into six sequence groups. In the four principal multi-member groups, cognate tailspikes showed group-specific conservation in proximal N-terminal regions but substantially greater downstream diversity, consistent with sequence constraint at the DT-tailspike attachment boundary. A complementary ProstT5/Foldseek search supported the established groups but revealed no convincing additional highly divergent family. Together with the experimentally characterized Sf11 attachment interface, these findings define a recurrent morphogenetic architecture linking conserved distal-tail scaffolds to more variable receptor-binding proteins across siphophages infecting Gram-negative bacteria. Although universal exchangeability is not established, the identified scaffold-receptor-binding boundaries provide a framework for molecular characterization and rational phage engineering. Accession-level information for the 165 curated phages is available through PhageTailDB.

Viral Tail Proteins

Comparison of paralog identification methods and their impact on species tree topologies in target capture phylogenomics within the Sindora clade (Detarioideae: Leguminosae).

Target capture is a common method of generating high throughput DNA sequencing data for phylogenetic reconstruction of species relationships, for which single copy genes are usually most informative. However, a pervasive problem with target capture is that putatively single copy genes may in fact be paralogs resulting from gene duplication, which are problematic for phylogenetic inference because their evolutionary history may differ from the divergence history of species. Here, we use as a case study a target enrichment dataset of 88 species of Detarioideae (Leguminosae) with a focus on the Sindora clade to examine approaches for handling paralogs, including the built-in paralog handling functions in HybPiper and CAPTUS, plus subsequent steps using Putative Paralog Detection and the tree-based Yang & Smith orthology inference approach. We compare the paralogs flagged using these methods and verify their performance with BLAST mapping against a reference genome sequence of Sindora glabra, and then subsequently compare the species tree topologies produced across these methods. Our comparisons of paralogs flagged across the Sindora clade show that the Putative Paralog Detection pipeline was the most accurate in identifying paralogs in terms of its similarity to the BLAST mapping, followed by the built-in paralog identification function of CAPTUS. However, the results we recovered for the Detarioideae subfamily suggest that the largest differences in species tree topology resulted from the use of paralog-filtered alignments (such as with the Putative Paralog Detection pipeline and the Yang & Smith orthology inference approaches) rather than just by removing the sequences of identified paralogous genes. This was the true for HybPiper-assembled datasets but was not seen in CAPTUS-assembled datasets. In all comparisons, the topological differences caused by different paralog handling methods tended to be confined to clades where processes such as hybridisation and introgression are prevalent. Our study provides a roadmap to establish the best approach to identify, eliminate or separate paralogs in the absence of a chromosomally contiguous reference genome for a study group, and highlights the importance of careful data inspection and processing in addition to understanding the extent of paralogy and paralog characteristics (e.g. sequence divergence between copies) for their study group.

Phylogeny