Search PubMedSearch

PubMed · 42728485

Molecular characterization and genome sequence analysis of Dichroa emaravirus, a putative novel member of the genus Emaravirus.

Abstract

Hydrangea febrifuga (syn. Dichroa febrifuga) is a traditional medicinal plant distributed in China and Southeast Asia, and febrifugine, one of its principal bioactive constituents, has served as an important lead compound for antimalarial drug development. Viral infections may adversely affect the quality of medicinal plants; however, no emaravirus has previously been reported from H. febrifuga. Here, high-throughput sequencing was performed on H. febrifuga leaves exhibiting mosaic symptoms collected in Yunnan Province, China. Combined with RT-PCR, Sanger sequencing, and 5'/3' rapid amplification of cDNA ends (RACE), five full-length genomic RNA segments of a putative novel emaravirus, tentatively designated Dichroa emaravirus (DEV), were identified and characterized. The five negative-sense single-stranded RNA (-ssRNA) segments have a combined length of 12,971 nt and encode an RNA-dependent RNA polymerase (RdRp), glycoprotein precursor (GP), nucleocapsid protein (NP), movement protein (MP), and an uncharacterized accessory protein, P5. The maximum amino acid sequence identities of DEV P1-P4 with recognized emaraviruses were 73.90%, 51.82%, 65.60%, and 81.30%, respectively, whereas P5 showed a maximum identity of 49.16% with its closest homolog. Thus, three of the four core proteins had maximum identities below 80%, consistent with the current ICTV species demarcation criterion for the genus Emaravirus. Maximum-likelihood phylogenetic analyses based on the four core proteins further supported the placement of DEV within the genus Emaravirus (family Fimoviridae). These results support DEV as a putative novel emaravirus and represent the first report of an emaravirus associated with H. febrifuga.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Fangmei Xiao, Yuanjing Chen, Shengming Liu, Xuehua Li, Daihong Yu, Lin Li, Yizhao Qu, Yuan Su, Yonghong Yang, Mingfu Zhao. 2026-09-11. Molecular characterization and genome sequence analysis of Dichroa emaravirus, a putative novel member of the genus Emaravirus.. https://doi.org/10.1007/s00705-026-06725-y

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

The complete genomic sequence of a novel member of the genus Caulimovirus isolated from Dregea volubilis.

A novel caulimovirus was identified from diseased leaves of Dregea volubilis exhibiting yellowing and vein-associated chlorosis in Yuanjiang County, Yunnan Province, China. The virus was tentatively named Dregea volubilis caulimovirus 1 (DVCaV1). The complete genome sequence of DVCaV1, determined by de novo assembly of high-throughput sequencing data, comprises 8,160 bp of circular double-stranded DNA containing two intergenic regions and seven open reading frames (ORFs). These ORFs encode (in order) a movement protein (MP), an aphid transmission factor (ATF), a virion-associated protein (VAP), a coat protein (CP), a polymerase polyprotein (Pol, containing protease, reverse transcriptase, and RNase H domains), a transactivator/viroplasmin (TAV) protein, and a hypothetical protein of unknown function. Sequence comparisons revealed the highest nucleotide similarity with strawberry vein banding virus (SVBV; NC_001725). Phylogenetic analysis confirmed DVCaV1 as a member of the genus Caulimovirus, with SVBV as its closest known relative. According to current ICTV species demarcation criteria for the genus Caulimovirus (host range and > 20% nucleotide sequence divergence in the polymerase region), DVCaV1 represents a novel species. This is, to our knowledge, the first report of a caulimovirus detected in naturally symptomatic Dregea volubilis.

Genome, Viral

Complete genome sequence of a novel alternavirus infecting Fusarium falciforme.

We present the complete genome sequence of a novel alternavirus, tentatively named "Fusarium falciforme alternavirus 1 (FfAV1)", isolated from Fusarium falciforme. The host, F. falciforme strain Fod375, was isolated from a soil sample in Spain in 2012 and was found to be infected with a virus containing a tetra-segmented double-stranded (ds) RNA genome. The genome segments, designated as dsRNA1 (3529 bp), dsRNA2 (2641 bp), dsRNA3 (2459 bp), and dsRNA4 (1471 bp), each possess a single open reading frame (ORF). The protein predicted from dsRNA1 contains the typical domains of an RNA-dependent RNA polymerase (RdRP) homologous to those of previously reported alternaviruses, while the protein predicted from dsRNA3 shows homology to alternavirus capsid proteins. The proteins encoded by dsRNA2 and dsRNA4 are of unknown function. All predicted proteins exhibited the highest sequence identity with their counterparts in Hebei alternavirus and Marquandomyces marquandii alternavirus 1. Phylogenetic analysis supported the placement of this FfAV1 isolate within the genus Alternavirus. Considering these results, we propose that FfAV1, along with the two closely related unassigned alternaviruses, represents a new species within the genus.

Genome, Viral

A unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification.

The rapid growth of genomic sequencing demands fast, accurate, and scalable analysis methods. In viral genomic classification, expanding labeled reference collections can make supervised models costly to update and dependent on fixed label sets, motivating retrieval-based genomic classification as a simpler, more flexible alternative. We present a unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification across three viral classification tasks: hepatitis C virus (HCV) genotyping, COVID-19 discrimination, and human papillomavirus (HPV) genotyping. We compare standard sequence encodings (one-hot, k-mers, FCGR) with dense embeddings (dna2vec, DNABERT). For each representation, we evaluate supervised classifiers (Random Forest, Decision Tree, XGBoost) and retrieval-based classification, where sequence vectors are indexed with FAISS and labels are assigned via similarity-weighted k-NN. Furthermore, we benchmark multiple FAISS index types (Flat, IVF, HNSW, IVFPQ, OPQ) to characterize accuracy-speed-memory trade-offs at scale. The results show that XGBoost and retrieval using Flat or IVF indexes achieve strong classification performance under different computational profiles. Compressed indexes such as IVFPQ and OPQ substantially reduce memory usage, although their accuracy loss depends on the dataset and representation. Overall, supervised XGBoost provides a favorable accuracy-size trade-off, while retrieval-based classification remains competitive and allows labeled reference sequences to be incorporated without retraining a global classifier. This benchmark provides practical guidance for selecting sequence representations, classifiers, and vector-search indexes under different accuracy, memory, and update requirements.

Genome, Viral