Search PubMedSearch

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Discovery of diverse anellovirus sequences in Thai human sequencing data.

UNLABELLED: Anelloviruses are part of the normal human viral flora. Although their diversity in humans has been investigated in many countries, and despite their initial detection in Thailand in 1999, knowledge of Thai anelloviruses remains very limited. This study analyzed 1,175 whole-genome sequencing data sets from Thai individuals to mine for potential anellovirus sequences. Our analyses detected anellovirus sequences in 149 data sets (12.68%), uncovering 434 partial anellovirus sequences and 77 complete genome sequences, characterized by the presence of terminal redundancy, complete orf1, and the conserved untranslated region upstream of the orf1 gene. Sequence analyses indicated that these viruses belong to seven genera, including Alphatorquevirus, Betatorquevirus, Gammatorquevirus, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus. Notably, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus had not previously been reported in Thailand. Phylogenetic analysis of ORF1 protein sequences showed that Thai anelloviruses form multiple phylogenetic clusters with non-Thai anelloviruses, indicating frequent cross-country transmission and multiple origins of the virus in Thailand. Furthermore, sequence similarity network analysis identified 33 potentially novel anellovirus species in our data set. Our findings greatly expand the knowledge of anellovirus diversity in Thailand and demonstrate the potential of human whole-genome sequencing data as a valuable resource for viral discovery. Lastly, we highlight and discuss some challenges with the use of the current pairwise sequence similarity-based classification scheme, in particular, how gaps can influence similarity calculation and potentially lead to inconsistencies with a phylogenetic-based classification scheme. IMPORTANCE: Anelloviruses are widespread in humans, yet their diversity remains poorly characterized in many regions, including Thailand. Here, we demonstrate that human sequencing data sets, originally generated without the intention for virome research, can be effectively mined for anellovirus sequences, including complete genomes. Our findings reveal a substantial number of previously unreported anelloviruses in Thailand, significantly expanding the known diversity of the virus. We also highlight potential limitations of the current anellovirus species classification scheme, which is based on pairwise orf1 sequence similarity analysis with a hard threshold cutoff at 69%. Our results reveal that the current scheme can sometimes yield taxonomic groupings that are inconsistent with phylogenetic relationships, particularly when significant alignment gaps are present. Overall, our results show that existing human sequencing data can be effectively repurposed for virus discovery research and suggest the need for more robust and phylogenetically informed classification frameworks as viral sequence databases continue to expand.

Humans

A 9.6 kb intervening sequence in D. virilis rDNA, and sequence homology in rDNA interruptions of diverse species of Drosophila and other diptera.

A large proportion of the 28S ribosomal RNA genes in Drosophila virilis are interrupted by a DNA sequence 9.6 kilobase pairs long. As regards both its presence and its position in the 28S gene (about two thirds of the way in), the D. virilis rDNA intervening sequence is similar to that found in D. melanogaster rDNA, but lengths differ markedly between the two species. Degrees of nucleotide sequence homology have been detected bewteen rDNA interruptions of the two species. This homology extends to putative rDNA intervening sequences in diverse higher diptera (other Drosophila species, the house fly and the flesh fly), but hybridization of cloned D. melanogaster and D. virilis rDNA interruption segments to DNA of several lower diptera has been negative. As is the case with melanogaster rDNA interruptions, segments of the virilis rDNA intervening sequence hybridize with non-rDNA components of the virilis genome, and interspecific homology may involve these non-rDNA sequences as well as rDNA interruptions. There is, however, evidence from buoyant density fractionation of DNA that the distributions of interruption-related sequences are distinct in D. melanogaster and D. virilis genomes. Moreover, thermal denaturation studies have indicated differing extents of homology between hybridizable sequences in D. virilis DNA and different segments of the D. melanogaster rDNA intervening sequence. We infer from our studies that rDNA intervening sequences are prevalent among higher diptera; that in the course of the evolution of these organisms, elements of the intervening sequences have been moderately to highly conserved; and that this conservation extends in at least two distantly related species of Drosophila to similar sequences found elsewhere in the genomes.

Animals

Deoxyribonucleic acid sequence homologies among bacterial insertion sequence elements and genomes of various organisms.

Plasmid and phage deoxyribonucleic acid (DNA) harboring bacterial insertion sequence (IS) elements IS1, IS2, and IS5 were characterized and used as probes to detect homologous sequences in various procaryotic and eucaryotic genomes. The hybridization method used permits the detection of sequences partially homologous to the elements. Hybridization of the IS-containing probes to each other revealed a region of limited homology between IS1 and IS2. Homologous sequences were then detected by computer analysis of the published IS1 and IS2 nucleotide sequences. The homologous sequence contains a tandemly repeated tetranucleotide sequence which resembles the repeated sequence at the hot spot for spontaneous mutations in the lacI gene (P. J. Farabaugh, U. Schmeissner, M. Hofer, and J. Miller, J. Mol. Biol. 126:847-863, 1978). Homology between the IS elements and various genomes was determined by hybridizing labeled DNA containing IS1, IS2, and IS5 sequences to Southern blots of chromosomal DNA cleaved with restriction endonucleases. IS1 and IS5 appear limited to the enteric bacteria, whereas IS2 sequences can also be detected in Pseudomonas putida, Pseudomonas aeruginosa, and Serratia marcescens. Bacteria which appear not to possess extrachromosomal elements, e.g., Caulobacter crescentus, did not show homology with any insertion sequences tested. In addition, sequences homologous to IS1, IS2, or IS5 were not detected in Saccharomyces cerevisiae, Dictyostelium discoideum, or calf thymus DNA.

Bacteria

Classification and sequencing of hepatitis D virus from a large cohort of chronically infected individuals paired with co-infecting hepatitis B virus sequencing: a genomic characterisation study.

BACKGROUND: The most severe form of viral hepatitis is caused by co-infection of hepatitis D virus (HDV) and hepatitis B virus (HBV). Phylogenetic analyses classify HBV and HDV into eight major genotypes: HBV GTA to GTH and HDV GT1 to GT8. Paired HBV and HDV sequencing data from participants with chronic hepatitis delta are scarce. We aimed to sequence and genotype HDV and HBV from a large cohort of participants from clinical studies and diverse countries of origin. METHODS: 407 participants with chronic hepatitis D from 24 countries were characterised (124 participants from MYR301 clinical trial, 93 from MYR204, 114 from MYR202, and an additional 76 participants from diverse geographical locations). HBV and HDV from participants were analysed using sequencing, enzyme immunoassay, or both to determine HBV and HDV genotypes. BLAST analysis and phylogenetics were used to determine HBV and HDV genotypes with reference sequence libraries. Bulevirtide treatment response (measured by HDV RNA decline and normalisation of alanine aminotransferase) was compared by genotype for MYR trial participants. FINDINGS: HDV sequencing assays were successful for 386 (95%) of 407 participants and HBV sequencing or serology-based HBV genotyping assays were successful for genotyping 395 (97%) participants. For individual genotypes, HBV GTD (336 [83%] participants) and HDV GT1 (364 [89%]) were the most prevalent. For paired HBV-HDV genotypes, HBV-HDV D/1 was most common (320 [79%] of 407) followed by A/1 (30 [7%]). Phylogenetic analyses of HDV full-genome sequences showed distinct clusters of sequences within HDV GT1, and four novel provisional HDV GT1 subgenotypes, HDV GT1fp to HDVGT1ip, were identified. For 218 MYR clinical trial participants, bulevirtide treatment response was similar across HDV GT1 subgenotypes (both established and newly identified). INTERPRETATION: Novel HDV subgenotypes identified in this study indicate a greater genetic diversity of HDV GT1 than previously recognised. This knowledge will be important for developing better diagnostics, and in understanding HDV genotype-specific biology and response to treatment. More extensive HDV sequencing from under-sampled regions, such as Africa, is needed to determine the true breadth of HDV sequence and genotype diversity. FUNDING: Gilead Sciences.

Hepatitis Delta Virus

Characterizatiion of rat genetic sequences of Kirsten sarcoma virus: distinct class of endogenous rat type C viral sequences.

The nucleic acid sequences found in DNA and RNA from rat cells which are homologous to Kirsten sarcoma virus have been characterized. The homologous sequences are present in multiple copies per diploid rat cellular genome in a variety of different rat cellular dna's. In certain cells that constitutively express only low levels of sequences homologous to Kirsten sarcoma virus, bromodeoxyuridine treatment leads to the expression of high levels of these sequences in RNA. Supernatants from cell lines producing the sequences homologous to Kirsten sarcoma virus contain high levels of these sequences which are purified to the same degree as the previously known rat type C viral nucleic acid sequences by type C particles being released from such cells. The results indicate that the sequences in rat cells homologous to Kisten sarcoma virus have three characteristics of known mammalian type C viruses, and suggest that at least part of Kirsten sarcoma virus rat-derived sequences represent a distinct class of endogenous rat type C virus that has no detectable homology to the other known class of endogenous rat type C virus.

Animals

Complete sequence of constant and 3' noncoding regions of an immunoglobulin mRNA using the dideoxynucleotide method of RNA sequencing.

Three synthetic oligonucleotides were prepared to be complementary to known regions of the mouse immunoglublin light chain mRNA, and their ability to prime the transcription of complementary DNA (cDNA) was studied. The sequence of the cDNA was determined by adapting for mRNA the DNA sequencing method of Sanger, Nicklen and Coulson (1977) which uses 2'3' dideoxy ribonucleotides. A continuous sequence of 532 nucleotides was obtained, 321 corresponding to the whole of the constant region of the mRNA and the remaining 211 being the complete 3' noncoding region of the mRNA. The termination codon U-A-G occurs at the expected position in the mRNA corresponding to the triplet following the C terminal cystine. The nucleotide sequence is partially corroborated by the sequence of fragments obtained previously from 32P-mRNA fingerprints and endonuclease IV digests of 32P-cDNA, and is in agreement with the amino acid sequence of the constant region, except for a rearrangement of four amino acids (between amino acid positions 163 and 166). A revision of the amino acid sequence confirms the nucleic acid sequence.

Amino Acid Sequence

Reference Sequence Browser: An R application with a user-friendly GUI to rapidly query sequence databases.

Land managers, researchers, and regulators increasingly utilize environmental DNA (eDNA) techniques to monitor species richness, presence, and absence. In order to properly develop a biological assay for eDNA metabarcoding or quantitative PCR, scientists must be able to find not only reference sequences (previously identified sequences in a genomics database) that match their target taxa but also reference sequences that match non-target taxa. Determining which taxa have publicly available sequences in a time-efficient and accurate manner currently requires computational skills to search, manipulate, and parse multiple unconnected DNA sequence databases. Our team iteratively designed a Graphic User Interface (GUI) Shiny application called the Reference Sequence Browser (RSB) that provides users efficient and intuitive access to multiple genetic databases regardless of computer programming expertise. The application returns the number of publicly accessible barcode markers per organism in the NCBI Nucleotide, BOLD, or CALeDNA CRUX Metabarcoding Reference Databases. Depending on the database, we offer various search filters such as min and max sequence length or country of origin. Users can then download the FASTA/GenBank files from the RSB web tool, view statistics about the data, and explore results to determine details about the availability or absence of reference sequences.

User-Computer Interface

Allele Level Sequencing of Killer Cell Immunoglobulin-Like Receptor Genes Using Oxford Nanopore Long Read Sequencing.

The human Killer cell Immunoglobulin-like Receptor (KIR) genes, found on chromosome 19, encode for cell surface protein receptors that, through interaction with their ligand, modulate the action of Natural Killer (NK) cells and some subsets of T lymphocytes. KIR genes exhibit extensive variation through variable gene content, copy number, and allele polymorphism. The combination of KIR genes and their ligands is implicated in various clinical settings including haematopoietic stem cell and solid organ transplant, and infectious disease progression. KIR gene content has been used in the selection of optimal stem cell donors with haplotype variations in recipient and donor giving differential clinical outcomes. With the introduction of massively parallel clonal next generation sequencing and single molecule long read third generation sequencing, allele level determination of KIR genotypes has become feasible. We describe a method for amplicon-based long read sequencing on the Oxford Nanopore Technologies platform that provides largely unambiguous allele level typing of KIR genes. The method was validated using DNA extracted from 48 10th International Histocompatibility Workshop (IHWS) cell lines with previously published allele level KIR genotypes and 176 Western Australian samples previously tested for the presence or absence of KIR genes. Our long-read sequencing method was able to accurately determine KIR alleles with an overall concordance of 97%-99% with the published data. Importantly, phasing ambiguity caused by the inability to phase heterozygous base positions over long stretches of gene sequence was resolved in several samples. Thus, our long read PCR sequencing strategy can be used to determine KIR genotypes at allele resolution level.

Humans

[Sequence complexity of transcribed unique DNA sequences in genome of mouse P815 mastocytoma cells (author's transl)].

The sequence complexity of nuclear RNA from mouse liver, mouse spleen and highly malignant P815 mastocytoma was measured by nRNA driven hybridization to unique DNA sequences of P815 cells. The unique DNA sequences represent 63% of the total nuclear DNA of P815 cells and their availibility in hybridization experiments was found to be 76%. Of these sequences 7.8% formed hybrids with nuclear RNA of this cell, about 11.5% with mouse spleen and about 14.5% with mouse liver nuclear RNA. Assuming an asymmetrical transcription, the complexities of these transcripts are 2.8 X 10(8) nucleotides for mouse P815 mastocytomas, 4.3 X 10(8) for mouse spleen and about 5.3 X 10(8) nucleotides for mouse liver. Cellular specifity of the transcribed information was analyzed in additivity experiments, in which unique DNA sequences, not complementary to the nuclear RNA of one cell were annealed to the nuclear RNAs of the two other tissues/cells. In these experiments most of the nuclear RNA sequences of P815 cells were found to be also present in the nucleus of mouse liver and spleen. Only a small portion of the unique DNA sequences of P815 mastocytoma (about 1.2% corresponding to 4.4 X 10(7) nucleotides) was found to be complementary only to P815 mastocytoma nuclear RNA.

Animals

Dihydrofolate reductase from amethopterin-resistant Lactobacillus casei. Sequences of the cyanogen bromide peptides and complete sequences of the enzyme.

The complete amino acid sequence of dihydrofolate reductase from an amethopterin-resistant strain of Lactobacillus casei has been determined by sequence analysis of peptides produced by cleavage with cyanogen bromide, trypsin, staphylococcal protease, and myxobacter protease. Comparison of this sequence with those of reductases from other bacterial sources shows that the enzymes are homologous. The Lactobacillus casei reductase sequences shows a 29% sequence identity with that of the Escherichia coli enzyme and a 34% identity with the sequence of the enzyme from Streptococcus faecium. The NH2-terminal 68 residues of the L. casei reductase show a 54% sequence identity with that of the enzyme from S. faecium.

Amino Acid Sequence

GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequences.

MOTIVATION: Short tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. RESULTS: Here, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653.

Microsatellite Repeats

Nucleotide sequence at the junction between the coding region of the adenovirus 2 hexon messenger RNA and its leader sequence.

We have determined a 139-base-pair sequence of adenovirus 2 DNA that is located immediately leftwards of the cleavage site for endonuclease Sma I at position 51.1. The established sequence includes the hexon AUG initiator codon, located 75--77 nucleotides leftwards of this cleavage site, and codons for the first 26 amino acids of the hexon polypeptide. By the use of purified hexon mRNA as a template and separated strands of small restriction enzyme fragments as specific primers, the complete 5' noncoding region of the hexon mRNA was synthesized and part of its sequence was determined. The tripartite leader sequence of the hexon mRNA starts 39 nucleotides upstream from the initiator AUG triplet and the total length of the 5' noncoding part of the hexon mRNA was estimated to be 235 nucleotides. The sequence at the junction of the leader sequence permits the formation of secondary structures that may be of importance for the splicing reaction.

Adenoviruses, Human

Immunological comparison of azurins of known amino acid sequence. Dependence of cross-reactivity upon sequence resemblance.

To examine further the dependence of immunological cross-reactivity on sequence resemblance among proteins, we carried out micro-complement fixation studies with rabbit antisera to bacterial azurins of known amino acid sequence. There is a strong correlation (r = 0.9) between number of amino acid substitutions and degree of antigenic difference (immunological distance) among these azurins. The antigenic effects of amino acid substitutions are thus approximately equal and approximately additive. Similar observations and inferences were made before with a series of bird lysozymes. Indeed, the same approximate relationship between immunological distance (y) and percent difference in amino acid sequence (x) holds for both azurins and lysozymes, namely y congruent to 5x. An explanation is given for the dependence of immunological cross-reactivity on sequence resemblance among proteins. This entails reviewing evidence regarding the nature and number of antigenic sites on globular protein antigens as well as evidence for the existence of evolutionary biases against substitutions that are internal or cause large conformational changes. The explanation we give may apply only to those naturally occurring, globular, monomeric, isofunctional proteins whose sequences differ substantially from that of any rabbit protein.

Amino Acid Sequence

A microcosting and cost consequence analysis from a randomized controlled trial comparing genome sequencing with exome sequencing for genetic diagnosis.

PURPOSE: Diagnosing rare diseases is costly. The objectives were to microcost exome (ES) and genome sequencing (GS) trios and estimate the incremental costs of GS per additional diagnosis from an institutional payer perspective. METHODS: Trios (proband plus biological parents) that are referred for sequencing were randomly assigned to ES or GS. Laboratory workflow and sequencing were microcosted. Total and category cost per trio were estimated probabilistically. Effectiveness was expressed as diagnostic yield (rates of diagnostic or partially diagnostic variants detected). Incremental costs and effectiveness were calculated. RESULTS: The mean total cost per trio was CAD 2888.79 (95% CI 2567.72, 3492.72) for ES (n = 329) and 4364.02 (95% CI 3984.94, 5013.67) for GS (n = 324). Reagents accounted for 34% and 61% of total costs for ES and GS, respectively. The incremental cost of GS was 1475.23. The diagnostic yield was 35.9% for ES and 32.7% for GS with a difference of 0.032 (95% CI: -0.041, 0.104, P value .397). CONCLUSION: GS demonstrated higher costs and a similar diagnostic yield to ES but was limited by technical capabilities at the time of the study. The study provides comprehensive costs for the economic evaluation comparing alternative diagnostic pathways and impetus for further evaluating variants uniquely detectable by GS.

Humans

MCALIGN: stochastic alignment of noncoding DNA sequences based on an evolutionary model of sequence evolution.

A method is described for performing global alignment of noncoding DNA sequences based on an evolutionary model parameterized by the frequency distribution of lengths of insertion/deletion events (indels) and their rate relative to nucleotide substitutions. A stochastic hill-climbing algorithm is used to search for the most probable alignment between a pair of sequences or three sequences of known phylogenetic relationship. The performance of the procedure, parameterized according to the empirical distribution of indel lengths in noncoding DNA of Drosophila species, is investigated by simulation. We show that there is excellent agreement between true and estimated alignments over a wide range of sequence divergences, and that the method outperforms other available alignment methods.

Algorithms

DNA sequence analysis. Terminal sequences of bacteriophage phi80.

Sequences of the cohesive ends and the 3'-terminal regions of phi80 DNA have been determined. Sequences of the cohesive ends were obtained through the use of two standard methods. The first method involved the incorporation of all four labeled deoxyribonucleotides into the phi80 cohesive ends using DNA polymerase I. The DNA was then partially digested with micrococcal nuclease or pancreatic DNase. The products were separated by two-dimensional electrophoresis and characterized by composition, 3'-terminal, and nearest neighbor analyses. The second method involved partial incorporation using one, two, or three labeled deoxyribonucleotides followed by similar analyses. Sequences of the double-stranded regions adjacent to the cohesive ends were determined by three new methods. These methods were: (a) the DNA was specifically labeled at the 3' terminus and then partially degraded. Labeled oligonucleotide products were sequenced by their mobilities on various separation systems. (b) The cohesive ends were enlarged by limited degradation with exonuclease III. After this treatment, the DNA was partially repaired with labeled nucleotides, digested, and the products were analyzed. (c) A synthetic ologonucleotide primer was bound to phi80 DNA which had been repaired with DNA polymerase I, and then partially digested with lambda-exonuclease. The primer was extended into the region of interest by partial repair with labeled nucleotides. The extended primer was isolated and analyzed.

Base Sequence

RNA Sequencing Protocols for Short-Read Sequencing.

RNA sequencing (RNA-seq) methodologies allow the discovery of novel variants and transcripts. These comprise three general steps: (1) capture of RNA species of interest, (2) conversion of RNA to complementary DNA (cDNA), and (3) modification of cDNA to fit the sequencing platform. Here we describe four different library preparation protocols for short-read sequencing: cDNA synthesis with poly(A) selection, library preparation with ribosomal depletion, and cDNA synthesis with SMART® (Switching Mechanism at 5' end of RNA Template) technology for low and Pico inputs.

Gene Library

Identification and full genome sequencing of previously unknown sandfly-borne phleboviruses using a newly established capture-based next-generation sequencing approach.

Sandfly-borne phleboviruses cause febrile illness and neuroinvasive disease in humans. While infections are reported in the Mediterranean region, the discovery of previously unknown phleboviruses in sandflies from Kenya suggests a wider geographic distribution. Detection and characterization of novel phleboviruses are often hindered by low-quality and low-viral-load samples. We developed a capture-based target enrichment next-generation sequencing approach that showed a 99%-100% fold enrichment of viral genomes from primary material and provides a robust tool for generating complete genomes of both known and previously unknown viruses. From a collection of 15,652 sandflies in Kenya, we recovered seven complete coding sequences of Embossos, Bogoria, and Kiborgoch viruses, and of two previously unknown phleboviruses, which were named Sosoik and Shable viruses. Sosoik virus shared 83% amino acid identity in its RdRp gene with that of Bogoria virus, while Shable virus shared ca. 88% amino acid identity with viruses of the Salehabad serocomplex. Additionally, a reassortant of Shable virus was detected that possessed an M segment from an undescribed Ponticelli-like virus. DNA barcoding of blood-fed sandflies revealed several potentially novel Sergentomyia species and evidence of host-feeding on humans, livestock, and reptiles, suggesting possibilities for zoonotic transmission. Overall, our findings increase the known genetic diversity of Old World sandfly-borne phlebovirus species from 18 to 25 (by 38.9%), including the detection of viruses from all pathogenic sandfly-borne phlebovirus serocomplexes in East Africa, opening new horizons in disease ecology research.IMPORTANCEKnowledge of the genetic diversity of circulating pathogens is crucial for providing appropriate diagnostics and disease management. This study established a novel capture-based target enrichment next-generation sequencing approach that enabled the near-complete viral genome recovery from primary samples, while native NGS yielded negative or poor-quality results. In addition to the five recently discovered sandfly-borne phleboviruses in Kenya, two previously unknown phleboviruses were detected in sandflies from the same region. The viruses were detected in several sandfly species, which showed diverse host-feeding behaviors, including mixed feeding on humans and chickens. The study significantly advances the understanding of sandfly-borne phleboviruses by uncovering their broader geographic distribution and genetic diversity, particularly in East Africa, highlighting the importance of expanding surveillance efforts beyond traditionally studied regions.

Phlebovirus