Search PubMedSearch

SEARCH · Search PubMed

Results for “sequence diversity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Discovery of diverse anellovirus sequences in Thai human sequencing data.

UNLABELLED: Anelloviruses are part of the normal human viral flora. Although their diversity in humans has been investigated in many countries, and despite their initial detection in Thailand in 1999, knowledge of Thai anelloviruses remains very limited. This study analyzed 1,175 whole-genome sequencing data sets from Thai individuals to mine for potential anellovirus sequences. Our analyses detected anellovirus sequences in 149 data sets (12.68%), uncovering 434 partial anellovirus sequences and 77 complete genome sequences, characterized by the presence of terminal redundancy, complete orf1, and the conserved untranslated region upstream of the orf1 gene. Sequence analyses indicated that these viruses belong to seven genera, including Alphatorquevirus, Betatorquevirus, Gammatorquevirus, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus. Notably, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus had not previously been reported in Thailand. Phylogenetic analysis of ORF1 protein sequences showed that Thai anelloviruses form multiple phylogenetic clusters with non-Thai anelloviruses, indicating frequent cross-country transmission and multiple origins of the virus in Thailand. Furthermore, sequence similarity network analysis identified 33 potentially novel anellovirus species in our data set. Our findings greatly expand the knowledge of anellovirus diversity in Thailand and demonstrate the potential of human whole-genome sequencing data as a valuable resource for viral discovery. Lastly, we highlight and discuss some challenges with the use of the current pairwise sequence similarity-based classification scheme, in particular, how gaps can influence similarity calculation and potentially lead to inconsistencies with a phylogenetic-based classification scheme. IMPORTANCE: Anelloviruses are widespread in humans, yet their diversity remains poorly characterized in many regions, including Thailand. Here, we demonstrate that human sequencing data sets, originally generated without the intention for virome research, can be effectively mined for anellovirus sequences, including complete genomes. Our findings reveal a substantial number of previously unreported anelloviruses in Thailand, significantly expanding the known diversity of the virus. We also highlight potential limitations of the current anellovirus species classification scheme, which is based on pairwise orf1 sequence similarity analysis with a hard threshold cutoff at 69%. Our results reveal that the current scheme can sometimes yield taxonomic groupings that are inconsistent with phylogenetic relationships, particularly when significant alignment gaps are present. Overall, our results show that existing human sequencing data can be effectively repurposed for virus discovery research and suggest the need for more robust and phylogenetically informed classification frameworks as viral sequence databases continue to expand.

Humans

Genomic sequencing in diverse and underserved pediatric populations: Parent perspectives on understanding, uncertainty, psychosocial impact, and personal utility of results.

PURPOSE: Limited evidence evaluates parents' perceptions of their child's clinical genome-scale sequencing (GS) results, particularly among individuals from medically underserved groups. Five Clinical Sequencing Evidence-Generating Research consortium studies performed GS in children with suspected genetic conditions with high proportions of individuals from underserved groups to address this evidence gap. METHODS: Parents completed surveys of perceived understanding, personal utility, and test-related distress after GS result disclosure. We assessed outcomes' associations with child- and parent-related factors: child age; type of GS finding; and parent health literacy, numeracy, and education. RESULTS: A total of 1763 parents completed surveys; 83% met "underserved" criteria based on race, ethnicity, and risk factors for barriers to access. We observed high perceived understanding and personal utility and low test-related distress. Outcomes were associated with the type of GS finding; parents of children with a pathogenic or likely pathogenic finding endorsed higher personal utility and more test-related distress than those whose children had a variant of uncertain significance or normal finding. Personal utility was higher in parents who met the criteria for "underserved." CONCLUSION: Our findings shed light on correlates of parents' cognitive and emotional responses to their child's GS findings and emphasize the need for tailored support in disclosure discussions.

Humans

Effect of inhaled interferon-β1a on SARS-CoV-2 diversity and evolution.

Interferon resistance has been implicated in SARS-CoV-2 escape from innate immunity, but exogenous interferon's impact on viral evolution and diversity is unknown. SNG001, an inhaled interferon-β1a treatment, was evaluated in the ACTIV-2/A5401 randomized controlled trial of therapeutics for COVID-19. We measured viral kinetics and performed whole-genome sequencing on longitudinal nasal swabs collected from ACTIV-2 participants who received either SNG001 or placebo to assess viral sequence diversity. No difference in nasal viral load decay was detected between study arms when stratifying by SARS-CoV-2 variant or by viral culture conversion. Compared to placebo participants, the SNG001-treated participants displayed significantly lower nonsynonymous amino acid average pairwise distance, indicating lower sequence diversity. Similarly, SNG001-treated individuals also developed numerically fewer nonsynonymous mutations during their infection in ORF1a, ORF1b, Spike, and Nucleocapsid. No specific emerging SARS-CoV-2 nonsynonymous amino acid changes indicating signatures of viral escape were enriched in those receiving SNG001. These in vivo data provide an intriguing signal that exogenous interferon-β1a may restrict SARS-CoV-2 viral diversity and add to growing evidence that interferon levels play a critical role in antiviral responses during COVID-19.IMPORTANCESARS-CoV-2 encodes several genes which can antagonize the interferon signaling cascade, preventing it from activating antiviral responses and thereby facilitating viral establishment and dissemination. It is unknown how the administration of exogenous interferon might affect viral evolution and immune escape. ACTIV-2/A5401 represents a unique opportunity to study the virologic effects of interferon treatment in a rigorous randomized, placebo-controlled clinical trial setting. Our characterization of longitudinal nasal samples shows that interferon-treated individuals had lower viral diversity and no evidence of viral escape mutations.CLINICAL TRIALSThis study is registered with ClinicalTrials.gov as NCT04518410.

Humans

Comparative Genomics of Sex-Determination-Related Genes Reveals Shared Evolutionary Patterns Between Bivalves and Mammals, but Not Fruit Flies.

The molecular basis of sex determination (SD), while being extensively studied in model organisms, remains poorly understood in many animal groups. Bivalves, a diverse class of molluscs with a variety of reproductive modes, represent an ideal yet challenging clade for investigating SD and the evolution of sexual systems. However, the absence of a comprehensive framework has limited progress in this field, particularly regarding the study of sex-determination-related genes (SRGs). In this study, we performed a genome-wide sequence evolutionary analysis of the Dmrt, Sox and Fox gene families in more than 40 bivalve species. For the first time, we provide an extensive and phylogenetically aware dataset of these SRGs, and we find support for the hypothesis that Dmrt-1L and Sox-H may act as primary sex-determining genes by showing their high levels of sequence diversity within the bivalve genomic context. To validate our findings, we studied the same gene families in two well-characterised systems, mammals and fruit flies (genus Drosophila). In the former, we found that the male sex-determining gene Sry exhibits a pattern of amino acid sequence diversity similar to that of Dmrt-1L and Sox-H in bivalves, consistent with its role as master SD regulator. In contrast, no such pattern was observed among genes of the fruit fly SD cascade, which is controlled by a chromosomic mechanism. Overall, our findings highlight similarities in the sequence evolution of some mammal and bivalve SRGs, possibly driven by a comparable architecture of SD cascades. This work underscores once again the importance of employing a comparative approach when investigating understudied and non-model systems.

Animals

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals

Fast, accurate construction of multiple sequence alignments from protein language embeddings.

Multiple sequence alignment (MSA) is a foundational task in computational biology, underpinning protein structure prediction, evolutionary analysis, and domain annotation. Traditional MSA algorithms rely on pairwise amino acid substitution matrices derived from conserved protein families. While effective for aligning closely related sequences, these scoring schemes struggle in the low-identity "twilight zone." Here, we present a new approach for constructing MSAs leveraging amino acid embeddings generated by protein language models (PLMs), which capture rich evolutionary and contextual information from massive and diverse sequence datasets. We introduce a windowed reciprocal-weighted embedding similarity metric that is surprisingly effective in identifying corresponding amino acids across sequences. Building on this metric, we develop ARIES (Alignment via RecIprocal Embedding Similarity), an algorithm that constructs a PLM-generated template embedding and aligns each sequence to this template via dynamic time warping in order to build a global MSA. Across diverse benchmark datasets, ARIES achieves higher accuracies than existing state-of-the-art approaches, especially in low-identity regimes where traditional methods degrade, while scaling almost linearly with the number of sequences to be aligned. Together, these results provide the first large-scale demonstration of the power of PLMs for accurate and scalable MSA construction across protein families of varying sizes and levels of similarity, highlighting the potential of PLMs to transform comparative sequence analysis.

Deep Learning

Genetic heterogeneity and pathogenic potential of historical Crimean-Congo hemorrhagic fever virus isolates in China.

The Crimean-Congo hemorrhagic fever virus (CCHFV) poses a significant public health threat. In China, CCHFV has been circulating for decades, yet the genomic diversity and pathogenic potential of the circulating strains remain poorly characterized, hindering risk assessment and countermeasure development. In this study, we recovered 24 historical CCHFV strains isolated between 1966 and 2004 from humans, ticks and jerboas in Xinjiang Uyghur Autonomous Region of China. Whole-genome sequencing was performed, followed by comprehensive analyses of their phylogenetic relationships, in vitro infectivity and in vivo pathogenicity. Phylogenetic analyses revealed high genetic heterogeneity, identifying seven genotypes for the L segment, nine for the M segment (including a novel Asia 4 genotype), and nine for the S segment. Amino acid mutation analysis revealed that the mucin-like domain (MLD) of the glycoprotein (GP) exhibited the highest mutation rate, contributing substantially to sequence diversity. In vitro, Asia 2 (75024) and Asia 3 (79121M18) strains exhibited robust replication in monkey-, hamster-, and human-derived cell lines. In C57BL/6 mice, all four representative strains induced viral replication and specific antibody responses (IgM and IgG), causing mild to moderate pathological damage in the liver, spleen, and kidneys. In IFNAR-/- mice, virulence varied markedly among representative strains: Asia 2 and Asia 3 strains were highly lethal (LD50 < 1 TCID50), Asia 1 was moderately virulent (LD50 = 142.5 TCID50), and Asia 4 exhibited atypical, non-dose-dependent mortality. Collectively, our work reports a novel Asia 4 genotype and suggests strain- and lineage-associated differences in virulence for CCHFV in China, providing critical insights for surveillance and targeted countermeasure development.

Animals

Genome mining reveals an architecturally expanded pyoluteorin-associated biosynthetic gene cluster and a divergent flavin-dependent halogenase-like sequence in deep-sea Pseudomonas Aeruginosa from the Gulf of Guinea.

BACKGROUND: Marine deep-sea environments harbour microorganisms with extraordinary biosynthetic potential, yet their secondary metabolite repertoires remain largely uncharacterised. RESULTS: This study reports the isolation, phenotypic characterisation, and whole-genome analysis of Pseudomonas aeruginosa strain E1, recovered from deep Atlantic seawater (Gulf of Guinea, ~2500&#xa0;m depth), which exhibits antifungal activity against multidrug-resistant Candida parapsilosis. Three presumptive P. aeruginosa isolates (E1, E17, and E44) showed >&#x2009;99% 16S rRNA gene sequence identity to P. aeruginosa reference sequences, while whole-genome dDDH analysis of strain E1 yielded 95.2% (95% CI: 93.6-96.4%; formula d4) relative to the P. aeruginosa type strain DSM 50071&#x1d40; (=&#x2009;ATCC 10145&#x1d40;), supporting its species-level assignment. Antifungal screening and PCR-based detection of flavin-dependent halogenase genes identified strain E1 as the primary candidate for genomic investigation. Illumina whole-genome sequencing produced a 6.33&#xa0;Mb draft genome assembly (113 contigs, 5862 protein-coding genes, 66.4% GC content). Genome mining with antiSMASH 8.0 identified 27 biosynthetic gene clusters (BGCs) spanning nonribosomal peptide synthetase (NRPS), polyketide synthase (PKS), phenazine, terpene, and metallophore pathways. Region 7.1 of strain E1 harbours a predicted 50.8&#xa0;kb pyoluteorin-associated BGC, comprising 34 genes, substantially larger than its terrestrial counterpart (~&#x2009;22&#xa0;kb, ~&#x2009;17 genes), and featuring nine transport genes and three regulatory elements. Phylogenetic analysis resolved three halogenase genes: ctg7_146 showed 98.7% amino acid identity to PltA, and ctg7_149 showed 99.2% amino acid identity to PltM, supporting their annotation as PltA-like and PltM-like components of the predicted pyoluteorin biosynthetic pathway. Among the characterised reference enzymes included in this analysis, ctg7_143 showed the highest amino acid identity to PltM from P. fluorescens Pf-5. However, the identity remained low at approximately 30.4%, supporting its placement as a divergent FDH-like sequence rather than a close PltM orthologue. CONCLUSION: This study provides the first comprehensive genomic characterisation of a pyoluteorin-BGC-harbouring marine P. aeruginosa strain, demonstrating conservation of the core biosynthetic machinery alongside an expanded transport architecture and a divergent FDH-like sequence that may represent a candidate for future biochemical investigation. These findings expand current knowledge of FDH-like sequence diversity in deep-sea bacteria and support further investigation of Gulf of Guinea microorganisms as a potential source of biosynthetic and enzymatic diversity.

Multigene Family

Genomic characterization of KPC-2 and NDM coproducing carbapenem-resistant Klebsiella pneumoniae in a hospital: discovery of ST1869 clone and a novel hybrid plasmid.

UNLABELLED: To characterize the plasmid architecture and molecular background of KPC-NDM coproducing carbapenem-resistant Klebsiella pneumoniae (KN-CRKP) in a South China hospital. Five KN-CRKP isolates were collected, including three from one patient. All underwent Illumina sequencing; two (ST11 and ST1869) additionally had Nanopore sequencing. Antimicrobial susceptibility testing strain sequence types, conjugation assays, resistance gene profiling, plasmid typing, genetic structure comparison, core-genome single nucleotide polymorphisms (SNPs) analysis, and plasmid clustering were performed. All isolates exhibited an imipenem minimum inhibitory concentration (MIC) of &#x2265;128 &#xb5;g/mL and harbored multiple resistance genes. One isolate (1/5) belonged to ST1869 and co-harbored blaKPC-2 and blaNDM-5. The blaNDM-5-carrying plasmid was a novel IncI1/X3 fusion plasmid that also carried blaCMY-42. Unlike several IncX3 plasmids carrying blaNDM in publicly available KN-CRKP genomes from South China, this IncI1/X3 hybrid lacked a complete conjugative transfer system. ST11 was the predominant clone (4/5), co-harboring blaKPC-2 and blaNDM-1. A rare genetic structure, &#x394;ISKpn6-blaKPC-2-ISKpn28, was identified on IncFII plasmids carrying blaKPC-2. Plasmid clustering analysis of 126 comparative KN-CRKP genomes showed diverse sequence types and plasmid backgrounds associated with the KPC/NDM co-production pattern. The observed plasmid diversity and structural variation in KN-CRKP support continued genomic surveillance, with particular attention to the ST1869 clone, the novel IncI1/X3 hybrid plasmid harboring blaNDM-5 and blaCMY-42, and the rare "&#x394;ISKpn6-blaKPC-2-ISKpn28" genetic structure. Expanded genomic data on KN-CRKP are needed to further elucidate its resistance mechanisms and plasmid evolutionary trajectories. IMPORTANCE: The co-production of KPC and NDM carbapenemases in Klebsiella pneumoniae poses a formidable threat to clinical antimicrobial therapy, as these enzymes confer resistance to virtually all &#x3b2;-lactam agents, including carbapenems. Here, we report novel genomic features of KN-CRKP in South China, including the emergence of the ST1869 clone, a unique IncI1/X3 hybrid plasmid harboring blaNDM-5 and blaCMY-42, and the rare &#x394;ISKpn6-blaKPC-2-ISKpn28 genetic structure. These findings substantially expand current understanding of plasmid evolution and resistance gene dissemination in this region. The identification of diverse resistance mechanisms and clonal backgrounds supports enhanced genomic surveillance and infection-control awareness for pan-resistant Enterobacterales.

Plasmids

UMI-nea: a fast, robust tool for reference-free UMI deduplication and accurate quantification.

MOTIVATION: One of the key applications of Unique Molecular Identifiers (UMIs) in high-throughput sequencing is to correct for PCR amplification bias and removal of PCR duplicates, thereby improving quantification in DNA-seq and RNA-seq applications. Accurately grouping error-bearing UMIs that originate from the same input molecule through a UMI deduplication method is a critical step in this process. However, many existing UMI deduplication tools rely on simple Hamming distance comparisons or suboptimal clustering algorithms, often resulting in erroneous UMI groupings, particularly in error-prone long-read sequencing or ultra-high-depth short-read sequencing. RESULTS: We introduce UMI-nea, a tool that utilizes Levenshtein distance comparisons and a novel clustering approach to optimize multithreading workflows. Compared against three other indel-aware UMI deduplication tools, UMI-nea achieves more accurate UMI groupings with efficient run time. It demonstrates robust performance across diverse sequencing platforms, depths, and UMI lengths. Additionally, UMI-nea incorporates a data-guided adaptive UMI filter, further enhancing quantification accuracy. AVAILABILITY AND IMPLEMENTATION: UMI-nea is available on github https://github.com/Qiaseq-research/UMI-nea.git or Zenodo https://doi.org/10.5281/zenodo.16745758. Sequencing data are stored at https://qiagenpublic.blob.core.windows.net/umi-nea-datasets/.

High-Throughput Nucleotide Sequencing

Structure-informed theoretical modeling defines principles governing avidity in bivalent protein interactions.

In signaling cascades, signaling proteins often encode multiple domains or motifs, which presents the possibility for avidity -- where multivalent binding drastically increases interaction strength and duration. However, predicting and validating multivalent interactions that interact with avidity is a challenge. Here, we integrate mechanistic modeling, structure-based analysis, and experimental approaches as a framework for defining the conditions under which avidity plays a role. We explore the tandem SH2 domain family of interactions with bisphosphorylated partners as a multivalent archetype, which encompasses key secondary messengers in tyrosine kinase signaling networks. Theoretical modeling suggests that maximum avidity occurs with closely spaced tyrosine phosphorylation sites combined with moderate monovalent affinities - exactly around the innate range of SH2 domain affinity - or with phosphorylation sites separated by sufficiently flexible linkers. Surprisingly, despite sequence diversity, structure-based analysis showed relatively conserved three-dimensional spacing between SH2 domains across all tandem SH2 families, which we corroborate experimentally, suggesting evolutionary optimization for avidity interactions. The combination of structure-based analysis of domain spacing with available monovalent experimental data appears, along with iterative experimental refinement of biophysical parameters, can identify high affinity interactions of tandem SH2 domain recruitment to the EGFR C-terminal tail. Using these principles, we extended bivalent predictions into the full phosphoproteome space and structural parameterization of other partners of SH2 domain binding, providing resources and methods for more rapid expansion of bivalent analysis. These approaches lay the groundwork for larger utility in multivalent prediction and testing to help better understand protein interactions that drive cell signaling.

BLI

Distinct YY dinucleotide periodicity in adeno-associated virus DNA.

Dinucleotide periodicity is a hallmark of genome organization, yet its role in single-stranded (ss)DNA viruses remains poorly understood. Here, we systematically analyzed dinucleotide spacing patterns in adeno-associated virus (AAV) genomes and other viruses. Across 13 primate AAV serotypes, we identified a pronounced and highly conserved &#x223c;15-bp periodicity specific to pyrimidine-pyrimidine (YY) dinucleotides and their reverse complements (RR). Comparative analyses across >25,000 viral sequences demonstrate that this 15-bp YY/RR periodicity is unique to the genus Dependoparvovirus and absent from other ssDNA viruses, satellite viruses, and helper viruses, which predominantly exhibit canonical &#x223c;10- to 11-bp periodicities. Upon disruption of the YY/RR pattern using DNA family shuffling of AAV capsid genes, and subsequent iterative selection for viral production or cell entry, we found that the pattern is under positive selection. Selected sequences display increased periodicity alongside reduced sequence diversity, supporting a functional role for this genomic feature. Finally, engineered recombinant AAV genomes containing YY periodic motifs exhibit enhanced production and, for some designs, improved transduction efficiency, demonstrating that YY periodicity can modulate viral replication and infectivity. Our findings uncover a unique DNA-encoded signal in dependoparvoviruses that contributes to AAV fitness, expands our knowledge of virus biology, and has implications for vector engineering.

Dependovirus

Integrated computational and experimental benchmarking of Bacillus phage endolysins reveals the relationship between peptidoglycan-fragment recognition descriptors and antibacterial performance.

Protein-based antibacterials such as bacteriophage endolysins offer a targeted therapeutic strategy against Gram-positive pathogens. However, prioritizing the most effective candidates from the large sequence diversity available remains a significant challenge. Here we present a standardized computational-experimental benchmarking framework that evaluates seven phage-derived endolysin variants (E1, E2, E3, E7, E10, E12, and E15) identified from Bacillus genomes. We combined molecular docking and residue-level interaction mapping against muramyl dipeptide (MDP), a minimal conserved peptidoglycan motif, with 1000-ns molecular dynamics simulations, MM/PBSA binding free-energy estimation, and matched functional inhibition assays against Staphylococcus aureus and Micrococcus luteus. Computational analyses revealed generally favorable MDP recognition across variants, albeit with notable differences in contact patterns and complex stability profiles. Experimental screening identified E2 as the most potent antibacterial agent against both species, while E7 and E1 performed strongly in selected computational metrics. Integrated analysis showed only modest correlations between computational descriptors of fragment recognition/stability and observed antibacterial performance. This study establishes a practical comparative benchmarking platform for endolysin candidate prioritization, nominates E2 and E7 as promising candidates for further development, and highlights E1 as a potential structural scaffold for rational engineering, while explicitly demonstrating both the utility and the current limitations of using minimal peptidoglycan fragments as proxies for full cell-wall recognition in lysin benchmarking.

Endopeptidases

Structure-informed theoretical modeling defines principles governing avidity in bivalent protein interactions.

In signaling cascades, where domain-motif interactions tend to interact with relatively low affinity (allowing for reversibility), signaling proteins often encode multiple domains or motifs, which present the possibility of avidity - drastically increasing the interaction strength and duration as a result of multivalent binding. However, given the large combinatorial space, predicting and validating multivalent interactions that interact with avidity is a challenge. Here, we integrate mechanistic modeling, structure-based analysis, and experimental approaches as a framework for defining the conditions under which avidity plays a role. We explore the tandem SH2 domain family of interactions with bisphosphorylated partners as a multivalent archetype, which encompasses key secondary messengers in tyrosine kinase signaling networks. While certain multivalent interactions have been shown to be necessary in immune receptor recruitment of partners, bivalent recruitment of tandem SH2 domains more broadly is poorly understood. Theoretical modeling suggests that maximum avidity occurs with closely spaced or flexibly linked phosphotyrosine sites, combined with moderate monovalent affinities - exactly around the innate range of SH2 domain affinity. Surprisingly, despite sequence diversity, structure-based analysis showed remarkably conserved three-dimensional spacing between SH2 domains across all tandem SH2 families, which we corroborate experimentally, suggesting evolutionary optimization for avidity interactions. The combination of structure-based analysis of domain spacing with available monovalent experimental data appears to be sufficiently accurate to predict and rank order high affinity interactions of tandem SH2 domain recruitment to the EGFR C-terminal tail. These approaches lay the groundwork for larger utility in multivalent prediction and testing to help better understand protein interactions that drive cell signaling.

BLI

Structural insights into adeno-associated virus serotype 5.

The adeno-associated viruses (AAVs) display differential cell binding, transduction, and antigenic characteristics specified by their capsid viral protein (VP) composition. Toward structure-function annotation, the crystal structure of AAV5, one of the most sequence diverse AAV serotypes, was determined to 3.45-&#xc5; resolution. The AAV5 VP and capsid conserve topological features previously described for other AAVs but uniquely differ in the surface-exposed HI loop between &#x3b2;H and &#x3b2;I of the core &#x3b2;-barrel motif and have pronounced conformational differences in two of the AAV surface variable regions (VRs), VR-IV and VR-VII. The HI loop is structurally conserved in other AAVs despite amino acid differences but is smaller in AAV5 due to an amino acid deletion. This HI loop is adjacent to VR-VII, which is largest in AAV5. The VR-IV, which forms the larger outermost finger-like loop contributing to the protrusions surrounding the icosahedral 3-fold axes of the AAVs, is shorter in AAV5, creating a smoother capsid surface topology. The HI loop plays a role in AAV capsid assembly and genome packaging, and VR-IV and VR-VII are associated with transduction and antigenic differences, respectively, between the AAVs. A comparison of interior capsid surface charge and volume of AAV5 to AAV2 and AAV4 showed a higher propensity of acidic residues but similar volumes, consistent with comparable DNA packaging capacities. This structure provided a three-dimensional (3D) template for functional annotation of the AAV5 capsid with respect to regions that confer assembly efficiency, dictate cellular transduction phenotypes, and control antigenicity.

Capsid Proteins

Whole Genome Sequencing and Genetic Diversity of Respiratory Viruses Detected in Children With Acute Respiratory Infections: A One-Year Cross-Sectional Study in Senegal.

Acute respiratory infections (ARI) are a health priority, especially in countries with limited resources. They are a major cause of morbidity and mortality, especially among children and the elderly. In Senegal, the endemic circulation of respiratory viruses other than influenza has been demonstrated. However, there is a paucity of data exploring the genetic diversity of these viruses based on whole-genome sequencing. In this study, we present data on the genetic diversity of respiratory viruses in children under 15 years old in Senegal, including an overview of the different pathogens detected. Between November 2022 and November 2023, we collected nasopharyngeal swabs from children seen in curative consultations for symptoms of acute respiratory infections. Of the 156 children included, 73.7% tested positive for at least one pathogen. The most frequently detected virus was rhinovirus (50.0%), followed by influenza B (41.6%) and human parainfluenza virus type 3 (7.6%). Combinations of rhinovirus/influenza B, human parainfluenza virus type 2/human parainfluenza virus type 4, and rhinovirus/influenza B/adenovirus were the most frequently identified. A statistically significant association was detected between some of the viruses detected. A high genetic diversity of respiratory viruses circulating in children was revealed. The strains were phylogenetically close to various strains circulating worldwide, suggesting a global circulation of respiratory viruses. Our study provides the first complete genome sequences of human parainfluenza viruses type 2, 3, 4 and human bocavirus from Senegal and thus contributes to the enrichment of international databases on sequences from Senegal and underlines the importance of sequencing in the dynamics of pathogen circulation.

Humans

Low-pass whole-genome sequencing reveals genomic diversity and ecotype-specific adaptation in indigenous Tigrayan chickens.

Indigenous chickens play a critical role in food security and climate resilience in smallholder systems, yet their genomic diversity and adaptive potential remain insufficiently characterised. This study employed low-pass whole-genome sequencing (LP-WGS; 0.2-1.99&#xd7;) to investigate genomic diversity, population structure, inbreeding and candidate environment-associated genomic variation in 33 chickens from highland, midland, and lowland agroecologies in the Tigray region of northern Ethiopia. After imputation and stringent filtering, 23.4 million high-confidence SNPs were retained, including&#x2009;~&#x2009;17% novel variants, indicating substantial uncharacterised genetic diversity in these populations. SNP density (13.8&#x2009;&#xb1;&#x2009;8.6 SNPs/kb) was comparable to values reported from high-coverage Ethiopian chicken datasets, demonstrating the suitability of LP-WGS for population genomics in resource-limited settings. Marked differences in genomic diversity were observed among ecotypes: midland chickens showed the highest nucleotide diversity (&#x3c0;&#x2009;=&#x2009;0.00267), followed by lowland (&#x3c0;&#x2009;=&#x2009;0.00233), whereas highland chickens showed the lowest diversity (&#x3c0;&#x2009;=&#x2009;0.00203) and elevated genomic inbreeding (FROH and FHOM &#x2248; 0.18). Population structure analyses revealed clear genetic separation among ecotypes. PCA (13.91% variation explained) distinguished lowland chickens along PC1 and separated highland from midland along PC2, while ADMIXTURE and FST patterns supported three major ancestral genomic backgrounds. Functional annotation of private missense variants uncovered distinct adaptive signatures reflecting the contrasting agroecological conditions. Highland chickens showed enrichment of candidate genes potentially involved in physiological processes relevant to high-altitude environments, including cold response, angiogenesis, cardiovascular regulation and metabolic homeostasis (eg., PARP1, ACOX2, ITGB3, EDNRB, SOX8, and SOX10). Midland chickens exhibited candidate signals of selection in genes with known roles in innate antiviral immunity, bacterial defence and inflammatory regulation (eg., BAK1, CLSTN1, CYSLTR1, CYSLTR2, CXCR7, GIPR, DSCAM, GDAP1, TLR3, TLR4, TLR7, IFIH1, ADORA1, EPHB1, and TMPRSS2). Lowland chickens displayed candidate variants associated with heat-stress response, DNA damage repair, oxidative balance and cardiovascular support under extreme temperatures (e.g., MLH1, BDKRB1, GPR19, FLT1, CCL18, TGM2, and RAMP3). Overall, the results indicate substantial genomic differentiation among ecotypes and suggest candidate environment-associated genetic divergence across Tigray's diverse agroecological zones. These populations may represent important reservoirs of adaptive genetic variation for climate-resilient poultry breeding, warranting further functional validation and conservation-oriented management.

Animals

Whole-genome sequencing identifies genetic diversity and adaptive signatures of hypoxia and ultraviolet radiation in Chinese chickens.

INTRODUCTION: Domestic chickens primarily descended from the wild red junglefowl, play a crucial role in global egg and meat production. China hosts diverse indigenous chicken populations that have adapted to various environmental conditions, including high-altitude with hypoxic and ultraviolet radiation stress. METHOD: We analyzed whole-genome sequences of 118 birds from five Indigenous Chinese chicken populations and 295 chicken genomes from publicly available databases to identify genomic diversity, admixture, and selection signatures of chickens adapted to high-altitude environments. Selection signatures were identified using nucleotide diversity (&#x3c0;), Tajima's D, XPEHH, and XP-CLR, selection scan methods. RESULTS: We observed a reduction in genetic diversity and historical declines in effective population size in high-altitude chicken, suggesting ongoing selection pressures shaping these populations. Selection scans identified nine genomic regions under strong positive selection, enriched for genes associated with hypoxia and ultraviolet radiation. Notably, five genes (TPK1, BAZ2B, MARCHF7, LLGL2, and RCAN3) were repeatedly detected across multiple selection signature analyses. RNA-seq analysis further confirmed the differential expression of these genes in the lung and heart tissues of chickens adapted to high and low altitudes, reinforcing their role in physiological adaptation to hypoxic environments. Altitude adaptation is driven by the selection of genes involved in oxygen metabolism, cellular stress response, and energy regulation. CONCLUSION: Our study provides compelling genetic evidence for differentiation between high and low and high-altitude Chinese chicken populations. These findings also ensure our understanding of local adaptation in poultry and establish a genomic framework for breeding strategies to improve environmental resilience to altitude-related stressors.

Animals