Search PubMedSearch

SEARCH · Search PubMed

Results for “synonymous substitution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Genomic regionality in rates of evolution is not explained by clustering of genes of comparable expression profile.

In mammalian genomes, linked genes show similar rates of evolution, both at fourfold degenerate synonymous sites (K4) and at nonsynonymous sites (KA). Although it has been suggested that the local similarity in the synonymous substitution rate is an artifact caused by the inclusion of disparately evolving gene pairs, we demonstrate here that this is not the case: after removal of disparately evolving genes, both (1) linked genes and (2) introns from the same gene have more similar silent substitution rates than expected by chance. What causes the local similarity in both synonymous and nonsynonymous substitution rates? One class of hypotheses argues that both may be related to the observed clustering of genes of comparable expression profile. We investigate these hypotheses using substitution rates from both human-mouse and mouse-rat comparisons, and employing three different methods to assay expression parameters. Although we confirm a negative correlation of expression breadth with both K4 and KA, we find no evidence that clustering of similarly expressed genes explains the clustering of genes of comparable substitution rates. If gene expression is not responsible, what about other causes? At least in the human-mouse comparison, the local similarity in KA can be explained by the covariation of KA and K4. As regards K4, our results appear consistent with the notion that local similarity is due to processes associated with meiotic recombination.

Animals

First nationwide full-genome characterisation of human-derived Andes virus in Chile: a retrospective genomic epidemiology study.

BACKGROUND: Andes virus (ANDV) is the only hantavirus known to transmit between humans and causes hantavirus cardiopulmonary syndrome in Chile and Argentina. In Chile, ANDV genomic diversity remains incompletely characterised. This study aimed to characterise the genetic diversity, geographical structure, and molecular signatures of ANDV using human clinical samples collected over a 13-year period (2011-24). METHODS: We conducted a retrospective genomic epidemiology study of ANDV infections in Chile. Clinical samples from patients with confirmed ANDV, collected between March 9, 2011, and June 27, 2024, were analysed and sequenced. Clinical and epidemiological data were obtained from diagnostic laboratories and surveillance programmes. Consensus sequences for the S, M, and L segments were generated, and genetic clustering and divergence were assessed using phylogenetic inference and variant calling. FINDINGS: We analysed clinical samples from 58 infected individuals and identified two major genomic variants of ANDV with distinct geographical distributions, defined by regionally structured patterns of nucleotide and amino acid substitutions across the S, M, and L segments: ANDV Chi-North (central Chile) and ANDV-South (southern Chile). No consistent clustering by clinical severity was observed, and no recurrent non-synonymous substitutions were uniquely associated with severe disease. Substitutions previously associated with person-to-person transmission in outbreaks in Argentina were not consistently observed in Chilean sequences, including in four person-to-person transmission cases. Although some substitutions described in ANDV-like viruses were present in the Chi-North lineage, this lineage remained phylogenetically distinct and geographically restricted to central Chile. INTERPRETATION: To our knowledge, this study provides the first nationwide genomic characterisation of human-derived ANDV in Chile. The identification of geographically structured variants indicates that ANDV diversity in Chile is driven by regional diversification rather than clinical outcome. The absence of consistent amino acid signatures associated with disease severity or person-to-person transmission suggests that these phenotypes are unlikely to be explained by viral genetic variation alone. These findings refine current understanding of ANDV evolution and highlight the need for continued integrated genomic surveillance in endemic regions. FUNDING: Agencia Nacional de Investigación y Desarrollo de Chile and National Institutes of Health.

Humans

Molecular Characterization and Epidemiology of Human Noroviruses in the Sverdlovsk Region, Russian Federation.

Human noroviruses (HuNoVs) stand as the primary cause of acute viral gastroenteritis outbreaks worldwide, particularly impacting children under the age of five. In Russia, reports of norovirus gastroenteritis have surged, especially in the post-COVID-19 era starting in 2022, with elevated infection rates reported into 2024. These viruses exhibit significant mutational variability, leading to the emergence of recombinant strains that can evade immune responses. A comprehensive examination of the complete genome is crucial for understanding the evolution of norovirus genes and for predicting potential outbreaks. This research focuses on analyzing the genotypic composition of HuNoVs circulating in the Sverdlovsk region during 2024, using Sanger sequencing and next-generation sequencing (NGS). Biological samples were collected (n = 384) from patients diagnosed with norovirus infection within the region. Bioinformatics analysis targeted the nucleotide sequences of the ORF1/ORF2 fragment and the assembly of complete genomes for the GII.4 and GII.7 genotypes. In total, 220 HuNoVs were characterized, representing 57.3% of the collected samples. The main capsid variants forming the predominant genotypic profile included GII.4 (n = 88, 40%), GII.7 (n = 86, 39%), and GII.17 (n = 14, 6%). Using NGS, we successfully assembled 8 out of 10 complete genomes for noroviruses GII.4[P16] and GII.7[P7]. Non-synonymous substitutions appeared at amino acid sites corresponding to the subdomains of VP1 in these strains. This molecular-genetic analysis provides contemporary insights into the genotypic composition, circulation patterns, and evolutionary dynamics associated with the dominant genovariants GII.4[P16] and GII.7[P7].

Norovirus

A novel regulation on the developmental checkpoint protein Sda that controls sporulation and biofilm formation in Bacillus subtilis.

UNLABELLED: Biofilm formation by Bacillus subtilis is triggered by an unusually simple environmental sensing mechanism. Certain serine codons, the four TCN codons (N for A, T, C, or G), in the gene for the biofilm repressor SinR caused lowered SinR translation and subsequent biofilm induction during transition from exponential to stationary growth. Global ribosome profiling showed that ribosomes pause when translating the four UCN (U for T on the mRNA) serine codons on mRNA, but not the two AGC/AGU serine codons. We proposed a serine codon hierarchy (AGC/AGT vs TCN) in that genes enriched in the TCN serine codons may experience reduced translation efficiency when serine is limited. In this study, we designed an algorithm to score all protein-coding genes in B. subtilis NCIB3610 based on the serine codon hierarchy. We generated a short list of 50 genes that could be subject to regulation by this novel mechanism. We further investigated one such gene from the list, sda, which encodes a developmental checkpoint protein regulating both sporulation and biofilm formation. We showed that synonymously switching the TCN serine codons to AGC in sda led to delayed biofilm formation and sporulation. This engineered strain also outgrew strains with other synonymously substituted sda alleles (TCN) in competition assays for biofilm formation and sporulation. Finally, we showed that the AGC serine codon substitutions in sda elevated the Sda protein levels. This serine codon hierarchy-based novel signaling mechanism could be exploited by bacteria in adapting to stationary phase and regulating important biological processes. IMPORTANCE: Genome-wide ribosome profiling in Bacillus subtilis shows that under serine limitation, ribosomes pause on the four TCN (N for A, C, G, and T), but not AGC/AGT serine codons, during translation at a global scale. This serine codon hierarchy (AGC/T vs TCN) differentially influences the translation efficiency of genes enriched in certain serine codons. In this study, we designed an algorithm to score all 4,000+ genes in the B. subtilis genome and generated a list of 50 genes that could be subject to this novel serine codon hierarchy-mediated regulation. We further investigated one such gene, sda, encoding a developmental checkpoint protein. We show that sda and cell developments controlled by Sda are also regulated by this novel mechanism.

Bacillus subtilis

Genomic Epidemiology and Clinical Characteristics of Mpox Lineage C.1 Outbreak in Thailand, 2023-2024.

Since 2022, human monkeypox virus (hMPXV) has emerged in non-endemic regions, including Thailand. However, the genomic dynamics and clinical correlates of local transmission remain incompletely defined. Whole-genome sequencing was performed on hMPXV from 16 patients in Thailand (2023-2024) using targeted amplicon NGS. Phylogenetic analyses integrated global reference sequences. Mutational profiles, specifically non-synonymous substitutions and APOBEC3-associated signatures, were analyzed in relation to clinical data. Phylogenetic reconstruction identified three temporal phases. Early 2022 cases (clade IIb lineages A and B) were interspersed with global sequences, consistent with multiple introductions. In contrast, 2023-2024 cases were dominated by lineage C.1. All 16 genomes belonged to C.1 (one C.1.1), and formed a distinct mid-2023 cluster, designated C.1/Thai/Cluster, supporting sustained local transmission. APOBEC3-associated mutations were pervasive across the C.1 lineage overall, including within C.1/Thai/Cluster, without evidence of significant enrichment specific to this cluster. The cohort comprised exclusively male patients (81% HIV-positive, MSM), with predominantly genital painful lesions and a median recovery time of 23 days. No significant associations were detected between viral genetic variation and clinical outcomes. Mpox transmission in Thailand evolved from multiple introductions to sustained C.1-dominated local spread, underscoring the importance of continued genomic surveillance.

Humans

Resurgence and molecular epidemiology of dengue virus serotype 4 amid the COVID-19 pandemic in Thailand.

INTRODUCTION: Dengue virus serotype 4 (DENV4) co-circulates with other serotypes in Thailand, a hyperendemic setting characterized by cyclical shifts in serotype predominance. During the COVID-19 pandemic, related public health interventions may have influenced dengue detection, transmission and epidemiological dynamics. MATERIALS AND METHODS: A total of 413 dengue NS1/PCR-positive samples collected from 2018 to 2024 at a tertiary hospital in Bangkok were serotyped using a real-time PCR DENV1-4 subtyping assay. Forty-seven DENV4-positive samples (Ct < 32) underwent whole-genome sequencing using the ATOPlex DENV1-4 panel on the DNBSEQ-G99RS platform. Sequencing reads were processed using CLC Genomics Workbench, followed by phylogenetic, mutation, codon-based selection-pressure, and epitope-mapping analyses. RESULTS: DENV4 was identified in 18.2% (75/413) of samples, with 48% of cases requiring hospitalization. Detection increased markedly from 11.8% (22/186) during 2018-2021-23.3% (53/227) during 2022-2024 (proportion ratio, 1.97; p&#x202f;=&#x202f;0.0025). Among the 47 whole-genome-sequenced samples, most strains clustered within genotype I lineage 4I_A.3 and showed lineage-associated non-synonymous substitutions. NS1-A90T, NS2A-L113F, and NS3-F523Y appeared exclusively during 2022-2024 (p&#x202f;<&#x202f;0.001), whereas NS2A-L113F, NS5-G223S, and NS5-Q631R showed concordant positive-selection signals by FEL and MEME. CONCLUSIONS: This hospital-based study highlights an increase in DENV4 detection in Bangkok during 2022-2024 compared with 2018-2021, accompanied by predominance of lineage 4I_A.3 distinct from Malaysian and Indonesian strains reported during a similar period. These findings support integrated clinical and genomic surveillance to monitor DENV serotype and genotype dynamics and inform public health and vaccine strategies.

COVID-19 pandemic

Chromosomal-level genome assembly of Trypanosoma carassii, the etiologic agent of a recent outbreak of trypanosomiasis in cage-cultured large yellow croaker (Larimichthys crocea) in China.

Trypanosoma carassii, a typical freshwater fish trypanosome, has recently been identified as the etiological agent of a trypanosomiasis outbreak in cage-cultured large yellow croaker (Larimichthys crocea) in China and has been designated as T. c. larimichthys. To date, publicly available genomic data for trypanosomes have been limited to terrestrial species, particularly those of medical importance. Here, we present a chromosome-level genome assembly of T. carassii, the first genome of an aquatic trypanosome, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding technologies. A preliminary genome survey based on Illumina sequencing data estimated the genome size at 56.38&#x2009;Mb with a heterozygosity of 1.17%. The final assembled genome spans 48.55&#x2009;Mb, with contig N50 and scaffold N50 values of 139.15 Kb, and achieves 100.00% BUSCO completeness. Hi-C data resolved the assembly into 34 chromosomes and 9 unanchored scaffolds. Repetitive elements account for 53.29% of the genome (approximately 25.87&#x2009;Mb). A total of 11,584 protein-coding genes were predicted, 95.36% of which were functionally annotated. Synonymous substitution rates analysis of paralogous genes indicates a recent burst of gene duplication, which likely corresponds to a whole-genome duplications. This high-quality genome assembly provides invaluable resources for understanding the evolution and host adaptation of aquatic trypanosomes.

Animals

Revisiting the genome assembly of Lupinus species reveals differential diploidization after a shared whole-genome duplication.

Accurate genome assemblies are essential for comparative genomics, yet Hi-C-guided scaffolding can introduce structural errors that misrepresent chromosome architecture and bias evolutionary inferences. Here, we identified pervasive scaffolding errors-including artificial fusions, internal inversions, and incomplete contig mounting-in 2 previously published Lupinus genomes (L. cosentinii and L. digitatus) using a segmentation method based on long terminal repeat (LTR) retrotransposon density. We reassembled both genomes, producing chromosome-level references of 472.7 Mb (16 chromosomes) and 427.2 Mb (21 chromosomes), with BUSCO completeness >98.5%. Synteny validation and reapplication of LTR profiling confirmed that all prior errors were resolved. Using these corrected genomes together with 4 additional Lupinus species and 2 outgroup legumes, we investigated postpolyploid evolution. Synonymous substitution rate (Ks) analysis revealed a genus-specific whole-genome duplication (WGD) event (Ks = 0.17) shared by all 6 Lupinus species. The proportion of WGD-derived genes varied markedly, from 60% in L. digitatus to only 36% in L. mutabilis, indicating differential diploidization. While all species retained a core set of WGD duplicates enriched in cytoskeleton organization, ion transport, and defense responses, each exhibited lineage-specific functional trajectories: cell wall modification in L. cosentinii and L. digitatus, nitrogen metabolism in L. albus and L. angustifolius, flower development in L. luteus, and stress/lipid metabolism in L. mutabilis. Our corrected assemblies provide optimal references for Lupinus comparative genomics, and our findings demonstrate that a shared WGD event can lead to both conserved and highly divergent postpolyploid fates, likely underpinning adaptive diversification within the genus.

Lupinus

Global Environmental Factors Impact the Evolution of Adult Hemoglobins in Squamata Reptiles (Lizards and Snakes) and Terrestrial Turtles.

Convergent evolution of oxygen transport mechanisms arises from respiratory proteins adapting to similar environmental pressures. We examined this relationship between adult hemoglobin subunits (Hbs: HBA1, HBAD, HBB1, and HBB2) found in land reptiles (lizards, snakes, and turtles) with their global distribution variables: Altitude, latitude, ambient temperature, and biomass production. We found that biomass was positively associated with the synonymous substitution rate (dS) of HBAD, while it showed the opposite trend for HBB2 in snakes. Additionally, latitude was negatively related to the dS of HBB2 in snakes, but nonsignificant with other Hbs. Altitude was negatively associated with &#x3c9; = dN/dS of HBA1 and HBAD, whereas temperature showed a similar negative trend with the &#x3c9; of HBAD across reptiles and in HBB2 of snakes. At amino acid sites, we found most were conserved except for 11 (two near the heme-binding pocket) across Hbs. These fast-changing sites shifted from polar to nonpolar residues, showing a pattern seen in high-altitude mammals. Our results highlight that in reptiles (i) Hbs are diversifying at individual amino acid sites while generally some subunits exhibiting lower &#x3c9; rates at higher altitudes and hotter temperatures, with the later and higher biomass ecosystems also linked to increases in dS; (ii) HBBs are the most conserved of the Hbs; (iii) latitudinal gradients only show a significant association with the dS of HBB2 in snakes; and (iv) gene conversion events occurred across HBBs in reptiles, which confound their homology assignation, except for snakes that evidenced a single major duplication in their HBBs.

Animals

The human IG heavy chain constant gene locus is enriched for large structural variants and coding polymorphisms that vary among human populations.

The human immunoglobulin heavy chain constant (IGHC) domain of antibodies (Ab) is responsible for effector functions critical to immunity. This domain is encoded by genes in the IGHC locus, where descriptions of genomic diversity remain incomplete. We utilized long-read sequencing to build an IGHC haplotype/variant catalog from 105 individuals of diverse ancestry. We discovered uncharacterized single nucleotide variants (SNV) and large structural variants (SVs, n=7), representing new genes and alleles enriched for non-synonymous substitutions, highlighting potential functional effects. Of the 221 identified IGHC alleles, 192 were novel. SNV, SV, and gene allele/genotype frequencies revealed population differentiation, including (i) hundreds of SNVs in African and East Asian populations exceeding a fixation index (FST) of 0.3, and (ii) an IGHG4 haplotype carrying coding variants uniquely enriched in Asian populations. Our results illuminate missing signatures of IGHC diversity and establish a new foundation for investigating IGHC germline variation in Ab function and disease.

Journal Article