Search PubMedSearch

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Chromosome-scale assembly with improved annotation provides insights into breed-wide genomic structure and diversity in domestic cats.

INTRODUCTION: Comprehensive genomic resources offer insights into biological features, including traits/disease-related genetic loci. The current reference genome assembly for the domestic cat (Felis catus), Felis_Catus_9.0 (felCat9), derived from sequences of the Abyssinian cat, may inadequately represent the general cat population, limiting the extent of deducible genetic variations. OBJECTIVES: The goal was to develop Anicom American Shorthair 1.0 (AnAms1.0), a reference-grade chromosome-scale cat genome assembly. METHODS: In contrast to prior assemblies relying on Abyssinian cat sequences, AnAms1.0 was constructed from the sequences of more popular American Shorthair breed, which is related to more breeds than the Abyssinian cat. By combining advanced genomics technologies, including PacBio long-read sequencing and Hi-C- and optical mapping data-based sequence scaffolding, we compared AnAms1.0 to existing Felidae genome assemblies (20 scaffolds, scaffolds N50 > 150 Mbp). Homology-based and ab initio gene annotation through Iso-Seq and RNA-Seq was used to identify new coding genes and splice variants. RESULTS: AnAms1.0 demonstrated superior contiguity and accuracy than existing Felidae genome assemblies. Using AnAms1.0, we identified over 1.5 thousand structural variants and 29 million repetitions compared to felCat9. Additionally, we identified > 1,600 novel protein-coding genes. Notably, olfactory receptor structural variants and cardiomyopathy-related variants were identified. CONCLUSION: AnAms1.0 facilitates the discovery of novel genes related to normal and disease phenotypes in domestic cats. The analyzed data are publicly accessible on Cats-I (https://cat.annotation.jp/), which we established as a platform for accumulating and sharing genomic resources to discover novel genetic traits and advance veterinary medicine.

Animals

Genetic structuring and estimation of reproductive adults in Onchocerca volvulus: A genome-wide analysis across hosts and regions.

Genomic analysis of parasites can deepen our understanding of their transmission, population structure, and important biological characteristics. Onchocerciasis (river blindness), caused by the parasitic nematode Onchocerca volvulus, involves adult worms residing in subcutaneous nodules that produce larval-stage microfilariae (mf), which are routinely detected in the skin for diagnosis. Whole-genome studies of mf are limited; most analyses have focused on the mitochondrial genome. We conducted a genome-wide analysis with 94% median nuclear genome coverage, analyzing 171, 37, and 98 mf from 16, 3, and 5 individuals from Ghana, Liberia, and the Democratic Republic of Congo, respectively. These data were used to investigate population differentiation, estimate the number of reproductive adult worms, and analyze genetic variation across chromosomes. Population genetic analyses across hosts and countries showed that nuclear genome diversity can reveal fine-scale genetic structure, even between geographically close countries, providing more resolution than mitochondrial haplotype data. By reconstructing maternal and paternal sibships, we estimated the number of reproductively active adult filariae. Comparisons between adult worm estimates from genetic data and nodule observations showed that genetics-based estimates were higher or equal to observed worm counts in 8 out of 9 hosts for female worms and 7 out of 9 hosts for male worms. Our analysis also revealed lower-than-expected X chromosome diversity, consistent with neo-X chromosome fusions in filarial species. This study represents an important step in using nuclear genome data from mf to support onchocerciasis elimination efforts and in developing genetic tools that could inform mass drug administration programs.

Onchocerca volvulus

Genomic and Structural Analysis of Gamete Recognition Proteins in a Broadcast Spawning Echinoderm Mesocentrotus franciscanus.

Gamete recognition proteins are expressed on the surfaces of sperm and eggs, where they mediate interactions between gametes. The genetic basis for gamete recognition proteins, as well as their structure and interactions, have yet to be fully resolved. Using a new high-quality de novo genome assembly for the sea urchin Mesocentrotus franciscanus, we investigated the genomic structure, expression, and protein forms of several gamete recognition proteins: sperm bindin, egg receptor for sperm (HSP110), and egg bindin receptor (EBR1), as well as the receptor for egg jelly (REJ) and its paralogs. To inform future population genetic and evolutionary studies, we resolve the genomic structure of the large EBR1 protein, identifying fewer tandem CUB-TSP1 repeats in EBR1 compared to the initial characterization of this protein. As expected for an egg receptor for sperm, EBR1 is highly expressed in female reproductive tissues (eggs and female gonad), compared to other tissues. In contrast, HSP110 shows similar levels of expression across male and female reproductive tissues, as well as across non-reproductive tissues and development stages. HSP110 might be a pleiotropic gene that in part influences fertilization. Using protein structural modeling and functional domain predictions, we propose hypotheses about potential interactions among EBR1, bindin, and HSP110 proteins that may provide insight into sperm-egg interactions in sea urchins. Resolving the genomic structure of genes encoding gamete recognition proteins, in combination with functional annotations and protein structural modeling, enables deeper investigation into the consequences of variation in gamete recognition proteins and the evolution of reproductive isolation.

Mesocentrotus franciscanus

Surprisingly frequent chromosomal instability in cultivated peanut.

This study, the third in a three-part series, investigates whether chromosomal instability persists in cultivated peanut. The allotetraploid peanut (Arachis hypogaea; genome type AABB) originated from the hybridization and polyploidization of A. duranensis (AA) and A. ipaënsis (BB). Our first study established that this was an extremely narrow genetic origin, likely from a single hybridization event. This raised a paradox: how did such narrow genetics give rise to the phenotypic diversity seen in cultivated peanut? The second study addressed this, showing that a single neoallotetraploid spontaneously generates striking diversity, and that homoeologous exchanges-abundant in early generations following polyploidy-are a key mechanism in creating this diversity. In contrast to this early-generation instability, cultivated peanut is generally considered to be genetically stable, presumably due to selection. This third study tests whether residual instability still occurs in modern peanut. From a single plant of the highly selfed 'genome stock' of the cultivar 'Tifrunner', we advanced lineages through seven generations in a pollinator-free greenhouse. Among 233 plants, we identified three new large-scale chromosomal instability events: a large deletion on chromosome B01, associated with reduced pod width and seed weight, and two ABBB compositions involving chromosomes A02/B02 and A05/B05. With these observations in hand, we reinterpreted previously published data from two recombinant inbred populations. Together, these results indicate that at least 1% of pure pedigree A. hypogaea plants exhibit spontaneous large-scale chromosomal changes-a surprising frequency of instability that likely contributes to peanut's long-term adaptability and evolution.

Arachis

Multi-locus allelic architecture underlying natural variation in leaf rolling in japonica rice.

Leaf rolling is a key component of rice canopy architecture that affects light interception, microclimate formation, and planting density. The contribution of naturally occurring allelic variation to quantitative variation in leaf rolling within cultivated rice remains poorly understood, while extreme leaf rolling caused by loss-of-function mutations often results in detrimental pleiotropic effects. Herein, we examined how multi-locus allelic variation contributes to natural variation in leaf rolling within japonica rice. Leaf rolling was quantified based on the leaf rolling index (LRI) using a panel of 201 japonica accessions. The phenotype was transformed using the Yeo-Johnson method to reduce strong right skewness and improve the distributional properties of the data, thereby facilitating subsequent regression modeling. Haplotype analyses were performed for previously reported leaf rolling-associated genes and genome-wide association study (GWAS) lead loci, leading to the identification of five loci exhibiting substantial haplotype-dependent phenotypic variation. Phenotypically defined allelic groups represented these loci were subsequently evaluated using multiple linear regression (MLR), with the first two principal components derived from genome-wide SNP data included as covariates to account for population structure. The final MLR model identified four loci (qALR1, OsYABBY1, OsSLL2, and OsSRL10) as the independent contributors to leaf rolling variation, collectively explaining 21% of the variance in the transformed phenotype after accounting for population structure. Model diagnostics and ten-fold cross-validation supported the statistical validity of the framework and indicated stable model performance across validation folds. Analysis of multi-locus allelic combinations showed 13 distinct configurations that clustered into three phenotypically differentiated groups. This reflected the cumulative dosage of high-leaf rolling alleles. Thus, the natural variation in leaf rolling in japonica rice is governed by the additive effects of multiple moderate-impact loci. The multi-locus allelic framework established here provides a statistically sound and biologically interpretable basis for dissecting polygenic canopy traits and practical guidance for developing genetic materials aimed at optimizing rice plant architecture.

cross-validation

Whole-genome sequencing of 490,640 UK Biobank participants.

Whole-genome sequencing provides an unbiased and complete view of the human genome and enables the discovery of genetic variation without the technical limitations of other genotyping technologies. Here we report on whole-genome sequencing of 490,640 UK Biobank participants, building on previous genotyping effort1. This advance deepens our understanding of how genetics associates with disease biology and further enhances the value of this open resource for the study of human biology and health. Coupling this dataset with rich phenotypic data, we surveyed within- and cross-ancestry genomic associations and identified novel genetic and clinical insights. Although most associations with disease traits were primarily observed in individuals of European ancestries, strong or novel signals were also identified in individuals of African and Asian ancestries. With the improved ability to accurately genotype structural variants and exonic variation in both coding and UTR sequences, we strengthened and revealed novel insights relative to whole-exome sequencing2,3 analyses. This dataset, representing a large collection of whole-genome sequencing data that is available to the UK Biobank research community, will enable advances of our understanding of the human genome, facilitate the discovery of diagnostics and therapeutics with higher efficacy and improved safety profile, and enable precision medicine strategies with the potential to improve global health.

Humans

[Modern variations of human influenza group A viruses at the molecular level].

The authors own results on the variety of the genomic primary structures in human influenza A viruses participating in the epidemic process, including the atypical viruses. The comparative studies revealed new trends in the HA gene antigenic drift on the late stages and the PB1 gene shift. Modifications occurring in the primary structure of the influenza A viruses native genomes during laboratory treatment (adaptation to new hosts, vaccine preparation, egg passaging) have been analyzed. Sequencing of several types of "antigenic anachronisms" revealed the direct links between some of such viruses and the anthropogenic pollution of the biosphere by vaccine strains. Modifications in the HA genes of influenza A viruses during the persistent infection have also been studied.

Amino Acid Sequence

Complex structural variation, phylogeny, and disease associations of the mucin pangenome.

Mucins are large glycoproteins that provide hydration and barrier function to epithelial tissues. Although genetically heterogeneous, all mucins harbor a large exon composed of variable number tandem repeats (VNTRs). Short-read sequencing has limited our understanding of mucin VNTR diversity and makes disease association studies challenging. We leverage 296 long-read phased genome assemblies to characterize 14 mucin family members, achieving &#x2265;97% accuracy across 572 haplotypes. Phylogenetic haplogroup analysis reveals extraordinary structural heterozygosity, with MUC4 harboring the greatest allelic diversity (n=240 distinct lengths) and MUC12 the greatest size range (&#x394; = 55,233 bp; 23,080 amino acids). Ten mucins show significant population stratification (pFDR < 0.05). At the MUC4/MUC20 locus, we characterize higher-order structural variation, including a recurrent inversion, copy number variation, and interlocus gene conversion. Optimized genotyping achieves &#x2265;95% haplogroup concordance across 10 loci. We apply this to 4,637 deeply phenotyped cystic fibrosis patients and identify a significant association between short MUC1 VNTRs and severe disease (p=0.0056), demonstrating the pangenome's utility for complex locus genotyping and disease discovery.

Journal Article

Molecular cloning and sequence analysis of duck hepatitis B virus genomes of a new variant isolated from Shanghai ducks.

The genomes of duck hepatitis B virus (DHBV) from a brown duck (S5) and a white duck (S31) kept independently in Shanghai, China, were cloned and the complete nucleotide sequence of each virion DNA (DHBV-S5 and DHBV-S31) was determined. DHBV-S5 and DHBV-S31 were both 3027 bp in length and 6 bp longer than the other two DHBVs analyzed previously, DHBV16 and DHBV3. The genomes of DHBV-S5 and DHBV-S31 encoded three long overlapping open reading frames designated as P, S, and C. A possible new open reading frame was found in a complementary strand of each viral genome, as 336 bp for DHBV-S5 and 306 bp for DHBV-S31, respectively. A pair of 3-bp insertions were found in the overlapping region of pre-S2 and P and so two amino acids were inserted in this region in DHBV-S5 and DHBV-S31. The nucleotide sequence variation between DHBV-S5 and DHBV-S31 (4.9%) was similar to that between DHBV16 and DHBV3 (5.6%), and less than the variations between either of these Shanghai clones and DHBV16 or DHBV3 (9.5-10.4%). The amino acid sequence was also conserved in the two Shanghai clones but showed group difference from DHBV16 or DHBV3. Thus these two independent Shanghai clones of DHBV showed geographical characteristics of genomic structure.

Amino Acid Sequence

Targeted population genomics uncovers demographic history and genetic divergence in north American wild cranberry.

Wild populations of North American cranberry (Vaccinium macrocarpon Aiton) are reservoirs of genetic variation that may contribute to the improvement of breeding-relevant traits. However, the extent to which wild genetic variation is geographically structured and represented in elite germplasm remains unclear. We analysed 179 wild cranberry accessions from the upper Midwest and Eastern North America to estimate nucleotide diversity (&#x3c0;), population structure, and loci associated with genetic differentiation and environmental variables using a genome-informed targeted genotyping panel. Additionally, 14 demographic scenarios were evaluated using site-frequency-spectrum-based inference to identify historical events that could explain current genetic diversity. We observed extremely low nucleotide diversity within the targeted panel (&#x3c0; = 5 &#xd7; 10-6). Rare allele distributions strongly influenced &#x3c0; and Tajima's D values, suggesting constrained diversity in the genomic regions assayed that is not captured by heterozygosity-based estimates alone. However, we interpreted these results as conservative lower bounds on genome-wide neutral diversity because the targeted panel is enriched for genic and conserved regions. A clear separation between the Midwest and East populations was observed, with inbreeding coefficients ranging from -0.13 to 0.15. Furthermore, site frequency spectrum inference from the targeted panel supported a demographic scenario consistent with a significant population reduction &#x2248;15-14 thousand years ago (kya), followed by a divergence between the two regions &#x2248;12 kya, and an asymmetric gene flow &#x2248;1.3 kya. We detected 254 candidate loci showing regional allele-frequency differentiation. Several of these loci colocalized with candidate genes linked to stress response, development, and metabolic processes. To evaluate the representation of geographically differentiated wild alleles in a breeding context, we analysed Rutgers breeding materials (n&#x2009;=&#x2009;484) and found that this panel is enriched for common alleles in Eastern wild populations. These findings indicate regionally structured allele-frequency variation in wild cranberry, with potential relevance to environmental response and breeding. This study extends prior wild cranberry population-genetic research by providing targeted-panel estimates of diversity, comparisons of demographic models, and breeding insights on geographically differentiated alleles, while highlighting the importance of conserving wild cranberry germplasm for use in modern breeding programs.

Journal Article

Mitogenomic Insights Into the Population Structure and Demographic History of Tree Shrews (Tupaia belangeri) in China.

The northern tree shrew (Tupaia belangeri) exhibits significant morphological and geographical variations, but its evolutionary history and subspecies boundaries remain controversial. Here, we analyzed the complete mitochondrial genomes of 63 individuals, representing 12 populations in China to study phylogenetic relationships, genetic diversity, and population history. Phylogenetic analysis consistently restored four mitochondrial branches with strong geographic structures and significant differences. The three lineages correspond to geographically restricted subspecies (T. b. tonquinia, T. b. modesta, and T. b. gaoligongensis), while individuals assigned to several traditional subspecies cluster in a broad mainland lineage (T. b. chinensis, T. b. yunalis, and T. b. yaoshanensis). The divergence time estimate places the origin of the main lineage in the Miocene, consistent with major tectonic and geomorphological events. Demographic analysis revealed different population histories, including varying degrees of expansion in recent continental and island lineages, as well as the long-term stability of T. b. gaoligongensis. Genetic diversity varied markedly among lineages, with the highest diversity observed in the T. b. gaoligongensis and the lowest diversity observed in the T. b. modesta. These findings demonstrate that landscape complexity and demographic history are key drivers of evolutionary diversification in T. belangeri, challenging classical morphology-based subspecies classifications and underscoring the need for comprehensive sampling across both domestic and international ranges.

Tupaia belangeri

Identification of a novel non-coding deletion in Allan-Herndon-Dudley syndrome by long-read HiFi genome sequencing.

BACKGROUND: Allan-Herndon-Dudley syndrome (AHDS) is an X-linked disorder caused by pathogenic variants in the SLC16A2 gene. Although most reported variants are found in protein-coding regions or adjacent junctions, structural variations (SVs) within non-coding regions have not been previously reported. METHODS: We investigated two male siblings with severe neurodevelopmental disorders and spasticity, who had remained undiagnosed for over a decade and were negative from exome sequencing, utilizing long-read HiFi genome sequencing. We conducted a comprehensive analysis including short-tandem repeats (STRs) and SVs to identify the genetic cause in this familial case. RESULTS: While coding variant and STR analyses yielded negative results, SV analysis revealed a novel hemizygous deletion in intron 1 of the SLC16A2 gene (chrX:74,460,691&#x2009;-&#x2009;74,463,566; 2,876&#xa0;bp), inherited from their carrier mother and shared by the siblings. Determination of the breakpoints indicates that the deletion probably resulted from Alu/Alu-mediated rearrangements between homologous AluY pairs. The deleted region is predicted to include multiple transcription factor binding sites, such as Stat2, Zic1, Zic2, and FOXD3, which are crucial for the neurodevelopmental process, as well as a regulatory element including an eQTL (rs1263181) that is implicated in the tissue-specific regulation of SLC16A2 expression, notably in skeletal muscle and thyroid tissues. CONCLUSIONS: This report, to our knowledge, is the first to describe a non-coding deletion associated with AHDS, demonstrating the potential utility of long-read sequencing for undiagnosed patients. Although interpreting variants in non-coding regions remains challenging, our study highlights this region as a high priority for future investigation and functional studies.

Humans

Genomic characterization and mutation rate of hepatitis C virus isolated from a patient who contracted hepatitis during an epidemic of non-A, non-B hepatitis in Japan.

To investigate the genomic characterization of hepatitis C virus (HCV) isolated from patient who contracted hepatitis during an epidemic of non-A, non-B (NANB) hepatitis in Shimizu city, Japan, we have cloned the nucleotide sequence of the viral genome (HCV-KF) spanning the structural domain. When compared to other previously reported HCV isolates, HCV-KF showed an overall identity at the amino acid level of 90.0 to 92.1% with Japanese isolates and 80.9 to 82.1% with American-like isolates. The HCV-KF genome displays an insertion of three nucleotides in-frame (corresponding to one amino acid) found at the junction between the E1 and E2/NS1 region. The mutation rate of the HCV-KF genome was assessed by comparing the nucleotide and deduced amino acid sequences of the viral RNA obtained from the serum of the original patient with viral sequences derived from the serum of a chimpanzee inoculated with the same serum 9 years previously. The substitution rate of the viral genome was estimated at 0.9 x 10(-3) nucleotides per site per year for the HCV structural region. The highest mutation rate was found in the hypervariable region within the E2/NS1 domain. It is suggested that the outbreak in Shimizu city was caused by a strain of HCV closely related to the Japanese-like subgroup of isolates.

Americas

Genotyping and sequence analysis of apolipoprotein E isoforms.

Apolipoprotein E (apoE), a polymorphic plasma protein, is essential for catabolism of lipoproteins by receptor-mediated endocytosis. One of the apoE isoforms (E2) differs in its binding affinity to specific receptors and contributes to variations in lipoprotein metabolism. Diagnosis of apoE isoforms is done by isoelectric focusing, but it is hindered by various degrees of post-translational sialylation of the apoE protein. Electrophoretically silent structural variations may also escape detection by this technique. We describe a method for genotyping apoE based on hybridization of allele-specific oligonucleotides with enzymatically amplified genomic DNA, which permits unambiguous diagnosis of six common apoE phenotypes within 24 h. Among 100 E2 alleles present in 81 unrelated individuals genotyped by this technique, we found two rare structural mutants of apoE in addition to the common E2 form, E2(158Arg----Cys). Automated sequencing of amplified DNA identified the rare mutants as E2(136Arg----Ser) and E2(145Arg----Cys). The genotypic method may complement or even replace isoelectric focusing for routine determination of apoE phenotypes and for identification of rare structural variants.

Alleles

A genome-wide assessment of the population structure of thirteen admixed and pure Australian beef cattle breeds.

Knowledge of population structure is a key factor for successful multi-breed genomic prediction, especially in single-step analysis when metafounders are considered. In Australia, current assessments mostly focus on single breeds using a single-step genomic prediction method. However, the effective integration of pedigree, phenotypic, and genomic data in a multi-breed framework still requires further research, especially for combined analyses including admixed and multi-breed populations. This study began with 602,952 genotyped individuals with 8K SNPs in common from 13 beef cattle breeds (Alexandria, Angus, Brahman, Brangus, Charolais, Droughtmaster, Hereford, Kynuna, Limousin, Santa Gertrudis, Shorthorn, Speckle Park, and Wagyu). Due to different numbers of animals being genotyped in each breed, a representative subset of animals was chosen by employing a validated sampling strategy using Gaussian Mixture Models (GMM) complemented by Principal Component Analysis (PCA) within each breed. Subsequently, a specific number of animals in each cluster were randomly selected to capture the entire genetic diversity per breed, with a total of 260 animals from each breed. The first three principal components explained 59.89% of the total variation, with PC1 (33.54%) clearly separating Bos indicus from Bos taurus lineages. Admixture analysis identified stable ancestral components and defined the genetic makeup of both pure and composite populations. The results showed extensive genetic diversity in some breeds and highlighted distinct genetic differences between Bos indicus and Bos taurus breeds. In addition, six composite breeds' admixture levels confirmed their origin and breed history, revealing a directional shift in ancestry proportions by a longitudinal increase in Brahman ancestry within tropical composites over time. Thus, the findings pave the way for more effective utilization of genetic diversity both within and across populations and provide a framework for designing multi-breed genetic evaluations and breeding programs to improve productivity and profitability in Australian beef production.

Animals

Osteoarthritis phenotypes: advancing precision medicine through clinical, structural, and molecular stratification.

PURPOSE: Osteoarthritis (OA) is now understood as a heterogeneous syndrome driven by diverse biological, biomechanical, metabolic, genetic, and molecular mechanisms. This variability explains differences in disease progression and treatment response, challenging the traditional "one-size-fits-all" approach. This review highlights OA phenotyping as a key step toward precision medicine, focusing on clinical, structural, and molecular classifications that inform individualized care. METHODS: A narrative review was conducted using a non-systematic search of major databases and Osteoarthritis Research Society International sources (2010-2026). Evidence was thematically synthesized across clinical, imaging, and molecular domains to characterize OA phenotypes and their potential relevance to precision medicine. RESULTS: Multiple OA phenotypes were identified: inflammatory, metabolic, biomechanical, cartilage-subchondral, pain-sensitization, and aging/senescence. These exhibit distinct clinical features, risk factors, and therapeutic responses. Imaging-based phenotypes (e.g., inflammatory, meniscus-cartilage, subchondral bone, atrophic, hypertrophic) and molecular endotypes (low turnover, structural damage, systemic inflammation) further refine stratification. Pain-structure discordance is notable in sensitization phenotypes and may predict poorer surgical outcomes. Joint-specific variations and emerging genomic and epigenetic insights underscore disease complexity. Advances in imaging, biomarkers, and machine learning may enable earlier detection and patient clustering, though clinical application remains limited. CONCLUSION: Phenotype- and endotype-based classification represents a critical advancement toward precision OA management. Tailored interventions based on stratification hold promise for improving outcomes; however, clinical translation remains limited by overlapping phenotypes, lack of validated biomarkers, and inconsistent results from phenotype-driven trials. Wider clinical adoption requires standardized definitions, validation across joints, and integration of multimodal diagnostic tools into routine practice.

Humans

Tandem duplication-driven expansion and UV-B stress adaptation of the LHC gene family in Artemisia annua L.

BACKGROUND: Artemisia annua L., is the primary natural source of the antimalarial drug artemisinin. In nature, fluctuating light is a major environmental stress that affects plant growth and artemisinin biosynthesis. Although the light-harvesting chlorophyll a/b-binding (LHC) superfamily plays a key role in mediating plant responses to fluctuating light, systematic research of this gene family in A. annua has not yet been conducted, limiting our understanding of light adaptation in this medicinally important species. RESULTS: This study investigated the evolutionary dynamics and functional adaptation of the light-harvesting chlorophyll a/b-binding (LHC) superfamily in A. annua, with a focus on the early light&#x2011;induced protein (ELIP) subfamily. Comparative genomics of 24 plant species showed that the LHC superfamily recently expanded in the examined Asteraceae lineages through duplication events. In A. annua, 229 LHC genes identified from four haplotype genomes comprised 205 allelic and 24 haplotype-specific loci, with the ELIP subfamily expanding significantly via tandem duplication. Notably, compared to non-Asteraceae plants, ELIPs exhibited a uniform single-exon architecture, indicating it is a genomic feature unique to Asteraceae plants. Population genomics of 41 individuals showed dynamic copy number variations ranging from 1 to 4 copies per locus. Interestingly, a structurally disrupted ELIP allele remained transcriptionally active and produced long aberrant transcripts, showing that this subfamily is still actively evolving. Under UV-B stress, AaELIP loci showed synchronized induction trend but differed in expression levels, suggesting a division into major and auxiliary roles within the expanded tandem cluster. Overall, while the response of ELIPs to light stress is evolutionarily conserved, this dramatic expansion and structural streamlining of AaELIPs may represent a key evolutionary adaptation that enhances the plant's ability to cope with intense light and radiation stress. CONCLUSIONS: Collectively, this study demonstrates a significant expansion of the LHC superfamily in A. annua, especially within the ELIP subfamily, as well as its robust response to UV-B treatment, underscoring the essential role of ELIPs in mediating light stress responses. These findings provide a valuable foundation for future research to uncover the molecular mechanisms underlying A. annua's adaptation to complex light environments.

Artemisia annua

Global diversity of integrating conjugative elements (ICEs) in Helicobacter pylori and their influence on genome architecture.

Integrating conjugative elements (ICEs) are mobile genetic elements conferring a wide range of beneficial functions upon their bacterial hosts. Generally, they can be activated from their integrated states to undergo horizontal gene transfer via conjugation. In the case of the human gastric pathogen Helicobacter pylori, a paradigm for extensive genetic diversity, highly efficient natural transformation and recombination processes may superimpose canonical transfer of its two ICEs termed ICEHptfs3 and ICEHptfs4, and thus shape their composition substantially. Here, as a part of the Helicobacter pylori Genome Project (HpGP) initiative, we have analyzed high-quality genome sequences from 1011 clinical strains with respect to their ICE content and variability. We show that both elements are highly prevalent in all H. pylori populations, but have a strong tendency for gene erosion. ICE sequence variations reflect the population structure and show a clear signature of increased horizontal transfer. A detailed map of ICE integration sites revealed local preferences, but also how recombination processes result in hybrid elements or genome rearrangements. Population-specific differences in ICE cargo genes might reflect distinct requirements in the biological functions provided by these mobile elements.

Journal Article