Search PubMedSearch

SEARCH · Search PubMed

Results for “nucleotide polymorphism patterns”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Mitochondrial DNA polymorphism in dogs.

Mitochondrial DNA (mtDNA) polymorphism was studied in 20 mongrel dogs using 14 restriction enzymes. The polymorphism was observed in the cleavage patterns of Apa I, EcoR I, EcoR V, Hinc II and Sty I. Three morphs using EcoR I and Apa I and two morphs using EcoR V, Hinc II and Sty I were found. However, no polymorphism was detected in the cleavage patterns of BamH I, Bgl II, Hae II, Hind III, Pst I, Sca I, Stu I and Xba I. The examined dogs were classified in seven types using five restriction endonucleases which recognized six nucleotide sequences. The value of nucleotide diversity was estimated to be 0.0055. A phylogenetic tree constructed by genetic distances among seven restriction types showed at least two clusters of mtDNA.

Animals

Geographic mitochondrial DNA variation in the rock hyrax, Procavia capensis.

Restriction-fragment-length polymorphisms in mitochondrial DNA (mtDNA) were used to evaluate geographic population genetic structure in the rock hyrax, Procavia capensis, a species which occurs widely, though restricted to rocky habitat, throughout South Africa. Ten restriction endonucleases were employed to assay mtDNAs from 55 specimens representing 10 localities. Haplotypes showed strong geographic patterning, and estimates of nucleotide sequence divergence indicate two major clades thought to be dispersing along separate routes. The divergence time of approximately 2 Myr between clades is relatively high for intraspecific variation. We speculate that the marked genetic break distinguishing the northwestern populations from those constituting the south/central clade may be indicative of two species in what has conventionally been regarded as P. capensis.

Animals

Protein C deficiency: identification of a novel two-base pair insertion and two point mutations in exon 7 of the protein C gene in Spanish families.

We have applied single-strand conformation polymorphism (SSCP) to the analysis of exon 7 of the anticoagulant protein C (PC) gene, in 13 PC-deficient Spanish families. Abnormal patterns were visualized in three samples from type I or quantitative PC deficient proposita. A previously undescribed mutation due to a TT insertion after nucleotide 6139, between codons Gly-142 and Arg-143 was found in one family. The mutation (6139,ins TT) should result in a frameshift with a stop at codon 156, which agrees with the presence of a type I or quantitative PC deficiency in the affected members of the family. The second mutation identified was a C to T transition at nucleotide 6274, 9 base pairs into intron G. This mutation (6274,C-->T), found for the first time in a Spanish family, is identical to the previously characterized PC Sant Louis. The third mutation was a G to A transition that replaces arginine 178 with glutamine (178,R-->Q). This is the third case of 178,R-->Q mutation in 17 apparently unrelated Spanish families with type I PC deficiency. Furthermore, SSCP analysis allowed the detection of another previously described mutation in a PC-deficient Spanish family (178,R-->W).

Base Sequence

Mono- and bi-phasic Salmonella typhi: genetic homogeneity and distinguishing characteristics.

Several lines of evidence indicate a relatively low genetic heterogeneity in the natural Salmonella typhi population. However, some S. typhi isolates found in Indonesia express, instead of the usual fliC-d flagellin gene, a different flagellar gene fliC-j. In addition, Indonesian strains may have a second flagellar antigen fliC-z66. We have previously suggested, on the basis of the flagellar antigen constitution, that S. typhi evolved in an isolated human population in Indonesia. In order to test this hypothesis, we have gathered S. typhi isolates from around the world and tested the genetic heterogeneity among them. In general, polymorphism was greater in isolates from the Far East, as was indicated by Southern hybridizations with rDNA and fliC DNA probes. Gene fliC-j was not found in S. typhi isolates, other than those from Indonesia. However, the one-clone origin of S. typhi was indicated by a common DNA fingerprint pattern and by the occurrence, in the 5' end region of the fliC gene, of 10 scattered nucleotides that differ from the corresponding 10 nucleotides in other fliC alleles studied. These nucleotides were present in all isolates tested but did not change the amino acid sequence of the flagellin polypeptide.

Africa

Evolutionary dynamics of the chloroplast genome in Abutilon (Malvoideae, Malvaceae).

The genus Abutilon Mill. (Malvaceae) comprises approximately 178 species distributed across tropical and subtropical regions, many of which hold significant ornamental, economic, and medicinal value; yet its taxonomic classification remains challenging. In this study, six species were sequenced from herbarium specimens, and the chloroplast (cp.) genomes of ten additional species were assembled de novo from publicly available raw data. Three previously reported cp. genomes were also incorporated to characterise cp. genome structure, identify polymorphic loci, and perform phylogenetic analyses. The cp. genomes ranged from 159,458 to 160,454 bp and exhibited the typical quadripartite structure, with each genome containing 112 unique genes (78 protein-coding, 30 tRNA, and 4 rRNA) that showed conserved content and organisation. These genomes exhibited high similarity in GC content, inverted repeat boundaries, relative synonymous codon usage, amino acid frequencies, and substitution patterns. However, notable variation was observed in the total number of simple sequence repeats, ranging from 70 to 97 per genome. Selection analyses indicated predominant purifying selection, with evidence of episodic positive selection detected in rpoC2, rbcL, and ycf1. Two codons in rbcL were clade-specific and provided phylogenetic signal distinguishing Australian and Old World pantropical species. Nucleotide diversity analysis identified six highly polymorphic intergenic spacers (trnH-psbA, rps19-rpl2, psbT-pbf1, psaC-ndhD, trnR-atpA, and ndhJ-ndhK) that may be suitable for taxonomic studies. The phylogeny from maximum likelihood (ML) and Bayesian inference (BI) resolved two major clades: one comprising an exclusively Australian lineage occurring predominantly in arid and semi-arid environments, and the other a pantropical lineage spanning multiple continents. Abutilon grandifolium was recovered as sister to the remaining sampled Abutilon taxa in both ML and BI analyses, although no biogeographic origin inference can be drawn from this placement pending broader taxon sampling and integration of nuclear genomic data. These findings provide insights into the evolutionary dynamics of the cp. genome in Abutilon and offer a foundational genomic framework for refining Abutilon taxonomy.

Genome, Chloroplast

Same Sex Chromosomes With Independent Origins in Haplochromine Cichlids.

Elucidating theories of sex chromosome evolution requires approaches that allow fine scale delimitations of sex-determining regions within a phylogenetic context. This can address whether shared sex chromosomes across related species are due to shared ancestry, or whether genetic sex-determining regions have repeatedly evolved. Haplochromine cichlids, as one of the most successful fish lineages on Earth, have been a focal study system of sex chromosome research, both because of their rapid rate of sex chromosome turnover and the repeated emergence of certain sex chromosomes across the lineage. Here, we newly describe sex chromosomes in members of the earliest branch of the modern haplochromines, the Tropheini, based on whole-genome sequencing data, using a combination of SNP- and kmers-based methods. We show that despite the repeated co-options of ancestral chromosomes LG5 and LG7 in these species, the origins of these sex chromosomes are independent. Investigation of gene functions, allele differences, and sex-biased gene expression within the discovered sex-linked regions provides no evidence that sexual antagonism has driven the repeated evolution of a region on LG5 that overlaps between four of these species. By comparing the sex-determining regions on LG5 and LG7 across haplochromines, we show that a common origin is unlikely, and that while sex chromosomes themselves may be shared between several Haplochromini, the sex-determining genes or mechanism likely differ. This study paves the way to explore newly emerging theories of sex chromosome evolution, such as the role of chromosomal fusion or recombination patterns across the genome.

Animals

Use of a dense single nucleotide polymorphism map for in silico mapping in the mouse.

Rapid expansion of available data, both phenotypic and genotypic, for multiple strains of mice has enabled the development of new methods to interrogate the mouse genome for functional genetic perturbations. In silico mapping provides an expedient way to associate the natural diversity of phenotypic traits with ancestrally inherited polymorphisms for the purpose of dissecting genetic traits. In mouse, the current single nucleotide polymorphism (SNP) data have lacked the density across the genome and coverage of enough strains to properly achieve this goal. To remedy this, 470,407 allele calls were produced for 10,990 evenly spaced SNP loci across 48 inbred mouse strains. Use of the SNP set with statistical models that considered unique patterns within blocks of three SNPs as an inferred haplotype could successfully map known single gene traits and a cloned quantitative trait gene. Application of this method to high-density lipoprotein and gallstone phenotypes reproduced previously characterized quantitative trait loci (QTL). The inferred haplotype data also facilitates the refinement of QTL regions such that candidate genes can be more easily identified and characterized as shown for adenylate cyclase 7.

Adenylyl Cyclases

Should neighbours of tuberculosis (TB) cases be prioritised for active case finding in high TB-burden settings? A prospective molecular epidemiological study.

INTRODUCTION: In high tuberculosis (TB)-burden countries, considerable transmission of Mycobacterium tuberculosis (M. tb) likely occurs outside of households. We aimed to estimate the TB prevalence and incidence in households and neighbourhoods around known TB cases and to understand transmission patterns. METHODS: Household and neighbourhood contacts of pulmonary TB index cases from contiguous areas in Bandung, Indonesia, were screened and followed up for 12 months. Sputum samples underwent smear microscopy, M. tb culture, Xpert MTB/RIF, DNA isolation and whole-genome sequencing (WGS). Pairwise single-nucleotide polymorphism (SNP) distance ≤12 defined transmission for pairs with known epidemiological links, or SNP≤3 for pairs without epidemiological link. An SNP=12 cut-off was used to characterise transmission clusters. RESULTS: From 213 index cases, 514 household and 4141 neighbourhood contacts underwent TB screening: 19 household (3.70%, 95% CI 2.24 to 5.71) and 45 neighbourhood (1.09%, 95% CI 0.79 to 1.45) contacts were identified with TB, of whom 18 (3.50%, 95% CI 2.20 to 5.48) and 38 (0.92%, 95% CI 0.65 to 1.13) respectively, were bacteriologically confirmed. During follow-up, 11 household and 13 neighbourhood contacts were identified with TB (incidence per 100 000 person-years: 2286 (95% CI 1286 to 4148) and 350 (95% CI 190 to 563)), of whom 6 and 8, respectively, were bacteriologically confirmed (incidence per 100 000 person-years: 1247 (95% CI 560 to 2776) and 201 (95% CI 101 to 402)). A total of 223 patient M. tb isolates underwent WGS. Of 15 intra-household pairs, 8 (53.3%) were transmission pairs. Of 24 neighbour to index case pairs, 1 (4.2%) was a transmission pair. 11 of 19 transmission pairs shared no epidemiological link. We identified 25 M. tb genetic clusters from 205 mono-TB isolates overall. CONCLUSION: Neighbours have lower prevalence and incidence of TB than household contacts, but twice as many cases. Very few received M. tb from their index case, suggesting uncontrolled community-wide transmission. Whole population active case finding may be necessary in high TB-burden settings.

Humans

Two distinct mechanisms alter p53 in breast cancer: mutation and nuclear exclusion.

Twenty-seven cases of inflammatory breast cancer were screened for the presence of the p53 protein by immunocytochemical methods using a monoclonal antibody directed against the p53 protein. Three groups were detected: 8 cases (30%) had high levels of p53 in the nucleus of the cancer cells; 9 cases (33%) had a complete lack of detectable staining; 10 cases (37%) showed a pattern of cytoplasmic staining with nuclear sparing. Nucleotide sequence analysis of p53 cDNAs derived from the samples with cytoplasmic staining revealed only wild-type p53 alleles in 6 out of 7 cases. An eighth case was determined to be wild type by a single-strand conformation polymorphism. In contrast, the samples containing nuclear p53 contained a variety of missense mutations and a nonsense mutation. The p53 cDNAs from 3 of the tumors that lacked detectable p53 staining were analyzed, and all 3 had wild-type nucleotide sequences. Interestingly, a case of normal lactating breast tissue also showed intense cytoplasmic staining for p53 with nuclear sparing. These data suggest that some breast cancers that contain the wild-type form of p53 protein may inactivate its tumor-suppressing activity by sequestering this protein in the cytoplasm, away from its site of action in the cell nucleus. The detection of cytoplasmic p53 in normal lactating breast tissue could suggest that this is the mechanism employed in specific physiological situations to permit transient cell proliferation. This observation could explain how some breast cancer tissues inactivate p53 function without mutation.

Amino Acid Sequence

Fine-Scale Landscape Genomics Show Asymmetric Patterns of Gene Flow for the Invasive Mosquito Aedes albopictus.

Mosquito-borne viruses like dengue, Zika, and chikungunya pose increasing health risks in the United States due to the expanding range of Aedes albopictus, a highly invasive mosquito species that now has a global distribution. Aedes albopictus thrive in artificial containers associated with anthropogenic land use, allowing populations to reach high numbers in urban and suburban environments. While the global spread of Ae. albopictus has been well characterized, the effects of heterogeneous urban landscapes on dispersal and gene flow at fine spatial scales remain unclear. This study analyzed the genetic connectivity of Aedes albopictus populations collected in Wake County, North Carolina in 2018. We used single nucleotide polymorphisms (SNP) data from double-digest restriction-enzyme associated DNA sequencing (ddRADseq) and examined genetic connectivity through principal component analysis (PCA) and genetic network analysis. We then evaluated migration and source-sink dynamics using a Bayesian approach for SNP data (BA3-SNP). We found little evidence of genetic clustering or isolated populations of Ae. albopictus in Wake County, suggesting high gene flow between sites. Migration analysis demonstrated asymmetric gene flow from rural to urban regions within Wake County, with greater gene flow occurring between and within urban regions. These findings suggest that the pattern of gene flow of Ae. albopictus populations within local metropolitan areas may involve urban city centers serving as genetic sinks and surrounding suburban and rural regions serving as sources. This study highlights how heterogeneous landscapes shape mosquito population connectivity and migration at fine spatial scales, which is critical for informing vector control and public health intervention strategies.

Aedes albopictus

Shared candidate genes associated with variation in egg size in cold-adapted and artificially selected Drosophila melanogaster.

The development of most multicellular organisms begins with oogenesis, the production of the egg. In D. melanogaster, egg size is a highly polygenic trait closely related to fitness. Elements of shifts in egg size have been widely studied and modeled, but the genes underlying this variation are still poorly understood. This study aimed to identify candidate genes associated with processes underlying egg-size variation using D. melanogaster as a model. In selection experiments, we generated large-egg populations from a shared ancestral population using both cold-adaptation and artificial selection, and identified candidate genes for the large-egg phenotype. Using whole-genome DNA sequencing and strict computational filtering, we uncovered single-nucleotide polymorphisms in 10 genes. Characterization of these candidates revealed functions in cytoskeletal dynamics, DNA replication and repair, intracellular signaling, and stem cell maintenance and differentiation. RT-PCR and qPCR were used to validate gene expression differences between cold-adapted lines and the Oregon R control (OrR) in a subset of the candidates. In RT-PCR, stathmin demonstrated a modified expression pattern in all cold-adapted lines relative to OrR controls. In qPCR experiments, Pde1c had significantly higher expression (p&#x202f;<&#x202f;0.05) in the cold-adapted flies compared to OrR controls for all three fly cages tested. For Ino80, significantly higher expression was observed for one of three cages while one cage showed lower expression. We have assembled a candidate list we hope will be a useful resource for researchers across specialties, from germ cells to cytoskeletal dynamics, to further investigate the genetic and developmental aspects of variation in egg size in D. melanogaster.

Animals

Dissemination of blaKPC-3-harbouring Klebsiella pneumoniae across ST48 and ST628 in multiple healthcare facilities in the Republic of Korea.

Klebsiella pneumoniae carbapenemase-3 (KPC-3) remains rare in South Korea, where KPC-2 is the dominant carbapenemase, making the repeated detection of a concentrated blaKPC-3 signal over five years notable. We performed genomic analyses of blaKPC-3-harbouring K. pneumoniae from a regional healthcare network. Two chromosomally distinct lineages with concordant capsule loci (ST628/KL15 and ST48/KL62) presented multidrug-resistant phenotypes, and the virulence-associated loci were confined to ST48. Single-nucleotide polymorphism (SNP) analyses revealed near-clonal relatedness within lineages, with 0-38 pairwise SNPs among ST628 isolates and 8 SNPs between the two ST48 isolates. Core-genome multilocus sequence typing (cgMLST) supported this structure, as ST628 isolates were assigned to complex type 19149 with 0-7 allelic differences, and ST48 isolates were assigned to complex type 19150 with 5 allelic differences. These patterns support vertical spread via clonal expansion across multiple facilities. Despite substantial chromosomal separation, most isolates carried the same IncFII(K) plasmid backbone and blaKPC-3, and they were nearly indistinguishable from a plasmid previously reported in South Korea. One isolate carried blaKPC-3 on a distinct multireplicon IncFIB(K)/IncFII(K) plasmid, indicating that the signal was not confined to a single plasmid backbone. In both plasmids, blaKPC-3 was embedded within Tn4401b. These findings indicate that a rare blaKPC-3 genotype can persist regionally through sustained clonal dissemination and that cross-lineage linkage is compatible with past horizontal transfer involving a conserved plasmid. These findings underscore the need for subtype-resolved, regionally coordinated genomic surveillance in connected healthcare networks to detect uncommon carbapenemase variants early.

Klebsiella pneumoniae

Multi-omics identification of ZFPM2 and CD44 as key candidates for testicular size in sheep.

Testicular size is a key determinant of ram fertility, yet its genetic architecture in sheep remains poorly understood. Crossbred sheep exhibit substantial testicular developmental variation and serve as ideal models for screening fertility-related genes. Here, we performed whole-genome sequencing (WGS)-based genome-wide association study (GWAS) on 115 rams, including seven crossbred populations (each derived from a distinct sire breed crossed with Hu sheep) and a purebred Hu sheep population. We identified two variants within ZFPM2, including an intronic single nucleotide polymorphism (SNP) rs421073404 and a missense SNP rs1086332841, both significantly correlated with testis weight, length, and width. Importantly, these variants exerted testis-specific effects, consistent with their weak correlation with body weight (r&#x202f;<&#x202f;0.2). Genomic selection scans (FST and &#x3c0;-ratio) between rams with large (>150&#x202f;g/side) and small (<75&#x202f;g/side) testes revealed multiple divergent genomic regions, with prominent haplotype differentiation at the CD44 locus on chromosome 15. Notably, two linked missense variants (rs160734053 and rs414318703) in CD44 exerted opposing effects on testicular size, suggesting allelic heterogeneity at this locus. Integrative RNA-seq and ATAC-seq across 0-12 months further uncovered prepubertal stage-specific expression and chromatin accessibility patterns of ZFPM2 and CD44, providing mechanistic insights into their regulatory roles. Collectively, this study provides candidate SNPs for sheep marker-assisted selection and multi-omic evidence for dissecting the genetic basis of testicular development in ovine breeding.

Animals

Dual-dimensional profiling of host genomic variations and HPV integration in PD-L1-stratified cervical cancer via Oxford Nanopore Technology.

BACKGROUND: The integration of human papillomavirus (HPV) DNA into the host genome is a key step in the development of HPV-associated cervical cancer (CC). However, the genomic characteristics of host genomic variations and HPV integration within the context of programmed death-ligand 1 (PD-L1) expression stratification have not been systematically investigated. METHODS: Whole-genome sequencing was performed using Oxford Nanopore Technology (ONT) on six samples (three from the high PD-L1 expression group and three from the low PD-L1 expression group). The characteristics of host genomic variations under different PD-L1 expression stratifications were explored, including structural variations (SV), copy number variations (CNV), single nucleotide polymorphisms (SNP), and insertion-deletions (Indel). Subsequently, the distribution features of HPV integration sites were analyzed, different integration types were identified, and pathway analysis was conducted. RESULTS: Whole-genome SV analysis revealed that the total number of SVs and the composition of mutation types were similar between the high and low PD-L1 expression groups, with insertions (INS) and deletions (DEL) predominating in both. These variations were primarily enriched in intergenic regions and introns. In the low PD-L1 expression group, integration events were observed at multiple chromosomal loci, with the most frequent integration occurring in the KLF5 gene region on chromosome 13. No frequently integrated loci were identified in the high PD-L1 expression group. Additionally, four distinct HPV integration breakpoint patterns were preliminarily identified and analyzed. CONCLUSION: PD-L1 expression stratification did not significantly alter the overall genomic instability of the host. However, differences were observed in the distribution patterns of HPV integration sites. These findings provide new insights into the genomic heterogeneity of CC under different PD-L1 expression backgrounds and may lay the groundwork for future research exploring stratified immunotherapy based on HPV integration features.

Humans

Patterns of naturally occurring restriction map variation, dopa decarboxylase activity variation and linkage disequilibrium in the Ddc gene region of Drosophila melanogaster.

Forty-six second-chromosome lines of Drosophila melanogaster isolated from five natural populations were surveyed for restriction map variation in a 65-kb region surrounding the gene (Ddc) encoding dopa decarboxylase (DDC). Sixty-nine restriction sites were scored, 13 of which were polymorphic. Average heterozygosity per nucleotide was estimated to be 0.005. Eight large (0.7-5.0 kb) inserts, two small inserts (100 and 200 bp) and three small deletions (100-300 bp) were also observed across the 65-kb region. We see no evidence for a reduction in either nucleotide heterozygosity or insertion/deletion variation in the central 26-kb segment containing Ddc and a dense cluster of lethal complementation groups and transcripts (greater than or equal to 9 genes) compared to that seen in the adjacent regions (totaling 39 kb) in which only a single gene and transcript has been detected, or to that observed for other gene regions in D. melanogaster. The distribution of restriction site variation shows no significant departure from that expected under an equilibrium neutral model. However insertions and deletions show a significant departure from neutrality in that they are too rare in frequency, consistent with them being deleterious on average. Significant linkage disequilibrium among variants exists across much of the 65-kb region. Lower regional rates of recombination combined with the influence of polymorphic chromosomal inversions, rather than epistatic selection among genes in the dense cluster, probably are sufficient explanations for the creation and/or maintenance of the linkage disequilibrium observed in the Ddc region. We have also assayed adult DDC enzyme activity in these same lines. Twofold variation in activity among lines is observed within our sample. Significant associations are observed between level of DDC enzyme activity and restriction map variants. Surprisingly, one line with a 5.0-kb insert within an intron and one line with a 1.5-kb insert near the 5' end of Ddc each show normal adult DDC activities.

Animals

Nested Admixture During and After the Trans-Atlantic Slave Trade on the Island of S&#xe3;o Tom&#xe9;.

Human genetic admixture, involving the contact between two or more previously isolated populations, can be a complex process influenced by social dynamics. In this study, we aim to reconstruct complex admixture histories in S&#xe3;o Tom&#xe9;, an island in the Gulf of Guinea where the Portuguese established one of the first plantation-based slave societies. Since the 15th century, migration waves from Africa and Europe, slavery, marooning, and indentured labour led to profound demographic shifts and social stratification on the island. Examining 2.5 million SNPs newly genotyped in 96 S&#xe3;o Tom&#xe9;ans, we observed patterns of genetic differentiation that were more complex than those of other populations descended from enslaved Africans on either side of the Atlantic. Using local ancestry inference and Identical-by-Descent methods, we identified five genetic clusters in S&#xe3;o Tom&#xe9; and reconstructed shared ancestries between each cluster and 70 African and European population samples, including an extensive sample from the Cabo Verde archipelago. Our findings align with historical records, retracing the major slave trade routes and labour-driven migrations after the abolition of slavery. We also identified gene flow between recently admixed groups that were previously isolated on the island. We call this process, creating multiple layers of genetic ancestry in admixed genomes, nested admixture. We suggest that changing social structures in S&#xe3;o Tom&#xe9; transformed the genetic structure of its population and influenced the admixture process. This study demonstrates how successive admixture and isolation events during and after the Trans-Atlantic Slave Trade shaped extant genetic diversity patterns at local scale in Africa.

Humans

Concordance and divergence between self-declared ancestry and genome-derived ancestry composition in 10&#x2009;250 participants from the HostSeq cohort.

Accurate characterization of human genetic diversity is essential for robust genomic analyses. We compared self-declared and genome-derived ancestry composition in 10&#x2009;250 participants from the pan-Canadian HostSeq cohort using whole-genome sequencing data. Global and local ancestry were inferred at the continental super-population level using the alignment-free ntRoot algorithm and evaluated through both hard-label concordance and multiclass Brier score analyses incorporating full ancestry fraction profiles. Strong agreement was observed among East Asian / Pacific Islander (mean Brier score&#xa0;&#xb1;&#xa0;SD: 0.012&#xa0;&#xb1;&#xa0;0.052), Black (0.013&#xa0;&#xb1;&#xa0;0.042), White (0.055&#xa0;&#xb1;&#xa0;0.022), and South Asian (0.057&#xa0;&#xb1;&#xa0;0.098) participants, whereas higher scores among Hispanic (0.083&#xa0;&#xb1;&#xa0;0.060) and Middle Eastern or Central Asian (0.122&#xa0;&#xb1;&#xa0;0.034) participants reflected broader and more admixed ancestry profiles. Principal component analysis of centered log-ratio-transformed ancestry fractions revealed overlapping ancestry gradients rather than discrete continental groupings. Entropy- and dominance margin-based analyses further indicated that many discordant cases reflected diffuse admixture rather than categorical mismatch. Together, these findings support representing ancestry as a continuous compositional spectrum rather than discrete categories. Genome-derived ancestry estimates describe patterns of genomic variation and should not be interpreted as proxies for race.

Humans

Lineage-specific transmission and spatial clustering of Mycobacterium tuberculosis in Kaohsiung, Taiwan, in 2019-23: a population-based genomic study.

BACKGROUND: The epidemiology of tuberculosis in Taiwan has been influenced by the introduction of multiple Mycobacterium tuberculosis lineages and by the ageing of the population. We conducted a population-based study to investigate M tuberculosis transmission in Kaohsiung, a city in southern Taiwan. METHODS: In this study, we performed whole-genome sequencing (WGS) of M tuberculosis isolates from all culture-positive cases of tuberculosis notified in Kaohsiung between Jan 1, 2019 and Dec 31, 2023. We obtained routine epidemiological data for each case collected through the national tuberculosis control programme. We characterised the lineage composition of the isolate collection and evaluated genomic clustering of isolates, defined as a difference of 12 or fewer single-nucleotide polymorphisms. Univariable and multivariable logistic regression analyses were performed to estimate the odds of a case belonging to a genomic cluster based on host factors (age, sex, sputum smear status, and residential region) and pathogen factors (drug resistance status and strain lineage). Spatial aggregation of large genomic clusters (including greater than or equal to ten isolates) was assessed using a non-parametric statistical clustering method. We used a Bayesian transmission tree inference method to explore the patterns of age-dependent transmission. FINDINGS: During the study period, 5667 tuberculosis cases were notified in Kaohsiung, 4916 (86&#xb7;7%) of which were culture-positive. Of these 4916 cases, whole-genome sequencing was successfully performed for 4168 (84&#xb7;8%) isolates. 1219 (29&#xb7;2%) of 4168 individuals were female and 2947 (70&#xb7;7%) were male; the median age was 69&#xb7;7 years (IQR 57&#xb7;4-80&#xb7;7). The dominant lineages were lineage 1 (1749 [42&#xb7;0%] of 4168 isolates), lineage 2 (1510 [36&#xb7;2%]), and lineage 4 (905 [21&#xb7;7%]). 1069 (25&#xb7;6%) of 4168 were genomically linked and formed 287 clusters. Lineage 2 isolates had higher odds (aOR 2&#xb7;15 [95% CI 1&#xb7;80-2&#xb7;52]) than lineage 1 isolates of genomic clustering across all regions, whereas lineage 4 isolates had a significantly higher risk (2&#xb7;75 [1&#xb7;16-6&#xb7;89]) of genomic clustering than lineage 1 only in the rural northeast region, inhabited primarily by indigenous populations. Spatial clustering analysis corroborated these lineage-region interactions. Although younger adults (<35 years) had the highest individual-level odds (5&#xb7;64 [4&#xb7;16-7&#xb7;68]) of clustering in the logistic regression analysis compared with those aged 80 years or older, the transmission inference indicated that individuals aged 55-74 years were responsible for a greater proportion of inferred transmission events, contributing 50&#xb7;8% of all transmission events. INTERPRETATION: This sequencing study revealed that older adults (aged &#x2265;65 years) might have played a substantial and under-recognised role in the transmission of tuberculosis in Taiwan. The lineage-specific clustering and spatial patterns suggested that both pathogen characteristics and host demographics shaped tuberculosis transmission dynamics. These findings support the use of integrated genomic surveillance to guide precision tuberculosis control and motivate further research on age-specific transmission pathways and targeted interventions to advance tuberculosis elimination efforts. FUNDING: Taiwan National Health Research Institutes and Taiwan National Science and Technology Council.

Mycobacterium tuberculosis