Search PubMedSearch

SEARCH · Search PubMed

Results for “Principal component”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

WinPCA: a package for windowed principal component analysis.

SUMMARY: With chromosomal reference genomes and population-scale whole genome-sequencing becoming increasingly accessible, contemporary studies often include characterizations of the genomic landscape as it varies along chromosomes, commonly termed genome scans. While traditional summary statistics like FST and dXY between pre-assigned populations remain integral to characterizing the genomic divergence profile, PCA differs by providing single-sample resolution, thereby supporting the identification of polymorphic inversions, introgression and other types of divergent sequence that may not be fully aligned with global population structure. Here, we introduce WinPCA, a user-friendly package to compute, polarize and visualize genetic principal components in windows along the genome. To accommodate low-coverage whole genome-sequencing datasets, WinPCA can optionally make use of PCAngsd methods to compute principal components in a genotype likelihood framework. WinPCA accepts variant data in either VCF or BEAGLE format and can generate rich plots for interactive data exploration and downstream presentation. AVAILABILITY AND IMPLEMENTATION: WinPCA is implemented in Python and freely available at https://github.com/MoritzBlumer/winpca and https://doi.org/10.5281/zenodo.15614979.

Software

Matching Heterogeneous Cohorts by Projected Principal Components Reveals Two Novel Alzheimer's Disease-Associated Genes in the Hispanic Population.

Alzheimer's disease (AD) is the most common form of dementia in elderly, affecting 6.9 million individuals in the United States. Some studies have suggested the prevalence of AD is greater in individuals who self-identify as Hispanic. Focused results are relevant for personalized and equitable clinical interventions. Ethnicity as a stratifying tool in genetic studies is often accompanied by genomic inflation due to heterogeneity. In this study, we report GWAS and meta-analyses conducted among NIAGADS subjects who self-identified as Hispanic and All of Us (AoU) sub-cohorts matched to that cohort, using projected genetically-derived principal components, with and without age and sex. In Hispanic NIAGADS subjects, we identified a common variant in PIEZO2 that was protective for AD with a p-value just beyond genome-wide significance (p = 5.4*10-8). Meta-analyses with genetically-matched AoU participants yielded three (two novel) genome-wide significant AD-associated loci based on rare lead variants: rs374043832 (RGS6/PSEN1), rs192423465 (ASPSCR1), and rs935208076 (GDAP2), which were also nominally significant in AoU sub-cohorts. We thus demonstrate an efficient way to select subjects from large heterogeneous biobank cohorts who are genetically similar to a smaller disease-specific cohort, yielding novel disease-relevant findings.

Journal Article

Serum Proteomic Profiling Implicates a Dysregulated Neurohormonal-Inflammatory Axis in Post-Fontan Sinus Tachycardia.

BACKGROUND: Postoperative sinus tachycardia is a poorly understood complication following the Fontan procedure. The molecular signaling cascades triggering acute tachycardia remain uncharacterized, limiting therapeutic innovation. Here, we present a retrospective study leveraging serum proteomics and machine learning to identify the molecular drivers of postoperative Fontan sinus tachycardia. METHODS: We integrated a clinically relevant ovine Fontan model with continuous telemetric heart rate monitoring and human patient data. Serum proteomics coupled with least absolute shrinkage and selection operator and Boruta machine learning algorithms were used to identify protein panels predictive of postoperative sinus tachycardia. Cross-species validation was performed by comparing proteomic signatures from sheep and pediatric patients undergoing Glenn or Fontan surgery. RESULTS: Ovine Fontan animals demonstrated significant heart rate elevation beginning on postoperative day 1, peaking at postoperative day 3 (159.4±11.7 bpm versus preoperative, 105.3±10.5 bpm; P=0.0002), before trending toward baseline by postoperative day 10. This pattern was mirrored in human patients with a more modest magnitude. Surgical controls did not exhibit tachycardia. The principal component most correlated with heart rate (principal component 1: r=0.78, P=2.2×10-4) was enriched for inflammatory and neural pathways. The Boruta algorithm identified an 11-protein panel with strong predictive power (area under the receiver operating characteristic curve, 0.963). Cross-species comparison demonstrated that angiotensinogen, angiotensin-converting enzyme, and pentraxin 3 were similarly dysregulated in both species postoperatively. CONCLUSIONS: This study provides molecular evidence implicating a dysregulated neurohormonal-inflammatory axis in acute postoperative Fontan sinus tachycardia and establishes a foundation for developing targeted diagnostics and therapeutics for this complication.

Animals

Genome-Wide SNP Characterisation of Three Kazakh Sheep Breeds: Kazakh Fat-Tailed Coarse-Wool, Degeres, and Etti Merino.

Kazakhstan's sheep portfolio underpins much of the country's mutton and wool production, yet several of its principal breeds remain genomically uncharacterised. The aim of this study was to characterise the genomic diversity, population structure, and global phylogenetic placement of three economically important Kazakh breeds and to determine whether they constitute separate gene pools requiring independent management. We present the first genome-wide SNP characterisation to include the Degeres (DE), the Etti Merino (EM), and the Kazakh fat-tailed coarse-wool (KKG) breeds simultaneously. A total of 1497 animals (DE = 354, EM = 642, KKG = 501) sampled across seven production households were genotyped and, after quality control, analysed at 42,279 SNPs, of which 22,766 LD-pruned markers were used for principal component analysis and AMOVA. We applied principal component analysis (PCA), pairwise FST, analysis of molecular variance (AMOVA), neighbour-joining phylogenetics, model-based ancestry estimation (ADMIXTURE), and Hill-number diversity profiling, and projected the breeds against the global Ovine SNP50 HapMap panel (74 reference breeds, 2819 animals; 37,685 shared SNPs). All three breeds retained uniformly high within-breed diversity (expected heterozygosity 0.413-0.417) with fixation indices at or near zero. AMOVA partitioned 94.03% of variance within breeds (&#x3a6;ST = 0.060, p < 0.001). PCA, phylogeny, and ADMIXTURE concordantly resolved three breed-specific clusters at K = 3, with a maximum interbreed FST of 0.038 within the study dataset. Against the global panel, EM was genetically closest to Merino and Merino-derived reference breeds (pooled FST = 0.017) and substantially more distant from Southwest Asian sheep (FST = 0.045), whereas DE and KKG showed the reciprocal pattern (FST = 0.027 and 0.020 to Southwest Asia, 0.052 to the Merino group). DE additionally displayed the heterozygote excess and partial admixture expected of an incompletely consolidated composite. These results delineate three distinct gene pools and carry direct implications for breed management and the conservation of genomic diversity in Kazakhstani sheep.

ADMIXTURE

jsPCA: fast, scalable, and interpretable identification of spatial domains and variable genes across multi-slice and multi-sample spatial transcriptomics data.

MOTIVATION: Spatial transcriptomics technologies record genome-wide measurements of gene expression with high spatial resolution. These technologies generate large and high-dimensional datasets requiring efficient automated methods for their analysis. We introduce joint spatial PCA (jsPCA), a novel, fast, scalable and interpretable method for the automatic identification of spatial domains and variable genes in multi-slice and multi-sample spatial transcriptomics data. RESULTS: jsPCA relies on a simple mathematical formulation of a spatial covariance defined as the product of the gene expression covariance with the spatial autocorrelation. The principal components of this spatial covariance yield a biologically meaningful low-dimensional representation. From this representation, spatial domains are derived by simple clustering and spatially variable genes are identified directly from the principal component coefficients. A joint representation of multiple slices and samples without spatial alignment is obtained by computing common principal components via joint diagonalization. By leveraging data sparsity and non-convex manifold optimization, jsPCA leads to computing time in the order of seconds to minutes, substantially outperforming state-of-the-art approaches. We benchmarked jsPCA against 10 state-of-the-art methods on two reference databases. Our approach demonstrated excellent performance, comparable or better than state-of-the-art methods, while being much faster, interpretable, and scalable to very large datasets.

Journal Article

Genomic analysis of breed composition and population structure in Montana composite cattle.

The Montana composite was developed in Brazil from crosses between Bos indicus and Bos taurus and structured into four biological types: Zebu (N), adapted taurine (A), British taurine (B), and continental taurine (C). This study aimed to characterize the genetic diversity and population structure of the Montana composite using genomic data through principal component analysis (PCA), admixture analysis, and Wright's FST statistic. The PCA revealed a clear separation between Bos indicus and Bos taurus groups, with Montana animals distributed in an intermediate position. The first two principal components explained 69.48% and 3.45% of the total variation, respectively. Supervised admixture estimates indicated a predominance of taurine contribution, with type A accounting for 34.47%, 52.64%, and 51.71% at K&#x2009;=&#x2009;4, 9, and 11, respectively. Increasing the ancestry resolution refined the contribution of individual founder breeds without changing the overall predominance of taurine ancestry. Comparisons between breed proportions obtained from pedigree and genomic data revealed significant differences, for most biological types and ancestry models (P&#x2009;<&#x2009;0.001), indicating that realized breed composition deviates from theoretical expectations. Estimates of genetic differentiation confirmed greater divergence between Zebu and taurine groups, as well as reduced distances among populations sharing common ancestry. Specific relationships were identified between the composite and some of its founder breeds, particularly Belmont Red, Senepol, and Tuli. Overall, the results demonstrate that the Montana composite has a complex genomic structure, with genomic ancestry varying according to the resolution adopted and differing from pedigree-based expectations.

Animals

Multi-level aggregation analysis of microbiome composition and host gene expression reveals associations with systemic and local immunity.

The human gut microbiome plays a critical role in immune regulation, yet the molecular links between microbiome composition and host gene expression remain incompletely understood. We analyzed associations between host gene expression and microbiome composition in a cohort of 315 healthy individuals, integrating microarray-based gene expression data from three intestinal sites (ileum, transverse colon, and rectum) and six immune cell types with microbiome sequencing data. Using a hierarchical feature aggregation strategy combining principal component analysis, clustering, and covariate correction, we discovered significant associations primarily related to immunity. While microbial profiles were similar across the three intestinal sites, the transverse colon yielded the most "microbiome-host gene expression" associations. Among the immune cell types, CD8+ cells showed the highest number of associations. The first principal component of microbiome composition, reflecting a gradient from commensals (e.g., Ruminococcaceae and Christensenellaceae) to proinflammatory taxa ([Ruminococcus] gnavus and Lachnoclostridium), correlated with the expression of TNF-&#x3b1;-linked genes (HMOX1, CPI17, HSD3B2, and SLC5A1). Among individual genera, Catenibacterium abundance was associated with gene expression in both intestinal and immune cells, including negative associations with MRPS21 (related to mitochondrial function) in the transverse colon and with CD8+ gene programs related to T cell differentiation. These findings align with emerging evidence implicating mitochondrial dysfunction in intestinal inflammation. Our results identify multi-level associations between the gut microbiome and host gene expression, suggesting potential mechanisms by which microbiota shape local and systemic immunity and vice versa. The implicated genes and taxa represent candidates for experimental validation to improve understanding of host-microbiome homeostasis and its disruption in disease.IMPORTANCEThe gut microbiome and immune system are engaged in a complex interplay throughout human life. While most associative studies focus on case-control comparisons-typically examining patients with conditions such as inflammatory bowel disease or metabolic diseases-less is known about the molecular links between the microbiome and immune system in healthy individuals. In this study of a large cohort of healthy individuals, we addressed this gap by applying multiscale modeling to tackle the high dimensionality of host-microbiome data. We identified multi-level associations between microbiome composition and host gene expression in both intestinal tissues and immune cells. These findings offer a valuable reference for understanding baseline host-microbiome communication and highlight molecular candidates-such as TNF-&#x3b1;-related genes and mitochondrial pathways-for future experimental validation.

Humans

A SuperLearner-based pipeline for the development of DNA methylation-derived predictors of phenotypic traits.

BACKGROUND: DNA methylation (DNAm) provides a window to characterize the impacts of environmental exposures and the biological aging process. Epigenetic clocks are often trained on DNAm using penalized regression of CpG sites, but recent evidence suggests potential benefits of training epigenetic predictors on principal components. METHODOLOGY/FINDINGS: We developed a pipeline to simultaneously train three epigenetic predictors; a traditional CpG Clock, a PCA Clock, and a SuperLearner PCA Clock (SL PCA). We gathered publicly available DNAm datasets to generate i) a novel childhood epigenetic clock, ii) a reconstructed Hannum adult blood clock, and iii) as a proof of concept, a predictor of polybrominated biphenyl exposure using the three developmental methodologies. We used correlation coefficients and median absolute error to assess fit between predicted and observed measures, as well as agreement between duplicates. The SL PCA clocks improved fit with observed phenotypes relative to the PCA clocks or CpG clocks across several datasets. We found evidence for higher agreement between duplicate samples run on alternate DNAm arrays when using SL PCA clocks relative to traditional methods. Analyses examining associations between relevant exposures and epigenetic age acceleration (EAA) produced more precise effect estimates when using predictions derived from SL PCA clocks. CONCLUSIONS: We introduce a novel method for the development of DNAm-based predictors that combines the improved reliability conferred by training on principal components with advanced ensemble-based machine learning. Coupling SuperLearner with PCA in the predictor development process may be especially relevant for studies with longitudinal designs utilizing multiple array types, as well as for the development of predictors of more complex phenotypic traits.

DNA Methylation

Host clustering of Campylobacter species and enteric pathogens in a longitudinal cohort of infants, family members and livestock in rural Eastern Ethiopia.

BACKGROUND: Livestock are recognized as major reservoirs for Campylobacter species and other enteric pathogens, posing infection risks to humans. High prevalence of Campylobacter during early childhood has been linked to environmental enteric dysfunction and stunting, particularly in low-resource settings. METHODS: A total of 280 samples from Campylobacter positive households with complete metadata were analyzed by shotgun metagenomic sequencing followed by bioinformatic analysis via the CZ-ID metagenomic pipeline (Illumina mNGS Pipeline v7.1). Further statistical analyses in JMP PRO 16 explored the microbiome, emphasizing Campylobacter and other enteric pathogens. Two-way hierarchical clustering and split k-mer analysis examined host structuring, patterns of co-infections and genetic relationships. Principal component analysis was used to characterize microbiome composition across the seven sample types. RESULTS: The study identified that microbiome composition was strongly host-driven, with more than 3844 genera detected, and two principal components explaining 62% of the total variation. Twenty-one dominant (based on relative abundance) Campylobacter species showed distinct clustering patterns for humans, ruminants, and broad hosts. The broad-host cluster included the most prevalent species, C. jejuni, C. concisus, and C. coli, present across sample types&#xa0;and a sub-cluster within C. jejuni involving humans, chickens, and ruminants. Campylobacter species from chickens showed strong positive correlations with mothers (r&#x2009;=&#x2009;0.76), siblings (r&#x2009;=&#x2009;0.61) and infants (r&#x2009;=&#x2009;0.54), while co-occurrence analysis found a higher likelihood (Pr&#x2009;>&#x2009;0.5) of pairs such as C. jejuni with C. coli, C. concisus, and C. showae. Analysis of the top 50 most abundant microbial taxa showed a distinct cluster uniquely present in human stool and absent in all livestock. The study also found frequent co-occurrence of C. jejuni with other enteric pathogens such as Salmonella, and Shigella, particularly in human and chicken. Additionally, instances of Candidatus Campylobacter infans (C. infans) were identified co-occurring with Salmonella and Shigella species in stool samples from infants, mothers, and siblings. CONCLUSIONS: A comprehensive analysis of Campylobacter diversity in humans and livestock in a low-resource setting revealed that infants can be exposed to multiple Campylobacter species early in life. C. jejuni is the dominant species with a propensity for co-occurrence with other notable enteric bacterial pathogens, including Salmonella, and Shigella, especially among infants. Video Abstract.

Animals

Pilot study identifying distinct circulating proteomic profiles associated with longitudinal CT-defined fibrotic and inflammatory sarcoidosis.

INTRODUCTION: Pulmonary sarcoidosis exhibits heterogeneous clinical trajectories ranging from self-limited disease resolution to chronic progressive fibrosis, yet reliable biomarkers capable of distinguishing these disease patterns remain lacking. Whether longitudinal CT-defined sarcoidosis phenotypes are associated with distinct circulating molecular signatures remains unknown. METHODS: We performed high-throughput plasma proteomics (SomaScan 11K) in participants with pulmonary sarcoidosis classified into longitudinal chest CT-defined progressive fibrosis, progressive nodular inflammatory disease, or resolving disease trajectories, along with healthy controls. CT phenotypes were assigned based on predefined longitudinal changes in reticulation, traction bronchiectasis, nodular involvement, and mediastinal lymphadenopathy across serial CT scans. One plasma sample per participant was selected from the study visit corresponding to the CT time point at which criteria for the assigned longitudinal phenotype were met. Principal component analysis, hierarchical clustering, pathway enrichment, and correlation-based analyses linking protein expression to quantitative CT features were used to evaluate whether distinct longitudinal CT phenotypes were associated with divergent proteomic signatures. RESULTS: Principal component analysis and hierarchical clustering suggested partial segregation by CT-defined phenotype. Longitudinal CT phenotypes were associated with distinct pathway-level proteomic signatures, with progressive fibrosis enriched for epithelial-mesenchymal transition signaling, and progressive nodular inflammatory disease enriched for mTORC1, MYC, oxidative phosphorylation, adipogenesis, and fatty acid metabolism pathways. Correlation analyses showed coordinated protein-expression patterns associated with fibrotic CT features and mediastinal lymph node enlargement. DISCUSSION: These findings suggest that longitudinal CT-defined fibrotic and inflammatory sarcoidosis phenotypes are associated with distinct pathway-level proteomic signatures. This pilot study provides preliminary proof-of-concept evidence that integrating longitudinal CT imaging phenotypes with plasma proteomics may serve as a framework for future mechanistic studies and biomarker discovery in pulmonary sarcoidosis.

Humans

A genome-wide assessment of the population structure of thirteen admixed and pure Australian beef cattle breeds.

Knowledge of population structure is a key factor for successful multi-breed genomic prediction, especially in single-step analysis when metafounders are considered. In Australia, current assessments mostly focus on single breeds using a single-step genomic prediction method. However, the effective integration of pedigree, phenotypic, and genomic data in a multi-breed framework still requires further research, especially for combined analyses including admixed and multi-breed populations. This study began with 602,952 genotyped individuals with 8K SNPs in common from 13 beef cattle breeds (Alexandria, Angus, Brahman, Brangus, Charolais, Droughtmaster, Hereford, Kynuna, Limousin, Santa Gertrudis, Shorthorn, Speckle Park, and Wagyu). Due to different numbers of animals being genotyped in each breed, a representative subset of animals was chosen by employing a validated sampling strategy using Gaussian Mixture Models (GMM) complemented by Principal Component Analysis (PCA) within each breed. Subsequently, a specific number of animals in each cluster were randomly selected to capture the entire genetic diversity per breed, with a total of 260 animals from each breed. The first three principal components explained 59.89% of the total variation, with PC1 (33.54%) clearly separating Bos indicus from Bos taurus lineages. Admixture analysis identified stable ancestral components and defined the genetic makeup of both pure and composite populations. The results showed extensive genetic diversity in some breeds and highlighted distinct genetic differences between Bos indicus and Bos taurus breeds. In addition, six composite breeds' admixture levels confirmed their origin and breed history, revealing a directional shift in ancestry proportions by a longitudinal increase in Brahman ancestry within tropical composites over time. Thus, the findings pave the way for more effective utilization of genetic diversity both within and across populations and provide a framework for designing multi-breed genetic evaluations and breeding programs to improve productivity and profitability in Australian beef production.

Animals

Charting the phenotypic landscape of mitochondrial diseases through a systematic evaluation of pathogenic mitochondrial DNA and nuclear gene variants.

PURPOSE: Primary mitochondrial diseases (PMD) arise from variants in the mitochondrial or nuclear genomes. Phenotype-based recognition of specific PMD genotypes remains difficult, prolonging the diagnostic odyssey. We expanded the MitoPhen database to characterize phenotypic variation across PMD more systematically. METHODS: Individual-level data on mitochondrial DNA disorders, nuclear-encoded mitochondrial diseases, and single large-scale mitochondrial DNA deletions were manually curated with Human Phenotype Ontology (HPO) terms to produce MitoPhen v2. Principal-component analysis summarized system-level abnormalities; HPO-level enrichment and mean phenotype-similarity scores were then used to distinguish common PMD genotypes. RESULTS: MitoPhen v2 adds 3940 individuals to the original release, now encompassing 1597 publications, 10,626 individuals, and 117 genotypes. Among 7586 affected cases, 72,861 HPO terms were recorded. Principal-component analysis revealed 6 phenotype dimensions capturing most system-level variance. At the HPO level, we observed genotype-specific enrichments and identified 111 gene-phenotype links absent from the current HPO database. Using MT-TL1, single large-scale mitochondrial DNA deletions, and POLG as exemplars, phenotype-similarity scores reliably separated individuals with these genotypes from those without. CONCLUSION: MitoPhen v2 enabled systematic, genotype-aware analysis of heterogeneous PMD phenotypes and highlighted the diagnostic value of structured, individual-level data. Phenotype-similarity metrics from such data sets can refine variant interpretation in large rare-disease cohorts and provide a transferable framework for other phenotypically complex genetic disorders.

Humans

[Genetic diversity analysis of Forsythia suspensa germplasm resources in Shanxi based on phenotypic traits and SNP molecular markers].

This study aimed to clarify the degree of fruit phenotypic variation and the characteristics of genetic diversity, population structure, and genetic differentiation of Forsythia suspensa resources in Shanxi, providing an important basis for germplasm conservation and breeding of superior varieties. A total of 46 F. suspensa fruits were collected, and 12 agronomic traits were measured and analyzed. The population genetic structure and genetic diversity of F. suspensa germplasm were evaluated using simplified genome sequencing technology. For the five quality traits of the 46 fruits, the Shannon-Wiener index ranged from 0.631 to 1.074, and the Simpson index ranged from 0.379 to 0.560. The seven quantitative traits exhibited abundant genetic variation, with coefficients of variation ranging from 9.764%(fruit shape index) to 45.494%(forsythin content). Principal component analysis reduced the 12 phenotypic traits to four factors, with a cumulative variance contribution of 74.547%. Sequencing data showed mean Q20 and Q30 values of 98.13% and 94.33%, respectively, with an average GC content of 35.95%. After filtering, a total of 12 347 327 high-quality single nucleotide polymorphism(SNP) loci were obtained. Based on these high-quality SNPs, principal component analysis, population structure analysis, and phylogenetic tree construction were carried out. The 46 germplasm resources were divided into four groups; however, grouping showed little relationship with geographic origin, and intermixing occurred among regions. Mantel test revealed a significant but weak positive correlation between phenotypic and genetic distances(r=0.159, P=0.001). At the molecular level, the four groups exhibited moderate genetic diversity overall, and the genetic differentiation index among populations ranged from 0.027 to 0.084, indicating low to moderate differentiation. The rich genetic diversity of the main phenotypic traits provides a solid material basis for screening superior germplasm and genetic breeding of F. suspensa.

Forsythia

Research on identification of key genes and immune-metabolic mechanisms in atrial fibrillation through integrated multi-cohort transcriptomic analysis and machine learning.

This study aimed to integrate multiple datasets for the identification of atrial fibrillation (AF)-related differentially expressed genes (DEGs), analyze their underlying mechanisms through functional enrichment and machine learning, construct diagnostic models, and explore immune-metabolic interactions to provide novel biomarkers and theoretical foundations. Gene expression datasets were integrated and normalized, with batch effects removed using principal component analysis. Differential expression analysis, functional enrichment analysis (Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathways), and machine learning-based feature gene selection and model construction were performed. Shapley additive explanations analysis was utilized to interpret the constructed models, while gene set enrichment analysis, gene set variation analysis, and immune cell infiltration analysis were conducted to investigate the associations between feature genes and immune infiltration. After integrating and normalizing gene expression data and eliminating batch effects via principal component analysis, 6 DEGs were identified, including 4 upregulated and 2 down-regulated ones. Functional enrichment analysis showed these DEGs were significantly enriched in neuro-related biological processes and pathways, indicating their key roles in AF pathogenesis. Five key feature genes were selected using LASSO, random forest, and support vector machine-recursive feature elimination algorithms. They had significant expression differences between the AF and control groups (P&#x2005;<&#x2005;.001) and were located on distinct chromosomes. The constructed random forest and support vector machine models performed excellently (area under the curve&#x2005;&#x2265;&#x2005;0.85). Shapley additive explanations analysis revealed TNNI1 contributed most to model prediction, with its expression significantly positively correlated with immune cell infiltration. Gene set enrichment analysis and gene set variation analysis analyses further showed feature genes participated in AF pathogenesis by regulating immune modulation, metabolic pathways, and autophagy. Immune cell infiltration analysis found altered proportions of T-cell subsets and M0 macrophages in the AF group, along with complex links between feature gene expression and immune cell function. This study systematically elucidated the unique gene expression patterns and key regulatory pathways associated with AF, clarifying the crucial roles of feature genes in immune regulation, metabolic imbalance, and cellular dysfunction. These findings provide a theoretical basis and potential therapeutic targets for understanding AF pathogenesis and developing targeted treatment strategies.

Atrial Fibrillation

MaxComp: Predicting single-cell chromatin compartments from 3D chromosome structures.

The genome is organized into distinct chromatin compartments with at least two main classes, a transcriptionally active A and an inactive B compartment, broadly corresponding to euchromatin and heterochromatin. Chromatin regions within the same compartment preferentially interact with each other over regions in the opposite compartment. A/B compartments are traditionally identified from ensemble Hi-C contact frequency matrices using principal component analysis of their covariance matrices. However, defining compartments at the single-cell level from sparse single-cell Hi-C data is challenging, especially since homologous copies are often not resolved. To address this, we present MaxComp, an unsupervised method, for inferring single-cell A/B compartments based on 3D geometric considerations in single-cell chromosome structures-derived either from multiplexed FISH-omics imaging or 3D structure models derived from Hi-C data. By representing each 3D chromosome structure as an undirected graph with edge-weights encoding structural information, MaxComp reformulates compartment prediction as a variant of the Max-cut problem, solved using semidefinite graph programming (SPD) to optimally partition the graph into two structural compartments. Our results show that the population average of MaxComp single-cell compartment annotations closely matches those derived from ensemble Hi-C principal component analysis, demonstrating that compartmentalization can be recovered from geometric principles alone, using only the 3D coordinates and nuclear microenvironment of chromatin regions. Our approach reveals widespread cell-to-cell variability in compartment organization, with substantial heterogeneity across genomic loci. When applied to multiplexed FISH imaging data, MaxComp also uncovers relationships between compartment annotations and transcriptional activity at the single-cell level. In summary, MaxComp offers a new framework for understanding chromatin compartmentalization in single cells, connecting 3D genome architecture, and transcriptional activity with the cell-to-cell variations of chromatin compartments.

Chromatin

Estimating population structure using epigenome-wide methylation data.

Population stratification is one of the source of inflation in epigenome-wide association studies (EWAS) when not properly accounted for. To address this, we developed methylation population scores (MPSs) to predict genetic principal components (GPCs) using a feature selection approach. We used multi-ethnic DNA methylation data from Illumina EPIC arrays across five cohorts, including MESA (n&#xa0;=&#xa0;929), CARDIA (n&#xa0;=&#xa0;1123), JHS (n&#xa0;=&#xa0;1365), ARIC (n&#xa0;=&#xa0;2338), and HCHS/SOL (n&#xa0;=&#xa0;1475), randomly splitting participants into training (85%) and test (15%) sets. Within each cohort, associations between GPCs and CpG sites were estimated using linear regression adjusting for age, sex, smoking and alcohol use, race/ethnicity, body mass index, and cell type proportions, followed by meta-analysis and selection of CpGs with FDR <0.05. We then applied a two-stage weighted least squares Lasso regression to construct MPSs, adjusting for the aforementioned covariates. In the test dataset, MPSs showed strong correlation with GPCs, with R&#xb2; ranging from 0.27 (MPS7 vs. GPC7) to 0.98 (MPS1 vs. GPC1). Visualization demonstrated that MPSs recapitulated the pattern shown by GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups and outperformed methylation-based principal components constructed using alternative published methods. Additionally, MPSs showed comparable performance to GPCs in reducing inflation in EWAS. Overall, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations, and provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent.

Humans

Morphological characterization, genetic diversity and population structure of the rice blast pathogen Magnaporthe oryzae in Northeast India.

The blast pathogen, Magnaporthe oryzae, is one of the most destructive fungal pathogens of rice worldwide, yet its morphological features, genetic diversity and population structure in Northeast India remain poorly understood. In this study, twenty&#x2012;two M. oryzae isolates collected from eight states of Northeast India were characterized using morphological, molecular, and population genetic analyses. Morphological characterization revealed whitish to greyish&#x2012;white mycelia with sparse sporulation and colony diameters ranged from 36 to 90&#xa0;mm, classifying the isolates into 14 fast and 8 slow&#x2012;growing groups. Whole genome sequencing was performed to enable both ITS&#x2012;based identification and SSR locus mining from the assembled genomes. Molecular identification using ITS rDNA sequences confirmed all isolates as M. oryzae, with 95.5-100% similarity. Phylogenetic analysis grouped the isolates into two major clades and identified seven ITS sequence types (GenBank Accessions: PX273287-PX273293). Genetic diversity assessed using 30 SSR markers revealed substantial polymorphism, with 1-7 alleles per locus and polymorphism information content (PIC) values ranging from 0.00 to 0.81. Heatmap clustering, dendrogram analysis, and distance metrics consistently identified two major genetic groups, with some isolates forming nearly identical clusters and others showing moderate divergence. Principal Component Analysis (PCA) and Principal Coordinates Analysis (PCoA) accounted for 87.8% of the total variance (PC1 and PC2 accounted for 54.4% and 33.4% respectively of the total variance) and revealed distinct outliers. Analysis of Molecular Variance (AMOVA) attributed 80% of the total genetic variation to differences among populations while only 20% was attributed to within population differences highlighting significant inter&#x2012;population divergence and clonal population structure. The study revealed substantial morphological and genetic diversity among M. oryzae populations in Northeast India, underscoring the need for region&#x2012;specific disease management strategies.

India

Genomic diversity, inbreeding, and selection signatures in duroc, landrace, and yorkshire pigs from a long-term closed breeding system.

Duroc (DD), Landrace (LL), and Yorkshire (YY) are among the most widely used commercial pig breeds, having undergone intense long-term selection within closed breeding systems. This study presents a comprehensive genomic analysis of genetic diversity, inbreeding patterns, and selection signatures in DD, LL, and YY populations that have been subject to close breeding for over 15 years. Genomic and pedigree data were available for 1,088 animals (DD&#x2009;=&#x2009;348, LL&#x2009;=&#x2009;276, YY&#x2009;=&#x2009;464), genotyped using the GenoBaits&#xae; Porcine 100&#xa0;K SNP panel. Principal component analysis and genetic diversity metrics revealed distinct population structures among the three breeds. Pairwise genetic differentiation supported this pattern, with DD showing the greatest divergence from LL (0.34&#x2009;&#xb1;&#x2009;0.24) and YY (0.33&#x2009;&#xb1;&#x2009;0.24), while LL and YY were more closely related (FST&#x2009;=&#x2009;0.22&#x2009;&#xb1;&#x2009;0.19). Linkage disequilibrium (LD) analysis further confirmed these differences, as DD exhibited the highest average r&#xb2; (0.34), followed by LL (0.28) and YY (0.25). Within-breed genetic diversity metrics, including observed heterozygosity (HO: 0.37 in DD, 0.39 in LL, 0.38 in YY), expected heterozygosity (HE: 0.36 in DD, 0.37 in LL, 0.38 in YY), and minor allele frequency (MAF: 0.27 in DD, 0.28 in LL, 0.29 in YY), indicated greater genetic variability in LL and YY compared to DD. Runs of homozygosity (ROH) analyses revealed different patterns of autozygosity, with DD exhibiting more long ROH indicative of recent inbreeding, while YY harbored a higher number of short ROH, suggestive of more ancient demographic events. ROH-based inbreeding coefficients (FROH) consistently exceeded pedigree-based estimates (FPED) across all breeds, highlighting the presence of recent or unrecorded inbreeding that pedigree data may not fully capture. According to Generation Proxy Selection Mapping (GPSM), 17, 1, and 12 significant SNPs were detected in DD, LL, and YY, respectively. Functional annotation of ROH islands and GPSM-significant loci revealed both breed-specific and overlapping QTLs related to traits such as growth, reproduction, and carcass. In general, the findings of this study contribute to a deeper understanding of the genomic consequences of long-term closed breeding and provide reference information to support consideration of breeding strategies that balance continued selection for productivity with the maintenance of genetic diversity in modern commercial pig populations.

Animals