Search PubMedSearch

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Chromosome-level genome assembly and annotation of the porcupine fish (Diodon hystrix).

The porcupinefish (Diodon hystrix), a coral reef teleost, is widely distributed in tropical/subtropical waters of the Pacific, Atlantic, Indian Oceans, and Mediterranean Sea. It shares easily recognizable features with pufferfish, such as body inflation and spines. Additionally, its culinary value makes D. hystrix a highly desirable species in many tropical coastal regions, with considerable market potential. However, lack of a high-quality genome hindered further studies on its reproduction, molecular biology, and genomic improvement. Here, we assembled the chromosome-scale genome using PacBio HiFi, ultra-long reads, and Hi-C. Of the 713.62 Mb genome, 98.63% anchored to 23 chromosomes (scaffold N50: 31.52 Mb) with 39.82% repetitive sequences. The assembled genome achieved a BUSCO completeness score of 97.7%, with 23,171 protein-coding genes predicted, 22,221 of which were functionally annotated. Phylogenetic analysis identified D. hystrix's evolutionary relationships with other species in the Tetraodontiformes. In summary, the high-quality genome of D. hystrix sheds light on valuable insights into genome size evolution, and provides a valuable resource for exploiting genomic study and breeding applications in this species.

Animals

Chromosome-level genome assembly and annotation of Spinibarbus caldwelli.

Spinibarbus caldwelli is an economically important freshwater species within the Cyprinidae family, abundant in the middle and lower reaches of the Yangtze River and its adjacent basins. As a promising species suitable for aquaculture in southern China, the lack of genomic resources has hampered the genetic breeding and conservation. Here, we release a chromosome-level genome assembly for S. caldwelli using PacBio HiFi long-reads, Illumina short-reads, and Hi-C sequencing data. The final genome assembly is 1.77 Gb in size, with a contig N50 of 24.27 Mb. Using Hi-C scaffolding, 99.14% of the contigs were successfully anchored to 50 chromosomes, resulting in a scaffold N50 of 35.29 Mb. The final genome assembly shows a BUSCO completeness of 98.27%. The assembled genome contains 49.41% repetitive sequences and 51,505 predicted genes, 90.83% of which have been functionally annotated. This genome provides a genetic basis for S. caldwelli, facilitating the exploration of Cyprinid phylogeny, genetic improvement, and conservation efforts.

Animals

Transcriptomic insights into the coordinated regulation of signaling, apoptosis, immunity, and metabolism during Sinonovacula constricta larval metamorphosis.

Metamorphosis is a critical ontogenetic transition for marine bivalves, marking the shift from planktonic to benthic lifestyles, where successful transformation dictates survival. The razor clam Sinonovacula constricta is economically important; however, low larval metamorphosis rates remain a major bottleneck in seedling production. To elucidate the mechanisms governing this process, we performed a comparative transcriptome analysis of S. constricta larvae at pre- and post-metamorphosis stages using Illumina sequencing. A total of 3701 differentially expressed genes (DEGs) were identified, including 3254 up-regulated and 447 down-regulated genes. Functional annotation of the respective top 20 significantly up-regulated and down-regulated DEGs indicated their potential pivotal roles in signal transduction (e.g., up-regulated: CAV1, CHRNA2; down-regulated: APP, NOTCH1), cellular proliferation and differentiation (e.g., up-regulated: TUBA, EGF1; down-regulated: KIF23, TTC25), transcriptional and epigenetic regulation (e.g., up-regulated: NFIL3; down-regulated: OVO, HMX1), substance transport (e.g., up-regulated: LRP2, LRP1B; down-regulated: SLC51A, Slc33a1), substance metabolism (e.g., up-regulated: CPK3, CYP26A1; down-regulated: RDMT1, ADAC), immunomodulation (e.g., up-regulated: CPN2, CRISP2), and protein homeostasis (e.g., up-regulated: HSP27, NAS-27). Functional enrichment analysis further revealed that DEGs were significantly enriched in pathways related to signal transduction and developmental regulation (e.g., Ras, TNF), cell death and homeostasis (e.g., apoptosis), immune responses (e.g., Toll-like receptor), energy metabolism (e.g., lipid), cardiovascular related (e.g., Fluid shear stress), cell junction and architecture (e.g., Tight junction), and infectious disease (e.g., measles). These results suggest a synergistic interplay between signaling, apoptosis, immunity, and metabolism during S. constricta metamorphosis. This study advances our understanding of marine bivalve metamorphosis and offers candidate genes for further mechanistic studies.

Animals

Meta-QTL Analysis Reveals Consensus Genomic Regions and Candidate Genes for Resistance to Sudden Death Syndrome in Soybean.

Sudden death syndrome (SDS), caused by Fusarium virguliforme, is one of the most economically important diseases limiting soybean production worldwide. Although numerous quantitative trait loci (QTL) associated with SDS resistance have been reported, inconsistencies among mapping populations, marker systems, and experimental conditions have hindered the identification of robust resistance loci for soybean improvement. In this study, a comprehensive meta-analysis was conducted to integrate published QTL and identify stable consensus genomic regions associated with SDS resistance. After a systematic literature survey and data curation, 153 QTL derived from 14 linkage-mapping studies were analyzed using a custom R-based workflow, resulting in the identification of 23 consensus meta-QTL (MQTL) distributed across 17 chromosomes. Several MQTL, particularly those located on chromosomes 6, 8, 18, and 20, were supported by multiple independent studies and represented major genomic hotspots for SDS resistance. Physical localization and functional annotation of these MQTL identified 217 candidate genes, including genes predicted to be involved in plant defense, signal transduction, transcriptional regulation, and secondary metabolism. Gene Ontology enrichment analysis identified response to salicylic acid as the only biological process that remained significant after FDR correction, whereas Kyoto Encyclopedia of Genes and Genomes pathway analysis did not identify significantly enriched pathways. Independent support using five published genome-wide association studies further supported several MQTL, especially those on chromosomes 6, 18, and 20, thereby increasing confidence in these genomic regions. The identified MQTL and prioritized candidate genes provide potential genomic resources for future marker development, improvement applications, and functional validation aimed at improving soybean resistance to SDS.

Fusarium virguliforme

MetagenomicKG: a knowledge graph for metagenomic applications.

MOTIVATION: The sheer volume and variety of genomic content within microbial communities makes metagenomics a field rich in biomedical knowledge. To traverse these complex communities and their vast unknowns, metagenomic studies often depend on distinct reference databases, such as the Genome Taxonomy Database (GTDB), the Kyoto Encyclopedia of Genes and Genomes (KEGG), and the Bacterial and Viral Bioinformatics Resource Center (BV-BRC), for various analytical purposes. These databases are crucial for the genetic and functional annotation of microbial communities. Nevertheless, the inconsistent nomenclature or identifiers of these databases present challenges for effective integration, representation, and utilization. Knowledge graphs (KGs) offer an appropriate solution by organizing biological entities from different databases to standardized identifiers, allowing their interrelations to be captured into a cohesive network regardless of the naming conventions used in each source. The graph structure not only facilitates the unveiling of hidden patterns but also enriches our biological understanding with deeper insights. Despite KGs having shown potential in various biomedical fields, their application in metagenomics remains underexplored. RESULTS: We present MetagenomicKG, a novel knowledge graph specifically tailored for metagenomic analysis. MetagenomicKG integrates taxonomic, functional, and pathogenesis-related information on the human microbiome sourced from various databases, and further connects these with existing biomedical KGs to expand the biological network. Through various case studies involving the human microbiome, we demonstrate its utility in enabling hypothesis generation regarding the relationships between microbes and diseases, generating sample-specific graph embeddings, and providing robust pathogen prediction. CODE AVAILABILITY: The source code and technical details for constructing the MetagenomicKG and reproducing all analyses are available on GitHub at https://github.com/KoslickiLab/MetagenomicKG. The data used in this manuscript, including the pre-built files and use case input data, are archived on Zenodo with DOI: 10.5281/zenodo.17546861.

Metagenomics

Whole genome sequence-based association analysis of African American individuals with bipolar disorder and schizophrenia.

In studies of individuals of primarily European genetic ancestry, common and low-frequency variants and rare coding variants have been found to be associated with the risk of bipolar disorder (BD) and schizophrenia (SZ). However, less is known for individuals of other genetic ancestries or the role of rare non-coding variants in BD and SZ risk. We performed whole genome sequencing of African American individuals: 1,598 with BD, 3,295 with SZ, and 2,651 unaffected controls (InPSYght study). We increased power by incorporating 14,812 jointly called psychiatrically unscreened ancestry-matched controls from the Trans-Omics for Precision Medicine (TOPMed) Program for a total of 17,463 controls. To identify variants and sets of variants associated with BD and/or SZ, we performed single-variant tests, gene-based tests for singleton protein truncating variants, and rare and low-frequency variant annotation-based tests with conservation and universal chromatin states and sliding windows. We found suggestive evidence of BD association with single-variants on chromosome 18 and of lower BD risk associated with rare and low-frequency variants on chromosome 11 in a region with multiple BD GWAS loci, using a sliding window approach. We also found that chromatin and conservation state tests can be used to detect differential calling of variants in controls sequenced at different centers and to assess the effectiveness of sequencing metric covariate adjustments. Our findings reinforce the need for continued whole genome sequencing in additional samples of African American individuals and more comprehensive functional annotation of non-coding variants.

Journal Article

A high-quality chromosome-level genome assembly and annotation of the giant freshwater prawn (Macrobrachium rosenbergii).

The giant freshwater prawn, Macrobrachium rosenbergii, is native to Southeast Asia and is used in aquacultural practices worldwide. It is considered advantageous because of its rapid growth, high nutritional value, and economic benefits. As one of the three major freshwater aquaculture shrimp sources in China, a high-quality genome resource is of great significance for promoting the germplasm improvement of varieties. This study presents a high-quality chromosome-level genome assembly of M. rosenbergii that was generated by combining PacBio, MGI, and Hi-C reads. The assembled genome was 2.96 Gb in size, with a contig N50 of 0.64 Mb and a scaffold N50 of 55.76 Mb, which was positioned on 59 pseudo-chromosomes. The Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis for genome assembly reached 94.37%. In total, 27,111 protein-coding genes were identified, of which 25,470 were functionally annotated. These results provide a foundation for future research into adaptive evolution, genomics, and molecular breeding in M. rosenbergii.

Animals

Integrating machine learning and GWAS for variant prioritization in the INCIPE cohort highlights ABC transporter genes in chronic kidney disease.

INTRODUCTION: Chronic kidney disease (CKD) is a major public health challenge, affecting approximately 674 million people worldwide and representing one of the fastest-growing causes of mortality. Since CKD is frequently asymptomatic in its early stages, the identification of novel genetic biomarkers may improve early detection and risk stratification. Genome-Wide Association Studies (GWAS) have identified numerous genetic loci associated with CKD and related traits; however, their performance is often limited in small and imbalanced cohorts, where reduced statistical power increases both false-positive and false-negative findings. Machine learning (ML) approaches can complement conventional GWAS by prioritizing biologically relevant genetic signals from high-dimensional genomic data. METHODS: In this study, we implemented a nested ensemble (NCBC) model composed of an undersampler and a CatBoostClassifier (CBC) to prioritize candidate genetic variants associated with CKD in the INCIPE cohort. Prioritized variants were functionally annotated and evaluated through enrichment analyses, GTEx gene expression profiling, and protein-protein interaction network analyses. Genes identified by the CKDGen Consortium were analysed as an external reference set and used to validate the biological relevance of the prioritized results. RESULTS: The NCBC model outperformed conventional ML classifiers, achieving a ROC AUC score of 87.77%, compared to 50%-53% for the other evaluated models. Among the prioritized genes, 56.25% showed protein-protein interactions with genes previously reported by the CKDGen Consortium, whereas only 1.9% of randomly generated gene sets showed interactions. DISCUSSION: Our study demonstrates that the NCBC model improves the prioritization of biologically plausible candidate variants in a small and imbalanced CKD cohort. Functional analyses suggested ABC transporter-related genes, including ABCA13, ABCA4, and ABCC4 genes, as promising candidate for future validation, with ABCA4 showing substantial expression in kidney tissues. Overall, these findings support the integration of ML with GWAS to prioritize candidate genes and investigate the genetic architecture of complex diseases.

SNP prioritization

Unraveling 'F' factor: towards a genetic-clinical framework for the musculoskeletal-heart crosstalk in metabolic aging.

BACKGROUND: The rising co-occurrence of cardiometabolic diseases and musculoskeletal degeneration poses a critical challenge to healthy aging, yet the shared biological mechanisms underlying this multimorbidity remain poorly defined. This study aimed to establish an integrative clinical-genetic framework to elucidate the common frailty factor, the 'F' factor, that captures the systemic vulnerability linking cardiometabolic multimorbidity (CMM) and musculoskeletal aging. METHODS: Utilizing the prospective China Health and Retirement Longitudinal Study (CHARLS) cohort, we developed and validated novel Frailty-Integrated Indices for CMM risk prediction, evaluated with machine learning models interpreted via SHapley Additive exPlanations (SHAP). Independently, we applied genomic structural equation modeling (Genomic-SEM) to integrate genome-wide association data from six traits-coronary artery disease, type 2 diabetes, hypertension, bone mineral density, frailty, and telomere length-to model a shared latent genetic factor ('F' factor). This was followed by multivariate GWAS, fine-mapping, transcriptome-wide association study (TWAS), gene-based analysis, and functional annotation to prioritize causal genes, pathways, and cell types. RESULTS: Clinically, several Frailty-Integrated Indices significantly improved CMM risk prediction, with the optimal model achieving an AUC of 0.727. Genetically, we modeled a significant shared latent genetic factor ('F' factor), pinpointing novel risk loci and implicating key genes such as APOE and SLC22A3. These genes were enriched in pathways including cellular senescence and cholesterol metabolism and showed specific expression patterns in developmental brain stages and across multi-organ endothelial cells. CONCLUSION: Our findings provide converging evidence for Musculoskeletal‑Heart crosstalk of metabolic aging and inferred the 'F' factor as a genetic correlate of a transdiagnostic state, which links genetic predisposition to metabolic dysregulation, and systemic functional decline. This work provides a multi-level biological characterization of multimorbidity liability, informing early-risk detection and preventive strategies for complex aging-related comorbidities.

Humans

A sequence-based classifier distinguishes phenotype-associated genes from other gene models in plants.

Only a small fraction of annotated plant genes possess experimentally validated associations with specific phenotypes. Phenotype-associated genes have distinct structural, molecular, and evolutionary characteristics compared with nonvalidated gene models. Here, we develop a simple classifier that uses sequence and evolutionary features, which can be generated for any species with an annotated reference genome assembly, to accurately distinguish phenotype-associated genes from both the overall population of annotated gene models and a specific set of genes identified as being tolerant of premature stop mutations. A model trained solely on genes from maize (Zea mays) identifies and prioritizes rice (Oryza sativa) and Arabidopsis (Arabidopsis thaliana) genes that are highly enriched in genes with experimentally validated links to phenotypes in both of these evolutionarily distant species. Gene models predicted to have a higher probability of being linked to phenotypes display patterns consistent with known biological properties of phenotype-associated genes. Notably, the sets of genes predicted to have a high probability of being linked to phenotype variation do not consist exclusively of well-characterized gene families but included many uncharacterized gene families carrying domains of unknown function. The quantitative scores generated by this model offer a valuable resource for prioritizing and exploring the vast number of uncharacterized gene models in plants, reducing the risk of failure in future reverse genetic efforts and potentially accelerating gene discovery and functional annotation in crops.

Phenotype

Genetic-epigenetic interactions (meQTLs) in orofacial clefts etiology.

OBJECTIVES: Nonsyndromic orofacial clefts (OFCs) involve complex genetic and environmental factors, with over 60 risk loci accounting for only a minority of estimated heritability and residing in non-coding regions with unclear functional relevance. We hypothesize that some genetic variants alter orofacial cleft risk by modifying DNA methylation (DNAm) at regulatory sequences essential for craniofacial development, acting as methylation quantitative trait loci (meQTLs). METHODS: We analyzed 10 well-established OFC-associated SNPs against genome-wide DNAm profiles in 409 cases and 456 controls, identifying 23 potential meQTLs. We validated findings using 358 cleft-discordant sibling pairs analyzed with quantitative MethyLight assays. Cross-referencing with the mQTL Database assessed temporal patterns across human development. Functional annotation used GeneHancer and craniofacial enhancer databases. RESULTS: Nine meQTLs were successfully replicated, including the highly significant rs987525 (8q24) - cg16561172 (MYC) association (P = 9.610E-6). This association mapped to a mesendoderm-active enhancer upstream of MYC, providing mechanistic explanation for the longstanding 8q24 cleft locus. Additional validated associations involved MAFB-PLCG1, NOG-PPM1E, FOXE1-FRZB, and SPRY2-LGR4 interactions. Independent differential methylation analysis revealed significant differences between discordant siblings at three CpG sites. Cross-referencing confirmed concordance with population-level methylation effects, with childhood representing the critical developmental window for most associations. CONCLUSIONS: This systematic meQTL characterization in OFCs demonstrates that genetic variants influence disease risk through epigenetic mechanisms. The 8q24-MYC regulatory pathway evidence provides crucial mechanistic insight into a major OFC risk locus. These findings bridge genetic associations with functional consequences, address missing heritability challenges, and suggest potential biomarkers and therapeutic targets for OFC prevention and treatment.

Journal Article

Association Between the Root Canal Microbiome and Apical Lesion Size: An Observational Shotgun Metagenomic Study.

AIM: The aim was to characterize the taxonomic and functional composition of the microbiome involved in primary endodontic infections and to evaluate their association with the periapical lesion size using shotgun metagenomic sequencing. METHODOLOGY: Samples from primary root canal infections diagnosed with apical periodontitis were analysed with shotgun sequencing. Samples were classified according to the lesion size as small (<&#x2009;3&#x2009;mm) or large (>&#x2009;7&#x2009;mm). The bacterial DNA copies in each group were quantified by qPCR. Taxonomic and functional annotations were made using Bracken/Kraken2 and HUMAnN3 software. Species richness, Shannon, Simpson and Pielou indices were used to measure alpha diversity. The similarity of the bacterial communities between study groups was evaluated by Principal Coordinate Analysis based on Bray-Curtis distances. The ALDEx2 package was used to infer the differences between species, and the edgeR package for KEGG pathways. For all statistical analyses, p&#x2009;<&#x2009;0.05 was considered as significant. RESULTS: A total of 49 samples were analysed, 27 with small lesions and 22 with large lesions. Species richness and Shannon indices showed differences between both groups, whereas no differences were seen according to Simpson and Pielou indices. A different community composition (PERMANOVA, p&#x2009;=&#x2009;0.0019) was observed between the two groups. Three species were significantly enriched in the large lesion samples, Filifactor alocis, Lachnospiraceae bacterium oral taxon 500 and Olsenella uli, while three others were enriched in small lesion samples, Acinetobacter baumannii, Acinetobacter pittii and Cutibacterium acnes. Functionally, benzoate, flavonoid and steroid degradation, the sphingolipid signalling pathway and proteasome function were enriched in samples with large lesions. Monoterpenoid biosynthesis, phospholipase D signalling, the sulphur relay system and staurosporine biosynthesis were enriched in small lesions. CONCLUSIONS: Teeth with large periapical lesions harbour greater bacterial loads and exhibit a more diverse microbial community than those with small lesions. Differences in species-level taxonomic composition were observed between both groups. Functionally, large lesions are enriched in pathways associated with immune evasion and pro-inflammatory activity, whereas small lesions are characterized by pathways related to apoptosis, metabolic adaptation and anti-inflammatory processes. These findings suggest that lesion severity is also shaped by the functional potential of the microbiome to modulate host inflammation.

Humans

Low-pass whole-genome sequencing reveals genomic diversity and ecotype-specific adaptation in indigenous Tigrayan chickens.

Indigenous chickens play a critical role in food security and climate resilience in smallholder systems, yet their genomic diversity and adaptive potential remain insufficiently characterised. This study employed low-pass whole-genome sequencing (LP-WGS; 0.2-1.99&#xd7;) to investigate genomic diversity, population structure, inbreeding and candidate environment-associated genomic variation in 33 chickens from highland, midland, and lowland agroecologies in the Tigray region of northern Ethiopia. After imputation and stringent filtering, 23.4 million high-confidence SNPs were retained, including&#x2009;~&#x2009;17% novel variants, indicating substantial uncharacterised genetic diversity in these populations. SNP density (13.8&#x2009;&#xb1;&#x2009;8.6 SNPs/kb) was comparable to values reported from high-coverage Ethiopian chicken datasets, demonstrating the suitability of LP-WGS for population genomics in resource-limited settings. Marked differences in genomic diversity were observed among ecotypes: midland chickens showed the highest nucleotide diversity (&#x3c0;&#x2009;=&#x2009;0.00267), followed by lowland (&#x3c0;&#x2009;=&#x2009;0.00233), whereas highland chickens showed the lowest diversity (&#x3c0;&#x2009;=&#x2009;0.00203) and elevated genomic inbreeding (FROH and FHOM &#x2248; 0.18). Population structure analyses revealed clear genetic separation among ecotypes. PCA (13.91% variation explained) distinguished lowland chickens along PC1 and separated highland from midland along PC2, while ADMIXTURE and FST patterns supported three major ancestral genomic backgrounds. Functional annotation of private missense variants uncovered distinct adaptive signatures reflecting the contrasting agroecological conditions. Highland chickens showed enrichment of candidate genes potentially involved in physiological processes relevant to high-altitude environments, including cold response, angiogenesis, cardiovascular regulation and metabolic homeostasis (eg., PARP1, ACOX2, ITGB3, EDNRB, SOX8, and SOX10). Midland chickens exhibited candidate signals of selection in genes with known roles in innate antiviral immunity, bacterial defence and inflammatory regulation (eg., BAK1, CLSTN1, CYSLTR1, CYSLTR2, CXCR7, GIPR, DSCAM, GDAP1, TLR3, TLR4, TLR7, IFIH1, ADORA1, EPHB1, and TMPRSS2). Lowland chickens displayed candidate variants associated with heat-stress response, DNA damage repair, oxidative balance and cardiovascular support under extreme temperatures (e.g., MLH1, BDKRB1, GPR19, FLT1, CCL18, TGM2, and RAMP3). Overall, the results indicate substantial genomic differentiation among ecotypes and suggest candidate environment-associated genetic divergence across Tigray's diverse agroecological zones. These populations may represent important reservoirs of adaptive genetic variation for climate-resilient poultry breeding, warranting further functional validation and conservation-oriented management.

Animals

Genome-wide scans reveal candidate genes associated with wing morph differentiation in Tetrix japonica.

Wing dimorphism is an important dispersal-related trait in insects, but its genomic basis remains poorly understood in pygmy grasshoppers. Here, we integrated genome-wide single-nucleotide polymorphism (SNP) analyses, population structure inference, selection scans, and functional annotation to investigate genomic differentiation between long- and short-winged Tetrix japonica. Principal component analysis (PCA), ADMIXTURE, and phylogenetic analyses revealed weak genome-wide separation between morphs, indicating differentiation on a largely shared genetic background. Genome-wide scans based on the fixation index (FST), nucleotide diversity ratios, and Tajima's D, using 50-kb non-overlapping windows and empirical top-5% outlier thresholds, identified multiple candidate regions across seven chromosomes. The broader long- and short-winged candidate sets spanned 9.35&#xa0;Mb and 9.37&#xa0;Mb and directly overlapped 82 and 77 genes, respectively. Candidate genes were associated with signaling/hormone regulation, membrane transport, metabolism, cytoskeletal organization, extracellular matrix structure, and development. Short-winged candidate genes were significantly enriched for ABC-type transporter activity and ATP hydrolysis activity. Because all individuals originated from a single laboratory-maintained population with weak genome-wide structure, these regions should be regarded as candidate loci from a screening-stage analysis that require validation in independent populations and by functional assays, rather than as confirmed targets of selection.

Animals

Role of functional genes for seed vigor related traits through genome-wide association mapping in finger millet (Eleusine coracana L. Gaertn.).

Finger millet (Eleusine coracana (L.) Gaertn.) is a calcium-rich, nutritious and resilient crop that thrives even in harsh environmental conditions. In such ecologies, seed longevity and seedling vigor are crucial for sustainable crop production amid climate change. The current study explores the genetics of accelerated aging on seed longevity traits across 221 diverse accessions of finger millet through genome-wide association approach (GWAS). A significant variation was identified in germination percentage, germination rate indices, mean germination time, seedling vigor indices and dry weight upon aging treatment. GWAS model from 11,832 high-quality SNPs identified through Genotyping-by-Sequencing (GBS) approach produced 491 marker-trait associations (MTAs) for 27 traits, of which 54 were FDR-corrected. A pleiotropic SNP, FM_SNP_9478 identified on chromosome 7B was associated with the traits viz., germination after aging, germination index after aging and their relative measures. Functional annotation revealed DET1 and expansin-A2 influenced seed coat integrity, critical for germination and aging resilience. Probable protein phosphatase 2C3 and piezo-type ion channels contributed to mechanical sensing and stress adaptation in seeds. Beta-amylase and acetyl-CoA carboxylase 2 were identified for seed metabolism and stress response. These insights lay the framework for targeted breeding efforts to improve seed quality and resilience under diverse production conditions.

Eleusine

Identification of a common secondary mutation in the Neurospora crassa knockout collection conferring a cell fusion-defective phenotype.

Gene-deletion mutants represent a powerful tool to study gene function. The filamentous fungus Neurospora crassa is a well-established model organism, and features a comprehensive gene knockout strain collection. While these mutant strains have been used in numerous studies, resulting in the functional annotation of many Neurospora genes, direct confirmation of gene-phenotype relationships is often lacking, which is particularly relevant given the possibility of background mutations, sample contamination, and/or strain mislabeling. Indeed, spontaneous mutations resulting in phenotypes resembling many cell fusion mutants have long been known to occur at relatively high frequency in N. crassa, and these secondary mutations are common in the Neurospora deletion collection. The identity of these mutations, however, is largely unknown. Here, we report that the &#x394;ada-3 strain from the N. crassa knockout collection, which exhibits a cell fusion defect, harbors a secondary mutation responsible for this phenotype. Through whole-genome sequencing and genetic analyses, we found a ~30-Kb deletion in this strain affecting a known cell fusion-related gene, so/ham-1, and show that it is the absence of this gene&#x2014;and not of ada-3&#x2014;that underlies its cell fusion defect. We additionally found three other knockout strains harboring the same deletion, suggesting that this mutation may be common in the collection and could have impacted previous studies. Our findings provide a cautionary note and highlight the importance of proper functional validation of strains from mutant collections. We discuss our results in the context of the spread of cell fusion-defective cheater variants in N. crassa cultures.

Neurospora crassa

Genetic Architecture of Idiopathic Inflammatory Myopathies From Meta-Analyses.

OBJECTIVE: Idiopathic inflammatory myopathies (IIMs, myositis) are rare systemic autoimmune disorders that lead to muscle inflammation, weakness, and extramuscular manifestations, with a strong genetic component influencing disease development and progression. Previous genome-wide association studies identified loci associated with IIMs. In this study, we imputed data from two prior genome-wide myositis studies and analyzed the largest myositis data set to date to identify novel risk loci and susceptibility genes associated with IIMs and its clinical subtypes. METHODS: We performed association analyses on 14,903 individuals (3,206 patients and 11,697 controls) with genotypes and imputed data from the Trans-Omics for Precision Medicine reference panel. Fine-mapping and expression quantitative trait locus colocalization analyses in myositis-relevant tissues indicated potential causal variants. Functional annotation and network analyses using the random walk with restart (RWR) algorithm explored underlying genetic networks and drug repurposing opportunities. RESULTS: Our analyses identified novel risk loci and susceptibility genes, such as FCRLA, NFKB1, IRF4, DCAKD, and ATXN2 in overall IIMs; NEMP2 in polymyositis; ACBC11 in dermatomyositis; and PSD3 in myositis with anti-histidyl-transfer RNA synthetase autoantibodies (anti-Jo-1). We also characterized effects of HLA region variants and the role of C4. Colocalization analyses suggested putative causal variants in DCAKD in skin and muscle, HCP5 in lung, and IRF4 in Epstein-Barr virus (EBV)-transformed lymphocytes, lung, and whole blood. RWR further prioritized additional candidate genes, including APP, CD74, CIITA, NR1H4, and TXNIP, for future investigation. CONCLUSION: Our study uncovers novel genetic regions contributing to IIMs, advancing our understanding of myositis pathogenesis and offering new insights for future research.

Humans

Revealing the Shared Genetic Architecture of Metabolic Dysfunction-Associated Steatotic Liver Disease-Related Traits Through Genomic Structural Equation Modeling.

Although individual traits related to metabolic dysfunction-associated steatotic liver disease (MASLD) have been investigated through large-scale genome-wide association studies (GWASs), the shared genetic susceptibility across these traits remains unclear. We therefore conducted a multivariate GWAS of key MASLD-related traits to elucidate their common genetic architecture. We applied genomic structural equation modeling to model a latent genetic factor (MASLD-F) underlying genetically correlated MASLD-related traits, leveraging their GWAS-derived genetic correlations. We then performed functional annotations, including fine-mapping, transcriptome-wide association study, and cell- and tissue-type-specific enrichment analyses, and conducted Mendelian randomization analyses to identify modifiable risk factors. Our multivariate MASLD-F GWAS identified 50 independent variants across 48 genomic loci. Transcriptomic imputation identified several MASLD-F-associated genes, including ARNTL, NPC1, BTBD10, VDAC2, TSKU, SFMBT1, and ABHD17C. We observed significant enrichment of MASLD-F-related genetic signals predominantly in brain tissues, pancreatic islets, and the adrenal gland. Additionally, six modifiable risk factors and four modifiable protective factors for MASLD-F were identified. These findings reveal a complex shared genetic architecture underlying MASLD components, thereby expanding our understanding of disease pathogenesis and providing novel insights for precision medicine and public health interventions.

Humans