Search PubMedSearch

SEARCH · Search PubMed

Results for “Integrative taxonomy”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Genome-resolved analysis of colonization factor repertoires reveals ecological stratification in cervid gut microbiomes.

INTRODUCTION: Colonization factors (CFs) are important microbial traits associated with persistence and host adaptation in the gut, yet their large-scale organization in cervid gut microbiomes remains unclear. METHODS: A total of 3,311 non-redundant high-quality metagenome-assembled genomes (MAGs), derived from 688 cervid gut metagenomic samples across 15 publicly available projects and one in-house dataset, were analyzed. CF-associated genes were identified by comparison against the GHA CF database, and CF repertoires were characterized at genome, host-species, and gastrointestinal-segment levels. RESULTS: A total of 138,729 CF-associated genes spanning 71 CF families were identified. MAGs from Cervinae contained richer CF repertoires than those from Caprinae, and CF47 (Peptidase_C69), CF24_29 (QueH), and CF18 (Glycos_transf_2) were among the most prevalent families. CF repertoires were strongly structured by taxonomy, showed a moderate association with bacterial phylogenetic distance, and formed two recurrent genome-level configurations with distinct KEGG functional profiles. Integration of sample metadata further revealed differentiation of CF repertoires across host species and gastrointestinal segments, representing the major ecological dimensions examined in this study. Segment-associated CF variation was accompanied by redistribution of broader functional profiles, including enrichment of carbohydrate and lipid metabolism in the jejunum, membrane transport in the ileum, xenobiotics biodegradation in the cecum, and environmental adaptation in the rumen. DISCUSSION: These findings provide a genome-resolved view of CF repertoire organization in cervid gut microbiomes and demonstrate that colonization-associated functions are structured across microbial lineages and ecological contexts. This study highlights the importance of considering microbial taxonomy and host-associated environments when interpreting the distribution of CF repertoires in mammalian gut ecosystems.

Cervidae

Stenotrophomonas maltophilia in the Antimicrobial Resistance Era: Species-Complex Taxonomy, Pathogenesis, Evolving Therapeutic Priorities, and Genomic Surveillance.

Stenotrophomonas maltophilia is a globally distributed, aerobic, non-fermenting Gram-negative bacillus increasingly recognized as an opportunistic pathogen in hospitalized and immunocompromised patients. Clinical interpretation is challenging because respiratory and device-associated isolates may represent colonization, polymicrobial infection, or true invasive disease. Recent genomic studies further suggest that organisms historically identified as S. maltophilia comprise a genetically diverse species complex, with implications for epidemiology, virulence, resistance surveillance, and susceptibility testing. Treatment is difficult because of biofilm formation, persistence in water-associated healthcare reservoirs, and intrinsic or acquired resistance mediated by L1 and L2 β-lactamases, multidrug efflux pumps, reduced permeability, mobile resistance determinants, and biofilm-associated tolerance. Current IDSA guidance identifies cefiderocol monotherapy as the preferred treatment for invasive S. maltophilia infection, whereas aztreonam-avibactam and agents such as trimethoprim-sulfamethoxazole, levofloxacin, and minocycline occupy alternative or combination-based roles. Nevertheless, the therapeutic evidence base remains uneven, and clinical decisions should integrate infection severity, source control, susceptibility findings, pharmacokinetic/pharmacodynamic (PK/PD) exposure, toxicity, infection site, and host-related factors. This review summarizes advances in taxonomy, epidemiology, pathogenesis, diagnostics, resistance, treatment, infection prevention, and genomic surveillance, and highlights the need for standardized identification, validated breakpoints, prospective comparative-effectiveness studies, and pragmatic or adaptive trial designs.

L1 β-lactamase

FANTASIA suite: a reproducible and configurable framework for embedding-based functional annotation of proteins.

Embedding-based annotation transfer is increasingly used for protein function inference due to protein language models capture sequence, structural, and functional signals that may extend beyond conventional pairwise similarity. However, systematic application of these approaches requires control over model choice, reference composition, lookup parameters, evidence traceability, and output formats. We developed the FANTASIA suite, a configurable framework for embedding-based functional annotation of proteins. The suite combines a database-backed implementation for reproducible and extensible analyses with a portable flat-file implementation for rapid local annotation and pipeline integration. Using non-model and model-organism proteomes, we show that larger neighbourhood sizes remain practical for proteome-scale analyses and that taxonomy and sequence-identity filtering support leakage-aware benchmarking. We also compare the supported models with baseline methods through external CAFA5 evaluation and provide practical guidance based on empirical evidence variables. FANTASIA provides a controlled, scalable, and reproducible framework for extending functional annotation across the rapidly expanding diversity of sequenced organisms.

Software

Evolutionary dynamics of the chloroplast genome in Abutilon (Malvoideae, Malvaceae).

The genus Abutilon Mill. (Malvaceae) comprises approximately 178 species distributed across tropical and subtropical regions, many of which hold significant ornamental, economic, and medicinal value; yet its taxonomic classification remains challenging. In this study, six species were sequenced from herbarium specimens, and the chloroplast (cp.) genomes of ten additional species were assembled de novo from publicly available raw data. Three previously reported cp. genomes were also incorporated to characterise cp. genome structure, identify polymorphic loci, and perform phylogenetic analyses. The cp. genomes ranged from 159,458 to 160,454 bp and exhibited the typical quadripartite structure, with each genome containing 112 unique genes (78 protein-coding, 30 tRNA, and 4 rRNA) that showed conserved content and organisation. These genomes exhibited high similarity in GC content, inverted repeat boundaries, relative synonymous codon usage, amino acid frequencies, and substitution patterns. However, notable variation was observed in the total number of simple sequence repeats, ranging from 70 to 97 per genome. Selection analyses indicated predominant purifying selection, with evidence of episodic positive selection detected in rpoC2, rbcL, and ycf1. Two codons in rbcL were clade-specific and provided phylogenetic signal distinguishing Australian and Old World pantropical species. Nucleotide diversity analysis identified six highly polymorphic intergenic spacers (trnH-psbA, rps19-rpl2, psbT-pbf1, psaC-ndhD, trnR-atpA, and ndhJ-ndhK) that may be suitable for taxonomic studies. The phylogeny from maximum likelihood (ML) and Bayesian inference (BI) resolved two major clades: one comprising an exclusively Australian lineage occurring predominantly in arid and semi-arid environments, and the other a pantropical lineage spanning multiple continents. Abutilon grandifolium was recovered as sister to the remaining sampled Abutilon taxa in both ML and BI analyses, although no biogeographic origin inference can be drawn from this placement pending broader taxon sampling and integration of nuclear genomic data. These findings provide insights into the evolutionary dynamics of the cp. genome in Abutilon and offer a foundational genomic framework for refining Abutilon taxonomy.

Genome, Chloroplast

Integrated multi-omics analysis and functional experiments reveals PPAP2C as a potential prognostic biomarker and therapeutic target in breast cancer.

BACKGROUND: This study aims to systematically elucidate the clinical significance and biological function of the phospholipid phosphatase (PLPP) family member (PPAP2C) phosphatidic acid phosphatase type 2C in breast cancer, and to evaluate its potential as a prognostic biomarker and therapeutic target. METHODS: Gene expression data from The Cancer Genome Atlas (TCGA), Genotype-Tissue Expression (GTEx), and Cancer Cell Line Encyclopedia (CCLE) databases were integrated to characterize the expression profile of PLPP family members, focusing on PPAP2C in breast cancer. The prognostic value of PPAP2C, initially identified at the mRNA level (TCGA, (METABRIC) Molecular Taxonomy of Breast Cancer International Consortium, Gene Expression Omnibus (GEO)), was confirmed at the protein level by immunohistochemistry (IHC) on tissue microarrays (TMA). The oncogenic functions of PPAP2C were investigated in triple-negative breast cancer (TNBC) cells through CRISPR-Cas9-mediated knockout and ectopic overexpression, with assessment of key phenotypes including proliferation, colony formation, migration, and invasion. In vivo validation was subsequently performed using an MDA-MB-231 xenograft model. RESULTS: PPAP2C exhibits the most significant overexpression pattern across 33 cancer types (upregulated in 16 cancers, downregulated in only 3). Compared with normal tissues, PPAP2C showed specific overexpression in breast cancer tissues and was significantly associated with advanced clinical stages and aggressive subtypes (HER2+ and TNBC). Survival analysis demonstrated that high PPAP2C expression correlated with significantly shorter overall survival and disease-free survival, which was further validated in METABRIC and GEO cohorts. Tissue microarray analysis confirmed higher PPAP2C protein positivity in tumor tissues (94.7%) than in adjacent normal tissues (59.7%), with worse OS and RFS in high-expression groups. Multivariate analysis identified PPAP2C as an independent prognostic factor for OS. Functional experiments revealed that PPAP2C knockout (via 5-bp/1-bp frameshift mutations) suppressed TNBC cell proliferation, colony formation, migration, and invasion, while overexpression enhanced these phenotypes. In vivo studies further demonstrated complete tumor regression in MDA-MB-231 xenografts upon PPAP2C knockout. CONCLUSION: This study identifies PPAP2C as a key oncogenic driver and a robust independent prognostic biomarker in breast cancer. The findings provide compelling evidence that PPAP2C represents a promising therapeutic target, offering a new strategic avenue for precision therapy, particularly for aggressive breast cancer subtypes.

PLPP2

Digital and computational morphology in hematology: current platforms, clinical evidence, and future requirements.

INTRODUCTION: Morphologic examination of peripheral blood and bone marrow remains central to the diagnosis and classification of hematologic disorders. Conventional optical microscopy, however, is labor-intensive, dependent on operator expertise, and affected by interobserver variability. Digital morphology has developed from automated image acquisition and cell pre-classification into a broader field that includes whole-slide imaging, remote review, quantitative morphometry, and artificial intelligence-based analysis. CONTENT: This review examines current applications of digital morphology in peripheral blood, bone marrow aspirates, malaria detection, and body-fluid analysis. Commercial platforms are evaluated with particular attention to the distinction between raw automated pre-classification, expert digital post-classification, and comparison with independent optical microscopy. Digital systems generally perform well for common mature leukocyte populations but remain less reliable for rare or diagnostically critical cells, including blasts, abnormal lymphoid cells, plasma cells, and intermediate maturation stages. Research systems increasingly extend analysis from individual-cell classification to whole-slide, specimen-level, and patient-level assessment. SUMMARY: Digital morphology can improve standardization, image traceability, remote consultation, education, proficiency testing, quality assurance, and selected aspects of laboratory workflow. Its clinical value depends on appropriate validation, transparent reporting of reference methods, recognition of algorithm-specific failure modes, and clearly defined criteria for expert review and conventional microscopy. Human expertise remains essential not only for validating results but also for adapting cell taxonomies and interpretive rules to evolving classifications of hematologic diseases. OUTLOOK: Future progress will require representative multicenter datasets, harmonized morphologic terminology, external validation, interoperability with laboratory information systems, and continuous monitoring after software or hardware updates. Integration of morphology with quantitative hematology, flow cytometry, cytogenetics, genomics, and clinical data may support more comprehensive computational diagnosis. Digital platforms may also broaden access to specialist expertise, training, and quality programs in resource-limited institutions and regions, provided that infrastructure, governance, and professional competency are adequately supported.

artificial intelligence

Advances in the Genus Ulva Research: From Structural Diversity to Applied Utility.

The green macroalgae Ulva Linnaeus, 1753, also known as sea lettuce, is one of the most ecologically and economically significant algal genera. Its representatives occur in marine, brackish, and freshwater environments worldwide and show high adaptability, rapid growth, and marked biochemical diversity. These traits support their ecological roles in nutrient cycling, primary productivity, and habitat provision, and they also explain their growing relevance to the blue bioeconomy. This review summarizes current knowledge of Ulva biodiversity, taxonomy, and physiology, and evaluates applications in food, feed, bioremediation, biofuel, pharmaceuticals, and biomaterials. Particular attention is given to molecular approaches that resolve taxonomic difficulties and to biochemical profiles that determine nutritional value and industrial potential. This review also considers risks and limitations. Ulva species can act as hyperaccumulators of heavy metals, microplastics, and organic pollutants, which creates safety concerns for food and feed uses and highlights the necessity of strict monitoring and quality control. Technical and economic barriers restrict large-scale use in energy and material production. By presenting both opportunities and constraints, this review stresses the dual role of Ulva as a promising bioresource and a potential ecological risk. Future research must integrate molecular genetics, physiology, and applied studies to support sustainable utilization and ensure safe contributions of Ulva to biodiversity assessment, environmental management, and bioeconomic development.

algal bloom

MetagenomicKG: a knowledge graph for metagenomic applications.

MOTIVATION: The sheer volume and variety of genomic content within microbial communities makes metagenomics a field rich in biomedical knowledge. To traverse these complex communities and their vast unknowns, metagenomic studies often depend on distinct reference databases, such as the Genome Taxonomy Database (GTDB), the Kyoto Encyclopedia of Genes and Genomes (KEGG), and the Bacterial and Viral Bioinformatics Resource Center (BV-BRC), for various analytical purposes. These databases are crucial for the genetic and functional annotation of microbial communities. Nevertheless, the inconsistent nomenclature or identifiers of these databases present challenges for effective integration, representation, and utilization. Knowledge graphs (KGs) offer an appropriate solution by organizing biological entities from different databases to standardized identifiers, allowing their interrelations to be captured into a cohesive network regardless of the naming conventions used in each source. The graph structure not only facilitates the unveiling of hidden patterns but also enriches our biological understanding with deeper insights. Despite KGs having shown potential in various biomedical fields, their application in metagenomics remains underexplored. RESULTS: We present MetagenomicKG, a novel knowledge graph specifically tailored for metagenomic analysis. MetagenomicKG integrates taxonomic, functional, and pathogenesis-related information on the human microbiome sourced from various databases, and further connects these with existing biomedical KGs to expand the biological network. Through various case studies involving the human microbiome, we demonstrate its utility in enabling hypothesis generation regarding the relationships between microbes and diseases, generating sample-specific graph embeddings, and providing robust pathogen prediction. CODE AVAILABILITY: The source code and technical details for constructing the MetagenomicKG and reproducing all analyses are available on GitHub at https://github.com/KoslickiLab/MetagenomicKG. The data used in this manuscript, including the pre-built files and use case input data, are archived on Zenodo with DOI: 10.5281/zenodo.17546861.

Metagenomics

The landscape of pruning for large language models: A systematic review and unified taxonomy.

Confronting the inherent tension between the exceptional capabilities and the immense computational costs of Large Language Models (LLMs), pruning has become a crucial technique for achieving efficient deployment. However, a systematic analytical framework dedicated specifically to LLM pruning remains absent. In this paper, we aim to bridge this gap. We first elucidate the theoretical foundations that underpin the effectiveness of pruning, namely overparameterization and redundancy, and then propose a multidimensional taxonomy that organizes existing approaches along the axes of granularity, timing, and criteria. Building upon this unified perspective, we further analyze performance recovery mechanisms and the broader evaluation ecosystem, while also exploring forward-looking challenges such as interpretability, automation, and hardware-algorithm co-design. Through this comprehensive synthesis, we seek to provide an integrated and coherent analytical lens for advancing both research and practice in LLM pruning.

Large Language Models

Tropilaelaps mercedesae: an emerging global threat to apiculture - a comprehensive review.

Honey bees (Apis spp.) are key pollinators in agricultural and natural ecosystems; however, their populations are declining due to multiple interacting stressors and their synergistic effects, including parasitic mites. While Varroa destructor is widely recognized as the primary global driver of colony losses, mites of the genus Tropilaelaps, particularly Tropilaelaps mercedesae, are emerging as a serious and still underestimated threat. Native to Asia and naturally associated with wild hosts such as Apis dorsata, T. mercedesae has successfully transitioned to managed Apis mellifera colonies and is now widespread across much of Asia. Recent reports from Central Asia and the Caucasus and western Eurasian regions (including Georgia and Russia) indicate that this species is undergoing an ongoing westward expansion toward Europe. Its biological traits-including an extremely short reproductive cycle, obligate dependence on sealed brood, high dispersal capacity, and the potential to transmit viruses such as deformed wing virus (DWV)-facilitate rapid population growth and severe colony-level damage, particularly in A. mellifera, which lacks effective behavioral defenses against this mite. This review synthesizes current knowledge on the taxonomy, morphology, life cycle, host-parasite interactions, geographic distribution, and spread of Tropilaelaps mites, with emphasis on T. mercedesae. It also evaluates available diagnostic approaches, including brood-based methods, adult bee-based methods, and natural mite-fall techniques. Furthermore, evidence on chemical and biotechnical control strategies is summarized, and their strengths, limitations, and integration within an Integrated Pest Management (IPM) framework are discussed. Overall, current findings highlight the urgent need to strengthen surveillance, standardize diagnostic protocols, and develop sustainable control strategies to prevent the global spread of Tropilaelaps mites.

A. mellifera

The diagnostic potential of combined quantitative polymerase chain reaction and next-generation sequencing using the same primers for periprosthetic joint infection.

Next-generation sequencing (NGS) enables the detection of specific pathogens unidentifiable by conventional cultures, but its application in orthopedics remains inconsistent due to background contamination and irreproducible findings. This study evaluated the diagnostic performance of a novel workflow combining broad-range 16S rRNA gene quantitative PCR (qPCR) screening with downstream NGS, focusing on bacterial biomass thresholds. The qPCR assay demonstrated excellent intrarater reliability, with an intraclass correlation coefficient (ICC) of 0.961 (95% confidence interval, 0.881 to 0.997). Based on serially diluted positive controls, a quantitative threshold of 10⁵ CFU/mL was established as the minimum concentration required for the consistent detection of fastidious taxa, such as Escherichia coli. When evaluated against conventional cultures using 95 sonicate fluid and 276 pre/intraoperative tissue samples, the qPCR assay achieved a sensitivity of 80% and a specificity of 72%. Subsequent NGS sequencing of 26 clinical samples and 9 controls showed concordance in 4 of 6 culture-positive infected cases with NGS taxonomy, whereas the remaining discrepancies were likely attributable to culture-based phenotypic misidentification. Notably, among the qPCR-positive cases, three were culture-negative, including two hip prosthesis loosening cases exhibiting polymicrobial profiles, and one post-traumatic osteoarthritis case harboring low-level Staphylococcus. Crucially, this post-traumatic patient developed delayed periprosthetic joint infection (PJI) 2 years post-surgery, with cultures identifying Staphylococcus previously detected by the initial NGS analysis. Integrating qPCR screening with targeted NGS effectively refines pathogen identification, filters environmental artifacts, and overcomes the diagnostic limitations of culture-negative infections in orthopedic practice.IMPORTANCENext-generation sequencing (NGS) enables the detection of specific pathogens in clinical samples that are not identifiable by conventional methods. However, NGS applications in orthopedics have not been quantitatively evaluated, and findings have been inconsistent owing to contaminants and the presence of non-credible causative organisms. These factors primarily stem from the failure to evaluate low-biomass samples and the absence of proper controls, such as negative controls or mock community DNA samples. This study demonstrates that interpreting results from low-biomass samples requires careful consideration because NGS relies on relative bacterial abundances; distinguishing likely pathogens from contaminants is particularly challenging when bacterial loads are low. We demonstrated that combining NGS with quantitative PCR (qPCR) and applying a Cq cutoff can reduce false positives.

Humans

Autism care: reimagining the spectrum.

Autism care policy is at a critical inflection point. Applied behavior analysis (ABA), long established as the "gold standard" through state insurance mandates in the US, has functioned as the default reimbursable intervention for autistic children. However, advances in genomics, neuroscience, developmental psychology, and scholarship on autistic lived experience have expanded understanding of autism as a heterogeneous neurotype characterized by meaningful differences in neural organization rather than a unitary disorder. Contemporary models emphasize neurodiversity, strengths-based perspectives, and the interaction between developmental processes and environmental contexts in shaping functional outcomes. Many autistic children also meet criteria for complex care needs, requiring coordinated, interdisciplinary services across health, educational, and community systems. This manuscript proposes reframing the "autism spectrum" from a hierarchy of symptom severity to a prevention-oriented "spectrum of care." Adapting a public health taxonomy, interventions are organized into universal, selective, and indicated levels, targeting the prevention of avoidable disability, distress, and participation barriers. This model aligns autism services with whole-child, neurodiversity-affirming, and developmentally informed care, emphasizing relational health, autonomy, and life-course participation.

applied behavior analysis

Unraveling the diversity, function, and virus-host interactions of archaeal proviruses.

Archaea, the third domain of life, play critical roles in global biogeochemical cycles. However, archaeal proviruses integrated into host genomes remain largely unexplored. To bridge this gap, we conducted a large-scale mining of genomes spanning all presently known 21 archaeal phyla for their proviruses. We identified 770 archaeal proviruses across 12 archaeal phyla and 84 families, which clustered into 655 viral operational taxonomic units (vOTUs). Among these, 86.1% of the vOTUs were novel at the species level, and 69.3% could not be classified at the family level, substantially expanding the known diversity of archaeal viruses. Additionally, phylogenomic analysis supported the proposal of 16 putative novel viral families, further extending the current taxonomy landscape of archaeal viruses. Notably, 21.8% of the identified proviruses were predicted to adopt a lytic lifestyle, suggesting that these proviruses may retain the capacity to enter the lytic cycle under appropriate conditions. Host prediction indicated only 14 out of the 655 vOTUs might have potential across-lineage infection abilities. We detected 63 anti-defense genes encoded by 61 provirus genomes, such as anti-CRISPR and anti-RM, suggesting an ongoing evolutionary arms race between hosts and proviruses. However, only 10 auxiliary metabolic genes (AMGs) were identified, suggesting a limited impact of proviruses in the modulation of host metabolism through AMGs. This study establishes a systematic global genomic atlas of archaeal proviruses, advancing our understanding of their distribution and diversity while providing a foundation for future research into how proviruses regulate archaeal metabolism and ecosystem functioning.

anti-defense system

Methylation profiling in CNS tumor diagnostics: a single-centre real-world experience from Central Europe.

Genome-wide DNA methylation profiling has transformed neuro-oncology by providing an objective, machine learning-based taxonomy that mitigates interobserver variability and refines the histo-molecular criteria of the current WHO classification. We evaluate the real-world diagnostic performance and clinical utility of this modality in a prospective, consecutively accrued three-year cohort of 291 central nervous system (CNS) tumors across a mixed adult-pediatric population. Successful profiling was completed in 95.9% of cases. Using the Epignostix classifier, a high-confidence diagnostic match (calibrated score [CS]&#x2009;&#x2265;&#x2009;0.84) was achieved in 70.3% of analyzable samples, while 26.5% returned lower-confidence scores (&#x2265;&#x2009;0.3 to <&#x2009;0.84) and only 3.2% remained completely unclassifiable (CS&#x2009;<&#x2009;0.3). When integrated into a comprehensive diagnostic framework, methylation profiling provided clinically useful results in 81.1% of cases, establishing diagnoses in 70 cases submitted for molecular subclassification and resolving diagnostic uncertainty or prompting major revisions in 149 histologically challenging tumors. Within truly ambiguous lesions, integration of methylome data dictated tumor grade modifications in 38.8% of cases (upgrading in 29.4% and downgrading in 9.4%), shifting patient risk stratification. Crucially, over half (52.7%) of the lower-confidence cases yielded meaningful clinical integration when supported by histomorphology and ancillary genetic or immunohistochemical markers, demonstrating that rigid score cutoffs should not dictate assay failure. Discrepant or misleading classifications occurred in 1.9%. Updating bioinformatic pipelines from version 11b4 to 12.8 rescued multiple ambiguous entries, increasing overall clinical utility to 84.1%. These findings demonstrate that integrating computational epigenomics with classical neuropathology enhances diagnostic precision, while highlighting the ongoing need for careful clinical-pathological correlation.

Central nervous system tumors

Integrative genomics elucidates the evolutionary, temporal, and developmental origins of a hydrocephalus risk gene.

INTRODUCTION: A prior integrative, multi-omics human genetics and functional genomics study identified maelstrom (MAEL), a gene involved in regulation of DNA transposon activity and genome structure, as a transcriptome-wide predictor of hydrocephalus (HC) in the brain cortex. Here we expand on this discovery and further characterize the evolutionary origin and expression of MAEL across developmental timescales and cell-lineages in the neonatal human brain towards a mechanistic understanding how variation in MAEL expression may cause HC. OBJECTIVE: To characterize the evolutionary, temporal, developmental, and lineages of MAEL expression in HC and the developing human brain. METHODS: Ensembl was used to delineate the evolution and taxonomy of MAEL across species. Analysis of single-cell RNA sequencing (scRNA-seq) of 49 brain regions across pre- and post-natal timescales from the Developing Human Brain Atlas (Allen Institute) identified temporal and spatial MAEL expression patterns. We quantified MAEL expression in primary cortical brain tissue obtained during the surgical treatment of HC. RESULTS: We performed taxonomic gene-mapping to define the evolutionary origin of MAEL to assess suitability for mechanistic characterization in vitro and in vivo across species. We find that MAEL is among the top 0.01% human-specific genes and < 50% sequence homology among commonly used model organisms with highly divergent functions, necessitating mechanistic validation in human tissue. scRNA-seq of the non-disease prenatal human brain identified MAEL expression enriched in cortical excitatory neurons, which was recapitulated in primary HC brain tissue obtained during surgery. Finally, using scRNA-seq of primary HC brain tissue, we functionally validated reduced MAEL expression, consistent with a prior human TWAS analysis. CONCLUSIONS: We identify the evolutionary, temporal, and developmental expression pattern of MAEL in the neonatal human brain. We also provide direct evidence for reduced MAEL expression in human HC brain tissue. These data, at least in part, implicate reduced MAEL expression underlying human HC across etiologies.

Journal Article

Resolving taxonomic complexity in the genus Boechera (Brassicaceae) using the Boechera Microsatellite Website: a case study of the rare triploid B. bodiensis.

BACKGROUND AND AIMS: The genus Boechera (rock cress) comprises &#x223c;75 sexual diploid taxa and >355 genetically distinct hybrid lineages, many of which reproduce asexually through apomixis. This complex reproductive landscape poses substantial challenges for taxonomy, similar to those encountered in genera such as Taraxacum, Hieracium, Poa and Rubus. The Boechera Microsatellite Website (BMW) offers an extensive database and analytical tools that are proving instrumental in resolving these difficulties. Here, we demonstrate the utility of the BMW through analysis of Boechera bodiensis, a rare and poorly understood species endemic to the western Great Basin of the USA. METHODS: First described as Arabis bodiensis by Rollins in 1982, this taxon is sparsely represented in herbaria and has long been considered a candidate for protection under the Endangered Species Act. However, its taxonomic identity has remained uncertain owing to morphological similarities with other 'Arabis' (Boechera) taxa. We integrate microsatellite DNA data from the BMW with morphological analyses to provide a clearer understanding of the taxonomic status, distribution and evolutionary origins of B. bodiensis. KEY RESULTS: Pollen studies reveal that B. bodiensis is a diplosporous apomict. Microsatellite genotyping of the holotype confirms it to be triploid, containing three subgenomes derived from Boechera cobrensis, B. fernaldiana and B. sparsiflora. Expanded microsatellite surveys detect this triploid genotype at 22 additional sites, primarily in Mono County, CA, USA. Morphological analyses of genetically verified specimens identify a consistent set of characters that distinguish B. bodiensis from co-occurring congeners. CONCLUSIONS: The BMW enables high-resolution analyses of genome composition, reproductive mode and hybrid origins, making it a powerful tool for resolving taxonomic complexity in Boechera. Our case study of B. bodiensis highlights the effectiveness of combining molecular and morphological data to clarify species boundaries, inform conservation assessments and refine nomenclatural understanding in this notoriously difficult genus.

Microsatellite Repeats

In silico analysis and comparison of the metabolic capabilities of different organisms by reducing metabolic complexity.

BACKGROUND: Understanding how metabolic capabilities diverge across microbial species is essential for deciphering community function, ecological interactions, and the design of synthetic microbiomes. Despite shared core pathways, microbial phenotypes can differ markedly due to evolutionary adaptations and metabolic specialization. Genome-scale metabolic models (GEMs) provide a systems-level framework to explore these differences; however, their complexity hinders direct comparison. RESULTS: We introduce NIS (Neidhardt-Ingraham-Schaechter), a computational workflow that integrates the redGEM, lumpGEM, and redGEMX algorithms to systematically reduce genome-scale models into biologically interpretable modules. This approach enables direct, quantitative comparison of fueling pathways, biomass biosynthetic routes, and environmental exchange processes while retaining essential metabolic information. We first demonstrate the utility of NIS by analyzing Escherichia coli and Saccharomyces cerevisiae, which revealed both conserved and divergent strategies in central metabolism, biosynthetic cost, and substrate utilization. We then applied NIS to the core honeybee gut microbiome, uncovering distinct metabolic traits, functional redundancy, and complementarity that help explain auxotrophy, cross-feeding interactions, and microbial coexistence. CONCLUSIONS: NIS provides an automated, scalable, and reproducible framework for dissecting microbial metabolic networks beyond gene content or taxonomy. By linking metabolism to ecological function, NIS offers new opportunities to interpret microbial community dynamics and to support the rational design of microbiomes in health, agriculture, and environmental applications. Video Abstract.

Metabolic Networks and Pathways

metaExpertPro: A Computational Workflow for Metaproteomics Spectral Library Construction and Data-Independent Acquisition Mass Spectrometry Data Analysis.

Analysis of large-scale data-independent acquisition mass spectrometry metaproteomics data remains a computational challenge. Here, we present a computational pipeline called metaExpertPro for metaproteomics data analysis. This pipeline encompasses spectral library generation using data-dependent acquisition MS, protein identification and quantification using data-independent acquisition mass spectrometry, functional and taxonomic annotation, as well as quantitative matrix generation for both microbiota and hosts. By integrating FragPipe and DIA-NN, metaExpertPro offers compatibility with both Orbitrap and timsTOF MS instruments. To evaluate the depth and accuracy of identification and quantification, we conducted extensive assessments using human fecal samples and benchmark tests. Performance tests conducted on human fecal samples indicated that metaExpertPro quantified an average of 45,000 peptides in a 60-min diaPASEF injection. Notably, metaExpertPro outperformed three existing software tools by characterizing a higher number of peptides and proteins. Importantly, metaExpertPro maintained a low factual false discovery rate of approximately 5% for protein groups across four benchmark tests. Applying a filter of five peptides per genus, metaExpertPro achieved relatively high accuracy (F-score&#xa0;=&#xa0;0.67-0.90) in genus diversity and showed a high correlation (rSpearman&#xa0;=&#xa0;0.73-0.82) between the measured and true genus relative abundance in benchmark tests. Additionally, the quantitative results at the protein, taxonomy, and function levels exhibited high reproducibility and consistency across the commonly adopted public human gut microbial protein databases IGC and UHGP. In a metaproteomic analysis of dyslipidemia patients, metaExpertPro revealed characteristic alterations in microbial functions and potential interactions between the microbiota and the host.

Proteomics