Search PubMedSearch

SEARCH · Search PubMed

Results for “Feature selection stability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

28 records · Page 2Linked to original sources

NCBoost v2: a classifier for non-coding single-nucleotide variants in Mendelian diseases.

MOTIVATION: The current diagnostic rate of rare diseases through whole-genome sequencing has stabilized at around 30% on average, highlighting the need for improved computational scores to identify pathogenic variants. In 2019, we developed NCBoost, a supervised-learning approach that mined a comprehensive set of sequence constraint features and proved particularly well suited to identifying high-effect pathogenic non-coding variants in genetic diseases. Since its first release, the substantial increase in the number of variants available for training, as well as the enhanced capacity to detect purifying selection signals from large-scale genome sequencing projects, motivated an update of NCBoost. RESULTS: We implemented NCBoost v2, a pathogenicity score for non-coding single-nucleotide variants, trained on the largest set of curated pathogenic variants in monogenic Mendelian diseases available to date. It leverages conservation features computed from recent large-scale genomic consortia such as Zoonomia and gnomAD, and incorporates recent splice-altering predictive scores. NCBoost v2 outperformed alternative state-of-the-art methods in a variety of scenarii, providing more consistent scores across non-coding genomic regions and fine-tuning the scoring of pathogenic splice-altering variants in Mendelian disease genes. AVAILABILITY AND IMPLEMENTATION: NCBoost v2 software is implemented in Python 3.10 and is freely available under the GNU General Public License Version 3 at https://doi.org/10.5281/zenodo.16029049 and https://github.com/RausellLab/NCBoost-2, together with precomputed scores for the human genome assembly GRCh38.

Polymorphism, Single Nucleotide

Electrophrenic respiration: report of six cases.

The development of electrophrenic respiration has permitted freedom from mechanical ventilation for patients who have irreversible respiratory failure in association with high-cervical spinal cord or brainstem lesions. There are three basic criteria for successful diaphragm pacing: (1) the need for long-term mechanical ventilatory assistance, (2) a functionally intact phrenic nerve-diaphragm axis, and (3) chest wall stability. Inability to achieve satisfactory pacing can be due to malfunction of equipment, instability of the chest wall, or inadequate neuromuscular responsiveness. These features of diaphragm pacing are exemplified in a series of six patients. Three achieved independence from mechanical ventilatory assistance with full-time phrenic pacing. In one patient, only limited electrophrenic respiration was achieved, and in another the method was entirely unsuccessful. Although functioning well, pacing systems were removed from the sixth patient because of infection. Diaphragm pacing can be a valuable form of respiratory support for carefully selected patients.

Adolescent

DNA replication fidelity.

DNA replication fidelity is a key determinant of genome stability and is central to the evolution of species and to the origins of human diseases. Here we review our current understanding of replication fidelity, with emphasis on structural and biochemical studies of DNA polymerases that provide new insights into the importance of hydrogen bonding, base pair geometry, and substrate-induced conformational changes to fidelity. These studies also reveal polymerase interactions with the DNA minor groove at and upstream of the active site that influence nucleotide selectivity, the efficiency of exonucleolytic proofreading, and the rate of forming errors via strand misalignments. We highlight common features that are relevant to the fidelity of any DNA synthesis reaction, and consider why fidelity varies depending on the enzymes, the error, and the local sequence environment.

Base Pair Mismatch

Megamimivirus double-stranded DNA linear genomes flanked by highly diverse terminal inverted repeats.

UNLABELLED: Giant viruses have fundamentally expanded our understanding of virology by challenging the conventional boundaries of both virion size and genome complexity. However, the scarcity of isolates has left many of their unique biological features unexplored. Here, we report the isolation and characterization of four new giant virus species belonging to the subfamily Megamimivirinae, sampled from distinct environments across China. Among these, Megavirus daqingense is the first giant virus isolated from an oil reservoir; it exhibits virion stability under high salinity, chloroform exposure, and elevated temperatures, suggesting fitness adaptations to subsurface conditions. Using a hybrid sequencing approach that integrates short- and long-read technologies, we assembled complete linear genomes for all four isolates, each flanked by long terminal inverted repeats (TIRs). Comparative genomic and synteny analyses identified 29 distinct TIRs from 46 megamimivirus genomes. Gene content within these TIRs was highly diverse, with no orthologous proteins conserved across all repeats. Furthermore, TIR genes experienced weaker purifying selection than those in non-TIR regions (i.e., the genomic regions excluding the TIRs), consistent with their role as drivers of genome plasticity. Notably, we discovered for the first time that identical tRNA genes are shared between TIRs and non-TIR regions of eukaryotic viruses. Collectively, our work provides insights into the structural and evolutionary complexity of megamimiviruses, revealing TIRs as reservoirs of genetic diversity and hotspots for gene transfer, thereby playing a pivotal role in shaping the dynamic architecture of giant virus genomes. IMPORTANCE: Terminal inverted repeats (TIRs) are critical structural elements at the termini of linear genomes essential for fundamental processes such as recombination, replication, and integration across diverse organisms. However, the inherent limitations of short-read sequencing technologies have left the complete structure, diversity, and evolutionary significance of long TIRs in giant viruses unexplored. In this study, we leverage hybrid sequencing and comparative genomic analyses to unveil the complexity of TIRs across the subfamily Megamimivirinae. We demonstrate that TIRs are dynamic genomic hotspots characterized by remarkable gene diversity and unexpected conservation of specific tRNA genes. These findings establish TIRs as key drivers of genome plasticity, serving as hotspots for horizontal gene transfer and genetic innovation. By resolving the long-hidden terminal structures of megamimivirus genomes, this work provides a foundational framework for understanding how TIRs shape the evolution of giant viruses and, more broadly, advances our understanding of genome architecture in large DNA viruses.

Megavirus

High early death rates, treatment resistance, and short survival of Black adolescents and young adults with AML.

Survival of patients with acute myeloid leukemia (AML) is inversely associated with age, but the impact of race on outcomes of adolescent and young adult (AYA; range, 18-39 years) patients is unknown. We compared survival of 89 non-Hispanic Black and 566 non-Hispanic White AYA patients with AML treated on frontline Cancer and Leukemia Group B/Alliance for Clinical Trials in Oncology protocols. Samples of 327 patients (50 Black and 277 White) were analyzed via targeted sequencing. Integrated genomic profiling was performed on select longitudinal samples. Black patients had worse outcomes, especially those aged 18 to 29 years, who had a higher early death rate (16% vs 3%; P=.002), lower complete remission rate (66% vs 83%; P=.01), and decreased overall survival (OS; 5-year rates: 22% vs 51%; P<.001) compared with White patients. Survival disparities persisted across cytogenetic groups: Black patients aged 18 to 29 years with non-core-binding factor (CBF)-AML had worse OS than White patients (5-year rates: 12% vs 44%; P<.001), including patients with cytogenetically normal AML (13% vs 50%; P<.003). Genetic features differed, including lower frequencies of normal karyotypes and NPM1 and biallelic CEBPA mutations, and higher frequencies of CBF rearrangements and ASXL1, BCOR, and KRAS mutations in Black patients. Integrated genomic analysis identified both known and novel somatic variants, and relative clonal stability at relapse. Reduced response rates to induction chemotherapy and leukemic clone persistence suggest a need for different treatment intensities and/or modalities in Black AYA patients with AML. Higher early death rates suggest a delay in diagnosis and treatment, calling for systematic changes to patient care.

Adolescent

Genomic Insights Into Convergent Evolution: Adaptation to Rocky Habitats in Rock-Inhabiting Fungi.

Rock-inhabiting fungi (RIF), obligate colonizers of bare rocks, are primarily distributed across two major phylogenetic classes: Dothideomycetes and Eurotiomycetes. These fungi display striking convergence in morphology and physiology, characterized by meristematic growth, melanized cell walls, and extreme stress tolerance. However, the genomic underpinnings of this adaptive convergence remain poorly understood. Here, through comparative genomic analysis of 9 RIF and 18 non-RIF fungi, we revealed that RIF possess compact, gene-dense genomes marked by contraction of genes involved in nutrient uptake and secondary metabolism, alongside expansions in cell wall biosynthesis, lipid metabolism, and stress-responsive pathways. We identified two genes under positive selection across multiple RIF lineages: Ino80 ATPase (chromatin remodeling) and the ER chaperone BiP (protein folding). Further evidence of convergence was found in the mannosyltransferase Mnn9, a key enzyme in cell wall assembly, where two RIF-specific amino acid substitutions were predicted to enhance protein stability. Additionally, a unique Mnn9-like clade has expanded exclusively in RIF. RNAi-mediated knockdown of an Mnn9-like gene in Rachicladosporium sp. confirmed its role in cell wall mannosylation, osmotic stress response, and the transition from meristematic to filamentous growth. Our findings elucidate a set of common genomic adaptations and highlight the specialized evolution of the Mnn9 family in driving the convergent success of phylogenetically diverse RIF in rocky environments.

Phylogeny

Discovery and validation of a multi-protein panel for predicting non-fatal major adverse cardiovascular events in diabetic kidney disease.

OBJECTIVE: To identify plasma protein biomarkers associated with incident non-fatal major adverse cardiovascular events (MACE) in diabetic kidney disease (DKD) patients. RESEARCH DESIGN AND METHODS: We analyzed 317 DKD patients from the UK Biobank. Plasma proteomics and clinical data (demographics, metabolism, renal function) were integrated. In an exploratory discovery phase, three sequential Cox regression models (crude, socio-demographic-adjusted, socio-demographic-metabolic adjusted) screened non-fatal MACE-associated proteins. To prevent information leakage, the cohort was then randomly split into training (70%) and testing (30%) sets; machine-learning feature selection, hyperparameter optimization, and final model development were performed exclusively within the training set. The associated proteins were input into the four-step machine-learning pipeline (LASSO-Cox, random survival forest, Boruta, XGBoost-Cox). Predictive performance was validated using Kaplan-Meier survival analyses, longitudinal trajectory modeling, and ROC benchmarking. An interactive web application was deployed for clinical implementation. RESULTS: Of 1,463 plasma proteins, 561 were associated with non-fatal MACE across Cox models, with 14 overlapping proteins. Nine core proteins (ANG, IL1R1, CXCL14, ESAM, PTGDS, HAVCR1, FGFR2, IGSF8, CCL3) were validated: ANG showed the strongest non-fatal MACE association (HR&#xa0;=&#xa0;3.88, 95%CI 2.33-6.48, p<0.001), and all high-expression groups had elevated non-fatal MACE risk. GO/KEGG enrichment highlighted inflammatory-immune pathways like positive regulation of MAPK cascade, Cytokine-cytokine receptor interaction and PI3K-Akt signaling pathway as key mechanisms. The model integrating proteins, demographic factors, and clinical variables achieved the highest predictive performance across non-fatal MACE (AUC&#xa0;=&#xa0;0.768), myocardial infarction (MI) (0.808), and stroke (0.816) outcomes, with superior stability in cross-validation. CoxBoost + Elastic Net framework was selected as the optimal framework via benchmarking of 101 algorithms. The model demonstrated favorable calibration in high-risk patients and yielded positive net clinical benefit across decision thresholds of 5% to 45%. The web tool (https://jiangli2941.github.io/MACE-prediction-v2/) enables input of 28 variables, outputs non-fatal MACE risk status, risk probability, and highlights abnormal indicators. CONCLUSION: Plasma proteomics combined with machine learning identifies robust non-fatal MACE predictors in DKD.

Humans

Causal circuit tracing reveals distinct computational architectures in single-cell foundation models: inhibitory dominance, biological coherence, and cross-model convergence.

MOTIVATION: Sparse autoencoders (SAEs) decompose foundation-model activations into interpretable features, but the model-internal causal interactions between those features (i.e. what ablating one feature does to the others, as distinct from the biological causal structure of the underlying cells)-and how those model-internal relationships relate to biological structure-are uncharacterized in single-cell foundation models. RESULTS: We introduce model-internal causal circuit tracing-zeroing one SAE feature at a source layer and measuring the resulting change in all downstream SAE features, for each of 120 source features-and apply it to Geneformer V2-316M and scGPT whole-human across four conditions (96&#xa0;892 ablation-derived edges, 80&#xa0;191 forward passes). On annotation-selected source features, edges share GO/KEGG/Reactome/STRING/TRRUST ontology terms at 50.9%-68.5%, a 2.9-6.2&#xd7; enrichment over a configuration-preserving permutation null (P<.002); on 20 randomly sampled source features this attenuates to 21.5%-26.3%-still 2.5-3.1&#xd7; above null-quantifying the annotation-selection contribution. Inhibitory dominance (fraction of ablation edges with d<0, i.e. source activation supports downstream target) is 65.5%-89.4%. scGPT produces larger raw per-edge effects (mean |d|=1.40 versus 1.05); after feature-share normalization, Geneformer is stronger (paired gene-pair ratio 0.64 on 33&#xa0;301 shared pairs). Cross-model consensus yields 1142 architecture-invariant domain pairs (ordered pairs of GO biological-process categories "A&#x2192;B" each connected by at least one ablation edge in both models; 10.6&#xd7; enrichment over permutation null; P<.001). Circuit edge magnitude explains <1% of the variance in marginal driver-gene coexpression on the same cells (R2=0.010, n=31&#xa0;176): the graph encodes structure beyond bivariate correlation. Against a matched-cell-type ENCODE ChIP-seq prior, circuit-predicted transcription factor (TF)&#x2192;target pairs are enriched 2.06&#xd7; (Fisher OR 5.84), markedly higher than 1.12&#xd7; against TRRUST; direct ChIP-seq-supported target pairs show 10-30&#xd7; larger CRISPRi sign-bias-corrected excess than indirect pairs. Gene-level CRISPRi validation on Replogle K562 and the noncancer RPE1 arm (and a true primary-T-cell control from Shifrut E, Carnevale J, Tobin V et&#xa0;al. Genome-wide CRISPR screens in primary human T cells reveal key regulators of immune function. Cell 2018; 175: 1958-71.e15) after sign-bias correction shows excess over baseline of +0.03 and +0.35 percentage points on K562 and RPE1, respectively (baseline already 52%-56% from sign marginals); effect-magnitude Spearman correlations &#x3c1;&#x2248;0. Bootstrap and per-cell-type stability (N&#x2208;{50,100,200}; B cell, CD4&#xa0;+ T, macrophage) give Pearson r&#x2265;0.97 on shared edges with 100% sign agreement; edge Jaccard grows monotonically with sample size. The circuit graph is therefore highly reproducible as an effect-size map, cell type specific in edge identity, consistent with coexpression encoding, and weakly but detectably enriched for ChIP-seq-supported direct regulatory edges. AVAILABILITY AND IMPLEMENTATION: https://github.com/Biodyn-AI/bio-sae-circuits (Python). Archival DOI: 10.5281/zenodo.19,633,166 (Zenodo).

Humans

Molecular characterization of pESI-like megaplasmids in Salmonella Infantis from poultry in Lebanon.

UNLABELLED: Salmonella enterica serovar Infantis has emerged as a globally disseminated multidrug-resistant (MDR) pathogen, largely driven by the spread of the plasmid of emerging Salmonella Infantis (pESI)-like megaplasmid. In our study, we investigated the prevalence, antimicrobial resistance (AMR) phenotypes, and genomic features of S. Infantis isolates collected from poultry farms in Lebanon. A total of 72 isolates were recovered during a nationwide surveillance effort, among which 67 (93%) were MDR based on antimicrobial susceptibility testing (disk diffusion and broth microdilution) results, including resistance to critically important agents such as quinolones, and highly important classes such as tetracyclines and sulfonamides. Whole-genome sequencing was performed on 19 isolates selected through a stratified approach to encompass all identified AMR phenotypes; this analysis revealed a conserved pESI-like backbone together with MDR-associated determinants, including sul1, tet(A), and aadA. Plasmid marker analysis confirmed the presence of pESI in the majority of isolates, with plasmid-associated genes (ardA and trbA) and replicon markers (IncP and IncFIB(pN55391)) among the most prevalent. Comparative plasmid alignments with representative pESI sequences from Italy, Turkey, and the United States revealed strong conservation of the backbone alongside regional variation in AMR gene content. These findings highlight the role of poultry production systems in Lebanon as reservoirs for pESI-like megaplasmids and MDR S. Infantis, underscoring the zoonotic and public health risks posed at the human-animal-environment interface. Strengthened surveillance, antimicrobial stewardship, and biosecurity interventions are urgently needed to mitigate the spread of MDR S. Infantis within agriculture and beyond. IMPORTANCE: The emergence of plasmid of emerging Salmonella Infantis (pESI)-like megaplasmids has transformed Salmonella Infantis into a globally distributed multidrug-resistant (MDR) clone with the capacity to persist in livestock and disseminate resistance genes across ecological boundaries. Our study provides the first genomic characterization of pESI-positive S. Infantis from poultry farms in Lebanon, a region with high antimicrobial usage and limited stewardship frameworks. By integrating phenotypic susceptibility testing and whole-genome sequencing, we demonstrate that Lebanese isolates harbor conserved pESI-like backbone markers together with antimicrobial resistance determinants, aligning them with internationally circulating lineages. Comparative analysis with isolates from Italy, Turkey, and the United States highlights both the evolutionary stability and geographic diversity of pESI. These findings emphasize the urgent need for integrated surveillance and stewardship strategies to curb the spread of MDR S. Infantis and reduce the zoonotic risk at the human-animal-environment interface.

Animals

MWENA: a novel sample re-weighting-based algorithm for disease classification and data interpretation using extracellular vesicles omics data.

BACKGROUND AND OBJECTIVE: Extracellular vesicles (EVs), considered as a form of liquid biopsy, have gained significant attention in recent years due to their stability and the preservation of disease markers. Research studies underscore the clinical significance of molecules found in EVs, highlighting their role as communicative mediators between cells. However, analyzing this data is challenging due to noisy measurements, having far more variables than samples, and some groups (e.g., disease subtypes or experimental conditions) having much less data than others. We therefore develop an algorithm to address aforementioned challenges for the classification of imbalanced EVs omics data. METHODS AND RESULTS: We propose the EV Meta-Weight Elastic Net Algorithm (MWENA), which utilizes logistic regression with elastic net regularization for the classification and identification of EV signatures, effectively addressing the challenges posed by high-dimensional small sample sizes. To mitigate issues related to class imbalance and high noise levels, MWENA incorporates an automatic sample re-weighting function, which uses a meta-net to adaptively learn generalizable patterns directly from the data itself. We validate the MWENA algorithm on both simulated data and EVs omics data, covering six classification tasks that involve four different types of diseases (pancreatic ductal adenocarcinoma, interstitial lung diseases, colorectal cancer, and ovarian cancer) and three clinical scenarios (disease diagnosis, disease-stage screening, and disease-subtype classification). Compared to other machine learning methods, MWENA demonstrates superiority in identifying small class samples and achieves the highest scores in both sensitivity and G-means. Biological analysis is also performed to further explore the significance of selected signatures as biological markers and their roles in disease mechanisms. CONCLUSIONS: We anticipate that our proposed approach will take a modest step in harnessing EV omics data to discover biomarkers, aiding researchers in gaining a comprehensive understanding of biological processes.

Extracellular Vesicles