Search PubMedSearch

SEARCH · Search PubMed

Results for “locus robustness”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Engineering a probiotic Bacillus subtilis for acetaldehyde removal: A hag locus integration to robustly express acetaldehyde dehydrogenase.

We have addressed critical challenges in probiotic design to develop a commercially viable bacterial strain capable of removing the intestinal toxin, acetaldehyde. In this study, we report the engineering of the hag locus, a σD-dependent flagellin expression site, as a stable location for robust enzyme production. We demonstrate constitutive gene expression in relevant conditions driven by the endogenous hag promoter, following a deletion of the gene encoding a post-translational regulator of σD, FlgM, and a point mutation to abrogate the binding of the translational inhibitor CsrA. Reporter constructs demonstrate activity at the hag locus after germination, with a steady increase in heterologous expression throughout outgrowth and vegetative growth. To evaluate the chassis as a spore-based probiotic solution, we identified the physiologically relevant ethanol metabolic pathway and the subsequent accumulation of gut-derived acetaldehyde following alcohol consumption. We integrated a Cupriavidus necator aldehyde dehydrogenase gene (acoD) into the hag locus under the control of the flagellin promoter and observed a rapid reduction in acetaldehyde levels in gut-simulated conditions post-germination. This work demonstrates a promising approach for the development of genetically engineered spore-based probiotics.

Acetaldehyde

Deconstructing empirical fitness seascapes across scales of granularity.

The fitness landscape metaphor remains resonant in evolutionary theory and has facilitated the birth of newer concepts, like the fitness seascape, that consider the role of environmental context in shaping the dynamics of evolution. Since its emergence, the seascape has appeared in numerous studies examining how different and fluctuating environments shape evolutionary outcomes. Despite growing interest, we lack comprehensive examinations of how environmental context shapes features of fitness seascapes. In this study, we address this gap by deconstructing empirical fitness seascapes across scales of granularity: loci, locus interactions (epistasis), alleles, trajectories, and entire seascapes. For each, we examine how environmental context influences qualitative and quantitative aspects of seascapes, and find that they change appreciably, with patterns specific to individual systems of study. We also quantify how much each scale varies across environments, and find that certain scales tend to be more sensitive to context than others. In summary, we reflect on the implications of the seascape metaphor for the incorporation of environmental effects into theoretical population genetics, for understanding how the environment shapes evolution in disease systems, and for contemporary bioengineering efforts.

Genetic Fitness

Characterization of a PRKCE::ETV6 fusion as a potential oncogenic driver in T-cell acute lymphoblastic leukemia.

BACKGROUND: T-cell acute lymphoblastic leukemia (T-ALL) is an aggressive hematologic malignancy caused by mutation accumulation during hematopoiesis. The characterization of chromosomal abnormalities may provide significant insights into genetic mechanisms of malignant transformation in hematopoietic cells. However, T-ALL is genetically very heterogenous and driving mutations as well as clonal markers for the assessment of minimal residual disease are not always identifiable. Hence, there is a clinical need to further refine the genetic landscape of T-ALL including previously unrecognized fusion partners of commonly translocated genes in T-ALL of childhood. RESULTS: In this study, we screened n = 229 T-ALL cases by our targeted genomic capture high-throughput sequencing (gc-HTS) approach. In total, we identified n = 60 gene–gene fusions, present in n = 57 (25%) of the patients. Nine rare or even unrecognized translocations were identified and validated. Furthermore, owing to its interesting chromosomal structure, we studied the oncogenic potential of the complex rearrangement of chromosome 2 and 12, found in a near-early T-cell progenitor (ETP) ALL that leads to the fusion events PRKCE::ETV6 and ETV6::INO80D. Exogenous expression of PRKCE::ETV6 in Ba/F3 pro-B and D1 T-cells caused interleukin-independent proliferation and enhanced survival upon interleukin withdrawal, respectively. CONCLUSION: Our study underlines the heterogenous mutational landscape in T-ALL. The previously unrecognized PRKCE::ETV6 resulting from a complex rearrangement involving chromosome 2 and 12 demonstrated transforming potential in cytokine-dependent cellular models support the notion of a driver mutation in near ETP-ALL. Our data reconfirm the relevance of ETV6-fusion proteins in the pathogenesis of undifferentiated T-ALL. Importantly, genomic breakpoints at the ETV6 locus represent potentially robust MRD markers for (near) ETP-ALL that lack IG/TR rearrangements.

ETV6::INO80D

Reduced R-loop abundance at proinflammatory loci: a shared epigenetic mechanism in inflammatory and metabolic diseases.

INTRODUCTION: R-loops, RNA-DNA hybrid structures with a displaced single-stranded DNA loop, are key regulators of transcriptional control, chromatin architecture, and genome stability and have emerging roles in inflammatory signaling. However, the relationship between R-loop abundance and strongly modulated inflammatory effector genes in metabolic inflammation and influenza virus infection remains underexplored. METHODS: We performed a locus-centric integrative analysis combining robust differentially expressed genes (DEGs) from multiple inflammatory and infection-related murine and human transcriptomic disease models with experimentally validated multi-cell R-loop annotations from the reference atlas RLoopBase. Our correlation framework evaluated the directional relationship between R-loop abundance and inflammatory gene expression rather than assuming disease-sample-matched R-loop measurements. We further analyzed R-loop regulatory proteins, NRF2-associated R-loop regulators, and overlaps between R-loop regulators and CRISPRi-identified mitochondrial and cellular reactive oxygen species (ROS) regulators. RESULTS: In angiotensin II-infused apolipoprotein E-deficient (ApoE-/-) mice, a model of abdominal aortic aneurysm (AAA), genomic regions encoding the top significantly upregulated genes exhibited significantly fewer R-loops than those encoding downregulated genes at days 14 and 28. Similarly, in atherosclerotic ApoE-/- mice fed a high-fat diet for 32 and 78 weeks, upregulated genes were associated with fewer R-loops than downregulated genes. Reduced R-loop abundance was also observed in genomic regions encoding the top significantly upregulated genes in liver tissues from patients with non-alcoholic steatohepatitis (NASH), as well as in monosodium urate (MSU)-stimulated lymphatic endothelial cells (LECs) and influenza virus-infected human umbilical vein endothelial cells (HUVECs). R-loop regulatory proteins upregulated during metabolic inflammation were enriched in immune and inflammatory pathways. NRF2 was identified as a regulator of 27 R-loop regulatory proteins, including 10 positively and 17 negatively regulated proteins. Furthermore, 54 R-loop regulatory proteins overlapped with CRISPRi-identified mitochondrial and cellular ROS regulators, suggesting potential reciprocal regulation between R-loop homeostasis and ROS signaling. Disease-associated changes in pro-ROS and anti-ROS R-loop regulatory proteins further linked R-loop regulation to inflammatory and oxidative stress pathways. DISCUSSION: These findings identify reduced R-loop abundance at genomic regions encoding strongly upregulated inflammatory genes as a shared feature across multiple models of metabolic inflammation and influenza virus infection. The results further suggest that immune-associated R-loop regulatory proteins and the NRF2-ROS axis may contribute to R-loop remodeling during inflammatory disease. This integrative framework provides new insight into the potential role of R-loops and ROS-sensitive R-loop regulators in inflammatory and metabolic diseases and identifies candidate pathways for future mechanistic investigation and therapeutic targeting.

R-loop regulatory proteins

Versatile, marker-free platform for life cycle-wide imaging of Plasmodium falciparum by integrating an exogenous gene cassette into a conserved intergenic locus.

The creation of transgenic Plasmodium falciparum lines with robust fluorescence across the entire life cycle is essential for advancing our understanding of parasite biology, which in turn informs the development of new drugs and vaccines. In this study, we utilized Plasmodium-optimized genome editing to integrate an mCherry expression cassette into a selected intergenic locus without gene disruption. The resulting marker-free line, NF54-mCh, exhibited intense fluorescence throughout all developmental stages, including asexual and sexual blood stages, as well as mosquito (ookinete, oocyst, and sporozoite) and liver stages. NF54-mCh showed normal proliferation, gametocytogenesis, and efficient transmission to mosquitoes. The ultra-high brightness in salivary gland sporozoites allowed for the non-invasive identification of infected mosquitoes. Sporozoites remained highly infectious to humanized mouse livers, thus enabling the completion of the full life cycle. NF54-mCh serves as a parental line for performing additional genetic modifications, because the CRISPR/Cas9-based genome editing method is free of introduced drug resistance markers. The broader applicability of this strategy was validated by generating similar reporter lines in Plasmodium species utilized in rodent malaria models. In summary, NF54-mCh represents a unique, versatile platform that will accelerate fundamental research and support the future development of malaria control strategies, including new vaccines and drugs.

Animals

Genome-wide association study of copy number variations in Parkinson's disease.

OBJECTIVE: To investigate the impact of copy number variations (CNVs) on Parkinson's disease (PD) pathogenesis using genome-wide data and explore their role in sporadic PD. METHODS: We analyzed CNV data from 11,035 PD patients (including 2,731 early-onset PD (EOPD)) and 8,901 controls from the COURAGE-PD consortium using a sliding window CNV-GWAS and genome-wide burden analysis. The independent dataset from the Global Parkinson Genetics Program (GP2) consisted of 23,089 cases and 18,824 controls were used to validate our initial findings. RESULTS: The exploratory dataset identifies multiple CNV regions associated with PD risk. The nominated CNV loci were not confirmed in an independent dataset, except that only a deletion in the PRKN gene, a well-established EOPD locus, remained genome-wide significant and robustly supported. CNV burden analysis showed a higher prevalence of CNVs in PD-related genes in patients compared to controls (OR=1.56 [1.18-2.09], p=0.0013), with PRKN showing the highest burden (OR=1.47 [1.10-1.98], p=0.026). Patients with CNVs in PRKN had an earlier disease onset. Burden analysis with controls and EOPD patients showed similar results. INTERPRETATION: The largest CNV-based GWAS on PD highlights both the promise and pitfalls of array-based CNV detection in PD and underscores the relevance of whole-genome sequencing approaches in resolving the role of CNV in PD. The array-based findings are prone towards false positive findings that might arise either from platform limitations and/or cohort biases. Future studies require improved genotyping resolution and rigorous cross-cohort validation to reliably assess CNV contributions to PD risk.

Journal Article

DNA-FISH Metaphase Spreads to Distinguish Extrachromosomal DNA from Homogeneously Staining Regions in Human Cancer Cell Lines.

UNLABELLED: Whole-genome sequencing identifies focal DNA amplifications with base-pair resolution but cannot determine whether amplified sequences reside on extrachromosomal DNA (ecDNA, also known as double minutes) or within chromosomally integrated homogeneously staining regions (HSRs). DNA fluorescence in situ hybridization (DNA-FISH) metaphase spreads remain the gold standard for distinguishing these amplification states at single-cell resolution. Here, we present a detailed protocol for DNA-FISH metaphase spreads using human cancer cell lines, encompassing cell culture, metaphase arrest, hypotonic treatment, fixation, chromosome spreading, fluorescent probe hybridization, and fluorescence imaging. The protocol incorporates intermediate quality-control steps to verify successful chromosome dispersion and optimize metaphase spread quality, making the workflow accessible to laboratories without specialized cytogenetics expertise. Results demonstrate clear visualization of ecDNA and HSR amplification states using locus-specific probes and illustrate common technical artifacts that can affect interpretation. This protocol provides a robust and reproducible approach for studying the structural organization of oncogene amplification in cancer cells. SUMMARY: We report a DNA-FISH metaphase spread protocol that visually detects locus copy number and location within the genome. This approach enables single-cell resolution of amplification states, specifically in cancer cell lines containing extrachromosomal DNA and homogeneously staining regions.

Journal Article

TPMM: three-component posterior mixture model enables robust inverton detection in low-depth metagenomes and suggests potential viral invertons.

SUMMARY: Bacterial phase variation enables reversible, locus-specific phenotypic switching, often driven by DNA inversion (invertons). To identify these events, researchers commonly rely on sequencing reads that provide orientation-specific support. Metagenomic sequencing, which captures total genetic material independent of cultivation, offers a powerful platform for the comprehensive study of invertons. However, computational inverton calling from metagenomic data is difficult at low sequencing depth: hard read-support cutoffs can miss true events, while sequence-only predictors lack read-backed interpretability and uncertainty quantification. To address this, we present TPMM, a three-component posterior mixture model for inverton calling in metagenomic data. TPMM explicitly incorporates sequencing depth to formulate inverton detection as a probabilistic mixture problem. Starting from candidates flanked by inverted repeats, the model classifies the candidates into noise, low-probability, or high-probability inversion signals using read evidence. Finally, TPMM assigns posterior probabilities as soft labels and applies cumulative Bayesian False Discovery Rate control to robustly identify true invertons. On two real gut metagenomic datasets, TPMM agrees well with PhaseFinder at high depth but recovers substantially more invertons under systematic downsampling, demonstrating superior performance in sparse-data regimes. We further examine potential reversible inversion elements in viral genomes and provide supporting analyses, suggesting a broader scope for inversion-mediated regulation. AVAILABILITY: The source code of TPMM is available via: https://github.com/KennyxxD/TPMM.

Metagenomics

Development and application of a novel beta-tubulin genotyping tool reveals host-specific transmission cluster in Balantioides coli.

Balantioides coli is a zoonotic ciliated protozoan that infects humans and other mammals. Conventional and ITS-based genotyping approaches have limitations that hinder precise molecular epidemiological investigations. The objective of this study was to develop a new β-tubulin gene-based approach to enhance the detection and genotyping of B. coli. We performed single-cell isolation and whole-genome sequencing on two B. coli isolates from pigs and two from guinea pigs. We then used the β-tubulin gene sequences to design PCR primers for the new genotyping assay. We validated the assay using 56 ITS-confirmed B. coli-positive fecal DNA samples from pigs, cattle, sheep, and guinea pigs. Phylogenetic analyses were conducted using both β-tubulin and ITS sequences. The β-tubulin-based nested PCR assay exhibited 100% detection efficiency and greater specificity than ITS-based methods. Phylogenetic analysis of the β-tubulin gene sequences classified B. coli into three genotypes (I-III). Genotype III appears to be specific to guinea pigs. Genotypes I and II were found across multiple hosts, indicating potential cross-species transmission. Of the five full-length B. coli β-tubulin sequences obtained in this study, 264 polymorphic sites (19.8%) were identified, including both synonymous and non-synonymous mutations. Frequent recombination events within the β-tubulin locus were detected, indicating substantial genetic diversity. Therefore, the β-tubulin gene is a robust marker for genotyping and epidemiological studies of B. coli. The novel nested PCR assay overcomes the limitations of ITS-based methods and has produced data revealing previously unrecognized genetic diversity and host specificity patterns of B. coli.

Tubulin

Genome-Wide Differentiation, Inbreeding, and Candidate Selection Loci in Local Vietnamese Pig Breeds.

Vietnam harbors exceptional genetic diversity among at least 26 indigenous pig breeds. We analyzed genome-wide single-nucleotide polymorphism (SNP) data from 90 animals representing 15 local Vietnamese breeds and six Landrace pigs using principal component analysis, the windowed fixation index (FST), cross-population extended haplotype homozygosity (XP-EHH), within-population integrated haplotype score (iHS), and runs of homozygosity (ROHs). The population structure was consistent with a north-south differentiation axis, and Ba Xuyen showed elevated heterozygosity, providing suggestive evidence of a European genetic contribution; the f3 statistic was positive (f3 = +0.015), and formal evidence of admixture requires a significantly negative f3, so this criterion was not met. Integration of FST and XP-EHH identified GPC5, E2F6, NOS1, and TLR4 as top Northern candidate loci and CRYM/ZP2 as the leading Central candidate locus, and these windows were recovered at both the 90th and 95th percentile thresholds, indicating analytical robustness rather than independent biological validation. iHS was elevated at E2F6 in Northern breeds (|iHS| = 3.04) and at NOS1 across all regional groups (|iHS| = 2.66-3.36). Breed-level phenotypic XP-EHH, based on published breed descriptions and coat color rather than individual body-composition measurements, identified GALNT2 as a candidate shared across breed groups; HCAR1 and ATG10 as candidates specific to the extreme-fat/prolific breed group; and EFNA5 and HIPK2 as candidates specific to the medium-bodied breed group. ROHs identified Soc, Co, and Hung as breeds warranting particular attention in conservation planning due to elevated autozygosity. Because each breed was represented by only six individuals, and because no individual-level phenotypic measurements were available, all findings are reported as exploratory population-genomic signals requiring replication in larger cohorts. Overall, we describe genomic differentiation and candidate selection signatures among local Vietnamese pig breeds and provide a foundation for further genomic studies of these breeds.

Animals

DNA-FISH Metaphase Spreads to Distinguish Extrachromosomal DNA from Homogeneously Staining Regions in Human Cancer Cell Lines.

Whole-genome sequencing identifies focal DNA amplifications with base-pair resolution but cannot determine whether amplified sequences reside on extrachromosomal DNA (ecDNA, also known as double minutes) or within chromosomally integrated homogeneously staining regions (HSRs). DNA fluorescence in situ hybridization (DNA-FISH) metaphase spreads remain the gold standard for distinguishing these amplification states at single-cell resolution. Here, we present a detailed protocol for DNA-FISH metaphase spreads using human cancer cell lines, encompassing cell culture, metaphase arrest, hypotonic treatment, fixation, chromosome spreading, fluorescent probe hybridization, and fluorescence imaging. The protocol incorporates intermediate quality-control steps to verify successful chromosome dispersion and optimize metaphase spread quality, making the workflow accessible to laboratories without specialized cytogenetics expertise. Results demonstrate clear visualization of ecDNA and HSR amplification states using locus-specific probes and illustrate common technical artifacts that can affect interpretation. This protocol provides a robust and reproducible approach for studying the structural organization of oncogene amplification in cancer cells.

Humans

Local ancestry inference identifies robust evidence of selection in Neolithic Europe.

During the European Neolithic, migrating Anatolian farmers admixed with local hunter-gatherers, coinciding with major shifts in diet, environment, and lifestyle that imposed strong selective pressures. Local ancestry inference is widely used to detect selection following admixture, but most methods were developed and validated on present-day populations. Their performance in ancient DNA - where reference panels are smaller, data are sparser, and admixture is more ancient - remains unresolved. We benchmark eight local ancestry inference methods on 176 imputed Neolithic genomes. While individual-level ancestry estimates are highly correlated across methods, inferred tract lengths and admixture time estimates vary by an order of magnitude. Overall, we recommend Gnomix or RFMix for general use. We also investigated our ability to detect natural selection using LAI. Integrating results across methods and replicating across methods and in two independent datasets (n=378 and 1,121) we identify a robust ancestry deviation at FADS1/2, consistent with adaptation on metabolism. We also identify IRAK4 (innate immunity) as a candidate locus, but with less consistent signal across methods. Finally, we replicate previous reports of excess hunter-gatherer ancestry at the HLA, but these results are inconsistent across methods and suggest that they may be affected by bias in local ancestry inference. Our findings demonstrate that while local ancestry inference recovers biologically meaningful signals in ancient genomes, results can be sensitive to the methods used for inference, particularly in complex regions like the HLA. Method choice critically influences inferred ancestry patterns and selection signals, underscoring the importance of multi-method validation.

Journal Article

Humanizing acidic mammalian chitinase variants establish lung immune conditioning and control environmentally driven inflammation and fibrosis.

Chitin, a widespread environmental particle constituent, triggers lung inflammation but is degraded by chitinases. In humans, single-nucleotide polymorphisms (SNPs) in CHIA (acidic mammalian chitinase; AMCase) are associated with lung disease, suggesting that chitinase variants influence responses to airborne particles. Here, we edit the mouse Chia1 locus to generate humanized (hChia) mice harboring common human SNPs. Compared with controls expressing disease-protective SNPs, hChia mice lack robust chitinase activity and fail to degrade natural chitin substrates. Lung-resident lymphocytes and macrophages are spontaneously primed and sensitive to inflammatory triggering by environmental chitin. Immune cell infiltration correlates with airway chitin following challenge, and hChia mice exhibit exacerbated inflammatory and fibrotic lung disease. In humans with acute respiratory failure, alveolar hemorrhage coincides with environmentally derived chitin particles that are susceptible to chitinase degradation, attenuating inflammatory cell responses. Thus, environmental chitin and chitinase activity are crucial determinants of lung immune conditioning with potential therapeutic applications.

AMCase

A standardized, genome-guided MLST scheme for Avibacterium paragallinarum: enhanced epidemiological typing and validation against existing methods.

Avibacterium paragallinarum, the causative agent of infectious coryza (IC), is an important respiratory pathogen of chickens with growing prevalence in commercial and backyard flocks. Current strain-typing methods, including classical serotyping and molecular approaches, such as ERIC-PCR or single-locus HPG2 typing, lack sufficient discriminatory power to investigate the epidemiology or population structure. To address this limitation, we developed a genome-guided multilocus sequence typing (MLST) scheme as a robust and portable tool for A. paragallinarum strain differentiation. Housekeeping genes were identified from 42 whole-genome sequences (WGS); 18 candidates were evaluated; and six were selected for the final MLST scheme. We used the scheme to differentiate 75 A. paragallinarum samples and compared its performance against classical HPG2-based typing, ad hoc core genome MLST (cgMLST), and the MLST scheme published by M. Guo, Y. Jin, H. Wang, X. Zhang, and Y. Wu (Vet Sci 11:208, 2024, https://doi.org/10.3390/vetsci11050208). The new MLST showed higher discriminatory power than HPG2 and outperformed Guo's scheme with higher discriminatory power, particularly for characterizing the samples originating from North and South America. It also showed strong concordance with cgMLST clustering while being more practical for routine use. Overall, the six-locus MLST identified 31 sequence types across 75 samples, revealing epidemiologically meaningful clustering at regional and national scales and capturing temporal persistence of lineages. All allele definitions and sequence types have been deposited in PubMLST, ensuring standardized nomenclature and global accessibility. This scheme represents a reproducible, cost-effective, and globally applicable tool that enhances outbreak investigation, surveillance, and population studies of A. paragallinarum, bridging the gap between low-resolution traditional methods and resource-intensive whole-genome sequencing.IMPORTANCEInfectious coryza (IC) caused by Avibacterium paragallinarum is a major respiratory disease of poultry that causes acute infection, reducing egg production and growth and resulting in significant economic losses in poultry production worldwide. Controlling IC depends on understanding how different strains spread and persist, yet current methods to differentiate strains are either unreliable or too costly for routine use. In this study, we developed a standardized multilocus sequence typing system that provides a simple, accurate, and globally accessible way to identify and compare strains of A. paragallinarum. This scheme identified important links between outbreaks at local and regional levels and showed that certain strains persisted over time. By making the scheme available through PubMLST, laboratories worldwide can use a common tool to track and investigate the pathogen. This accessible tool improves disease surveillance, supports outbreak investigations, and helps poultry producers and veterinarians respond more effectively to IC.

Multilocus Sequence Typing

Deciphering the Genetic Underpinnings of Liver Cirrhosis-Heart Failure Comorbidity Through Multi-Omics: CRIM1 as a Key Endothelial Mediator.

The co-occurrence of liver cirrhosis (LC) and heart failure (HF) poses considerable clinical challenges, yet the cellular and molecular determinants of this comorbidity remain poorly characterized. To address this, we developed an integrative multi-omics pipeline encompassing GWAS meta-analysis, gsMap-based spatial transcriptomic projection, GeneEnrich functional annotation, single-cell atlas construction, seismicGWAS and ECLIPSER cell-type scoring, eCAVIAR and fastenloc colocalization, hdWGCNA network inference, scTenifoldKnk in silico gene perturbation, and GCTA-COJO fine-mapping. Quality-controlled meta-analysis yielded 12,347,758 and 9,256,862 variant-level associations for LC and HF, respectively. Spatial projection confirmed preferential enrichment of disease signals within embryonic hepatic and cardiac compartments. Pathway analyses disclosed that LC-linked loci were concentrated in lipid metabolic programs, whereas HF-linked loci implicated mitochondrial bioenergetics and lysosomal degradation. At the cellular level, endothelial cells emerged as the dominant HF-associated population. Convergent evidence from five orthogonal algorithms pinpointed CRIM1 as the sole robustly supported shared gene, selectively enriched in HF endothelial cells; virtual perturbation further identified LCP1 and PTPRC as downstream regulatory nodes. Fine-mapping of the chromosome 2 locus harboring rs12476437 revealed multiple statistically independent signals in the vicinity of CRIM1. Collectively, these findings computationally prioritize the endothelial-CRIM1 axis as a previously unappreciated candidate mechanistic bridge between LC and HF requiring experimental validation.

Humans

A decentralized future for the open-science databases.

The continuous and reliable open access to curated biological data repositories is indispensable for accelerating rigorous scientific inquiry and fostering reproducible research outcomes. However, the current paradigm, which relies heavily on centralized infrastructure for the storage and distribution of foundational biomedical datasets, inherently introduces significant vulnerabilities. This centralized model is susceptible to single points of failure, including cyberattacks, technical malfunctions, natural disasters, and even political or funding uncertainties. Such disruptions can lead to widespread data unavailability, data loss, integrity compromises, and substantial delays in critical research, ultimately impeding scientific progress. The downstream effect of such interruptions can be the widespread paralysis of diverse research activities, including computational, clinical, molecular, and climate studies. This scenario vividly illustrates the inherent dangers of consolidating essential scientific resources within a single geopolitical or institutional locus. As data generation is accelerating and the global landscape continues to fluctuate, the sustainability of centralized models must be critically re-evaluated. A shift toward federated and decentralized architectures may offer a robust and forward-looking approach to enhancing the resilience of scientific data infrastructures by reducing exposure to governance instability, infrastructural fragility, and funding volatility, while also promoting equity and global accessibility. Inspired by established models such as ELIXIR's federated infrastructure and the policy and funding frameworks developed by CODATA and the Global Biodata Coalition (GBC), emerging Decentralized Science (DeSci) initiatives can contribute to building more resilient, fair, and incentive-aligned data ecosystems. The future of open science depends on integrating these complementary approaches to establish a globally distributed, economically sustainable, and institutionally robust infrastructure that safeguards scientific data as a public good, further ensuring continued accessibility, interoperability, and preservation for generations to come. Here, we examine the structural limitations of centralized repositories, evaluate federated and decentralized models, and propose a hybrid framework for resilient, fair, and sustainable scientific data stewardship.

data accessibility

Linkage analysis in the presence of errors III: marker loci and their map as nuisance parameters.

In linkage and linkage disequilibrium (LD) analysis of complex multifactorial phenotypes, various types of errors can greatly reduce the chance of successful gene localization. The power of such studies-even in the absence of errors-is quite low, and, accordingly, their robustness to errors can be poor, especially in multipoint analysis. For this reason, it is important to deal with the ramifications of errors up front, as part of the analytical strategy. In this study, errors in the characterization of marker-locus parameters-including allele frequencies, haplotype frequencies (i.e., LD between marker loci), recombination fractions, and locus order-are dealt with through the use of profile likelihoods maximized over such nuisance parameters. It is shown that the common practice of assuming fixed, erroneous values for such parameters can reduce the power and/or increase the probability of obtaining false positive results in a study. The effects of errors in assumed parameter values are generally more severe when a larger number of less informative marker loci, like the highly-touted single nucleotide polymorphisms (SNPs), are analyzed jointly than when fewer but more informative marker loci, such as microsatellites, are used. Rather than fixing inaccurate values for these parameters a priori, we propose to treat them as nuisance parameters through the use of profile likelihoods. It is demonstrated that the power of linkage and/or LD analysis can be increased through application of this technique in situations where parameter values cannot be specified with a high degree of certainty.

Alleles

Achromobacter species in cystic fibrosis and chronic lung disease: a review of virulence, antibiotic resistance, diagnostic challenges, and emerging therapies.

Achromobacter species (spp) is an emerging opportunistic organism more frequently isolated from immunocompromised patients' and hospital settings. This bacterium was once considered an environmental bacterium, but now it is recognized as a serious cause of respiratory infections, bloodstream infections, and urinary tract infections, particularly among patients with cystic fibrosis (CF), chronic illnesses, and medical devices. The purpose of this review is to highlight Achromobacte's clinical significance, pathogenic mechanism, and recent approaches for diagnosis and treatment. By utilizing specific keywords relevant to Achromobacter spp., a comprehensive literature search was performed in PubMed and Google Scholar. To summarize existing knowledge and highlight gaps in the literature, peer-reviewed studies on clinical relevance, pathogenicity, antimicrobial resistance, and therapeutic approaches were gathered, screened, and narratively assembled. Among the 19 identified species, Achromobacter xylosoxidans (A. xylosoxidans) is the most prevalent and clinically relevant, especially in CF settings. This review explores the organism's microbiological characteristics, virulence strategies-including robust biofilm formation, motility, and secretion systems-and its alarming intrinsic and acquired resistance to antibiotics. Misidentification due to phenotypic overlap with other non-fermenting Gram-negative bacilli complicates diagnosis, while limited MALDI-TOF MS and database representation hinders species-level identification. Genotyping methods, including multi-locus sequence analysis and housekeeping gene sequencing, offer superior resolution but remain underutilized in clinical diagnostics. With rising resistance mediated by β-lactamases, efflux pumps, and adaptive genomic traits, Achromobacter spp presents a growing challenge for treatment and infection control. This review highlights the urgent need for improved diagnostic strategies, species-level clinical and microbiological data, and tailored therapeutic approaches to manage Achromobacter spp. infections effectively.

Humans