Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Combining prediction of secondary structure and solvent accessibility in proteins.

Owing to the use of evolutionary information and advanced machine learning protocols, secondary structures of amino acid residues in proteins can be predicted from the primary sequence with more than 75% per-residue accuracy for the 3-state (i.e., helix, beta-strand, and coil) classification problem. In this work we investigate whether further progress may be achieved by incorporating the relative solvent accessibility (RSA) of an amino acid residue as a fingerprint of the overall topology of the protein. Toward that goal, we developed a novel method for secondary structure prediction that uses predicted RSA in addition to attributes derived from evolutionary profiles. Our general approach follows the 2-stage protocol of Rost and Sander, with a number of Elman-type recurrent neural networks (NNs) combined into a consensus predictor. The RSA is predicted using our recently developed regression-based method that provides real-valued RSA, with the overall correlation coefficients between the actual and predicted RSA of about 0.66 in rigorous tests on independent control sets. Using the predicted RSA, we were able to improve the performance of our secondary structure prediction by up to 1.4% and achieved the overall per-residue accuracy between 77.0% and 78.4% for the 3-state classification problem on different control sets comprising, together, 603 proteins without homology to proteins included in the training. The effects of including solvent accessibility depend on the quality of RSA prediction. In the limit of perfect prediction (i.e., when using the actual RSA values derived from known protein structures), the accuracy of secondary structure prediction increases by up to 4%. We also observed that projecting real-valued RSA into 2 discrete classes with the commonly used threshold of 25% RSA decreases the classification accuracy for secondary structure prediction. While the level of improvement of secondary structure prediction may be different for prediction protocols that implicitly account for RSA in other ways, we conclude that an increase in the 3-state classification accuracy may be achieved when combining RSA with a state-of-the-art protocol utilizing evolutionary profiles. The new method is available through a Web server at http://sable.cchmc.org.

Amino Acid Sequence↗

On the nature of cavities on protein surfaces: application to the identification of drug-binding sites.

In this article we introduce a new method for the identification and the accurate characterization of protein surface cavities. The method is encoded in the program SCREEN (Surface Cavity REcognition and EvaluatioN). As a first test of the utility of our approach we used SCREEN to locate and analyze the surface cavities of a nonredundant set of 99 proteins cocrystallized with drugs. We find that this set of proteins has on average about 14 distinct cavities per protein. In all cases, a drug is bound at one (and sometimes more than one) of these cavities. Using cavity size alone as a criterion for predicting drug-binding sites yields a high balanced error rate of 15.7%, with only 71.7% coverage. Here we characterize each surface cavity by computing a comprehensive set of 408 physicochemical, structural, and geometric attributes. By applying modern machine learning techniques (Random Forests) we were able to develop a classifier that can identify drug-binding cavities with a balanced error rate of 7.2% and coverage of 88.9%. Only 18 of the 408 cavity attributes had a statistically significant role in the prediction. Of these 18 important attributes, almost all involved size and shape rather than physicochemical properties of the surface cavity. The implications of these results are discussed. A SCREEN Web server is available at http://interface.bioc.columbia.edu/screen.

Binding Sites↗

Predicting transmembrane beta-barrels and interstrand residue interactions from sequence.

Transmembrane beta-barrel (TMB) proteins are embedded in the outer membrane of Gram-negative bacteria, mitochondria, and chloroplasts. The cellular location and functional diversity of beta-barrel outer membrane proteins (omps) makes them an important protein class. At the present time, very few nonhomologous TMB structures have been determined by X-ray diffraction because of the experimental difficulty encountered in crystallizing transmembrane proteins. A novel method using pairwise interstrand residue statistical potentials derived from globular (nonouter membrane) proteins is introduced to predict the supersecondary structure of transmembrane beta-barrel proteins. The algorithm transFold employs a generalized hidden Markov model (i.e., multitape S-attribute grammar) to describe potential beta-barrel supersecondary structures and then computes by dynamic programming the minimum free energy beta-barrel structure. Hence, the approach can be viewed as a "wrapping" component that may capture folding processes with an initiation stage followed by progressive interaction of the sequence with the already-formed motifs. This approach differs significantly from others, which use traditional machine learning to solve this problem, because it does not require a training phase on known TMB structures and is the first to explicitly capture and predict long-range interactions. TransFold outperforms previous programs for predicting TMBs on smaller (<or=200 residues) proteins and matches their performance for straightforward recognition of longer proteins. An exception is for multimeric porins where the algorithm does perform well when an important functional motif in loops is initially identified. We verify our simulations of the folding process by comparing them with experimental data on the functional folding of TMBs. A Web server running transFold is available and outputs contact predictions and locations for sequences predicted to form TMBs.

Amino Acid Sequence↗

FiberID--a technique to identify fibrous protein subclasses.

Fibrous proteins such as collagen, silk, and elastin play critical biological roles, yet they have been the subject of few projects that use computational techniques to predict either their class or their structure. In this article, we present FiberID, a simple yet effective method for identifying and distinguishing three fibrous protein subclasses from their primary sequences. Using a combination of amino acid composition and fast Fourier measurements, FiberID can classify fibrous proteins belonging to these subclasses with high accuracy by using two standard machine learning techniques (decision trees and Naïve Bayesian classifiers). After presenting our results, we present several fibrous sequences that are regularly misclassified by FiberID as sequences of potential interest for further study. Finally, we analyze the decision trees developed by FiberID for potential insights regarding the structure of these proteins.

Algorithms↗

A boosting approach to flexible semiparametric mixed models.

In linear mixed models the influence of covariates is restricted to a strictly parametric form. With the rise of semi- and non-parametric regression also the mixed model has been expanded to allow for additive predictors. The common approach uses the representation of additive models as mixed models. An alternative approach that is proposed in the present paper is likelihood based boosting. Boosting originates in the machine learning community where it has been proposed as a technique to improve classification procedures by combining estimates with reweighted observations. Likelihood based boosting is a general method which may be seen as an extension of L2 boost. In additive mixed models the advantage of boosting techniques in the form of componentwise boosting is that it is suitable for high dimensional settings where many explanatory variables are present. It allows to fit additive models for many covariates with implicit selection of relevant variables and automatic selection of smoothing parameters. Moreover, boosting techniques may be used to incorporate the subject-specific variation of smooth influence functions by specifying 'random slopes' on smooth effects. This results in flexible semiparametric mixed models which are appropriate in cases where a simple random intercept is unable to capture the variation of effects across subjects.

Cohort Studies↗

Automated discovery of structural signatures of protein fold and function.

There are constraints on a protein sequence/structure for it to adopt a particular fold. These constraints could be either a local signature involving particular sequences or arrangements of secondary structure or a global signature involving features along the entire chain. To search systematically for protein fold signatures, we have explored the use of Inductive Logic Programming (ILP). ILP is a machine learning technique which derives rules from observation and encoded principles. The derived rules are readily interpreted in terms of concepts used by experts. For 20 populated folds in SCOP, 59 rules were found automatically. The accuracy of these rules, which is defined as the number of true positive plus true negative over the total number of examples, is 74% (cross-validated value). Further analysis was carried out for 23 signatures covering 30% or more positive examples of a particular fold. The work showed that signatures of protein folds exist, about half of rules discovered automatically coincide with the level of fold in the SCOP classification. Other signatures correspond to homologous family and may be the consequence of a functional requirement. Examination of the rules shows that many correspond to established principles published in specific literature. However, in general, the list of signatures is not part of standard biological databases of protein patterns. We find that the length of the loops makes an important contribution to the signatures, suggesting that this is an important determinant of the identity of protein folds. With the expansion in the number of determined protein structures, stimulated by structural genomics initiatives, there will be an increased need for automated methods to extract principles of protein folding from coordinates.

Algorithms↗

A novel method of protein secondary structure prediction with high segment overlap measure: support vector machine approach.

We have introduced a new method of protein secondary structure prediction which is based on the theory of support vector machine (SVM). SVM represents a new approach to supervised pattern classification which has been successfully applied to a wide range of pattern recognition problems, including object recognition, speaker identification, gene function prediction with microarray expression profile, etc. In these cases, the performance of SVM either matches or is significantly better than that of traditional machine learning approaches, including neural networks.The first use of the SVM approach to predict protein secondary structure is described here. Unlike the previous studies, we first constructed several binary classifiers, then assembled a tertiary classifier for three secondary structure states (helix, sheet and coil) based on these binary classifiers. The SVM method achieved a good performance of segment overlap accuracy SOV=76.2 % through sevenfold cross validation on a database of 513 non-homologous protein chains with multiple sequence alignments, which out-performs existing methods. Meanwhile three-state overall per-residue accuracy Q(3) achieved 73.5 %, which is at least comparable to existing single prediction methods. Furthermore a useful "reliability index" for the predictions was developed. In addition, SVM has many attractive features, including effective avoidance of overfitting, the ability to handle large feature spaces, information condensing of the given data set, etc. The SVM method is conveniently applied to many other pattern classification tasks in biology.

Computer Simulation↗

Nanopore Sequencing for Chikungunya Virus: Principles and Application.

Nanopore sequencing is transforming viral genomics through real-time, portable, long-read analysis of RNA and DNA. Unlike traditional short-read platforms, it detects nucleotide sequences by measuring ionic current changes as nucleic acids pass through nanoscale pores, enabling direct single-molecule sequencing and base modification detection. Its simplicity, flexibility, and capacity for ultra-long reads make it ideal for resolving complex genomic regions, structural variants, and full viral genomes. These advantages have accelerated its use in pathogen surveillance and outbreak response, especially in resource-limited settings. For chikungunya virus (CHIKV), nanopore sequencing allows rapid, culture-independent recovery of complete genomes from clinical and vector samples, enabling real-time tracking of viral diversity, evolution, and spread. Experiences from Ebola, Zika, and COVID-19 have demonstrated the power of portable sequencing, now applied to CHIKV monitoring. Advances in tools such as Guppy, Dorado, Minimap2, and Medaka enhance read quality, consensus accuracy, and downstream analyses. Despite challenges in basecalling and error correction, robust quality control pipelines ensure reliable results. Ongoing improvements in chemistry, flow cell design, and machine learning will further enhance fidelity and throughput, establishing nanopore sequencing as a cornerstone of CHIKV genomic surveillance and epidemic preparedness.

Chikungunya virus↗

Finite state control of functional electrical stimulation for the rehabilitation of gait.

Finite state control is an established technique for the implementation of intention detection and activity co-ordination levels of hierarchical control in neural prostheses, and has been used for these purposes over the last thirty years. The first finite state controllers (FSC) in the functional electrical stimulation of gait were manually crafted systems, based on observations of the events occurring during the gait cycle. Subsequent systems used machine learning to automatically learn finite state control behaviour directly from human experts. Recently, fuzzy control has been utilised as an extension of finite state control, resulting in improved state detection over standard finite state control systems in some instances. Clinical experience over the last thirty years has been positive, and has shown finite state control to be an effective and intuitive method for the control of functional electrical stimulation (FES) in neural prostheses. However, while finite state controlled neural prostheses are of interest in the research community, they are not widely used outside of this setting. This is largely due to the cumbersome nature of many neural prostheses which utilise externally mounted gait sensors and FES electrodes. FES-based control of movement has been subject to the constraints of artificial sensor and FES actuator technologies. However, continued advances in natural sensors and implanted multi-channel stimulators are broadening the boundaries of artificial control of movement, driving an evolutionary process towards increasingly human-like control of FES-based gait rehabilitation systems.

Electric Stimulation Therapy↗

Unsupervised clustering of evoked potentials by waveform.

A procedure for clustering evoked potentials (EPs) according to their waveforms is presented. Clustering is performed without a priori selection of basis waveforms, the number of basis waveforms or the number of clusters. The method uses the principal-component-analysis coefficients of EP records as features for unsupervised optimal fuzzy clustering (UOFC) of the records. The validity of the procedure is demonstrated in two instances: visual evoked potentials (VEPs) and cognitive event-related potentials (ERPs) from humans in a memory-scanning task. In the clustering of VEPs, the procedure differentiates between waveforms judged to be clinically normal and abnormal. In the clustering of ERPs, the procedure correctly differentiates between waveforms evoked by the same stimuli which differ in their context to the performance of a memory-scanning task (memorised items against probes). Within this classification, the procedure detects two subgroups to probe-evoked waveforms, which are not obvious from visual inspection of the waveforms. The advantage of the procedure, which conducts clustering by UOFC, is the adaptive and machine-learning nature of its operation.

Evoked Potentials↗

[Necessity and usefulness of bioinformatic methods for microarray data analysis].

Data emerging from DNA microarray experiments are usually difficult to interpret. While the level of expression of several thousand genes can be measured in a single experiment, only a few dozen experiments are normally carried out, leading to data sets of very high dimensionality and low cardinality. The computational analysis of gene expression data makes significant usage of machine learning and statistical methods. Nevertheless, caution should be used in the blind adoption of these methods, as this usually leads to an over-interpretation of the expression profiles. The following presentation provides an overview of up-to-date principles of biostatistical analysis. A potential application for the analysis of high-dimensional expression profiles of prostate cancer is given.

Chromosome Aberrations↗

Osteoarthritis phenotypes: advancing precision medicine through clinical, structural, and molecular stratification.

PURPOSE: Osteoarthritis (OA) is now understood as a heterogeneous syndrome driven by diverse biological, biomechanical, metabolic, genetic, and molecular mechanisms. This variability explains differences in disease progression and treatment response, challenging the traditional "one-size-fits-all" approach. This review highlights OA phenotyping as a key step toward precision medicine, focusing on clinical, structural, and molecular classifications that inform individualized care. METHODS: A narrative review was conducted using a non-systematic search of major databases and Osteoarthritis Research Society International sources (2010-2026). Evidence was thematically synthesized across clinical, imaging, and molecular domains to characterize OA phenotypes and their potential relevance to precision medicine. RESULTS: Multiple OA phenotypes were identified: inflammatory, metabolic, biomechanical, cartilage-subchondral, pain-sensitization, and aging/senescence. These exhibit distinct clinical features, risk factors, and therapeutic responses. Imaging-based phenotypes (e.g., inflammatory, meniscus-cartilage, subchondral bone, atrophic, hypertrophic) and molecular endotypes (low turnover, structural damage, systemic inflammation) further refine stratification. Pain-structure discordance is notable in sensitization phenotypes and may predict poorer surgical outcomes. Joint-specific variations and emerging genomic and epigenetic insights underscore disease complexity. Advances in imaging, biomarkers, and machine learning may enable earlier detection and patient clustering, though clinical application remains limited. CONCLUSION: Phenotype- and endotype-based classification represents a critical advancement toward precision OA management. Tailored interventions based on stratification hold promise for improving outcomes; however, clinical translation remains limited by overlapping phenotypes, lack of validated biomarkers, and inconsistent results from phenotype-driven trials. Wider clinical adoption requires standardized definitions, validation across joints, and integration of multimodal diagnostic tools into routine practice.

Humans↗

Esketamine multi-omic biomarker evaluation in major depressive disorder (EMBER-MDD): concept, objectives and methodologies of a non-clinical investigator-initiated study.

Treatment resistance (TR) in major depressive disorder (MDD) affects a substantial minority of patients and is hard to recognize early, delaying intensified care. The Esketamine multi-omic biomarker evaluation in MDD (EMBER-MDD) is a non-interventional, investigator-initiated, in-vitro study within the EU Psych-STRATA programme, analyzing biospecimens collected in the randomized INTENSIFY study and the mirror OBS-TR cohort after participants complete treatment. EMBER-MDD aims to discover individual-omic and integrated multi-omic (hypothesis-free) biomarkers and signatures associated with TR risk, and molecular correlates of clinical response to esketamine nasal spray versus treatment as usual (TAU). Biomaterials will derive from approximately 420 adults with MDD (estimated n&#x2009;=&#x2009;210 esketamine; n&#x2009;=&#x2009;210 TAU) and include whole blood, RNA-stabilized whole blood, plasma and serum, sampled at baseline and, when feasible, during and after treatment (up to ~&#x2009;5,040 aliquots stored at -&#x2009;80&#xa0;&#xb0;C). Genomics will use baseline DNA genotyping on Illumina Infinium GSA v3.0+MD arrays; epigenomics will profile genome-wide DNA methylation across time points using MethylationEPIC v2.0; transcriptomics will employ mRNA-seq (NovaSeq X/ X Plus); and proteomics/ metabolomics will be generated using high-throughput Olink and/ or Biocrates platforms. Each layer will undergo state-of-the-art preprocessing and analyses (e.g., GWAS/ PRS, EWAS, differential expression, WGCNA, pathway and network analyses), followed by integrative strategies including QTL mapping (meQTL/ eQTL/ pQTL/ mQTL) and intermediate-fusion machine learning with nested cross-validation, explainable AI (SHAP/ LIME) and treatment-effect modelling. All outputs are research-only and will not support individual efficacy, tolerability, or clinical decision-making. The study will deliver robust biosignatures and mechanistic hypotheses to guide future validation and inform stratified, molecularly guided intervention strategies in subsequent prospective trials. Trial registration number: 2023-506617-21-00 and 2025-178-f-S.

Humans↗

Proteomic profiling of bone for the estimation of post-mortem interval and post-mortem submersion interval: a systematic review.

Accurate estimation of the Post-Mortem Interval (PMI) and Post-Mortem Submersion Interval (PMSI) remains a persistent challenge in forensic science, especially when traditional morphological and entomological methods fail due to advanced decomposition or in aquatic environments. Proteomic profiling of bone tissues has recently emerged as a promising approach, leveraging the predictable degradation patterns of bone proteins to estimate time since death more reliably. This systematic review, conducted in accordance with PRISMA guidelines, analyzed 24 peer-reviewed studies focusing on the application of proteomic techniques to bone tissue for PMI and PMSI estimation. The included studies were evaluated based on sample type, analytical techniques used, identified biomarkers, environmental conditions assessed, and the overall reliability and reproducibility of the findings. The review found that specific bone proteins, particularly collagen, osteocalcin, fetuin-A, etc. exhibited consistent degradation patterns that correlated strongly with elapsed post-mortem time. Cortical bone was identified as a more stable and informative matrix compared to trabecular bone. Mass spectrometry, especially LC-MS/MS, emerged as the predominant analytical technique due to its high sensitivity and accuracy in detecting low-abundance proteins over extended PMIs and PMSIs. However, protein degradation rates were significantly influenced by environmental variables such as temperature, humidity, soil pH, and microbial activity. This review also emphasizes the transformative role of bone proteomics in advancing forensic science while identifying key gaps that must be addressed to achieve global standardization and practical implementation in diverse forensic contexts. The integration of proteomics with other emerging technologies, such as machine learning algorithms and computational modeling, may further enhance the precision of PMI and PMSI estimation in future applications.

Postmortem Changes↗

The in vitro influence of eight hormones and growth factors on the proliferation of eight sarcoma cell lines.

Little is known about the regulation of sarcoma proliferation by hormones and/or growth factors. We therefore characterised the in vitro proliferative influence on eight sarcoma cell lines of the platelet-derived growth factor, the insulin-like growth factor 1, triiodothyronine, the epidermal growth factor, the luteinising-hormone-releasing hormone, progesterone, gastrin and 17 beta-oestradiol. The influence of the different factors on the proliferation of sarcoma cell lines was measured by the colorimetric 3-(4,5-dimethylthiazol-2-yl)-2,5-diphenyltetrazolium bromide test. Two culture media were studied: (1) a nutritionally poor medium containing 2% of fetal calf serum and (2) a nutritionally rich one containing 5% or 10% FCS both with and without the addition of non-essential amino acids. The results were analysed either by conventional statistical analyses or by a classification method based on a decision-tree approach developed in Machine Learning. This latter method was also compared to other classifiers (such as logistic regression and k nearest neighbours) with respect to its accuracy of classification. Monovariate statistical analysis showed that each of the eight cell lines exhibited sensitivity to at least one factor, and each factor significantly modified the proliferation of five or six of the eight cell lines under study. Of these eight lines one of fibrosarcoma origin was the most "factor-sensitive". Decision-tree-related data analysis enabled the specific pattern of factor sensitivity to be characterised for the three histological types of cell line under study. The effects of hormone and growth factors are significantly influenced by the type of culture medium used. The method used appeared to be an accurate classifier for the kind of data analysed. Sarcoma proliferation can be modulated, at least in vitro, by various hormones and growth factors, and the proliferation of each histopathological type exhibited a distinct sensitivity to different hormone and/or growth-factors.

Cell Division↗

Decoding the distribution, structure-function-redox potential relationship and recent advances in fungal laccases: a systematic approach.

Laccases, categorized as multicopper oxidases, are recognized for their multifaceted roles in ecosystems and their utility in diverse industrial applications. Laccases from higher fungi, specifically Ascomycota and Basidiomycota, have garnered significant research interest due to their elevated redox potentials and their capacity to degrade lignin in decaying wood, alongside other industrial uses. Here, we have conducted a comprehensive and systematic analysis on fungal laccases using Web of Science, Scopus, PubMed, and ScienceDirect. The genomic distribution, phylogenetic affiliation, and structural organization of laccase-encoding genes in higher fungal species were investigated, as were the catalytic mechanisms of the corresponding enzymes. Additionally, the study explores the correlation between structural domains and redox potential, as well as the impact of post-translational modifications like glycosylation on enzyme activity. Furthermore, the recent advancements in laccase engineering, employing strategies such as rational design, directed evolution, and heterologous expression are discussed. The review also explores the scope of "artificial intelligence and machine learning" in deducing the structure-function relationships, optimizing codon usage, predicting signal peptides, enhancing enzymatic performance, and developing host-specific genetic engineering techniques is also discussed for tailoring fungal laccases to meet the demands of industrial biocatalysis for improved activity and stability.

Laccase↗

Investigation of Fatty Acid Metabolism-Associated Molecular CPOX and the Underlying Mechanism in Follicular Lymphoma.

Dysregulated lipid metabolism is a key driver of follicular lymphoma (FL). This study aimed to explore the lipid metabolism-related genes (LMRGs) and clarify the underlying roles and mechanisms in FL. Bioinformatics methods, including differential analysis, WGCNA, machine learning, and Mendelian randomization, were utilized to select the LMRGs in FL. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) analyses were conducted to investigate the function of the key LMRG. Receiver operator characteristic (ROC) was used to evaluate the diagnostic value of the key gene CPOX. A pan-cancer analysis investigated CPOX's expression level and immune correlations. In vitro experiments using FL cell lines (WSU-FSCCL, DOHH2) validated CPOX expression, and CPOX knockdown in DOHH2 cells was used to assess its impact on viability, migration, invasion, and fatty acid metabolism. CPOX was confirmed to be a risk factor, significantly overexpressed in FL, and exhibited effective diagnostic ability in FL (AUC&#x2009;=&#x2009;0.731). Functional analysis linked CPOX to mitochondrial function, oxidative phosphorylation, and heme metabolic process. Pan-cancer indicated the dysregulated CPOX across multiple cancers and closely correlation with immune characteristics. Experimentally, CPOX was higher in the more invasive DOHH2 cells; and CPOX knockdown suppressed FL progression and reduced lipid droplet formation, triglyceride, total cholesterol, and free fatty acid levels. In conclusion, this study fills the gap in understanding the significance of lipid metabolism-related molecules in FL, and innovatively proposes that CPOX is a risk factor for FL. Knockdown of CPOX inhibits the FL progression, which is regulated by fatty acid metabolism.

Lymphoma, Follicular↗

Ethical Governance of Open Data Across Biomedical Research, Healthcare, and Public Health: Privacy, Equity, Trust, and Controlled Access.

Open data has become central to biomedical research and public health, but health information is uniquely sensitive and difficult to share responsibly. In this narrative review, open data is considered as a spectrum of health-data sharing arrangements, ranging from public aggregate datasets to controlled-access repositories, federated analysis, and synthetic data. This narrative review synthesizes the scientific and societal rationale for greater openness with the ethical, legal, and governance constraints that shape what "open" can realistically mean in healthcare. We examine how data sharing supports reproducibility, machine learning, and more efficient research, while also enabling public health surveillance and learning health systems. Against these benefits, we analyze privacy and re-identification risks, consent challenges in large-scale secondary use, inequities including data colonialism, and tensions introduced by commercialization. We integrate lessons from prominent case examples spanning pandemic data sharing, genomic initiatives, population registries, patient-led rare disease infrastructures, and regional data spaces. Across these domains, experience suggests that durable progress depends less on unrestricted openness than on calibrated access, privacy-preserving architectures, clear accountability, and sustained public engagement. We conclude by proposing a pragmatic ethical orientation for healthcare open data: treat openness as a spectrum of controlled sharing arrangements, embed equity and reciprocity into governance, and institutionalize trust-building measures that can persist beyond emergencies and political cycles.

Data colonialism↗