Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transfer machine learning”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Integrating explainable AI with multiomics systems biology and EHR data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health record (EHR) data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; nine tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations (SHAP) identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct "subtissues" (clusters of samples); and gene-gene co-expression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six FDA-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large U.S. de-identified insurance-claims database (n = 364733), exposure to promethazine, one of the candidate drugs, was associated with a 57-62 % lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both p < 0.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multi-omics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Computational Biology↗

Senescent fibroblasts drive CD8+ T cell dysfunction in colorectal cancer via CD36-mediated lipid transfer and peroxidation.

BACKGROUND: Functional exhaustion of tumor-infiltrating CD8+ T cells represents a hallmark of colorectal cancer (CRC) immunosuppression, though its mechanistic drivers remain elusive. Given the established correlation between CRC progression and stromal senescence characterized by pathological lipid accumulation and impaired immunity, we investigated whether and how senescent fibroblasts actively regulate CD8+ T cell dysfunction. METHODS: Single-cell RNA sequencing (scRNA-seq) analysis was conducted to unveil the diverse fibroblast populations and the significant lipid metabolism changes between senescent fibroblasts and non-senescent fibroblasts in human CRC specimens and adjacent normal mucosa. Machine-learning identified senescent fibroblasts with a distinct gene signature. Cell-cell communication analysis was used to evaluate the interactions between senescent fibroblasts and CD8+ T cells in colorectal cancer. Co-culture experiments were conducted among senescent fibroblasts, CD8+ T cells and patient-derived organoids of CRC (CRC-PDOs), with the results evaluated with high-content imaging and propidium iodide/Hoechst 33,342 staining. Flow cytometry, ELISA and lipid pulse-chase with BODIPY FL C16 were performed to detect the alterations of CD8+ T cell cytotoxic function and metabolic status. AOM/DSS-induced CRC mouse model was used to conduct in vivo validation to evaluate whether senolytics could suppress CRC progression. Patients from the Cancer Genome Atlas colorectal cancer cohort were stratified into CD36-high and CD36-low groups by median expression, and drug sensitivity for GDSC2 compounds was predicted computationally using the oncoPredict R package. RESULTS: ScRNA-seq demonstrated the specific cell population presence and divergence of senescent fibroblasts between neoplastic and histologically normal adjacent cell clusters in CRC. Random Forest was employed for cell senescence classification. Feature importance analysis identified five genes as key contributors to the model&#x2019;s decision process. Cell-cell communication analysis revealed enhanced interactions between senescent fibroblasts and CD8+ T cells in CRC. Co-culture of senescent fibroblasts significantly impaired the cytotoxic functions of CD8+ T cells on CRC-PDOs, which was reflected by the declined proportions of granzyme B (GZMB) + and interferon gamma (IFN&#x3b3;) + CD8+ T cells and enhanced viability of CRC-PDOs. Mechanistically, the co-culture with senescent fibroblasts promoted the lipid shuttling into CD8+ T cells to induce lipid peroxidation and downstream impairment of cytotoxicity. Furthermore, the inhibition of CD36, the specific scavenger receptor for lipid uptake of CD8+ T cells, effectively suppressed lipid transfer and peroxidation thereby preserving the effector functions of CD8+ T cells and ultimately promoting tumor apoptosis. Complementarily, in vivo senolytic treatment significantly suppressed CRC progression in AOM-DSS CRC mouse models. Top 12 therapeutic agents were identified significantly enhanced predicted efficacy in CD36-high tumors. CONCLUSIONS: Our study identified a substantial population of senescent fibroblasts in human CRC through single cell transcriptomics, machine-learning and clinical biopsies. These senescent fibroblasts impair CD8+ T cell-mediated killing of CRC-PDOs via CD36-dependent lipid transfer, suggesting senolytic targeting of stromal cells as a promising immunotherapeutic strategy for CRC.

Colorectal Neoplasms↗

An approach to biological computation: unicellular core-memory creatures evolved using genetic algorithms.

A novel machine language genetic programming system that uses one-dimensional core memories is proposed and simulated. The core is compared to a biochemical reaction space, and in imitation of biological molecules, four types of data words (Membrane, Pure data, Operator, and Instruction) are prepared in the core. A program is represented by a sequence of Instructions. During execution of the core, Instructions are transcribed into corresponding Operators, and Operators modify, create, or transfer Pure data. The core is hierarchically partitioned into sections by the Membrane data, and the data transfer between sections by special channel Operators constitutes a tree data-flow structure among sections in the core. In the experiment, genetic algorithms are used to modify program information. A simple machine learning problem is prepared for the environment data set of the creatures (programs), and the fitness value of a creature is calculated from the Pure data excreted by the creature. Breeding of programs that can output the predefined answer is successfully carried out. Several future plans to extend this system are also discussed.

Algorithms↗

Use of the Doppler ultrasonic flowmeter for pedicle flaps.

Summary--The Doppler ultrasonic flowmeter is presented as a method to help in flap outlining, transfer and return. When a directional flowmeter is added, this machine is extremely valuable for returning a flap at the earliest possible time. The Doppler is a safe, inexpensive, atraumatic, reliable instrument that can be learned in a very short time. We have found the Doppler to be very helpful in head and neck flaps.

Aged↗

Metagenomic polymorphic toxin effector and immunity profiling predicts microbiome development and disease-related dysbiosis.

Bacteria use antagonistic interbacterial weapons, such as polymorphic toxin secretion systems (TSS), to compete for niches in the human gut microbiome. We hypothesized that TSS influence gut microbiome development and disease-related dysbiosis. We developed a bioinformatic marker gene approach (PolyProf) to quantify TSS including ~200 effector and immunity genes and applied it to ~15,000 publicly available human metagenomes. PolyProf alpha and beta diversity readily distinguished 12 different human disease states and enabled the construction of highly accurate linear regression classifier machine learning models. Elastic net machine learning models integrating bacterial taxonomy with PolyProf had strong predictive value for 12 disease states, outperforming models utilizing taxonomy alone. During microbiome development in the first year of life, PolyProf alpha diversity increases, and beta diversity becomes increasingly like the maternal microbiome, influenced by vertical transfer, delivery mode, and breastfeeding. PolyProf is related to strain sharing among adults through social interactions. In summary, TSS genes strongly correlate with microbiome development and interpersonal strain sharing, suggesting roles for interbacterial antagonism. Since PolyProf distinguishes diverse adult disease statuses, these dynamics may contribute to non-genetic inheritance.IMPORTANCEPrevious research has demonstrated that bacteria compete within the gut microbiome using toxin secretion systems (TSS). How TSS contribute to human microbiome development and the microbiome alterations observed in human diseases is not known. This study develops a new bioinformatic tool for profiling TSS-related genes in metagenomic data. Application of this approach to large-scale human fecal metagenomic data demonstrates the dynamic association of TSS during microbiome development, including the exchange of strains among social contacts. TSS gene abundance patterns are highly predictive of 12 disease states. This study advances the field by enabling TSS profiling in metagenomes and by identifying disease and microbiome development biomarkers that provide hypotheses for future mechanistic studies and may be useful for disease diagnosis.

Dysbiosis↗

Seeing and Feeling DNA Methylation: Single-Molecule Biophysics Meets Machine Learning.

DNA methylation at 5-methylcytosine (5mC) is crucial for embryonic development and cellular function, while aberrant patterns strongly drive disease onset and progression. Its reversible nature offers substantial therapeutic potential, emphasizing the need for precise, context-specific genome wide 5mC mapping. Conventional techniques such as bisulfite sequencing and ensemble biosensor assays are hindered by DNA degradation, amplification bias, high cost, and inability to resolve single-molecule structural and mechanical effects of methylation. This review examines advances in single-molecule biophysical methods (nanopore sensing, smFRET, optical/magnetic tweezers, and AFM) that provide direct, label-free/minimally invasive 5mC detection, along with quantitative insights into DNA conformation, mechanics, and protein-DNA interactions. These techniques complement traditional methylome mapping by linking genomic localization to molecular mechanisms. Emerging machine-learning approaches are revolutionizing analysis, particularly in nanopore sensing, while promising applications in smFRET, tweezers, and AFM address throughput and reproducibility challenges. Their convergence promises scalable, high-resolution epigenetic profiling, advancing precision epigenomics toward clinical application.

DNA Methylation↗

Integrating explainable artificial intelligence with multiomics systems biology and electronic health record data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health records data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; 9 tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct subtissues (defined as clusters of samples within a brain tissue that share a specific expression pattern); and gene-gene coexpression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six Food and Drug Administration (FDA)-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large US de-identified insurance-claims database (n&#x2009;=&#x2009;364&#xa0;733), exposure to promethazine, one of the candidate drugs, was associated with a 57%-62% lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both P&#x2009;<&#x2009;.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multiomics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Alzheimer Disease↗

OpenSpliceAI: An efficient, modular implementation of SpliceAI enabling easy retraining on non-human species.

The SpliceAI deep learning system is currently one of the most accurate methods for identifying splicing signals directly from DNA sequences. However, its utility is limited by its reliance on older software frameworks and human-centric training data. Here we introduce OpenSpliceAI, a trainable, open-source version of SpliceAI implemented in PyTorch to address these challenges. OpenSpliceAI supports both training from scratch and transfer learning, enabling seamless retraining on species-specific datasets and mitigating human-centric biases. Our experiments show that it achieves faster processing speeds and lower memory usage than the original SpliceAI code, allowing large-scale analyses of extensive genomic regions on a single GPU. Additionally, OpenSpliceAI's flexible architecture makes for easier integration with established machine learning ecosystems, simplifying the development of custom splicing models for different species and applications. We demonstrate that OpenSpliceAI's output is highly concordant with SpliceAI. In silico mutagenesis (ISM) analyses confirm that both models rely on similar sequence features, and calibration experiments demonstrate similar score probability estimates.

Journal Article↗

[Increase in strength after active therapy in chronic low back pain (CLBP) patients: muscular adaptations and clinical relevance].

INTRODUCTION: Active treatments are advocated for the management of non-specific chronic low back pain (CLBP), although few studies have documented the relative efficacy of differing types of programme. A number of the available treatments comprise exercise routines on specially designed training machines, which are ostensibly better disposed to reverse the compromised trunk muscle function displayed by these patients than are 'free exercise' programmes. However, in using these muscle-training programmes, the physiological or anatomical adaptations that might account for the improved performance are rarely investigated, let alone identified. This is an important issue, because if the 'newly-acquired strength' is mostly specific to performance on the devices on which the patient has trained and been tested, and reflects the skill in executing these particular tasks, this will not necessarily assist the patient during performance of his/her everyday activities. The aims of the present study were (1) to quantify the changes in back muscle performance in chronic LBP patients following 3 months active therapy, and (2) to analyse the corresponding changes in activation and cross-sectional area of the paraspinal muscles. METHODS: 148 individuals (57% women) with CLBP (age 45.0+/-10.0 years; duration of LBP 10.9+/-9.5 years) were randomised to a treatment which they attended 2/week for 3 months: active physiotherapy, muscle reconditioning on training devices, or low-impact aerobics. Pre- and post-therapy, assessments were made of isometric trunk muscle strength in each plane of movement and of erector spinae activation (using surface electromyography) during back extension. In a sub-group of 56 patients, the cross-sectional area of the paravertebral muscles was determined using magnetic resonance imaging (MRI). In all patients, self-rated pain intensity, pain frequency and disability were assessed before and after therapy. RESULTS: 132/148 patients completed the therapy. Isometric strength in each movement plane increased significantly in all groups post-therapy. Apart from trunk extension, the changes were significantly greater in the devices group than in the other two groups (Fig 1). Activation of the paraspinal muscles during back extension also increased significantly in all groups (Fig 2) and was weakly, but significantly (r = 0.37; p = 0.0001) correlated with increased strength in back extension. Although, at baseline, highly significant correlations were observed between the size of the paraspinal muscles (at L3/4 and at L4/5) and isometric back extension strength (r=0.75; p< 0.0001), post-training increases in strength were not accompanied by corresponding changes in muscle size. None of the improvements in strength showed any relationship with the clinical changes in pain and disability, regardless of whether the latter were examined on an individual basis or in relation to 'outcome groups'. CONCLUSION: The superior trunk strength shown by the devices group post-therapy was considered to be attributable, in part, to a 'learning effect', of the type often seen when training and testing are carried out on the same machines. These gains are considered to be mostly 'task-specific'. However, part of the improvement in strength after active therapy (in all groups) also appeared to be due to an increased neural activation of the trunk muscles. These positive effects should be transferable to the performance of everyday activities for which the same muscles are employed, although the percentage improvement is probably not as high as the measured increase in strength might suggest. Possible roles for improved co-ordination and changes in motivation and/or pain tolerance after therapy cannot be excluded. No differences in the clinical outcome were observed between the three therapy groups, and the changes in physical performance after therapy did not correlate with the clinical outcome. It is therefore questionable whether strength measurements have any clinical significance in documenting the success of rehabilitation programmes, other than on a motivational basis. The results of the present study suggest that the value of supervised active therapy programmes does not reside in the reversal of specific muscular deficiencies, but rather in the provision of a source of confirmation/encouragement for the patient, that movement is not harmful, and a foundation upon which to further build. Whether the utilisation of specific training devices, or individual instruction, is necessary to elicit these particular effects is questionable.

Adaptation, Physiological↗

A robust transfer learning approach for high-dimensional linear regression to support integration of multi-source gene expression data.

Transfer learning aims to integrate useful information from multi-source datasets to improve the learning performance of target data. This can be effectively applied in genomics when we learn the gene associations in a target tissue, and data from other tissues can be integrated. However, heavy-tail distribution and outliers are common in genomics data, which poses challenges to the effectiveness of current transfer learning approaches. In this paper, we study the transfer learning problem under high-dimensional linear models with t-distributed error (Trans-PtLR), which aims to improve the estimation and prediction of target data by borrowing information from useful source data and offering robustness to accommodate complex data with heavy tails and outliers. In the oracle case with known transferable source datasets, a transfer learning algorithm based on penalized maximum likelihood and expectation-maximization algorithm is established. To avoid including non-informative sources, we propose to select the transferable sources based on cross-validation. Extensive simulation experiments as well as an application demonstrate that Trans-PtLR demonstrates robustness and better performance of estimation and prediction when heavy-tail and outliers exist compared to transfer learning for linear regression model with normal error distribution. Data integration, Variable selection, T distribution, Expectation maximization algorithm, Genotype-Tissue Expression, Cross validation.

Linear Models↗

Integration of radiotherapy planning systems and radiotherapy treatment equipment: 11 years experience.

PURPOSE: We have investigated the requirements, design, implementation, and operation of a computer-controlled medical accelerator with multileaf collimator (MLC), integrated with a radiation treatment-planning system (RTPS), and we report on the performance, benefits, and lessons learned from this experience. METHODS AND MATERIALS: In 1984 the University of Washington installed a computer-controlled radiation therapy machine (the Clinical Neutron Therapy System, or CNTS) with a multileaf collimator. Since the beginning of operation the control system computer has been connected by commercially available network hardware and software to three generations of radiation treatment-planning systems. Semiautomated setup and completely computerized check and confirm were incorporated into the system from the beginning of clinical operation in 1984. The system cannot deliver a patient treatment without a computer-prepared treatment plan. RESULTS: The CNTS has been in use for routine patient treatments for over 11 years. The cost of the network connection and software was an insignificant fraction of the facility cost. Operation has been efficient and reliable. Of the 441 machine-related session reschedulings (out of 18,432 sessions total) during the past 9 years, only 20 were due to problems with data transfer between the RTPS and CNTS, associated primarily with two incidents. Close integration with the treatment-planning system allows complex treatments to be delivered. Dramatic evolution of the departmental treatment-planning system has not required any changes or redesign of either the accelerator control system or the network connection. CONCLUSIONS: Our experience shows that a large degree of automation is possible with reasonable effort, by using well-known software and hardware design strategies. The lessons we have learned from this can be carried over into photon therapy now that photon accelerators with MLC facilities are commercially available.

Computer Communication Networks↗

Can host genetics transform the sustainable control of tropical theileriosis? Insights from the Tick-Theileria interface.

Tropical theileriosis, caused by the tick-transmitted apicomplexan parasite Theileria annulata, remains a major constraint on cattle production across North Africa, the Mediterranean basin, the Middle East and South Asia. Current control depends on acaricides, the theilericidal drug buparvaquone and live attenuated schizont vaccines, but acaricide resistance, buparvaquone-resistance mutations and the logistical demands of vaccination are eroding the sustainability of these tools. Host genetics offers a complementary and durable alternative. Indigenous Bos indicus breeds are consistently more resistant to ticks and tolerate T. annulata infection better than exotic Bos taurus cattle, and this advantage has a measurable heritable component. Unlike previous reviews, which treat tick resistance, T. annulata immunobiology and livestock genomic selection as separate subjects, we integrate all three and assess host genetics specifically against the failure modes of current control. We review the tick, parasite and host interface, the evidence for natural resistance, and the genetic and immunological mechanisms involved, including signal-regulatory protein, bovine major histocompatibility complex class II and inflammatory pathway genes. We then assess whether genomic selection, multi-omics, machine learning and gene editing can translate these mechanisms into resistant cattle, and we weigh the biological, economic and infrastructural barriers to implementation. The evidence indicates that host genetics will not replace existing control but could reduce reliance on acaricides and chemotherapy. That contribution remains prospective rather than demonstrated: no resistance marker for T. annulata has yet been validated, prediction accuracies are moderate and transfer poorly between breeds, and no endemic production system has implemented selection for resistance.

Animals↗

Evolving mobile robots in simulated and real environments.

The problem of the validity of simulation is particularly relevant for methodologies that use machine learning techniques to develop control systems for autonomous robots, as, for instance, the artificial life approach known as evolutionary robotics. In fact, although it has been demonstrated that training or evolving robots in real environments is possible, the number of trials needed to test the system discourages the use of physical robots during the training period. By evolving neural controllers for a Khepera robot in computer simulations and then transferring the agents obtained to the real environment we show that (a) an accurate model of a particular robot-environment dynamics can be built by sampling the real world through the sensors and the actuators of the robot; (b) the performance gap between the obtained behaviors in simulated and real environments may be significantly reduced by introducing a "conservative" form of noise; (c) if a decrease in performance is observed when the system is transferred to a real environment, successful and robust results can be obtained by continuing the evolutionary process in the real environment for a few generations.

Algorithms↗

Water source, latrine type, and rainfall are associated with detection of non-optimal and enteric bacteria in the vaginal microbiome: a prospective observational cohort study nested within a cluster randomized controlled trial.

BACKGROUND: Less than one-third of sub-Saharan Africans have access to improved water sources. In US, Indian, and African studies, Bacterial vaginosis (BV) is increased among women with poor water, sanitation, and hygiene (WASH). We examined water source, sanitation (latrine type), and rainfall in relation to the vaginal microbiome (VMB). METHODS: In a cluster randomized controlled trial of menstrual cups and cash transfer, we measured the impact of cups on VMB via 16S rRNA gene amplicon sequencing in a subset of 436 adolescent girls. We analyzed how self-reported water source and latrine type at home related to VMB over 18-months, examining community state type I (CST-I, L. crispatus dominant) vs. other CST; alpha diversity; targeted taxa (coliform and other water-related pathogens); and non-targeted taxa via machine learning approaches. Mixed effects multivariable longitudinal models were adjusted for intervention arm, age, socioeconomic status, sexual activity, and cluster-level school WASH and rainfall (in millimeters). RESULTS: Adjusting for all covariates in all models: (1) the odds of CST-I were increased among participants with piped water (vs. pond), and decreased with traditional pit latrine vs. flush toilet. (2) Alpha diversity varied by water source and latrine type without consistent trends. (3) Coliform bacteria relative abundance (RA) was higher among participants with traditional pit or ventilated improved pit latrines vs. flush toilet, and higher among participants relying on stream vs. pond water. Streptococcus agalactiae RA was higher among participants with non-flush toilets, while Bacteroides fragilis RA was lower with non-flush toilets. (4) Key taxa from non-targeted analyses associated with water source and latrine type included typical vaginal bacteria, opportunistic pathogens, and urinary tract pathobionts. (6) Increased rainfall was associated with decreased odds of CST-I. TRIAL REGISTRATION: ClinicalTrials.gov NCT03051789, February 14, 2017.

Adolescent↗

The cerebellum: a neuronal learning machine?

Comparison of two seemingly quite different behaviors yields a surprisingly consistent picture of the role of the cerebellum in motor learning. Behavioral and physiological data about classical conditioning of the eyelid response and motor learning in the vestibulo-ocular reflex suggests that (i) plasticity is distributed between the cerebellar cortex and the deep cerebellar nuclei; (ii) the cerebellar cortex plays a special role in learning the timing of movement; and (iii) the cerebellar cortex guides learning in the deep nuclei, which may allow learning to be transferred from the cortex to the deep nuclei. Because many of the similarities in the data from the two systems typify general features of cerebellar organization, the cerebellar mechanisms of learning in these two systems may represent principles that apply to many motor systems.

Animals↗

Crosstalk between cysteine and lysine modifications: Integrating redox and metabolic regulation.

Protein post-translational modifications (PTMs) on amino acid residues enable dynamic cellular responses to changes in metabolic and redox state. Cysteine and lysine are among the most extensively modified amino acid residues, with both undergoing a diversity of acylation and oxidative modifications. Indeed, proximal (<10&#x202f;&#xc5;) cysteine and lysine residues may form integration nodes for crosstalk between metabolism and redox homeostasis pathways. This review highlights the interaction of proximal Cys-Lys residues, including influence on residue pKa by local electrostatics, cysteine-to-lysine transfer of PTM moieties, and covalent crosslinking. We discuss candidate Cys-Lys regulatory pairs in proteins involved in redox regulation, proteostasis, metabolic adaptation and inflammation. We further utilize computational modeling to identify proximity between cysteine and lysine residues in proteins known to be regulated by acylation and oxidative PTMs, and to demonstrate changes in these distances and local electrostatic potential due to lysine acetylation. Finally, we review how mass spectrometry-based proteomics and machine-learning PTM predictive tools can enable the identification, validation, and interpretation of proximal Cys-Lys interactions that regulate cellular responses to oxidative challenge and metabolic flux.

Cysteine↗

Deep learning-based annotation of plant abiotic stress resistance genes for crops.

The declining costs of DNA sequencing have expanded genomic data, crucial for understanding plant abiotic stress responses and crop improvement. However, accurate gene annotation remains challenging. To address this limitation, we propose the PASRGA, a deep learning approach that leverages transfer learning and contrastive learning to annotate genes related to drought, salt, cold, and UV resistance. PASRGA achieves high F1-scores, area under the receiver operating characteristic (AUROC), area under the precision-recall curve (AUPRC), and Matthews correlation coefficient (MCC) in annotating stress resistance genes, significantly outperforming the general protein annotation model CLEAN, the plant phosphatase gene annotation model PF-NET, the top-ranked model in the CAFA5 challenge NetGO 4.0, and four traditional machine learning methods. Its effectiveness was further validated with a salt stress treatment experiment in Eutrema salsugineum. To facilitate crop breeding practices, we utilized PASRGA to annotate the genomes of 17 major crops. To improve accessibility and utility, we incorporated both manually curated and PASRGA-predicted gene data, together with the PASRGA tool, into the PlantASRG database (https://bioinfor.nefu.edu.cn/PlantASRG/). This comprehensive resource aims to support crop breeding initiatives and ensure food security.

Crops, Agricultural↗

The archaeal molecular chaperone machine: peculiarities and paradoxes.

A major finding within the field of archaea and molecular chaperones has been the demonstration that, while some species have the stress (heat-shock) gene hsp70(dnaK), others do not. This gene encodes Hsp70(DnaK), an essential molecular chaperone in bacteria and eukaryotes. Due to the physiological importance and the high degree of conservation of this protein, its absence in archaeal organisms has raised intriguing questions pertaining to the evolution of the chaperone machine as a whole and that of its components in particular, namely, Hsp70(DnaK), Hsp40(DnaJ), and GrpE. Another archaeal paradox is that the proteins coded by these genes are very similar to bacterial homologs, as if the genes had been received via lateral transfer from bacteria, whereas the upstream flanking regions have no bacterial markers, but instead have typical archaeal promoters, which are like those of eukaryotes. Furthermore, the chaperonin system in all archaea studied to the present, including those that possess a bacterial-like chaperone machine, is similar to that of the eukaryotic-cell cytosol. Thus, two chaperoning systems that are designed to interact with a compatible partner, e.g., the bacterial chaperone machine physiologically interacts with the bacterial but not with the eucaryal chaperonins, coexist in archaeal cells in spite of their apparent functional incompatibility. It is difficult to understand how these hybrid characteristics of the archaeal chaperoning system became established and work, if one bears in mind the classical ideas learned from studying bacteria and eukaryotes. No doubt, archaea are intriguing organisms that offer an opportunity to find novel molecules and mechanisms that will, most likely, enhance our understanding of the stress response and the protein folding and refolding processes in the three phylogenetic domains.

Archaea↗