Search PubMedSearch

SEARCH · Search PubMed

Results for “Empirical data training”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

15 recordsLinked to original sources

IQ-NET: fast and accurate quartet phylogenetic inference using deep learning trained on empirical DNA alignments.

Phylogenetic inference is fundamental to modern biology, with many applications including evolutionary biology, epidemiology, and comparative genomics. While maximum likelihood and Bayesian methods remain the gold standard for phylogenetic analysis, they rely on simplifying assumptions and are computationally intensive. Recent machine learning approaches for phylogenetics offer speed advantages, but have several limitations: exclusive reliance on simulated data for training, inadequate handling of gaps, and sensitivity to input sequence order. Here, we introduce IQ-NET (Intelligent Quartet NETwork), a deep learning framework that solves these limitations to infer four-taxon trees. IQ-NET estimates both tree topology and branch lengths directly from gapped alignments. IQ-NET outperforms existing machine learning methods in terms of accuracy, and obtained a 24-fold speedup compared with the widely used maximum likelihood software, IQ-TREE. We finally introduce a pipeline using IQ-NET and the ASTRAL software to reconstruct a larger species tree, i.e., with more than four taxa.

Empirical data training

NanoSSL: attention mechanism-based self-supervised learning method for protein identification using nanopores.

MOTIVATION: Nanopores are cutting-edge interdisciplinary tools that can analyze biomolecules at the single-molecule level for many applications, e.g. DNA sequencing. Efforts are underway to extend nanopores to proteomics, including the development of machine learning algorithms for protein sequencing and identification. However, single-molecule data are intrinsically noisy and hard to process. Moreover, the development and performance of machine learning for nanopore is jeopardized by data scarcity. Self-supervised learning is an emerging method that may yield advantages in nanopore scenarios. RESULTS: We propose and experimentally validate Nanopore analysis using Self-Supervised Learning (NanoSSL), a generative self-supervised learning framework based on attention mechanisms for the identification of protein signals from nanopores. Leveraging a two-step approach consisting of self-supervised pre-training and supervised fine-tuning, NanoSSL learns useful feature representations from empirical data to facilitate downstream classification tasks. Inspired by the concept of fragmentation in conventional protein sequencing technologies, during pretraining each translocation event is split into multiple non-overlapping fragments of equal size, some of which are randomly masked and reconstructed using a masked autoencoder. Learning the feature representations of the reconstructed nanopore events facilitates molecular identification in fine-tuning. In this study, we retested a publicly available nanopore multiplexed protein sensing dataset for model iteration, and subsequently measured Alzheimer's disease biomarker Aβ1-42 using homemade solid-state nanopores. Empirical results indicated NanoSSL achieved an unprecedented performance across four metrics: accuracy, precision, recall, and F1 score, when classifying two mutated Aβ1-42, E22G and G37R. The self-supervised learning and attention mechanism were verified as the source of performance gains. AVAILABILITY AND IMPLEMENTATION: The main program is available at https://doi.org/10.5281/zenodo.17172822.

Nanopores

Computer separation of unitary spikes from whole-nerve recordings.

A practical and efficient off-line computer technique is described for automatically separating unitary waveforms from multiunit whole-nerve spike train data with no a priori knowledge about the number of units or their waveforms. The procedure requires two recording eletrodes which provide 3 measures on each putative unit: (1) peak-to-peak amplitude on the proximal channel, (2) amplitude on the distal channel, and (3) temporal offset (depends on conduction velocity) between proximal and distal spikes. On the basis of these 3 measurements, individual unitary spikes are automatically separated into clusters according to empirically-determined limits of variability. The results of the program are displayed in 3-D plots of the 3 measures on each unitary spike and in plots of superimposed waveforms from each cluster. These plots can be used to interactively correct clustering errors. The procedure is illustrated with a 1-min segment of spike train data recorded in vivo from the siphon nerve of a freely-behaving Aplysia. We routinely obtain about 10 relatively well-isolated units in such segments. By utilizing the average waveforms and conduction velocities for individual clusters, it may eventually be possible to separate unitary spikes from compound waveforms resulting from simultaneous of two or more units.

Animals

T-SMmOTE: tweaked synthetic majority minority oversampling technique for data scarcity issue in multi omics studies.

MOTIVATION: Multiomics data offer a rich data mine for modeling complex as well as day-to-day diseases, but their practical deployment is constrained by the limited sample availability. To this end, generating synthetic samples is a viable remedy. Extant schemes operating along this line, however, are mostly limited to augmenting the minority class in imbalanced datasets and often produce synthetic samples that lack sufficient diversity and fail to faithfully capture the underlying data distribution. As a result, the full potential of synthetic augmentation in multi-omics learning remains underexplored. The aim is to address the data scarcity problem in multi-omics domain. We propose a synthetic oversampling framework, which is dedicated to addressing overall data scarcity in multi-omics datasets and the lack of diversity in synthetic samples. Contrary to conventional methods that restrict augmentation to minority classes and rely on interpolation of two neighbors, our method generates diverse yet distribution-aligned synthetic samples by interpolating three neighbors and extends this augmentation paradigm to the majority class. The framework first balances the dataset by generating synthetic minority samples, and subsequently augments the balanced dataset by oversampling both majority and minority classes. RESULTS: Empirical evaluation on multi-omics data obtained from three heterogeneous health scenarios-inflammatory bowel disease, multi-organ dysfunction syndrome, and colorectal cancer-substantiates the utility of the proposed scheme in improving the predictive performance. The models trained on T-SMmOTE-augmented data achieve higher Matthews correlation coefficient values, along with improvedscores for both majority and minority classes. Notably, oversampling of the majority class improves the cognition of the minority class as well. We also explore the consistency of the class distributions between the original and augmented class-specific datasets. These findings confirm the capability of our scheme to learn from small, high-dimensional multi-omics datasets and highlight its potential for non-invasive disease detection. AVAILABILITY AND IMPLEMENTATION: https://github.com/payelu/TSMm.

Journal Article

MO-GCAN: multi-omics integration based on graph convolutional and attention networks.

MOTIVATION: Cancer subtypes play a critical role in disease progression, prognosis, and treatment, making their detection essential for tailoring precision medicine. Studies have shown that multi-omics integration outperforms single-omics approaches in cancer subtyping tasks. However, due to the high-dimensionality of multi-omics data, many existing studies either fail to capture the correlation between true labels and learned features, or lack sufficient capacity to model complex biological representations. These limitations hinder the full potential of leveraging the rich and complementary information embedded in multi-omics datasets. RESULT: We propose a framework that leverages supervised feature learning and classification based on a graph-based learning approach with attention mechanism for cancer subtyping. More specifically, we train graph convolutional network models on each omics dataset to extract latent representations, which are then concatenated to form a comprehensive multi-omics feature embedding. We further develop sample fusion network based on the omics-specific graphs, incorporating the derived features and feeding them into a graph attention model for subtype classification. This two-stage multi-omics framework is applied to eight cancer types, with performance evaluated in terms of test accuracy, training time, macro-averaged precision, recall, and F-score. Experimental results show that the proposed method outperforms state-of-the-art approaches across various cancer types. Additionally, we provide empirical evidence supporting the hypothesis that retaining a limited number of high-confidence edges and utilizing enriched embeddings from intermediate graph neural network layers can improve predictive performance. AVAILABILITY AND IMPLEMENTATION: Data and the code are available at https://github.com/YD-00/MO-GCAN-Updated.git.

Neoplasms

Proteomics at scale: Bottlenecks and opportunities for early-career researchers in a fast developing field.

The field of proteomics has rapidly evolved over the last five years enabled by rapid advances in instrumentation and computation. At the same time, the proteomics community is also growing. This is reflected by the increasing participation in international conferences such as those organized by the European Proteomics Association and the Human Proteome Organization. These events provide early-career researchers with unique opportunities to exchange ideas, develop collaborations, and build networks that support professional development. One such network is the Young Proteomics Investigators Club, a European initiative supported by European Proteomics Association and led by early-career researchers. In this Community-Driven project, we investigate recent trends in proteomics by screening conference abstracts and evaluating the session attendance at Human Proteome Organization Congresses and European Proteomics Association conferences. Based on these analyses, we identified five areas that, from our perspective, are shaping the current trends in proteomics: clinical proteomics, proteomics of post-translational modifications, single-cell proteomics, systems biology and multi-omics, and computational proteomics. For each area, we highlight both unique challenges and identify a common theme: a shift from exploratory studies with manageable sample numbers towards large screenings and cohorts and the generation of big data, which often comes with the lack of computational support, organizational networks, and infrastructure. In this light, we describe the unique challenges and opportunities faced by early-career researchers. We point to actionable directions for enabling reproducible and transparent proteomics as well as community-driven projects and initiatives, which are often providing training and support. SIGNIFICANCE: In this perspective, the Young Proteomics Investigators Club (YPIC) discusses advances in analytical developments and computational approaches in proteomics research. Based on empirical analysis of recent European Proteomics Association conference and Human Proteome Organization congresses contributions, we identify clinical, single-cell, post-translational and systems-level proteomics as the research areas that have gained most momentum in the last three to five years. What makes this work distinctive is that it is written by and for early-career researchers, thereby uniquely identifying where momentum, challenges, and unmet needs converge for the newest generation of proteomics researchers. Rather than cataloguing advances, we examine the widening gap between what modern proteomics can generate and what individual researchers can realistically process, validate, and interpret. We describe specific structural barriers including access to high performance computing, limited formal training in scalable data analysis, the need for unified benchmarking standards and navigating clinical collaboration frameworks. We then highlight opportunities for the field, such as community-curated benchmarks, interdisciplinary mentorship models, and shared computational infrastructure. By making these challenges explicit from an early-career researchers standpoint, we aim to inform how training, funding, and community initiatives can be shaped to support the next generation of proteomics researchers.

Proteomics

Spectral-Proteomic Integration Analysis (SPIA) Deciphers Molecular Trajectories of Breast Cancer and Enables Multitarget Therapeutic Assessment.

Raman spectroscopy and mass spectrometry-based proteomics offer deeply complementary yet largely disconnected views of cancer biology: the former provides a label-free, real-time biochemical phenotype, while the latter delivers a quantitative inventory of specific protein effectors. Bridging this gap remains a fundamental challenge in analytical biomedicine. Here, we introduce Spectral-Proteomic Integration Analysis (SPIA)─a novel, data-driven integrative framework that systematically links Raman spectroscopic phenotypes with quantitative proteomic profiles through machine learning and statistical correlation. Using a DMBA-induced rat breast cancer model with and without Toremifene (TOR) intervention, SPIA dynamically maps tumor microenvironment remodeling, capturing progressive collagen deposition and lipid metabolic reprogramming. An SVM classifier trained on Raman spectra achieves exceptional diagnostic accuracy (AUC ≥ 99.0%) and successfully predicts TOR therapeutic response. Proteomic analysis identifies 1,350 differentially expressed proteins, with convergent machine learning feature selection (LASSO, Random Forest, XGBoost) pinpointing core regulators including Luc7l2, Nucb1, Cbx3, and Csnk2a1. Crucially, Spearman correlation analysis between key Raman bands and core DEPs reveals strong, statistically robust associations (median ρ ∼ 0.75 in the 1533-1669 cm-1 region), empirically validating SPIA's core integrative logic. Leveraging this multimodal map, we elucidate a multitarget mechanism for TOR involving concurrent suppression of collagen deposition and correction of aberrant lipid metabolism. SPIA establishes a powerful, generalizable paradigm for integrating phenotypic and molecular data, with broad implications for biomarker discovery, drug mechanism elucidation, and precision oncology.

Animals

DNABERT-S: Pioneering Species Differentiation with Species-Aware DNA Embeddings.

We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e., DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 23 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. Model, codes, and data is publicly available at https://github.com/MAGlCS-LAB/DNABERT_S.

Journal Article

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings.

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Sequence Analysis, DNA

Improving spliced alignment by modeling splice sites with deep learning.

MOTIVATION: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alignment demands sophisticated modeling of splice sites, but current aligners use simple models, which may affect their accuracy given dissimilar sequences. RESULTS: We implemented minisplice to learn splice signals with a one-dimensional convolutional neural network (1D-CNN) and trained a model with 7,026 parameters for vertebrate and insect genomes. It captures conserved splice signals across phyla and reveals GC-rich introns specific to mammals and birds. We used this model to estimate the empirical splicing probability for every GT and AG in genomes, and modified minimap2 and miniprot to leverage pre-computed splicing probability during alignment. Evaluation on human long-read RNA-seq data and cross-species protein datasets showed our method greatly improves the junction accuracy especially for noisy long RNA-seq reads and proteins of distant homology. AVAILABILITY AND IMPLEMENTATION: https://github.com/lh3/minisplice.

Journal Article

The effects of mood upon imaginal thought.t.

The effects of mood upon imaginal thought were explored with a highly trained undergraduate female hypnotic subject. She was hypnotically programmed to experience free-floating anxiety or pleasure in varying degrees just before the exposure of combinations of three Blacky Pictures, and to produce dreamlike imagery in response to the Blacky stimuli while under sway of the mood. Data from 98 dream trials, separated by amnesia, indicated that the affective states clearly influenced imaginal processes. Blind ratings by a psychoanalyst showed anxiety moods to be more closely associated with primary-process features characteristic of nocturnal dreams, whereas pleasure had a relatively higher incidence of daydreamlike ratings. Empirical analysis of themes yielded significant relationships of anxiety to physical injury to the self and verbal aggression toward others; pleasure was associated with circular movements and overt sex themes.

Adult

Assessment of genomic prediction capabilities of transcriptome data in a barley multi-parent RIL population.

Low-cost and high-throughput RNA sequencing data for barley RILs achieved GP performance comparable to or better than traditional SNP array datasets when combined with parental whole-genome sequencing SNP data. The field of genomic selection (GS) is advancing rapidly on many fronts including the utilization of multi-omics datasets with the goal of increasing prediction ability and becoming an integral part of an increasing number of breeding programs ensuring future food security. In this study, we used RNA sequencing (RNA-Seq) data to perform genomic prediction (GP) on three related barley RIL populations. We investigated the potential of increasing prediction ability by combining genomic and transcriptomic datasets, adding whole-genome sequencing (WGS) SNP data, functional annotation-based filtering, and empirical quality filtering. Our RNA-Seq data were generated cost-efficiently using small-footprint plant cultivation, high-throughput RNA extraction, and Library preparation miniaturization. We also examined sequencing depth reduction as an additional cost-saving measure. We used fivefold cross-validation to evaluate the prediction ability of the gene expression dataset, the RNA-Seq SNP dataset, and the consensus SNP dataset between the RNA-Seq and parental WGS data, resulting in prediction abilities between 0.73 and 0.78. The consensus SNP dataset performed best, with five out of eight traits performing significantly better compared to a 50K SNP array, which served as a benchmark. The advantage of the consensus SNP dataset was most prominent in the inter-population predictions, in which the training and validation sets originated from different RIL sub-populations. We were therefore able to not only show that RNA-Seq data alone are able to predict various complex traits in barley using RILs, but also that the performance can be further increased with WGS data for which the public availability will steadily increase.

Hordeum

What do we know about medical invalidation and related concepts? - A scoping review and thematic analysis about the definitions, measurements, causes, consequences and potential solutions for medical invalidation.

BACKGROUND: Medical invalidation, medical gaslighting, and related constructs have gained visibility in public discourse but remain inconsistently defined in scientific literature. Despite growing research-often focused on specific diseases- to date, no single review has comprehensively synthesized their definitions, causes, consequences, or methods of measurement. This scoping review addresses this gap by examining medical invalidation and related constructs. METHODS: Using a preregistered protocol, we systematically searched PubMed, CINAHL, Web of Science, Google Scholar, and ProQuest (dissertations) without year restrictions. Eligible sources included peer-reviewed empirical, theoretical, and conceptual work in English addressing invalidation, gaslighting, or closely related notions within healthcare. A total of 158 studies were identified through database searches and citation tracking. Data extraction followed a standardized schema, and findings were synthesized descriptively and through thematic analysis to clarify terminology, map determinants and outcomes, and identify existing measurement approaches. RESULTS: The results showed substantial inconsistency in how "invalidation," "not being taken seriously," and "gaslighting" were defined. Medical invalidation emerged as a multifactorial phenomenon driven by diagnostic challenges, structural and societal factors, provider and patient characteristics, stigma, misattribution, interactional dynamics, academic knowledge gaps, and disease-related complexity. Invalidation was associated with wide-ranging behavioural, emotional, cognitive, physical, relational, and systemic harms, while validation had consistently beneficial effects. Proposed solutions in the summarized studies included communication improvements, clinician training, patient support, targeted research, and structural and systemic changes. DISCUSSIONS: Medical invalidation represents a complex, systemic issue with significant implications for patient safety. The discussion highlights its multifactorial origins, its potential to cause both psychological and physical harm, and the need for clearer conceptualisation within the field. Advancing research requires validated instruments and longitudinal designs to examine underlying mechanisms and consequences. Addressing medical invalidation will demand multi-level interventions to improve communication, reduce structural barriers, and promote equitable, patient-centred care. OSF PREREGISTRATION: https://doi.org/10.17605/OSF.IO/MPE6U.

Humans

Temporal patterns, their distribution and redundancy in trains of spontaneous neuronal spike intervals of the feline hippocampus studied with a non-parametric technique.

A modification of the non-parametric technique for the analysis of temporal patterns in long trains of single neuronal spike intervals has been described and tested empiracally. The technique is based on inequality testing of sequential pairs of intervals. If the second interval in a pair is longer or shorter than or equal to the first interval, a(+), a(-), and a (0) is recorded respectively in sequential bins of the computer memory. Subsequently, the long sequences of signs are arranged into transition frequency matrices which are then converted into transition probability matrices of various complexity. In this manner, the sign permutations composed of 4, 5, 6, etc. signs were studied. First of all, the theoretical distribution of various sign permutations was derived, assuming that the arrangement of intervals that generate the signs is totally independent. The theoretical distribution of signs permutations in tetragrams, pentagrams and hexagrams constitute the 'controls' with which the empirical data can be compared. In this manner, using the chi-square test, the total deviation of a studied neuronal output from an independent state can be quantified. The empirical data showed a consistent deviation from the theoretical distribution of sign permutations during REM sleep, as compared to slow wave sleep which was characterized by an almost perfect theoretical distribution of sign permutations. This indicates that slow wave sleep is associated with relaxation of constraints that are responsible for the emergence of specific patterns. In addition, redundancies in the occurrence of sign permutations, and the linear relationships between them, have been defined and tested empirically. The apparent discrepancies between the redundancies, based on theoretical symmetry in sign distribution and the linear redundancy that fits the empirical data, have been defined and discussed.

Action Potentials

Preparation for labor: a historical perspective.

A historical analysis of the literature pertaining to psychoprophylaxis demonstrates that contemporary treatment methods have diverse and complex origins. Although many training manuals are presented as outlines of the "Lamaze" method, historical evidence indicates that Grantly Dick-Read (Natural Childbirth. London, Heinemann, 1933; Childbirth Without Fear. New York, Harper and Brothers, 1944), an English obstetrician, made the most substantive contributions to this area. Although Fernand Lamaze is generally regarded as the pre-eminent authority on psychoprophylaxis, a comparison of his 1958 text with the original Soviet source (I. Velvovsky et al., (Eds.), Painless Childbirth Through Psychoprophylaxis, Moscow, Foreign Languages Publ. House, 1960) demonstrates that he deleted and modified substantial portions of the treatment regimen and failed to keep abreast of developments in Soviet theory. Neither Dick-Read, Velvovsky et al. or Lamaze (Painless Childbirth. London, Burke, 1958) present data which permit cause and effect conclusions regarding treatment and outcome. By the same token, none of these authors demonstrated interest in the empirical validation of their theories regarding pain, anxiety, or fear reduction.

Female