Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

Data mining as a tool for research and knowledge development in nursing.

The ability to collect and store data has grown at a dramatic rate in all disciplines over the past two decades. Healthcare has been no exception. The shift toward evidence-based practice and outcomes research presents significant opportunities and challenges to extract meaningful information from massive amounts of clinical data to transform it into the best available knowledge to guide nursing practice. Data mining, a step in the process of Knowledge Discovery in Databases, is a method of unearthing information from large data sets. Built upon statistical analysis, artificial intelligence, and machine learning technologies, data mining can analyze massive amounts of data and provide useful and interesting information about patterns and relationships that exist within the data that might otherwise be missed. As domain experts, nurse researchers are in ideal positions to use this proven technology to transform the information that is available in existing data repositories into useful and understandable knowledge to guide nursing practice and for active interdisciplinary collaboration and research.

Algorithms↗

Bioinformatics and genomic medicine.

Bioinformatics is a rapidly emerging field of biomedical research. A flood of large-scale genomic and postgenomic data means that many of the challenges in biomedical research are now challenges in computational science. Clinical informatics has long developed methodologies to improve biomedical research and clinical care by integrating experimental and clinical information systems. The informatics revolution in both bioinformatics and clinical informatics will eventually change the current practice of medicine, including diagnostics, therapeutics, and prognostics. Postgenome informatics, powered by high-throughput technologies and genomic-scale databases, is likely to transform our biomedical understanding forever, in much the same way that biochemistry did a generation ago. This paper describes how these technologies will impact biomedical research and clinical care, emphasizing recent advances in biochip-based functional genomics and proteomics. Basic data preprocessing with normalization and filtering, primary pattern analysis, and machine-learning algorithms are discussed. Use of integrative biochip informatics technologies, including multivariate data projection, gene-metabolic pathway mapping, automated biomolecular annotation, text mining of factual and literature databases, and the integrated management of biomolecular databases, are also discussed.

Computational Biology↗

Diagnosing glaucoma progression: current practice and promising technologies.

PURPOSE OF REVIEW: An update on recent work is provided that has broadened our understanding of the evaluation of visual function and structure, and their use in evaluating glaucoma progression. RECENT FINDINGS: The challenge of determining visual-field progression and the implications of long-term fluctuation are reviewed and data to support the magnitude of the fluctuation are cited. The use of confirmatory testing can limit the over diagnosis of glaucoma progression. Focusing visual-field testing on the locations of present scotomas or using frequency doubling technology may provide new approaches to assessing visual function. New standardized techniques to interpret visual fields, including neural networks, unsupervised machine learning and pointwise linear regression, may provide more quantitative means for visual-field interpretation. These techniques, along with structural evaluation of the optic nerve and nerve fiber layer, are essential in glaucoma management. Optic-nerve-head photography is still a mainstay in evaluating glaucoma progression, although many technologies including scanning laser tomography, scanning laser polarimetry and optical coherence tomography offer more quantitative means to follow structural change. These modalities, in different ways, show promise in providing additional information regarding the stability of glaucoma. SUMMARY: Identifying the functional visual component as well as structural changes is essential in evaluating glaucoma progression. New techniques of testing and evaluating visual fields, the optic-nerve head, and the retinal nerve fiber layer offer exciting opportunities to more accurately identify glaucoma progression, and are likely to become more central as imaging devices and software support develop further.

Diagnostic Techniques, Ophthalmological↗

Improving glaucoma diagnosis by the combination of perimetry and HRT measurements.

PURPOSE: The aim of this study was to determine, whether the combination of morphologic data of the optic nerve head and visual field (VF) data would improve diagnosis of glaucoma, on the basis of the measurements alone. PATIENTS AND METHODS: Eighty-eight perimetric glaucomatous and 88 normal optic discs from the Erlangen Glaucoma Registry were matched for age. All normals and patients were examined in a standardized manner (Slitlamp biomicroscopy, gonioscopy, 24 h-applanation tonometry, automated VF testing, 15-degree optic disc stereographs, and Heidelberg Retina Tomograph (HRT)-scanning of the optic disc). The HRT variables were calculated in 4 optic disc sectors. All variables were calculated with the software's standard reference plane. To gain the same allocation of sectors as provided by the HRT software, the VF responses were averaged within 4 sectors. Classification results of these VF responses were compared with the summarized results within 4 sectors. Six different combinations of morphologic and VF data were used to assess their suitability to diagnose the disease. HRT measurements, and the standard output of the Octopus (HRT/PERI1), HRT measurements and the summarized sectors and their standard deviations (HRT/PERI2), HRT measurements, standard output of the octopus and the summarized sectors and their standard deviations (HRT/PERI1/PERI2), standard output of the Octopus (PERI1), summarized sectors of the Octopus and their standard deviations (PERI2) and HRT measurements. To assess the diagnostic value of the different data sets machine learning classifiers, stabilized linear discriminant analysis, classification trees, bagging, and double-bagging were applied. RESULTS: Combination of morphologic and VF data improved the automated classification rules. The accuracy to diagnose glaucoma just by VF and HRT indices was maximized for double-bagging using both diagnostic tools. An estimated misclassification probability of less than 0.07 could be achieved for the primary open angle glaucoma patients combining HRT and VF sectors by double bagging. So highest sensitivity was 95% and specificity 91%, achieved by double-bagging and combination of HRT, PERI1, and PERI2. CONCLUSIONS: The combination of optic disc measurements and VF data could not only improve glaucoma diagnosis in future, but could also help to find an objective way to diagnose glaucomatous optic atrophy. The limitation of the topographic relationship between structure and function is the individual variability of the optic disc morphology and the subjective variability of VF testing.

Female↗

Automated decision tree classification of corneal shape.

PURPOSE: The volume and complexity of data produced during videokeratography examinations present a challenge of interpretation. As a consequence, results are often analyzed qualitatively by subjective pattern recognition or reduced to comparisons of summary indices. We describe the application of decision tree induction, an automated machine learning classification method, to discriminate between normal and keratoconic corneal shapes in an objective and quantitative way. We then compared this method with other known classification methods. METHODS: The corneal surface was modeled with a seventh-order Zernike polynomial for 132 normal eyes of 92 subjects and 112 eyes of 71 subjects diagnosed with keratoconus. A decision tree classifier was induced using the C4.5 algorithm, and its classification performance was compared with the modified Rabinowitz-McDonnell index, Schwiegerling's Z3 index (Z3), Keratoconus Prediction Index (KPI), KISA%, and Cone Location and Magnitude Index using recommended classification thresholds for each method. We also evaluated the area under the receiver operator characteristic (ROC) curve for each classification method. RESULTS: Our decision tree classifier performed equal to or better than the other classifiers tested: accuracy was 92% and the area under the ROC curve was 0.97. Our decision tree classifier reduced the information needed to distinguish between normal and keratoconus eyes using four of 36 Zernike polynomial coefficients. The four surface features selected as classification attributes by the decision tree method were inferior elevation, greater sagittal depth, oblique toricity, and trefoil. CONCLUSION: Automated decision tree classification of corneal shape through Zernike polynomials is an accurate quantitative method of classification that is interpretable and can be generated from any instrument platform capable of raw elevation data output. This method of pattern classification is extendable to other classification problems.

Cornea↗

Digital pathology and spatial omics in steatohepatitis: Clinical applications and discovery potentials.

Steatohepatitis with diverse etiologies is the most common histological manifestation in patients with liver disease. However, there are currently no specific histopathological features pathognomonic for metabolic dysfunction-associated steatotic liver disease, alcohol-associated liver disease, or metabolic dysfunction-associated steatotic liver disease with increased alcohol intake. Digitizing traditional pathology slides has created an emerging field of digital pathology, allowing for easier access, storage, sharing, and analysis of whole-slide images. Artificial intelligence (AI) algorithms have been developed for whole-slide images to enhance the accuracy and speed of the histological interpretation of steatohepatitis and are currently employed in biomarker development. Spatial biology is a novel field that enables investigators to map gene and protein expression within a specific region of interest on liver histological sections, examine disease heterogeneity within tissues, and understand the relationship between molecular changes and distinct tissue morphology. Here, we review the utility of digital pathology (using linear and nonlinear microscopy) augmented with AI analysis to improve the accuracy of histological interpretation. We will also discuss the spatial omics landscape with special emphasis on the strengths and limitations of established spatial transcriptomics and proteomics technologies and their application in steatohepatitis. We then highlight the power of multimodal integration of digital pathology augmented by machine learning (ML)algorithms with spatial biology. The review concludes with a discussion of the current gaps in knowledge, the limitations and premises of these tools and technologies, and the areas of future research.

Humans↗

A grounded theory of abstraction in artificial intelligence.

In artificial intelligence, abstraction is commonly used to account for the use of various levels of details in a given representation language or the ability to change from one level to another while preserving useful properties. Abstraction has been mainly studied in problem solving, theorem proving, knowledge representation (in particular for spatial and temporal reasoning) and machine learning. In such contexts, abstraction is defined as a mapping between formalisms that reduces the computational complexity of the task at stake. By analysing the notion of abstraction from an information quantity point of view, we pinpoint the differences and the complementary role of reformulation and abstraction in any representation change. We contribute to extending the existing semantic theories of abstraction to be grounded on perception, where the notion of information quantity is easier to characterize formally. In the author's view, abstraction is best represented using abstraction operators, as they provide semantics for classifying different abstractions and support the automation of representation changes. The usefulness of a grounded theory of abstraction in the cartography domain is illustrated. Finally, the importance of explicitly representing abstraction for designing more autonomous and adaptive systems is discussed.

Artificial Intelligence↗

Plasticity of functional connectivity in the adult spinal cord.

This paper emphasizes several characteristics of the neural control of locomotion that provide opportunities for developing strategies to maximize the recovery of postural and locomotor functions after a spinal cord injury (SCI). The major points of this paper are: (i) the circuitry that controls standing and stepping is extremely malleable and reflects a continuously varying combination of neurons that are activated when executing stereotypical movements; (ii) the connectivity between neurons is more accurately perceived as a functional rather than as an anatomical phenomenon; (iii) the functional connectivity that controls standing and stepping reflects the physiological state of a given assembly of synapses, where the probability of these synaptic events is not deterministic; (iv) rather, this probability can be modulated by other factors such as pharmacological agents, epidural stimulation and/or motor training; (v) the variability observed in the kinematics of consecutive steps reflects a fundamental feature of the neural control system and (vi) machine-learning theories elucidate the need to accommodate variability in developing strategies designed to enhance motor performance by motor training using robotic devices after an SCI.

Aging↗

Combining the performance strengths of the logistic regression and neural network models: a medical outcomes approach.

The assessment of medical outcomes is important in the effort to contain costs, streamline patient management, and codify medical practices. As such, it is necessary to develop predictive models that will make accurate predictions of these outcomes. The neural network methodology has often been shown to perform as well, if not better, than the logistic regression methodology in terms of sample predictive performance. However, the logistic regression method is capable of providing an explanation regarding the relationship(s) between variables. This explanation is often crucial to understanding the clinical underpinnings of the disease process. Given the respective strengths of the methodologies in question, the combined use of a statistical (i.e., logistic regression) and machine learning (i.e., neural network) technology in the classification of medical outcomes is warranted under appropriate conditions. The study discusses these conditions and describes an approach for combining the strengths of the models.

Artificial Intelligence↗

vcfsim: flexible simulation of all-sites VCFs with missing data.

BACKGROUND |: VCFs are the most widely used data format for encoding genetic variation. By design, standard VCFs do not include data from sites where all individuals are homozygous for the reference allele ("invariant sites") and thus do not differentiate these from sites where data are completely missing. However, missing data are a key feature of biological datasets across all domains of genomics, and many recent studies have shown that missing data can introduce a variety of statistical biases in the estimation of key population genetic parameters. A solution to this limitation is to include invariant sites in a standard VCF, creating an "all-sites VCF", exposing missing and invariant sites explicitly. One hurdle to the wider adoption of all-sites VCFs is a reliable parameterized simulation framework for generating biologically realistic all-sites VCFs. RESULTS |: Here, we introduce an open-source command line tool, vcfsim, that interfaces with the popular coalescent simulation platform msprime and provides convenience functions for simulating all-sites VCFs with variable levels of ploidy and missing data. We show that the post-processed VCFs generated using vcfsim align precisely with population genetic expectations (i.e. are statistically identical to raw msprime output), accurately introduce missing data, and permit the simulation of data with varying ploidy levels, including the simulation of intraindividual ploidy variation (e.g. heterogametic sex chromosomes) and population structures. CONCLUSIONS |: Our results vcfsim is a useful and easy-to-use tool for the benchmarking of new software tools, performing population genetic inference, training of machine learning models, and the exploration of the effects of missing data in genomics data sets.

Benchmarking↗

Demographics, Overlap, and Latency of Severe Cutaneous Adverse Reactions in an FDA Database.

IMPORTANCE: Severe cutaneous adverse reactions (SCARs), including Stevens-Johnson syndrome/toxic epidermal necrolysis (SJS-TEN), drug reaction with eosinophilia and systemic symptoms (DRESS), acute generalized exanthematous pustulosis (AGEP), and generalized bullous fixed drug eruption (GBFDE), are rare but life-threatening drug hypersensitivity syndromes. Due to their low incidence and diagnostic complexity, large-scale characterization of SCAR is challenging. OBJECTIVE: To characterize the demographics, causative agents, trends, latency, and phenotypic overlap of SCAR using a large-scale, sanitized pharmacovigilance dataset from FAERS (FDA Adverse Event Reporting System). DESIGN: Cross-sectional study of spontaneous adverse event reports. Cases were drawn from the U.S. Food and Drug Administration Adverse Event Reporting System (FDA FAERS) from January 2004 to December 2023 and subjected to sanitization and deduplication. Disproportionality analysis was used to characterize causative agents. Machine learning (random forest classifiers) was used to analyze predictors of drug latency and mortality. SETTING: Global pharmacovigilance reports submitted to FAERS. PARTICIPANTS: A total of 56,683 deduplicated SCAR reports were identified, representing 0.33% of reports during the study period. EXPOSURES: Suspected causative drugs, including both small molecules and biologics. MAIN OUTCOMES AND MEASURES: Main outcomes included the frequency and distribution of SCAR syndromes, reporting trends over time, latency from drug start to reaction onset, drug-specific disproportionality (PRR, ROR, IC), and co-reporting between SCAR types and related conditions. RESULTS: A total of 56,683 unique SCAR reports were identified, including SJS-TEN (28,871), DRESS (22,444), AGEP (6,183), and GBFDE (150). We identified 237 drugs with significant disproportionality for SCAR overall. Co-reporting between SCARs was significantly enriched (p < 1e-200), suggesting overlapping phenotypes. Latency varied by drug and syndrome (median: GBFDE 3 days, AGEP 4 days, SJS-TEN 12 days, DRESS 20 days). CONCLUSIONS AND RELEVANCE: SCAR syndromes display distinct but overlapping phenotypes, with variable latency and diverse causative agents. These findings, based on the largest SCAR dataset to date, highlight the need for improved classification frameworks and molecular validation. Large-scale pharmacovigilance, integrated with genomic and histopathologic data, will be critical to improving diagnosis, mechanistic understanding, and clinical management of SCAR.

Acute Generalized Exanthematous Pustulosis↗

Genomic Language Model for Predicting Enhancers and Their Allele-Specific Activity in the Human Genome.

Predicting and deciphering the regulatory logic of enhancers is a challenging problem, due to the intricate sequence features and lack of consistent genetic or epigenetic signatures that can accurately discriminate enhancers from other genomic regions. Recent machine-learning based methods have spotlighted the importance of extracting nucleotide composition of enhancers but failed to learn the sequence context and perform suboptimally. Motivated by advances in genomic language models, we developed DNABERT-Enhancer, a novel enhancer prediction method, by applying DNABERT pre-trained language model on the human genome. We trained two different models, using large collection of enhancers curated from the ENCODE registry of candidate cis-Regulatory Elements. The best fine-tuned model achieved 88.05% accuracy with Matthews correlation coefficient of 76% on independent set aside data. Further, we present the analysis of the predicted enhancers for all chromosomes of the human genome by comparing with the enhancer regions reported in publicly available databases. Finally, we applied DNABERT-Enhancer along with other DNABERT based regulatory genomic region prediction models to predict candidate SNPs with allele-specific enhancer and transcription factor binding activity. The genome-wide enhancer annotations and candidate loss-of-function genetic variants predicted by DNABERT-Enhancer provide valuable resources for genome interpretation in functional and clinical genomics studies.

Journal Article↗

PATTY corrects open chromatin bias for improved bulk and single-cell CUT&Tag profiling.

Precise profiling of epigenomes is essential for better understanding chromatin biology and gene regulation. Cleavage Under Targets & Tagmentation (CUT&Tag) is an efficient epigenomic profiling technique that can be performed on a low number of cells and at the single-cell level. With its growing adoption, CUT&Tag datasets spanning diverse biological systems are rapidly accumulating in the field. CUT&Tag assays use the hyperactive transposase Tn5 for DNA tagmentation. Tn5's preference toward accessible chromatin alters CUT&Tag sequence read distributions in the genome and introduces open chromatin bias that can confound downstream analysis, an issue more substantial in sparse single-cell data. We show that open chromatin bias extensively exists in published CUT&Tag datasets, including those generated with recently optimized high-salt protocols. To address this challenge, we present PATTY (Propensity Analyzer for Tn5 Transposase Yielded bias), a comprehensive computational method that corrects open chromatin bias in CUT&Tag data by leveraging accompanying ATAC-seq. By integrating transcriptomic and epigenomic data using machine learning and integrative modeling, we demonstrate that PATTY enables accurate and robust detection of occupancy sites for both active and repressive histone modifications, including H3K27ac, H3K27me3, and H3K9me3, with experimental validation. We further develop a single-cell CUT&Tag analysis framework built on PATTY and show improved cell clustering when using bias-corrected single-cell CUT&Tag data compared to using uncorrected data. Beyond CUT&Tag, PATTY sets a foundation for further development of bias correction methods for improving data analysis for all Tn5-based high-throughput assays.

Journal Article↗

Gene Specific Pathogenicity Predictor for Chromatin-Remodeling BAF Complex-Associated Neurodevelopmental Disorders.

Advancements in whole genome sequencing have increased the number of variants of uncertain significance (VUS) identified in patient genomes. This has created a diagnostic bottleneck for genetic counselors tasked with sifting through these variants and determining those most likely to be causative for a patient's clinical presentation. Machine learning (ML) tools can aid in identifying pathogenic variants from VUS, but there is a need for gene-specific algorithms that predict pathogenic variants with high accuracy. To address this need, we present a workflow for developing gene-specific, ensemble-learning ML tools, that leverage outputs from other algorithms, locations of variants within the gene, and evolutionary conservation data to make a prediction of pathogenicity. Variants in SMARCA2 and SMARCA4 that are associated with rare neurodevelopmental diseases were used to screen 15 ML algorithms. A random forest learner was tuned to yield a final accuracy of 0.93 on holdout data. Generalizing this predictor to other BAF complex proteins resulted in a sharp decline in performance. We trained a final predictor for all genes in the study to create a predictor that identifies pathogenic variants in these BAF subunits with an accuracy of 0.91 on holdout data. This predictor specific to BAF complex proteins performs with higher accuracy and AUROC than any other predictor. The decline in performance when generalized to other proteins emphasizes the need for the gene-specific calibration of predictors. Our workflow for the development of such models provides a quick, computationally inexpensive route for improving the ML tools available to genetic counselors.

Journal Article↗

Discovery and performance of DNA methylation panels for cancer detection and classification in blood.

Examining DNA in a liquid biopsy for non-invasive cancer detection relies on identifying dilute signal in a high background. This study aims to identify DNA methylation biomarkers for multi-cancer detection. Utilizing large tissue datasets, we apply novel search algorithms to discover confined biomarker panels capable of distinguishing tumor from normal and determining the tissue of origin. We explore the applicability to blood-based testing using targeted methylation sequencing followed by machine learning classification. We present an 8-marker panel, which successfully predicts tumors across 14 types with a 91% average sensitivity, maintaining a low false positive rate (< 0.04%). Additionally, a panel of 39 CpG sites exhibits accuracies ranging from 69% to 98% for identifying tissue of origin. When tested on 114 patient plasma samples (colon, liver, pancreatic, prostate, and stomach cancer), the 8-marker panel obtains an AUC of 0.78 with a 78% sensitivity among 32 early-stage patients (stage I-II), and 60% overall. Using the 39-marker panel in a multi-class classification model selecting only the best match, 54% of tumor samples were on average correctly assigned to the tissue of origin, and up to 80% when allowing more inclusive criteria. Using a limited set of biomarkers, our work contributes to advancing non-invasive cancer diagnostics.

DNA methylation↗

Architectural logic of the 3D genome: mechanisms of dysregulation and emerging cancer therapeutics.

The three-dimensional (3D) genome provides an essential layer of organization that shapes genome function in space and time. Chromatin compartments and topologically associating domains (TADs) arise from the interplay between intrinsic properties of chromatin and architectural factors, including cohesin and CTCF. Despite substantial progress in defining these structural features, whether 3D genome architecture plays a causal role in regulating processes such as transcription, DNA replication, and DNA repair, or instead reflects underlying regulatory activity, remains unresolved. Here, we use the distinction between chromatin-intrinsic features and architectural factors as a framework to evaluate evidence for causality in genome structure-function relationships. We extend this framework to cancer, where both intrinsic alterations (including noncoding mutations, structural variants, and changes in chromatin state) and architectural factor perturbations (such as mutations in architectural proteins and dysregulation of transcriptional machinery) disrupt genome organization and contribute to disease progression. These findings suggest that alterations in genome structure can, in some contexts, actively reshape oncogenic programs. A major limitation in applying 3D genome insights to cancer biology is the cost and complexity of omics assays. Recent advances in artificial intelligence (AI) and machine learning (ML) enable inference and prediction of 3D genome organization from sequence and epigenomic features, providing insight into the extent to which genome folding is encoded intrinsically versus dynamically regulated in architectural factors. This perspective provides a unified view of how genome structure is established, how it relates to function, and how its disruption contributes to tumorigenesis.

3D genome↗

Sequence information for the splicing of human pre-mRNA identified by support vector machine classification.

Vertebrate pre-mRNA transcripts contain many sequences that resemble splice sites on the basis of agreement to the consensus,yet these more numerous false splice sites are usually completely ignored by the cellular splicing machinery. Even at the level of exon definition,pseudo exons defined by such false splices sites outnumber real exons by an order of magnitude. We used a support vector machine to discover sequence information that could be used to distinguish real exons from pseudo exons. This machine learning tool led to the definition of potential branch points,an extended polypyrimidine tract,and C-rich and TG-rich motifs in a region limited to 50 nt upstream of constitutively spliced exons. C-rich sequences were also found in a region extending to 80 nt downstream of exons,along with G-triplet motifs. In addition,it was shown that combinations of three bases within the splice donor consensus sequence were more effective than consensus values in distinguishing real from pseudo splice sites; two-way base combinations were optimal for distinguishing 3' splice sites. These data also suggest that interactions between two or more of these elements may contribute to exon recognition,and provide candidate sequences for assessment as intronic splicing enhancers.

Artificial Intelligence↗

Computational detection and location of transcription start sites in mammalian genomic DNA.

Transcription, the process whereby RNA copies are made from sections of the DNA genome, is directed by promoter regions. These define the transcription start site, and also the set of cellular conditions under which the promoter is active. At least in more complex species, it appears to be common for genes to have several different transcription start sites, which may be active under different conditions. Eukaryotic promoters are complex and fairly diffuse structures, which have proven hard to detect in silico. We show that a novel hybrid machine-learning method is able to build useful models of promoters for >50% of human transcription start sites. We estimate specificity to be >70%, and demonstrate good positional accuracy. Based on the structure of our learned models, we conclude that a signal resembling the well known TATA box, together with flanking regions of C-G enrichment, are the most important sequence-based signals marking sites of transcriptional initiation at a large class of typical promoters.

Animals↗