Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Prospective recruitment of patients with congestive heart failure using an ad-hoc binary classifier.

This paper addresses a very specific problem of identifying patients diagnosed with a specific condition for potential recruitment in a clinical trial or an epidemiological study. We present a simple machine learning method for identifying patients diagnosed with congestive heart failure and other related conditions by automatically classifying clinical notes dictated at Mayo Clinic. This method relies on an automatic classifier trained on comparable amounts of positive and negative samples of clinical notes previously categorized by human experts. The documents are represented as feature vectors, where features are a mix of demographic information as well as single words and concept mappings to MeSH and HICDA classification systems. We compare two simple and efficient classification algorithms (Naïve Bayes and Perceptron) and a baseline term spotting method with respect to their accuracy and recall on positive samples. Depending on the test set, we find that Naïve Bayes yields better recall on positive samples (95 vs. 86%) but worse accuracy than Perceptron (57 vs. 65%). Both algorithms perform better than the baseline with recall on positive samples of 71% and accuracy of 54%.

Artificial Intelligence↗

SVM-based feature selection for characterization of focused compound collections.

Artificial neural networks, the support vector machine (SVM), and other machine learning methods for the classification of molecules are often considered as a "black box", since the molecular features that are most relevant for a given classifier are usually not presented in a human-interpretable form. We report on an SVM-based algorithm for the selection of relevant molecular features from a trained classifier that might be important for an understanding of ligand-receptor interactions. The original SVM approach was extended to allow for feature selection. The method was applied to characterize focused libraries of enzyme inhibitors. A comparison with classical Kolmogorov-Smirnov (KS)-based feature selection was performed. In most of the applications the SVM method showed sustained classification accuracy, thereby relying on a smaller number of molecular features than KS-based classifiers. In one case both methods produced comparable results. Limiting the calculation of descriptors to only the most relevant ones for a certain biological activity can also be used to speed up high-throughput virtual screening.

Algorithms↗

Prediction of protein solvent accessibility using support vector machines.

A Support Vector Machine learning system has been trained to predict protein solvent accessibility from the primary structure. Different kernel functions and sliding window sizes have been explored to find how they affect the prediction performance. Using a cut-off threshold of 15% that splits the dataset evenly (an equal number of exposed and buried residues), this method was able to achieve a prediction accuracy of 70.1% for single sequence input and 73.9% for multiple alignment sequence input, respectively. The prediction of three and more states of solvent accessibility was also studied and compared with other methods. The prediction accuracies are better than, or comparable to, those obtained by other methods such as neural networks, Bayesian classification, multiple linear regression, and information theory. In addition, our results further suggest that this system may be combined with other prediction methods to achieve more reliable results, and that the Support Vector Machine method is a very useful tool for biological sequence analysis.

Bayes Theorem↗

Plasticity in the vestibulo-ocular reflex: a new hypothesis.

The vestibulo-ocular reflex functions to prevent head movements from disturbing retinal images by generating compensatory eye movements to offset the head movements. In the monkey--the species mainly under consideration here--this reflex is machine-like and very effective. In the short-term, the VOR operates as an open-loop control system without the benefit of feedback and its performance is fixed and immutable: No matter what pattern of eye-head coordination the animal uses to view external objects, there is a continuing need for the VOR and it continues to operate; however, should the VOR consistently fail to stabilize the retinal images during head turns, it will gradually undergo long-term adaptive gain changes that restore, that stability. This adaptive capability is ultimately dependent upon vision, and a variety of optical devices that disturb the visual input normally associated with lead turns have been used to induce large changes in the reflex. Insofar as the monkey is concerned, all of the available evidence suggests to us that the modifiable elements underlying these long-term adjustments are located in the brainstem vestibular pathways and not, as previously suggested by others, in the floccular lobes of the cerebellum. However, the flocculus does appear to have an important, inductive role in the adaptive process providing at least part of the error signal guiding the long-term adjustments in the brainstem. In our view, the VOR is a particularly well-defined example of a plastic system and promises to be a most useful model for studying the cellular mechanisms underlying memory and learning the central nervous system.

Animals↗

Multi-omics identification and functional validation of signal regulatory protein gamma as a prognostic biomarker and immune regulator in head and neck squamous cell carcinoma.

BACKGROUND: Head and neck squamous cell carcinoma (HNSCC) comprises biologically diverse tumors, and durable responses to immune-checkpoint blockade are achieved by only a subset of patients. There remains a need for markers that connect clinical outcome with malignant-cell phenotypes and tissue-level immune organization. METHODS: We integrated The Cancer Genome Atlas HNSCC cohort (TCGA-HNSC), five Gene Expression Omnibus (GEO) validation cohorts, single-cell RNA sequencing, Visium spatial transcriptomics, cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq)-informed protein-potential inference, pharmacogenomic screening, genetic-risk analysis and experimental validation. A reconstructed 296-pipeline survival modelling framework was used to prioritize prognostic hub genes across validation-cohort-specific analyses. RESULTS: SIRPG was repeatedly ranked among the top ten selected genes in all five validation cohorts. At single-cell resolution, SIRPG-high tumor cells showed stronger malignant-cell features, immune-inhibitory and metabolic programs, Scissor-positive risk association, CLCA2/P53-related perturbation signals and inferred SIRPG-CD47/signal regulatory protein (SIRP) communication. Spatial analyses placed this axis within an immune-checkpoint-coupled niche, supported by Maxspin/multiview intercellular spatial modelling (MISTy) spatial coupling, communication analysis by optimal transport (COMMOT)-inferred CD47-SIRPG communication and scProTrans-inferred CD47/SIRPG protein-potential overlap. Functionally, SIRPG knockdown reduced HNSCC cell viability and increased apoptosis, whereas re-expression of short hairpin RNA (shRNA)-resistant SIRPG restored the CLCA2-BAX/BCL2 protein response. CONCLUSION: Together, these findings identify SIRPG as an immune-related prognostic hub and context-dependent tumor-cell regulator associated with apoptosis, immune communication and spatial microenvironmental organization in HNSCC.

Humans↗

Spike sorting based upon machine learning algorithms (SOMA).

We have developed a spike sorting method, using a combination of various machine learning algorithms, to analyse electrophysiological data and automatically determine the number of sampled neurons from an individual electrode, and discriminate their activities. We discuss extensions to a standard unsupervised learning algorithm (Kohonen), as using a simple application of this technique would only identify a known number of clusters. Our extra techniques automatically identify the number of clusters within the dataset, and their sizes, thereby reducing the chance of misclassification. We also discuss a new pre-processing technique, which transforms the data into a higher dimensional feature space revealing separable clusters. Using principal component analysis (PCA) alone may not achieve this. Our new approach appends the features acquired using PCA with features describing the geometric shapes that constitute a spike waveform. To validate our new spike sorting approach, we have applied it to multi-electrode array datasets acquired from the rat olfactory bulb, and from the sheep infero-temporal cortex, and using simulated data. The SOMA sofware is available at http://www.sussex.ac.uk/Users/pmh20/spikes.

Action Potentials↗

Integrative analysis and experiment validation of SLC12A8 as a biomarker for the malignant transition from endometriosis to endometriosis associated ovarian cancer.

Endometriosis (EM) is a chronic inflammatory, estrogen‑dependent benign gynecological disorder. A subset of patients with EM may subsequently develop endometriosis‑associated ovarian cancer (EAOC), implying a biological continuum between these two conditions. Nevertheless, the molecular events underlying the progression from benign endometriotic lesions toward EAOC remain incompletely characterized. In this study, transcriptomic datasets retrieved from the GEO database were interrogated through differentially expressed gene screening, functional enrichment analysis, and weighted gene co‑expression network analysis (WGCNA) to identify key genes and pathways relevant to EM and EAOC. Candidate genes were further prioritized by integrating survival analysis via the Kaplan‑Meier Plotter, LASSO regression, random‑forest modeling, and CIBERSORT immune‑infiltration profiling. Loss and gain‑of‑function cellular models were established using siRNA and overexpression plasmids, and in‑vitro functional assays were performed to characterize the phenotypic effects of target genes.We identified several candidate genes associated with EM and EAOC and evaluated their discriminatory performance. Among them, SLC12A8 elevated expression across EM and EAOC tissues and exhibited moderate diagnostic capacity. Higher SLC12A8 expression was also associated with poorer prognosis in EAOC patients. In‑vitro experiments further demonstrated that SLC12A8 modulates proliferation, invasion, and migration in both EM and EAOC cell lines. Collectively, our exploratory research findings support SLC12A8 as a candidate functional mediator and potential biomarker linked to EM‑EAOC pathological progression, thereby extending the mechanistic understanding of these disorders.

Female↗

LSAT: learning about alternative transcripts in MEDLINE.

MOTIVATION: Generation of alternative transcripts from the same gene is an important biological event due to their contribution in creating functional diversity in eukaryotes. In this work, we choose the task of extracting information around this complex topic using a two-step procedure involving machine learning and information extraction. RESULTS: In the first step, we trained a classifier that inductively learns to identify sentences about physiological transcript diversity from the MEDLINE abstracts. Using a large hand-built corpus, we compared the sentence classification performance of various text categorization methods. Support vector machines (SVMs) followed by the maximum entropy classifier outperformed other methods for the sentence classification task. The SVM with the radial basis function kernel and optimized parameters achieved Fbeta-measure of 91% during the 4-fold cross validation and of 74% when applied to all sentences in more than 12 million abstracts of MEDLINE. In the second step, we identified eight frequently present semantic categories in the sentences and performed a limited amount of semantic role labeling. The role labeling step also achieved very high Fbeta-measure for all eight categories. AVAILABILITY: The results of our two-step procedure are summarized in the LSAT database of alternative transcripts. LSAT is available at http://www.bork.embl.de/LSAT CONTACT: shah@embl.de SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Benchmarking↗

Improved protein secondary structure prediction using support vector machine with a new encoding scheme and an advanced tertiary classifier.

Prediction of protein secondary structures is an important problem in bioinformatics and has many applications. The recent trend of secondary structure prediction studies is mostly based on the neural network or the support vector machine (SVM). The SVM method is a comparatively new learning system which has mostly been used in pattern recognition problems. In this study, SVM is used as a machine learning tool for the prediction of secondary structure and several encoding schemes, including orthogonal matrix, hydrophobicity matrix, BLOSUM62 substitution matrix, and combined matrix of these, are applied and optimized to improve the prediction accuracy. Also, the optimal window length for six SVM binary classifiers is established by testing different window sizes and our new encoding scheme is tested based on this optimal window size via sevenfold cross validation tests. The results show 2% increase in the accuracy of the binary classifiers when compared with the instances in which the classical orthogonal matrix is used. Finally, to combine the results of the six SVM binary classifiers, a new tertiary classifier which combines the results of one-versus-one binary classifiers is introduced and the performance is compared with those of existing tertiary classifiers. According to the results, the Q3 prediction accuracy of new tertiary classifier reaches 78.8% and this is better than the best result reported in the literature.

Algorithms↗

An SVM classifier to separate false signals from microcalcifications in digital mammograms.

In this paper we investigate the feasibility of using an SVM (support vector machine) classifier in our automatic system for the detection of clustered microcalcifications in digital mammograms. SVM is a technique for pattern recognition which relies on the statistical learning theory. It minimizes a function of two terms: the number of misclassified vectors of the training set and a term regarding the generalization classifier capability. We compare the SVM classifier with an MLP (multi-layer perceptron) in the false-positive reduction phase of our detection scheme: a detected signal is considered either microcalcification or false signal, according to the value of a set of its features. The SVM classifier gets slightly better results than the MLP one (Az value of 0.963 against 0.958) in the presence of a high number of training data; the improvement becomes much more evident (Az value of 0.952 against 0.918) in training sets of reduced size. Finally, the setting of the SVM classifier is much easier than the MLP one.

Algorithms↗

Graph kernels for chemical informatics.

Increased availability of large repositories of chemical compounds is creating new challenges and opportunities for the application of machine learning methods to problems in computational chemistry and chemical informatics. Because chemical compounds are often represented by the graph of their covalent bonds, machine learning methods in this domain must be capable of processing graphical structures with variable size. Here, we first briefly review the literature on graph kernels and then introduce three new kernels (Tanimoto, MinMax, Hybrid) based on the idea of molecular fingerprints and counting labeled paths of depth up to d using depth-first search from each possible vertex. The kernels are applied to three classification problems to predict mutagenicity, toxicity, and anti-cancer activity on three publicly available data sets. The kernels achieve performances at least comparable, and most often superior, to those previously reported in the literature reaching accuracies of 91.5% on the Mutag dataset, 65-67% on the PTC (Predictive Toxicology Challenge) dataset, and 72% on the NCI (National Cancer Institute) dataset. Properties and tradeoffs of these kernels, as well as other proposed kernels that leverage 1D or 3D representations of molecules, are briefly discussed.

Anticarcinogenic Agents↗

A machine learning information retrieval approach to protein fold recognition.

MOTIVATION: Recognizing proteins that have similar tertiary structure is the key step of template-based protein structure prediction methods. Traditionally, a variety of alignment methods are used to identify similar folds, based on sequence similarity and sequence-structure compatibility. Although these methods are complementary, their integration has not been thoroughly exploited. Statistical machine learning methods provide tools for integrating multiple features, but so far these methods have been used primarily for protein and fold classification, rather than addressing the retrieval problem of fold recognition-finding a proper template for a given query protein. RESULTS: Here we present a two-stage machine learning, information retrieval, approach to fold recognition. First, we use alignment methods to derive pairwise similarity features for query-template protein pairs. We also use global profile-profile alignments in combination with predicted secondary structure, relative solvent accessibility, contact map and beta-strand pairing to extract pairwise structural compatibility features. Second, we apply support vector machines to these features to predict the structural relevance (i.e. in the same fold or not) of the query-template pairs. For each query, the continuous relevance scores are used to rank the templates. The FOLDpro approach is modular, scalable and effective. Compared with 11 other fold recognition methods, FOLDpro yields the best results in almost all standard categories on a comprehensive benchmark dataset. Using predictions of the top-ranked template, the sensitivity is approximately 85, 56, and 27% at the family, superfamily and fold levels respectively. Using the 5 top-ranked templates, the sensitivity increases to 90, 70, and 48%.

Algorithms↗

Structural semantic interconnections: a knowledge-based approach to word sense disambiguation.

Word Sense Disambiguation (WSD) is traditionally considered an Al-hard problem. A break-through in this field would have a significant impact on many relevant Web-based applications, such as Web information retrieval, improved access to Web services, information extraction, etc. Early approaches to WSD, based on knowledge representation techniques, have been replaced in the past few years by more robust machine learning and statistical techniques. The results of recent comparative evaluations of WSD systems, however, show that these methods have inherent limitations. On the other hand, the increasing availability of large-scale, rich lexical knowledge resources seems to provide new challenges to knowledge-based approaches. In this paper, we present a method, called structural semantic interconnections (SSI), which creates structural specifications of the possible senses for each word in a context and selects the best hypothesis according to a grammar G, describing relations between sense specifications. Sense specifications are created from several available lexical resources that we integrated in part manually, in part with the help of automatic procedures. The SSI algorithm has been applied to different semantic disambiguation problems, like automatic ontology population, disambiguation of sentences in generic texts, disambiguation of words in glossary definitions. Evaluation experiments have been performed on specific knowledge domains (e.g., tourism, computer networks, enterprise interoperability), as well as on standard disambiguation test sets.

Algorithms↗

Prediction of torsade-causing potential of drugs by support vector machine approach.

In an effort to facilitate drug discovery, computational methods for facilitating the prediction of various adverse drug reactions (ADRs) have been developed. So far, attention has not been sufficiently paid to the development of methods for the prediction of serious ADRs that occur less frequently. Some of these ADRs, such as torsade de pointes (TdP), are important issues in the approval of drugs for certain diseases. Thus there is a need to develop tools for facilitating the prediction of these ADRs. This work explores the use of a statistical learning method, support vector machine (SVM), for TdP prediction. TdP involves multiple mechanisms and SVM is a method suitable for such a problem. Our SVM classification system used a set of linear solvation energy relationship (LSER) descriptors and was optimized by leave-one-out cross validation procedure. Its prediction accuracy was evaluated by using an independent set of agents and by comparison with results obtained from other commonly used classification methods using the same dataset and optimization procedure. The accuracies for the SVM prediction of TdP-causing agents and non-TdP-causing agents are 97.4 and 84.6% respectively; one is substantially improved against and the other is comparable to the results obtained by other classification methods useful for multiple-mechanism prediction problems. This indicates the potential of SVM in facilitating the prediction of TdP-causing risk of small molecules and perhaps other ADRs that involve multiple mechanisms.

Algorithms↗

Electronic van der Waals surface property descriptors and genetic algorithms for developing structure-activity correlations in olfactory databases.

A methodology to facilitate the intelligent design of new odorants (e.g., musks) with specialized properties has been developed as part of an ongoing research effort in machine learning. In a traditional framework, the introduction of a new odorant is a lengthy, costly, and laborious discovery, development, and testing process. We propose to streamline this process utilizing large existing olfactory databases available through the open scientific literature as input for a new structure/activity correlation methodology. The first step in this process is to characterize each molecule in the database by an appropriate set of descriptors. To accomplish this task, an enhanced version of Breneman's Transferable Atom Equivalent (TAE) descriptor methodology will be used to create a large set of electron density derived shape/property hybrid (PEST), wavelet coefficient (WCD), and TAE histogram descriptors. We have chosen these molecular property descriptors to represent the problem because they have been shown to contain pertinent shape and electronic properties of the molecule and correlate with key modes of intermolecular interactions. Traditional QSAR methodologies, which employ fragment based descriptors, have been shown to be effective for QSAR development within homologous sets of molecules but are less effective when applied to data sets containing a great deal of structural variation. In contrast to previous attempts at SAR, our use of shape-aware electron density based molecular property descriptors has removed many of the limitations brought about by the use of descriptors based on substructure fragments, molecular surface properties, or other whole molecule descriptors. Another reason for the mixed success of past QSAR efforts can be traced to the nature of the underlying modeling problem, which is often quite complex. To meet these challenges, a genetic algorithm for pattern recognition analysis has been developed that selects descriptors which create class separation in a plot of the two largest principal components of the data while simultaneously searching for features that increase clustering of the data.

Journal Article↗

Representation learning for multi-modal spatially resolved transcriptomics data.

MOTIVATION: Spatial transcriptomics enables in-depth molecular characterization of samples on a morphology and RNA level while preserving spatial location. Integrating the resulting multi-modal data is an unsolved problem, and developing new solutions in precision medicine depends on improved methodologies. RESULTS: We introduce AESTETIK, a convolutional deep learning model that jointly integrates spatial, transcriptomics, and morphology information to learn accurate spot representations. AESTETIK yielded substantially improved cluster assignments on widely adopted technology platforms (e.g. 10x Genomics™, NanoString™) across multiple datasets. We achieved performance enhancement on structured tissues (e.g. brain) with a 21% increase in median ARI over previous state-of-the-art methods. Notably, AESTETIK also demonstrated superior performance on cancer tissues with heterogeneous cell populations, showing a 2-fold increase in breast cancer, 79% in melanoma, and 21% in liver cancer. We expect that these advances will enable a multi-modal understanding of key biological processes. AVAILABILITY AND IMPLEMENTATION: AESTETIK is implemented in Python 3 and is available as open source software at http://www.github.com/ratschlab/aestetik. The Snakemake pipeline for reproducing the results is available at http://www.github.com/ratschlab/st-rep.

Spatial Transcriptomics↗

ANN-Spec: a method for discovering transcription factor binding sites with improved specificity.

This work describes ANN-Spec, a machine learning algorithm and its application to discovering un-gapped patterns in DNA sequence. The approach makes use of an Artificial Neural Network and a Gibbs sampling method to define the Specificity of a DNA-binding protein. ANN-Spec searches for the parameters of a simple network (or weight matrix) that will maximize the specificity for binding sequences of a positive set compared to a background sequence set. Binding sites in the positive data set are found with the resulting weight matrix and these sites are then used to define a local multiple sequence alignment. Training complexity is O(lN) where l is the width of the pattern and N is the size of the positive training data. A quantitative comparison of ANN-Spec and a few related programs is presented. The comparison shows that ANN-Spec finds patterns of higher specificity when training with a background data set. The program and documentation are available from the authors for UNIX systems.

Algorithms↗

SNP-based analysis of genetic substructure in the German population.

OBJECTIVE: To evaluate the relevance and necessity to account for the effects of population substructure on association studies under a case-control design in central Europe, we analysed three samples drawn from different geographic areas of Germany. Two of the three samples, POPGEN (n = 720) and SHIP (n = 709), are from north and north-east Germany, respectively, and one sample, KORA (n = 730), is from southern Germany. METHODS: Population genetic differentiation was measured by classical F-statistics for different marker sets, either consisting of genome-wide selected coding SNPs located in functional genes, or consisting of selectively neutral SNPs from 'genomic deserts'. Quantitative estimates of the degree of stratification were performed comparing the genomic control approach [Devlin B, Roeder K: Biometrics 1999;55:997-1004], structured association [Pritchard JK, Stephens M, Donnelly P: Genetics 2000;155:945-959] and sophisticated methods like random forests [Breiman L: Machine Learning 2001;45:5-32]. RESULTS: F-statistics showed that there exists a low genetic differentiation between the samples along a north-south gradient within Germany (F(ST)(KORA/POPGEN): 1.7 . 10(-4); F(ST)(KORA/SHIP): 5.4 . 10(-4); F(ST)(POPGEN/SHIP): -1.3 . 10(-5)). CONCLUSION: Although the F(ST )-values are very small, indicating a minor degree of population structure, and are too low to be detectable from methods without using prior information of subpopulation membership, such as STRUCTURE [Pritchard JK, Stephens M, Donnelly P: Genetics 2000;155:945-959], they may be a possible source for confounding due to population stratification.

Case-Control Studies↗