Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Discovery and performance of DNA methylation panels for cancer detection and classification in blood.

Examining DNA in a liquid biopsy for non-invasive cancer detection relies on identifying dilute signal in a high background. This study aims to identify DNA methylation biomarkers for multi-cancer detection. Utilizing large tissue datasets, we apply novel search algorithms to discover confined biomarker panels capable of distinguishing tumor from normal and determining the tissue of origin. We explore the applicability to blood-based testing using targeted methylation sequencing followed by machine learning classification. We present an 8-marker panel, which successfully predicts tumors across 14 types with a 91% average sensitivity, maintaining a low false positive rate (< 0.04%). Additionally, a panel of 39 CpG sites exhibits accuracies ranging from 69% to 98% for identifying tissue of origin. When tested on 114 patient plasma samples (colon, liver, pancreatic, prostate, and stomach cancer), the 8-marker panel obtains an AUC of 0.78 with a 78% sensitivity among 32 early-stage patients (stage I-II), and 60% overall. Using the 39-marker panel in a multi-class classification model selecting only the best match, 54% of tumor samples were on average correctly assigned to the tissue of origin, and up to 80% when allowing more inclusive criteria. Using a limited set of biomarkers, our work contributes to advancing non-invasive cancer diagnostics.

DNA methylation↗

Selection of patient samples and genes for outcome prediction.

Gene expression profiles with clinical outcome data enable monitoring of disease progression and prediction of patient survival at the molecular level. We present a new computational method for outcome prediction. Our idea is to use an informative subset of original training samples. This subset consists of only short-term survivors who died within a short period and long-term survivors who were still alive after a long follow-up time. These extreme training samples yield a clear platform to identify genes whose expression is related to survival. To find relevant genes, we combine two feature selection methods -- entropy measure and Wilcoxon rank sum test -- so that a set of sharp discriminating features are identified. The selected training samples and genes are then integrated by a support vector machine to build a prediction model, by which each validation sample is assigned a survival/relapse risk score for drawing Kaplan-Meier survival curves. We apply this method to two data sets: diffuse large-B-cell lymphoma (DLBCL) and primary lung adenocarcinoma. In both cases, patients in high and low risk groups stratified by our risk scores are clearly distinguishable. We also compare our risk scores to some clinical factors, such as International Prognostic Index score for DLBCL analysis and tumor stage information for lung adenocarcinoma. Our results indicate that gene expression profiles combined with carefully chosen learning algorithms can predict patient survival for certain diseases.

Biomarkers, Tumor↗

Evaluating variable selection methods for diagnosis of myocardial infarction.

This paper evaluates the variable selection performed by several machine-learning techniques on a myocardial infarction data set. The focus of this work is to determine which of 43 input variables are considered relevant for prediction of myocardial infarction. The algorithms investigated were logistic regression (with stepwise, forward, and backward selection), backpropagation for multilayer perceptrons (input relevance determination), Bayesian neural networks (automatic relevance determination), and rough sets. An independent method (self-organizing maps) was then used to evaluate and visualize the different subsets of predictor variables. Results show good agreement on some predictors, but also variability among different methods; only one variable was selected by all models.

Algorithms↗

Finding relevant biomolecular features.

Many methods for analyzing biological problems are constrained by problem size. The ability to distinguish between relevant and irrelevant features of a problem may allow a problem to be reduced in size sufficiently to make it tractable. The issue of learning in the presence of large numbers of irrelevant features is an important one in machine learning, and recently, several methods have been proposed to address this issue. A combination of machine learning approaches and statistical analysis methods can be used to identify a set of relevant attributes for currently intractable biological problems. We call our framework F/I/E (Focus-Induce-Extract). As an example of this methodology, this paper reports on the identification of the features of mutations in collagen that are likely to be relevant in the bone disease Osteogenesis imperfecta.

Amino Acid Sequence↗

Protein ranking by semi-supervised network propagation.

BACKGROUND: Biologists regularly search DNA or protein databases for sequences that share an evolutionary or functional relationship with a given query sequence. Traditional search methods, such as BLAST and PSI-BLAST, focus on detecting statistically significant pairwise sequence alignments and often miss more subtle sequence similarity. Recent work in the machine learning community has shown that exploiting the global structure of the network defined by these pairwise similarities can help detect more remote relationships than a purely local measure. METHODS: We review RankProp, a ranking algorithm that exploits the global network structure of similarity relationships among proteins in a database by performing a diffusion operation on a protein similarity network with weighted edges. The original RankProp algorithm is unsupervised. Here, we describe a semi-supervised version of the algorithm that uses labeled examples. Three possible ways of incorporating label information are considered: (i) as a validation set for model selection, (ii) to learn a new network, by choosing which transfer function to use for a given query, and (iii) to estimate edge weights, which measure the probability of inferring structural similarity. RESULTS: Benchmarked on a human-curated database of protein structures, the original RankProp algorithm provides significant improvement over local network search algorithms such as PSI-BLAST. Furthermore, we show here that labeled data can be used to learn a network without any need for estimating parameters of the transfer function, and that diffusion on this learned network produces better results than the original RankProp algorithm with a fixed network. CONCLUSION: In order to gain maximal information from a network, labeled and unlabeled data should be used to extract both local and global structure.

Algorithms↗

A statistical problem for inference to regulatory structure from associations of gene expression measurements with microarrays.

MOTIVATION: One approach to inferring genetic regulatory structure from microarray measurements of mRNA transcript hybridization is to estimate the associations of gene expression levels measured in repeated samples. The associations may be estimated by correlation coefficients or by conditional frequencies (for discretized measurements) or by some other statistic. Although these procedures have been successfully applied to other areas, their validity when applied to microarray measurements has yet to be tested. RESULTS: This paper describes an elementary statistical difficulty for all such procedures, no matter whether based on Bayesian updating, conditional independence testing, or other machine learning procedures such as simulated annealing or neural net pruning. The difficulty obtains if a number of cells from a common population are aggregated in a measurement of expression levels. Although there are special cases where the conditional associations are preserved under aggregation, in general inference of genetic regulatory structure based on conditional association is unwarranted

Algorithms↗

Learning neural dynamics through instructive signals.

Rapid learning is essential for flexible behavior, but its basis in the brain remains unknown. Here we introduce the PRISM plasticity rule, a unifying mechanistic model of three well-established, fast-acting synaptic plasticity rules-in hippocampus, cerebellum and mushroom body-which relies exclusively on pre-synaptic activity and an "instructive signal" from another brain area. Using a multi-region network model we show that guiding PRISM plasticity with instructive signals enables the network to quickly learn extremely flexible nonlinear dynamics underlying behaviorally relevant computations, as well as to emulate unknown external system dynamics from real-time error signals, which we demonstrate with comprehensive simulations supported by exact mathematical theory. Thus, PRISM plasticity guided by instructive signals is well-suited to rapidly learn general-purpose neural computations-in contrast to canonical Hebbian rules. Finally, we show how including this plasticity rule in artificial learning algorithms can solve long-range temporal credit assignment, a long-standing challenge in machine learning.

cerebellum↗

Intelligent software for laboratory automation.

The automation of laboratory techniques has greatly increased the number of experiments that can be carried out in the chemical and biological sciences. Until recently, this automation has focused primarily on improving hardware. Here we argue that future advances will concentrate on intelligent software to integrate physical experimentation and results analysis with hypothesis formulation and experiment planning. To illustrate our thesis, we describe the 'Robot Scientist' - the first physically implemented example of such a closed loop system. In the Robot Scientist, experimentation is performed by a laboratory robot, hypotheses concerning the results are generated by machine learning and experiments are allocated and selected by a combination of techniques derived from artificial intelligence research. The performance of the Robot Scientist has been evaluated by a rediscovery task based on yeast functional genomics. The Robot Scientist is proof that the integration of programmable laboratory hardware and intelligent software can be used to develop increasingly automated laboratories.

Algorithms↗

Dynamic On-line Clustering and State Extraction: An Approach to Symbolic Learning.

Although recurrent neural nets have been moderately successful in learning to emulate finite-state machines (FSMs), the continuous internal state dynamics of a neural net are not well matched to the discrete behavior of an FSM. We describe an architecture, called DOLCE, that allows discrete states to evolve in a net as learning progresses. DOLCE consists of a standard recurrent neural net trained by gradient descent and an adaptive clustering technique that quantizes the state space. We describe two implementations of DOLCE. The first implementation, called DOLCE(u), uses an adaptive clustering scheme in an unsupervised mode to determine both the number of clusters and the partitioning of the state space as learning progresses. The second model, DOLCE(s), uses a Gaussian Mixture Model in a supervised learning framework to infer the states of an FSM. DOLCE(s) is based on the assumption that a finite set of discrete internal states is required for the task, and that the actual network state belongs to this set but has been corrupted by noise due to inaccuracy in the weights. DOLCE(s) learns to recover the discrete state with maximum a posteriori probability from the noisy state. Simulations show that both implementations of DOLCE lead to a significant improvement in generalization performance over earlier neural net approaches to FSM induction. The idea of adaptive quantization is not just applicable to DOLCE but can be applied to other domains as well.

Journal Article↗

An integrated machine learning system to computationally screen protein databases for protein binding peptide ligands.

A fairly large set of protein interactions is mediated by families of peptide binding domains, such as Src homology 2 (SH2), SH3, PDZ, major histocompatibility complex, etc. To identify their ligands by experimental screening is not only labor-intensive but almost futile in screening low abundance species due to the suppression by high abundance species. An ideal way of studying protein-protein interactions is to use high throughput computational approaches to screen protein sequence databases to direct the validating experiments toward the most promising peptides. Predictors with only good cross-validation were not good enough to screen protein databases. In the current study we built integrated machine learning systems using three novel coding methods and screened the Swiss-Prot and GenBank protein databases for potential ligands of 10 SH3 and three PDZ domains. A large fraction of predictions has already been experimentally confirmed by other independent research groups, indicating a satisfying generalization capability for future applications in identifying protein interactions.

Amino Acid Motifs↗

Selection and combination of machine learning classifiers for prediction of linear B-cell epitopes on proteins.

Recently, new machine learning classifiers for the prediction of linear B-cell epitopes were presented. Here we show the application of Receiver Operator Characteristics (ROC) convex hulls to select optimal classifiers as well as possibilities to improve the post test probability (PTP) to meet real world requirements such as high throughput epitope screening of whole proteomes. The major finding is that ROC convex hulls present an easy to use way to rank classifiers based on their prediction conservativity as well as to select candidates for ensemble classifiers when validating against the antigenicity profile of 10 HIV-1 proteins. We also show that linear models are at least equally efficient to model the available data when compared to multi-layer feed-forward neural networks.

Algorithms↗

Breast Cancer Recurrence Status Assessment in 5 Years Using Multimodal Integrated Learning: A Feasibility Study.

Despite advances in breast cancer detection and treatment, recurrence after curative therapy continues to impact long-term survival and quality of life. Therefore, early identification of high-risk patients is crucial to guide personalized treatment and follow-up strategies. Although genomic assays provide valuable prognostic insights, their high cost and limited accessibility hinder widespread adoption in clinical practice. Recent machine learning or deep learning approaches leveraging clinical, imaging, or multimodal data have shown promise but do not reflect real-world clinical scenarios. This study proposes a deep learning-based multimodal framework for predicting 5-year breast cancer recurrence using routinely collected clinical data. The framework consists of three main components. First, we adopted automated tumor segmentation with MedSAM to extract the tumor region from ultrasound images. The radiomics features are extracted from those tumor regions. Second, report features are extracted using a Med-Contrastive Pre-trained Transformers (MedCPT)-based approach incorporating predefined, clinically informed queries. Third, a multimodal integration model jointly processes image, radiomics, clinical features, and report features through modality-specific branches. The image branch employs the Ultrasound Foundation Model (USFM) as the backbone, while structured tabular data is processed using the FT-Transformer architecture. The features of all branches are fused using a mixture-of-experts (MoE)-based classifier, and the entire model is trained using a progressive fusion training strategy. Experimental results confirm the feasibility of using ultrasound images with tumor mask integration for recurrence prediction and demonstrate the additive value of integrating multiple data modalities through the proposed multimodal integration model. The final model for recurrence prediction achieved an AUC of 0.7540, accuracy of 74.61%, sensitivity of 70.41%, and specificity of 76.44%. This feasibility study's findings underscore the potential of the proposed multimodal deep learning framework to provide accessible, accurate, and generalizable recurrence risk prediction using routinely available clinical data, potentially supporting more informed treatment decisions and personalized post-treatment monitoring in real-world clinical practice.

Breast cancer recurrence↗

EEG source localization: a neural network approach.

Functional activity in the brain is associated with the generation of currents and resultant voltages which may be observed on the scalp as the electroencephelogram. The current sources may be modeled as dipoles. The properties of the current dipole sources may be studied by solving either the forward or inverse problems. The forward problem utilizes a volume conductor model for the head, in which the potentials on the conductor surface are computed based on an assumed current dipole at an arbitrary location, orientation, and strength. In the inverse problem, on the other hand, a current dipole, or a group of dipoles, is identified based on the observed EEG. Both the forward and inverse problems are typically solved by numerical procedures, such as a boundary element method and an optimization algorithm. These approaches are highly time-consuming and unsuitable for the rapid evaluation of brain function. In this paper we present a different approach to these problems based on machine learning. We solve both problems using artificial neural networks which are trained off-line using back-propagation techniques to learn the complex source-potential relationships of head volume conduction. Once trained, these networks are able to generalize their knowledge to localize functional activity within the brain in a computationally efficient manner.

Algorithms↗

Using data mining to explore complex clinical decisions: A study of hospitalization after a suicide attempt.

BACKGROUND: Medical education is moving toward developing guidelines using the evidence-based approach; however, controlled data are missing for answering complex treatment decisions such as those made during suicide attempts. A new set of statistical techniques called data mining (or machine learning) is being used by different industries to explore complex databases and can be used to explore large clinical databases. METHOD: The study goal was to reanalyze, using data mining techniques, a published study of which variables predicted psychiatrists' decisions to hospitalize in 509 suicide attempters over the age of 18 years who were assessed in the emergency department. Patients were recruited for the study between 1996 and 1998. Traditional multivariate statistics were compared with data mining techniques to determine variables predicting hospitalization. RESULTS: Five analyses done by psychiatric researchers using traditional statistical techniques classified 72% to 88% of patients correctly. The model developed by researchers with no psychiatric knowledge and employing data mining techniques used 5 variables (drug consumption during the attempt, relief that the attempt was not effective, lack of family support, being a housewife, and family history of suicide attempts) and classified 99% of patients correctly (99% sensitivity and 100% specificity). CONCLUSIONS: This reanalysis of a published study fundamentally tries to make the point that these new multivariate techniques, called data mining, can be used to study large clinical databases in psychiatry. Data mining techniques may be used to explore important treatment questions and outcomes in large clinical databases and to help develop guidelines for problems where controlled data are difficult to obtain. New opportunities for good clinical research may be developed by using data mining analyses.

Adult↗

Monitoring of complex industrial bioprocesses for metabolite concentrations using modern spectroscopies and machine learning: application to gibberellic acid production.

Two rapid vibrational spectroscopic approaches (diffuse reflectance-absorbance Fourier transform infrared [FT-IR] and dispersive Raman spectroscopy), and one mass spectrometric method based on in vacuo Curie-point pyrolysis (PyMS), were investigated in this study. A diverse range of unprocessed, industrial fed-batch fermentation broths containing the fungus Gibberella fujikuroi producing the natural product gibberellic acid, were analyzed directly without a priori chromatographic separation. Partial least squares regression (PLSR) and artificial neural networks (ANNs) were applied to all of the information-rich spectra obtained by each of the methods to obtain quantitative information on the gibberellic acid titer. These estimates were of good precision, and the typical root-mean-square error for predictions of concentrations in an independent test set was <10% over a very wide titer range from 0 to 4925 ppm. However, although PLSR and ANNs are very powerful techniques they are often described as "black box" methods because the information they use to construct the calibration model is largely inaccessible. Therefore, a variety of novel evolutionary computation-based methods, including genetic algorithms and genetic programming, were used to produce models that allowed the determination of those input variables that contributed most to the models formed, and to observe that these models were predominantly based on the concentration of gibberellic acid itself. This is the first time that these three modern analytical spectroscopies, in combination with advanced chemometric data analysis, have been compared for their ability to analyze a real commercial bioprocess. The results demonstrate unequivocally that all methods provide very rapid and accurate estimates of the progress of industrial fermentations, and indicate that, of the three methods studied, Raman spectroscopy is the ideal bioprocess monitoring method because it can be adapted for on-line analysis.

Algorithms↗

[Difficulties of describing human relationships with the help of psychoanalytic descriptive models: a critique of ego-psychology].

Progress of psychotherapy and of related behaviour sciences makes evident the importance of a better understanding of human relations. But psychoanalysis finds it hard to describe interpersonal processes without transference. In order to remain within the conceptional frame of metapsychology it has to see interaction between individuals as the oral, aggressive or sexual cathexis of an object or as satisfaction or denial of the subject by the object. The structure of the "ego", which--in analogy to medical thinking--is conceived as an organ with its functions, is considered to have no interpersonal activities. The "ego" of the classic psychoanalytic theory is chiefly occupied with itself. It has to care for its egoistical interests and to guarantee its self-preservation. As an auxiliary and meanwhile popular concept the "self" has been introduced to describe object-relations. This concept is not sharply defined. Due to its metapsychological implications it produces additional theoretical difficulties. Linguistic studies show that every inventory of words implies a certain insight into reality. For this reason the metapsychological machine-like concept of psychic structures does not permit new ideas about interpersonal relations. If we leave metapsychology and base on colloquial speech we see that the experience of "I" is much more related to persons than the rather autistic concept of the "ego" shows. Further we learn that self-preservation cannot be an egoistical interest; it depends on the attachment to others. All feelings of self-esteem depend much more on interpersonal relations than on "narcissistic regulations". From these experiences three conclusions are derived: a) One of the main qualities of the ego is the relatedness to persons. b) The concept of narcissistic regulation as a successor of primary narcissism is no longer useful. Narcissistic traits develop as the secundary compensations if the individual failed to build up satisfactory interpersonal relations. c) The revision of (a) ego-psychology and (b) theory of narcissism asks for modifications of the therapeutic technique, where now the interest is especially concentrated on interpersonal problems instead on the pathology of the ego.

Drive↗

Bounds on error expectation for support vector machines.

We introduce the concept of span of support vectors (SV) and show that the generalization ability of support vector machines (SVM) depends on this new geometrical concept. We prove that the value of the span is always smaller (and can be much smaller) than the diameter of the smallest sphere containing the support vectors, used in previous bounds (Vapnik, 1998). We also demonstrate experimentally that the prediction of the test error given by the span is very accurate and has direct application in model selection (choice of the optimal parameters of the SVM).

Learning↗

Multi-sensor integration for on-line tool wear estimation through radial basis function networks and fuzzy neural network.

On-line tool wear estimation plays a very critical role in industry automation for higher productivity and product quality. In addition, appropriate and timely decision for tool change is significantly required in the machining systems. Thus, this paper is dedicated to develop an estimation system through integration of two promising technologies, artificial neural networks (ANN) and fuzzy logic. An on-line estimation system consisting of five components: (1) data collection; (2) feature extraction; (3) pattern recognition; (4) multi-sensor integration; and (5) tool/work distance compensation for tool flank wear, is proposed herein. For each sensor, a radial basis function (RBF) network is employed to recognize the extracted features. Thereafter, the decisions from multiple sensors are integrated through a proposed fuzzy neural network (FNN) model. Such a model is self-organizing and self-adjusting, and is able to learn from the experience. Physical experiments for the metal cutting process are implemented to evaluate the proposed system. The results show that the proposed system can significantly increase the accuracy of the product profile.

Journal Article↗