Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Perceptual adaptive insensitivity for support vector machine image coding.

Support vector machine (SVM) learning has been recently proposed for image compression in the frequency domain using a constant epsilon-insensitivity zone by Robinson and Kecman. However, according to the statistical properties of natural images and the properties of human perception, a constant insensitivity makes sense in the spatial domain but it is certainly not a good option in a frequency domain. In fact, in their approach, they made a fixed low-pass assumption as the number of discrete cosine transform (DCT) coefficients to be used in the training was limited. This paper extends the work of Robinson and Kecman by proposing the use of adaptive insensitivity SVMs [2] for image coding using an appropriate distortion criterion [3], [4] based on a simple visual cortex model. Training the SVM by using an accurate perception model avoids any a priori assumption and improves the rate-distortion performance of the original approach.

Algorithms↗

Gaussian processes for classification: mean-field algorithms.

We derive a mean-field algorithm for binary classification with gaussian processes that is based on the TAP approach originally proposed in statistical physics of disordered systems. The theory also yields an approximate leave-one-out estimator for the generalization error, which is computed with no extra computational cost. We show that from the TAP approach, it is possible to derive both a simpler "naive" mean-field theory and support vector machines (SVMs) as limiting cases. For both mean-field algorithms and support vector machines, simulation results for three small benchmark data sets are presented. They show that one may get state-of-the-art performance by using the leave-one-out estimator for model selection and the built-in leave-one-out estimators are extremely precise when compared to the exact leave-one-out estimate. The second result is taken as strong support for the internal consistency of the mean-field approach.

Algorithms↗

Mapping of neural networks onto the memory-processor integrated architecture.

In this paper, an effective memory-processor integrated architecture, called memory-based processor array for artificial neural networks (MPAA), is proposed. The MPAA can be easily integrated into any host system via memory interface. Specifically, the MPA system provides an efficient mechanism for its local memory accesses allowed by row and column bases, using hybrid row and column decoding, which is suitable for computation models of ANNs such as the accessing and alignment patterns given for matrix-by-vector operations. Mapping algorithms to implement the multilayer perceptron with backpropagation learning on the MPAA system are also provided. The proposed algorithms support both neuron and layer level parallelisms which allow the MPAA system to operate the learning phase as well as the recall phase in the pipelined fashion. Performance evaluation is provided by detailed comparison in terms of two metrics such as the cost and number of computation steps. The results show that the performance of the proposed architecture and algorithms is superior to those of the previous approaches, such as one-dimensional single-instruction multiple data (SIMD) arrays, two-dimensional SIMD arrays, systolic ring structures, and hypercube machines.

Journal Article↗

Transcriptomic analysis identifies novel ferroptosis-related biomarkers and therapeutic targets in pulmonary arterial hypertension.

BACKGROUND: Ferroptosis plays a significant role in pulmonary arterial hypertension (PAH), although its underlying mechanisms and key pathogenic genes remain unclear. METHODS: Transcriptomic data from human PAH and control lung tissue were obtained from the Gene Expression Omnibus (GEO) database, whereas ferroptosis-related genes (FRGs) were sourced from the MsigDb and FerrDb databases. Differentially expressed FRGs (DE-FRGs) were identified through the intersection of FRGs with differentially expressed genes (DEGs). Functional enrichment analysis was performed using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways. Key hub genes were identified through Least Absolute Shrinkage and Selection Operator (LASSO), support vector machine-recursive feature elimination (SVM-RFE), and weighted correlation network analysis (WGCNA). Gene set enrichment analysis (GSEA) was conducted to explore the functional roles and associated pathways of hub genes. The relationship between hub genes and immune infiltration was investigated. Expression levels of potential biomarkers were validated via Quantitative real-time polymerase chain reaction (qRT-PCR) and immunohistochemistry (IHC) in two PAH animal models (monocrotaline-induced and Sugen5416 plus hypoxia-induced PAH). Finally, molecular docking was employed to screen potential therapeutic compounds. RESULTS: A total of 133 DE-FRGs were identified, with KEGG and GO analyses highlighting their involvement in intracellular iron homeostasis and ferroptosis. Hub genes, notably FZD7 and NFE2, were identified using LASSO, SVM-RFE, and WGCNA. Immune infiltration analysis suggested that monocytes and neutrophils play key roles in PAH pathogenesis. Validation in PAH animal models showed significant upregulation of Fzd7 and downregulation of Nfe2 in lung tissues of both MCT- and SuHx-induced PAH models. Molecular docking identified tetrachlorodibenzodioxin (TCDD) has good binding affinity. CONCLUSION: In summary, we investigated two ferroptosis-related biomarkers, FZD7 and NFE2, in PAH using transcriptomics, offering new insights into molecular mechanisms and potential targeted therapies for the disease.

Ferroptosis↗

Prediction of standard Gibbs energies of the transfer of peptide anions from aqueous solution to nitrobenzene based on support vector machine and the heuristic method.

Quantitative structure-property relationship (QSPR) method was performed for the prediction of the standard Gibbs energies (DeltaGtheta) of the transfer of peptide anions from aqueous solution to nitrobenzene. Descriptors calculated from the molecular structures alone were used to represent the characteristics of the peptides. The four molecular descriptors selected by the heuristic method (HM) in COmprehensive DEscriptors for Structural and Statistical Analysis (CODESSA) were used as inputs for support vector machine (SVM) and radial basis function neural networks (RNFNN). The results obtained by the novel machine learning technique, SVM, were compared with those obtained by HM and RBFNN. The root mean squared errors (RMS) of the training, predicted and overall data sets are 2.192, 2.541 and 2.267 unit (kJ/mol) for HM, 1.604, 2.478 and 1.817 unit (kJ/mol) for RBFNN and 1.5621, 2.364 and 1.756 unit (kJ/mol) for SVM, respectively. The prediction results were in agreement with the experimental values. This paper provided a potential method for predicting the physiochemical property (DeltaGtheta) of various small peptides.

Anions↗

Selective networks and recognition automata.

The results we have presented demonstrate that a network based on a selective principle can function in the absence of forced learning or an a priori program to give recognition, classification, generalization, and association. While Darwin II is not a model of any actual nervous system, it does set out to solve one of the same problems that evolution had to solve--the need to form categories in a bottom-up manner from information in the environment, without incorporating the assumptions of any particular observer. The key features of the model that make this possible are (1) Darwin II incorporates selective networks whose initial specificities enable them to respond without instruction to unfamiliar stimuli; (2) degeneracy provides multiple possibilities of response to any one stimulus, at the same time providing functional redundancy against component failure; (3) the output of Darwin II is a pattern of response, making use of the simultaneous responses of multiple degenerate groups to avoid the need for very high specificity and the combinatorial disaster that would imply; (4) reentry within individual networks vitiates the limitations described by Minsky and Papert for a class of perceptual automata lacking such connections; and (5) reentry between intercommunicating networks with different functions gives rise to new functions, such as association, that either one alone could not display. The two kinds of network are roughly analogous to the two kinds of category formation that people use: Darwin, corresponding to the exemplar description of categories, and Wallace, corresponding to the probabilistic matching description of categories. These principles lead to a new class of pattern-recognizing machine of which Darwin II is just an example. There are a number of obvious extensions to this work that we are pursuing. These include giving Darwin II the capability to deal with stimuli that are in motion, an ability that probably precedes the ability of biological organisms to deal with stationary stimuli, giving it the capability to deal with multiple stimulus objects through some form of attentional mechanism, and giving it a means to respond directly and to receive feedback from the world so that it can learn conventionally. Already, however, we have shown that a working pattern-recognition automaton can be built based on a selective principle. This development promises ultimately to show us how to build recognizing machines without programs and to provide a sound basis for the study of both natural and artificial intelligence.

Automation↗

Effect of selection of molecular descriptors on the prediction of blood-brain barrier penetrating and nonpenetrating agents by statistical learning methods.

The ability or inability of a drug to penetrate into the brain is a key consideration in drug design. Drugs for treating central nervous system (CNS) disorders need to be able to penetrate the blood-brain barrier (BBB). BBB nonpenetration is desirable for non-CNS-targeting drugs to minimize potential CNS-related side effects. Computational methods have been employed for the prediction of BBB-penetrating (BBB+) and -nonpenetrating (BBB-) agents at impressive accuracies of 75-92% and 60-80%, respectively. However, the majority of these studies give a substantially lower BBB- accuracy, and thus overall accuracy, than the BBB+ accuracy. This work examined whether proper selection of molecular descriptors can improve both the BBB- and the overall accuracies of statistical learning methods. The methods tested include logistic regression, linear discriminate analysis, k nearest neighbor, C4.5 decision tree, probabilistic neural network, and support vector machine. Molecular descriptors were selected by using a feature selection method, recursive feature elimination (RFE). Results by using 415 BBB+ and BBB- agents show that RFE substantially improves both the BBB- and the overall accuracy for all of the methods studied. This suggests that statistical learning methods combined with proper feature selection is potentially useful for facilitating a more balanced and improved prediction of BBB+ and BBB- agents.

Artificial Intelligence↗

Pathway analysis using random forests classification and regression.

MOTIVATION: Although numerous methods have been developed to better capture biological information from microarray data, commonly used single gene-based methods neglect interactions among genes and leave room for other novel approaches. For example, most classification and regression methods for microarray data are based on the whole set of genes and have not made use of pathway information. Pathway-based analysis in microarray studies may lead to more informative and relevant knowledge for biological researchers. RESULTS: In this paper, we describe a pathway-based classification and regression method using Random Forests to analyze gene expression data. The proposed methods allow researchers to rank important pathways from externally available databases, discover important genes, find pathway-based outlying cases and make full use of a continuous outcome variable in the regression setting. We also compared Random Forests with other machine learning methods using several datasets and found that Random Forests classification error rates were either the lowest or the second-lowest. By combining pathway information and novel statistical methods, this procedure represents a promising computational strategy in dissecting pathways and can provide biological insight into the study of microarray data. AVAILABILITY: Source code written in R is available from http://bioinformatics.med.yale.edu/pathway-analysis/rf.htm.

Algorithms↗

Identifying genes related to chemosensitivity using support vector machine.

In an effort to identify genes involved in chemosensitivity and to evaluate the functional relationships between genes and anticancer drugs acting by the same mechanism, a supervised machine learning approach called support vector machine (SVM) is used to associate genes with any of five predefined anticancer drug mechanistic categories. The drug activity profiles are used as training examples to train the SVM and then the gene expression profiles are used as test examples to predict their associated mechanistic categories. This method of correlating drugs and genes provides a strategy for finding novel biologically significant relationships for molecular pharmacology.

Algorithms↗

[Application of support vector machine in the detection of early cancer].

Support Vector Machine (SVM) is an efficient novel method originated from the statistical learning theory. It is powerful in machine learning to solve problems with finite samples. Due to the deficiency of cancer cells, character of patient and noise in the raw data, it is very difficult to diagnose early cancer accurately. In this paper, SVM is employed in detecting early cancer and the results are encouraged compared with conventional methods. The accuracy of Non-linear SVM classifier is especially high in all kinds of classifiers, which indicates the potential application of SVM in early cancer detection.

Algorithms↗

Simple recurrent networks learn context-free and context-sensitive languages by counting.

It has been shown that if a recurrent neural network (RNN) learns to process a regular language, one can extract a finite-state machine (FSM) by treating regions of phase-space as FSM states. However, it has also been shown that one can construct an RNN to implement Turing machines by using RNN dynamics as counters. But how does a network learn languages that require counting? Rodriguez, Wiles, and Elman (1999) showed that a simple recurrent network (SRN) can learn to process a simple context-free language (CFL) by counting up and down. This article extends that to show a range of language tasks in which an SRN develops solutions that not only count but also copy and store counting information. In one case, the network stores information like an explicit storage mechanism. In other cases, the network stores information more indirectly in trajectories that are sensitive to slight displacements that depend on context. In this sense, an SRN can learn analog computation as a set of interdependent counters. This demonstrates how SRNs may be an alternative psychological model of language or sequence processing.

Models, Neurological↗

Data mining as a tool for research and knowledge development in nursing.

The ability to collect and store data has grown at a dramatic rate in all disciplines over the past two decades. Healthcare has been no exception. The shift toward evidence-based practice and outcomes research presents significant opportunities and challenges to extract meaningful information from massive amounts of clinical data to transform it into the best available knowledge to guide nursing practice. Data mining, a step in the process of Knowledge Discovery in Databases, is a method of unearthing information from large data sets. Built upon statistical analysis, artificial intelligence, and machine learning technologies, data mining can analyze massive amounts of data and provide useful and interesting information about patterns and relationships that exist within the data that might otherwise be missed. As domain experts, nurse researchers are in ideal positions to use this proven technology to transform the information that is available in existing data repositories into useful and understandable knowledge to guide nursing practice and for active interdisciplinary collaboration and research.

Algorithms↗

Maximizing sensitivity in medical diagnosis using biased minimax probability machine.

The challenging task of medical diagnosis based on machine learning techniques requires an inherent bias, i.e., the diagnosis should favor the "ill" class over the "healthy" class, since misdiagnosing a patient as a healthy person may delay the therapy and aggravate the illness. Therefore, the objective in this task is not to improve the overall accuracy of the classification, but to focus on improving the sensitivity (the accuracy of the "ill" class) while maintaining an acceptable specificity (the accuracy of the "healthy" class). Some current methods adopt roundabout ways to impose a certain bias toward the important class, i.e., they try to utilize some intermediate factors to influence the classification. However, it remains uncertain whether these methods can improve the classification performance systematically. In this paper, by engaging a novel learning tool, the biased minimax probability machine (BMPM), we deal with the issue in a more elegant way and directly achieve the objective of appropriate medical diagnosis. More specifically, the BMPM directly controls the worst case accuracies to incorporate a bias toward the "ill" class. Moreover, in a distribution-free way, the BMPM derives the decision rule in such a way as to maximize the worst case sensitivity while maintaining an acceptable worst case specificity. By directly controlling the accuracies, the BMPM provides a more rigorous way to handle medical diagnosis; by deriving a distribution-free decision rule, the BMPM distinguishes itself from a large family of classifiers, namely, the generative classifiers, where an assumption on the data distribution is necessary. We evaluate the performance of the model and compare it with three traditional classifiers: the k-nearest neighbor, the naive Bayesian, and the C4.5. The test results on two medical datasets, the breast-cancer dataset and the heart disease dataset, show that the BMPM outperforms the other three models.

Algorithms↗

Active learning of enhancers and silencers in the developing neural retina.

Deep learning is a promising strategy for modeling cis-regulatory elements. However, models trained on genomic sequences often fail to explain why the same transcription factor can activate or repress transcription in different contexts. To address this limitation, we developed an active learning approach to train models that distinguish between enhancers and silencers composed of binding sites for the photoreceptor transcription factor cone-rod homeobox (CRX). After training the model on nearly all bound CRX sites from the genome, we coupled synthetic biology with uncertainty sampling to generate additional rounds of informative training data. This allowed us to iteratively train models on data from multiple rounds of massively parallel reporter assays. The ability of the resulting models to discriminate between CRX sites with identical sequence but opposite functions establishes active learning as an effective strategy to train models of regulatory DNA. A record of this paper's transparent peer review process is included in the supplemental information.

Retina↗

Symbiogenesis in learning classifier systems.

Symbiosis is the phenomenon in which organisms of different species live together in close association, resulting in a raised level of fitness for one or more of the organisms. Symbiogenesis is the name given to the process by which symbiotic partners combine and unify, that is, become genetically linked, giving rise to new morphologies and physiologies evolutionarily more advanced than their constituents. The importance of this process in the evolution of complexity is now well established. Learning classifier systems are a machine learning technique that uses both evolutionary computing techniques and reinforcement learning to develop a population of cooperative rules to solve a given task. In this article we examine the use of symbiogenesis within the classifier system rule base to improve their performance. Results show that incorporating simple rule linkage does not give any benefits. The concept of (temporal) encapsulation is then added to the symbiotic rules and shown to improve performance in ambiguous/non-Markov environments.

Algorithms↗

Geometrical properties of nu support vector machines with different norms.

By employing the L1 or Linfinity norms in maximizing margins, support vector machines (SVMs) result in a linear programming problem that requires a lower computational load compared to SVMs with the L2 norm. However, how the change of norm affects the generalization ability of SVMs has not been clarified so far except for numerical experiments. In this letter, the geometrical meaning of SVMs with the Lp norm is investigated, and the SVM solutions are shown to have rather little dependency on p.

Algorithms↗

CRNPRED: highly accurate prediction of one-dimensional protein structures by large-scale critical random networks.

BACKGROUND: One-dimensional protein structures such as secondary structures or contact numbers are useful for three-dimensional structure prediction and helpful for intuitive understanding of the sequence-structure relationship. Accurate prediction methods will serve as a basis for these and other purposes. RESULTS: We implemented a program CRNPRED which predicts secondary structures, contact numbers and residue-wise contact orders. This program is based on a novel machine learning scheme called critical random networks. Unlike most conventional one-dimensional structure prediction methods which are based on local windows of an amino acid sequence, CRNPRED takes into account the whole sequence. CRNPRED achieves, on average per chain, Q3 = 81% for secondary structure prediction, and correlation coefficients of 0.75 and 0.61 for contact number and residue-wise contact order predictions, respectively. CONCLUSION: CRNPRED will be a useful tool for computational as well as experimental biologists who need accurate one-dimensional protein structure predictions.

Algorithms↗

A case study where biology inspired a solution to a computer science problem.

This paper describes how the biological theory of gene duplication described in Susumu Ohno's provocative book, Evolution by Means of Gene Duplication, was brought to bear on a vexatious problem from the domain of automated machine learning, namely the problem of architecture discovery. Six new architecture-altering operations for genetic programming were motivated by the way that new biological structures, functions, and behaviors arise in nature using gene duplication. Genetic programming with the new architecture-altering operations was then applied to the transmembrane protein segment identification problem. The out-of-sample error rate for the best genetically-evolved program achieved was slightly better than that of previously-reported human-written algorithms for this problem.

Amino Acid Sequence↗