Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

Gradient-based optimization of hyperparameters.

Many machine learning algorithms can be formulated as the minimization of a training criterion that involves a hyperparameter. This hyperparameter is usually chosen by trial and error with a model selection criterion. In this article we present a methodology to optimize several hyperparameters, based on the computation of the gradient of a model selection criterion with respect to the hyperparameters. In the case of a quadratic training criterion, the gradient of the selection criterion with respect to the hyperparameters is efficiently computed by backpropagating through a Cholesky decomposition. In the more general case, we show that the implicit function theorem can be used to derive a formula for the hyperparameter gradient involving second derivatives of the training criterion.

Algorithms↗

Optimization of the kernel functions in a probabilistic neural network analyzing the local pattern distribution.

This article proposes a procedure for the automatic determination of the elements of the covariance matrix of the gaussian kernel function of probabilistic neural networks. Two matrices, a rotation matrix and a matrix of variances, can be calculated by analyzing the local environment of each training pattern. The combination of them will form the covariance matrix of each training pattern. This automation has two advantages: First, it will free the neural network designer from indicating the complete covariance matrix, and second, it will result in a network with better generalization ability than the original model. A variation of the famous two-spiral problem and real-world examples from the UCI Machine Learning Repository will show a classification rate not only better than the original probabilistic neural network but also that this model can outperform other well-known classification techniques.

Neural Networks, Computer↗

On the asymptotic distribution of the least-squares estimators in unidentifiable models.

In order to analyze the stochastic property of multilayered perceptrons or other learning machines, we deal with simpler models and derive the asymptotic distribution of the least-squares estimators of their parameters. In the case where a model is unidentified, we show different results from traditional linear models: the well-known property of asymptotic normality never holds for the estimates of redundant parameters.

Algorithms↗

An approach to biological computation: unicellular core-memory creatures evolved using genetic algorithms.

A novel machine language genetic programming system that uses one-dimensional core memories is proposed and simulated. The core is compared to a biochemical reaction space, and in imitation of biological molecules, four types of data words (Membrane, Pure data, Operator, and Instruction) are prepared in the core. A program is represented by a sequence of Instructions. During execution of the core, Instructions are transcribed into corresponding Operators, and Operators modify, create, or transfer Pure data. The core is hierarchically partitioned into sections by the Membrane data, and the data transfer between sections by special channel Operators constitutes a tree data-flow structure among sections in the core. In the experiment, genetic algorithms are used to modify program information. A simple machine learning problem is prepared for the environment data set of the creatures (programs), and the fitness value of a creature is calculated from the Pure data excreted by the creature. Breeding of programs that can output the predefined answer is successfully carried out. Several future plans to extend this system are also discussed.

Algorithms↗

Evolving mobile robots in simulated and real environments.

The problem of the validity of simulation is particularly relevant for methodologies that use machine learning techniques to develop control systems for autonomous robots, as, for instance, the artificial life approach known as evolutionary robotics. In fact, although it has been demonstrated that training or evolving robots in real environments is possible, the number of trials needed to test the system discourages the use of physical robots during the training period. By evolving neural controllers for a Khepera robot in computer simulations and then transferring the agents obtained to the real environment we show that (a) an accurate model of a particular robot-environment dynamics can be built by sampling the real world through the sensors and the actuators of the robot; (b) the performance gap between the obtained behaviors in simulated and real environments may be significantly reduced by introducing a "conservative" form of noise; (c) if a decrease in performance is observed when the system is transferred to a real environment, successful and robust results can be obtained by continuing the evolutionary process in the real environment for a few generations.

Algorithms↗

A probabilistic Classifier System and its application in data mining.

The article is about a new Classifier System framework for classification tasks called BYP-CS (for BaYesian Predictive Classifier System). The proposed CS approach abandons the focus on high accuracy and addresses a well-posed Data Mining goal, namely, that of uncovering the low-uncertainty patterns of dependence that manifest often in the data. To attain this goal, BYP-CS uses a fair amount of probabilistic machinery, which brings its representation language closer to other related methods of interest in statistics and machine learning. On the practical side, the new algorithm is seen to yield stable learning of compact populations, and these still maintain a respectable amount of predictive power. Furthermore, the emerging rules self-organize in interesting ways, sometimes providing unexpected solutions to certain benchmark problems.

Algorithms↗

A maximum likelihood approach to density estimation with semidefinite programming.

Density estimation plays an important and fundamental role in pattern recognition, machine learning, and statistics. In this article, we develop a parametric approach to univariate (or low-dimensional) density estimation based on semidefinite programming (SDP). Our density model is expressed as the product of a nonnegative polynomial and a base density such as normal distribution, exponential distribution, and uniform distribution. When the base density is specified, the maximum likelihood estimation of the polynomial is formulated as a variant of SDP that is solved in polynomial time with the interior point methods. Since the base density typically contains just one or two parameters, computation of the maximum likelihood estimate reduces to a one- or two-dimensional easy optimization problem with this use of SDP. Thus, the rigorous maximum likelihood estimate can be computed in our approach. Furthermore, such conditions as symmetry and unimodality of the density function can be easily handled within this framework. AIC is used to choose the best model. Through applications to several instances, we demonstrate flexibility of the model and performance of the proposed procedure. Combination with a mixture approach is also presented. The proposed approach has possible other applications beyond density estimation. This point is clarified through an application to the maximum likelihood estimation of the intensity function of a nonstationary Poisson process.

Data Interpretation, Statistical↗

Second-order cone programming formulations for robust multiclass classification.

Multiclass classification is an important and ongoing research subject in machine learning. Current support vector methods for multiclass classification implicitly assume that the parameters in the optimization problems are known exactly. However, in practice, the parameters have perturbations since they are estimated from the training data, which are usually subject to measurement noise. In this article, we propose linear and nonlinear robust formulations for multiclass classification based on the M-SVM method. The preliminary numerical experiments confirm the robustness of the proposed method.

Artificial Intelligence↗

Reconstructing mental object representations: a machine vision approach to human visual recognition.

This paper introduces a new approach to assess visual representations underlying the recognition of objects. Human performance is modeled by CLARET, a machine learning and matching system, based on inductive logic programming and graph matching principles. The model is applied to data of a learning experiment addressing the role of prior experience in the ontogenesis of mental object representations. Prior experience was varied in terms of sensory modality, i.e. visual versus haptic versus visuohaptic. The analysis revealed distinct differences between the representational formats used by subjects with haptic versus those with no prior object experience. These differences suggest that prior haptic exploration stimulates the evolution of object representations which are characterized by an increased differentiation between attribute values and a pronounced structural encoding.

Computer Simulation↗

The IPRS Image Processing and Pattern Recognition System.

IPRS is a freely available software system which consists of about 250 library functions in C, and a set of application programs. It is designed to run under UNIX and comes with full source code, system manual pages, and a comprehensive user's and programmer's guide. It is intended for use by researchers in human vision, pattern recognition, image processing, machine vision and machine learning.

Humans↗

RAS signaling in lung adenocarcinoma is defined by lineage context and DUSP4 loss.

BACKGROUNDThe molecular landscape of lung adenocarcinoma (LUAD) is often illustrated as a driver-oncogene pie chart, but identical mutations exhibit heterogeneous signaling shaped by comutations, transcriptional programs, and lineage context. We propose a lineage-integrated signaling framework using an EGFR mutation signature (mSig).METHODSWe defined EGFR mSig using differentially expressed genes in EGFR-mutant (EGFR-mt) LUADs. Semisupervised clustering and machine learning models were used to test reproducibility in different combinations of datasets. We analyzed molecular subtypes, lineage markers, co-occurring mutations, and EGFR copy number alterations in EGFR mSig-defined subtypes of LUAD.RESULTSEGFR mSig showed robust classification performance (area under receiver operating characteristic curve = 0.83-0.95; mean negative predictive value = 96.3%). Validated gene expression subtypes and lung lineage markers were closely aligned with EGFR mSig status. Most EGFR mSig+ tumors, including many without EGFR mutations, belonged to the bronchioid subtype. A subset of canonical RAS mutations were mSig+ and mirrored the EGFR mutation pattern. EGFR WT/mSig- tumors were enriched for nonbronchioid subtypes and had comutations in TP53 or RAS/RAF/RTKs. We highlight a parsimonious collection of coordinated mutations, including RAS, KEAP1, STK11, TP53, and CDKN2A, that taken together suggest coordination of tumor signaling previously suggested but now reproduced and expanded.CONCLUSIONA potentially novel EGFR mSig that captures the transcriptional footprint of EGFR activation revealed a subset of EGFR WT LUADs with mt-like features. mSig refines LUAD taxonomy beyond mutation-only pie-chart models by incorporating lineage and comutation context. Lineage-directed stratification with coalteration identifies clinically relevant groups across EGFR and RAS states and highlights treatment opportunities for patients currently considered oncogene-negative.FUNDINGNational Cancer Institute (NCI) U01CA272541, R01CA262296, U24CA264021, UG1CA233333, R01CA211939.

Humans↗

A validated, modifiable proteomic score from the EXSCEL trial predicts cardiovascular events in diabetes.

BACKGROUNDAdults with type 2 diabetes mellitus (T2DM) are at increased risk for stroke, myocardial infarction, and cardiovascular death, yet individual risk is heterogeneous and incompletely captured by clinical models.METHODSIn the Exenatide Study of Cardiovascular Event Lowering (EXSCEL), adults with T2DM were randomized to a GLP-1 RA (exenatide) or a placebo and followed longitudinally for major adverse cardiovascular events (MACE). High-throughput discovery proteomics was done in plasma collected at baseline and 12 months. Proteins associated with time to MACE were identified using multivariable regression and incorporated into supervised machine learning models. A multi-protein score was developed and externally validated in 2 independent population-based and trial cohorts.RESULTSThe proteomic score showed incremental improvement in cardiovascular risk discrimination beyond clinical factors alone, and several proteins were consistently prioritized across modeling approaches. The protein score and a top-ranked protein, tetranectin, were modified by GLP-1 RA treatment, and a decrease in protein score was associated with improved outcomes, supporting modifiability of MACE risk.CONCLUSIONExternal validation confirmed generalizability across cohorts with and without diabetes. Together, these findings demonstrate that plasma proteomic signatures can enhance cardiovascular risk stratification and identify treatment-responsive biomarkers in T2DM, supporting their potential role in precision prevention strategiesFUNDINGThe EXSCEL study was funded by Amylin Pharmaceuticals. This research was supported by contracts HHSN268201200036C, HHSN268200800007C, HHSN268201800001C, N01HC55222, N01HC85079, N01HC85080, N01HC85081, N01HC85082, N01HC85083, N01HC85086, 75N92021D00006, and grants R01HL146145, U01HL080295, U01HL130114, R01HL172803, and R01HL144483 from the National Heart, Lung, and Blood Institute, with additional contribution from the National Institute of Neurological Disorders and Stroke. Additional support was provided by R01AG023629 from the National Institute on Aging.

Aged↗

Using the ID3 algorithm to find discrepant diagnoses from laboratory databases of thyroid patients.

Rare cases are a central problem when an expert system is constructed from example cases with machine learning techniques. It is difficult to make a decision support system (DSS) to cover all possible clinical cases. An inductive learning program can be used to construct an expert system for detecting cases that differ from routine cases. The ID3 algorithm and the pessimistic pruning algorithm were tested in this study: a DSS was built directly from the data of patient records. A decision tree was generated, and the cases misclassified by the decision tree as compared with the classifications of a clinician were listed on a checklist, which formed the feedback to the clinician. In clinical situations about 5-10% of functional thyroid disorders may be misclassified. At this error level, the method found over 90% of the errors with a specificity of 95%. In simple medical classification tasks this dynamic self-learning system can be used to create a DSS that can assist in the quality control of clinical decision making.

Adult↗

Knowledge discovery and data mining in toxicology.

Knowledge discovery and data mining tools are gaining increasing importance for the analysis of toxicological databases. This paper gives a survey of algorithms, capable to derive interpretable models from toxicological data, and presents the most important application areas. The majority of techniques in this area were derived from symbolic machine learning, one commercial product was developed especially for toxicological applications. The main application area is presently the detection of structure-activity relationships, very few authors have used these techniques to solve problems in epidemiological and clinical toxicology. Although the discussed algorithms are very flexible and powerful, further research is required to adopt the algorithms to the specific learning problems in this area, to develop improved representations of chemical and biological data and to enhance the interpretability of the derived models for toxicological experts.

Algorithms↗

In silico estimation of DMSO solubility of organic compounds for bioscreening.

Solubility of organic compounds in DMSO is an important issue for commercial and academic organizations handling large compound collections or performing biological screening. In particular, solubility data are critical for the optimization of storage conditions and for the selection of compounds for bioscreening compatible with the assay protocol. Solubility is largely determined by the solvation energy and the crystal disruption energy, and these molecular phenomena should be assessed in structure-solubility correlation studies. The authors summarize our long-term experimental observations and theoretical studies of physicochemical determinants of DMSO solubility of organic substances. They compiled a comprehensive reference database of proprietary data on compound solubility (55,277 compounds with good DMSO solubility and 10,223 compounds with poor DMSO solubility), calculated specific molecular descriptors (topological, electromagnetic, charge, and lipophilicity parameters), and applied an advanced machine-learning approach for training neural networks to address the solubility. Both supervised (feed-forward, back-propagated neural networks) and unsupervised (Kohonen neural networks) learning methods were used. The resulting neural network models were validated by successfully predicting DMSO solubility of compounds in independent test selections.

Dimethyl Sulfoxide↗

Biomarker discovery, disease classification, and similarity query processing on high-throughput MS/MS data of inborn errors of metabolism.

In newborn errors of metabolism, biomarkers are urgently needed for disease screening, diagnosis, and monitoring of therapeutic interventions. This article describes a 2-step approach to discover metabolic markers, which involves (1) the identification of marker candidates and (2) the prioritization of them based on expert knowledge of disease metabolism. For step 1, the authors developed a new algorithm, the biomarker identifier (BMI), to identify markers from quantified diseased versus normal tandem mass spectrometry data sets. BMI produces a ranked list of marker candidates and discards irrelevant metabolites based on a quality measure, taking into account the discriminatory performance, discriminatory space, and variance of metabolites' concentrations at the state of disease. To determine the ability of identified markers to classify subjects, the authors compared the discriminatory performance of several machine-learning paradigms and described a retrieval technique that searches and classifies abnormal metabolic profiles from a screening database. Seven inborn errors of metabolism-- phenylketonuria (PKU), glutaric acidemia type I (GA-I), 3-methylcrotonylglycinemia deficiency (3-MCCD), methylmalonic acidemia (MMA), propionic acidemia (PA), medium-chain acylCoAdehydrogenase deficiency (MCADD), and 3-OH long-chain acyl CoA dehydrogenase deficiency (LCHADD)-were investigated. All primarily prioritized marker candidates could be confirmed by literature. Some novel secondary candidates were identified (i.e., C16:1 and C4DC for PKU, C4DC for GA-I, and C18:1 forMCADD), which require further validation to confirm their biochemical role during health and disease.

Acyl-CoA Dehydrogenase↗

Computational models of oral and craniofacial development, growth, and repair.

This paper illustrates how biological and clinical problems stimulate research in biomedical informatics and how such research contributes to their solution. The computational models described use techniques from Logic Programming, Machine Learning, Computer Vision, and Biomathematics. They address problems in the development, growth, and repair of oral and craniofacial tissues arising in cell biology, clinical genetics, and dentistry. At the micro-level, the dynamic interaction of cells in the oral epithelium is modeled. At the macro-level, models are constructed of either the craniofacial shape of an individual or the craniofacial shape differences within and between healthy and congenitally abnormal populations. In between, in terms of scale, there are models of normal dentition and the use of computerized expert knowledge to guide the design of dental prostheses used to restore function in partially edentulous patients.

Adult↗

Unveiling the Molecular Secrets of Seaweeds: A Comprehensive Review of Bioinformatics Applications in Algal Research.

Recent advances in high-throughput sequencing, bioinformatics, and multi-omics technologies have transformed seaweed research by overcoming long-standing challenges associated with complex genomes, diverse life cycles, and limited genomic resources. This review provides a comprehensive overview of bioinformatics approaches used to investigate seaweed genomics, transcriptomics, proteomics, metabolomics, microbiomes, and functional genomics, with emphasis on the computational tools and databases that support these analyses. Applications of bioinformatics in phylogenetics, drug discovery, microbiome characterization, and the development of biofuels, nutraceuticals, pharmaceuticals, and sustainable agriculture are also discussed. Particular attention is given to emerging strategies involving multi-omics integration, genome editing, artificial intelligence, machine learning, and synthetic biology that are reshaping seaweed research. The review further examines current challenges, including incomplete genomic resources, data standardization, and the need for experimental validation of computational predictions. Collectively, these advances highlight the growing role of bioinformatics in enabling systems-level understanding of seaweed biology and accelerating their translation into sustainable biotechnological and marine bioeconomy applications.

macroalgal genomics↗