Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58Linked to original sources

New controller for functional electrical stimulation systems.

A novel, self-contained controller for functional electrical stimulation systems has been designed. The development was motivated by the need to have a general purpose, easy to use controller capable of stimulating many muscle groups, thus restoring complex motor functions (e.g. standing, walking, reaching, and grasping). The designed controller can regulate the frequency, pulse duration, and charge balance on up to 16 channels, and execute pre-programmed and sensory-driven control operations. The controller supports up to eight analog and six digital sensors, and comprises a memory block for including history of the sensory data (time series). Five independent timers provide the basis for the multi-modal and multi-level control of movement. The PC compatible interface is realised via an IR serial communication channel. The PC based software is user friendly and fully menu driven. This paper also presents a case study where the controller was implemented to restore walking in a paraplegic subject. The assistive system comprised the novel controller, the power and output stages of an eight-channel FES system (IEEE Trans Rehabil Eng, TRE-2 (1994) 234), ankle-foot orthoses, and a rolling walker. Stimulation was applied with surface electrodes positioned over the motoneurons that innervate muscles responsible for the hip and knee flexion and extension. The sensory system included goniometers at knee and hip joints, force-sensing resistors built in the shoe insoles, and digital accelerometers at the hips. A rule-based control algorithm was generated following a two-step procedure: (1) simulation and (2) machine learning as described in earlier studies (IEEE Trans Rehab Eng, TRE-7 (1999) 69). The paraplegic subject walked faster, and with less physiological effort, when automatic control was applied as compared to hand-control. This case study, as well as a previous one for assisting grasping (The design and testing of a new programmable electronic stimulator. N. Fisekovic, MS thesis. University of Belgrade, Belgrade, 2000) indicate that the novel control unit is effectively applicable to FES systems.

Adult↗

Feature mining and predictive model construction from severe trauma patient's data.

In management of severe trauma patients, trauma surgeons need to decide which patients are eligible for damage control. Such decision may be supported by utilizing models that predict the patient's outcome. The study described in this paper investigates the possibility to construct patient outcome prediction models from retrospective patient's data at the end of initial damage control surgery by using feature mining and machine learning techniques. As the data used comprises rather excessive number of features, special attention was paid to the problem of selecting only the most relevant features. We show that a small subset of features may carry enough information to construct reasonably accurate prognostic models. Furthermore, the techniques used in our study identified two factors, namely the pH value when admitted to ICU and the worst partial active thromboplastin time, to be of highest importance for prediction. This finding is pathophysiologically reasonable and represents two of three major problems with severe trauma patients, metabolic acidosis, hypothermia, and coagulopathy.

Algorithms↗

Information retrieval: an overview of system characteristics.

The paper gives an overview of characteristics of information retrieval (IR) systems. The characteristics are identified from the descriptions of 23 IR systems. Four IR models are discussed: the Boolean model, the vector model, the probabilistic model and the connectionistic model. Twelve other characteristics of IR models are identified: search intermediary, domain knowledge, relevance feedback, natural language interface, graphical query language, conceptual queries, full-text IR, field searching, fuzzy queries, hypertext integration, machine learning, and ranked output. Finally, the relevance of IR systems for the World Wide Web is established.

Algorithms↗

Logistic regression and artificial neural network classification models: a methodology review.

Logistic regression and artificial neural networks are the models of choice in many medical data classification tasks. In this review, we summarize the differences and similarities of these models from a technical point of view, and compare them with other machine learning algorithms. We provide considerations useful for critically assessing the quality of the models and the results based on these models. Finally, we summarize our findings on how quality criteria for logistic regression and artificial neural network models are met in a sample of papers from the medical literature.

Classification↗

Computer-assisted classification of HEp-2 immunofluorescence patterns in autoimmune diagnostics.

Indirect immunofluorescence with HEp-2 cells presents the major screening method for detection of autoantibodies in systemic autoimmune diseases. Hereby, a large variety of autoantibody entities can be detected and recognized by at least partially typic fluorescence patterns. Currently, this method requires highly specialized technicians and resists automatization. Nevertheless, requirements of good laboratory practice, especially standardization and documentation are hampered by the common microscopic technique. Here, we present a computer-assisted system for classification of interphase HEp-2 immunofluorescence patterns in autoimmune diagnostics. Designed as an assisting system, representative patterns are acquired by an operator with a digital microscope camera and transferred to a personal computer. By use of a novel software package based on image analysis, feature extraction and machine learning algorithms, relevant characteristics describing patterns could be found out. Our results show that identification of positive fluorescence and pre-differentiation between most important HEp-2 staining patterns can be performed by this system. Results and documentation of fluorescence patterns can be integrated into the laboratory system. To enable the usage of such a system in routine diagnostics, accuracy of this system and correct recognition of interferring patterns must be further improved.

Algorithms↗

Closed-loop, multiobjective optimization of analytical instrumentation: gas chromatography/time-of-flight mass spectrometry of the metabolomes of human serum and of yeast fermentations.

The number of instrumental parameters controlling modern analytical apparatus can be substantial, and varying them systematically to optimize a particular chromatographic separation, for example, is out of the question because of the astronomical number of combinations that are possible (i.e., the "search space" is very large). However, heuristic methods, such as those based on evolutionary computing, can be used to explore such search spaces efficiently. We here describe the implementation of an entirely automated (closed-loop) strategy for doing this and apply it to the optimization of gas chromatographic separations of the metabolomes of human serum and of yeast fermentation broths. Without human intervention, the Robot Chromatographer system (i) initializes the settings on the instrument, (ii) controls the analytical run, (iii) extracts the variables defining the analytical performance (specifically the number of peaks, signal/noise ratio, and run time), (iv) chooses (via the PESA-II multiobjective genetic algorithm), and (v) programs the next series of instrumental settings, the whole continuing in an iterative cycle until suitable sets of optimal conditions have been established. Genetic programming was used to remove noise peaks and to establish the basis for the improvements observed. The system showed that the number of peaks observable depended enormously on the conditions used and served to increase them by as much as 3-fold (e.g., to over 950 in human serum) while in many cases maintaining or reducing the run time and preserving excellent signal/noise ratios. The evolutionary closed-loop machine learning strategy we describe is generic to any type of analytical optimization.

Fermentation↗

Sample handling for mass spectrometric proteomic investigations of human sera.

Proteomic investigations of sera are potentially of value for diagnosis, prognosis, choice of therapy, and disease activity assessment by virtue of discovering new biomarkers and biomarker patterns. Much debate focuses on the biological relevance and the need for identification of such biomarkers while less effort has been invested in devising standard procedures for sample preparation and storage in relation to model building based on complex sets of mass spectrometric (MS) data. Thus, development of standardized methods for collection and storage of patient samples together with standards for transportation and handling of samples are needed. This requires knowledge about how sample processing affects MS-based proteome analyses and thereby how nonbiological biased classification errors are avoided. In this study, we characterize the effects of sample handling, including clotting conditions, storage temperature, storage time, and freeze/thaw cycles, on MS-based proteomics of human serum by using principal components analysis, support vector machine learning, and clustering methods based on genetic algorithms as class modeling and prediction methods. Using spiking to artificially create differentiable sample groups, this integrated approach yields data that--even when working with sample groups that differ more than may be expected in biological studies--clearly demonstrate the need for comparable sampling conditions for samples used for modeling and for the samples that are going into the test set group. Also, the study emphasizes the difference between class prediction and class comparison studies as well as the advantages and disadvantages of different modeling methods.

Humans↗

Controlled protein precipitation in combination with chip-based nanospray infusion mass spectrometry. An approach for metabolomics profiling of plasma.

Liquid chromatography-mass spectrometry (LC-MS) is a common method for profiling biological samples in metabolomics. However, LC-MS data of metabolomic studies are often affected by high noise levels, retention time shifts, and high variability in signal intensities. With a new chip-based nanoelectrospray source it becomes possible to directly infuse complex biological samples such as plasma without any chromatographic separation beforehand. In combination with highly diluted samples and long data acquisition times, the parallel analysis of hundreds of compounds is now possible. In a proof-of-concept study, 10 human plasma samples from females and males were analyzed with the intention to separate the two groups by their different metabolomes. The reproducibility was so high that statistical analysis of the data could be performed without prior normalization. Two groups of female and male samples were separated by a supervised machine learning algorithm, principal component analysis, and hierarchical clustering. Peaks contributing to the group separation were characterized by accurate mass measurement and MS-MS fragmentation and by spiking experiments. The feasibility of direct sample infusion using the new chip-based nanoelectrospray device opens a new dimension for the rapid parallel analysis of complex biological mixtures.

Adult↗

Ecological Restoration of the Soil-Like Function in the Bauxite Residue: Natural Microbiomes Mediated Molecular Transformation of Dissolved Organic Matter.

Soilization of bauxite residues offers a scalable route for long-term carbon management and ecological restoration. However, the microbial processes that transform exogenous organic inputs into stable soil-like carbon pools remain poorly resolved. Here, we combined cross-ecosystem meta-analysis, machine-learning prediction, native synthetic community (SynCom) construction, 13C-labeled straw microcosms, field validation, Fourier transform ion cyclotron resonance mass spectrometry, and genome-resolved metagenomics to unravel microbiome-mediated carbon transformation at the dissolved organic matter (DOM) molecular scale. Our meta-analysis revealed that alkaline industrial wastes retained soil-like DOM signatures but were enriched in microbial humic- and protein-like components, indicating active yet incomplete carbon processing. Guided by these patterns, native SynCom inoculation increased 13C incorporation into total organic carbon (TOC) and dissolved organic carbon (DOC), enlarged biodegradable and adsorbable DOC fractions, and shifted DOM from recalcitrant aromatic pools toward oxygenated carbohydrate-, tannin-, and phenolic-like molecular classes. Genome-resolved analyses linked this transformation to complementary polymer degradation and nutrient-cycling functions across fungal and bacterial guilds, including enriched carbohydrate-active enzymes in straw-carbon-utilizing metagenome-assembled genomes. Null model and thermodynamic analyses further showed that microbial communities were constrained by homogeneous selection, whereas DOM molecules were diversified through variable selection and redox-dependent transformation. Field-scale validation confirmed that SynCom promoted TOC and DOC accumulation and humic-like, high-density DOM fractions under alkaline conditions. Together, these findings establish a mechanistic framework in which functional microbiomes couple plant carbon depolymerization, DOM molecular diversification, and mineral-interactive carbon stabilization, providing a microbiome-guided strategy for carbon sequestration and soilization in the bauxite residue.

Soil↗

Accelerate Your Science: Direct-to-Biology Strategies in Medicinal Chemistry.

Direct-to-biology (D2B) is a powerful strategy that accelerates early drug discovery. It enables compounds to be synthesized in miniaturized formats and evaluated directly as crude reaction mixtures. This bypasses the need for purification during the initial design-make-test cycle. Advances in robust synthetic methodologies, automation, reaction miniaturization, and biological screening have transformed D2B from a proof-of-concept approach into a versatile medicinal chemistry platform. This platform is applicable to fragment optimization, covalent ligands, macrocycles, proteolysis-targeting chimeras (PROTACs), molecular glues, and cellular phenotypic screening. This perspective focuses on the synthetic transformations, assay technologies, and platform implementations that drive modern D2B workflows. It emphasizes reaction robustness, assay compatibility, and practical implementation. Analysis of the current literature revealed that D2B is more governed by reaction reliability than synthetic diversity. Amide coupling and click chemistry dominate reported workflows, while more complex transformations remain underexplored. We discuss the complementary strengths and limitations of biochemical, biophysical, and cellular readouts, identify current bottlenecks in reaction scope and data management, and highlight emerging opportunities arising from reaction miniaturization, machine learning, automated experimentation, and advanced synthetic methodologies. Rather than replacing conventional medicinal chemistry, D2B fundamentally shifts experimental effort from purification toward early biological validation and is poised to become an integral component of future medicinal chemistry workflows.

Humans↗

Immunoproteomic Profiling of Autoantibodies and Antibodies against Infectious Agents in Autoimmune Diseases.

Prior research investigated limited antibody sets within individual autoimmune diseases. Using the Nucleic-Acid Programmable Protein Array platform, we measured antibodies against 280 human, 40 viral, and 15 bacterial antigens in serum from 237 patients with 8 autoimmune diseases, including autoimmune gastritis (AG), autoimmune thyroiditis (AT), celiac disease (CD), idiopathic inflammatory myopathies (IIM), type 1 diabetes mellitus (T1D), rheumatoid arthritis (RA), Sjögren's disease (SjD), and systemic lupus erythematosus (SLE), and 112 controls. Candidate antibodies were identified by combining Firth logistic regression and machine learning. We identified disease-specific antibodies, ranging from 3 in IIM to 13 in SLE for IgG and 1 in CD to 13 in AG for IgA. Additionally, 63 IgG and 44 IgA antibodies were shared across two or more diseases. Notably, two IgG autoantibodies overlapped in up to five diseases: directed against STNM4 (SLE, SjD, T1D, CD, and RA) and TRIM21 (SLE, SjD, IIM, CD, and RA); and three IgA antibodies in up to seven diseases: directed against H1N1 Influenza A virus NP (IIM, SjD, AG, T1D, CD, RA, and AT) and Coxsackievirus B3MK012537 and Enterovirus C PVgp1 (IIM, SjD, AG, T1D, CD, SLE, and RA). These findings underscore the potential of antibody profiling in autoimmune disease characterization and biomarker discovery.

Humans↗

Support vector machines for predictive modeling in heterogeneous catalysis: a comprehensive introduction and overfitting investigation based on two real applications.

This works provides an introduction to support vector machines (SVMs) for predictive modeling in heterogeneous catalysis, describing step by step the methodology with a highlighting of the points which make such technique an attractive approach. We first investigate linear SVMs, working in detail through a simple example based on experimental data derived from a study aiming at optimizing olefin epoxidation catalysts applying high-throughput experimentation. This case study has been chosen to underline SVM features in a visual manner because of the few catalytic variables investigated. It is shown how SVMs transform original data into another representation space of higher dimensionality. The concepts of Vapnik-Chervonenkis dimension and structural risk minimization are introduced. The SVM methodology is evaluated with a second catalytic application, that is, light paraffin isomerization. Finally, we discuss why SVMs is a strategic method, as compared to other machine learning techniques, such as neural networks or induction trees, and why emphasis is put on the problem of overfitting.

Alkenes↗

Application of pattern recognition techniques to mass spectrometric data for sequencing C-terminal peptide residue series.

The application of pattern recognition to sequence elucidation from fast atom bombardment (FAB), collisionally activated dissociation (CAD), and tandem mass spectrometric data of peptides was investigated. Learning machine techniques for pattern recognition were applied to detect C-terminal series amino acid sequences up to the pentapeptide Try-Gly-Gly-Phe-Leu (YGGFL). The approach conditions the data by building upon known fragmentation pathways of peptides in FAB/CAD-related analysis. The intensities of critical sequence ion peaks are used to describe each pattern in the training set. A well-defined training set is then made use of to classify unknown species. The FORTRAN-77 program is adapted from one first developed by P.C. Jurs. It requires only the input of the critical peak intensities and the length of the unit to be tested. The method has potential in applications to larger peptides provided a database for the training set can be constructed.

Amino Acid Sequence↗

Analysis of calcium, oxalate, and citrate interaction in idiopathic calcium urolithiasis in children.

The majority of urinary stones in children are composed of calcium oxalate. To investigate the interaction between urinary calcium, oxalate, and citrate as major risk factors for calcium stones formation, their 24-h urinary excretion was determined in 30 children with urolithiasis and 15 normal healthy children. The cutoff points between children with urolithiasis and healthy children, accuracy, sensitivity, and specificity for each risk factor alone as well as for all three taken together were determined. OneR and J4.8 classifiers as parts of the larger data mining software Weka, based on machine learning algorithms, were used for the determination of the cutoff points for differentiation of the children. The decision tree based on J4.8 classifier analysis of all three risk factors together proved to be the best for differentiating stone formers from normal children. In comparison to the accuracy of the differentiation after calcium and oxalate of 80% and 75.6%, respectively, the decision tree showed an accuracy of 97.8%. Even when its stability was tested by the leave-one-out cross-validation procedure, the accuracy remained at a very acceptable percentage of 93.2% correctly classified patients. J4.8 classifier analysis gave a look inside urinary calcium, oxalate, and citrate interaction. Urinary calcium excretion was shown as the most informative in discrimination of the children with urolithiasis from healthy children. However, it was shown that oxalate and citrate excretions might influence the stone formation in a subpopulation of the stone formers. In patients with low urinary calcium, a major role in lithogenesis belongs to oxalate, in some of them alone and in others in conjunction with citrate. Decreased urinary citrate excretion in the presence of increased oxalate excretion may lead to stone formation.

Adolescent↗

Electronic van der Waals surface property descriptors and genetic algorithms for developing structure-activity correlations in olfactory databases.

A methodology to facilitate the intelligent design of new odorants (e.g., musks) with specialized properties has been developed as part of an ongoing research effort in machine learning. In a traditional framework, the introduction of a new odorant is a lengthy, costly, and laborious discovery, development, and testing process. We propose to streamline this process utilizing large existing olfactory databases available through the open scientific literature as input for a new structure/activity correlation methodology. The first step in this process is to characterize each molecule in the database by an appropriate set of descriptors. To accomplish this task, an enhanced version of Breneman's Transferable Atom Equivalent (TAE) descriptor methodology will be used to create a large set of electron density derived shape/property hybrid (PEST), wavelet coefficient (WCD), and TAE histogram descriptors. We have chosen these molecular property descriptors to represent the problem because they have been shown to contain pertinent shape and electronic properties of the molecule and correlate with key modes of intermolecular interactions. Traditional QSAR methodologies, which employ fragment based descriptors, have been shown to be effective for QSAR development within homologous sets of molecules but are less effective when applied to data sets containing a great deal of structural variation. In contrast to previous attempts at SAR, our use of shape-aware electron density based molecular property descriptors has removed many of the limitations brought about by the use of descriptors based on substructure fragments, molecular surface properties, or other whole molecule descriptors. Another reason for the mixed success of past QSAR efforts can be traced to the nature of the underlying modeling problem, which is often quite complex. To meet these challenges, a genetic algorithm for pattern recognition analysis has been developed that selects descriptors which create class separation in a plot of the two largest principal components of the data while simultaneously searching for features that increase clustering of the data.

Journal Article↗

Data mining the NCI cancer cell line compound GI(50) values: identifying quinone subtypes effective against melanoma and leukemia cell classes.

Using data mining techniques, we have studied a subset (1400) of compounds from the large public National Cancer Institute (NCI) compounds data repository. We first carried out a functional class identity assignment for the 60 NCI cancer testing cell lines via hierarchical clustering of gene expression data. Comprised of nine clinical tissue types, the 60 cell lines were placed into six classes-melanoma, leukemia, renal, lung, and colorectal, and the sixth class was comprised of mixed tissue cell lines not found in any of the other five classes. We then carried out supervised machine learning, using the GI(50) values tested on a panel of 60 NCI cancer cell lines. For separate 3-class and 2-class problem clustering, we successfully carried out clear cell line class separation at high stringency, p < 0.01 (Bonferroni corrected t-statistic), using feature reduction clustering algorithms embedded in RadViz, an integrated high dimensional analytic and visualization tool. We started with the 1400 compound GI(50) values as input and selected only those compounds, or features, significant in carrying out the classification. With this approach, we identified two small sets of compounds that were most effective in carrying out complete class separation of the melanoma, non-melanoma classes and leukemia, non-leukemia classes. To validate these results, we showed that these two compound sets' GI(50) values were highly accurate classifiers using five standard analytical algorithms. One compound set was most effective against the melanoma class cell lines (14 compounds), and the other set was most effective against the leukemia class cell lines (30 compounds). The two compound classes were both significantly enriched in two different types of substituted p-quinones. The melanoma cell line class of 14 compounds was comprised of 11 compounds that were internal substituted p-quinones, and the leukemia cell line class of 30 compounds was comprised of 6 compounds that were external substituted p-quinones. Attempts to subclassify melanoma or leukemia cell lines based upon their clinical cancer subtype met with limited success. For example, using GI(50) values for the 30 compounds we identified as effective against all leukemia cell lines, we could subclassify acute lymphoblastic leukemia (ALL) origin cell lines from non-ALL leukemia origin cell lines without significant overlap from non-leukemia cell lines. Based upon clustering using GI(50) values for the 60 cancer cell lines laid out by the RadViz algorithm, these two compound subsets did not overlap with clusters containing any of the NCI's 92 compounds of known mechanism of action, a few of which are quinones. Given their structural patterns, the two p-quinone subtypes we identified would clearly be expected to possess different redox potentials/substrate specificities for enzymatic reduction in vivo. These two p-quinone subtypes represent valuable information that may be used in the elucidation of pharmacophores for the design of compounds to treat these two cancer tissue types in the clinic.

Algorithms↗

Drug discovery using support vector machines. The case studies of drug-likeness, agrochemical-likeness, and enzyme inhibition predictions.

Support Vector Machines (SVM) is a powerful classification and regression tool that is becoming increasingly popular in various machine learning applications. We tested the ability of SVM, in comparison with well-known neural network techniques, to predict drug-likeness and agrochemical-likeness for large compound collections. For both kinds of data, SVM outperforms various neural networks using the same set of descriptors. We also used SVM for estimating the activity of Carbonic Anhydrase II (CA II) enzyme inhibitors and found that the prediction quality of our SVM model is better than that reported earlier for conventional QSAR. Model characteristics and data set features were studied in detail.

Agrochemicals↗

Nonlinear prediction of quantitative structure-activity relationships.

Predicting the log of the partition coefficient P is a long-standing benchmark problem in Quantitative Structure-Activity Relationships (QSAR). In this paper we show that a relatively simple molecular representation (using 14 variables) can be combined with leading edge machine learning algorithms to predict logP on new compounds more accurately than existing benchmark algorithms which use complex molecular representations.

Journal Article↗