Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Categorization of sentence types in medical abstracts.

This study evaluated the use of machine learning techniques in the classification of sentence type. 7253 structured abstracts and 204 unstructured abstracts of Randomized Controlled Trials from MedLINE were parsed into sentences and each sentence was labeled as one of four types (Introduction, Method, Result, or Conclusion). Support Vector Machine (SVM) and Linear Classifier models were generated and evaluated on cross-validated data. Treating sentences as a simple "bag of words", the SVM model had an average ROC area of 0.92. Adding a feature of relative sentence location improved performance markedly for some models and overall increasing the average ROC to 0.95. Linear classifier performance was significantly worse than the SVM in all datasets. Using the SVM model trained on structured abstracts to predict unstructured abstracts yielded performance similar to that of models trained with unstructured abstracts in 3 of the 4 types. We conclude that classification of sentence type seems feasible within the domain of RCT's. Identification of sentence types may be helpful for providing context to end users or other text summarization techniques.

Abstracting and Indexing↗

Relevance vector machine for optical diagnosis of cancer.

BACKGROUND AND OBJECTIVES: A probability-based, robust diagnostic algorithm is an essential requirement for successful clinical use of optical spectroscopy for cancer diagnosis. This study reports the use of the theory of relevance vector machine (RVM), a recent Bayesian machine-learning framework of statistical pattern recognition, for development of a fully probabilistic algorithm for autofluorescence diagnosis of early stage cancer of human oral cavity. It also presents a comparative evaluation of the diagnostic efficacy of the RVM algorithm with that based on support vector machine (SVM) that has recently received considerable attention for this purpose. STUDY DESIGN/MATERIALS AND METHODS: The diagnostic algorithms were developed using in vivo autofluorescence spectral data acquired from human oral cavity with a N(2) laser-based portable fluorimeter. The spectral data of both patients as well as normal volunteers, enrolled at Out Patient department of the Govt. Cancer Hospital, Indore for screening of oral cavity, were used for this purpose. The patients selected had no prior confirmed malignancy and were diagnosed of squamous cell carcinoma (SCC), Grade-I on the basis of histopathology of biopsy taken from abnormal site subsequent to acquisition of spectra. Autofluorescence spectra were recorded from a total of 171 tissue sites from 16 patients and 154 healthy squamous tissue sites from 13 normal volunteers. Of 171 tissues sites from patients, 83 were SCC and the rest were contralateral uninvolved squamous tissue. Each site was treated separately and classified via the diagnostic algorithm developed. Instead of the spectral data from uninvolved sites of patients, the data from normal volunteers were used as the normal database for the development of diagnostic algorithms. RESULTS: The diagnostic algorithms based on RVM were found to provide classification performance comparable to the state-of-the-art SVMs, while at the same time explicitly predicting the probability of class membership. The sensitivity and specificity towards cancer were up to 88% and 95% for the training set data based on leave- one-out cross validation and up to 91% and 96% for the validation set data. When implemented on the spectral data of the uninvolved oral cavity sites from the patients, it yielded a specificity of up to 91%. CONCLUSIONS: The Bayesian framework of RVM formulation makes it possible to predict the posterior probability of class membership in discriminating early SCC from the normal squamous tissue sites of the oral cavity in contrast to dichotomous classification provided by the non-Bayesian SVM. Such classification is very helpful in handling asymmetric misclassification costs like assigning different weights for having a false negative result for identifying cancer compared to false positive. The results further demonstrate that for comparable diagnostic performances, the RVM-based algorithms use significantly fewer kernel functions and do not need to estimate any hoc parameters associated with the learning or the optimization technique to be used. This implies a considerable saving in memory and computation in a practical implementation.

Algorithms↗

Parallel man-machine training in development of EEG-based cursor control.

A new parallel man-machine training approach to brain-computer interface (BCI) succeeded through a unique application of machine learning methods. The BCI system could train users to control an animated cursor on the computer screen by voluntary electroencephalogram (EEG) modulation. Our BCI system requires only two to four electrodes, and has a relatively short training time for both the user and the machine. Moving the cursor in one dimension, our subjects were able to hit 100% of randomly selected targets, while in two dimensions, accuracies of approximately 63% and 76% was achieved with our two subjects.

Biofeedback, Psychology↗

Induction of rules for biological macromolecule crystallization.

X-ray crystallography is the method of choice for determining the 3-D structure of large macromolecules at a high enough resolution. The rate limiting step in structure determination is the crystallization itself. It takes anywhere between a few weeks to several years to obtain macromolecular crystals that yield good diffraction patterns. The theory of forces that promote and maintain crystal growth is preliminary, and crystallographers systematically search a large parameter space of experimental settings to grow good crystals. There is a wealth of experimental data on crystal growth most of which is in paper laboratory notebooks. Some of the data has been gathered in electronic form, e.g., the Biological Macromolecular Crystallization Database (BMCD) which is a repository of successful experimental conditions for growing over 800 different macromolecules (Gilliland 1987). Crystallographers are in need of computational tools to gather and analyze past data to design new crystal growth trails. We are building the Crystallographer's Assistant (CA) to help crystallographers record and maintain experimental context in electronic form, offer suggestions on experimental conditions that are likely to be successful, and provide explanations for failed experiments. As an initial step in this project, we have applied RL, an inductive learning program, to the BMCD. In this paper we report initial experiments and findings in applying RL to the BMCD. From the point of view of crystallography, we have discovered possibly significant new empirical relationships in crystal growth. From the point of view of machine learning, our work suggests refinements of existing methods for incorporating detailed domain knowledge into inductive analysis techniques.

Animals↗

Optimal classification of long echo time in vivo magnetic resonance spectra in the detection of recurrent brain tumors.

We describe the optimal high-level postprocessing of single-voxel (1)H magnetic resonance spectra and assess the benefits and limitations of automated methods as diagnostic aids in the detection of recurrent brain tumor. In a previous clinical study, 90 long-echo-time single-voxel spectra were obtained from 52 patients and classified during follow-up (30/28/32 normal/non-progressive tumor/tumor). Based on these data, a large number of evaluation strategies, including both standard resonance line quantification and algorithms from pattern recognition and machine learning, were compared in a quantitative evaluation. Results from linear and non-linear feature extraction, including ICA, PCA and wavelet transformations, and also the data from resonance line quantification were combined systematically with different classifiers such as LDA, chemometric methods (PLS, PCR), support vector machines and ensemble methods. Classification accuracy was assessed using a leave-one-out cross-validation scheme and the area under the curve (AUC) of the receiver operator characteristic (ROC). A regularized linear regression on spectra with binned channels reached 91% classification accuracy compared with 83% from quantification. Interpreting the loadings of these regressions, we find that lipid and lactate signals are too unreliable to be used in a simple machine rule. Choline and NAA are the main source of relevant information. Overall, we find that fully automated pattern recognition algorithms perform as well as, or slightly better than, a manually controlled and optimized resonance line quantification.

Algorithms↗

Automatic synthesis of synergies for control of reaching--hierarchical clustering.

In this paper we describe a novel method for determining synergies between joint motions in reaching movements by hierarchical clustering. A set of recorded elbow and shoulder trajectories is used in a learning algorithm to determine the relationships between angular velocities at elbow and shoulder joints. The learning algorithm is based on optimal criteria for obtaining the hierarchy of descriptions of movement trajectories. We show that this method finds complex synergism between optimal joint trajectories for a given set of data and angular velocities at the shoulder and elbow joints. Three other machine learning techniques (ML) are used for comparison with our method of hierarchical clustering of trajectories. These MLs are: (1) radial basis functions (RBF), (2) inductive learning (IL), and (3) adaptive-network-based fuzzy inference system (ANFIS). Better error characteristics were obtained using the method of hierarchical clustering in comparison with the other techniques. The advantage of the method of hierarchical clustering with respect to the other MLs is in integrating the spatial and temporal elements of reaching movements. Determination and analysis of spatio-temporal events of movement trajectories is a useful tool in designing control systems for functional electrical stimulation (FES) assisted manipulation.

Algorithms↗

Protein solubility: sequence based prediction and experimental verification.

MOTIVATION: Obtaining soluble proteins in sufficient concentrations is a recurring limiting factor in various experimental studies. Solubility is an individual trait of proteins which, under a given set of experimental conditions, is determined by their amino acid sequence. Accurate theoretical prediction of solubility from sequence is instrumental for setting priorities on targets in large-scale proteomics projects. RESULTS: We present a machine-learning approach called PROSO to assess the chance of a protein to be soluble upon heterologous expression in Escherichia coli based on its amino acid composition. The classification algorithm is organized as a two-layered structure in which the output of primary support vector machine (SVM) classifiers serves as input for a secondary Naive Bayes classifier. Experimental progress information from the TargetDB database as well as previously published datasets were used as the source of training data. In comparison with previously published methods our classification algorithm possesses improved discriminatory capacity characterized by the Matthews Correlation Coefficient (MCC) of 0.434 between predicted and known solubility states and the overall prediction accuracy of 72% (75 and 68% for positive and negative class, respectively). We also provide experimental verification of our predictions using solubility measurements for 31 mutational variants of two different proteins.

Amino Acid Sequence↗

N6-methyladenine identification using deep learning and discriminative feature integration.

N6-methyladenine (6 mA) is a pivotal DNA modification that plays a crucial role in epigenetic regulation, gene expression, and various biological processes. With advancements in sequencing technologies and computational biology, there is an increasing focus on developing accurate methods for 6 mA site identification to enhance early detection and understand its biological significance. Despite the rapid progress of machine learning in bioinformatics, accurately detecting 6 mA sites remains a challenge due to the limited generalizability and efficiency of existing approaches. In this study, we present Deep-N6mA, a novel Deep Neural Network (DNN) model incorporating optimal hybrid features for precise 6 mA site identification. The proposed framework captures complex patterns from DNA sequences through a comprehensive feature extraction process, leveraging k-mer, Dinucleotide-based Cross Covariance (DCC), Trinucleotide-based Auto Covariance (TAC), Pseudo Single Nucleotide Composition (PseSNC), Pseudo Dinucleotide Composition (PseDNC), and Pseudo Trinucleotide Composition (PseTNC). To optimize computational efficiency and eliminate irrelevant or noisy features, an unsupervised Principal Component Analysis (PCA) algorithm is employed, ensuring the selection of the most informative features. A multilayer DNN serves as the classification algorithm to identify N6-methyladenine sites accurately. The robustness and generalizability of Deep-N6mA were rigorously validated using fivefold cross-validation on two benchmark datasets. Experimental results reveal that Deep-N6mA achieves an average accuracy of 97.70% on the F. vesca dataset and 95.75% on the R. chinensis dataset, outperforming existing methods by 4.12% and 4.55%, respectively. These findings underscore the effectiveness of Deep-N6mA as a reliable tool for early 6 mA site detection, contributing to epigenetic research and advancing the field of computational biology.

Deep Learning↗

A new classifier based on information theoretic learning with unlabeled data.

Supervised learning is conventionally performed with pairwise input-output labeled data. After the training procedure, the adaptive system's weights are fixed while the testing procedure with unlabeled data is performed. Recently, in an attempt to improve classification performance unlabeled data has been exploited in the machine learning community. In this paper, we present an information theoretic learning (ITL) approach based on density divergence minimization to obtain an extended training algorithm using unlabeled data during the testing. The method uses a boosting-like algorithm with an ITL based cost function. Preliminary simulations suggest that the method has the potential to improve the performance of classifiers in the application phase.

Algorithms↗

Cerebellar circuitry as a neuronal machine.

Shortly after John Eccles completed his studies of synaptic inhibition in the spinal cord, for which he was awarded the 1963 Nobel Prize in physiology/medicine, he opened another chapter of neuroscience with his work on the cerebellum. From 1963 to 1967, Eccles and his colleagues in Canberra successfully dissected the complex neuronal circuitry in the cerebellar cortex. In the 1967 monograph, "The Cerebellum as a Neuronal Machine", he, in collaboration with Masao Ito and Janos Szentágothai, presented blue-print-like wiring diagrams of the cerebellar neuronal circuitry. These stimulated worldwide discussions and experimentation on the potential operational mechanisms of the circuitry and spurred theoreticians to develop relevant network models of the machinelike function of the cerebellum. In following decades, the neuronal machine concept of the cerebellum was strengthened by additional knowledge of the modular organization of its structure and memory mechanism, the latter in the form of synaptic plasticity, in particular, long-term depression. Moreover, several types of motor control were established as model systems representing learning mechanisms of the cerebellum. More recently, both the quantitative preciseness of cerebellar analyses and overall knowledge about the cerebellum have advanced considerably at the cellular and molecular levels of analysis. Cerebellar circuitry now includes Lugaro cells and unipolar brush cells as additional unique elements. Other new revelations include the operation of the complex glomerulus structure, intricate signal transduction for synaptic plasticity, silent synapses, irregularity of spike discharges, temporal fidelity of synaptic activation, rhythm generators, a Golgi cell clock circuit, and sensory or motor representation by mossy fibers and climbing fibers. Furthermore, it has become evident that the cerebellum has cognitive functions, and probably also emotion, as well as better-known motor and autonomic functions. Further cerebellar research is required for full understanding of the cerebellum as a broad learning machine for neural control of these functions.

Action Potentials↗

Breast Cancer Recurrence Status Assessment in 5 Years Using Multimodal Integrated Learning: A Feasibility Study.

Despite advances in breast cancer detection and treatment, recurrence after curative therapy continues to impact long-term survival and quality of life. Therefore, early identification of high-risk patients is crucial to guide personalized treatment and follow-up strategies. Although genomic assays provide valuable prognostic insights, their high cost and limited accessibility hinder widespread adoption in clinical practice. Recent machine learning or deep learning approaches leveraging clinical, imaging, or multimodal data have shown promise but do not reflect real-world clinical scenarios. This study proposes a deep learning-based multimodal framework for predicting 5-year breast cancer recurrence using routinely collected clinical data. The framework consists of three main components. First, we adopted automated tumor segmentation with MedSAM to extract the tumor region from ultrasound images. The radiomics features are extracted from those tumor regions. Second, report features are extracted using a Med-Contrastive Pre-trained Transformers (MedCPT)-based approach incorporating predefined, clinically informed queries. Third, a multimodal integration model jointly processes image, radiomics, clinical features, and report features through modality-specific branches. The image branch employs the Ultrasound Foundation Model (USFM) as the backbone, while structured tabular data is processed using the FT-Transformer architecture. The features of all branches are fused using a mixture-of-experts (MoE)-based classifier, and the entire model is trained using a progressive fusion training strategy. Experimental results confirm the feasibility of using ultrasound images with tumor mask integration for recurrence prediction and demonstrate the additive value of integrating multiple data modalities through the proposed multimodal integration model. The final model for recurrence prediction achieved an AUC of 0.7540, accuracy of 74.61%, sensitivity of 70.41%, and specificity of 76.44%. This feasibility study's findings underscore the potential of the proposed multimodal deep learning framework to provide accessible, accurate, and generalizable recurrence risk prediction using routinely available clinical data, potentially supporting more informed treatment decisions and personalized post-treatment monitoring in real-world clinical practice.

Breast cancer recurrence↗

Predicting genetic regulatory response using classification.

MOTIVATION: Studying gene regulatory mechanisms in simple model organisms through analysis of high-throughput genomic data has emerged as a central problem in computational biology. Most approaches in the literature have focused either on finding a few strong regulatory patterns or on learning descriptive models from training data. However, these approaches are not yet adequate for making accurate predictions about which genes will be up- or down-regulated in new or held-out experiments. By introducing a predictive methodology for this problem, we can use powerful tools from machine learning and assess the statistical significance of our predictions. RESULTS: We present a novel classification-based method for learning to predict gene regulatory response. Our approach is motivated by the hypothesis that in simple organisms such as Saccharomyces cerevisiae, we can learn a decision rule for predicting whether a gene is up- or down-regulated in a particular experiment based on (1) the presence of binding site subsequences ('motifs') in the gene's regulatory region and (2) the expression levels of regulators such as transcription factors in the experiment ('parents'). Thus, our learning task integrates two qualitatively different data sources: genome-wide cDNA microarray data across multiple perturbation and mutant experiments along with motif profile data from regulatory sequences. We convert the regression task of predicting real-valued gene expression measurements to a classification task of predicting +1 and -1 labels, corresponding to up- and down-regulation beyond the levels of biological and measurement noise in microarray measurements. The learning algorithm employed is boosting with a margin-based generalization of decision trees, alternating decision trees. This large-margin classifier is sufficiently flexible to allow complex logical functions, yet sufficiently simple to give insight into the combinatorial mechanisms of gene regulation. We observe encouraging prediction accuracy on experiments based on the Gasch S.cerevisiae dataset, and we show that we can accurately predict up- and down-regulation on held-out experiments. We also show how to extract significant regulators, motifs and motif-regulator pairs from the learned models for various stress responses. Our method thus provides predictive hypotheses, suggests biological experiments, and provides interpretable insight into the structure of genetic regulatory networks. AVAILABILITY: The MLJava package is available upon request to the authors. Supplementary: Additional results are available from http://www.cs.columbia.edu/compbio/geneclass

Binding Sites↗

Boosting naïve Bayesian learning on a large subset of MEDLINE.

We are concerned with the rating of new documents that appear in a large database (MEDLINE) and are candidates for inclusion in a small specialty database (REBASE). The requirement is to rank the new documents as nearly in order of decreasing potential to be added to the smaller database as possible, so as to improve the coverage of the smaller database without increasing the effort of those who manage this specialty database. To perform this ranking task we have considered several machine learning approaches based on the naï ve Bayesian algorithm. We find that adaptive boosting outperforms naï ve Bayes, but that a new form of boosting which we term staged Bayesian retrieval outperforms adaptive boosting. Staged Bayesian retrieval involves two stages of Bayesian retrieval and we further find that if the second stage is replaced by a support vector machine we again obtain a significant improvement over the strictly Bayesian approach.

Algorithms↗

A comparative study on feature selection methods for drug discovery.

Feature selection is frequently used as a preprocessing step to machine learning. The removal of irrelevant and redundant information often improves the performance of learning algorithms. This paper is a comparative study of feature selection in drug discovery. The focus is on aggressive dimensionality reduction. Five methods were evaluated, including information gain, mutual information, a chi2-test, odds ratio, and GSS coefficient. Two well-known classification algorithms, Naïve Bayesian and Support Vector Machine (SVM), were used to classify the chemical compounds. The results showed that Naïve Bayesian benefited significantly from the feature selection, while SVM performed better when all features were used. In this experiment, information gain and chi2-test were most effective feature selection methods. Using information gain with a Naïve Bayesian classifier, removal of up to 96% of the features yielded an improved classification accuracy measured by sensitivity. When information gain was used to select the features, SVM was much less sensitive to the reduction of feature space. The feature set size was reduced by 99%, while losing only a few percent in terms of sensitivity (from 58.7% to 52.5%) and specificity (from 98.4% to 97.2%). In contrast to information gain and chi2-test, mutual information had relatively poor performance due to its bias toward favoring rare features and its sensitivity to probability estimation errors.

Algorithms↗

Microarray expression profiling in melanoma reveals a BRAF mutation signature.

We have used microarray gene expression profiling and machine learning to predict the presence of BRAF mutations in a panel of 61 melanoma cell lines. The BRAF gene was found to be mutated in 42 samples (69%) and intragenic mutations of the NRAS gene were detected in seven samples (11%). No cell line carried mutations of both genes. Using support vector machines, we have built a classifier that differentiates between melanoma cell lines based on BRAF mutation status. As few as 83 genes are able to discriminate between BRAF mutant and BRAF wild-type samples with clear separation observed using hierarchical clustering. Multidimensional scaling was used to visualize the relationship between a BRAF mutation signature and that of a generalized mitogen-activated protein kinase (MAPK) activation (either BRAF or NRAS mutation) in the context of the discriminating gene list. We observed that samples carrying NRAS mutations lie somewhere between those with or without BRAF mutations. These observations suggest that there are gene-specific mutation signals in addition to a common MAPK activation that result from the pleiotropic effects of either BRAF or NRAS on other signaling pathways, leading to measurably different transcriptional changes.

Amino Acid Substitution↗

Predicting CNS permeability of drug molecules: comparison of neural network and support vector machine algorithms.

Two different machine-learning algorithms have been used to predict the blood-brain barrier permeability of different classes of molecules, to develop a method to predict the ability of drug compounds to penetrate the CNS. The first algorithm is based on a multilayer perceptron neural network and the second algorithm uses a support vector machine. Both algorithms are trained on an identical data set consisting of 179 CNS active molecules and 145 CNS inactive molecules. The training parameters include molecular weight, lipophilicity, hydrogen bonding, and other variables that govern the ability of a molecule to diffuse through a membrane. The results show that the support vector machine outperforms the neural network. Based on over 30 different validation sets, the SVM can predict up to 96% of the molecules correctly, averaging 81.5% over 30 test sets, which comprised of equal numbers of CNS positive and negative molecules. This is quite favorable when compared with the neural network's average performance of 75.7% with the same 30 test sets. The results of the SVM algorithm are very encouraging and suggest that a classification tool like this one will prove to be a valuable prediction approach.

Algorithms↗

A transcription factor regulatory atlas for activity inference and perturbation prediction.

Inferring transcription factor (TF) activity from transcriptomes and predicting transcriptome-wide responses to TF perturbations remain challenging, in part because available TF-mRNA resources often face a trade-off between precision and coverage and typically lack signed regulatory information. Here, we present TFActProfiler, a TF-mRNA resource and computational framework that learns signed, quantitative TF-mRNA regulatory coefficients by integrating heterogeneous prior evidence (ChIP-based, motif-based, and curated TF-mRNA annotations) with large-scale bulk and single-cell RNA-seq atlases. TFActProfiler contains 2 606 176 signed TF-mRNA interactions and improves TF activity inference in TF knockdown benchmarks relative to widely used regulon resources while retaining broad TF and target coverage. In addition, because the same learned regulatory coefficients can be used to model downstream transcriptional effects, TFActProfiler enables prediction of transcriptome-wide gene expression responses to TF knockdown without training on task-matched perturbation data. When perturbation datasets are available, TFActProfiler can be further refined to achieve performance comparable to state-of-the-art machine-learning baselines. By providing a direction-aware representation of TF-mRNA regulation for both activity inference and perturbation-response modeling, TFActProfiler supports systematic dissection of gene regulatory programs across diverse cellular contexts.

Transcription Factors↗

PostDOCK: a structural, empirical approach to scoring protein ligand complexes.

In this work we introduce a postprocessing filter (PostDOCK) that distinguishes true binding ligand-protein complexes from docking artifacts (that are created by DOCK 4.0.1). PostDOCK is a pattern recognition system that relies on (1) a database of complexes, (2) biochemical descriptors of those complexes, and (3) machine learning tools. We use the protein databank (PDB) as the structural database of complexes and create diverse training and validation sets from it based on the "families of structurally similar proteins" (FSSP) hierarchy. For the biochemical descriptors, we consider terms from the DOCK score, empirical scoring, and buried solvent accessible surface area. For the machine-learners, we use a random forest classifier and logistic regression. Our results were obtained on a test set of 44 structurally diverse protein targets. Our highest performing descriptor combinations obtained approximately 19-fold enrichment (39 of 44 binding complexes were correctly identified, while only allowing 2 of 44 decoy complexes), and our best overall accuracy was 92%.

Ligands↗