Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

Computation of conformational entropy from protein sequences using the machine-learning method--application to the study of the relationship between structural conservation and local structural stability.

A complete protein sequence can usually determine a unique conformation; however, the situation is different for shorter subsequences--some of them are able to adopt unique conformations, independent of context; while others assume diverse conformations in different contexts. The conformations of subsequences are determined by the interplay between local and nonlocal interactions. A quantitative measure of such structural conservation or variability will be useful in the understanding of the sequence-structure relationship. In this report, we developed an approach using the support vector machine method to compute the conformational variability directly from sequences, which is referred to as the sequence structural entropy. As a practical application, we studied the relationship between sequence structural entropy and the hydrogen exchange for a set of well-studied proteins. We found that the slowest exchange cores usually comprise amino acids of the lowest sequence structural entropy. Our results indicate that structural conservation is closely related to the local structural stability. This relationship may have interesting implications in the protein folding processes, and may be useful in the study of the sequence-structure relationship.

Amino Acid Sequence↗

Binary halftone image resolution increasing by decision tree learning.

This paper presents a new, accurate, and efficient technique to increase the spatial resolution of binary halftone images. It makes use of a machine learning process to automatically design a zoom operator starting from pairs of input-output sample images. To accurately zoom a halftone image, a large window and large sample images are required. Unfortunately, in this case, the execution time required by most of the previous techniques may be prohibitive. The new solution overcomes this difficulty by using decision tree (DT) learning. Original DT learning is modified to obtain a more efficient technique (WZDT learning). It is useful to know, a priori, sample complexity (the number of training samples needed to obtain, with probability 1 - delta, an operator with accuracy epsilon): we use the probably approximately correct (PAC) learning theory to compute the sample complexity. Since the PAC theory usually yields an overestimated sample complexity, statistical estimation is used to evaluate, a posteriori, a tight error bound. Statistical estimation is also used to choose an appropriate window and to show that DT learning has good inductive bias. The new technique is more accurate than a zooming method based on simple inverse halftoning techniques. The quality of the proposed solution is very close to the theoretical optimal obtainable quality for a neighborhood-based zooming process using the Hamming distance to quantify the error.

Algorithms↗

Selective inhibition of adaptor complex-mediated vesiculation.

A short while ago, we could only inhibit post-Golgi membrane traffic with crude, unselective tools, such as low temperature or high extracellular sucrose. Molecular dissection of vesiculation steps has revealed unexpected complexity in the coating machinery that has initiated a search for more specific inhibitors. We have learned that membrane vesiculation is driven by a tightly regulated multicomponent, membrane-associated protein machine held together by carefully specified interaction domains. An experimental advantage of such complex interacting machinery is that it is very susceptible to disruption by dominant negative inhibitors or by overexpression. As a result, we now have much more specific inhibitors of post-Golgi membrane traffic. Some, such as dynamin K44A, may be general inhibitors, whereas others can distinguish classes of endocytotic events (10), binding events that require clathrin from those that do not (42), or specific steps of endocytosis (62). Ligand-mediated uptake of EGF and numerous, but not all, GPCRs can be inhibited by overexpression of an ARF GTPase-activating protein that has no effect on transferrin uptake (67). We can look forward to increasingly powerful and selective inhibitors that should help us to navigate successfully the complex routes of post-Golgi membrane traffic.

Adaptor Protein Complex alpha Subunits↗

Application of metabolomics to plant genotype discrimination using statistics and machine learning.

MOTIVATION: Metabolomics is a post genomic technology which seeks to provide a comprehensive profile of all the metabolites present in a biological sample. This complements the mRNA profiles provided by microarrays, and the protein profiles provided by proteomics. To test the power of metabolome analysis we selected the problem of discrimating between related genotypes of Arabidopsis. Specifically, the problem tackled was to discrimate between two background genotypes (Col0 and C24) and, more significantly, the offspring produced by the crossbreeding of these two lines, the progeny (whose genotypes would differ only in their maternally inherited mitichondia and chloroplasts). OVERVIEW: A gas chromotography--mass spectrometry (GCMS) profiling protocol was used to identify 433 metabolites in the samples. The metabolomic profiles were compared using descriptive statistics which indicated that key primary metabolites vary more than other metabolites. We then applied neural networks to discriminate between the genotypes. This showed clearly that the two background lines can be discrimated between each other and their progeny, and indicated that the two progeny lines can also be discriminated. We applied Euclidean hierarchical and Principal Component Analysis (PCA) to help understand the basis of genotype discrimination. PCA indicated that malic acid and citrate are the two most important metabolites for discriminating between the background lines, and glucose and fructose are two most important metabolites for discriminating between the crosses. These results are consistant with genotype differences in mitochondia and chloroplasts.

Algorithms↗

Regularized Least Squares Cancer classifiers from DNA microarray data.

BACKGROUND: The advent of the technology of DNA microarrays constitutes an epochal change in the classification and discovery of different types of cancer because the information provided by DNA microarrays allows an approach to the problem of cancer analysis from a quantitative rather than qualitative point of view. Cancer classification requires well founded mathematical methods which are able to predict the status of new specimens with high significance levels starting from a limited number of data. In this paper we assess the performances of Regularized Least Squares (RLS) classifiers, originally proposed in regularization theory, by comparing them with Support Vector Machines (SVM), the state-of-the-art supervised learning technique for cancer classification by DNA microarray data. The performances of both approaches have been also investigated with respect to the number of selected genes and different gene selection strategies. RESULTS: We show that RLS classifiers have performances comparable to those of SVM classifiers as the Leave-One-Out (LOO) error evaluated on three different data sets shows. The main advantage of RLS machines is that for solving a classification problem they use a linear system of order equal to either the number of features or the number of training examples. Moreover, RLS machines allow to get an exact measure of the LOO error with just one training. CONCLUSION: RLS classifiers are a valuable alternative to SVM classifiers for the problem of cancer classification by gene expression data, due to their simplicity and low computational complexity. Moreover, RLS classifiers show generalization ability comparable to the ones of SVM classifiers also in the case the classification of new specimens involves very few gene expression levels.

Algorithms↗

Escalation of aggression: experimental studies.

A finding commonly obtained in research using the Buss "aggression machine" is a main effect for trail blocks, indicating an escalation in shock intensity over trails. Theoretical explanations for this effect were tested in a modified verbal operant-conditioning situation. In Experiment 1, subjects could administer any of 10 levels of positive reinforcement to a "learner" for correct verbal responses or any of 10 levels of negative reinforcement to a learner for incorrect responses. Half of the subjects were required to begin with weak, half with strong, reinforcements. Results indicated that, regardless of condition, subjects gave more intense reinforcements as the learning trails progressed. Those who administered negative reinforcements devalued the learner relative to those who administered positive reinforcements. In Experiment 2, a role-playing procedure was used in which subjects administered either positive or negative reinforcements to a learner whose performance either did or did not improve over trials. Again, in all experimental groups, subjects administered increasingly intense reinforcements over trials. The results are interpreted as supporting a disinhibition theory of anti- and prosocial behavior.

Aggression↗

Fairness-aware supervised hierarchical contrastive semantic learning for sexual dimorphism analysis.

MOTIVATION: Sexual dimorphism is a fundamental biological determinant driving systematic differences in disease susceptibility, progression, and clinical outcomes. However, current sex-combined AI-based genomic models often exhibit algorithmic bias and fail to capture these sex-specific mechanisms, creating a critical barrier to unbiased precision medicine. Ensuring fairness in the context of sexual dimorphism requires understanding and addressing the distinct biological mechanisms functioning in each sex, rather than focusing solely on equalizing predictive performance. RESULTS: We propose a fairness-aware supervised hierarchical contrastive learning approach, called FairHICON, to discover unbiased sex-common and sex-specific predictive features. Evaluations on cancer and asthma transcriptomic datasets demonstrate that FairHICON significantly outperforms state-of-the-art benchmarks, improving predictive performance by up to 9% while effectively reducing the performance gap between male and female sexes. Furthermore, prognostic validation confirms that the identified sex-specific pathways stratify patient survival significantly better within their corresponding sex groups. This validates FairHICON to elucidate the molecular heterogeneity of sexual dimorphism, advancing inclusive precision medicine. AVAILABILITY AND IMPLEMENTATION: The source code and data is available at https://github.com/datax-lab/FairHICON.

Sex Characteristics↗

Physicochemical descriptors to discriminate protein-protein interactions in permanent and transient complexes selected by means of machine learning algorithms.

Analyzing protein-protein interactions at the atomic level is critical for our understanding of the principles governing the interactions involved in protein-protein recognition. For this purpose, descriptors explaining the nature of different protein-protein complexes are desirable. In this work, the authors introduced Epic Protein Interface Classification as a framework handling the preparation, processing, and analysis of protein-protein complexes for classification with machine learning algorithms. We applied four different machine learning algorithms: Support Vector Machines, C4.5 Decision Trees, K Nearest Neighbors, and Naïve Bayes algorithm in combination with three feature selection methods, Filter (Relief F), Wrapper, and Genetic Algorithms, to extract discriminating features from the protein-protein complexes. To compare protein-protein complexes to each other, the authors represented the physicochemical characteristics of their interfaces in four different ways, using two different atomic contact vectors, DrugScore pair potential vectors and SFCscore descriptor vectors. We classified two different datasets: (A) 172 protein-protein complexes comprising 96 monomers, forming contacts enforced by the crystallographic packing environment (crystal contacts), and 76 biologically functional homodimer complexes; (B) 345 protein-protein complexes containing 147 permanent complexes and 198 transient complexes. We were able to classify up to 94.8% of the packing enforced/functional and up to 93.6% of the permanent/transient complexes correctly. Furthermore, we were able to extract relevant features from the different protein-protein complexes and introduce an approach for scoring the importance of the extracted features.

Algorithms↗

Predicting deleterious nsSNPs: an analysis of sequence and structural attributes.

BACKGROUND: There has been an explosion in the number of single nucleotide polymorphisms (SNPs) within public databases. In this study we focused on non-synonymous protein coding single nucleotide polymorphisms (nsSNPs), some associated with disease and others which are thought to be neutral. We describe the distribution of both types of nsSNPs using structural and sequence based features and assess the relative value of these attributes as predictors of function using machine learning methods. We also address the common problem of balance within machine learning methods and show the effect of imbalance on nsSNP function prediction. We show that nsSNP function prediction can be significantly improved by 100% undersampling of the majority class. The learnt rules were then applied to make predictions of function on all nsSNPs within Ensembl. RESULTS: The measure of prediction success is greatly affected by the level of imbalance in the training dataset. We found the balanced dataset that included all attributes produced the best prediction. The performance as measured by the Matthews correlation coefficient (MCC) varied between 0.49 and 0.25 depending on the imbalance. As previously observed, the degree of sequence conservation at the nsSNP position is the single most useful attribute. In addition to conservation, structural predictions made using a balanced dataset can be of value. CONCLUSION: The predictions for all nsSNPs within Ensembl, based on a balanced dataset using all attributes, are available as a DAS annotation. Instructions for adding the track to Ensembl are at http://www.brightstudy.ac.uk/das_help.html.

Algorithms↗

An in-silico method for prediction of polyadenylation signals in human sequences.

This paper presents a machine learning method to predict polyadenylation signals (PASes) in human DNA and mRNA sequences by analysing features around them. This method consists of three sequential steps of feature manipulation: generation, selection and integration of features. In the first step, new features are generated using k-gram nucleotide acid or amino acid patterns. In the second step, a number of important features are selected by an entropy-based algorithm. In the third step, support vector machines are employed to recognize true PASes from a large number of candidates. Our study shows that true PASes in DNA and mRNA sequences can be characterized by different features, and also shows that both upstream and downstream sequence elements are important for recognizing PASes from DNA sequences. We tested our method on several public data sets as well as our own extracted data sets. In most cases, we achieved better validation results than those reported previously on the same data sets. The important motifs observed are highly consistent with those reported in literature.

Base Sequence↗

Feature selection in MLPs and SVMs based on maximum output information.

This paper presents feature selection algorithms for multilayer perceptrons (MLPs) and multiclass support vector machines (SVMs), using mutual information between class labels and classifier outputs, as an objective function. This objective function involves inexpensive computation of information measures only on discrete variables; provides immunity to prior class probabilities; and brackets the probability of error of the classifier. The maximum output information (MOI) algorithms employ this function for feature subset selection by greedy elimination and directed search. The output of the MOI algorithms is a feature subset of user-defined size and an associated trained classifier (MLP/SVM). These algorithms compare favorably with a number of other methods in terms of performance on various artificial and real-world data sets.

Algorithms↗

Organization of the state space of a simple recurrent network before and after training on recursive linguistic structures.

Recurrent neural networks are often employed in the cognitive science community to process symbol sequences that represent various natural language structures. The aim is to study possible neural mechanisms of language processing and aid in development of artificial language processing systems. We used data sets containing recursive linguistic structures and trained the Elman simple recurrent network (SRN) for the next-symbol prediction task. Concentrating on neuron activation clusters in the recurrent layer of SRN we investigate the network state space organization before and after training. Given a SRN and a training stream, we construct predictive models, called neural prediction machines, that directly employ the state space dynamics of the network. We demonstrate two important properties of representations of recursive symbol series in the SRN. First, the clusters of recurrent activations emerging before training are meaningful and correspond to Markov prediction contexts. We show that prediction states that naturally arise in the SRN initialized with small random weights approximately correspond to states of Variable Memory Length Markov Models (VLMM) based on individual symbols (i.e. words). Second, we demonstrate that during training, the SRN reorganizes its state space according to word categories and their grammatical subcategories, and the next-symbol prediction is again based on the VLMM strategy. However, after training, the prediction is based on word categories and their grammatical subcategories rather than individual words. Our conclusion holds for small depths of recursions that are comparable to human performances. The methods of SRN training and analysis of its state space introduced in this paper are of a general nature and can be used for investigation of processing of any other symbol time series by means of SRN.

Artificial Intelligence↗

Neural-network-based adaptive UPFC for improving transient stability performance of power system.

This paper uses the recently proposed H(infinity)-learning method, for updating the parameter of the radial basis function neural network (RBFNN) used as a control scheme for the unified power flow controller (UPFC) to improve the transient stability performance of a multimachine power system. The RBFNN uses a single neuron architecture whose input is proportional to the difference in error and the updating of its parameters is carried via a proportional value of the error. Also, the coefficients of the difference of error, error, and auxiliary signal used for improving damping performance are depicted by a genetic algorithm. The performance of the newly designed controller is evaluated in a four-machine power system subjected to different types of disturbances. The newly designed single-neuron RBFNN-based UPFC exhibits better damping performance compared to the conventional PID as well as the extended Kalman filter (EKF) updating-based RBFNN scheme, making the unstable cases stable. Its simple architecture reduces the computational burden, thereby making it attractive for real-time implementation. Also, all the machines are being equipped with the conventional power system stabilizer (PSS) to study the coordinated effect of UPFC and PSS in the system.

Algorithms↗

Computer prediction of allergen proteins from sequence-derived protein structural and physicochemical properties.

BACKGROUND: Computational methods have been developed for predicting allergen proteins from sequence segments that show identity, homology, or motif match to a known allergen. These methods achieve good prediction accuracies, but are less effective for novel proteins with no similarity to any known allergen. METHODS: This work tests the feasibility of using a statistical learning method, support vector machines, as such a method. The prediction system is trained and tested by using 1005 allergen proteins from the Allergome database and 22,469 non-allergen proteins from 7871 Pfam families. RESULTS: Testing results by an independent set of 229 allergen and 6717 non-allergen proteins from 7871 Pfam families show that 93.0% and 99.9% of these are correctly predicted, which are comparable to the best results of other methods. Of the 18 novel allergen proteins non-homologous to any other proteins in the Swissprot database, 88.9% is correctly predicted. A further screening of 168,128 proteins in the Swissprot database finds that 2.9% of the proteins are predicted as allergen proteins, which is consistent with the estimated numbers from motif-based methods. CONCLUSIONS: Our study suggests that SVM is a potentially useful method for predicting allergen proteins and it has certain capability for predicting novel allergen proteins. Our software can be accessed at .

Allergens↗

Learning interpretable SVMs for biological sequence classification.

BACKGROUND: Support Vector Machines (SVMs)--using a variety of string kernels--have been successfully applied to biological sequence classification problems. While SVMs achieve high classification accuracy they lack interpretability. In many applications, it does not suffice that an algorithm just detects a biological signal in the sequence, but it should also provide means to interpret its solution in order to gain biological insight. RESULTS: We propose novel and efficient algorithms for solving the so-called Support Vector Multiple Kernel Learning problem. The developed techniques can be used to understand the obtained support vector decision function in order to extract biologically relevant knowledge about the sequence analysis problem at hand. We apply the proposed methods to the task of acceptor splice site prediction and to the problem of recognizing alternatively spliced exons. Our algorithms compute sparse weightings of substring locations, highlighting which parts of the sequence are important for discrimination. CONCLUSION: The proposed method is able to deal with thousands of examples while combining hundreds of kernels within reasonable time, and reliably identifies a few statistically significant positions.

Algorithms↗

Conditional visuo-motor learning in primates: a key role for the basal ganglia.

Sensory guidance of behavior often involves standard visuo-motor mapping of body movements onto objects and spatial locations. For example, looking at and reaching to grasp a glass of wine requires the mapping of the eyes and hand to the location of the glass in space, as well as the formation of a hand configuration appropriate to the shape of the glass. But our brain is far more than just a standard sensorimotor mapping machine. Through evolution, the brain of advanced mammals, in particular human and non-human primates, has acquired a formidable capacity to construct non-standard, arbitrary mapping using associations between external events and behavioral responses that bear no direct relationship. For example, we have all learned to stop at a red traffic light and to go at a green one, or to wait for a specific tone before dialing a phone number and to hang up when hearing a busy signal. These arbitrary associations are acquired through experience, thereby providing primates with a rich and flexible sensorimotor repertoire. Understanding how they are learned, and how they are recalled and used when the context requires them, has been one of the challenging issues for cognitive neuroscience. Valuable insights have been gained over the last two decades through the convergence of multiple complementary approaches. Human neuropsychology and experimental lesions in monkeys have identified a network of brain structures important for conditional sensorimotor associations, whereas imaging studies in healthy human subjects and electrophysiological recordings in awake monkeys have sought to identify the different functional processes underlying the overall function. The present review focuses on the contribution of a network linking the prefrontal cortex, basal ganglia, and dorsal premotor cortex, with special emphasis on results from recording experiments in monkeys. We will first review data pointing to a specific contribution of each component of the network to the performance of well-learned arbitrary visuo-motor associations, as well as data suggesting how novel associations are formed. Then we will propose a model positing that each component of the fronto-striatal network makes a specific contribution to the formation and/or execution of sensorimotor associations. In this model, the basal ganglia are thought to play a key role in linking the sensory, motor, and reward information necessary for arbitrary mapping.

Animals↗

Optimization of rifamycin B fermentation in shake flasks via a machine-learning-based approach.

Rifamycin B is an important polyketide antibiotic used in the treatment of tuberculosis and leprosy. We present results on medium optimization for Rifamycin B production via a barbital insensitive mutant strain of Amycolatopsis mediterranei S699. Machine-learning approaches such as Genetic algorithm (GA), Neighborhood analysis (NA) and Decision Tree technique (DT) were explored for optimizing the medium composition. Genetic algorithm was applied as a global search algorithm while NA was used for a guided local search and to develop medium predictors. The fermentation medium for Rifamycin B consisted of nine components. A large number of distinct medium compositions are possible by variation of concentration of each component. This presents a large combinatorial search space. Optimization was achieved within five generations via GA as well as NA. These five generations consisted of 178 shake-flask experiments, which is a small fraction of the search space. We detected multiple optima in the form of 11 distinct medium combinations. These medium combinations provided over 600% improvement in Rifamycin B productivity. Genetic algorithm performed better in optimizing fermentation medium as compared to NA. The Decision Tree technique revealed the media-media interactions qualitatively in the form of sets of rules for medium composition that give high as well as low productivity.

Actinomycetales↗

Protein structure and fold prediction using Tree-Augmented naïve Bayesian classifier.

Due to the large volume of protein sequence data, computational methods to determine the structure class and the fold class of a protein sequence have become essential. Several techniques based on sequence similarity, Neural Networks, Support Vector Machines (SVMs), etc. have been applied. Since most of these classifiers use binary classifiers for multi-classification, there may be (N) c2 classifiers required. This paper presents a framework using the Tree-Augmented Bayesian Networks (TAN) which performs multi-classification based on the theory of learning Bayesian Networks and using improved feature vector representation of (Ding et al., 2001). In order to enhance TAN's performance, pre-processing of data is done by feature discretization and post-processing is done by using Mean Probability Voting (MPV) scheme. The advantage of using Bayesian approach over other learning methods is that the network structure is intuitive. In addition, one can read off the TAN structure probabilities to determine the significance of each feature (say, hydrophobicity) for each class, which helps to further understand the complexity in protein structure. The experiments on the datasets used in three prominent recent works show that our approach is more accurate than other discriminative methods. The framework is implemented on the BAYESPROT web server and it is available at http://www-appn.comp.nus.edu.sg/~bioinfo/bayesprot/Default.htm. More detailed results are also available on the above website.

Algorithms↗