Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58Linked to original sources

Genetic algorithm learning as a robust approach to RNA editing site prediction.

BACKGROUND: RNA editing is one of several post-transcriptional modifications that may contribute to organismal complexity in the face of limited gene complement in a genome. One form, known as C --> U editing, appears to exist in a wide range of organisms, but most instances of this form of RNA editing have been discovered serendipitously. With the large amount of genomic and transcriptomic data now available, a computational analysis could provide a more rapid means of identifying novel sites of C --> U RNA editing. Previous efforts have had some success but also some limitations. We present a computational method for identifying C --> U RNA editing sites in genomic sequences that is both robust and generalizable. We evaluate its potential use on the best data set available for these purposes: C --> U editing sites in plant mitochondrial genomes. RESULTS: Our method is derived from a machine learning approach known as a genetic algorithm. REGAL (RNA Editing site prediction by Genetic Algorithm Learning) is 87% accurate when tested on three mitochondrial genomes, with an overall sensitivity of 82% and an overall specificity of 91%. REGAL's performance significantly improves on other ab initio approaches to predicting RNA editing sites in this data set. REGAL has a comparable sensitivity and higher specificity than approaches which rely on sequence homology, and it has the advantage that strong sequence conservation is not required for reliable prediction of edit sites. CONCLUSION: Our results suggest that ab initio methods can generate robust classifiers of putative edit sites, and we highlight the value of combinatorial approaches as embodied by genetic algorithms. We present REGAL as one approach with the potential to be generalized to other organisms exhibiting C --> U RNA editing.

Algorithms↗

Support vector machines for learning to identify the critical positions of a protein.

A method for identifying the positions in the amino acid sequence, which are critical for the catalytic activity of a protein using support vector machines (SVMs) is introduced and analysed. SVMs are supported by an efficient learning algorithm and can utilize some prior knowledge about the structure of the problem. The amino acid sequences of the variants of a protein, created by inducing mutations, along with their fitness are required as input data by the method to predict its critical positions. To investigate the performance of this algorithm, variants of the beta-lactamase enzyme were created in silico using simulations of both mutagenesis and recombination protocols. Results from literature on beta-lactamase were used to test the accuracy of this method. It was also compared with the results from a simple search algorithm. The algorithm was also shown to be able to predict critical positions that can tolerate two different amino acids and retain function.

Algorithms↗

A neural-network-based method for predicting protein stability changes upon single point mutations.

MOTIVATION: One important requirement for protein design is to be able to predict changes of protein stability upon mutation. Different methods addressing this task have been described and their performance tested considering global linear correlation between predicted and experimental data. Neither is direct statistical evaluation of their prediction performance available, nor is a direct comparison among different approaches possible. Recently, a significant database of thermodynamic data on protein stability changes upon single point mutation has been generated (ProTherm). This allows the application of machine learning techniques to predicting free energy stability changes upon mutation starting from the protein sequence. RESULTS: In this paper, we present a neural-network-based method to predict if a given mutation increases or decreases the protein thermodynamic stability with respect to the native structure. Using a dataset consisting of 1615 mutations, our predictor correctly classifies >80% of the mutations in the database. On the same task and using the same data, our predictor performs better than other methods available on the Web. Moreover, when our system is coupled with energy-based methods, the joint prediction accuracy increases up to 90%, suggesting that it can be used to increase also the performance of pre-existing methods, and generally to improve protein design strategies. AVAILABILITY: The server is under construction and will be available at http://www.biocomp.unibo.it

Algorithms↗

An incremental approach to genetic-algorithms-based classification.

Incremental learning has been widely addressed in the machine learning literature to cope with learning tasks where the learning environment is ever changing or training samples become available over time. However, most research work explores incremental learning with statistical algorithms or neural networks, rather than evolutionary algorithms. The work in this paper employs genetic algorithms (GAs) as basic learning algorithms for incremental learning within one or more classifier agents in a multiagent environment. Four new approaches with different initialization schemes are proposed. They keep the old solutions and use an "integration" operation to integrate them with new elements to accommodate new attributes, while biased mutation and crossover operations are adopted to further evolve a reinforced solution. The simulation results on benchmark classification data sets show that the proposed approaches can deal with the arrival of new input attributes and integrate them with the original input space. It is also shown that the proposed approaches can be successfully used for incremental learning and improve classification rates as compared to the retraining GA. Possible applications for continuous incremental training and feature selection are also discussed.

Algorithms↗

A gene mapping expert system.

Expert systems are now commonly developed to solve practical problems. Nevertheless, genetics has just begun to benefit from this new technology, since genetic expert systems are extremely rare and often purely experimental. A prototype for risk calculation in pedigrees was developed at the University of Utah, using a commercial frames/rules developmental shell (Intelligence Compiler), which runs on an IBM PC. When small data sets were used, the implementation functioned well, but it could not handle larger data sets. Performance became a major issue, with two possible solutions. The first possibility would have been to port the system to a more powerful machine, and the second would have been to use several different shells or languages, each efficiently representing a specific type of knowledge. Neither of these solutions was applicable in this case. From this experience, we learned that performance, portability, and modifiability were three major requirements for genetic expert systems. To achieve these goals, we implemented the gene mapping expert system GMES: (GMES is unrelated to the gene mapping system, GMS in Lisp combined with a frame/object shell (FROBS). We were able to efficiently represent, control, and optimize a gene mapping experiment, achieving portability by building GMES on top of a C-based version of Common Lisp. Lisp combined with the FROBS expert system shell permitted a declarative representation of each of the components of the experiment, resulting in a transplant specification of the problem within a maintainable system.

Algorithms↗

Next-Generation Disease Profiling by Integrating Histopathology with Spatial Multi-Omics Data.

The field of pathology has experienced several transformative changes in recent years with the advent of digital pathology and spatial multi-omics. These technologies have enhanced every aspect of pathology practice, from streamlining daily workflows to generating high-fidelity multi-omics data that provide pathologists with novel tools to refine disease profiling and clinical diagnosis. Each layer of multimodal data (genomic, metabolomic, proteomic, or transcriptomic) has uncovered a distinct facet of disease pathologies, and combined with machine learning/artificial intelligence-based data analysis and pattern recognition models, has provided holistic understanding of regulatory mechanisms underpinning them. However, high-dimensional data have far exceeded the volume, scale, and complexity of immunostaining methods implemented by pathologists and, thus, have generated significant challenges related to deconvolution, interpretation, and clinical translation. Furthermore, these multimodal studies have predominantly relied on computational methods to process data and extract disease-relevant insights, thus raising questions around relevance or role of a pathologist in this new era of multi-omics. This review will provide a perspective on the evolving fields of molecular histopathology and spatial -omics, leveraging them to approach disease profiling, and redefining the role of a pathologist during this process.

Humans↗

[Application of support vector machines to classification of blood cells].

The support vector machine (SVM) is a new learning technique based on the statistical learning theory. It was originally developed for two-class classification. In this paper, the SVM approach is extended to multi-class classification problems, a hierarchical SVM is applied to classify blood cells in different maturation stages from bone marrow. Based on stepwise decomposition, a hierarchical clustering method is presented to construct the architecture of the hierarchical (tree-like) SVM, then the optimal control parameters of SVM are determined by some criterion for each discriminant step. To verify the performances of classifiers, the SVM method is compared with three classical classifiers using 3-fold cross validation. The preliminary results indicate that the proposed method avoids the curse of dimensionality and has greater generalization. Thus, the method can improve the classification correctness for blood cells from bone marrow.

Algorithms↗

Computerized radiographic mass detection--part II: Decision support by featured database visualization and modular neural networks.

Based on the enhanced segmentation of suspicious mass areas, further development of computer-assisted mass detection may be decomposed into three distinctive machine learning tasks: 1) construction of the featured knowledge database; 2) mapping of the classified and/or unclassified data points in the database; and 3) development of an intelligent user interface. A decision support system may then be constructed as a complementary machine observer that should enhance the radiologists performance in mass detection. We adopt a mathematical feature extraction procedure to construct the featured knowledge database from all the suspicious mass sites localized by the enhanced segmentation. The optimal mapping of the data points is then obtained by learning the generalized normal mixtures and decision boundaries, where a is developed to carry out both soft and hard clustering. A visual explanation of the decision making is further invented as a decision support, based on an interactive visualization hierarchy through the probabilistic principal component projections of the knowledge database and the localized optimal displays of the retrieved raw data. A prototype system is developed and pilot tested to demonstrate the applicability of this framework to mammographic mass detection.

Artificial Intelligence↗

Stable encoding of finite-state machines in discrete-time recurrent neural nets with sigmoid units.

There has been a lot of interest in the use of discrete-time recurrent neural nets (DTRNN) to learn finite-state tasks, with interesting results regarding the induction of simple finite-state machines from input-output strings. Parallel work has studied the computational power of DTRNN in connection with finite-state computation. This article describes a simple strategy to devise stable encodings of finite-state machines in computationally capable discrete-time recurrent neural architectures with sigmoid units and gives a detailed presentation on how this strategy may be applied to encode a general class of finite-state machines in a variety of commonly used first- and second-order recurrent neural networks. Unlike previous work that either imposed some restrictions to state values or used a detailed analysis based on fixed-point attractors, our approach applies to any positive, bounded, strictly growing, continuous activation function and uses simple bounding criteria based on a study of the conditions under which a proposed encoding scheme guarantees that the DTRNN is actually behaving as a finite-state machine.

Models, Neurological↗

Reinforcement learning-based dynamic ensemble for missense variant effect prediction and tiered prioritization of VUS.

BACKGROUND: Accurate classification of missense variants remains a challenging task despite major advances in genomics. Numerous computational models have been developed to assist in variant classification, but often require repeated integration and benchmarking efforts. Ensemble methods have been proposed to overcome the limitations of single predictors, but mostly rely on fixed, predefined weights that constrain their ability to capture interactions among predictive signals. METHODS: We present GenixRL, a dynamic ensemble framework that reformulates model fusion as a reinforcement learning optimization problem. GenixRL uses a Q-learning agent to learn a policy that dynamically weights the probabilistic outputs of complementary predictors, including BayesDel (addAF and noAF), ClinPred, and MetaRNN. Replacing static weighting with policy learning allows GenixRL to adaptively identify optimal weightings and substantially improve classification accuracy. RESULTS: In benchmark evaluation against 25 state-of-the-art predictors, GenixRL achieved an AUROC of 0.9644 on an independent ClinVar dataset. On saturation genome editing assays for BRCA1 and BRCA2, GenixRL achieved the best performance and ranked highest on 14 of 17 clinically significant genes in a zero-shot evaluation. Applied to uncertain and conflicting ClinVar variants, GenixRL enabled tiered, evidence-based prioritization of hundreds of thousands of variants as likely pathogenic or pathogenic with high confidence, supported by orthogonal population evidence from gnomAD. CONCLUSION: GenixRL advances pathogenicity prediction for missense variants and provides an adaptive ensemble that sorts variants of uncertain significance into tiered candidates for expert curation and functional validation.

Mutation, Missense↗

Children using cellular phones: the effects of shortcomings in user interface design.

Differences in cellular phones' complexity and their impact on children's performance are under study in this experiment. Twenty children (age 9-14 years) solved tasks on two phones that were simulated according to existing models on a PC with a touch screen, holding constant display size, fonts, and colors. Actions were logged and analyzed regarding execution time, detour steps, and specific errors. Results show that children using the Siemens C35i with 25% higher complexity (with regard to number of required production rules) spent double the time solving tasks and undertook three times as many detour steps as children using the less complex Nokia 3210. A detailed analysis of user actions revealed that the number of production rules to be learned fails to account for most difficulties. Instead, ambiguous naming, poor categorization of functions, and unclear functionality of keys undermined performance. Actual or potential applications of this research include guidelines to improve the usability of all devices with small displays and hierarchical menu structures, such as cellular phones.

Adolescent↗

Using support vector machines to optimally classify rotator cuff strength data and quantify post-operative strength in rotator cuff tear patients.

Shoulder strength data are important for post-operative assessment of shoulder function and have been used in diagnosis of rotator cuff pathology. Support vector machines (SVM) employ complex analysis techniques to solve classification and regression problems. A SVM, a machine learning technique, can be used for analysis and classification of shoulder strength data. The goals of this study were to determine the diagnostic competency of SVM based on shoulder strength data and to apply SVM analysis in efforts to derive a single representative shoulder strength score. Data were taken from fourteen isometric shoulder strength measurements of each shoulder (involved and uninvolved) in 45 rotator cuff tear patients. SVM diagnostic proficiency was found to be comparable to reported ultrasound values. Improvement of shoulder function was accurately represented by a single score in pairwise comparison of the pre-operative and the 12 month post-operative group (P < 0.004). Thus, the SVM-based score may be a promising metric for summarizing rotator cuff strength data.

Artificial Intelligence↗

Efficiently mining gene expression data via a novel parameterless clustering method.

Clustering analysis has been an important research topic in the machine learning field due to the wide applications. In recent years, it has even become a valuable and useful tool for in-silico analysis of microarray or gene expression data. Although a number of clustering methods have been proposed, they are confronted with difficulties in meeting the requirements of automation, high quality, and high efficiency at the same time. In this paper, we propose a novel, parameterless and efficient clustering algorithm, namely, Correlation Search Technique (CST), which fits for analysis of gene expression data. The unique feature of CST is it incorporates the validation techniques into the clustering process so that high quality clustering results can be produced on the fly. Through experimental evaluation, CST is shown to outperform other clustering methods greatly in terms of clustering quality, efficiency, and automation on both of synthetic and real data sets.

Algorithms↗

From latent disseminated cells to overt metastasis: genetic analysis of systemic breast cancer progression.

According to the present view, metastasis marks the end in a sequence of genomic changes underlying the progression of an epithelial cell to a lethal cancer. Here, we aimed to find out at what stage of tumor development transformed cells leave the primary tumor and whether a defined genotype corresponds to metastatic disease. To this end, we isolated single disseminated cancer cells from bone marrow of breast cancer patients and performed single-cell comparative genomic hybridization. We analyzed disseminated tumor cells from patients after curative resection of the primary tumor (stage M0), as presumptive progenitors of manifest metastasis, and from patients with manifest metastasis (stage M1). Their genomic data were compared with those from microdissected areas of matched primary tumors. Disseminated cells from M0-stage patients displayed significantly fewer chromosomal aberrations than primary tumors or cells from M1-stage patients (P < 0.008 and P < 0.0001, respectively), and their aberrations appeared to be randomly generated. In contrast, primary tumors and M1 cells harbored different and characteristic chromosomal imbalances. Moreover, applying machine-learning methods for the classification of the genotypes, we could correctly identify the presence or absence of metastatic disease in a patient on the basis of a single-cell genome. We suggest that in breast cancer, tumor cells may disseminate in a far less progressed genomic state than previously thought, and that they acquire genomic aberrations typical of metastatic cells thereafter. Thus, our data challenge the widely held view that the precursors of metastasis are derived from the most advanced clone within the primary tumor.

Algorithms↗

Control of FES in paraplegia: modeling voluntary arm forces.

Skilled behavior is difficult or impossible to articulate explicitly by the performers. Likewise biomechanical models of skilled motor actions are often limited by the lack of knowledge of the underlying mechanisms. A 'behavioral cloning' technique is described, based on a trained artificial neural network (ANN), that precisely mimics an individual's learned skill. In this paper the motor skill considered is that of paraplegics using their upper limbs whilst standing-up with FES. In a group of eight paraplegics with complete spinal injuries, it was possible to develop clones that followed closely the observed behavior of the subjects. Each subject used a unique and consistent voluntary control strategy. Subjects with more experience in using FES were more consistent in the use of their arms from trial to trial. Comparison of the clones revealed features suggestive of some common underlying voluntary control strategies.

Adolescent↗

Learning to lead at Toyota.

Many companies have tried to copy Toyota's famous production system--but without success. Why? Part of the reason, says the author, is that imitators fail to recognize the underlying principles of the Toyota Production System (TPS), focusing instead on specific tools and practices. This article tells the other part of the story. Building on a previous HBR article, "Decoding the DNA of the Toyota Production System," Spear explains how Toyota inculcates managers with TPS principles. He describes the training of a star recruit--a talented young American destined for a high-level position at one of Toyota's U.S. plants. Rich in detail, the story offers four basic lessons for any company wishing to train its managers to apply Toyota's system: There's no substitute for direct observation. Toyota employees are encouraged to observe failures as they occur--for example, by sitting next to a machine on the assembly line and waiting and watching for any problems. Proposed changes should always be structured as experiments. Employees embed explicit and testable assumptions in the analysis of their work. That allows them to examine the gaps between predicted and actual results. Workers and managers should experiment as frequently as possible. The company teaches employees at all levels to achieve continuous improvement through quick, simple experiments rather than through lengthy, complex ones. Managers should coach, not fix. Toyota managers act as enablers, directing employees but not telling them where to find opportunities for improvements. Rather than undergo a brief period of cursory walk-throughs, orientations, and introductions as incoming fast-track executives at most companies might, the executive in this story learned TPS the long, hard way--by practicing it, which is how Toyota trains any new employee, regardless of rank or function.

Administrative Personnel↗

Using automatically learnt verb selectional preferences for classification of biomedical terms.

In this paper, we present an approach to term classification based on verb selectional patterns (VSPs), where such a pattern is defined as a set of semantic classes that could be used in combination with a given domain-specific verb. VSPs have been automatically learnt based on the information found in a corpus and an ontology in the biomedical domain. Prior to the learning phase, the corpus is terminologically processed: term recognition is performed by both looking up the dictionary of terms listed in the ontology and applying the C/NC-value method for on-the-fly term extraction. Subsequently, domain-specific verbs are automatically identified in the corpus based on the frequency of occurrence and the frequency of their co-occurrence with terms. VSPs are then learnt automatically for these verbs. Two machine learning approaches are presented. The first approach has been implemented as an iterative generalisation procedure based on a partial order relation induced by the domain-specific ontology. The second approach exploits the idea of genetic algorithms. Once the VSPs are acquired, they can be used to classify newly recognised terms co-occurring with domain-specific verbs. Given a term, the most frequently co-occurring domain-specific verb is selected. Its VSP is used to constrain the search space by focusing on potential classes of the given term. A nearest-neighbour approach is then applied to select a class from the constrained space of candidate classes. The most similar candidate class is predicted for the given term. The similarity measure used for this purpose combines contextual, lexical, and syntactic properties of terms.

Abstracting and Indexing↗

Artificial Intelligence in Diagnosing Depression Through Behavioural Cues: A Diagnostic Accuracy Systematic Review and Meta-Analysis.

AIM: To synthesise existing evidence concerning the application of AI methods in detecting depression through behavioural cues among adults in healthcare and community settings. DESIGN: This is a diagnostic accuracy systematic review. METHODS: This review included studies examining different AI methods in detecting depression among adults. Two independent reviewers screened, appraised and extracted data. Data were analysed by meta-analysis, narrative synthesis and subgroup analysis. DATA SOURCES: Published studies and grey literature were sought in 11 electronic databases. Hand search was conducted on reference lists and two journals. RESULTS: In total, 30 studies were included in this review. Twenty of which demonstrated that AI models had the potential to detect depression. Speech and facial expression showed better sensitivity, reflecting the ability to detect people with depression. Text and movement had better specificity, indicating the ability to rule out non-depressed individuals. Heterogeneity was initially high. Less heterogeneity was observed within each modality subgroup. CONCLUSIONS: This is the first systematic review examining AI models in detecting depression using all four behavioural cues: speech, texts, movement and facial expressions. IMPLICATIONS: A collaborative effort among healthcare professionals can be initiated to develop an AI-assisted depression detection system in general healthcare or community settings. IMPACT: It is challenging for general healthcare professionals to detect depressive symptoms among people in non-psychiatric settings. Our findings suggested the need for objective screening tools, such as an AI-assisted system, for screening depression. Therefore, people could receive accurate diagnosis and proper treatments for depression. REPORTING METHOD: This review followed the PRISMA checklist. PATIENTS OR PUBLIC CONTRIBUTION: No patients or public contribution.

Humans↗