Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Survival prediction for clear cell renal cell carcinoma based on deep multimodal synergistic survival network.

Objective.To propose a deep multimodal synergistic survival analysis framework (Deep Multimodal Synergistic Survival Network, DMSSN) to achieve accurate prognostic analysis for clear cell renal cell carcinoma (ccRCC).Methods.This study (DMSSN) utilized matched multimodal data from the Cancer Genome Atlas-KIRC database, including CT imaging data, whole slide images, copy number variation (CNV) features, and clinical data. Deep Canonical Correlation Analysis was employed to map heterogeneous modalities into a shared latent space. Contrastive learning was introduced to enhance semantic consistency across multimodal features, and a gating network was utilized for the adaptive fusion of multimodal information to achieve precise survival risk prediction for patients.Results.Experimental results demonstrated that DMSSN achieved a Concordance Index (C-index) of 0.8153 ± 0.0994, with a Log-rank testp-value of 1.6553×10-11. DMSSN exhibited significant performance advantages over traditional statistical methods like Log-rank-Cox (0.7055 ± 0.0670) and machine learning methods such as Random Survival Forest (RSF) (0.6836 ± 0.1048). Furthermore, in comparison with similar deep learning approaches, DMSSN outperformed late fusion strategies (0.7493 ± 0.1211) and discrete-time survival models such as DeepHit (0.7655 ± 0.1041) and Nnet-surv (0.7694 ± 0.0635). Notably, DMSSN still achieved the best predictive performance when compared to the classic deep survival model DeepSurv (0.7919 ± 0.0978) and advanced state-of-the-art multimodal fusion frameworks like Context-Aware Transformer (0.7735 ± 0.0818) and Multimodal Co-Attention Transformer (0.8102 ± 0.0972). Ablation studies showed that removing any single modality led to a decline in performance, with the largest numerical decrease occurring after removing CT imaging features (C-index decreased to 0.7327), validating the complementarity of multimodal data and the pivotal role of radiomic features in prognostic assessment. Module ablation experiments further confirmed the effectiveness of the core components.Conclusion:By effectively integrating imaging, pathology, genomic, and clinical features, the DMSSN framework demonstrates superior performance and robustness in the survival prediction of ccRCC.

Carcinoma, Renal Cell↗

Nanopore cheminformatics.

A cheminformatics method is described for classification, and biophysical examination, of individual molecules. A novel molecular detector is used--one based on current blockade measurements through a nanometer-scale ion channel (alpha-hemolysin). Classification results are described for blockades caused by DNA molecules in the alpha-hemolysin nanopore detector, with signal analysis and pattern recognition performed using a combination of methods from bioinformatics and machine learning. Due to the size of the alpha-hemolysin protein channel, the blockade events report on one DNA molecule at a time, which enables a variety of reproducible, single-molecule biophysical experiments. To capture the full sensitivity of the nanopore detector's blockade signal, Hidden Markov Models (HMMs) were used with Expectation/Maximization for denoising and for associating a feature vector with the ionic current blockade of each captured DNA molecule. Support Vector Machines (SVMs) that employ novel kernel designs were then used as discriminators. With SVM training performed off-line, and economical HMM processing on-line, blockade classification was possible during capture. HMMs were also used in conjunction with a time-domain finite state automaton (off-line) for feature discovery and kinetics analysis. Analysis of the DNA data indicates a variety of binding (DNA-protein), fraying, and conformational shifts that are consistent with data obtained from thermodynamic analyses (melting curves), X-ray crystallography, and NMR studies. The software tools are designed for analysis of generic blockades in ionic channels, including those in other biological pore-forming toxins, other biological channels in general, and semiconductor-based channels.

Base Sequence↗

Deep learning-based annotation of plant abiotic stress resistance genes for crops.

The declining costs of DNA sequencing have expanded genomic data, crucial for understanding plant abiotic stress responses and crop improvement. However, accurate gene annotation remains challenging. To address this limitation, we propose the PASRGA, a deep learning approach that leverages transfer learning and contrastive learning to annotate genes related to drought, salt, cold, and UV resistance. PASRGA achieves high F1-scores, area under the receiver operating characteristic (AUROC), area under the precision-recall curve (AUPRC), and Matthews correlation coefficient (MCC) in annotating stress resistance genes, significantly outperforming the general protein annotation model CLEAN, the plant phosphatase gene annotation model PF-NET, the top-ranked model in the CAFA5 challenge NetGO 4.0, and four traditional machine learning methods. Its effectiveness was further validated with a salt stress treatment experiment in Eutrema salsugineum. To facilitate crop breeding practices, we utilized PASRGA to annotate the genomes of 17 major crops. To improve accessibility and utility, we incorporated both manually curated and PASRGA-predicted gene data, together with the PASRGA tool, into the PlantASRG database (https://bioinfor.nefu.edu.cn/PlantASRG/). This comprehensive resource aims to support crop breeding initiatives and ensure food security.

Crops, Agricultural↗

An in-silico method for prediction of polyadenylation signals in human sequences.

This paper presents a machine learning method to predict polyadenylation signals (PASes) in human DNA and mRNA sequences by analysing features around them. This method consists of three sequential steps of feature manipulation: generation, selection and integration of features. In the first step, new features are generated using k-gram nucleotide acid or amino acid patterns. In the second step, a number of important features are selected by an entropy-based algorithm. In the third step, support vector machines are employed to recognize true PASes from a large number of candidates. Our study shows that true PASes in DNA and mRNA sequences can be characterized by different features, and also shows that both upstream and downstream sequence elements are important for recognizing PASes from DNA sequences. We tested our method on several public data sets as well as our own extracted data sets. In most cases, we achieved better validation results than those reported previously on the same data sets. The important motifs observed are highly consistent with those reported in literature.

Base Sequence↗

MAP: an iterative experimental design methodology for the optimization of catalytic search space structure modeling.

One of the main problems in high-throughput research for materials is still the design of experiments. At early stages of discovery programs, purely exploratory methodologies coupled with fast screening tools should be employed. This should lead to opportunities to find unexpected catalytic results and identify the "groups" of catalyst outputs, providing well-defined boundaries for future optimizations. However, very few new papers deal with strategies that guide exploratory studies. Mostly, traditional designs, homogeneous covering, or simple random samplings are exploited. Typical catalytic output distributions exhibit unbalanced datasets for which an efficient learning is hardly carried out, and interesting but rare classes are usually unrecognized. Here is suggested a new iterative algorithm for the characterization of the search space structure, working independently of learning processes. It enhances recognition rates by transferring catalysts to be screened from "performance-stable" space zones to "unsteady" ones which necessitate more experiments to be well-modeled. The evaluation of new algorithm attempts through benchmarks is compulsory due to the lack of past proofs about their efficiency. The method is detailed and thoroughly tested with mathematical functions exhibiting different levels of complexity. The strategy is not only empirically evaluated, the effect or efficiency of sampling on future Machine Learning performances is also quantified. The minimum sample size required by the algorithm for being statistically discriminated from simple random sampling is investigated.

Algorithms↗

Experiments with AdaBoost.RT, an improved boosting scheme for regression.

The application of boosting technique to regression problems has received relatively little attention in contrast to research aimed at classification problems. This letter describes a new boosting algorithm, AdaBoost.RT, for regression problems. Its idea is in filtering out the examples with the relative estimation error that is higher than the preset threshold value, and then following the AdaBoost procedure. Thus, it requires selecting the suboptimal value of the error threshold to demarcate examples as poorly or well predicted. Some experimental results using the M5 model tree as a weak learning machine for several benchmark data sets are reported. The results are compared to other boosting methods, bagging, artificial neural networks, and a single M5 model tree. The preliminary empirical comparisons show higher performance of AdaBoost.RT for most of the considered data sets.

Algorithms↗

Predicting the efficacy of short oligonucleotides in antisense and RNAi experiments with boosted genetic programming.

MOTIVATION: Both small interfering RNAs (siRNAs) and antisense oligonucleotides can selectively block gene expression. Although the two methods rely on different cellular mechanisms, these methods share the common property that not all oligonucleotides (oligos) are equally effective. That is, if mRNA target sites are picked at random, many of the antisense or siRNA oligos will not be effective. Algorithms that can reliably predict the efficacy of candidate oligos can greatly reduce the cost of knockdown experiments, but previous attempts to predict the efficacy of antisense oligos have had limited success. Machine learning has not previously been used to predict siRNA efficacy. RESULTS: We develop a genetic programming based prediction system that shows promising results on both antisense and siRNA efficacy prediction. We train and evaluate our system on a previously published database of antisense efficacies and our own database of siRNA efficacies collected from the literature. The best models gave an overall correlation between predicted and observed efficacy of 0.46 on both antisense and siRNA data. As a comparison, the best correlations of support vector machine classifiers trained on the same data were 0.40 and 0.30, respectively.

Algorithms↗

Adaptive hybrid learning for neural networks.

A robust locally adaptive learning algorithm is developed via two enhancements of the Resilient Propagation (RPROP) method. Remaining drawbacks of the gradient-based approach are addressed by hybridization with gradient-independent Local Search. Finally, a global optimization method based on recursion of the hybrid is constructed, making use of tabu neighborhoods to accelerate the search for minima through diversification. Enhanced RPROP is shown to be faster and more accurate than the standard RPROP in solving classification tasks based on natural data sets taken from the UCI repository of machine learning databases. Furthermore, the use of Local Search is shown to improve Enhanced RPROP by solving the same classification tasks as part of the global optimization method.

Algorithms↗

An SVM scorer for more sensitive and reliable peptide identification via tandem mass spectrometry.

Tandem mass spectrometry (MS/MS) has become increasingly important and indispensable in high-throughput proteomics for identifying complex protein mixtures. Database searching is the standard method to accomplish this purpose. A key sub-routine, peptide identification, is used to generate a list of candidate peptides from a protein database according to an experimental MS/MS spectrum, and then validate these candidate peptides for protein identification. Although currently there are many algorithms for peptide identification, most of them either lack an effective validation module or only validate the first-ranked peptide, thus leading to a low identification reliability or sensitivity. This paper proposes a new algorithm, named pepReap, to overcome the above drawbacks. It consists of a two-layered scoring scheme based on machine learning. The first layer is a rough scoring function which uses some simple and heuristic factors to measure the degree of the matches between an experimental MS/MS spectrum and the candidate peptides; thus a ranked list of candidate peptides is generated at a relatively low computational cost. The second layer is a fine scoring function which re-ranks the candidate peptides generated in the first layer and determines which one among them is the true positive. The fine scoring function was designed based on support vector machines (SVMs) using more comprehensive factors, such as the correlations between ions, the mass matching errors of fragment and peptide ions, etc. Consequently, the SVM classifier serves as not only a scorer but also a validation module. Experimental comparison with the popular SEQUEST algorithm coupled with threshold validation criteria on a reported dataset demonstrates that the pepReap algorithm achieves higher performance in terms of identification sensitivity with comparable precision.

Algorithms↗

PCP: a program for supervised classification of gene expression profiles.

UNLABELLED: PCP (Pattern Classification Program) is an open-source machine learning program for supervised classification of patterns (vectors of measurements). The principal use of PCP in bioinformatics is design and evaluation of classifiers for use in clinical diagnostic tests based on measurements of gene expression. PCP implements leading pattern classification and gene selection algorithms and incorporates cross-validation estimation of classifier performance. Importantly, the implementation integrates gene selection and class prediction stages, which is vital for computing reliable performance estimates in small-sample scenarios. Additionally, the program includes automated and efficient model selection (optimization of parameters) for support vector machine (SVM) classifier. The distribution includes Linux and Windows/Cygwin binaries. The program can easily be ported to other platforms. AVAILABILITY: Free download at http://pcp.sourceforge.net

Algorithms↗

Structure-based prediction of bZIP partnering specificity.

Predicting protein interaction specificity from sequence is an important goal in computational biology. We present a model for predicting the interaction preferences of coiled-coil peptides derived from bZIP transcription factors that performs very well when tested against experimental protein microarray data. We used only sequence information to build atomic-resolution structures for 1711 dimeric complexes, and evaluated these with a variety of functions based on physics, learned empirical weights or experimental coupling energies. A purely physical model, similar to those used for protein design studies, gave reasonable performance. The results were improved significantly when helix propensities were used in place of a structurally explicit model to represent the unfolded reference state. Further improvement resulted upon accounting for residue-residue interactions in competing states in a generic way. Purely physical structure-based methods had difficulty capturing core interactions accurately, especially those involving polar residues such as asparagine. When these terms were replaced with weights from a machine-learning approach, the resulting model was able to correctly order the stabilities of over 6000 pairs of complexes with greater than 90% accuracy. The final model is physically interpretable, and suggests specific pairs of residues that are important for bZIP interaction specificity. Our results illustrate the power and potential of structural modeling as a method for predicting protein interactions and highlight obstacles that must be overcome to reach quantitative accuracy using a de novo approach. Our method shows unprecedented performance in predicting protein-protein interaction specificity accurately using structural modeling and suggests that predicting coiled-coil interactions generally may be within reach.

Basic-Leucine Zipper Transcription Factors↗

Human mutations in high-confidence Tourette disorder genes affect sensorimotor behavior, reward learning, and striatal dopamine in mice.

Tourette disorder (TD) is poorly understood, despite affecting 1/160 children. A lack of animal models possessing construct, face, and predictive validity hinders progress in the field. We used CRISPR/Cas9 genome editing to generate mice with mutations orthologous to human de novo variants in two high-confidence Tourette genes, CELSR3 and WWC1. Mice with human mutations in Celsr3 and Wwc1 exhibit cognitive and/or sensorimotor behavioral phenotypes consistent with TD. Sensorimotor gating deficits, as measured by acoustic prepulse inhibition, occur in both male and female Celsr3 TD models. Wwc1 mice show reduced prepulse inhibition only in females. Repetitive motor behaviors, common to Celsr3 mice and more pronounced in females, include vertical rearing and grooming. Sensorimotor gating deficits and rearing are attenuated by aripiprazole, a partial agonist at dopamine type II receptors. Unsupervised machine learning reveals numerous changes to spontaneous motor behavior and less predictable patterns of movement. Continuous fixed-ratio reinforcement shows that Celsr3 TD mice have enhanced motor responding and reward learning. Electrically evoked striatal dopamine release, tested in one model, is greater. Brain development is otherwise grossly normal without signs of striatal interneuron loss. Altogether, mice expressing human mutations in high-confidence TD genes exhibit face and predictive validity. Reduced prepulse inhibition and repetitive motor behaviors are core behavioral phenotypes and are responsive to aripiprazole. Enhanced reward learning and motor responding occur alongside greater evoked dopamine release. Phenotypes can also vary by sex and show stronger affection in females, an unexpected finding considering males are more frequently affected in TD.

Animals↗

Human mutations in high-confidence Tourette disorder genes affect sensorimotor behavior, reward learning, and striatal dopamine in mice.

UNLABELLED: Tourette disorder (TD) is poorly understood, despite affecting 1/160 children. A lack of animal models possessing construct, face, and predictive validity hinders progress in the field. We used CRISPR/Cas9 genome editing to generate mice with mutations orthologous to human de novo variants in two high-confidence Tourette genes, CELSR3 and WWC1 . Mice with human mutations in Celsr3 and Wwc1 exhibit cognitive and/or sensorimotor behavioral phenotypes consistent with TD. Sensorimotor gating deficits, as measured by acoustic prepulse inhibition, occur in both male and female Celsr3 TD models. Wwc1 mice show reduced prepulse inhibition only in females. Repetitive motor behaviors, common to Celsr3 mice and more pronounced in females, include vertical rearing and grooming. Sensorimotor gating deficits and rearing are attenuated by aripiprazole, a partial agonist at dopamine type II receptors. Unsupervised machine learning reveals numerous changes to spontaneous motor behavior and less predictable patterns of movement. Continuous fixed-ratio reinforcement shows Celsr3 TD mice have enhanced motor responding and reward learning. Electrically evoked striatal dopamine release, tested in one model, is greater. Brain development is otherwise grossly normal without signs of striatal interneuron loss. Altogether, mice expressing human mutations in high-confidence TD genes exhibit face and predictive validity. Reduced prepulse inhibition and repetitive motor behaviors are core behavioral phenotypes and are responsive to aripiprazole. Enhanced reward learning and motor responding occurs alongside greater evoked dopamine release. Phenotypes can also vary by sex and show stronger affection in females, an unexpected finding considering males are more frequently affected in TD. SIGNIFICANCE STATEMENT: We generated mouse models that express mutations in high-confidence genes linked to Tourette disorder (TD). These models show sensorimotor and cognitive behavioral phenotypes resembling TD-like behaviors. Sensorimotor gating deficits and repetitive motor behaviors are attenuated by drugs that act on dopamine. Reward learning and striatal dopamine is enhanced. Brain development is grossly normal, including cortical layering and patterning of major axon tracts. Further, no signs of striatal interneuron loss are detected. Interestingly, behavioral phenotypes in affected females can be more pronounced than in males, despite male sex bias in the diagnosis of TD. These novel mouse models with construct, face, and predictive validity provide a new resource to study neural substrates that cause tics and related behavioral phenotypes in TD.

Preprint↗

Prediction of oxidoreductase-catalyzed reactions based on atomic properties of metabolites.

MOTIVATION: Our knowledge of metabolism is far from complete, and the gaps in our knowledge are being revealed by metabolomic detection of small-molecules not previously known to exist in cells. An important challenge is to determine the reactions in which these compounds participate, which can lead to the identification of gene products responsible for novel metabolic pathways. To address this challenge, we investigate how machine learning can be used to predict potential substrates and products of oxidoreductase-catalyzed reactions. RESULTS: We examined 1956 oxidation/reduction reactions in the KEGG database. The vast majority of these reactions (1626) can be divided into 12 subclasses, each of which is marked by a particular type of functional group transformation. For a given transformation, the local structures of reaction centers in substrates and products can be characterized by patterns. These patterns are not unique to reactants but are widely distributed among KEGG metabolites. To distinguish reactants from non-reactants, we trained classifiers (linear-kernel Support Vector Machines) using negative and positive examples. The input to a classifier is a set of atomic features that can be determined from the 2D chemical structure of a compound. Depending on the subclass of reaction, the accuracy of prediction for positives (negatives) is 64 to 93% (44 to 92%) when asking if a compound is a substrate and 71 to 98% (50 to 92%) when asking if a compound is a product. Sensitivity analysis reveals that this performance is robust to variations of the training data. Our results suggest that metabolic connectivity can be predicted with reasonable accuracy from the presence or absence of local structural motifs in compounds and their readily calculated atomic features. AVAILABILITY: Classifiers reported here can be used freely for noncommercial purposes via a Java program available upon request.

Algorithms↗

Leveraging protein language models for cross-variant CRISPR/Cas9 sgRNA activity prediction.

MOTIVATION: Accurate prediction of single-guide RNA (sgRNA) activity is crucial for optimizing the CRISPR/Cas9 gene-editing system, as it directly influences the efficiency and accuracy of genome modifications. However, existing prediction methods mainly rely on large-scale experimental data of a single Cas9 variant to construct Cas9 protein (variants)-specific sgRNA activity prediction models, which limits their generalization ability and prediction performance across different Cas9 protein (variants), as well as their scalability to the continuously discovered new variants. RESULTS: In this study, we proposed PLM-CRISPR, a novel deep learning-based model that leverages protein language models to capture Cas9 protein (variants) representations for cross-variant sgRNA activity prediction. PLM-CRISPR uses tailored feature extraction modules for both sgRNA and protein sequences, incorporating a cross-variant training strategy and a dynamic feature fusion mechanism to effectively model their interactions. Extensive experiments demonstrate that PLM-CRISPR outperforms existing methods across datasets spanning seven Cas9 protein (variants) in three real-world scenarios, demonstrating its superior performance in handling data-scarce situations, including cases with few or no samples for novel variants. Comparative analyses with traditional machine learning and deep learning models further confirm the effectiveness of PLM-CRISPR. Additionally, motif analysis reveals that PLM-CRISPR accurately identifies high-activity sgRNA sequence patterns across diverse Cas9 protein (variants). Overall, PLM-CRISPR provides a robust, scalable, and generalizable solution for sgRNA activity prediction across diverse Cas9 protein (variants). AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/CSUBioGroup/PLM-CRISPR.

CRISPR-Cas Systems↗

Decoding cancer with artificial intelligence: Transforming research, diagnosis, and therapy with future insights.

Cancer remains one of the leading global health burdens, with increasing complexity in genomic, imaging, and clinical datasets presenting significant challenges for effective management. Artificial intelligence (AI) has emerged as a powerful tool to address these challenges by enabling pattern recognition, knowledge integration, and data-driven decision-making. This review highlights recent advances in the application of AI across cancer research, diagnosis, and therapy. In research, AI accelerates drug discovery and repurposing, enhances genomic data interpretation, and facilitates biomarker identification through multi-omics integration. In diagnosis, AI has demonstrated high technical performance in radiology for lesion detection and image segmentation, in pathology for tumour grading and molecular prediction, and in liquid biopsy for non-invasive biomarker analysis. In therapy, AI supports precision medicine by predicting treatment responses, monitoring disease progression, and optimizing clinical trial design. Despite these advances, barriers such as data heterogeneity, algorithmic bias, interpretability, and regulatory challenges remain. Future directions, including explainable AI, federated learning, multimodal modelling, and digital twins, hold promise for translating AI-driven innovations into routine oncology practice. Significance Statement This review provides a timely synthesis of recent (2020-2025) advances in artificial intelligence across cancer research, diagnosis, and therapy, highlighting applications in drug discovery, genomics, multi-omics biomarker identification, and clinical decision-making. By integrating technological progress with translational and clinical relevance, this work serves as a valuable resource for bridging AI innovation with precision oncology practice. As a narrative review, the literature was identified through targeted PubMed, Scopus, and Google Scholar searches, combining terms for artificial intelligence, machine learning, and deep learning with cancer-related keywords, with priority given to peer-reviewed studies published between 2020 and 2025, seminal earlier works, and official regulatory or guideline documents. Within each domain, representative studies were selected to illustrate methodological diversity, clinical context, and current translational readiness rather than to provide exhaustive coverage of an extremely rapidly evolving field.

Artificial intelligence↗

Privacy-Enhancing Sequential Learning under Heterogeneous Selection Bias in Multi-Site EHR Data.

OBJECTIVE: To develop privacy-enhancing statistical methods for estimation of binary disease risk model association parameters across multiple electronic health record (EHR) sites with heterogeneous selection mechanisms, without sharing raw individual-level data. We illustrate their utility through a cross-biobank analysis of smoking and 97 cancer subtypes using data from the NIH All of Us (AOU) and the Michigan Genomics Initiative (MGI). MATERIALS AND METHODS: Large-scale biobanks often follow heterogeneous recruitment strategies and store data in separate cloud-based platforms, making centralized algorithms infeasible. To address this, we propose two decentralized sequential estimators namely, Sequential Pseudo-likelihood (SPL) and Sequential Augmented Inverse Probability Weighting (SAIPW) that leverage external population-level information to adjust for selection bias, with valid variance estimation. SAIPW additionally protects against misspecification of the selection model using flexible machine learning based auxiliary outcome models. We compare SPL and SAIPW with the existing Sequential Unweighted (SUW) estimator and with centralized and meta learning extensions of IPW and AIPW in simulations under both correctly specified and misspecified selection mechanisms. We apply the methods to harmonized data from MGI ( n = 50,935) and AOU ( n = 241,563) to estimate smoking-cancer associations. RESULTS: In simulations, SUW exhibited substantial bias and poor coverage. SPL and SAIPW yielded unbiased estimates with valid coverage probabilities under correct model specification, with SAIPW remaining robust under selection model misspecification. Both approaches showed no notable efficiency loss relative to centralized methods. Meta-learning methods were efficient for large sites but failed in settings with small cohort sizes and rare outcome prevalence. In real-data analysis, strong associations were consistently identified between smoking and cancers of the lung, bladder, and larynx, aligning with established epidemiological evidence. CONCLUSION: Our framework enables valid, privacy-enhancing inference across EHR cohorts with heterogeneous selection, supporting scalable, decentralized research using real-world data.

Journal Article↗

Genetic algorithm learning as a robust approach to RNA editing site prediction.

BACKGROUND: RNA editing is one of several post-transcriptional modifications that may contribute to organismal complexity in the face of limited gene complement in a genome. One form, known as C --> U editing, appears to exist in a wide range of organisms, but most instances of this form of RNA editing have been discovered serendipitously. With the large amount of genomic and transcriptomic data now available, a computational analysis could provide a more rapid means of identifying novel sites of C --> U RNA editing. Previous efforts have had some success but also some limitations. We present a computational method for identifying C --> U RNA editing sites in genomic sequences that is both robust and generalizable. We evaluate its potential use on the best data set available for these purposes: C --> U editing sites in plant mitochondrial genomes. RESULTS: Our method is derived from a machine learning approach known as a genetic algorithm. REGAL (RNA Editing site prediction by Genetic Algorithm Learning) is 87% accurate when tested on three mitochondrial genomes, with an overall sensitivity of 82% and an overall specificity of 91%. REGAL's performance significantly improves on other ab initio approaches to predicting RNA editing sites in this data set. REGAL has a comparable sensitivity and higher specificity than approaches which rely on sequence homology, and it has the advantage that strong sequence conservation is not required for reliable prediction of edit sites. CONCLUSION: Our results suggest that ab initio methods can generate robust classifiers of putative edit sites, and we highlight the value of combinatorial approaches as embodied by genetic algorithms. We present REGAL as one approach with the potential to be generalized to other organisms exhibiting C --> U RNA editing.

Algorithms↗