Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Integration of single-cell transcriptomics and genomic mutation analysis identifies an immunotherapy-resistant tumor subcluster and validates ARNTL2 as a malignant driver in lung adenocarcinoma.

BACKGROUND: Immunotherapy resistance in lung adenocarcinoma (LUAD) remains a critical clinical challenge, and the mechanisms underlying resistance-associated intratumoral heterogeneity are poorly characterized. METHODS: We performed single-cell RNA sequencing of LUAD patients receiving neoadjuvant immunotherapy (responders vs. non-responders), integrating inferCNV, GSVA, and differential expression analyses. Cluster-specific genes were validated across seven independent cohorts (TCGA-LUAD, GSE13213, GSE26939, GSE29016, GSE30219, GSE31210, GSE42127). A multi-algorithm machine learning framework was used to construct a prognostic model, and the immune microenvironment was characterized using TCIA scoring, seven infiltration algorithms, and ESTIMATE. ARNTL2 function was assessed by CCK-8 and Transwell assays in A549 and H1299 cells. RESULTS: Non-responders showed significant enrichment of epithelial cells, depletion of cytotoxic T/NK cells, and elevated copy number variation burden versus responders (p < 0.0001). A resistance-enriched malignant subcluster (Cluster 2) exhibited hyperproliferative and metabolic reprogramming signatures with upregulated KRT17, S100A2, and CST6, which showed tumor-specific overexpression, adverse prognostic value, and genomic amplification across cohorts. CoxBoost combined with survivalSVM achieved optimal predictive performance (C-index = 0.686), yielding robust risk stratification (HR: 2.54-10.51, all p < 0.05). Low-risk patients showed greater immune infiltration and higher TCIA immunophenoscores. ARNTL2 was an independent prognostic factor (HR: 2.07-4.64) strongly correlated with risk score (r = 0.69), and its knockdown suppressed proliferation and invasion in both LUAD cell lines (all p < 0.05). CONCLUSION: This study identifies a resistance-associated malignant subcluster in LUAD, constructs a validated CoxBoost + survivalSVM prognostic model with robust immune stratification, and establishes ARNTL2 as a core oncogenic driver and therapeutic target.

ARNTL2↗

Can host genetics transform the sustainable control of tropical theileriosis? Insights from the Tick-Theileria interface.

Tropical theileriosis, caused by the tick-transmitted apicomplexan parasite Theileria annulata, remains a major constraint on cattle production across North Africa, the Mediterranean basin, the Middle East and South Asia. Current control depends on acaricides, the theilericidal drug buparvaquone and live attenuated schizont vaccines, but acaricide resistance, buparvaquone-resistance mutations and the logistical demands of vaccination are eroding the sustainability of these tools. Host genetics offers a complementary and durable alternative. Indigenous Bos indicus breeds are consistently more resistant to ticks and tolerate T. annulata infection better than exotic Bos taurus cattle, and this advantage has a measurable heritable component. Unlike previous reviews, which treat tick resistance, T. annulata immunobiology and livestock genomic selection as separate subjects, we integrate all three and assess host genetics specifically against the failure modes of current control. We review the tick, parasite and host interface, the evidence for natural resistance, and the genetic and immunological mechanisms involved, including signal-regulatory protein, bovine major histocompatibility complex class II and inflammatory pathway genes. We then assess whether genomic selection, multi-omics, machine learning and gene editing can translate these mechanisms into resistant cattle, and we weigh the biological, economic and infrastructural barriers to implementation. The evidence indicates that host genetics will not replace existing control but could reduce reliance on acaricides and chemotherapy. That contribution remains prospective rather than demonstrated: no resistance marker for T. annulata has yet been validated, prediction accuracies are moderate and transfer poorly between breeds, and no endemic production system has implemented selection for resistance.

Animals↗

ERCC2 mutations alter the genomic distribution pattern of somatic mutations and are independently prognostic in bladder cancer.

Excision repair cross-complementation group 2 (ERCC2) encodes the DNA helicase xeroderma pigmentosum group D, which functions in transcription and nucleotide excision repair. Point mutations in ERCC2 are putative drivers in around 10% of bladder cancers (BLCAs) and a potential positive biomarker for cisplatin therapy response. Nevertheless, the prognostic significance directly attributed to ERCC2 mutations and its pathogenic role in genome instability remain poorly understood. We first demonstrated that mutant ERCC2 is an independent predictor of prognosis in BLCA. We then examined its impact on the somatic mutational landscape using a cohort of ERCC2 wild-type (n&#xa0;= 343) and mutant (n&#xa0;= 39) BLCA whole genomes. The genome-wide distribution of somatic mutations is significantly altered in ERCC2 mutants, including T[C>T]N enrichment, altered replication time correlations, and CTCF-cohesin binding site mutation hotspots. We leverage these alterations to develop a machine learning model for predicting pathogenic ERCC2 mutations, which may be useful to inform treatment of patients with BLCA.

Humans↗

A corpus of GA4GH phenopackets: Case-level phenotyping for genomic diagnostics and discovery.

The Global Alliance for Genomics and Health (GA4GH) Phenopacket Schema was released in 2022 and approved by ISO as a standard for sharing clinical and genomic information about an individual, including phenotypic descriptions, numerical measurements, genetic information, diagnoses, and treatments. A phenopacket can be used as an input file for software that supports phenotype-driven genomic diagnostics and for algorithms that facilitate patient classification and stratification for identifying new diseases and treatments. There has been a great need for a collection of phenopackets to test software pipelines and algorithms. Here, we present Phenopacket Store. Phenopacket Store v.0.1.19 includes 6,668 phenopackets representing 475 Mendelian and chromosomal diseases associated with 423 genes and 3,834 unique pathogenic alleles curated from 959 different publications. This represents the first large-scale collection of case-level, standardized phenotypic information derived from case reports in the literature with detailed descriptions of the clinical data and will be useful for many purposes, including the development and testing of software for prioritizing genes and diseases in diagnostic genomics, machine learning analysis of clinical phenotype data, patient stratification, and genotype-phenotype correlations. This corpus also provides best-practice examples for curating literature-derived data using the GA4GH Phenopacket Schema.

Humans↗

De novo Genes in Plants: Origins, Mechanisms, and Functional Implications.

De novo genes originate from previously non-coding genomic regions. They provide an important source of lineage-specific innovation. In plants, these genes may contribute to adaptation, trait diversity and crop evolution. This review summarizes recent progress in plant de novo gene research. It first discusses major routes of gene birth, including transcription-first, open reading frame (ORF)-first and concurrent models. It also examines how nascent loci acquire regulatory control and enter existing biological networks. The review then summarizes their evolutionary features, including weak early constraint, rapid molecular change, restricted expression and structural refinement. It further discusses plant de novo genes involved in stress responses, seed germination, kernel dehydration, subspecies divergence, reproductive isolation and floral scent diversification. Current methods for identifying de novo genes remain limited by rapid sequence evolution, genome annotation quality, polyploidy and transposable elements. Whole-genome synteny alignment, multi-omics evidence and machine-learning approaches can improve candidate discovery. However, each method has important limitations. Finally, this review highlights key future questions in functional validation, latent coding potential in long non-coding RNAs, epigenetic activation, regulatory-network integration and crop improvement. These perspectives clarify how de novo genes shape plant adaptation and how they may be used in precision breeding and synthetic biology.

adaptive evolution↗

On preprocessing of protein sequences for neural network prediction of polyproline type II secondary structures.

Polyproline type II stretches are somewhat rare on proteins. The backbone of this secondary structural element folds to a triangular form instead of the normal alpha-helix with 3.6 residues per turn. It is a very challenging task to try to detect them computationally from protein sequence. Here, we have studied the preprocessing phase in particular, which is important for any machine learning method. Preprocessing included selection of relevant data from the Protein Data Bank and investigation of learnability properties. These properties show whether the material is suitable for neural network computing. The complexity of algorithms in connection with preprocessing was briefly considered. We found that feedforward perceptron neural networks were appropriate for the prediction of polyproline type II and also relatively efficient in this task. The problem is very difficult because of the great similarity of the two classes present in the classification. Nevertheless, neural networks were able to recognize and predict about 75% of secondary structures.

Algorithms↗

In silico prediction of the peroxisomal proteome in fungi, plants and animals.

In an attempt to improve our abilities to predict peroxisomal proteins, we have combined machine-learning techniques for analyzing peroxisomal targeting signals (PTS1) with domain-based cross-species comparisons between eight eukaryotic genomes. Our results indicate that this combined approach has a significantly higher specificity than earlier attempts to predict peroxisomal localization, without a loss in sensitivity. This allowed us to predict 430 peroxisomal proteins that almost completely lack a localization annotation. These proteins can be grouped into 29 families covering most of the known steps in all known peroxisomal pathways. In general, plants have the highest number of predicted peroxisomal proteins, and fungi the smallest number.

Amino Acid Sequence↗

Distinguishing enzyme structures from non-enzymes without alignments.

The ability to predict protein function from structure is becoming increasingly important as the number of structures resolved is growing more rapidly than our capacity to study function. Current methods for predicting protein function are mostly reliant on identifying a similar protein of known function. For proteins that are highly dissimilar or are only similar to proteins also lacking functional annotations, these methods fail. Here, we show that protein function can be predicted as enzymatic or not without resorting to alignments. We describe 1178 high-resolution proteins in a structurally non-redundant subset of the Protein Data Bank using simple features such as secondary-structure content, amino acid propensities, surface properties and ligands. The subset is split into two functional groupings, enzymes and non-enzymes. We use the support vector machine-learning algorithm to develop models that are capable of assigning the protein class. Validation of the method shows that the function can be predicted to an accuracy of 77% using 52 features to describe each protein. An adaptive search of possible subsets of features produces a simplified model based on 36 features that predicts at an accuracy of 80%. We compare the method to sequence-based methods that also avoid calculating alignments and predict a recently released set of unrelated proteins. The most useful features for distinguishing enzymes from non-enzymes are secondary-structure content, amino acid frequencies, number of disulphide bonds and size of the largest cleft. This method is applicable to any structure as it does not require the identification of sequence or structural similarity to a protein of known function.

Algorithms↗

Chemometric discrimination of unfractionated plant extracts analyzed by electrospray mass spectrometry.

Metabolic fingerprints were obtained from unfractionated Pharbitis nil leaf sap samples by direct infusion into an electrospray ionization mass spectrometer. Analyses took less than 30 s per sample and yielded complex mass spectra. Various chemometric methods, including discriminant function analysis and the machine-learning methods of artificial neural networks and genetic programming, could discriminate the metabolic fingerprints of plants subjected to different photoperiod treatments. This rapid automated analytical procedure could find use in a variety of phytochemical applications requiring high sample throughput.

Artificial Intelligence↗

Metabolic fingerprinting of salt-stressed tomatoes.

The aim of this study was to adopt the approach of metabolic fingerprinting through the use of Fourier transform infrared (FT-IR) spectroscopy and chemometrics to study the effect of salinity on tomato fruit. Two varieties of tomato were studied, Edkawy and Simge F1. Salinity treatment significantly reduced the relative growth rate of Simge F1 but had no significant effect on that of Edkawy. In both tomato varieties salt-treatment significantly reduced mean fruit fresh weight and size class but had no significant affect on total fruit number. Marketable yield was however reduced in both varieties due to the occurrence of blossom end rot in response to salinity. Whole fruit flesh extracts from control and salt-grown tomatoes were analysed using FT-IR spectroscopy. Each sample spectrum contained 882 variables, absorbance values at different wavenumbers, making visual analysis difficult and therefore machine learning methods were applied. The unsupervised clustering method, principal component analysis (PCA) showed no discrimination between the control and salt-treated fruit for either variety. The supervised method, discriminant function analysis (DFA) was able to classify control and salt-treated fruit in both varieties. Genetic algorithms (GA) were applied to identify discriminatory regions within the FT-IR spectra important for fruit classification. The GA models were able to classify control and salt-treated fruit with a typical error, when classifying the whole data set, of 9% in Edkawy and 5% in Simge F1. Key regions were identified within the spectra corresponding to nitrile containing compounds and amino radicals. The application of GA enabled the identification of functional groups of potential importance in relation to the response of tomato to salinity.

Algorithms↗

Functional genomics and proteomics in the clinical neurosciences: data mining and bioinformatics.

The goal of this chapter is to introduce some of the available computational methods for expression analysis. Genomic and proteomic experimental techniques are briefly discussed to help the reader understand these methods and results better in context with the biological significance. Furthermore, a case study is presented that will illustrate the use of these analytical methods to extract significant biomarkers from high-throughput microarray data. Genomic and proteomic data analysis is essential for understanding the underlying factors that are involved in human disease. Currently, such experimental data are generally obtained by high-throughput microarray or mass spectrometry technologies among others. The sheer amount of raw data obtained using these methods warrants specialized computational methods for data analysis. Biomarker discovery for neurological diagnosis and prognosis is one such example. By extracting significant genomic and proteomic biomarkers in controlled experiments, we come closer to understanding how biological mechanisms contribute to neural degenerative diseases such as Alzheimers' and how drug treatments interact with the nervous system. In the biomarker discovery process, there are several computational methods that must be carefully considered to accurately analyze genomic or proteomic data. These methods include quality control, clustering, classification, feature ranking, and validation. Data quality control and normalization methods reduce technical variability and ensure that discovered biomarkers are statistically significant. Preprocessing steps must be carefully selected since they may adversely affect the results of the following expression analysis steps, which generally fall into two categories: unsupervised and supervised. Unsupervised or clustering methods can be used to group similar genomic or proteomic profiles and therefore can elucidate relationships within sample groups. These methods can also assign biomarkers to sub-groups based on their expression profiles across patient samples. Although clustering is useful for exploratory analysis, it is limited due to its inability to incorporate expert knowledge. On the other hand, classification and feature ranking are supervised, knowledge-based machine learning methods that estimate the distribution of biological expression data and, in doing so, can extract important information about these experiments. Classification is closely coupled with feature ranking, which is essentially a data reduction method that uses classification error estimation or other statistical tests to score features. Biomarkers can subsequently be extracted by eliminating insignificantly ranked features. These analytical methods may be equally applied to genetic and proteomic data. However, because of both biological differences between the data sources and technical differences between the experimental methods used to obtain these data, it is important to have a firm understanding of the data sources and experimental methods. At the same time, regardless of the data quality, it is inevitable that some discovered biomarkers are false positives. Thus, it is important to validate discovered biomarkers. The validation process may be slow; yet, the overall biomarker discovery process is significantly accelerated due to initial feature ranking and data reduction steps. Information obtained from the validation process may also be used to refine data analysis procedures for future iteration. Biomarker validation may be performed in a number of ways - bench-side in traditional labs, web-based electronic resources such as gene ontology and literature databases, and clinical trials.

Animals↗

Identification and ranking of genetic and laboratory environment factors influencing a behavioral trait, thermal nociception, via computational analysis of a large data archive.

Laboratory conditions in biobehavioral experiments are commonly assumed to be 'controlled', having little impact on the outcome. However, recent studies have illustrated that the laboratory environment has a robust effect on behavioral traits. Given that environmental factors can interact with trait-relevant genes, some have questioned the reliability and generalizability of behavior genetic research designed to identify those genes. This problem might be alleviated by the identification of the most relevant environmental factors, but the task is hindered by the large number of factors that typically vary between and within laboratories. We used a computational approach to retrospectively identify and rank sources of variability in nociceptive responses as they occurred in a typical research laboratory over several years. A machine-learning algorithm was applied to an archival data set of 8034 independent observations of baseline thermal nociceptive sensitivity. This analysis revealed that a factor even more important than mouse genotype was the experimenter performing the test, and that nociception can be affected by many additional laboratory factors including season/humidity, cage density, time of day, sex and within-cage order of testing. The results were confirmed by linear modeling in a subset of the data, and in confirmatory experiments, in which we were able to partition the variance of this complex trait among genetic (27%), environmental (42%) and genetic x environmental (18%) sources.

Animals↗

Protein therapeutics: promises and challenges for the 21st century.

Recent advances in massively parallel experimental and computational technologies are leading to radically new approaches to the early phases of the drug production pipeline. The revolution in DNA microarray technologies and the imminent emergence of its analogue for proteins, along with machine learning algorithms, promise rapid acceleration in the identification of potential drug targets, and in high-throughput screens for subpopulation-specific toxicity. Similarly, advances in structural genomics in conjunction with in vitro and in silico evolutionary methods will rapidly accelerate the number of lead drug candidates and substantially augment their target specificity. Taken collectively, these advances will usher in an era of predictive medicine, which will move medical practice from reactive therapy after disease onset, to proactive prevention.

Animals↗

A development environment for predictive modelling in foods.

Waikato Environment for Knowledge Analysis (WEKA) is a comprehensive suite of Java class libraries that implement many state-of-the-art machine learning/data mining algorithms. Non-programmers interact with the software via a user interface component called the Knowledge Explorer. Applications constructed from the WEKA class libraries can be run on any computer with a web-browsing capability, allowing users to apply machine learning techniques to their own data regardless of computer platform. This paper describes the user interface component of the WEKA system in reference to previous applications in the predictive modelling of foods.

Algorithms↗

Knowledge discovery with classification rules in a cardiovascular dataset.

In this paper we study an evolutionary machine learning approach to data mining and knowledge discovery based on the induction of classification rules. A method for automatic rules induction called AREX using evolutionary induction of decision trees and automatic programming is introduced. The proposed algorithm is applied to a cardiovascular dataset consisting of different groups of attributes which should possibly reveal the presence of some specific cardiovascular problems in young patients. A case study is presented that shows the use of AREX for the classification of patients and for discovering possible new medical knowledge from the dataset. The defined knowledge discovery loop comprises a medical expert's assessment of induced rules to drive the evolution of rule sets towards more appropriate solutions. The final result is the discovery of a possible new medical knowledge in the field of pediatric cardiology.

Algorithms↗

CORA--a knowledge-based system for the analysis of case-control studies.

Carrying out a statistical analysis, the researcher is concerned with the problem of choosing an appropriate statistical technique from a large number of competing methods. Most common statistical software offer different methods for analysing the data without giving any support regarding the adequacy of a method for a particular data set. This paper outlines the main features of the computer system CORA which provides a statistical analysis of stratified contingency tables and additionally supports the researcher at the different steps of this analysis. Here, the support given by the system consists of two different aspects. On the one hand, the help system of CORA contains general information on the implemented statistical methods which can be obtained on request. On the other hand, an advice tool recommends an adequate statistical method which depends on the actual empirical case-control data to be analysed. To build up the advice tool, a set of rules being discovered by machine learning from simulation studies is integrated into the system CORA.

Case-Control Studies↗

Prediction of 'drug-likeness'.

Recent developments in combinatorial chemistry and high-throughput screening have dramatically increased the scale on which drug discovery programs are carried out. Along with these advances has come a need for automated methods of determining which compounds from a library should be synthesized and screened. These methods range from simple counting schemes to sophisticated machine learning techniques such as neural networks. While many of these methods have performed well in validation studies, the field is still in its formative stage. This paper reviews a number of computational techniques for identifying drug-like molecules and examines challenges facing the field.

Artificial Intelligence↗

Support vector machines for prediction of protein signal sequences and their cleavage sites.

Given a nascent protein sequence, how can one predict its signal peptide or "Zipcode" sequence? This is an important problem for scientists to use signal peptides as a vehicle to find new drugs or to reprogram cells for gene therapy (see, e.g. K.C. Chou, Current Protein and Peptide Science 2002;3:615-22). In this paper, support vector machines (SVMs), a new machine learning method, is applied to approach this problem. The overall rate of correct prediction for 1939 secretary proteins and 1440 nonsecretary proteins was over 91%. It has not escaped our attention that the new method may also serve as a useful tool for further investigating many unclear details regarding the molecular mechanism of the ZIP code protein-sorting system in cells.

Genetic Vectors↗