Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “unsupervised clustering”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

High throughput quantitative analysis of serum proteins using glycopeptide capture and liquid chromatography mass spectrometry.

It is expected that the composition of the serum proteome can provide valuable information about the state of the human body in health and disease and that this information can be extracted via quantitative proteomic measurements. Suitable proteomic techniques need to be sensitive, reproducible, and robust to detect potential biomarkers below the level of highly expressed proteins, generate data sets that are comparable between experiments and laboratories, and have high throughput to support statistical studies. Here we report a method for high throughput quantitative analysis of serum proteins. It consists of the selective isolation of peptides that are N-linked glycosylated in the intact protein, the analysis of these now deglycosylated peptides by liquid chromatography electrospray ionization mass spectrometry, and the comparative analysis of the resulting patterns. By focusing selectively on a few formerly N-linked glycopeptides per serum protein, the complexity of the analyte sample is significantly reduced and the sensitivity and throughput of serum proteome analysis are increased compared with the analysis of total tryptic peptides from unfractionated samples. We provide data that document the performance of the method and show that sera from untreated normal mice and genetically identical mice with carcinogen-induced skin cancer can be unambiguously discriminated using unsupervised clustering of the resulting peptide patterns. We further identify, by tandem mass spectrometry, some of the peptides that were consistently elevated in cancer mice compared with their control littermates.

Animals↗

A software suite for the generation and comparison of peptide arrays from sets of data collected by liquid chromatography-mass spectrometry.

There is an increasing interest in the quantitative proteomic measurement of the protein contents of substantially similar biological samples, e.g. for the analysis of cellular response to perturbations over time or for the discovery of protein biomarkers from clinical samples. Technical limitations of current proteomic platforms such as limited reproducibility and low throughput make this a challenging task. A new LC-MS-based platform is able to generate complex peptide patterns from the analysis of proteolyzed protein samples at high throughput and represents a promising approach for quantitative proteomics. A crucial component of the LC-MS approach is the accurate evaluation of the abundance of detected peptides over many samples and the identification of peptide features that can stratify samples with respect to their genetic, physiological, or environmental origins. We present here a new software suite, SpecArray, that generates a peptide versus sample array from a set of LC-MS data. A peptide array stores the relative abundance of thousands of peptide features in many samples and is in a format identical to that of a gene expression microarray. A peptide array can be subjected to an unsupervised clustering analysis to stratify samples or to a discriminant analysis to identify discriminatory peptide features. We applied the SpecArray to analyze two sets of LC-MS data: one was from four repeat LC-MS analyses of the same glycopeptide sample, and another was from LC-MS analysis of serum samples of five male and five female mice. We demonstrate through these two study cases that the SpecArray software suite can serve as an effective software platform in the LC-MS approach for quantitative proteomics.

Amino Acid Sequence↗

Variation in the hepatic gene expression in individual male Fischer rats.

A new tool beginning to have wider application in toxicology studies is transcript profiling using microarrays. Microarrays provide an opportunity to directly compare transcript populations in the tissues of chemical-exposed and unexposed animals. While several studies have addressed variation between microarray platforms and between different laboratories, much less effort has been directed toward individual animal differences especially among control animals where RNA samples are usually pooled. Estimation of the variation in gene expression in tissues from untreated animals is essential for the recognition and interpretation of subtle changes associated with chemical exposure. In this study hepatic gene expression as well as standard toxicological parameters were evaluated in 24 rats receiving vehicle only in 2 independent experiments. Unsupervised clustering demonstrated some individual variation but supervised clustering suggested that differentially expressed genes were generally random. The level of hepatic gene expression under carefully controlled study conditions is less than 1.5-fold for most genes. The impact of individual animal variability on microarray data can be minimized through experimental design.

Animals↗

Gene expression profile of cytokines and chemokines in microdissected primary Hodgkin and Reed-Sternberg (HRS) cells: high expression of interleukin-11 receptor alpha.

We microdissected Hodgkin and Reed-Sternberg (HRS) cells from 14 Hodgkin's lymphoma tissue samples (nodular sclerosis = 5; mixed cellularity = 9), and after isolation and amplification of mRNA, analyzed the expression profile of 140 genes of chemokines, cytokines and their receptors by cDNA microarray methods. We also compared the profile with those of germinal center (GC) cells in reactive lymphadenitis. Unsupervised clustering revealed a relatively homogeneous expression profile in HRS cells. HRS cells tended to express mainly Th2 T cell-associated molecules rather than those of Th1, compared with GC cells. Interleukin-11 receptor alpha (IL-11Ralpha), a previously unknown HRS cell-specific gene, was detected in addition to known genes. Immunohistochemical staining confirmed the expression of IL-11Ralpha at the protein level. In contrast, only few cases were positive for IL-11Ralpha in B cell lymphoma, diffuse large cell lymphoma and follicular lymphoma. This is the first analysis report of tissue HRS cells with cDNA microarray technique.

Adolescent↗

Integrating explainable artificial intelligence with multiomics systems biology and electronic health record data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health records data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; 9 tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct subtissues (defined as clusters of samples within a brain tissue that share a specific expression pattern); and gene-gene coexpression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six Food and Drug Administration (FDA)-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large US de-identified insurance-claims database (n&#x2009;=&#x2009;364&#xa0;733), exposure to promethazine, one of the candidate drugs, was associated with a 57%-62% lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both P&#x2009;<&#x2009;.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multiomics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Alzheimer Disease↗

A comprehensive set of protein complexes in yeast: mining large scale protein-protein interaction screens.

MOTIVATION: The analysis of protein-protein interactions allows for detailed exploration of the cellular machinery. The biochemical purification of protein complexes followed by identification of components by mass spectrometry is currently the method, which delivers the most reliable information--albeit that the data sets are still difficult to interpret. Consolidating individual experiments into protein complexes, especially for high-throughput screens, is complicated by many contaminants, the occurrence of proteins in otherwise dissimilar purifications due to functional re-use and technical limitations in the detection. A non-redundant collection of protein complexes from experimental data would be useful for biological interpretation, but manual assembly is tedious and often inconsistent. RESULTS: Here, we introduce a measure to define similarity within collections of purifications and generate a set of minimally redundant, comprehensive complexes using unsupervised clustering. AVAILABILITY: Programs and results are freely available from http://www.bork.embl-heidelberg.de/Docu/purclust/

Algorithms↗

Cellular diversity in mouse neocortex revealed by multispectral analysis of amino acid immunoreactivity.

Cortical cells were classified using an unsupervised cluster analysis based upon their quantitative and combinatorial immunoreactivity for glutamate, gamma-aminobutyric acid (GABA), aspartate, glutamine and taurine. Overall, cell class-specific amino acid signatures were found for 12 cellular types; seven GABA-immunoreactive (GABA-IR) populations (GABA1--7), three classes containing high glutamate levels (GLUT1--3) and two putative glial (GLIA1, 2) cell types. From their large somata, associated vertical processes and high glutamate content, the GLUT classes most probably correspond to pyramidal neurons. Two of the GLUT classes demonstrated complementary distributions in different cortical layers, suggesting spatial separation of cells differing in amino acid immunoreactivity. Of the seven GABA classes, two comprised cells with large somata and displayed medium to low glutamate levels. On the basis of size, these two populations may correspond to large basket cell interneurons. Glial populations could be divided into two classes: GLIA1 cells were more frequently associated with blood vessels and GLIA2 cells were more commonly seen in the lower cortical layers. This work demonstrates that signature recognition based upon amino acid content can be used to separate cortical cells into different categories and reveal further subclasses within these categories. This approach is complementary to other methods using physiological and molecular tools and ultimately will enhance our understanding of neuronal heterogeneity.

Amino Acids↗

Cortical sources of CRF, NKB, and CCK and their effects on pyramidal cells in the neocortex.

In order to investigate how neuropeptide transmission can modulate the neocortical network, we mapped the expression of neurokinin (NK) B, cholecystokinin (CCK), and corticotropin-releasing factor (CRF) and their receptors to neuronal types using patch-clamp and single-cell reverse transcription-polymerase chain reaction in acute slices of rat neocortex. Classification of neurons by unsupervised clustering based on the analysis of multiple electrophysiological and molecular properties disclosed 3 GABAergic interneuron clusters and 1 pyramidal cell cluster. The 3 neuropeptides were expressed in a cluster of interneurons characteristically expressing vasoactive intestinal peptide. CRF was additionally found in a cluster containing almost exclusively somatostatin-expressing interneurons, whereas CCK was present in all clusters. The respective receptors of these peptides, NK-3, CCK-B, and CRF-1, were essentially expressed in pyramidal cells. At -60 mV, pyramidal cells were weakly depolarized by each of these peptides. When pyramidal neurons were maintained to about 5 mV below spike threshold, depolarization induced by each peptide resulted in a long-lasting action potential discharge. Neuropeptide effects were prevented by selective antagonists of NK-3, CCK-B, and CRF-1 receptors. These results suggest that pyramidal neurons are the primary target of NKB, CCK, and CRF in the neocortex. They further indicate that specific interneuron types coordinate the release of these peptides and can induce a long-lasting increase of the excitability of the neocortical network.

Action Potentials↗

PreDigs: A Database of Context-specific Cell Type Markers and Precise Cell Subtypes for Digestive Cell Annotation.

Research on cell type markers helps investigators explore the diverse cellular composition of gastrointestinal tumors, thereby enhancing our understanding of tumor heterogeneity and its impact on disease progression and treatment response. However,&#xa0;the integration of large-scale datasets and the standardization of cell type identification remain challenging. Here, we developed PreDigs, a user-friendly database of predicted signatures for the digestive system, which offers 124 curated single-cell RNA sequencing datasets, covering over 3.4 million cells, all available for download. After unsupervised clustering, we unified the identification and nomenclature of cell subtype labels, constructing a cell ontology tree with 142 cell types across 8 hierarchical levels. Meanwhile, we calculated three different context-specific cell type markers, including "Cell Markers", "Subtype Markers", and "TPN Markers", based on various application requirements within or across tissues. Through the integrated analysis of PreDigs data, we identified distinct cell subpopulations exclusive to tumors, one of which corresponds to tumor-specific endothelial cells. Additionally, PreDigs offers online cell annotation tools, allowing users to classify single cells with greater flexibility. PreDigs is accessible at https://www.biosino.org/predigs/.

Humans↗

DNA copy number patterns reveal prognostic markers and elucidate mechanisms of evolution in IDH-mutant astrocytoma.

BACKGROUND: Current literature suggestsisocitrate dehydrogenase (IDH)-mutant astrocytoma contains several molecular subgroups. In this study, we are interested in determining the connection between different molecular subgroups with grade and/or survival. METHODS: A cohort of 470 Mayo Clinic adult patients (&#x2265;18 years, 56.2% male) with primary IDH-mutant astrocytoma diagnosed by World Health Organization (WHO) 2021 criteria were examined. Results were validated in an independent cohort of 614 Mayo Clinic Neuropathology consult patients and 235 The Cancer Genome Atlas (TCGA) patients. RESULTS: The Mayo Clinic Practice cohort confirmed the association of CDKN2A/B deletion with overall survival (OS, homozygous vs hemizygous vs intact, 2.7 vs 9.6 vs 17.2 years, P&#x2009;<&#x2009;.001). Phosphatase and tensin homolog (PTEN) deletion was also associated with poor OS (7.3 vs 17.4 years, P&#x2009;<&#x2009;.001). Increased number of copy number alterations was associated with OS (continuous variable, HR&#x2009;=&#x2009;1.027, P&#x2009;<&#x2009;.001). Carrying one or more copies of the germline risk allele at rs55705857 was associated with earlier age of onset (median age 33 vs 35 years, P&#x2009;=&#x2009;.01), and a shorter OS after adjusting for age, grade, sex and treatment (HR&#x2009;=&#x2009;1.81, P&#x2009;=&#x2009;.007). The Mayo Clinic Neuropathology Consult cohort and TCGA were utilized to validate age of onset and survival, respectively. Unsupervised clustering of the copy number alterations identified several clinically significant groups that may define pathways to disease progression. Losses of chromosomes 11p, 13q, 1p, and 10q were all associated with reduced overall survival in the Mayo Clinic cohort. CONCLUSIONS: Patients with hemizygous loss of CDKN2A/B, loss of PTEN, increased number of copy number alterations, specific chromosomal arm losses or rs55705857 germline risk allele have reduced overall survival.

Humans↗

Identification of a unique gene expression signature that differentiates hepatocellular adenoma from well-differentiated hepatocellular carcinoma.

It is often difficult to distinguish hepatocellular adenoma (HCA) from well-differentiated hepatocellular carcinoma (WDHCC) when limited tissue from a needle biopsy is evaluated. The aim of this study was to identify gene expression patterns that can distinguish HCA from WDHCC, with the ultimate goal of discovering novel diagnostic markers. Gene expression profile analysis was performed using Affymetrix U133Plus2 GeneChip microarrays on RNA isolated from frozen tissue of 6 HCA and 8 WDHCC specimens. Statistical analysis of microarray data identified 63 genes whose expression levels were significantly different between HCA and WDHCC. These included 57 genes overexpressed by HCA and 6 overexpressed by WDHCC. Eight genes were chosen for further analysis by quantitative RT-PCR on RNA derived from archived, paraffin-embedded tissue blocks of an independent validation set comprising 9 HCAs and 9 HCCs. Seven of the 8 genes demonstrated average expression differences between HCA and HCC that were concordant with the microarray findings, and their expression pattern correctly classified the 18 tumors into HCA and HCC using unsupervised clustering analysis. Furthermore, immunohistochemical staining performed on a third, independent set of 27 HCAs and 33 HCCs confirmed the expression differences at protein levels for 5 of the genes. Taken together, our data demonstrate significant molecular differences between HCA and WDHCC, despite their morphologic similarity. More importantly, we have identified a unique set of genes whose expression pattern can discriminate between these two types of hepatocellular neoplasms, suggesting the possibility of future development of ancillary molecular and immunohistochemical diagnostic methods.

Adenoma, Liver Cell↗

Distinctive gene expression profiles by cDNA microarrays in endometrioid and serous carcinomas of the endometrium.

Endometrial carcinomas are classified by their morphology into two major subtypes. Endometrioid carcinomas (type I) are generally estrogen dependent, well-differentiated, superficially invasive, and have a good outcome. Serous carcinomas (type II) are hormone independent, frequently deeply invasive and widely metastatic, and have a poor prognosis. Microarray technology and analysis allows us to determine if the global gene expression profiles of these two subtypes correlate with their morphologic phenotype. Fresh tissue from 18 endometrial carcinomas was studied: 7 well-, 2 moderately, and one poorly differentiated endometrioid, 4 serous carcinomas, and 4 high-grade mixed endometrioid-serous carcinomas. Labeled cDNA probes were synthesized (Cy5 for tumor, Cy3 for reference) and applied to microarrays containing 18,098 cDNA clones or ESTs. A pool of equal amounts of total RNA from each tumor served as the reference RNA. By unsupervised cluster analysis, the endometrioid carcinomas clustered together and were separate from the serous carcinomas. The high-grade mixed carcinomas clustered with the serous carcinomas. Using a statistical algorithm based on gene expression pattern and conducting a supervised analysis of the two defined groups, we have identified 315 genes that statistically differentiate type I from type II endometrial carcinomas. In addition to corroborating the predicted overexpression of known markers (e.g., ras and catenin in endometrioid carcinomas), the cDNA microarray technique has revealed novel alterations in gene expression relevant to cell cycle, cell adhesion, signal transduction, apoptosis, and tumor progression not previously implicated in endometrial carcinomas. For serous carcinomas, these include aldolase, desmoplakin, integrin-linked kinase, PKC, and metallopeptidase. In conclusion, the gene expression profiles of type I and type II endometrial carcinomas are different. Refinement of these profiles will permit more accurate diagnostic tumor classification and the development of prognosis assays.

Carcinoma, Endometrioid↗

Metabolism pathway-based subtyping in pancreatic adenocarcinoma: an integrated study by bulk RNA-sequence and machine learning algorithms.

BACKGROUND: Pancreatic adenocarcinoma (PAAD) is highly aggressive, and its tumor microenvironment has significant metabolic and immune microenvironment complexity and genomic instability. In this study, by integrating the metabolic pathway activity score and clinical data, we constructed a novel risk assessment model to reveal the unique biological behavior and clinical significance behind different PAAD subtypes. METHODS: In this study, the transcriptome and clinical data of TCGA and GSE57495 databases were integrated to explore the interaction between metabolic pathways. Based on unsupervised clustering analysis of pathway activity and survival prognosis, patients with PAAD were classified into metabolic subtypes with significant prognostic differences. Subsequently, we assessed the heterogeneity of these subtypes in terms of clinical outcomes, genomic characteristics, and immune microenvironment composition. Based on the differentially expressed genes (DEGs) among metabolic subtypes, a clinical prognostic risk model and nomogram were constructed, which were double-validated by GSE57495-independent cohort and GSE57495&#xa0;+&#xa0;TCGA-PAAD combined cohort. Finally, the correlations between risk scores (RSs) and signaling pathway activity and tumor immune microenvironment characteristics were evaluated. RESULTS: Based on metabolic pathway correlation and prognostic information, 240 patients in the TCGA-PAAD and GSE57495 datasets were divided into three subgroups. There were significant differences between subgroups in gene expression, pathway activity, clinical prognosis, and immune infiltration characteristics among the subtypes. Using machine learning algorithms, an RS model was constructed from DEGs among the subgroups, with the random forest method showing the best performance. A nomogram integrating the RS and clinical indicators demonstrated excellent predictive accuracy for 1-, 3-, and 5-year survival rates, confirming the RS as an independent prognostic factor. High- and low-risk groups exhibited significant differences in immune infiltration, pathway activity, and gene mutations. Drug sensitivity analysis showed that the high-risk group was more sensitive to AZD6244, ABT737, and other drugs. CONCLUSION: This study stratified patients with PAAD into three subgroups based on metabolic pathways and prognostic information, revealing significant differences in clinical outcomes, immune characteristics, and genetic mutations. The robust RS model developed from these findings demonstrated strong predictive power for patient survival and identified promising therapeutic strategies, providing valuable insights for advancing precision medicine in PAAD.

immune microenvironment↗

A multifaceted investigation into the impact of m6A methylation-related genes on pancreatic cancer, integrating insights from various databases and foundational experimental research.

BACKGROUND: Despite advances in surgical techniques, immunotherapy, the mortality rate associated with pancreatic cancer (PC) has been on the rise in recent years. Understanding the importance of RNA N6-methyladenosine (m6A) in PC is critical for prognosis, tumor microenvironment, and immunotherapy efficacy. The study aims to identify m6A methylation regulators that play an important role in the development and progression of PC by mining databases. The effect of insulin-like growth factor-binding protein 3 (IGFBP3) on pancreatic tumors was explored, and the related mechanisms were explored. METHODS: We analyzed the expression of m6A regulators in PC by digging deeper into the datasets of The Cancer Genome Atlas and Gene Expression Omnibus (GEO) databases, and analyzed its relationship with the prognosis of patients with PC, looking for m6A methylation regulators that play an important role in the development and progression of PC. Reuse the ConsensusClusterPlus package, Cox analysis, and unsupervised clustering to delineate three distinct m6A clusters - designated as m6A cluster A, m6A cluster B, and m6A cluster C single-sample gene set enrichment analysis, gene set variation analysis, Gene Ontology, and Kyoto Encyclopedia of Genes and Genomes (KEGG) analyses evaluated the different pathway roles of these clusters in the development and progression of PC. Finally, the cell lines with IGFBP3 overexpression and knockdown were constructed by lentivirus transfection, the transfection effect was identified by WB, and the effects of IGFBP3 overexpression/knockdown on the survival and growth of PC cell lines were verified by cell cloning experiments and cell counting kit-8 experiments, and the possible related pathways were explored by KEGG. RESULTS: Most m6A regulatory factors are highly expressed in PC, and their high expression is negatively correlated with the prognosis of patients with PC. Furthermore, m6A regulatory factors may influence the occurrence and development of PC through metabolic pathways, stroma activation pathways, immune regulatory processes, and the immune microenvironment. Finally, the overexpression of IGFBP3 promoted the growth of PC cells, and vice versa. CONCLUSIONS: Most m6A regulatory factors are differentially expressed in PC and are associated with the prognosis of patients with PC, potentially influencing the occurrence and development of PC through pathways such as the immune microenvironment. The overexpression of IGFBP3 can promote the growth of PC cells and vice versa.

IGFBP3↗

Utility of Plasma Cell-free Chromatin Immunoprecipitation to Detect Cardiac Allograft Rejection.

BACKGROUND: Antibody-mediated rejection (AMR) remains the major risk factor for allograft loss across all solid organ transplantation. Unfortunately, its diagnosis relies on biopsy, an invasive gold standard that often sample unaffected allograft tissue leading to missed diagnosis. Plasma donor-derived cell-free DNA (dd-cfDNA) is noninvasive biomarker that has high sensitivity but low specificity for AMR diagnosis. This proof-of-concept study assessed the utility of cell-free chromatin immunoprecipitation (cfChIP) as a surrogate for gene expression to detect cardiac AMR and the associated pathobiology. METHODS: The discovery GRAfT multicenter cohort of heart transplant patients (NCT02423070) identified AMR, acute cellular rejection (ACR), and stable controls based on biopsy and ddcfDNA results. Plasma cfChIP-sequencing was performed to identify peaks, associated genes and pathobiological pathways. Plasma from an external cohort (GTD, NCT01985412) was also analyzed to verify pathways identified. Digital droplet PCR (ddPCR) assays targeting differential regions were constructed to test the diagnostic performance of cfDNA to detect AMR/ACR from stable controls (rejection-specific assays) or AMR from ACR (AMR-specific assays). RESULTS: The cohort included 21 AMR, 28 ACR, and 45 stable controls from GRAfT and GTD, and 23 healthy controls. cfChIP detected expected active genes, including housekeeping genes and gene targets of transplant immunosuppressive drugs but not inactive genes. Unsupervised clustering of the discovery GRAfT cohort assigned 95% of samples correctly as AMR, ACR or stable control. Differential analysis identified pathobiological pathways of AMR such as neutrophil degranulation and complement activation. The pathways were consistent in GTD samples. Rejection-specific assays detected AMR/ACR from controls with AUC of 0.78 - 0.95. AMR-specific assays detected AMR from ACR with AUC of 0.71 - 0.85, sensitivities of 0.73 - 0.94 and specificities of 0.73 - 0.80. CONCLUSION: This study provides valuable preliminary data supporting the use of cfChIP to detect AMR and the associated pathobiological pathways.

Allograft rejection↗

Integrating explainable AI with multiomics systems biology and EHR data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health record (EHR) data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; nine tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations (SHAP) identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct "subtissues" (clusters of samples); and gene-gene co-expression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six FDA-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large U.S. de-identified insurance-claims database (n = 364733), exposure to promethazine, one of the candidate drugs, was associated with a 57-62 % lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both p < 0.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multi-omics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Computational Biology↗

A single-cell transcriptomic atlas of the pigtail macaque placenta in late gestation.

The placenta is a complex organ with multiple immune and non-immune cell types that promote fetal tolerance and facilitate the transfer of nutrients and oxygen. The nonhuman primate (NHP) is a key experimental model for studying human pregnancy complications, in part due to similarities in placental structure, which makes it essential to understand how single-cell populations compare across the human and NHP maternal-fetal interface. We constructed a single-cell RNA-Seq (scRNA-Seq) atlas of the placenta from the pigtail macaque ( Macaca nemestrina ) in the third trimester, comprising three different tissues at the maternal-fetal interface: the chorionic villi (placental disc), chorioamniotic membranes, and the maternal decidua. Each tissue was separately dissociated into single cells and processed through the 10X Genomics and Seurat pipeline, followed by aggregation, unsupervised clustering, and cluster annotation. Next, we determined the maternal-fetal origins of cell populations and analyzed single-cell RNA trajectory, Gene Ontology enrichment, and cell-cell communication. Single-cell populations in the pigtail macaque were strikingly similar in their identity and frequency to those found in the human placenta, including cells from trophoblast, stromal cell, immune, and macrophage lineages. An advantage of our approach was the deep sequencing of three tissues at the maternal-fetal interface, which yielded a rich diversity of common and rare single-cell populations. The third-trimester pigtail macaque single-cell atlas enables the identification of cellular subclusters analogous to those in humans and provides a powerful resource for understanding experimental perturbations on the NHP placenta.

Journal Article↗

Systematic learning of gene functional classes from DNA array expression data by using multilayer perceptrons.

Recent advances in microarray technology have opened new ways for functional annotation of previously uncharacterised genes on a genomic scale. This has been demonstrated by unsupervised clustering of co-expressed genes and, more importantly, by supervised learning algorithms. Using prior knowledge, these algorithms can assign functional annotations based on more complex expression signatures found in existing functional classes. Previously, support vector machines (SVMs) and other machine-learning methods have been applied to a limited number of functional classes for this purpose. Here we present, for the first time, the comprehensive application of supervised neural networks (SNNs) for functional annotation. Our study is novel in that we report systematic results for ~100 classes in the Munich Information Center for Protein Sequences (MIPS) functional catalog. We found that only ~10% of these are learnable (based on the rate of false negatives). A closer analysis reveals that false positives (and negatives) in a machine-learning context are not necessarily "false" in a biological sense. We show that the high degree of interconnections among functional classes confounds the signatures that ought to be learned for a unique class. We term this the "Borges effect" and introduce two new numerical indices for its quantification. Our analysis indicates that classification systems with a lower Borges effect are better suitable for machine learning. Furthermore, we introduce a learning procedure for combining false positives with the original class. We show that in a few iterations this process converges to a gene set that is learnable with considerably low rates of false positives and negatives and contains genes that are biologically related to the original class, allowing for a coarse reconstruction of the interactions between associated biological pathways. We exemplify this methodology using the well-studied tricarboxylic acid cycle.

Algorithms↗