Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Cell annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Pathway proteomics: global and focused approaches.

Biological pathways represent the relationships (reactions and interactions) between biological molecules in the context of normal cellular functions and disease mechanisms. Understanding the roles of proteins and signaling pathways expressed within disease, and their link to drug discovery and drug development are central in today's target-driven pharmaceutical processes. This article gives an overview of proteomics strategies, including global expression analysis as well as focused approaches using multidimensional separation by both gel- and liquid-phase techniques linked to mass spectrometry, as applied to two of the pathways involved in inflammatory diseases. In primary human cell studies, our group has annotated and identified thousands of proteins using both electrospray ionization and matrix-assisted laser desorption ionization (MALDI)-sequencing technology. Annotations made from gel images and chromatography fractionation, interfaced to high-end mass spectrometry sequence and structure identity, are cornerstones in cutting-edge protein expression profiling. Regarding phosphorylation mechanisms of kinases, the quantitative stoichiometry can be determined using affinity probe isolations. Another strategy involves micro-preparative sample processing, which has been used to analyze single-target phosphoproteins and their relative phospho-stoichiometry.

Electrophoresis, Gel, Two-Dimensional↗

Multiomic study of cutaneous T-cell lymphoma reveals single-cell clonal evolution in progression and therapy resistance.

Cutaneous T-cell lymphoma (CTCL) remains a challenging disease due to its significant heterogeneity, therapy resistance, and relentless progression. Multiomics technologies offer the potential to provide uniquely precise views of disease progression and response to therapy. Here, we present a comprehensive multiomics view of CTCL clonal evolution, incorporating exome, whole-genome, epigenome, bulk, single-cell T-cell receptor, and single-cell RNA sequencing of 99 clinically annotated serial skin, peripheral blood, and lymph node samples from 34 patients with CTCL. We leveraged this extensive data set to define the molecular underpinnings of CTCL progression in individual patients at single-cell resolution with the goal of identifying clinically useful biomarkers and therapeutic targets. Our studies identified recurrent progression-associated clonal genomic alterations; we highlight mutation of CCR4, phosphoinositide 3-kinase inhibitor signaling, and programmed cell death protein 1 (PD-1) checkpoint pathways as evasion tactics deployed by malignant T cells. We identified a gain-of-function mutation in STAT3 (D661Y) and demonstrated, using cleavage under targets and release using nuclease (CUT&RUN) and RNA sequencing, that it enhances binding to and transcription of genes in Rho GTPase pathways. With our previous work implicating this pathway in histone deacetylase inhibitor-resistant CTCL, these data provide further support for a previously unrecognized role for Rho GTPase pathway dysregulation in CTCL progression. Recurrent progression-associated mutations were common in the epigenetic modifier EZH2, suggesting that EZH2 inhibition may benefit patients with CTCL. Our findings support an approach in which genomic analysis is widely used for improved disease monitoring, biomarker-informed clinical trial design, and genome-guided therapeutic decision-making. Moreover, these molecular changes present new opportunities for therapeutic targeting in this challenging and incurable cancer.

Multiomics↗

Genome-Wide Identification and Characterization of the TBL Gene Family and Temporal Expression Dynamics During Powdery Mildew Infection in Cucumber (Cucumis sativus).

Cell-wall polysaccharide O-acetylation contributes to cell-wall assembly, organ development, and plant-pathogen interactions, but the cucumber TBL gene family remains poorly characterized. Here, 37 CsTBL genes were identified genome-wide and analyzed using phylogenetic, syntenic, conserved-motif, gene-structure, promoter, protein-structure, Gene Ontology, and transcriptome approaches, followed by RT-qPCR analysis after powdery mildew inoculation. All CsTBL proteins contained the conserved GDS and DxxH motifs, whereas accessory motifs and predicted structural features varied among clades. Intraspecific analysis identified dispersed, WGD/segmental, and tandem duplication categories, and cross-species synteny was more extensive with melon than with Arabidopsis. Homology-derived annotations associated CsTBL genes with cell-wall polysaccharide metabolism, Golgi/endomembrane compartments, and O-acetyltransferase activity, including six genes assigned to xylan O-acetyltransferase-related annotations. Expression profiling revealed tissue- and developmental-stage-dependent patterns, whereas the publicly available powdery mildew RNA-seq dataset provided descriptive temporal expression profiles in Podosphaera xanthii-inoculated samples. Independent RT-qPCR analysis using time-matched mock controls revealed distinct post-inoculation responses among six selected genes. Relative to the corresponding mock controls, CsTBL2 was consistently repressed; CsTBL15 showed transient induction at 1 dpi followed by repression; CsTBL24 exhibited a biphasic response; CsTBL25 was induced at all sampled post-inoculation time points; CsTBL26 showed progressive induction; and CsTBL30 reached its highest observed expression level at 3 dpi. Integrated functional annotation and expression evidence highlighted CsTBL26 as a priority candidate for further functional characterization, while CsTBL24 and CsTBL25 represented fruit-associated candidates with distinct powdery mildew responses; CsTBL30 remained an additional strongly infection-responsive candidate. These findings provide an evolutionary and expression-based framework for the functional characterization of the cucumber TBL gene family.

O-acetylation↗

Protein analysis by mass spectrometry and sequence database searching: a proteomic approach to identify human lymphoblastoid cell line proteins.

Lymphoblastoid cell lines correspond to in vitro EBV-immortalized lymphocyte B-cells. These cells display a suitable model for experiments dealing with changes in protein expression occurring upon B-cell differentiation, after drug treatment, or after inhibition of some transcription factors. For all these reasons we have undertaken an effort aimed at developing a hematopoietic cell line protein two-dimensional electrophoresis (2-DE) database, containing B-lymphoblastoid 2-DE maps. In this work, matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF-MS) peptide mass fingerprinting analysis was adopted for protein identification. The peptide mass fingerprinting identification and the sequence coverage obtained on colloidal Coomassie blue (CBB) stained gel was close to that obtained using zinc-imidazole staining. Everything considered, CBB being more comfortable for subsequent spot manipulations, CBB staining was chosen for identification of a larger number of polypeptides. The results suggest that reticulation of the gel can interfere preventing the uptake of the enzyme during the in-gel digestion step. Consequently, low molecular mass proteins appear more difficult to identify by mass fingerprinting. Finally, the information provided in this study allows the construction of a new annoted reference map of human lymphoblastoid cell proteins. Among the identified proteins 60% were not yet positioned on 2-DE maps in three of the most important well-documented databases. The annoted map will be accessible via Internet on the LBPP server at URL:http:// www-smbh.univ-paris13.fr/lbtp/index.htm.

Acrylic Resins↗

Shared genetic basis and spatial cellular atlas of psoriasis and metabolic syndrome.

BACKGROUND: Psoriasis (PS) and metabolic syndrome (MetS) frequently co-occur. Characterizing their shared genetic architecture and spatially enriched cellular populations may clarify the context of their co-occurrence and generate hypotheses for functional validation. METHODS: We integrated genome-wide association study (GWAS) summary statistics for PS, MetS, and five related components with spatially resolved single-cell transcriptomic data. Global and local genetic correlations were assessed using linkage disequilibrium score regression, genetic covariance analysis, high-definition likelihood, and local analysis of variant association. A bivariate causal mixture model quantified polygenic overlap. Conditional/conjunctional false discovery rate and composite-null pleiotropy analyses identified shared susceptibility loci. Finally, gsMap evaluated trait-associated enrichment across annotated embryonic tissues at single-cell resolution. RESULTS: Genetic approaches identified significant genome-wide correlations and polygenic sharing between PS, MetS, and its components. Local and cross-trait analyses identified region-specific signals and cross-validated shared loci. gsMap revealed trait-specific tissue enrichment. PS showed the strongest enrichment in the epidermis (pCauchy = 1.0573 × 10  - ⁴), adipose tissue (pCauchy = 1.5366 × 10 - ⁴), and liver (pCauchy = 1.0167 × 10 - ³). Across MetS, FBG, HDL-C, hypertension, and TG, enriched regions mainly involved the liver, adipose tissue, and epidermis. WC enrichment was predominantly observed in adipose tissue (pCauchy = 1.7823 × 10 - ⁴), with no significant liver or epidermal enrichment. CONCLUSION: Integrating GWAS with single-cell transcriptomic and spatial information characterized shared genetic architecture between PS and MetS-related phenotypes and their spatial enrichment patterns. These findings provide a framework for generating testable hypotheses about comorbidity biology and guiding future functional and clinical validation.

Psoriasis↗

Cellular response to gene expression profiles of different hepatitis C virus core proteins in the Huh-7 cell line with microarray analysis.

In the current investigation, we constructed recombinants of expression of HCV genotype 1b, 2a, and 4d whole core proteins and established a human hepatoma (Huh-7) cell line which expressed different core proteins constitutively. In the Affymetrix human gene chip, the HG-U133 A and B were employed for identification of variant core protein gene expression in the Huh-7 cell line. In data analysis, we applied a threshold that eliminated all genes that were not increased or decreased by at least a 3-fold change in a comparison between transfected cells and control cells. All of these genes were annotated by using NetAffx analysis through the Affymetrix website and categorized on the basis of their biological processes. The microarray analysis result suggested that the gene expression profiles caused by three kinds of core proteins were mainly shown in metabolism, signal transduction, protease activity, immune responses, etc. and that some pathogenesis/oncogenesis, apoptosis, or anti-apoptosis gene expression were up/down-regulated simultaneously in the Huh-7 cell line. In conclusion, the gene expression profiles of variant core proteins were implicated in HCV replication, pathogenesis, or oncogenesis in the Huh-7 cell line, which is useful for our understanding of HCV variant core protein biological function and its pathogenic mechanism.

Carcinoma, Hepatocellular↗

Functional annotation of IFN-alpha-stimulated gene expression profiles from sensitive and resistant renal cell carcinoma cell lines.

The antiproliferative, antiviral, and immunomodulatory properties of interferons (IFNs) have led to its therapeutic implementation. IFNs effects are mediated by a complex network of signal transducers, culminating in IFN-stimulated gene (ISG) induction. This complexity leads to diverse clinical responses to IFN, from no response to complete regression of disease. Elucidation of ISG induction patterns is, therefore, essential to understand and maximize its therapeutic potential. To correlate ISG expression profiles with IFN responsiveness, two renal cell carcinoma (RCC) cell lines differing in antiviral and apoptotic response to IFN were treated with IFN-alpha for different times, and expression profiles were analyzed using a customized microarray containing 850 unique putative ISGs. Genes with similar kinetics of induction in both cell lines were clustered and analyzed for gene function. Seven sets of coordinately regulated genes were identified by k-means cluster analysis, and significant functional similarities were identified for five of the seven sets. Strikingly, expression of genes associated with transcription temporally preceded expression of those involved in signal transduction. Enhanced antiviral sensitivity to IFN was coincident with sustained expression of ISGs involved in transcriptional regulation. However, no difference in Stat1 activation was observed between the cell lines. Analysis of ISG expression patterns suggests that subtle differences in transcription profiles contribute to differences in IFN responsiveness.

Antineoplastic Agents↗

Identifying independent causal cell types for human diseases and risk variants.

The SNP-heritability of human diseases is extremely enriched in candidate regulatory elements (cREs) from disease-relevant cell types. Critical next steps are to understand whether these enrichments are driven by multiple causal cell types and whether individual variants impact disease risk via a single or multiple of cell types. Here, we propose CT-FM and CT-FM-SNP, 2 methods accounting for cREs shared across cell types to identify independent sets of causal cell types for a trait and its candidate causal variants, respectively. We applied CT-FM to 63 GWAS summary statistics (average N = 417K) using 924 cRE annotations, primarily from ENCODE4. CT-FM inferred 79 sets of causal cell types, with corresponding SNP-annotations explaining 39.0 ± 1.8% of trait SNP-heritability. It identified 14 traits with independent causal cell types, uncovering previously unexplored cellular mechanisms in height, schizophrenia and autoimmune diseases. We applied CT-FM-SNP to 39 UK Biobank traits and predicted high-confidence causal cell types for 3,091 candidate causal non-coding SNPs-trait pairs. Our results suggest that most SNPs affect a phenotype via a single set of cell types, whereas pleiotropic SNPs might target different cell types depending on the phenotype context. Altogether, CT-FM and CT-FM-SNP shed light on how genetic variants act collectively and individually at the cellular level to affect disease risk.

Journal Article↗

A two-dimensional electrophoresis proteomic reference map and systematic identification of 1367 proteins from a cell suspension culture of the model legume Medicago truncatula.

The proteome of a Medicago truncatula cell suspension culture was analyzed using two-dimensional electrophoresis and nanoscale HPLC coupled to a tandem Q-TOF mass spectrometer (QSTAR Pulsar i) to yield an extensive protein reference map. Coomassie Brilliant Blue R-250 was used to visualize more than 1661 proteins, which were excised, subjected to in-gel trypsin digestion, and analyzed using nanoscale HPLC/MS/MS. The resulting spectral data were queried against a custom legume protein database using the MASCOT search engine. A total of 1367 of the 1661 proteins were identified with high rigor, yielding an identification success rate of 83% and 907 unique protein accession numbers. Functional annotation of the M. truncatula suspension cell proteins revealed a complete tricarboxylic acid cycle, a nearly complete glycolytic pathway, a significant portion of the ubiquitin pathway with the associated proteolytic and regulatory complexes, and many enzymes involved in secondary metabolism such as flavonoid/isoflavonoid, chalcone, and lignin biosynthesis. Proteins were also identified from most other functional classes including primary metabolism, energy production, disease/defense, protein destination/storage, protein synthesis, transcription, cell growth/division, and signal transduction. This work represents the most extensive proteomic description of M. truncatula suspension cells to date and provides a reference map for future comparative proteomic and functional genomic studies of the response of these cells to biotic and abiotic stress.

Amino Acid Sequence↗

Molecular control of the oocyte to embryo transition.

The elucidation of the molecular control of the initiation of mammalian embryogenesis is possible now that the transcriptomes of the full-grown oocyte and two-cell stage embryo have been prepared and analysed. Functional annotation of the transcriptomes using gene ontology vocabularies, allows comparison of the oocyte and two-cell stage embryo between themselves, and with all known mouse genes in the Mouse Genome Database. Using this methodology one can outline the general distinguishing features of the oocyte and the two-cell stage embryo. This, when combined with oocyte-specific targeted deletion of genes, allows us to dissect the molecular networks at play as the differentiated oocyte and sperm transit into blastomeres with unlimited developmental potential.

Animals↗

The rat liver epithelial (RLE) cell protein database.

Computer databases of rat liver epithelial (RLE) cellular polypeptides have been established using high resolution two-dimensional gel electrophoresis and computer-assisted analysis. Databases have been constructed utilizing both [35S]methionine- and [32P]orthophosphate-labeled as well as silver-stained polypeptides from normal RLE cells. The RLE database, which contains both qualitative and quantitative annotations, includes experiments with normal, chemically and oncogene transformed as well as spontaneously transformed cell lines. A total of 2537 [35S]methionine-labeled polypeptides from whole cell lysates (1920 acidic and 617 basic, separated in the first dimension using isoelectric focusing and nonequilibrium pH gradient electrophoresis, respectively) were analyzed and databases constructed using the Elsie 5 gel analysis system. To increase the "viewing window" and hence the usefulness of the RLE database, subcellular fractionation of whole cell preparations was performed and high resolution two-dimensional maps of the individual subcellular components were constructed. Databases representing 1229 cytosolic, 1539 acidic and 674 basic nuclear, 1746 membrane-associated, 415 mitochondrial, 773 in vitro translated and 350 phosphoproteins were established from these maps. The RLE databases contain the Elsie 5 identification number, protein name (if known), molecular weight and pI information, quantitative and spot shape data, and specific information regarding transformation-sensitive, growth-related (exponentially proliferating versus confluent) cell populations as well as those polypeptides modulated by specific growth factors. The RLE databases represent initial efforts toward the establishment of comprehensive databases of rat liver proteins and serve as a vital resource for on-going as well as future studies regarding the regulation of growth and differentiation as well as transformation of RLE cells.

Animals↗

Investigating cross-organism prediction of prokaryotic essential proteins using unsupervised language model and ensemble strategy.

Cross-organism prediction of essential proteins is a critical task for drug discovery and microbial engineering, yet the generalizability of existing machine learning models across diverse species remains a significant challenge. In this study, we propose DeepPEP, a large language model-based framework designed to reliably transfer essential protein annotations between distantly related organisms. Utilizing 66 curated prokaryotic datasets, we systematically evaluated DeepPEP's cross-organism performance under various conditions. Initial pairwise predictions revealed a correlation between performance and evolutionary distance; however, further investigation demonstrated that integrating training data from multiple organisms yields superior predictive power. In a benchmark scenario designed to simulate real-world applications, DeepPEP outperformed the state-of-the-art tool Geptop 2.0, showcasing a robust ability to identify species-specific essential proteins. Finally, a case study on novel genomes confirmed the model's practical effectiveness. Our results suggest that DeepPEP is a powerful strategy for prokaryotic essential protein prediction, and the rigorous evaluation framework established in this study provides a new benchmark for the field.

Large Language Models↗

Comprehensive proteomic analysis of breast cancer cell membranes reveals unique proteins with potential roles in clinical cancer.

Proteins associated with cancer cell plasma membranes are rich in known drug and antibody targets as well as other proteins known to play key roles in the abnormal signal transduction processes required for carcinogenesis. We describe here a proteomics process that comprehensively annotates the protein content of breast tumor cell membranes and defines the clinical relevance of such proteins. Tumor-derived cell lines were used to ensure an enrichment for cancer cell-specific plasma membrane proteins because it is difficult to purify cancer cells and then obtain good membrane preparations from clinical material. Multiple cell lines with different molecular pathologies were used to represent the clinical heterogeneity of breast cancer. Peptide tandem mass spectra were searched against a comprehensive data base containing known and conceptual proteins derived from many public data bases including the draft human genome sequences. This plasma membrane-enriched proteome analysis created a data base of more than 500 breast cancer cell line proteins, 27% of which were of unknown function. The value of our approach is demonstrated by further detailed analyses of three previously uncharacterized proteins whose clinical relevance has been defined by their unique cancer expression profiles and the identification of protein-binding partners that elucidate potential functionality in cancer.

Amino Acid Sequence↗

ChickGCE: a novel germ cell EST database for studying the early developmental stage in chickens.

We established a database to study germ cells during the early developmental stage in the chicken. The ChickGCE database provides integrated expressed sequence tag (EST) data from chicken testis, ovary, embryonic gonads, and primordial germ cells. We gathered data on 10,294 ESTs from approximately 1000 embryonic gonads, and we experimentally determined 10,851 ESTs from primordial germ cells purified from 7955 embryonic gonads by magnetically activated cell sorting. The EST testis and ovary datasets were retrieved from the public database of The Institute for Genomic Research (TIGR). The EST data were clustered and assembled into unique sequences, contigs, and singletons. The ChickGCE database provides functional annotation, identification, and putative embryonic germ-cell-specific novel transcripts based on the Gene Ontology database, as well as statistical analyses of expression patterns and pair-wise comparisons of two types of tissue- and germ-cell-specific alternative splicing events in the chicken. The new database is accessible online and queries can be answered using several search options, including tissue database searches, keywords, clone IDs, expected values, and BLAST search scores.

Animals↗

Identification and phenotypic characterization of the cell-division protein CdpA.

Analysis of the automated computer annotation of the early draft phase genome of Lactobacillus acidophilus NCFM revealed the previously discovered S-layer gene slpA and an additional partial ORF with weak similarities to S-layer proteins. The entire gene was sequenced to reveal a 1799-bp gene coding for 599 amino acids with a calculated molecular mass of 64.8 kDa. No transcription or translation signals could be determined in close proximity to the 5'-region. However, a strong putative terminator with a free energy of -16.84 kcal/mol was identified directly downstream of the gene. A PSI-Blast analysis showed similarities to members of S-layer proteins, cell-wall associated proteinases and hexosyl-transferases. Calculation of an unrooted phylogenetic tree with other examples of S-layer proteins and proteinases placed the deduced protein separately from both groups. A derivative of L. acidophilus NCFM was constructed by targeted integration into the gene. SDS-PAGE analysis of non-covalently linked proteins of the cell wall of the mutant, compared to the wild type, revealed the loss of a cell-surface protein. Phenotypic analyses of the mutant revealed significant changes in cell morphology, altered responses to various environmental stresses, and lowered cell adhesion. Based on the in silico and functional analyses, we ascertained that this protein plays a role in cell-wall processing during the growth and cell-cell separation and designated the gene as cell-division protein, cdpA.

Bacterial Adhesion↗

Rare variant contribution to the heritability of coronary artery disease.

Whole genome sequences (WGS) enable discovery of rare variants which may contribute to missing heritability of coronary artery disease (CAD). To measure their contribution, we apply the GREML-LDMS-I approach to WGS of 4949 cases and 17,494 controls of European ancestry from the NHLBI TOPMed program. We estimate CAD heritability at 34.3% assuming a prevalence of 8.2%. Ultra-rare (minor allele frequency ≤ 0.1%) variants with low linkage disequilibrium (LD) score contribute ~50% of the heritability. We also investigate CAD heritability enrichment using a diverse set of functional annotations: i) constraint; ii) predicted protein-altering impact; iii) cis-regulatory elements from a cell-specific chromatin atlas of the human coronary; and iv) annotation principal components representing a wide range of functional processes. We observe marked enrichment of CAD heritability for most functional annotations. These results reveal the predominant role of ultra-rare variants in low LD on the heritability of CAD. Moreover, they highlight several functional processes including cell type-specific regulatory mechanisms as key drivers of CAD genetic risk.

Humans↗

Flower proteome: changes in protein spectrum during the advanced stages of rose petal development.

Flowering is a unique and highly programmed process, but hardly anything is known about the developmentally regulated proteome changes in petals. Here, we employed proteomic technologies to study petal development in rose (Rosa hybrida). Using two-dimensional polyacrylamide gel electrophoresis, we generated stage-specific (closed bud, mature flower and flower at anthesis) petal protein maps with ca. 1,000 unique protein spots. Expression analyses of all resolved protein spots revealed that almost 30% of them were stage-specific, with ca. 90 protein spots for each stage. Most of the proteins exhibited differential expression during petal development, whereas only ca. 6% were constitutively expressed. Eighty-two of the resolved proteins were identified by mass spectrometry and annotated. Classification of the annotated proteins into functional groups revealed energy, cell rescue, unknown function (including novel sequences) and metabolism to be the largest classes, together comprising ca. 90% of all identified proteins. Interestingly, a large number of stress-related proteins were identified in developing petals. Analyses of the expression patterns of annotated proteins and their corresponding RNAs confirmed the importance of proteome characterization.

Flowers↗