Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “single cell transcriptomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Deletion in a (T)8 microsatellite abrogates expression regulation by 3'-UTR.

A high level of genetic instability might cause mutations to accumulate in tumours. Microsatellite instability (MSI), due to defects of the DNA mismatch repair system, affects in particular repeat sequences (microsatellites) scattered throughout the genome. By scanning transcriptome databases, we found that microsatellites in the human genome are less numerous in coding DNA than in the 3'-untranslated region (UTR), known to mediate control of gene expression. By mutation analysis, we identified a 1 bp deletion in a (T)(8) microsatellite embedded in the 1801 nucleotide long 3'-UTR of CEACAM1 gene, thought to be involved in tumour onset and progression. By Lentiviral Vector- mediated gene transfer, we showed that the wild-type but not the mutated CEACAM1 3'-UTR greatly decreased transgene expression at both mRNA and protein level. Messenger RNA abundance was fully regulated by the most 3' region of CEACAM1 3'-UTR. This region includes the (T)(8) microsatellite but not any known classified regulatory element. These data show that CEACAM1 3'-UTR contains non-canonical elements contributing to mRNA regulation, among which a short repeat sequence could play a critical regulatory function. This suggests that, in cancer cells, a single mutation in a 3'-UTR short microsatellite might strongly affect gene expression.

3' Untranslated Regions↗

Use of Scots pine seedling roots as an experimental model to investigate gene expression during interaction with the conifer pathogen Heterobasidion annosum (P-type).

The root-rot fungus Heterobasidion annosum is a major pathogen of woody trees in temperate regions of the world. In this study, seedling root of Scots pine was used as an experimental model to investigate gene expression in conifer trees during challenge with H. annosum. Initial cellular and histochemical studies have established the systems and indicated the key sequence of events during the infection process. Also, to correlate histochemical observations with the time-dependent pattern of events in host gene expression, a transcriptome profiling of a selected set of host genes from a pine-root subtraction cDNA library was conducted. Differential screening of the subset of genes arrayed on nylon membrane with cDNA probes made from seedling roots infected for 1, 3, 7 and 15 days revealed a number of up-regulated genes [disease-resistance gene analog, antimicrobial peptide (AMP) gene homolog etc.] following inoculation. The results also showed strong expression of genes involved in cell defense and protein synthesis at the early stages of the infection (3-7 days) with a decline at late stages of infection (15 days). The decline in expression of key defense genes at late stages of infection correlated well with the period of vascular colonization and subsequent loss of root turgor. Northern analyses with two of the major induced genes (AMP homolog and disease-resistance gene analog) indicated a several-fold increase in host gene expression following infection. In addition, a particular single gene (thaumatin-like protein) was consistently expressed throughout the four sampling periods of the experiment. BlastX analyses revealed that the Scots-pine thaumatin-like gene shared 51-77% sequence homology with other thaumatin-like proteins in GenBank. The importance of these results in tree defense and use of conifer seedling root in host-parasite interaction in forest trees is discussed.

Amino Acid Sequence↗

The dual specificity phosphatase transcriptome of the murine thymus.

Properly regulated mitogen-activated protein (MAP) kinase activity is critical for normal thymocyte development. MAP kinases are activated by phosphorylation of tyrosine and threonine, and dual specificity phosphatases (DUSPs) can inactivate MAP kinases by dephosphorylating both tyrosine and threonine. However, a role for DUSPs in thymocyte development has not been described. In this study, we have defined the subset of DUSP genes expressed in the murine thymus, and how their expression varies in different thymocyte subsets. Of the murine DUSP genes screened that could potentially dephosphorylate MAP kinases, we found 10 transcribed in the thymus. Seven of these 10 thymic DUSPs are true MAP kinase phosphatases based on the presence of a MAP kinase binding domain and demonstrated phosphatase activity against MAP kinases. Six of the seven thymic MAP kinase phosphatases have been shown to dephosphorylate extracellular regulated kinase (ERK). Quantitative PCR analysis of thymocyte populations isolated from different developmental stages revealed significant changes in DUSP expression as thymocytes progressed through development. Specifically, DUSPs 1, 4, and 5 significantly increase in expression as cells go from small, resting CD4/CD8 double positive cells to the CD4 single positive stage. Additionally, in vitro experiments showed that DUSPs could respond to TCR signaling, as anti-CD3 stimulation of thymocytes transiently increased transcription of six of the 10 thymic DUSP genes within 30 min. Notably, the ERK-specific phosphatase DUSP5 was upregulated 43-fold within 30 min, and returned to baseline within 24 h. Overall, we have identified a subset of DUSPs that could potentially regulate ERK activation in response to TCR signals in thymocytes.

Animals↗

Moderate expression and activity of flocculins underlie the characteristic flocculation phenotype of Saccharomyces pastorianus.

Flocculation is a key technological trait in lager brewing, governing fermentation performance, yeast recovery, and beer quality. In the allo-aneuploid hybrid yeast Saccharomyces pastorianus, the genetic basis of flocculation remains poorly resolved due to its complex dual sub-genome architecture. Here, we systematically re-annotated and functionally characterized the complete FLO gene repertoire of the Group II strain CBS 1483. Thirteen FLO genes were identified, including allelic variants and a previously uncharacterized adhesin, Flo12, containing a Hyphal_reg_CWP domain instead of the canonical PA14 lectin-binding domain. Structural modeling revealed strong conservation of Ca²+-binding residues in PA14 domains, alongside repeat-region diversification likely contributing to functional variability. Using optogenetic expression in a FLO-null background, we demonstrated that SpcI-FLO9-1 and SpcI-FLO9-2_1 are the strongest drivers of flocculation, exhibiting NewFlo-like sugar sensitivity. Transcriptomic analysis during 17°P wort fermentation showed dynamic induction of these genes coinciding with flocculation onset. Surprisingly, deletion of both loci in CBS 1483 did not abolish but only delayed sedimentation in wort, accompanied by improved maltose utilization and attenuation. These findings reveal functional redundancy and compensatory mechanisms within the FLO network of lager yeast, highlighting the genetic complexity underlying flocculation, and providing a molecular framework to inform yeast selection, strain development, and optimization of the lager fermentation processes.IMPORTANCEFlocculation, the process by which yeast cells aggregate and settle, is essential for producing clear, high-quality lager beer, and for efficient yeast recovery during brewing. However, the genetic basis of this trait in lager yeast has remained poorly understood because these strains possess unusually complex hybrid genomes. In this study, we systematically identified and characterized the complete set of flocculation genes in the industrial lager yeast Saccharomyces pastorianus CBS 1483. We demonstrated that lager yeast flocculation is not controlled by a single dominant gene, but instead emerges from the combined action of several moderately active adhesion proteins that are expressed at low levels during fermentation. Surprisingly, deleting the two strongest candidate genes only delayed, rather than eliminated, sedimentation, revealing a robust compensatory network that preserves brewing performance. These findings refine the current understanding of yeast flocculation and provide a molecular framework for developing brewing strains with improved fermentation efficiency, product consistency, and flavor quality.

Saccharomyces pastorianus↗

The current and future perspective of ChickenGTEx project and its applications in precision breeding.

The Chicken Genotype-Tissue Expression (ChickenGTEx) project was established to systematically characterize the regulatory landscape of the chicken genome and to accelerate the translation of functional genomics into precision breeding. By integrating whole-genome sequencing with multi-tissue transcriptomic profiling, ChickenGTEx provides a comprehensive atlas of gene expression regulation across diverse tissues and physiological systems. Current findings demonstrate that complex production traits are governed by coordinated regulatory networks rather than isolated loci, with substantial contributions from tissue-specific gene expression, structural variation, and genotype-by-sex interactions. Sex-dependent regulatory effects further refine the genetic architecture of metabolic, immune, and reproductive traits, highlighting the importance of incorporating sex as a biological variable in genomic analyses. Application of integrative omics frameworks within elite layer populations has revealed multilayer regulatory mechanisms underlying extended laying performance, feed efficiency, metabolic health, and eggshell quality. By partitioning phenotypic variance into genetic, regulatory, and host-microbiome components, these approaches move beyond association-based mapping toward causal inference and biological interpretation. Importantly, validated regulatory loci identified through ChickenGTEx and related analyses provide actionable markers for genomic selection and rational targets for precision genome modification. Looking forward, continued expansion of regulatory atlases, incorporation of single-cell and longitudinal data in diverse environmental conditions, and integration of functional annotation into breeding pipelines will further enhance prediction accuracy and sustainable genetic improvement. The ChickenGTEx project thus represents a foundational platform bridging functional genomics and practical poultry breeding.

Animals↗

High-throughput functional genomic methods to analyze the effects of dietary lipids.

The applications of 'omics' (genomics, transcriptomics, proteomics and metabolomics) technologies in nutritional studies have opened new possibilities to understand the effects and the action of different diets both in healthy and diseased states and help to define personalized diets and to develop new drugs that revert or prevent the negative dietary effects. Several single nucleotide polymorphisms have already been investigated for potential gene-diet interactions in the response to different lipid diets. It is also well-known that besides the known cellular effects of lipid nutrition, dietary lipids influence gene expression in a tissue, concentration and age-dependent manner. Protein expression and post-translational changes due to different diets have been reported as well. To understand the molecular basis of the effects and roles of dietary lipids high-throughput functional genomic methods such as DNA- or protein microarrays, high-throughput NMR and mass spectrometry are needed to assess the changes in a global way at the genome, at the transcriptome, at the proteome and at the metabolome level. The present review will focus on different high-throughput technologies from the aspects of assessing the effects of dietary fatty acids including cholesterol and polyunsaturated fatty acids. Several genes were identified that exhibited altered expression in response to fish-oil treatment of human lung cancer cells, including protein kinase C, natriuretic peptide receptor-A, PKNbeta, interleukin-1 receptor associated kinase-1 (IRAK-1) and diacylglycerol kinase genes by using high-throughput quantitative real-time PCR. Other results will also be mentioned obtained from cholesterol and polyunsaturated fatty acid fed animals by using DNA- and protein microarrays.

Animals↗

DNA microarrays in pediatric cancer.

Childhood cancer, like all cancer, is at heart a genetic disease. Consequently, fundamental understanding of the oncogenic process is likely to be beneficially addressed by genetic methodology. Current methods have largely focused on single-gene defects, like chimeric genes, which are present in many sarcomas and leukemias. Real understanding is more likely to derive from a genome-wide analysis of these malignancies. Recent technologic advances have made it possible to simultaneously assess the entire expressed gene profile, or transcriptome, of a given cancer. Foremost among these methods is gene expression profiling using DNA microarrays. Two basic approaches predominate: spotted arrays and photolithography arrays. Regardless of the method, the resulting information can be used to create disease profiles, but only if appropriate bioinformatic solutions are employed. Common analytic approaches include two-way expression comparisons, or scatter analyses; outlier gene analysis, to identify significantly dysregulated genes; dendrogram analyses, as pioneered by Eisen; cluster analyses to identify diagnostic or biologic groups; and various forms of functional analyses to identify relevant genes and biologic pathways. Studies of both adult and pediatric cancer have demonstrated the feasibility of such analyses to identify both diagnostic and prognostic groups of tumors. Acute childhood leukemias have been grouped into myelogenous and lymphoid, and even B- and T-cell subsets. Breast cancer prognostic groups have been identified on the basis of a small subset of expressed genes. In addition, preliminary data on childhood sarcomas appear to identify both diagnostic and prognostic subsets. Specifically, embryonal rhabdomyosarcoma could be distinguished from alveolar rhabdomyosarcoma, and even morphologically mixed embryonal and alveolar rhabdomyosarcoma showed similar gene expression profiles in both histologies. Further, collaborative studies using clustering analyses appear to identify prognostic groups of diverse sarcomas. Larger institutional and cooperative group studies are currently underway to validate these preliminary findings.

Animals↗

Identifying cytotoxic T cell epitopes from genomic and proteomic information: "The human MHC project.".

Complete genomes of many species including pathogenic microorganisms are rapidly becoming available and with them the encoded proteins, or proteomes. Proteomes are extremely diverse and constitute unique imprints of the originating organisms allowing positive identification and accurate discrimination, even at the peptide level. It is not surprising that peptides are key targets of the immune system. It follows that proteomes can be translated into immunogens once it is known how the immune system generates and handles peptides. Recent advances have identified many of the basic principles involved. The single most selective event is that of peptide binding to MHC, making it particularly important to establish accurate descriptions and predictions of peptide binding for the most common MHC variants. These predictions should be integrated with those of other steps involved in antigen processing, as these become available. The ability to translate the accumulating primary sequence databases in terms of immune recognition should enable scientists and clinicians to analyze any protein of interest for the presence of potentially immunogenic epitopes. The computational tools to scan entire proteomes should also be developed, as this would enable a rational approach to vaccine development and immunotherapy. Thus, candidate vaccine epitopes might be predicted from the various microbial genome projects, tumor vaccine candidates from mRNA expression profiling of tumors ("transcriptomes") and auto-antigens from the human genome.

Antigen Presentation↗

Gene expression profiles of T lymphocytes are sensitive to the influence of heavy smoking: A pilot study.

Cigarette smoke components have a proven negative influence on human health. Adverse metabolic effects were observed in tissues and single cells. T lymphocytes get in contact with affected organs (e.g., lung) or cells (e.g., erythrocytes), as well as with smoke components and bioactive molecules, whose production is triggered by tobacco smoke. We therefore compared the gene expression profiles in these cells of the adaptive immune system of three male heavy smokers and three male nonsmokers using rapid T cell isolation and Affymetrix GeneChip HG U133A 2.0 microarray analysis. Eighty-eight genes were found to be significantly (t test) differentially expressed by a factor of 1.5-fold or larger between smokers and nonsmokers. Using the gene function groups of the gene ontology consortium to categorize the functions of the differentially expressed genes, the group termed "response to stimulus" was found to be most significantly affected by smoking. Our data indicate a prominent role of cytotoxic T lymphocytes in response to smoking. Several genes that are typically expressed in these cells were found regulated although the ratio of cytotoxic and helper T lymphocytes remained unchanged in smokers. Our data show that, in principle, it might be possible to identify health-related biomarkers in the transcriptome of T lymphocytes.

Adult↗

EWS::WT1 Isoform-Dependent Regulation of Neogenes in Desmoplastic Small Round Cell Tumors.

Desmoplastic small round cell tumor (DSRCT) is a rare, aggressive sarcoma characterized by the pathognomonic EWS::WT1 fusion protein (FP), an oncogenic chimeric transcription factor (OCTF) resulting from the t(11;22)(p13;q12) translocation. Recent studies have identified "neogenes" (NGs), genes normally silent in normal tissues but transcriptionally activated by OCTFs, as potential tumor-specific markers in fusion-driven cancers. In this study, we investigated the expression and regulation of DSRCT-specific NGs (DSRCT_NGs) using multimodal data across different cohorts of patients, PDX, and cell line data. We evaluated bulk and single-nucleus RNA sequencing of patient specimens from MD Anderson Cancer Center, revealing the robust ability for DSRCT_NGs to distinguish FP-positive DSRCT from samples failing detection of the EWS::WT1 FP. To elucidate the regulatory role of the EWS::WT1 FP in driving NG expression, we performed knockdown experiments in four DSRCT cell lines. This consistently resulted in a reduction of DSRCT_NG expression. Isoform-specific expression of EWS::WT1 in LP9 and MeT-5A mesothelial cells revealed that the E-KTS isoform of EWS::WT1 predominantly drives DSRCT_NG expression. Mechanistically, ATAC-seq and ChIP-seq analyses demonstrated that EWS::WT1 directly binds to accessible chromatin regions near NG transcription start sites, enriched for WT1 motifs and active histone marks. Integration of Hi-ChIP data further revealed that EWS::WT1 facilitates long-range enhancer-promoter looping at DSRCT_NG loci, promoting the expression of nearby genes. Collectively, these findings establish DSRCT_NGs as direct transcriptional outputs of the EWS::WT1 FP and implicate their loci as regulatory regions of the DSRCT transcriptome. Their fusion-dependent expression, chromatin accessibility, and promoter-enhancer connectivity underscore their potential utility as highly specific biomarkers and therapeutic targets in DSRCT.

DSRCT↗

Gene discovery and annotation using LCM-454 transcriptome sequencing.

454 DNA sequencing technology achieves significant throughput relative to traditional approaches. More than 261,000 ESTs were generated by 454 Life Sciences from cDNA isolated using laser capture microdissection (LCM) from the developmentally important shoot apical meristem (SAM) of maize (Zea mays L.). This single sequencing run annotated >25,000 maize genomic sequences and also captured approximately 400 expressed transcripts for which homologous sequences have not yet been identified in other species. Approximately 70% of the ESTs generated in this study had not been captured during a previous EST project conducted using a cDNA library constructed from hand-dissected apex tissue that is highly enriched for SAMs. In addition, at least 30% of the 454-ESTs do not align to any of the approximately 648,000 extant maize ESTs using conservative alignment criteria. These results indicate that the combination of LCM and the deep sequencing possible with 454 technology enriches for SAM transcripts not present in current EST collections. RT-PCR was used to validate the expression of 27 genes whose expression had been detected in the SAM via LCM-454 technology, but that lacked orthologs in GenBank. Significantly, transcripts from approximately 74% (20/27) of these validated SAM-expressed "orphans" were not detected in meristem-rich immature ears. We conclude that the coupling of LCM and 454 sequencing technologies facilitates the discovery of rare, possibly cell-type-specific transcripts.

Base Sequence↗

Revealing the Shared Genetic Architecture of Metabolic Dysfunction-Associated Steatotic Liver Disease-Related Traits Through Genomic Structural Equation Modeling.

Although individual traits related to metabolic dysfunction-associated steatotic liver disease (MASLD) have been investigated through large-scale genome-wide association studies (GWASs), the shared genetic susceptibility across these traits remains unclear. We therefore conducted a multivariate GWAS of key MASLD-related traits to elucidate their common genetic architecture. We applied genomic structural equation modeling to model a latent genetic factor (MASLD-F) underlying genetically correlated MASLD-related traits, leveraging their GWAS-derived genetic correlations. We then performed functional annotations, including fine-mapping, transcriptome-wide association study, and cell- and tissue-type-specific enrichment analyses, and conducted Mendelian randomization analyses to identify modifiable risk factors. Our multivariate MASLD-F GWAS identified 50 independent variants across 48 genomic loci. Transcriptomic imputation identified several MASLD-F-associated genes, including ARNTL, NPC1, BTBD10, VDAC2, TSKU, SFMBT1, and ABHD17C. We observed significant enrichment of MASLD-F-related genetic signals predominantly in brain tissues, pancreatic islets, and the adrenal gland. Additionally, six modifiable risk factors and four modifiable protective factors for MASLD-F were identified. These findings reveal a complex shared genetic architecture underlying MASLD components, thereby expanding our understanding of disease pathogenesis and providing novel insights for precision medicine and public health interventions.

Humans↗

Comprehensive analysis of differential gene expression profiles on D-galactosamine-induced acute mouse liver injury and regeneration.

Microarray analysis of RNA from d-galactosamine (GalN)-administered mouse livers was performed to establish a global gene expression profile during injury and regeneration stages at two different doses. A single dose of GalN at 266 or 26.6 mg/kg body weight was given intraperitoneally, and the liver samples were obtained after 6, 24, and 72 h. Histopathologic studies enabled the classification of the D-galactosamine effect into injury (6, 24 h) and regeneration (72 h) stages. By using the Applied Biosystems mouse genome survey microarray, a total of 7267 out of 33,315 (21.8%) genes were found to be statistically reliable at p<0.05 by two-way ANOVA, and 1469 (4.4%) probes at false discovery rate <5% by significance analysis of microarray. Among the statistically reliable clones by both analytical methods, 389 genes were differentially expressed when compared with non-treated control, with more than a 1.625-fold difference (which equals 0.7 in log(2) scale) at one or more GalN treatment conditions and with less than 1.625-fold difference at all three vehicle-treated conditions. Three hundred thirty six genes and 13 genes were identified as injury- and regeneration-specific genes, respectively, showing that most of the transcriptomic changes were seen during the injury stage. Furthermore, multiple genes involved in protein synthesis and degradation, mRNA processing and binding, and cell cycle regulation showed variable transcript levels upon acute GalN administration.

Acute Disease↗

Enterocutaneous Fistula-Associated Sepsis and Mortality: Development and Validation of a Multimodal Artificial Intelligence Prediction Model.

BACKGROUND: Predicting enterocutaneous fistula (ECF)-associated sepsis and mortality poses significant challenges in digital health care due to the disease's complexity and heterogeneous clinical manifestations. Current approaches that rely on single-modal data or traditional scoring systems often fail to capture the intricate immune-inflammatory dynamics and multisystem involvement in patients with ECF. OBJECTIVE: This study aims to develop an artificial intelligence (AI)-driven multimodal fusion model integrating clinical, imaging, and transcriptomic data for early prediction of ECF-associated sepsis and 28-day mortality, addressing the limitations of conventional single-dimensional models. METHODS: This study leveraged publicly available datasets (Medical Information Mart for Intensive Care III [MIMIC-III], electronic Intensive Care Unit [eICU], and The Cancer Genome Atlas) to construct a multimodal framework. Clinical parameters were processed using Extreme Gradient Boosting, abdominal imaging features were extracted via convolutional neural networks, and transcriptomic profiles were analyzed with variational autoencoders. A Transformer-based fusion network was employed for joint prediction and validated through cross-validation and external testing. Key features were identified using Shapley Additive Explanations and Local Interpretable Model-Agnostic Explanations interpretability algorithms, while immune regulatory mechanisms were explored via weighted gene co-expression network analysis. RESULTS: The multimodal model achieved an area under the curve (AUC) of 0.89 for predicting sepsis and 28-day mortality, outperforming unimodal models (clinical-only model, AUC 0.72, and imaging-only model, AUC 0.78). Critical predictors included Sequential Organ Failure Assessment score, lactate levels, intra-abdominal free fluid on imaging, and immunoregulatory genes (programmed death-ligand 1 [PD-L1] and indoleamine 2,3-dioxygenase 1 [IDO1]). Mechanistic analysis revealed distinct immune reprogramming in patients with sepsis, characterized by increased regulatory T cells and M2 macrophages, along with downregulated cluster of differentiation 8+ (CD8+) T cells. CONCLUSIONS: This multimodal AI model offers an innovative digital solution in medical informatics, enabling precise early risk stratification for ECF-associated sepsis. By integrating multisource data and providing interpretable insights into immune-inflammatory pathways, the model enhances health care quality for patients with ECF and paves the way for personalized intervention strategies.

Humans↗

DEQOR: a web-based tool for the design and quality control of siRNAs.

RNA interference (RNAi) is a powerful tool for inhibiting the expression of a gene by mediating the degradation of the corresponding mRNA. The basis of this gene-specific inhibition is small, double-stranded RNAs (dsRNAs), also referred to as small interfering RNAs (siRNAs), that correspond in sequence to a part of the exon sequence of a silenced gene. The selection of siRNAs for a target gene is a crucial step in siRNA-mediated gene silencing. According to present knowledge, siRNAs must fulfill certain properties including sequence length, GC-content and nucleotide composition. Furthermore, the cross-silencing capability of dsRNAs for other genes must be evaluated. When designing siRNAs for chemical synthesis, most of these criteria are achievable by simple sequence analysis of target mRNAs, and the specificity can be evaluated by a single BLAST search against the transcriptome of the studied organism. A different method for raising siRNAs has, however, emerged which uses enzymatic digestion to hydrolyze long pieces of dsRNA into shorter molecules. These endoribonuclease-prepared siRNAs (esiRNAs or 'diced' RNAs) are less variable in their silencing capabilities and circumvent the laborious process of sequence selection for RNAi due to a broader range of products. Though powerful, this method might be more susceptible to cross-silencing genes other than the target itself. We have developed a web-based tool that facilitates the design and quality control of siRNAs for RNAi. The program, DEQOR, uses a scoring system based on state-of-the-art parameters for siRNA design to evaluate the inhibitory potency of siRNAs. DEQOR, therefore, can help to predict (i) regions in a gene that show high silencing capacity based on the base pair composition and (ii) siRNAs with high silencing potential for chemical synthesis. In addition, each siRNA arising from the input query is evaluated for possible cross-silencing activities by performing BLAST searches against the transcriptome or genome of a selected organism. DEQOR can therefore predict the probability that an mRNA fragment will cross-react with other genes in the cell and helps researchers to design experiments to test the specificity of esiRNAs or chemically designed siRNAs. DEQOR is freely available at http://cluster-1.mpi-cbg.de/Deqor/deqor.html.

Internet↗

Alternative tandem transcription initiation links noncoding variants to human disease through translational control.

Alternative tandem transcription initiation is a pervasive mechanism of gene regulation, yet its genetic impact on human disease remains largely unknown. Here, we systematically quantify the genetic regulation of alternative tandem transcription initiation across 25,859 samples from 49 normal human&#xa0;tissues and 33 tumor tissues. We identify approximately 0.4 million genetic variants associated with alternative transcription initiation in 5295 genes, with 32% operating independently of gene expression. Moreover, we discover 2238 multi-tissue alternative tandem transcription initiation outliers enriched for rare deleterious promoter and 5' UTR variants, demonstrating that both common and rare variants modulate transcription initiation. Strikingly, 74% of disease variants that colocalize with genetic variants regulating alternative transcription initiation cannot be identified through expression quantitative trait loci. Transcriptome-wide association studies identify 614 disease susceptibility genes associated with alternative transcription initiation, including known cancer drivers such as MAFF and MLLT10. Functional validation uncovers OSGEP as a breast cancer risk gene, where the alternative allele lengthens the 5' UTR and reduces protein abundance through upstream open reading frame-mediated translation repression, and suppresses breast cancer cell proliferation. Our findings establish alternative transcription initiation as a major, underappreciated mechanism associating noncoding variation with disease, providing a critical resource for interpreting disease risk loci.

Humans↗

Transcriptomic analysis of extensive changes in metabolic regulation in Kluyveromyces lactis strains.

Genome-wide analysis of transcriptional regulation is generally carried out on well-characterized reference laboratory strains; hence, the characteristics of industrial isolates are therefore overlooked. In a previous study on the major cheese yeast Kluyveromyces lactis, we have shown that the reference strain and an industrial strain used in cheese making display a differential gene expression when grown on a single carbon source. Here, we have used more controlled conditions, i.e., growth in a fermentor with pH and oxygen maintained constant, to study how these two isolates grown in glucose reacted to an addition of lactose. The observed differences between sugar consumption and the production of various metabolites, ethanol, acetate, and glycerol, correlated with the response were monitored by the analysis of the expression of 482 genes. Extensive differences in gene expression between the strains were revealed in sugar transport, glucose repression, ethanol metabolism, and amino acid import. These differences were partly due to repression by glucose and another, yet-unknown regulation mechanism. Our results bring to light a new type of K. lactis strain with respect to hexose transport gene content and repression by glucose. We found that a combination of point mutations and variation in gene regulation generates a biodiversity within the K. lactis species that was not anticipated. In contrast to S. cerevisiae, in which there is a massive increase in the number of sugar transporter and fermentation genes, in K. lactis, interstrain diversity in adaptation to a changing environment is based on small changes at the level of key genes and cell growth control.

Acetates↗

Foxi2 and Sox3 are master regulators controlling ectoderm germ layer specification.

In vertebrates, germ layer specification represents a critical transition where pluripotent cells acquire lineage-specific identities. We identify the maternal transcription factors Foxi2 and Sox3 to be pivotal master regulators of ectodermal germ layer specification in Xenopus. Ectopic co-expression of Foxi2 and Sox3 in prospective endodermal tissue induces the expression of ectodermal markers while suppressing mesendodermal markers. Transcriptomics analyses reveal that Foxi2 and Sox3 jointly and independently regulate hundreds of ectodermal target genes. During early cleavage stages, Foxi2 and Sox3 pre-bind to key cis-regulatory modules (CRMs), marking sites that later recruit Ep300 and facilitate H3K27ac deposition, thereby shaping the epigenetic landscape of the ectodermal genome. These CRMs are highly enriched within ectoderm-specific super-enhancers (SEs). Our findings highlight the pivotal role of ectodermal SE-associated CRMs in precise and robust ectodermal gene activation, establishing Foxi2 and Sox3 as central architects of ectodermal lineage specification.

Ep300↗