Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptome annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

The UTRs of Leishmania donovani vary in length and are enriched in potential regulatory structures.

Leishmania spp. regulate gene expression largely post-transcriptionally, yet untranslated regions (UTRs) remain poorly delineated. We generated high-quality genome and transcriptome datasets for Leishmania donovani strain 1S2D (Ld1S) by combining PacBio HiFi de novo assembly with Oxford Nanopore direct RNA sequencing of promastigotes and axenic amastigotes. The genome assembly consists of 65 scaffolds totaling ~33.3 Mb. Structural comparisons to LdBPK282A1 revealed numerous rearrangements, including some reshuffling genes among polycistronic transcription units and validated by polycistronic reads from RNA sequencing. Promastigote and amastigote RNA sequencing produced 469,010 and 46,729 monocistronic reads containing a spliced-leader and a polyA tail sequences, defining 8,479 transcripts and supporting 7,415 of the 7,969 annotated protein coding genes, as well as 604 putative long non-coding RNAs. We annotated UTRs for 4,921 genes and observed that putative RNA G-quadruplexes were markedly enriched in UTRs. We also noted that 31.9% and 11.5% were expressed into multiple isoforms in promastigotes and amastigotes, respectively. Collectively, these data provide a comprehensive annotation of L. donovani genes and their UTRs and reveal widespread and stage-specific UTR length polymorphisms, and, overall, points to an important role of 3' UTR in post-transcriptional regulation in L. donovani.

Journal Article↗

A dolphin peripheral blood leukocyte cDNA microarray for studies of immune function and stress reactions.

A microarray focused on stress response and immune function genes of the bottlenosed dolphin has been developed. Random expressed sequence tags (ESTs) were isolated and sequenced from two dolphin peripheral blood leukocyte (PBL) cDNA libraries biased towards T- and B-cell gene expression by stimulation with IL-2 and LPS, respectively. A total of 2784 clones were sequenced and contig analysis yielded 1343 unigenes (archived and annotated at ). In addition, 52 dolphin genes known to be important in innate and adaptive immune function and stress responses of terrestrial mammals were specifically targeted, cloned and added to the unigene collection. The set of dolphin sequences printed on a cDNA microarray comprised the 1343 unigenes, the 52 targeted genes and 2305 randomly selected (but unsequenced) EST clones. This set was printed in duplicate spots, side by side, and in two replicates per slide, such that the total number of features per microarray slide was 19,200, including controls. The dolphin arrays were validated and transcriptomic profiles were generated using PBL from a wild dolphin, a captive dolphin and dolphin skin cells. The results demonstrate that the array is a reproducible and informative tool for assessing differential gene expression in dolphin PBL and in other tissues.

Animals↗

High-Content CRISPR Screening: Methods and Applications.

Clustered regularly interspaced short palindromic repeats (CRISPR)-Cas9 screening has become a central technology in functional genomics, enabling genome-scale interrogation via pooled perturbations. Early CRISPR screens employed survival or simple phenotypic readouts to identify essential genes and drug resistance mechanisms. However, as biological questions have shifted toward understanding regulatory networks, cellular heterogeneity, and context-dependent gene functions, there has been increasing demand for screening strategies capable of capturing complex cellular phenotypes beyond cell fitness. Recent advances in single-cell sequencing, high-content imaging, and spatial transcriptomics have expanded the resolution of CRISPR screening by enabling multidimensional phenotypic characterization following genetic perturbation. By integrating pooled perturbations with diverse readouts, these approaches systematically map targeted gene edits to transcriptional states, cellular phenotypes, and microenvironmental contexts. Meanwhile, innovations in library design, delivery, and computational pipelines have further improved the robustness and interpretability of high-content screening platforms. This review synthesizes the methodological evolution of CRISPR screening, emphasizing advances in perturbation strategies, delivery systems, and multimodal readouts. Representative applications spanning oncology, immunotherapy, developmental biology, neurobiology, and infectious diseases are delineated to demonstrate refined gene network annotations. Additionally, existing technical bottlenecks, such as scalability, cost constraints, and in vivo limitations, are critically assessed. Finally, future directions are proposed to facilitate the development of precise medicine.

CRISPR screening↗

Earliest changes in the left ventricular transcriptome postmyocardial infarction.

We report a genome-wide survey of early responses of the mouse heart transcriptome to acute myocardial infarction (AMI). For three regions of the left ventricle (LV), namely, ischemic/infarcted tissue (IF), the surviving LV free wall (FW), and the interventricular septum (IVS), 36,899 transcripts were assayed at six time points from 15 min to 48 h post-AMI in both AMI and sham surgery mice. For each transcript, temporal expression patterns were systematically compared between AMI and sham groups, which identified 515 AMI-responsive genes in IF tissue, 35 in the FW, 7 in the IVS, with three genes induced in all three regions. Using the literature, we assigned functional annotations to all 519 nonredundant AMI-induced genes and present two testable models for central signaling pathways induced early post-AMI. First, the early induction of 15 genes involved in assembly and activation of the activator protein-1 (AP-1) family of transcription factors implicates AP-1 as a dominant regulator of earliest post-ischemic molecular events. Second, dramatic increases in transcripts for arginase 1 (ARG1), the enzymes of polyamine biosynthesis, and protein inhibitor of nitric oxide synthase (NOS) activity indicate that NO production may be regulated, in part, by inhibition of NOS and coordinate depletion of the NOS substrate, L: -arginine. ARG1: was the single-most highly induced transcript in the database (121-fold in IF region) and its induction in heart has not been previously reported.

Acute Disease↗

EST database for early flower development in California poppy (Eschscholzia californica Cham., Papaveraceae) tags over 6,000 genes from a basal eudicot.

The Floral Genome Project (FGP) selected California poppy (Eschscholzia californica Cham. ssp. Californica) to help identify new florally-expressed genes related to floral diversity in basal eudicots. A large, non-normalized cDNA library was constructed from premeiotic and meiotic floral buds and sequenced to generate a database of 9,079 high quality Expressed Sequence Tags (ESTs). These sequences clustered into 5,713 unigenes, including 1,414 contigs and 4,299 singletons. Homologs of genes regulating many aspects of flower development were identified, including those for organ identity and development, cell and tissue differentiation, cell cycle control, and secondary metabolism. Over 5% of the transcriptome consisted of homologs to known floral gene families. Most are the first representatives of their respective gene families in basal eudicots and their conservation suggests they are important for floral development and/or function. App. 10% of the transcripts encoded transcription factors and other regulatory genes, including nine genes from the seven major lineages of the important MADS-box family of developmental regulators. Homologs of alkaloid pathway genes were also recovered, providing opportunities to explore adaptive evolution in secondary products. Furthermore, comparison of the poppy ESTs with the Arabidopsis genome provided support for putative Arabidopsis genes that previously lacked annotation. Finally, over 1,800 unique sequences had no observable homology in the public databases. The California poppy EST database and library will help bridge our understanding of flower initiation and development among higher eudicot and monocot model plants and provide new opportunities for comparative analysis of gene families across angiosperm species.

DNA, Complementary↗

GeneTide--Terra Incognita Discovery Endeavor: a new transcriptome focused member of the GeneCards/GeneNote suite of databases.

GeneCards is an automatically mined database of human genes that strives to create, along with its auxiliary databases--GeneLoc, GeneNote and GeneAnnot--the most inclusive resource of gene-centered information of the human genome. GeneTide, the Gene Terra Incognita Discovery Endeavor (http://genecards.weizmann.ac.il/genetide/), the newest addition to this family, is a transcriptome-focused database which aims to enhance GeneCards with additional expressed sequence tag (EST)-based genes. This is achieved by comprehensively mapping >85% of the approximately 5.6 million human ESTs currently available at dbEST to known genes by means of data mining and integration of genomic resources including UniGene, DoTS, AceView and in-house resources. GeneTide thus creates comprehensive links between ESTs and GeneCards genes. Furthermore, groups of unassociated transcripts serve as a basis for defining novel EST-based GeneCards Candidates (EGCs). These EGCs, nearly 25,000 of which were defined in version 0.3 of GeneTide, are further annotated with various parameters, including splicing evidence and expression data extracted from the GeneNote database, to determine their validity as possible de novo genes.

Databases, Genetic↗

Expressed sequence tag profiling identifies developmental and anatomic partitioning of gene expression in the mouse prostate.

BACKGROUND: The prostate gland is an organ with highly specialized functional attributes that serves to enhance the fertility of mammalian species. Much of the information pertaining to normal and pathological conditions affecting the prostate has been obtained through extensive developmental, biochemical and genetic analyses of rodent species. Although important insights can be obtained through detailed anatomical and histological assessments of mouse and rat models, further mechanistic explanations are greatly aided through studies of gene and protein expression. RESULTS: In this article we characterize the repertoire of genes expressed in the normal developing mouse prostate through the analysis of 50,562 expressed sequence tags derived from 14 mouse prostate cDNA libraries. Sequence assemblies and annotations identified 15,009 unique transcriptional units of which more than 600 represent high quality assemblies without corresponding annotations in public gene expression databases. Quantitative analyses demonstrate distinct anatomical and developmental partitioning of prostate gene expression. This finding may assist in the interpretation of comparative studies between human and mouse and guide the development of new transgenic murine disease models. The identification of several novel genes is reported, including a new member of the beta-defensin gene family with prostate-restricted expression. CONCLUSIONS: These findings suggest a potential role for the prostate as a defensive barrier for entry of pathogens into the genitourinary tract and, further, serve to emphasize the utility of the continued evaluation of transcriptomes from a diverse repertoire of tissues and cell types.

Amino Acid Sequence↗

Structural variant discovery and diagnostic impact in rare diseases from short-read and long-read sequencing.

Rare diseases collectively affect 1 in 10 individuals, yet current genetic testing fails to identify a causal variant for most cases. At present, cytogenetic methods and/or sequencing approaches such as exome (ES) or short-read genome sequencing (srGS) represent the state-of-the-art for comprehensive clinical discovery of sequence and structural variants (SVs), including copy number variants, balanced SVs, complex SVs, and tandem repeats (TRs). Recently, long-read genome sequencing (lrGS), coupled with multiomics data, has presented great promise to resolve variation in genomic regions recalcitrant to characterization by srGS such as highly repetitive simple repeat sequences and segmental duplications. However, there are few guidelines to enable clinical interpretation of genetic variation in these highly repetitive genomic regions, and the enthusiasm of the field in adopting lrGS has made it difficult to assess the true added diagnostic yield of this technology due to widely variable and inconsistently applied analytic pipelines and variable degrees of pre-screening by ES or srGS. Here, we investigated the contribution of SVs to rare diseases using srGS as a front-line strategy when paired with highly sensitive SV discovery and evaluate the added diagnostic yield of incorporating lrGS for a subset of cases. Our srGS analysis encompassed 1,462 families (3,450 individuals) recruited through the Broad Institute Center for Mendelian Genetics and the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) programs. Diagnostic SVs were identified in 5.4% of cases (79/1,462), of which 80% were uniquely detectable by srGS compared to standard cytogenetic techniques. For 96 families (including 10 families with a heterozygous variant observed in a known recessive gene of clinical relevance), we performed lrGS with methylation profiling, as well as long-read transcriptomic analyses in a subset of 20 trios. Analyses with lrGS yielded over 25,000 SVs per genome, 63% of which were not captured by srGS, along with an additional ~200 rare SNV/indels per genome not previously captured and 12 differentially methylated regions per genome. Among these, we identified only one diagnostic variant not interpreted by srGS, an apparently mosaic de novo SNV in CASK that was absent in the srGS callset due to allelic imbalance. No new diagnoses were supported by long-read transcriptomics or episignatures. In this well characterized rare disease cohort, the added diagnostic yield was thus 1.04% (1/96 families). Following a systematic literature review of prior lrGS studies, we find that most reported diagnoses were detectable by srGS and that our added diagnostic yield is consistent with those prior studies. These studies emphasize the significant impact of comprehensive SV discovery in rare disease cases and further demonstrate the power for increased discovery of novel genomic variation and episignatures from lrGS. Nonetheless, they also serve to temper expectations of dramatic diagnostic advances in rare disease patients until there is more extensive annotation of the functional and clinical impact of all coding and noncoding variation uniquely accessible to lrGS with extensive reference databases spanning highly repetitive genomic sequencing that could be enabled by this transformative technology.

Journal Article↗

An Annotated Biobank of Triple-Negative Breast Cancer Patient-Derived Xenografts Features Treatment-Naïve and Longitudinal Samples during Neoadjuvant Chemotherapy.

UNLABELLED: Triple-negative breast cancer (TNBC) that fails to respond to neoadjuvant chemotherapy (NACT) can be lethal. Developing effective strategies to eradicate chemoresistant disease requires experimental models that recapitulate the heterogeneity characteristic of TNBC. To that end, we established a biobank of 92 orthotopic patient-derived xenograft (PDX) models of TNBC from the tumors of 75 patients enrolled in A Robust TNBC Evaluation fraMework to Improve Survival clinical trial (ARTEMIS, NCT02276443), including 12 longitudinal sets generated from serial patient biopsies collected throughout NACT treatment and from metastatic disease. Models were established from both chemosensitive and chemoresistant tumors, and nearly 30% of the PDX models were capable of metastasizing to the lungs. Comprehensive molecular profiling demonstrated conservation of genomes and transcriptomes between patient and corresponding PDX tumors, with representation of all major transcriptional subtypes. Transcriptional changes observed in the longitudinal PDX models highlighted dysregulation in pathways associated with DNA integrity, extracellular matrix interactions, the ubiquitin-proteasome system, epigenetics, and inflammatory signaling. These alterations revealed a complex network of adaptations associated with chemoresistance. Overall, this PDX biobank provides a valuable tool for tackling the most pressing issues facing the clinical management of TNBC. SIGNIFICANCE: The development of a patient-derived xenograft biobank that comprehensively captures the genomic and transcriptional diversity of triple-negative breast cancer promises to be a robust resource to investigate and overcome chemoresistance and metastasis.

Animals↗

Moderate expression and activity of flocculins underlie the characteristic flocculation phenotype of Saccharomyces pastorianus.

Flocculation is a key technological trait in lager brewing, governing fermentation performance, yeast recovery, and beer quality. In the allo-aneuploid hybrid yeast Saccharomyces pastorianus, the genetic basis of flocculation remains poorly resolved due to its complex dual sub-genome architecture. Here, we systematically re-annotated and functionally characterized the complete FLO gene repertoire of the Group II strain CBS 1483. Thirteen FLO genes were identified, including allelic variants and a previously uncharacterized adhesin, Flo12, containing a Hyphal_reg_CWP domain instead of the canonical PA14 lectin-binding domain. Structural modeling revealed strong conservation of Ca²+-binding residues in PA14 domains, alongside repeat-region diversification likely contributing to functional variability. Using optogenetic expression in a FLO-null background, we demonstrated that SpcI-FLO9-1 and SpcI-FLO9-2_1 are the strongest drivers of flocculation, exhibiting NewFlo-like sugar sensitivity. Transcriptomic analysis during 17°P wort fermentation showed dynamic induction of these genes coinciding with flocculation onset. Surprisingly, deletion of both loci in CBS 1483 did not abolish but only delayed sedimentation in wort, accompanied by improved maltose utilization and attenuation. These findings reveal functional redundancy and compensatory mechanisms within the FLO network of lager yeast, highlighting the genetic complexity underlying flocculation, and providing a molecular framework to inform yeast selection, strain development, and optimization of the lager fermentation processes.IMPORTANCEFlocculation, the process by which yeast cells aggregate and settle, is essential for producing clear, high-quality lager beer, and for efficient yeast recovery during brewing. However, the genetic basis of this trait in lager yeast has remained poorly understood because these strains possess unusually complex hybrid genomes. In this study, we systematically identified and characterized the complete set of flocculation genes in the industrial lager yeast Saccharomyces pastorianus CBS 1483. We demonstrated that lager yeast flocculation is not controlled by a single dominant gene, but instead emerges from the combined action of several moderately active adhesion proteins that are expressed at low levels during fermentation. Surprisingly, deleting the two strongest candidate genes only delayed, rather than eliminated, sedimentation, revealing a robust compensatory network that preserves brewing performance. These findings refine the current understanding of yeast flocculation and provide a molecular framework for developing brewing strains with improved fermentation efficiency, product consistency, and flavor quality.

Saccharomyces pastorianus↗

Enhanced chromatin compaction is associated with de novo expression of a nuclear microprotein, global loss of H3 acetylation and local transcriptional changes in retinal rod photoreceptors.

We have limited understanding of how aging alters gene expression and remodels cellular architecture in post-mitotic neurons. The inverted nuclear organization of mouse rod photoreceptors provides a unique model to gain mechanistic insights into age-associated decline in neuronal function. We have generated and integrated multi-omic datasets including 3D-genome topology, histone modifications, chromatin accessibility, DNA methylation and transcriptome of rod photoreceptors from young- and aged-mice. We show that aging drives global chromatin compaction, with regional alterations enriched at active chromatin. Epigenomic and transcriptional changes broadly correlate with chromatin dynamics as validated by high resolution microscopy. We uncover a megabase-sized genomic region with multi-level alterations, including de novo transcription of Gm7239, which encodes a functional microprotein carrying histone acetyltransferase-inhibitor domain. Overexpression of Gm7239 is associated with global loss of histone H3 acetylation, highlighting a potential new axis of genomic regulation in aging. Finally, we identify multiple significant local transcriptional alterations in non-annotated regions and genes associated with age-related macular degeneration. Our studies link age-related chromatin landscape changes with gene expression that may influence rod function and vulnerability to diseases.

Journal Article↗

Impact of alternative initiation, splicing, and termination on the diversity of the mRNA transcripts encoded by the mouse transcriptome.

We analyzed the FANTOM2 clone set of 60,770 RIKEN full-length mouse cDNA sequences and 44,122 public mRNA sequences. We developed a new computational procedure to identify and classify the forms of splice variation evident in this data set and organized the results into a publicly accessible database that can be used for future expression array construction, structural genomics, and analyses of the mechanism and regulation of alternative splicing. Statistical analysis shows that at least 41% and possibly as much as 60% of multiexon genes in mouse have multiple splice forms. Of the transcription units with multiple splice forms, 49% contain transcripts in which the apparent use of an alternative transcription start (stop) is accompanied by alternative splicing of the initial (terminal) exon. This implies that alternative transcription may frequently induce alternative splicing. The fact that 73% of all exons with splice variation fall within the annotated coding region indicates that most splice variation is likely to affect the protein form. Finally, we compared the set of constitutive (present in all transcripts) exons with the set of cryptic (present only in some transcripts) exons and found statistically significant differences in their length distributions, the nucleotide distributions around their splice junctions, and the frequencies of occurrence of several short sequence motifs.

Alternative Splicing↗

A single-cell transcriptomic atlas of the pigtail macaque placenta in late gestation.

The placenta is a complex organ with multiple immune and non-immune cell types that promote fetal tolerance and facilitate the transfer of nutrients and oxygen. The nonhuman primate (NHP) is a key experimental model for studying human pregnancy complications, in part due to similarities in placental structure, which makes it essential to understand how single-cell populations compare across the human and NHP maternal-fetal interface. We constructed a single-cell RNA-Seq (scRNA-Seq) atlas of the placenta from the pigtail macaque ( Macaca nemestrina ) in the third trimester, comprising three different tissues at the maternal-fetal interface: the chorionic villi (placental disc), chorioamniotic membranes, and the maternal decidua. Each tissue was separately dissociated into single cells and processed through the 10X Genomics and Seurat pipeline, followed by aggregation, unsupervised clustering, and cluster annotation. Next, we determined the maternal-fetal origins of cell populations and analyzed single-cell RNA trajectory, Gene Ontology enrichment, and cell-cell communication. Single-cell populations in the pigtail macaque were strikingly similar in their identity and frequency to those found in the human placenta, including cells from trophoblast, stromal cell, immune, and macrophage lineages. An advantage of our approach was the deep sequencing of three tissues at the maternal-fetal interface, which yielded a rich diversity of common and rare single-cell populations. The third-trimester pigtail macaque single-cell atlas enables the identification of cellular subclusters analogous to those in humans and provides a powerful resource for understanding experimental perturbations on the NHP placenta.

Journal Article↗

Identification of "pathologs" (disease-related genes) from the RIKEN mouse cDNA dataset using human curation plus FACTS, a new biological information extraction system.

BACKGROUND: A major goal in the post-genomic era is to identify and characterise disease susceptibility genes and to apply this knowledge to disease prevention and treatment. Rodents and humans have remarkably similar genomes and share closely related biochemical, physiological and pathological pathways. In this work we utilised the latest information on the mouse transcriptome as revealed by the RIKEN FANTOM2 project to identify novel human disease-related candidate genes. We define a new term "patholog" to mean a homolog of a human disease-related gene encoding a product (transcript, anti-sense or protein) potentially relevant to disease. Rather than just focus on Mendelian inheritance, we applied the analysis to all potential pathologs regardless of their inheritance pattern. RESULTS: Bioinformatic analysis and human curation of 60,770 RIKEN full-length mouse cDNA clones produced 2,578 sequences that showed similarity (70-85% identity) to known human-disease genes. Using a newly developed biological information extraction and annotation tool (FACTS) in parallel with human expert analysis of 17,051 MEDLINE scientific abstracts we identified 182 novel potential pathologs. Of these, 36 were identified by computational tools only, 49 by human expert analysis only and 97 by both methods. These pathologs were related to neoplastic (53%), hereditary (24%), immunological (5%), cardio-vascular (4%), or other (14%), disorders. CONCLUSIONS: Large scale genome projects continue to produce a vast amount of data with potential application to the study of human disease. For this potential to be realised we need intelligent strategies for data categorisation and the ability to link sequence data with relevant literature. This paper demonstrates the power of combining human expert annotation with FACTS, a newly developed bioinformatics tool, to identify novel pathologs from within large-scale mouse transcript datasets.

Animals↗

Chemical effects in biological systems--data dictionary (CEBS-DD): a compendium of terms for the capture and integration of biological study design description, conventional phenotypes, and 'omics data.

A critical component in the design of the Chemical Effects in Biological Systems (CEBS) Knowledgebase is a strategy to capture toxicogenomics study protocols and the toxicity endpoint data (clinical pathology and histopathology). A Study is generally an experiment carried out during a period of time for the purpose of obtaining data, and the Study Design Description captures the methods, timing, and organization of the Study. The CEBS Data Dictionary (CEBS-DD) has been designed to define and organize terms in an attempt to standardize nomenclature needed to describe a toxicogenomics Study in a structured yet intuitive format and provide a flexible means to describe a Study as conceptualized by the investigator. The CEBS-DD will organize and annotate information from a variety of sources, thereby facilitating the capture and display of toxicogenomics data in biological context in CEBS, i.e., associating molecular events detected in highly-parallel data with the toxicology/pathology phenotype as observed in the individual Study Subjects and linked to the experimental treatments. The CEBS-DD has been developed with a focus on acute toxicity studies, but with a design that will permit it to be extended to other areas of toxicology and biology with the addition of domain-specific terms. To illustrate the utility of the CEBS-DD, we present an example of integrating data from two proteomics and transcriptomics studies of the response to acute acetaminophen toxicity (A. N. Heinloth et al., 2004, Toxicol. Sci. 80, 193-202).

Acetaminophen↗

Evaluation of monocot and eudicot divergence using the sugarcane transcriptome.

Over 40,000 sugarcane (Saccharum officinarum) consensus sequences assembled from 237,954 expressed sequence tags were compared with the protein and DNA sequences from other angiosperms, including the genomes of Arabidopsis and rice (Oryza sativa). Approximately two-thirds of the sugarcane transcriptome have similar sequences in Arabidopsis. These sequences may represent a core set of proteins or protein domains that are conserved among monocots and eudicots and probably encode for essential angiosperm functions. The remaining sequences represent putative monocot-specific genetic material, one-half of which were found only in sugarcane. These monocot-specific cDNAs represent either novelties or, in many cases, fast-evolving sequences that diverged substantially from their eudicot homologs. The wide comparative genome analysis presented here provides information on the evolutionary changes that underlie the divergence of monocots and eudicots. Our comparative analysis also led to the identification of several not yet annotated putative genes and possible gene loss events in Arabidopsis.

Arabidopsis↗

The transcriptome of the sea urchin embryo.

The sea urchin Strongylocentrotus purpuratus is a model organism for study of the genomic control circuitry underlying embryonic development. We examined the complete repertoire of genes expressed in the S. purpuratus embryo, up to late gastrula stage, by means of high-resolution custom tiling arrays covering the whole genome. We detected complete spliced structures even for genes known to be expressed at low levels in only a few cells. At least 11,000 to 12,000 genes are used in embryogenesis. These include most of the genes encoding transcription factors and signaling proteins, as well as some classes of general cytoskeletal and metabolic proteins, but only a minor fraction of genes encoding immune functions and sensory receptors. Thousands of small asymmetric transcripts of unknown function were also detected in intergenic regions throughout the genome. The tiling array data were used to correct and authenticate several thousand gene models during the genome annotation process.

Animals↗

Quantitative analysis of wine yeast gene expression profiles under winemaking conditions.

Wine fermentation is a dynamic and complex process in which the yeast cell is subjected to multiple stress conditions. A successful adaptation involves changes in gene expression profiles where a large number of genes are up- or downregulated. Functional genomic approaches are commonly used to obtain global gene expression profiles, thereby providing a comprehensive view of yeast physiology. We used SAGE to quantify gene expression profiles in an industrial strain of Saccharomyces cerevisiae under winemaking conditions. The transcriptome of wine yeast was analysed at three stages during the fermentation process, mid-exponential phase, and early- and late-stationary phases. Upon correlation with the yeast genome, we found three classes of transcripts: (a) sequences that corresponded to ORFs; (b) expressed sequences from intergenic regions; and (c) messengers that did not match the published reference yeast genome. In all fermentation phases studied, the most highly expressed genes related to energy production and stress response. For many pathways, including glycolysis, different transcript levels were observed during each phase. Different isoenzymes, including hexose transporters (HXT), were differentially induced, depending on the growth phase. About 10% of transcripts matched non-annotated ORF regions within the yeast genome and could correspond to small novel genes originally omitted in the first gene annotation effort. Up to 22% of transcripts, particularly at late-stationary phase, did not match any known location within the genome. As the available reference yeast genome was obtained from a laboratory strain, these expressed sequences could represent genes only expressed by an industrial yeast strain. Further studies are necessary to identify the role of these potential genes during wine fermentation.

Cluster Analysis↗