Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptome sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Spatiotemporal mapping of tertiary lymphoid structure heterogeneity shapes immune niches and clinical outcomes in intrahepatic cholangiocarcinoma.

Intrahepatic cholangiocarcinoma (iCCA) is a highly lethal malignancy with limited therapeutic options. The spatial architecture and functional diversity of tertiary lymphoid structures (TLSs) in iCCA remain unclear. Here, we present a multimodal spatial atlas of TLSs and identified intratumoral TLSs (iTLSs) as independent prognostic markers. Bulk proteomic profiling of 214 discovery and 155 validation cases identified a four-tier TLS-based tumor microenvironment classification system and supported development of a TLS-predictive random forest classifier. Imaging mass cytometry revealed that iTLS+ tumors harbor structured immune architectures, where M1-like tissue-resident macrophages (RTMs), dendritic cells, and CXCL13+ CD4+ T cells colocalize to form antigen-presenting neighborhoods (apc-CNs) spatially coupled to TLS core regions (TLScore-CNs). Single-cell spatial transcriptomics further resolved 61 TLSs into 14 spatial niches and defined a pseudotemporal maturation continuum: aggregated, activated, and postactivated. Intraniche communication, primarily mediated by ifnCAFs, iCAFs, and CXCL12+ macrophages, evolved dynamically with maturation. Single-nucleus RNA sequencing combined with Tangram-based spatial mapping revealed CXCL12+ macrophages and iCAFs forming a peripheral band in aggregated TLSs, whereas ifnCAFs infiltrated TLS interiors during activation. These findings define TLS heterogeneity and provide insights for stroma-directed immunotherapy.

Cholangiocarcinoma↗

Transcriptional and structural impact of TATA-initiation site spacing in mammalian core promoters.

BACKGROUND: The TATA box, one of the most well studied core promoter elements, is associated with induced, context-specific expression. The lack of precise transcription start site (TSS) locations linked with expression information has impeded genome-wide characterization of the interaction between TATA and the pre-initiation complex. RESULTS: Using a comprehensive set of 5.66 x 10(6) sequenced 5' cDNA ends from diverse tissues mapped to the mouse genome, we found that the TATA-TSS distance is correlated with the tissue specificity of the downstream transcript. To achieve tissue-specific regulation, the TATA box position relative to the TSS is constrained to a narrow window (-32 to -29), where positions -31 and -30 are the optimal positions for achieving high tissue specificity. Slightly larger spacings can be accommodated only when there is no optimally spaced initiation signal; in contrast, the TATA box like motifs found downstream of position -28 are generally nonfunctional. The strength of the TATA binding protein-DNA interaction plays a subordinate role to spacing in terms of tissue specificity. Furthermore, promoters with different TATA-TSS spacings have distinct features in terms of consensus sequence around the initiation site and distribution of alternative TSSs. Unexpectedly, promoters that have two dominant, consecutive TSSs are TATA depleted and have a novel GGG initiation site consensus. CONCLUSION: In this report we present the most comprehensive characterization of TATA-TSS spacing and functionality to date. The coupling of spacing to tissue specificity at the transcriptome level provides important clues as to the function of core promoters and the choice of TSS by the pre-initiation complex.

Animals↗

A picture of gene sampling/expression in model organisms using ESTs and KOG proteins.

The expressed sequence tag (EST) is an instrument of gene discovery. When available in large numbers, ESTs may be used to estimate gene expression. We analyzed gene expression by EST sampling, using the KOG database, which includes 24,154 proteins from Arabidopsis thaliana (Ath), 17,101 from Caenorhabditis elegans (Cel), 10,517 from Drosophila melanogaster (Dme), and 26,324 from Homo sapiens (Hsa), and 178,538 ESTs for Ath, 215,200 for Cel, 261,404 for Dme, and 1,941,556 for Hsa. BLAST similarity searches were performed to assign KOG annotation to all ESTs. We determined the amount of gene sampling or expression dedicated to each KOG functional category by each model organism. We found that the 25% most-expressed genes are frequently shared among these organisms. The KOG protein classification allowed the EST sampling calculation throughout the glycolysis pathway. We calculated the KOG cluster coverage and inferred that 50 to 80 K ESTs would efficiently cover 80-85% of the KOG database clusters in a transcriptome project. Since KOG is a database biased towards housekeeping genes, this is probably the number of ESTs needed to include the more commonly expressed genes in these organisms. We also examined a still unaddressed question: what is the minimum number of ESTs that should be produced in a transcriptome project?

Animals↗

Functional genomics as applied to mapping transcription regulatory networks.

The sequencing of the human genome and the entire genomes of many model organisms has resulted in the identification of many genes. Many large-scale experiments for generating gene disruptions and analyzing the phenotypes are underway to ascertain gene function. A future challenge will be to determine interaction and regulation of all the genes of an organism. Recent advances in functional genomic technology have begun to shine light on such gene network problems at both transcriptomic and proteomic levels. Functional genomics will not only elucidate what the genes do, but will also help determine when, where and how they are expressed as an orchestrated system. In this review, we discuss the functional genomics approaches to extract knowledge about transcription regulatory mechanisms from combinations of sequence data, microarray data and ChIP data. We focus in particular on the budding yeast Saccharomyces cerevisiae.

Gene Expression Regulation, Fungal↗

Single primer amplification (SPA) of cDNA for microarray expression analysis.

The potential of expression analysis using cDNA microarrays to address complex problems in a wide variety of biological contexts is now being realised. A limiting factor in such analyses is often the amount of RNA required, usually tens of micrograms. To address this problem researchers have turned to methods of improving detection sensitivity, either through increasing fluorescent signal output per mRNA molecule or increasing the amount of target available for labelling by use of an amplification procedure. We present a novel DNA-based method in which an oligonucleotide is incorporated into the 3' end of cDNA during second-strand cDNA synthesis. This sequence provides an annealing site for a single complementary heel primer that directs Taq DNA polymerase amplification of cDNA following multiple cycles of denaturation, annealing and extension. The utility of this technique for transcriptome-wide screening of relative expression levels was compared to two alternative methodologies for production of labelled cDNA target, namely incorporation of fluorescent nucleotides by reverse transcriptase or the Klenow fragment. Labelled targets from two distinct mouse tissues, adult liver and kidney, were compared by hybridisation to a set of cDNA microarrays containing 6500 mouse cDNA probes. Here we demonstrate, through a dilution series of cDNA derived from 10 micro g of total RNA, that it is possible to produce datasets comparable to those produced with unamplified targets with the equivalent of 30 ng of total RNA. The utility of this technique for microarray analysis in cases where sample is limited is discussed.

Animals↗

A latent activated olfactory stem cell state revealed by single-cell transcriptomic and epigenomic profiling.

The olfactory epithelium is one of the few regions of the nervous system that sustains neurogenesis throughout life. Its experimental accessibility makes it especially tractable for studying molecular mechanisms that drive neural regeneration in response to injury. In this study, we used single-cell sequencing to identify transcriptional and epigenetic processes involved in determining olfactory epithelial stem cell fate during injury-induced regeneration. By combining gene expression and accessible chromatin profiles of individual lineage-traced olfactory stem cells, we identified transcriptional heterogeneity among activated stem cells at a stage when cell fates are being specified. We further identified a subset of resting cells that appears poised for activation, characterized by accessible chromatin around silent genes prior to their expression in response to injury. These results provide evidence for a latent activated stem cell state in which a subset of quiescent olfactory epithelial stem cells are epigenetically primed to support injury-induced regeneration.

Animals↗

A recurrent CCDC82 frameshift variant associated with syndromic neurodevelopmental disorder in a consanguineous Pakistani family.

BACKGROUND: Intellectual disabilities (IDs) are part of neurodevelopmental disorders (NDDs) and are genetically heterogeneous conditions characterized by impairments in cognition, learning, and adaptive functioning. Despite advances in gene discovery, many individuals, particularly those from understudied populations, remain without a molecular diagnosis. Recent reports implicate CCDC82 (HGNC: 26282) as an autosomal recessive ID gene, although the phenotypic spectrum and biological context remain incompletely defined. METHODS: Exome sequencing (ES) was performed in a consanguineous Pakistani family (PKMR06A) with four affected individuals presenting with moderate to severe ID. Variant segregation was confirmed by Sanger sequencing. In silico analyses, including pathogenicity prediction, protein structural modeling, and domain intolerance assessment, were used to evaluate the functional consequences of the identified variant. Spatiotemporal gene expression patterns were examined using bulk and single-cell human brain transcriptomic datasets. RESULTS: Clinically, affected individuals of family PKMR06A presented with early childhood global developmental delay, speech delay, hypotonia, gait abnormalities, spasticity, and mild facial dysmorphism. Genetic screening revealed a recurrent rare homozygous frameshift variant in CCDC82 (NM_024725.4): c.373del; p.(Asp125Ilefs*6), segregating with disease in all available affected individuals of the family. The identified c.373del variant was absent from the gnomAD database and was classified as pathogenic (PVS1, PM2, and PP1) based on ACMG/AMP criteria. The c.373del variant is predicted to introduce a premature termination codon, p.(Asp125Ilefs*6), leading to deletion of essential coiled-coil domains from the encoded protein, supporting a loss-of-function mechanism. In silico, transcriptomic analyses demonstrated preferential CCDC82 expression during prenatal human brain development, providing developmental context for the neurodevelopmental phenotype associated with the identified truncating variant. CONCLUSIONS: This study expands the mutational landscape of CCDC82 and provides additional clinical and molecular evidence supporting its role in autosomal recessive NDD. The findings reinforce the importance of CCDC82 in human neurodevelopment and highlight the value of genomic investigation in underrepresented populations.

Autosomal recessive↗

Computational models with thermodynamic and composition features improve siRNA design.

BACKGROUND: Small interfering RNAs (siRNAs) have become an important tool in cell and molecular biology. Reliable design of siRNA molecules is essential for the needs of large functional genomics projects. RESULTS: To improve the design of efficient siRNA molecules, we performed a comparative, thermodynamic and correlation analysis on a heterogeneous set of 653 siRNAs collected from the literature. We used this training set to select siRNA features and optimize computational models. We identified 18 parameters that correlate significantly with silencing efficiency. Some of these parameters characterize only the siRNA sequence, while others involve the whole mRNA. Most importantly, we derived an siRNA position-dependent consensus, and optimized the free-energy difference of the 5' and 3' terminal dinucleotides of the siRNA antisense strand. The position-dependent consensus is based on correlation and t-test analyses of the training set, and accounts for both significantly preferred and avoided nucleotides in all sequence positions. On the training set, the two parameters' correlation with silencing efficiency was 0.5 and 0.36, respectively. Among other features, a dinucleotide content index and the frequency of potential targets for siRNA in the mRNA added predictive power to our model (R = 0.55). We showed that our model is effective for predicting the efficiency of siRNAs at different concentrations. We optimized a neural network model on our training set using three parameters characterizing the siRNA sequence, and predicted efficiencies for the test siRNA dataset recently published by Novartis. On this validation set, the correlation coefficient between predicted and observed efficiency was 0.75. Using the same model, we performed a transcriptome-wide analysis of optimal siRNA targets for 22,600 human mRNAs. CONCLUSION: We demonstrated that the properties of the siRNAs themselves are essential for efficient RNA interference. The 5' ends of antisense strands of efficient siRNAs are U-rich and possess a content similarity to the pyrimidine-rich oligonucleotides interacting with the polypurine RNA tracks that are recognized by RNase H. The advantage of our method over similar methods is the small number of parameters. As a result, our method requires a much smaller training set to produce consistent results. Other mRNA features, though expensive to compute, can slightly improve our model.

Artificial Intelligence↗

Admission whole-blood transcriptomic characterization of a neutrophil-predominant systemic immune response in patients with acute traumatic brain injury.

BACKGROUND: Acute traumatic brain injury (TBI) is accompanied by systemic immune responses, but their whole-blood transcriptomic features at hospital arrival remain incompletely characterized. We aimed to characterize these features in patients with acute TBI compared with healthy controls. METHODS: In this single-center prospective observational study, we performed whole-blood RNA sequencing on hospital-arrival samples from 42 patients with acute TBI and 21 healthy controls. Analyses included differential expression (limma-voom; FDR < 0.05, |log2FC| > 0.7), functional enrichment, Ingenuity Pathway Analysis, CIBERSORTx LM22 deconvolution, and per-sample neutrophil degranulation signature scoring. RESULTS: Differential expression analysis identified 996 upregulated and 863 downregulated genes, with marked upregulation of inflammation-, innate immunity-, and neutrophil-related genes including DUSP1, HMGB2, MMP9, and S100A8. Canonical pathways with positive IPA z-scores included Neutrophil degranulation, Neutrophil Extracellular Trap Signaling Pathway, and Toll-like Receptor Signaling; upstream regulators included TNF, IL1B, IFNG, and STAT3. Deconvolution identified 7 of 22 differing subsets (q < 0.05), with relatively higher myeloid and lower lymphoid fractions in TBI. The Neutrophil degranulation signature score correlated with Injury Severity Score within TBI (Spearman &#x3c1; = +0.55; q < 0.001). CONCLUSIONS: Admission whole-blood transcriptomics characterized a neutrophil-predominant systemic transcriptional response in patients with acute TBI. This response was also evident among patients without major extracranial injury and was associated with total ISS. However, because the study lacked an appropriately matched non-TBI trauma comparator, the findings should be interpreted as a descriptive characterization of a systemic injury response accompanying TBI and do not establish a TBI-specific molecular signature or mechanism.

gene expression↗

Gene expression profiles of T lymphocytes are sensitive to the influence of heavy smoking: A pilot study.

Cigarette smoke components have a proven negative influence on human health. Adverse metabolic effects were observed in tissues and single cells. T lymphocytes get in contact with affected organs (e.g., lung) or cells (e.g., erythrocytes), as well as with smoke components and bioactive molecules, whose production is triggered by tobacco smoke. We therefore compared the gene expression profiles in these cells of the adaptive immune system of three male heavy smokers and three male nonsmokers using rapid T cell isolation and Affymetrix GeneChip HG U133A 2.0 microarray analysis. Eighty-eight genes were found to be significantly (t test) differentially expressed by a factor of 1.5-fold or larger between smokers and nonsmokers. Using the gene function groups of the gene ontology consortium to categorize the functions of the differentially expressed genes, the group termed "response to stimulus" was found to be most significantly affected by smoking. Our data indicate a prominent role of cytotoxic T lymphocytes in response to smoking. Several genes that are typically expressed in these cells were found regulated although the ratio of cytotoxic and helper T lymphocytes remained unchanged in smokers. Our data show that, in principle, it might be possible to identify health-related biomarkers in the transcriptome of T lymphocytes.

Adult↗

Analysis of gene expression profiles of normal human nasal mucosa and nasal polyp tissues by SAGE.

BACKGROUND: A systemic determination of gene expression profiles in nasal polyp compared with normal nasal mucosa would contribute considerably to the investigation of the disease marker in various rhinopathy and general knowledge on the formation of human nasal polyp. OBJECTIVE: The aim of this study was to identify the transcriptome of the normal human nasal mucosa and nasal polyp by the serial analysis of gene expression. METHODS: mRNA was extracted from normal inferior turbinate mucosa and nasal polyp. Short sequences (tags), each one usually corresponding to a distinct transcript, was isolated and concatemerized into long DNA molecules, which were cloned and sequenced. RESULTS: We detected 65,305 tags for normal nasal mucosa and 55,829 tags for nasal polyp, representing 20,629 and 17,636 potential transcripts species, respectively. Of the unique tags encountered more than once, 92% (normal nasal mucosa) and 90% (nasal polyp) matched known genes or expressed sequence tags, whereas the remainder did not match any GenBank sequences. Therefore, 504 and 539 novel transcripts were identified in normal nasal mucosa and nasal polyp, respectively. Although the expression levels of most transcripts in both libraries were similar, 114 transcripts were differentially expressed at a statistically significant level (P < .05)-that is, 65 and 49 transcripts among them were expressed at a higher level in normal nasal mucosa and nasal polyp, respectively. CONCLUSION: This information should be very useful for basic knowledge as well as for future studies on pathophysiological conditions of human nasal mucosa, providing some clues to evaluate the altered factors in a variety of rhinopathies. CLINICAL IMPLICATIONS: The results of this study might contribute to general knowledge on the formation of human nasal polyp.

Adult↗

Catalog of gene expression in adult neural stem cells and their in vivo microenvironment.

Stem cells generally reside in a stem cell microenvironment, where cues for self-renewal and differentiation are present. However, the genetic program underlying stem cell proliferation and multipotency is poorly understood. Transcriptome analysis of stem cells and their in vivo microenvironment is one way of uncovering the unique stemness properties and provides a framework for the elucidation of stem cell function. Here, we characterize the gene expression profile of the in vivo neural stem cell microenvironment in the lateral ventricle wall of adult mouse brain and of in vitro proliferating neural stem cells. We have also analyzed an Lhx2-expressing hematopoietic-stem-cell-like cell line in order to define the transcriptome of a well-characterized and pure cell population with stem cell characteristics. We report the generation, assembly and annotation of 50,792 high-quality 5'-end expressed sequence tag sequences. We further describe a shared expression of 1065 transcripts by all three stem cell libraries and a large overlap with previously published gene expression signatures for neural stem/progenitor cells and other multipotent stem cells. The sequences and cDNA clones obtained within this framework provide a comprehensive resource for the analysis of genes in adult stem cells that can accelerate future stem cell research.

Animals↗

A web-accessible complete transcriptome of normal human and DMD muscle.

We present an assessment of the complete transcriptome of human skeletal muscle in Duchenne muscular dystrophy patient muscle and non-dystrophic controls (36 RNAs analyzed from ten Duchenne dystrophy and eight controls; approximately 65,000 gene/expressed sequence tag/probe sets queried on U95 five-GeneChip series and MuscleChip). The use of the multiple chip types allowed us to compare results from different probe sets for the same gene: we found excellent concordance between different probe sets on different microarrays. We found 30% of human genes expressed in muscle at detectable levels. Three percent of these showed differential regulation in dystrophin deficiency. Among 1,882 dysregulated probe sets, 1,324 corresponded to characterized genes/proteins (891 non-redundant transcript units), and 588 to expressed sequence tags or predicted genes. Data interpretation was limited to the insulin-like growth factor pathway members, an investigation of possible de-regulation towards a cardiac lineage, and identification of male- and female-specific transcripts. We found transcriptional upregulation of both IGF-I and IGF-II in dystrophic muscle, however the possible beneficial effects of the growth factors appear offset by transcriptional upregulation of inhibitory IGF-binding proteins and regulators (IGFBP-2, -4, -6 and -7; and PRSS11 [IGFBP-5 protease]). We hypothesize that the beneficial effects of IGF-I or IGF-II supplementation in dystrophic muscle may be the result of dose-dependent sequestration of inhibitory IGF-binding proteins. We also focused on six 'cardiac' genes expressed in muscle (alpha-cardiac actin, CARP, CASQ2, troponin T2 cardiac [TNNT2], CUGBP2, and connexin 43). Comparison to a 27 time point murine muscle regeneration series and mdx muscle profiles showed that CARP and Cx43 were macrophage-associated, and TNNT2 activated-myoblast-associated. Upregulation of cardiac actin and CUGBP2 was not associated with muscle regeneration profiles, suggesting a more specific dysregulation induced by dystrophin deficiency. We found two Y-linked genes expressed solely in male muscle (RPS4Y, DDX3Y), and two autosomal genes expressed much more highly in female muscle (GRO2, ZNF91) (all comparisons P<0.01). Finally, we present the first web-accessible expression profiling database for all data, including image files (.dat), processed image files (.cel), and complete comparison files which are publicly available through a novel queriable web site, that permits query-by-gene across all profiles (http://microarray.cnmcresearch.org/pga). These data enumerate the full range of molecular changes associated downstream of dystrophin deficiency, and provide a web-accessible platform to study the specificity of transcriptional pathway alterations in muscle disease.

Animals↗

A genomic perspective on nutrient provisioning by bacterial symbionts of insects.

Many animals show intimate interactions with bacterial symbionts that provision hosts with limiting nutrients. The best studied such association is that between aphids and Buchnera aphidicola, which produces essential amino acids that are rare in the phloem sap diet. Genomic studies of Buchnera have provided a new means for inferring metabolic capabilities of the symbionts and their likely contributions to hosts. Despite evolutionary reduction of genome size, involving loss of most ancestral genes, Buchnera retains capabilities for biosynthesis of all essential amino acids. In contrast, most genes duplicating amino acid biosynthetic capabilities of hosts have been eliminated. In Buchnera of many aphids, genes for biosynthesis of leucine and tryptophan have been transferred from the chromosome to distinctive plasmids, a feature interpreted as a mechanism for overproducing these amino acids through gene amplification. However, the extent of plasmid-associated amplification varies between and within species, and plasmid-borne genes are sometimes fewer in number than single copy genes on the (polyploid) main chromosome. This supports the broader interpretation of the plasmid location as a means of achieving regulatory control of gene copy number and/or transcription. Buchnera genomes have eliminated most regulatory sequences, raising the question of the extent to which gene expression is moderated in response to changing demands imposed by host nutrition or other factors. Microarray analyses of the Buchnera transcriptome reveal only slight changes in expression of nutrition-related genes in response to shifts in host diet, with responses less dramatic than those observed for the related nonsymbiotic species, Escherichia coli.

Animal Nutritional Physiological Phenomena↗

A specific promoter of the sensory cells of the inner ear defined by transgenesis.

To date, no promoter sequence specific to the inner ear sensory cells (hair cells) has been reported. In an effort to understand the molecular mechanisms that determine hair cell fate in the inner ear, and with the goal of developing a valuable tool for gene therapy and for the generation of conditional knockouts, we initiated a search for cis-acting DNA sequences that regulate the expression of the murine Myo7a and human MYO7A genes. These genes encode the unconventional myosin VIIA which is expressed in hair cells and in some other epithelial cells. We generated lines of transgenic mice expressing the green fluorescent protein (GFP ) reporter gene under the control of several 5'-truncated versions of the Myo7a/MYO7A promoter region and intron 1. We obtained transgenic mice with a GFP expression restricted to the hair cells of the inner ear, cochlea and vestibule. Successive deletions of the promoter allowed us to define a minimal sequence of 118 bp that is sufficient, in the presence of intron 1, to target the transgene expression to hair cells. In addition, the deletion of intron 1 from the transgenes abolished hair cell expression, thus indicating the presence of a strong enhancer in the intron. This is the first report of regulatory sequences sufficient to target the expression of a gene exclusively in sensory cells of the inner ear. It also opens up the possibility for the analysis of the hair cell transcriptome.

Animals↗

A complete and near-perfect rhesus macaque reference genome: lessons from subtelomeric repeats and sequencing bias.

A truly complete, telomere-to-telomere (T2T), and error-free reference genome remains a foundational resource-and long-standing goal-for unbiased comparative and functional genomics. While recent T2T assemblies of humans and other primates have made substantial progress, most still contain thousands of base-level errors, particularly within highly repetitive regions. Here, we present T2T-MMU8v2.0, a near-perfect T2T assembly of the rhesus macaque (Macaca mulatta), representing the highest base-level accuracy reported in a primate genome to date. By employing an optimized ONT-only assembly strategy, we identify subtelomeric satellite-rich regions as the principal bottleneck to improving assembly quality, owing to technological biases in long-read platforms and limitations in current hybrid assembly frameworks. We discover 268 previously unannotated repeat families and resolve ~8 Mbp of SATR satellite arrays, with over 99-fold enrichment in historically misassembled subtelomeric regions. These satellites form four distinct genomic architectures, each with unique SATR satellite composition, segmental duplication organization, and epigenetic signatures, distinct from the subtelomeric architectures observed in hominid genomes. Notably, in contrast to the largely gene-poor subtelomeric regions in African hominids, the SATR architectures in macaques harbor 58 actively transcribed genes, supported by open chromatin and expression data, suggesting gene innovation within these repetitive regions. Functionally, T2T-MMU8v2.0 improves read mappability and accuracy across sequencing platforms, and results in a 19% improvement of transcription start site enrichment scores and 5,821 additional chromatin accessibility peaks on average, thereby enhancing variant detection, regulatory annotation, and transcriptomic resolution in population genetics or single-nucleus studies. Together, this work establishes a new benchmark for genomics, offers a roadmap for resolving complex repetitive regions, and reveals previously unrecognized features of subtelomeric genome structure and evolution.

Journal Article↗

Combining Annotation Software to Identify Orthologous Genes (CASIO) Provides a New Dataset of Orthologous Genes for Swallowtail Butterflies.

With the massive increase in genomic resources, it is becoming increasingly popular to analyse thousands of loci across many species. However, many of the available genomes are not annotated, which hinders an efficient search for orthologous protein-coding genes. Here, we aim to develop a semi-automated pipeline and compare four genomic annotation methods (BRAKER2, BUSCO, Miniprot and Scipio). Our results highlight the importance of integrating multiple annotation tools to optimise ortholog detection and improve genomic studies. Each annotation method showed different strengths. BRAKER2 annotated a substantial number of genes. BUSCO, despite limitations inherent to its reference database, identified a higher number of orthologs. Miniprot exhibited notable flexibility in accommodating diverse protein datasets, whereas Scipio successfully recovered a considerable set of genes that were not detected by the other tools. The combination of these tools allowed for more comprehensive ortholog detection. Taking advantage of this pipeline, we developed a comprehensive dataset of orthologous genes for swallowtail butterflies (Lepidoptera: Papilionidae), called Papilionidae_odb, which will facilitate future studies, especially for a non-model group with abundant genomic data and few transcriptomic resources. We tested Papilionidae_odb by inferring a robust phylogenetic framework for Leptocircini using 142 complete genomes, which improved branch support for some phylogenetic relationships, although challenges remained in resolving relationships within certain species groups, likely due to rapid radiations. Our results highlight the complementary nature of the annotation methods and suggest that combining these tools can yield more accurate results in genomic research. This approach was implemented in a Snakemake workflow called CASIO (Combining Annotation Software to Identify Orthologous genes) and can easily be applied to other non-model groups to improve genomic datasets in diverse taxa where transcriptomic resources are still limited.

Animals↗

Global changes in transcription orchestrate metabolic differentiation during symbiotic nitrogen fixation in Lotus japonicus.

Research on legume nodule metabolism has contributed greatly to our knowledge of primary carbon and nitrogen metabolism in plants in general, and in symbiotic nitrogen fixation in particular. However, most previous studies focused on one or a few genes/enzymes involved in selected metabolic pathways in many different legume species. We utilized the tools of transcriptomics and metabolomics to obtain an unprecedented overview of the metabolic differentiation that results from nodule development in the model legume, Lotus japonicus. Using an array of more than 5000 nodule cDNA clones, representing 2500 different genes, we identified approximately 860 genes that were more highly expressed in nodules than in roots. One-third of these are involved in metabolism and transport, and over 100 encode proteins that are likely to be involved in signalling, or regulation of gene expression at the transcriptional or post-transcriptional level. Several metabolic pathways appeared to be co-ordinately upregulated in nodules, including glycolysis, CO(2) fixation, amino acid biosynthesis, and purine, haem, and redox metabolism. Insight into the physiological conditions that prevail within nodules was obtained from specific sets of induced genes. In addition to the expected signs of hypoxia, numerous indications were obtained that nodule cells also experience P-limitation and osmotic stress. Several potential regulators of these stress responses were identified. Metabolite profiling by gas chromatography coupled to mass spectrometry revealed a distinct metabolic phenotype for nodules that reflected the global changes in metabolism inferred from transcriptome analysis.

Expressed Sequence Tags↗