Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptome annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

scGPA: an LLM-assisted workflow for directional virtual gene perturbation analysis from single-cell transcriptomes.

BACKGROUND: Existing virtual perturbation methods can often infer directional changes by comparing predicted post-perturbation expression profiles with control cells. However, workflows that directly return direction-specific downstream candidate genes together with confidence scores, evidence support and interpretable summaries remain limited. We developed scGPA, an LLM-assisted workflow system for directional single-cell virtual gene perturbation analysis. METHODS: scGPA starts from raw single-cell RNA sequencing data and performs quality control, normalization, dimensionality reduction, clustering and cell-group selection. It then constructs cell-group-specific wild-type regulatory networks using repeated subsampling, principal component regression (PCR)/Ridge-based network inference and CP tensor denoising. Based on these networks, scGPA simulates dose-aware virtual knockdown of the target gene and applies signed perturbation propagation to estimate the magnitude and direction of downstream transcriptional responses. LLM assistance is used for marker-based cell-type annotation, evidence-guided candidate prioritization and user-facing biological summarization. RESULTS: We benchmarked scGPA across five public Perturb-seq datasets and compared its performance with GEARS, scGPT and a random baseline. The overall correct prediction rate of scGPA was 23.0%, exceeding those of GEARS (20.7%), scGPT (15.1%) and the random baseline (13.6%). These results indicate that scGPA achieved a higher correct prediction rate than the two comparator models and the random baseline. We subsequently evaluated scGPA using a public osteosarcoma single-cell dataset and performed qRT-PCR validation in 143B osteosarcoma cells. Among genes with significant experimental changes, scGPA achieved a directional concordance of 76.9%. When all tested downstream genes were counted, 37.0% were directionally correct, 51.9% showed no significant change and 11.1% changed in the opposite direction. CONCLUSIONS: scGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation. By integrating single-cell regulatory network inference, signed virtual perturbation and LLM-assisted interpretation, scGPA supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Single-Cell Gene Expression Analysis↗

Chromosome-level genome assembly of the hemiparasitic Taxillus sutchuenensis (Loranthaceae).

Taxillus sutchuenensis, an ecologically and medicinally important hemiparasitic plant that parasitizes diverse woody hosts, was sequenced to generate a high-quality chromosome-level genome assembly. PacBio HiFi long reads, RNA-seq transcriptome data, and Hi-C data were used to assemble a 406.32 Mb genome anchored onto nine pseudo-chromosomes, with a scaffold N50 of 45.59 Mb. The assembly showed high completeness and accuracy, supported by BUSCO (93.6%) and Merqury QV (70.6) assessments. The LTR Assembly Index (LAI) of 13.98 indicated excellent continuity. A total of 21,795 protein-coding genes were predicted, with 94.46% functionally annotated. Repetitive sequences accounted for 50.05% of the genome, primarily LTR retrotransposons. This genome provides a valuable resource for investigating the evolution, functional genomics, and parasitic mechanisms of hemiparasitic plants.

Genome, Plant↗

Isolation, characterization, and pericycle-specific transcriptome analyses of the novel maize lateral and seminal root initiation mutant rum1.

The monogenic recessive maize (Zea mays) mutant rootless with undetectable meristems 1 (rum1) is deficient in the initiation of the embryonic seminal roots and the postembryonic lateral roots at the primary root. Lateral root initiation at the shoot-borne roots and development of the aerial parts of the mutant rum1 are not affected. The mutant rum1 displays severely reduced auxin transport in the primary root and a delayed gravitropic response. Exogenously applied auxin does not induce lateral roots in the primary root of rum1. Lateral roots are initiated in a specific cell type, the pericycle. Cell-type-specific transcriptome profiling of the primary root pericycle 64 h after germination, thus before lateral root initiation, via a combination of laser capture microdissection and subsequent microarray analyses of 12k maize microarray chips revealed 90 genes preferentially expressed in the wild-type pericycle and 73 genes preferentially expressed in the rum1 pericycle (fold change >2; P-value <0.01; estimated false discovery rate of 13.8%). Among the 51 annotated genes predominately expressed in the wild-type pericycle, 19 genes are involved in signal transduction, transcription, and the cell cycle. This analysis defines an array of genes that is active before lateral root initiation and will contribute to the identification of checkpoints involved in lateral root formation downstream of rum1.

Biological Transport↗

Identifying JAK2 and ANXA5 as Key Genes Linking Obstructive Sleep Apnea and Oxidative Stress via Machine Learning and Multilayer Transcriptomic Integration With Functional Validation.

Obstructive sleep apnea (OSA) is a common and severe sleep disorder closely associated with oxidative stress (OS). This study aims to identify and validate potential OS-related genes associated with OSA through bioinformatics methods. We successfully identified OS-related differentially expressed genes (OS-DEGs) by combining the limma test, weighted correlation network analysis (WGCNA), and OS-related genes from the GeneCards database. Key genes and potential biological roles were further identified using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG), enrichment analysis, protein-protein interaction (PPI) network analysis, Lasso regression analysis, random forest algorithm, and support vector machine recursive feature elimination (SVM-RFE) method. Evaluate and validate the accuracy of key genes through receiver operating characteristic (ROC) curve analysis. The human single-cell RNA sequencing (scRNA-seq) dataset is used for cell classification annotation, analysis of key gene single-cell expression profiles, and virtual gene knockout experiments based on the scTenifoldKnk algorithm. Integrating scRNA-seq sequencing, pseudotime trajectory inference, cell-cell communication analysis, and bulk immune infiltration deconvolution reveals monocyte subtype remodeling in OSA. Finally, the expression levels of key genes in clinical samples were validated using real-time quantitative PCR (RT-qPCR) and Western blotting. A total of 57 common DEGs, indicating significant enrichment in OS, inflammation, and tumor pathways, particularly prominent in the immunometabolism pathway. By integrating DEGs, WGCNA, PPI results, and machine learning methods, key genes Janus kinase 2 (JAK2) and ANXA5 were screened out. JAK2 was significantly upregulated under disease conditions, while ANXA5 was significantly downregulated. ROC curve exhibited high accuracy (area under the curve [AUC] >&#x2009;0.85). Human scRNA-seq analysis revealed that key genes were predominantly highly expressed in monocytes. Virtual knockout experiments demonstrated that these key genes play a crucial role in regulating immune responses and inflammatory reactions. PPI networks and enrichment analysis verified that downstream genes S100P, ALOX5AP, PROK2, and PADI4 may collaboratively participate in immune response and inflammation regulation. Finally, clinical sample experiment further validated the results of bioinformatics analysis. This study provides new research insights for the diagnosis, mechanism research, and treatment development of OSA in the future by integrating multilayer transcriptomic and machine learning techniques.

Humans↗

Transcriptome analysis of the diseased intervertebral disc tissue in patients with spinal tuberculosis.

OBJECTIVE: To investigate the differential expression genes (DEGs) in spinal tuberculosis using transcriptomics, with the aim of identifying novel therapeutic targets and prognostic indicators for the clinical management of spinal tuberculosis. METHODS: Patients who visited the Department of Orthopedics at the Second Hospital, Lanzhou University from January 2021 to May 2023 were enrolled. Based on the inclusion and exclusion criteria, there were 5 patients in the test group and 5 patients in the control group. Total RNA was extracted and paired-end sequencing was conducted on the sequencing platform. After processing the sequencing data with clean reads and annotating the reference genome, FPKM normalization and differential expression analysis were performed. The DEGs and long non-coding RNAs (LncRNAs) were analyzed for Kyoto Encyclopedia of Genes and Genomes (KEGG) and Gene Ontology (GO) enrichment. The cis-regulation of differentially expressed mRNAs (DE mRNAs) by LncRNAs was predicted and analyzed to establish a co-expression network. RESULTS: This study identified 2366 DEGs, with 974 genes significantly upregulated and 1392 genes significantly downregulated. The upregulated genes are associated with cytokine-cytokine receptor interactions, tuberculosis, and TNF-&#x3b1; signaling pathways, primarily enriched in biological processes such as immunity and inflammation. The downregulated genes are related to muscle development, contraction, fungal defense response, and collagen metabolism processes. Analysis of LncRNAs from bone tuberculosis RNA-seq data detected a total of 3652 LncRNAs, with 356 significantly upregulated and 184 significantly downregulated. Further analysis identified 311 significantly different LncRNAs that could cis-regulate 777 target genes, enriched in pathways such as muscle contraction, inflammatory response, and immune response, closely related to bone tuberculosis. There are 51 genes enriched in the immune response pathway regulated by cis-acting LncRNAs. LncRNAs that regulate immune response-related genes, such as upregulated RP11-451G4.2, RP11-701P16.5, AC079767.4, AC017002.1, LINC01094, CTA-384D8.35, and AC092484.1, as well as downregulated RP11-2C24.7, may serve as potential prognostic and therapeutic targets. CONCLUSION: The DE mRNAs and LncRNAs in spinal tuberculosis are both associated with immune regulatory pathways. These pathways promote or inhibit the tuberculosis infection and development at the mechanistic level and play an important role in the process of tuberculosis transferring to bone tissue.

Humans↗

An expressed sequence tag analysis of the chicken reproductive tract transcriptome.

Analysis of the chicken reproductive tract transcriptome is important in comparative biology for analysis of reproductive tract development and evolution. In addition, molecular analysis of the reproductive tract is important for identification of genes affecting fertility in the poultry industry. We sampled the chicken reproductive tract (ovary, oviduct, and testis) transcriptome, generating 5,328 expressed sequence tags that assembled into 4,518 contigs. We identified 475 contigs with no match in the current expressed sequence tag databases or in GenBank. The novel contigs included 31 with no match to the current assembly of the chicken genome, 119 representing spliced transcripts, and 309 that were unspliced. More detailed molecular characterization of the 428 novel contigs present in the assembly will be important to gene discovery and annotation of the chicken and other vertebrate genomes.

Aging↗

Transcriptome variations in human CaCo-2 cells: a model for enterocyte differentiation and its link to iron absorption.

Complete clinical expression of the HFE1 hemochromatosis is very likely modulated by genes linked to duodenal iron absorption, whose level is conditioned by unknown processes taking place during enterocyte differentiation. We carried out a transcriptomic study on CaCo-2 cells used as a model of enterocyte differentiation in vitro. Of the 720 genes on the microarrays, 80, 50, and 56 were significantly down-regulated up-regulated, and invariant during differentiation. With regard to iron metabolism, we showed that HEPH, SLC11A2, SLC11A3, and TF are significantly up-regulated, while ATP7B and SLC39A1 (and SFT) are down-regulated and ACO1, dCYTb, FECH, and FTH1 show constant expression. Ontological annotations highlight the decrease in the expression of cell cycle and DNA metabolism associated genes as well as transcription, protein metabolism, signal transduction, and nucleocytoplasmic transport associated genes, whereas there are increases in the expression of genes linked to cell adhesion, lipid and xenobiotic metabolism, iron transport and homeostasis, and immune response.

Caco-2 Cells↗

Insights into dill (Anethum graveolens) flavor formation via integrative analysis of chromosomal-scale genome, metabolome and transcriptome.

INTRODUCTION: Dill (Anethum graveolens) is a significant medicinal herb belonging to the Apiaceae family. Owing to its high levels of volatile organic compounds (VOCs), dill is commonly utilized for essential oil extraction and medicine purpose. However, the biosynthesis of the crucial VOC in dill remains obscure. OBJECTIVES: Identify the key VOCs related to the flavor formation in dill and dissect the regulatory mechanism of their synthesis. METHODS: The dill chromosomal-level genome was constructed by PacBio HiFi, Hi-C, and BGISEQ second generation sequencing and assembly. The VOCs in dill leaves were identified through GC-MS. The potential mechanism involved in regulating the VOC accumulation in dill flavor formation was analyzed by multi-omics analysis. RESULTS: A 1.17&#xa0;Gb chromosome-scale genome of dill with a contig N50 of 10.78&#xa0;Mb was constructed. A total of 46,538 genes were annotated across 11 assembled chromosomes. Comparative genomics analysis suggested that transposable element insertions, especially LTR-Gypsy, have contributed to the evolution and expansion of the dill genome. The flavor formation of dill was mainly attributed to terpenoids, especially &#x3b1;-phellandrene, &#x3b2;-ocimene, and o-cymene. The contribution of expansion and replication of terpenoid synthesis pathway genes, especially terpene synthase (TPS), to the abundant terpenoid production of dill was identified. Differential gene expression patterns observed at various developmental stages and tissues provided key candidate genes for the regulation of terpenoid synthesis, as well as transcription factors. The different accumulation of esters and aromatics also affected the flavor formation of dill. The key genes implicated in the synthesis of anethole, namely AIS and AMT were further identified. CONCLUSION: This study constructed the chromosome level genome and identified the main VOCs and related key genes in flavor formation of dill, shedding lights on our understanding of terpenoid biosynthesis but also offered guidance for future genetic research on molecular breeding in Anethum graveolens.

Transcriptome↗

ceRNA network of lncRNAs and mRNAs in OSF-to-OSCC progression: Diagnostic biomarkers and functional pathways.

BACKGROUND: Oral submucous fibrosis (OSF) is a chronic potentially malignant disorder that can progress to oral squamous cell carcinoma (OSCC). Although dysregulated non-coding RNAs have been implicated in oral carcinogenesis, the competing endogenous RNA (ceRNA)-mediated regulatory mechanisms underlying OSF-to-OSCC progression remain poorly understood. This study aimed to identify candidate regulatory molecules and construct a putative lncRNA-miRNA-mRNA network associated with malignant transformation. METHODS: Publicly available microarray datasets (GSE117973 and GSE125866) were analyzed to identify differentially expressed genes between OSF and OSCC. Differentially expressed transcripts were classified into mRNAs and lncRNAs based on public transcript annotations. Highly correlated lncRNA-mRNA pairs were identified using Pearson correlation analysis and integrated with multiMiR-supported miRNA-mRNA interactions obtained from public databases to construct a putative ceRNA regulatory network. Functional characterization focused on apoptosis, epithelial-mesenchymal transition (EMT), and immune checkpoint-related pathways. Receiver operating characteristic (ROC) analysis was performed to evaluate diagnostic performance, and selected biomarkers were externally validated using The Cancer Genome Atlas (TCGA) OSCC cohort. RESULTS: Integrated transcriptomic analysis identified several dysregulated mRNAs and lncRNAs associated with OSF-to-OSCC progression. Network analysis highlighted TBC1D3B, RREB1, TEAD3, SREBF1, TMEM41B, FOXK2, and KIAA1958 as prominent hub genes within the putative regulatory network. Functional analyses demonstrated significant associations with apoptosis-, EMT-, and immune checkpoint-related genes, suggesting potential involvement in multiple biological processes contributing to malignant transformation. Several hub genes exhibited strong diagnostic performance, with ROC analysis yielding AUC values ranging from 0.891 to 1.000, indicating excellent discrimination between OSF and OSCC samples. External validation using TCGA further supported the relevance of the identified biomarkers in OSCC. CONCLUSIONS: This study provides a comprehensive transcriptomic framework describing putative lncRNA-miRNA-mRNA regulatory interactions associated with OSF progression to OSCC. The identified hub genes and regulatory networks represent candidate biomarkers for early detection and provide a foundation for future mechanistic and experimental validation. As the proposed ceRNA interactions are computationally inferred, further biological validation is required before clinical application.

RNA, Long Noncoding↗

Genome-wide analysis of transcriptional hierarchy and feedback regulation in the flagellar system of Helicobacter pylori.

The flagellar system of Helicobacter pylori, which comprises more than 40 mostly unclustered genes, is essential for colonization of the human stomach mucosa. In order to elucidate the complex transcriptional circuitry of flagellar biosynthesis in H. pylori and its link to other cell functions, mutants in regulatory genes governing flagellar biosynthesis (rpoN, flgR, flhA, flhF, HP0244) and whole-genome microarray technology were used in this study. The regulon controlled by RpoN, its activator FlgR (FleR) and the cognate histidine kinase HP0244 (FleS) was characterized on a genome-wide scale for the first time. Seven novel genes (HP1076, HP1233, HP1154/1155, HP0366/367, HP0869) were identified as belonging to RpoN-associated flagellar regulons. The hydrogenase accessory gene HP0869 was the only annotated non-flagellar gene in the RpoN regulon. Flagellar basal body components FlhA and FlhF were characterized as functional equivalents to master regulators in H. pylori, as their absence led to a general reduction of transcripts in the RpoN (class 2) and FliA (class 3) regulons, and of 24 genes newly attributed to intermediate regulons, under the control of two or more promoters. FlhA- and FlhF-dependent regulons comprised flagellar and non-flagellar genes. Transcriptome analysis revealed that negative feedback regulation of the FliA regulon was dependent on the antisigma factor FlgM. FlgM was also involved in FlhA- but not FlhF-dependent feedback control of the RpoN regulon. In contrast to other bacteria, chemotaxis and flagellar motor genes were not controlled by FliA or RpoN. A true master regulator of flagellar biosynthesis is absent in H. pylori, consistent with the essential role of flagellar motility and chemotaxis for this organism.

Bacterial Proteins↗

Large-scale analysis of the barley transcriptome based on expressed sequence tags.

To provide resources for barley genomics, 110,981 expressed sequence tags (ESTs) were generated from 22 cDNA libraries representing tissues at various developmental stages. This EST collection corresponds to approximately one-third of the 380,000 publicly available barley ESTs. Clustering and assembly resulted in 14,151 tentative consensi (TCs) and 11 073 singletons, altogether representing 25 224 putatively unique sequences. Of these, 17.5% showed no significant similarity to other barley ESTs present in dbEST. More than 41% of all barley genes are supposed to belong to multigene families and approximately 4% of the barley genes undergo alternative splicing. Based on the functional annotation of the set of unique sequences, the functional category 'Energy' was further analysed to reveal tissue- and stage-specific differences in gene expression. Hierarchical clustering of 362 differentially expressed TCs resulted in the identification of seven major clusters. The clusters reflect biochemical pathways predominantly activated in specific tissues and at various developmental stages. During seed germination glycolysis could be identified as the most predominant biochemical pathway. Germination-specific glycolysis is characterized by the coordinated expression of phosphoenolpyruvate carboxylase and phosphoenolpyruvate carboxykinase, whose antagonistic actions possibly regulate the flux of amino acids into protein biosynthesis and gluconeogenesis respectively. The expression of defence-related and antioxidant genes during germination might be controlled by the ethylene-signalling pathway as concluded from the coordinated expression of those genes and the transcription factors (TF) EIN3 and EREBPG. Moreover, because of their predominant expression in germinating seeds, TF of the AP2 and MYB type are presumably major regulators of germination.

Expressed Sequence Tags↗

Evaluation of the chicken transcriptome by SAGE of B cells and the DT40 cell line.

BACKGROUND: The understanding of whole genome sequences in higher eukaryotes depends to a large degree on the reliable definition of transcription units including exon/intron structures, translated open reading frames (ORFs) and flanking untranslated regions. The best currently available chicken transcript catalog is the Ensembl build based on the mappings of a relatively small number of full length cDNAs and ESTs to the genome as well as genome sequence derived in silico gene predictions. RESULTS: We use Long Serial Analysis of Gene Expression (LongSAGE) in bursal lymphocytes and the DT40 cell line to verify the quality and completeness of the annotated transcripts. 53.6% of the more than 38,000 unique SAGE tags (unitags) match to full length bursal cDNAs, the Ensembl transcript build or the genome sequence. The majority of all matching unitags show single matches to the genome, but no matches to the genome derived Ensembl transcript build. Nevertheless, most of these tags map close to the 3' boundaries of annotated Ensembl transcripts. CONCLUSIONS: These results suggests that rather few genes are missing in the current Ensembl chicken transcript build, but that the 3' ends of many transcripts may not have been accurately predicted. The tags with no match in the transcript sequences can now be used to improve gene predictions, pinpoint the genomic location of entirely missed transcripts and optimize the accuracy of gene finder software.

Animals↗

Ashbya Genome Database 3.0: a cross-species genome and transcriptome browser for yeast biologists.

BACKGROUND: The Ashbya Genome Database (AGD) 3.0 is an innovative cross-species genome and transcriptome browser based on release 40 of the Ensembl developer environment. DESCRIPTION: AGD 3.0 provides information on 4726 protein-encoding loci and 293 non-coding RNA genes present in the genome of the filamentous fungus Ashbya gossypii. A synteny viewer depicts the chromosomal location and orientation of orthologous genes in the budding yeast Saccharomyces cerevisiae. Genome-wide expression profiling data obtained with high-density oligonucleotide microarrays (GeneChips) are available for nearly all currently annotated protein-coding loci in A. gossypii and S. cerevisiae. CONCLUSION: AGD 3.0 hence provides yeast- and genome biologists with comprehensive report pages including reliable DNA annotation, Gene Ontology terms associated with S. cerevisiae orthologues and RNA expression data as well as numerous links to external sources of information. The database is accessible at http://agd.vital-it.ch/.

Databases, Genetic↗

A picture of gene sampling/expression in model organisms using ESTs and KOG proteins.

The expressed sequence tag (EST) is an instrument of gene discovery. When available in large numbers, ESTs may be used to estimate gene expression. We analyzed gene expression by EST sampling, using the KOG database, which includes 24,154 proteins from Arabidopsis thaliana (Ath), 17,101 from Caenorhabditis elegans (Cel), 10,517 from Drosophila melanogaster (Dme), and 26,324 from Homo sapiens (Hsa), and 178,538 ESTs for Ath, 215,200 for Cel, 261,404 for Dme, and 1,941,556 for Hsa. BLAST similarity searches were performed to assign KOG annotation to all ESTs. We determined the amount of gene sampling or expression dedicated to each KOG functional category by each model organism. We found that the 25% most-expressed genes are frequently shared among these organisms. The KOG protein classification allowed the EST sampling calculation throughout the glycolysis pathway. We calculated the KOG cluster coverage and inferred that 50 to 80 K ESTs would efficiently cover 80-85% of the KOG database clusters in a transcriptome project. Since KOG is a database biased towards housekeeping genes, this is probably the number of ESTs needed to include the more commonly expressed genes in these organisms. We also examined a still unaddressed question: what is the minimum number of ESTs that should be produced in a transcriptome project?

Animals↗

Characterization of antirrhinum petal development and identification of target genes of the class B MADS box gene DEFICIENS.

The class B MADS box transcription factors DEFICIENS (DEF) and GLOBOSA (GLO) of Antirrhinum majus together control the organogenesis of petals and stamens. Toward an understanding of how the downstream molecular mechanisms controlled by DEF contribute to petal organogenesis, we conducted expression profiling experiments using macroarrays comprising >11,600 annotated Antirrhinum unigenes. First, four late petal developmental stages were compared with sepals. More than 500 ESTs were identified that comprise a large number of stage-specifically regulated genes and reveal a highly dynamic transcriptional regulation. For identification of DEF target genes that might be directly controlled by DEF, we took advantage of the temperature-sensitive def-101 mutant. To enhance the sensitivity of the profiling experiments, one petal developmental stage was selected, characterized by increased transcriptome changes that reflect the onset of cell elongation processes replacing cell division processes. Upon reduction of the DEF function, 49 upregulated and 52 downregulated petal target genes were recovered. Eight target genes were further characterized in detail by RT-PCR and in situ studies. Expression of genes responding rapidly toward an altered DEF activity is confined to different petal tissues, demonstrating the complexity of the DEF function regulating diverse basic processes throughout petal morphogenesis.

Antirrhinum↗

Gene expression analysis of Tek/Tie2 signaling.

The elaboration of the vasculature during embryonic development involves restructuring of the early vessels into a more complex vascular network. Of particular importance to this vascular remodeling process is the requirement of the Tek/Tie2 receptor tyrosine kinase. Mouse gene-targeting studies have shown that the Tie2-deficient embryos succumb to embryonic death at midgestation due to insufficient sprouting and remodeling of the primary capillary plexus. To identify the underlying genetic mechanisms regulating the process of vascular remodeling, transcriptomes modulated by Tie2 signaling were analyzed utilizing serial analysis of gene expression (SAGE). Two libraries were constructed and sequenced using embryonic day 8.5 yolk sac tissues from Tie2 wild-type and the Tie2-null littermates. After tag extraction, 45,689 and 45,275 SAGE tags were obtained for the Tie2 wild-type and Tie2-null libraries, respectively, yielding a total of 21,376 distinct tags. Close to 62% of the tags were uniquely annotated, whereas 10% of the total tags were unknown. Using semiquantitative PCR, the differential expression of eight genes was confirmed that included Elk3, an important angiogenic switch gene which was upregulated in the absence of Tie2 signaling. The results of this study provide valuable insight into the potential association between Tie2 signaling and other known angiogenic pathways as well as genes that might have novel functions in vascular remodeling.

Animals↗

FREP: a database of functional repeats in mouse cDNAs.

The FREP database (http://facts.gsc.riken.go.jp/FREP/) contains 31 396 RepeatMasker-identified non-redundant variant repeat sequences derived from 16,527 mouse cDNAs with protein-coding potential. The repeats were computationally associated with potential effects on transcriptional variation, translation, protein function or involvement in disease to identify Functional REPeats (FREPs). FREPs are defined by the (i) occurrence of exon-exon boundaries in repeats, (ii) presence of polyadenylation sites in 3'UTR-located repeats, (iii) effect on translation, (iv) position in the protein- coding region or protein domains or (v) conditional association with disease MeSH terms. Currently the database contains 9261 (29.5%) inferred FREPs derived from 6861 (41.5%) mouse cDNAs. Integrated evidence of the functional assignments and dynamically generated sequence similarity search results support the exploration and annotation of functional, ancestral or taxon-specific repeats. Keyword and pre-selected feature searches (e.g. coding sequence-repeat or splice site-repeat relations) support intuitive database querying as well as the retrieval of repeat sequences. Integrated sequence search and alignment tools allow the analysis of known or identification of new functional repeat candidates. FREP is a unique resource for illuminating the role of transposons and repetitive sequences in shaping the coding part of the mouse transcriptome and for selecting the appropriate experimental model to study diseases with suspected repeat etiology contributions.

Animals↗

Comprehensive post-genomic data analysis approaches integrating biochemical pathway maps.

Post-genomic era research is focusing on studies to attribute functions to genes and their encoded proteins, and to describe the regulatory networks controlling metabolic, protein synthesis and signal transduction pathways. To facilitate the analysis of experiments using post-genomic technologies, new concepts for linking the vast amount of raw data to a biological context have to be developed. Visual representations of pathways help biologists to understand the complex relationships between components of metabolic networks, and provide an invaluable resource for the integration of transcriptomics, proteomics and metabolomics data sets. Besides providing an overview of currently available bioinformatic tools for plant scientists, we introduce BioPathAt, a newly developed visual interface that allows the knowledge-based analysis of genome-scale data by integrating biochemical pathway maps (BioPathAtMAPS module) with a manually scrutinized gene-function database (BioPathAtDB) for the model plant Arabidopsis thaliana. In addition, we discuss approaches for generating a biochemical pathway knowledge database for A. thaliana that includes, in addition to accurate annotation, condensed experimental information regarding in vitro and in vivo gene/protein function.

Computational Biology↗