Search PubMedSearch

SEARCH · Search PubMed

Results for “Expressed Sequence Tags”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

prot4EST: translating expressed sequence tags from neglected genomes.

BACKGROUND: The genomes of an increasing number of species are being investigated through generation of expressed sequence tags (ESTs). However, ESTs are prone to sequencing errors and typically define incomplete transcripts, making downstream annotation difficult. Annotation would be greatly improved with robust polypeptide translations. Many current solutions for EST translation require a large number of full-length gene sequences for training purposes, a resource that is not available for the majority of EST projects. RESULTS: As part of our ongoing EST programs investigating these "neglected" genomes, we have developed a polypeptide prediction pipeline, prot4EST. It incorporates freely available software to produce final translations that are more accurate than those derived from any single method. We show that this integrated approach goes a long way to overcoming the deficit in training data. CONCLUSIONS: prot4EST provides a portable EST translation solution and can be usefully applied to >95% of EST projects to improve downstream annotation. It is freely available from http://www.nematodes.org/PartiGene.

Animals

IMAGE cDNA clones, UniGene clustering, and ACeDB: an integrated resource for expressed sequence information.

In this study we describe a new information resource that provides integrated access to information on IMAGE (integrated molecular analysis of genomes and their expression) cDNA library clones and derived expressed sequence tags (ESTs). We have developed an automated procedure that collates data from various public sources into a single ACeDB database. This database is a valuable tool for electronic cloning experiments and gene expression studies. It allows researchers to find information about cDNA libraries, plate addresses, insert sizes, and sequence data for IMAGE clones, the assignment of ESTs to UniGene clusters, and the chromosomal location of those genes in an efficient, graphically oriented manner.

Cloning, Molecular

Gene finding in the chicken genome.

BACKGROUND: Despite the continuous production of genome sequence for a number of organisms, reliable, comprehensive, and cost effective gene prediction remains problematic. This is particularly true for genomes for which there is not a large collection of known gene sequences, such as the recently published chicken genome. We used the chicken sequence to test comparative and homology-based gene-finding methods followed by experimental validation as an effective genome annotation method. RESULTS: We performed experimental evaluation by RT-PCR of three different computational gene finders, Ensembl, SGP2 and TWINSCAN, applied to the chicken genome. A Venn diagram was computed and each component of it was evaluated. The results showed that de novo comparative methods can identify up to about 700 chicken genes with no previous evidence of expression, and can correctly extend about 40% of homology-based predictions at the 5' end. CONCLUSIONS: De novo comparative gene prediction followed by experimental verification is effective at enhancing the annotation of the newly sequenced genomes provided by standard homology-based methods.

Animals

AnoEST: toward A. gambiae functional genomics.

Here, we present an analysis of 215,634 EST and cDNA sequences of a major vector of human malaria Anopheles gambiae structured into the AnoEST database. The expressed sequences are grouped into clusters using genomic sequence as template and associated with inferred functional annotation, including the following: corresponding Ensembl gene prediction, putative orthologous genes in other species, homology to known proteins, protein domains, associated Gene Ontology terms, and corresponding classification into broad GO-slim functional groups. AnoEST is a vital resource for interpretation of expression profiles derived using recently developed A. gambiae cDNA microarrays. Using these cDNA microarrays, we have experimentally confirmed the expression of 7961 clusters during mosquito development. Of these, 3100 are not associated with currently predicted genes. Moreover, we found that clusters with confirmed expression are nonbiased with respect to the current gene annotation or homology to known proteins. Consequently, we expect that many as yet unconfirmed clusters are likely to be actual A. gambiae genes. [AnoEST is publicly available at http://komar.embl.de, and is also accessible as a Distributed Annotation Service (DAS).].

Animals

Conserved function of medaka pink-eyed dilution in melanin synthesis and its divergent transcriptional regulation in gonads among vertebrates.

Medaka is emerging as a model organism for the study of vertebrate development and genetics, and its effectiveness in forward genetics should prove equal to that of zebrafish. Here, we identify by positional cloning a gene responsible for the medaka i-3 albino mutant. i-3 larvae have weakly tyrosinase-positive cells but lack strongly positive and dendritic cells, suggesting loss of fully differentiated melanophores. The region surrounding the i-3 locus is syntenic to human 19p13, but a BAC clone covering the i-3 locus contained orthologs located at 15q11-13, including OCA2 (P). Medaka P consists of 842 amino acids and shares approximately 65% identity with mammalian P proteins. The i-3 mutation is a four-base deletion in exon 13, which causes a frameshift and truncation of the protein. We detected medaka P transcripts in melanin-producing eyeballs and (putative) skin melanophores on embryos and an alternatively spliced form in the non-melanin-producing ovary or oocytes. The mouse p is similarly expressed in gonads, but not alternatively spliced. This is the first isolation of nonmammalian P, the functional mechanism of action of which has not yet been elucidated, even in mammals. Further investigation of the functions of P proteins and the regulation of their expression will provide new insight into body color determination and gene evolution.

Amino Acid Sequence

EST-SSR based genetic polymorphism among Lablab (Lablab purpureus L. Sweet) accessions contrasting for drought stress at seedling stage.

Lablab is a multipurpose and the most drought-tolerant (DT) crop compared with its relatives. Despite its potential, Lablab is still an underutilized crop with a lack of improved varieties in many countries. The DT (D349, D147, HA4, D363, D352, D359, D348, D311, D55 and D250) and drought-susceptible (DS) (D271, D66, D106, D6, D26, D255, D28, D186, D95, and D258) accessions were earlier identified according to their morphological and biochemical responses to moisture stress at the seedling stage. These accessions were used to establish genetic polymorphism among the accessions contrasting for drought stress based on the Expressed Sequence Tag-Simple Sequence Repeats (EST-SSR) markers. The CTAB protocol was employed for the genomic DNA extraction. After DNA quality and quantity verification, the PCR was conducted using 16 EST-SSR primer pairs specific to the Lablab. The products were separated through the horizontal polyacrylamide gel electrophoresis (hPAGE). Discriminating ability of the markers and primers' efficiency were evaluated based on various genetic parameters. Principal Coordinate Analysis (PCoA) was performed to estimate the distance matrix among the population and among the accessions. While cluster analysis was processed to trace the genetic relationship among the accessions, dendrogram was constructed to decipher their genetic relationship. Analysis of Molecular Variance (AMOVA) was finally computed to quantify the diversity level and genetic relationship among the population, and among the accessions. A low polymorphism (GD = 0.19) was observed between the DT and DS accessions, likely due to limited discriminatory power of the EST-SSR markers. However, the PCoA, cluster analysis and AMOVA identified DT (D147, HA4, and D349) and DS (D106, D95, and D271) accessions as strongly contrasting populations under drought stress, with D147, HA4, D349, D363, D359, D352, and D348 further recommended as DT accessions. Given the low polymorphism observed, further validation using more informative molecular markers and advanced genomic approaches is recommended to improve the identification of drought-tolerance genes and related QTLs to support Lablab breeding programs.

Expressed Sequence Tags

Genomic regionality in rates of evolution is not explained by clustering of genes of comparable expression profile.

In mammalian genomes, linked genes show similar rates of evolution, both at fourfold degenerate synonymous sites (K4) and at nonsynonymous sites (KA). Although it has been suggested that the local similarity in the synonymous substitution rate is an artifact caused by the inclusion of disparately evolving gene pairs, we demonstrate here that this is not the case: after removal of disparately evolving genes, both (1) linked genes and (2) introns from the same gene have more similar silent substitution rates than expected by chance. What causes the local similarity in both synonymous and nonsynonymous substitution rates? One class of hypotheses argues that both may be related to the observed clustering of genes of comparable expression profile. We investigate these hypotheses using substitution rates from both human-mouse and mouse-rat comparisons, and employing three different methods to assay expression parameters. Although we confirm a negative correlation of expression breadth with both K4 and KA, we find no evidence that clustering of similarly expressed genes explains the clustering of genes of comparable substitution rates. If gene expression is not responsible, what about other causes? At least in the human-mouse comparison, the local similarity in KA can be explained by the covariation of KA and K4. As regards K4, our results appear consistent with the notion that local similarity is due to processes associated with meiotic recombination.

Animals

Molecular investigation of the progenitors, origin and domestication patterns of diploid Chinese old garden roses.

BACKGROUND AND AIMS: Chinese old garden roses are major contributors to the genetic development of modern roses. The RoKSN gene is associated with continuous flowering in roses and is proposed to have originated from Chinese wild roses. However, the wild roses that are implicated in the breeding of Chinese old garden roses and the origin of the RoKSN locus remain unidentified. We collected 25 of the most renowned and classic diploid Chinese old garden roses along with all related wild roses from East Asia. These roses were analysed with the aim of identifying the wild species that contributed to the genetic composition of Chinese old garden roses. In addition, we aimed to infer the geographical origin of the RoKSN gene and to develop a schematic overview of hybrid domestication of Chinese old garden roses. METHODS: We compared the haplotypes of internal transcribed spacers (nrITS), six nuclear single-copy genes and three chloroplast genes between Chinese old garden roses and wild roses. Additionally, we assessed genetic organization using 21 expressed sequence tag-simple sequence repeats to identify potential donor species that contributed to the emergence of these cultivars. Primers were designed for RoKSN to allow comparison of the gene across the entire distribution range of Rosa sect. Chinenses. KEY RESULTS: Our findings confirmed that the majority of rose cultivars are descendants of early hybridization events. Rosa chinensis var. spontanea, R. odorata var. gigantea and R. multiflora var. cathayensis were the primary donors for the 25 cultivar roses. Chinese old garden roses were categorized into four groups. Ten cultivars were hybrids between R. chinensis var. spontanea and R. multiflora var. cathayensis, thereby forming the 'Old Blush' group. Five cultivars were hybrids between 'Old Blush' and the R. kwangtungensis species complex, thereby forming the 'Slater's crimson' group. Six cultivars were hybrids between 'Old Blush' and R. odorata var. gigantea, thereby forming the 'Tea Rose' group, and three cultivars were hybrids that evolved from more than three donors. Moreover, we observed relatively close genetic proximity among Chinese old garden roses with an identical RoKSN-copia gene that is responsible for continuous flowering, which indicates a single origin for this retrotransposon-containing allele. Additionally, we determined that the haplotypes of the RoKSN-copia gene predominantly occurred in the Sichuan Basin region. In contrast, R. chinensis cultivated in the Ya'an region showed no markers of hybridization and displayed a genetic composition that was close to that of the wild species R. chinensis var. spontanea. This cultivar may represent the earliest mutated individual that bears the RoKSN-copia gene and may have served as a bridge from wild species to continuous-flowering old rose cultivars. CONCLUSIONS: The study provides crucial evidence that elucidates the origin of cultivated roses and lays the groundwork for further analysis of the breeding history of Chinese old garden roses using genomic data.

Domestication

Laforin targets malin to glycogen in Lafora progressive myoclonus epilepsy.

Glycogen is the largest cytosolic macromolecule and is kept in solution through a regular system of short branches allowing hydration. This structure was thought to solely require balanced glycogen synthase and branching enzyme activities. Deposition of overlong branched glycogen in the fatal epilepsy Lafora disease (LD) indicated involvement of the LD gene products laforin and the E3 ubiquitin ligase malin in regulating glycogen structure. Laforin binds glycogen, and LD-causing mutations disrupt this binding, laforin-malin interactions and malin's ligase activity, all indicating a critical role for malin. Neither malin's endogenous function nor location had previously been studied due to lack of suitable antibodies. Here, we generated a mouse in which the native malin gene is tagged with the FLAG sequence. We show that the tagged gene expresses physiologically, malin localizes to glycogen, laforin and malin indeed interact, at glycogen, and malin's presence at glycogen depends on laforin. These results, and mice, open the way to understanding unknown mechanisms of glycogen synthesis critical to LD and potentially other much more common diseases due to incompletely understood defects in glycogen metabolism.

Animals

Bioprospecting microbial genomes to expand the biocatalytic toolbox of rubber oxygenases.

A set of rubber oxygenases was discovered through phylogenetic analysis and AI-based structural modeling of complexes of the putative enzymes with a substrate mimicking cis-1,4-polyisoprene. Sixteen candidate proteins were selected from thermophilic microorganisms, all sequence-related to the Latex clearing protein from Streptomyces sp. K30 (LcpK30). Sequence truncation and solubility tags were then evaluated to enhance protein expression, with the SUMO tag proving to be the most effective. Including LcpK30, nine heme-containing oxygenases were successfully expressed in E. coli NEB 10-beta cells, purified (35-157 mg L-1 yield) and characterized. Steady-state kinetics revealed significant rubber latex-degrading properties for six of them, with the truncated SUMO-fused LcpK30 (SUMO-LcpK30T) showing activity in agreement with literature. Notably, the catalytic efficiencies of all the expressed homologs lay within one order of magnitude and the oxygenase from Thermomonospora echinospora was found to be particularly promising in terms of activity, especially at high latex concentrations (more than 1% w/v). The analysis of reaction mixtures by both HPLC and HPLC-MS confirmed the oxidation of cis-1,4-polyisoprene to form the expected isoprenoid oligomers (n = 2-12), whose distribution was consistent with the usual endo-type cleavage pattern in all but one case. This bioprospecting effort afforded a platform of new rubber-degrading enzymes with diverse efficiencies and product profiles, capable of adapting to targeted applications.

Oxygenases

Gene Cloning, Expression, and Purification of Kunitz Trypsin Inhibitor from Glycine max Using Halo Tag.

Soybean Kunitz Trypsin Inhibitor (SKTI) is one of the most extensively studied protease inhibitors, with applications in pest management, medicine, the food processing industry, and the leather industry. In this study, SKTI was cloned into the pFN29A Flexi vector containing a barnase gene. Genomic DNA was isolated from tender soybean leaves, and SKTI was amplified by PCR to obtain a 671 bp product. After cloning, an internal 380 bp sequence was amplified using specific primers to confirm that the cloned sequence was a functional SKTI, as non-functional SKTI genes also exist in Glycine max. The amplified PCR product, containing an AsiSI site at the 5' end and a PmeI site at the 3' end, was cloned into the pFN29A vector. The resulting colonies were screened by colony PCR, and the insert sequence was confirmed by Sanger sequencing. The recombinant protein, containing a His-tag, Halo-tag, and a TEV protease cleavage site, was expressed in Escherichia coli BL21 cells. Maximum expression was achieved 5 h after induction with 0.5 mM IPTG at 37 °C. The expressed SKTI was purified using affinity chromatography on HaloLink resin, and the bound SKTI was cleaved with HaloTEV protease to obtain pure SKTI. The purified inhibitor effectively inhibited bovine trypsin, with an IC₅₀ of 0.6 ± 0.003 µg/µl, yielding 1.6 mg per gram of bacterial pellet. The 24 kDa inhibitor remained stable up to a temperature of 50 °C. Kinetic analysis revealed that recombinant SKTI competitively inhibits trypsin, with a Kᵢ value of 14 µM.

Cloning, Molecular

Generation of Hoxa11-3XFLAG and Hoxd11-3XFLAG alleles to investigate Hox11 genome-wide binding.

Hox genes encode for evolutionary conserved transcription factors that direct the proper development of the body plan. Despite decades of research, little is known regarding their downstream target genes, especially in vertebrates. The strong evolutionary conservation of their DNA-binding homeodomain, their generic AT-rich binding sites, and the lack of specific antibodies has precluded rigorous examination. To circumvent these limitations, we have generated two mouse models in which a 3XFLAG epitope tag has been inserted into the 5' end of the coding sequence of both Hoxa11 and Hoxd11 loci via Cas9/CRISPR. The alleles have been validated by sequencing, PCR genotyping, western blotting, and protein expression analyses, demonstrating proper targeting and expression. Breeding these alleles in combination produces viable and fertile Hoxa11FLAG/FLAG; Hoxd11FLAG/FLAG animals, with no overt patterning defects unlike Hoxa11/Hoxd11 mutants that are infertile and have severe kidney and limb defects. By performing CUT&RUN and CUT&Tag analyses, we have confirmed DNA binding to a known Six2 enhancer in the developing kidney. These novel alleles will allow characterization of the genome-wide binding profile of Hox11 proteins in vivo.

Animals

Multi-Omics and Integrative Analytics in Natural Products Discovery.

Natural products (NPs) have long been an essential source of new bioactive compounds for drug discovery; however, traditional methods for screening and isolating these compounds can be slow and often yield diminishing returns. Fortunately, advanced multi-omics and computational approaches present powerful solutions to these challenges. This review highlights innovative methodologies that integrate metabolomics, genomics, transcriptomics, and proteomics with bioinformatics and analytical chemistry to accelerate NP discovery. For instance, untargeted metabolomics platforms like high-resolution liquid chromatography-tandem mass spectrometry (LC-MS/MS) and Global Natural Products Social (GNPS) molecular networking allow for comprehensive profiling of new compounds, while targeted isotope-labeling strategies enhance this process. Additionally, genome and metagenome mining tools such as antibiotics and secondary metabolite analysis shell (antiSMASH), Deep Biosynthetic Gene Cluster (DeepBGC), and Pipeline for Reconstructing Integrated Syntheses of Metabolites (PRISM) quickly identify biosynthetic gene clusters (BGCs) in both cultured and uncultured organisms, often using heterologous expression to validate products. Transcriptomic analyses, including RNA sequencing (RNA-seq), co-expression networks, and fluxomics, help clarify how pathways are regulated, while quantitative proteomics techniques like tandem mass tags/isobaric tags for relative and absolute quantitation (TMT/iTRAQ) and label-free methods, along with chemoproteomics approaches such as cellular thermal shift assay and thermal proteome profiling (TPP), uncover molecular targets and their mechanisms of action. This review also places significant emphasis on the role of artificial intelligence (AI) and machine learning (ML) in integrating multi-omics data, spanning activities from constructing gene-metabolite correlation networks to leveraging knowledge graphs and graph neural networks for data fusion and functional prediction. Finally, this review concludes by discussing the synergistic benefits of multi-omics for natural-product discovery, addressing current technical challenges, and exploring future directions toward high-throughput, intelligent data integration for next-generation NP research.

Biological Products

Marine-Inspired Antimicrobial Peptides Disrupt Gene Expression at the DNA Level.

Genome mining of Streptomyces sp. H-KF8 combined with sequence engineering yielded two serum-stable, noncytotoxic, nonlytic antimicrobial peptides, L3 and L3-K. Initial studies in uropathogenic Escherichia coli suggested membrane effects and nucleoid relaxation, prompting a comprehensive investigation of their mode of action. In this study tandem mass tag (TMT)-based quantitative proteomics revealed extensive proteome remodeling, with 175 and 120 differentially expressed proteins (DEPs) after treatment with L3 and L3-K, respectively. L3 induced predominantly upregulated responses linked to metabolism, RNA processing, transport, and homeostasis, whereas L3-K mainly caused the downregulation of proteins involved in metabolism, transport, and cell structure. Both peptides disrupted ABC transporter-mediated nutrient uptake and elicited stress responses, while L3 specifically perturbed the mal regulon, indicative of broader transcriptional dysregulation. Complementary fluorescent dye displacement and in vitro transcription/translation assays demonstrated nonspecific DNA binding, stronger for L3 than L3-K, and potent inhibition of transcriptional and translational processes. Strikingly, inhibitory concentrations paralleled their minimum inhibitory concentrations, directly linking DNA binding and interference with central information processing to antimicrobial activity. These findings reveal that L3 and L3-K primarily act by targeting DNA and interfering with the transcription-translation machinery. Beyond offering mechanistic insights, this study underscores peptides' potential to act as scaffolds for next-generation antimicrobial peptides with DNA-binding and nonmembrane-lytic activity.

Antimicrobial Peptides

Designing an optimized strategy for extracellular expression of recombinant human TNF-α in Escherichia coli.

Extracellular protein expression in Escherichia coli is an elegant solution that addresses the complex issue of protein misfolding while simultaneously simplifying downstream processing steps. Human TNF-α was chosen as the target protein for export since it is a therapeutically important cytokine. Different genomic knockouts were tested for the ability to sustain and enhance protein expression, and BW25113 Δ(elaA + cysW) knockout was found to give a sustained and high level of expression. To improve secretion, various tags were tested, and the MBP tag at the N-terminal end was found to give maximum enhancement in the export of hTNF-α. Even the linker peptide was found to play a critical role in export, with the Ek linker giving the highest extracellular secretion, while the intein sequence completely blocked export. The co-expression of pSecAB, which is involved in protein transport to the periplasm, was also found to be helpful in enhancing extracellular protein titers. Interestingly, pelB performed poorly as compared to the native signal sequence of MBP, which gave better results. Culture conditions were optimized, and it was observed that growing cells in TB medium at a temperature of 25 °C, coupled with a pulse of concentrated nutrients at 24 h, led to a very high extracellular accumulation of ∼1.3 g/L of MBP-hTNF-α in shake flask culture. The protein was purified and tested using L929 cells for bioactivity. Thus, a combination of genomic and bioprocess strategies allowed us to obtain high levels of soluble and active extracellular expression of hTNF-α, making this a very attractive strategy for protein production.

E. coli

The piRNA pathway mediates transcriptional silencing of LTR retrotransposons in ovaries and somatic tissues of Aedes mosquitoes.

The PIWI-interacting RNA (piRNA) pathway preserves genomic integrity by suppressing transposable elements in animal germlines. Despite its well-established function in the animal germline, piRNAs and PIWI proteins are expressed in somatic tissues across arthropod species, and their functions outside the gonads remain poorly understood. Aedes albopictus mosquitoes express four PIWI genes, Piwi4, Piwi5, Piwi6, and Ago3, in both gonadal and somatic tissues. Here, we generated Piwi6 knockout (KO) Ae. albopictus cell lines and observed a substantial upregulation of long terminal repeat retrotransposons, including a full-length endogenous retrovirus that we named Aedes albopictus Endogenous Retrovirus-1 (AalERV1). Nascent RNA sequencing and Cleavage Under Targets and Tagmentation (CUT&Tag) analyses revealed that Piwi6 silences AalERV1 transcriptionally by guiding the deposition of the repressive H3K9me3 histone mark. Consistently, Piwi6 localized to both the cytoplasm and nucleus, with sequences in the intrinsically disordered region guiding nuclear translocation. Reintroduction of full-length GFP-Piwi6, but not a mutant GFP-Piwi6 defective in nuclear localization, rescued AalERV1 repression in Piwi6 KO cells. Importantly, Piwi6-mediated control of AalERV1 was recapitulated in vivo as Piwi6 knockdown increased AalERV1 expression in both ovaries and somatic tissues of Ae. albopictus mosquitoes. These results establish Aedes mosquitoes as a model to study nuclear PIWI functions and suggest that somatic piRNA-mediated transposon silencing is evolutionarily conserved across arthropod species.

Animals

MLL4 protects cardiomyocytes against ischemia-reperfusion injury through STAT3-mediated mitochondrial function.

Myocardial ischemia-reperfusion injury (MIRI) is an inevitable pathophysiological response during the revascularization process following myocardial ischemia. Despite its clinical significance, effective targeted therapies for MIRI remain an unmet medical need. Mixed-lineage leukemia 4 (MLL4), a member of the SET family of histone methyltransferases, exhibits particular methyltransferase action toward histone H3 lysine 4 (H3K4). This study establishes a protective role for MLL4 in MIRI pathogenesis. Utilizing cardiomyocyte-specific Mll4 knockout mice and an in vivo ischemia-reperfusion (I/R) model induced by left anterior descending coronary artery ligation, we observed significant upregulation of MLL4 expression in cardiac tissue following I/R. Genetic ablation of Mll4 in cardiomyocytes markedly exacerbated both acute and chronic phases of MIRI. In vitro, Mll4 knockdown in neonatal rat cardiomyocytes (NRCMs) amplified mitochondrial dysfunction and apoptosis under hypoxia/reoxygenation (H/R) conditions. Integrated analysis of Cleavage Under Targets and Tagmentation sequencing (CUT&Tag-seq) and RNA sequencing (RNA-seq) revealed that Mll4 deficiency induces a pronounced reduction in H3K4 monomethylation (H3K4me1) and histone H3 lysine 27 acetylation (H3K27ac) enrichment at the Stat3 genomic locus. Mechanistically, MLL4 functions as a transcriptional activator of Stat3 by depositing H3K4me1 and H3K27ac, thereby facilitating STAT3 transcription. This regulatory cascade ultimately governs STAT3-dependent mitochondrial homeostasis. Collectively, these findings identify MLL4 as a critical epigenetic regulator of MIRI and suggest its therapeutic targeting may offer a promising strategy for mitigating reperfusion injury.

Animals