Search PubMedSearch

SEARCH · Search PubMed

Results for “Gene Library”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Expansion of the functional genomics GRACE library reveals genes relevant for temperature-dependent fitness in Candida albicans.

A small percentage of species in the fungal kingdom can cause devastating infections in humans, with Candida albicans reigning as a leading cause of systemic disease. One of the key virulence phenotypes for pathogenic fungi is the ability to survive at host body temperature; however, a comprehensive understanding of the mechanisms that orchestrate thermal adaptation in fungi remains incomplete. In this study, we expand the largest functional genomics resource in C. albicans, reaching 71.3% coverage of the entire genome, and perform screens under six different temperatures to identify genes important for temperature-dependent fitness. We describe the function of genes involved in translation (GAR1), splicing (C1_11680C or YSF3), and cell cycle progression (C6_00110C or RHT1) in enabling fungal survival at both low and high temperatures. Through experimental evolution, we also show that C. albicans can rapidly overcome deleterious mutations and adapt to extreme temperature environments. Overall, our study highlights the transformative potential of genome-wide functional genomics to uncover critical vulnerabilities in pathogenic fungi.

Genomics

High-throughput recovery of integron cassettes for gene discovery screens.

Integrons capture functional genes in mobile genetic elements called integron cassettes, which represent an untapped source of genes of biotechnological interest. Here we present two tools, cassette gatherer and cassette hunter, that enable high-throughput establishment of gene libraries either from genetically tractable strains or directly from DNA. We re-engineered a class 1 integron into counterselection markers on a plasmid or on the chromosome of a naturally competent Vibrio cholerae, which enabled capture of single cassettes in a sequence- and function-independent manner. When applied to Vibrio strains and genomic libraries, our tools recovered hundreds of single cassettes per assay with more than 99% specificity. We further subjected the library of cassettes generated by the hunter and gatherer tools to screens against phages ICP2 and T4, and identified nine phage-defence systems, including five previously undescribed. These tools enable rapid and large-scale recovery of integron cassettes that could be leveraged for functional gene discovery.

Journal Article

Broadening the heterologous cross-neutralizing antibody inducing ability of porcine reproductive and respiratory syndrome virus by breeding the GP4 or M genes.

Porcine reproductive and respiratory syndrome virus (PRRSV) is one of the most economically important swine pathogens, which causes reproductive failure in sows and respiratory disease in piglets. A major hurdle to control PRRSV is the ineffectiveness of the current vaccines to confer protection against heterologous strains. Since both GP4 and M genes of PRRSV induce neutralizing antibodies, in this study we molecularly bred PRRSV through DNA shuffling of the GP4 and M genes, separately, from six genetically different strains of PRRSV in an attempt to identify chimeras with improved heterologous cross-neutralizing capability. The shuffled GP4 and M genes libraries were each cloned into the backbone of PRRSV strain VR2385 infectious clone pIR-VR2385-CA. Three GP4-shuffled chimeras and five M-shuffled chimeras, each representing sequences from all six parental strains, were selected and further characterized in vitro and in pigs. These eight chimeric viruses showed similar levels of replication with their backbone strain VR2385 both in vitro and in vivo, indicating that the DNA shuffling of GP4 and M genes did not significantly impair the replication ability of these chimeras. Cross-neutralization test revealed that the GP4-shuffled chimera GP4TS14 induced significantly higher cross-neutralizing antibodies against heterologous strains FL-12 and NADC20, and similarly that the M-shuffled chimera MTS57 also induced significantly higher levels of cross-neutralizing antibodies against heterologous strains MN184B and NADC20, when compared with their backbone parental strain VR2385 in infected pigs. The results suggest that DNA shuffling of the GP4 or M genes from different parental viruses can broaden the cross-neutralizing antibody-inducing ability of the chimeric viruses against heterologous PRRSV strains. The study has important implications for future development of a broadly protective vaccine against PRRSV.

Animals

PTC is a novel rearranged form of the ret proto-oncogene and is frequently detected in vivo in human thyroid papillary carcinomas.

We recently detected a novel activated oncogene by transfection analysis on NIH 3T3 cells in five out of 20 primary human thyroid papillary carcinomas and in the available lymph node metastases. We designated this transforming gene PTC (for papillary thyroid carcinoma). Here we describe the molecular cloning and sequencing of the gene. The new oncogene resulted from the rearrangement of an unknown amino-terminal sequence to the tyrosine kinase domain of the ret proto-oncogene. This gene rearrangement was detected in all of the transfectants and in all of the original tumor DNAs, but not in normal DNA of the same patients, thus indicating that this genetic lesion occurred in vivo and is specific to somatic tumors. Moreover, the transcript coded for by the fused gene was detected in an additional PTC-positive human papillary carcinoma for which mRNA was available.

Amino Acid Sequence

SUMO modification of the Ets-related transcription factor ERM inhibits its transcriptional activity.

A variety of transcription factors are post-translationally modified by SUMO, a 97-residue ubiquitin-like protein bound covalently to the targeted lysine. Here we describe SUMO modification of the Ets family member ERM at positions 89, 263, 293, and 350. To investigate how SUMO modification affects the function of ERM, Ets-responsive intercellular adhesion molecule 1 (ICAM-1) and E74 reporter plasmids were employed to demonstrate that SUMO modification causes inhibition of ERM-dependent transcription without affecting the subcellular localization, stability, or DNA-binding capacity of the protein. When the adenoviral protein Gam1 or the SUMO protease SENP1 was used to inhibit the SUMO modification pathway, ERM-dependent transcription was de-repressed. These results demonstrate that ERM is subject to SUMO modification and that this post-translational modification causes inhibition of transcription-enhancing activity.

Adenoviridae

OligoSeq: Rapid nanopore-sequencing of single-stranded oligonucleotides.

Nanopore-based DNA sequencing technology has achieved remarkable success in sequencing increasingly long DNA strands (e.g., over a million nucleotides long) for genomics research and biotechnology applications. However, the same level of progress has not been achieved for DNA oligonucleotides (usually ≤ 300 nucleotides long). Oligonucleotides play a crucial role in genome engineering efforts through oligo library generation and in DNA data storage, where they are used to encode computer information, such as binary (digital) data in DNA libraries. To enable these applications, accurate sequencing of oligonucleotides in a way that allows to assess for sequence variability, quality and length is essential. But sequencing solutions for oligonucleotides - particularly DNA primers for PCR, oligo DNA libraries used for mutagenesis or cDNA libraries used in gene expression analysis - remain inadequate. To address this gap, OligoSeq is presented as an innovative approach that integrates two complementary techniques: AmpliSeq (based on PCR) and RevSeq (based on reverse complementation with sequence-specific or random primers) to facilitate sequencing of single-stranded oligonucleotides using reference sequence anchor matches of more than ≥ 90% identity spanning from about 70% to 10% with AmpliSeq or RevSeq with random nonamers, respectively, and resolving the final reference sequence based on the most likely candidate from basecall frequencies, regardless of length and double-stranding method. OligoSeq can be integrated with nanopore sequencing technology pipelines and can be used as a reference for other sequencing platforms requiring double-stranded adapters, offering a practical and scalable alternative for standard quality control in single-stranded oligonucleotide synthesis. The use of nanopore technology, compatible with the double-stranding methods showcased, is shown to be the most cost-effective method for resolving original DNA sequences of different length and quality, and to assess its sequence variability, compared to other methods such as Illumina, PacBio or HPLC/MS.

Sequence Analysis, DNA

Host-aware Identification of Intrinsic Gene Expression Biopart Parameters using Combinatorial Libraries.

Model-based design in synthetic biology is limited because bioparts are typically characterised by relative metrics that vary across genetic and physiological contexts. To address this, we introduce a host-aware framework for quantitatively characterising bioparts in combinatorial libraries of plasmid-based constitutive expression constructs. The approach integrates a digital twin of Escherichia coli, conditioned on measured growth rate, with model-in-the-loop parameter identification to separate biopart-associated properties from host-dependent effects. Using structured combinatorial libraries, we identify mechanistically interpretable, transferable parameters for plasmid origins, promoters and ribosome binding sites. In particular, we define an intrinsic translation initiation capacity that captures the dominant RBS-associated contribution to translation while context-dependent expression emerges from host physiology and local sequence context. The resulting parameterisation accurately predicts protein synthesis across physiological conditions, supports incremental library expansion, and reveals localised failures of modularity, providing a scalable foundation for predictive host-aware design in synthetic biology.

Escherichia coli

Using Prime Editing Guide Generator (PEGG) for high-throughput generation of prime editing sensor libraries.

Prime editing enables the generation of nearly any small genetic variant. However, the process of prime editing guide RNA (pegRNA) design is challenging and requires automated computational design tools. We developed Prime Editing Guide Generator (PEGG), a fast, flexible, and user-friendly Python package that enables the rapid generation of pegRNA and pegRNA-sensor libraries. Here, we describe the installation and use of PEGG (https://pegg.readthedocs.io) to rapidly generate custom pegRNA-sensor libraries for use in high-throughput prime editing screens.

Gene Editing

Using Chromosome Conformation Capture Combined with Deep Sequencing (Hi-C) to Study Genome Organization in Bacteria.

Genome organization is fundamental to all living organisms. Long DNA molecules are organized in hierarchical orders to be accommodated into eukaryotic nuclei or bacterial cells, which are thousands of folds shorter. Over the past two decades, chromosome conformation capture (3C) techniques substantially advanced our understanding of genome folding inside cells. 3C involves crosslinking and proximity ligation, and quantifies the physical contacts between two DNA regions within the genome. Coupled with high-throughput sequencing, 3C-seq and Hi-C techniques detect genome-wide DNA interactions, providing a comprehensive view of global genome organization. Here, we describe a detailed method to prepare Hi-C libraries using Bacillus subtilis, which includes procedures of crosslinking chromatin, digesting the crosslinked genome, labeling DNA ends with biotin, ligating DNA, and preparing the DNA library for sequencing using an Illumina platform.

High-Throughput Nucleotide Sequencing

HiChIP for Plant Tissues.

While most epigenomics studies are based on a linear view of genome organization, the necessity to take the three-dimensional chromatin folding into account to understand transcriptional regulation is now clearly recognized. In the past years, approaches combining proximity-based ligation with high-throughput sequencing have opened the way to study long/short-range chromatin interactions and, thus, to analyze 3D chromatin organization. Among them, HiChIP, a protein-based method to capture chromatin interactions, gave rise to the most comprehensive view of the chromatin contacts involving specific chromatin components in a given system. Here, we describe a detailed procedure to produce HiChIP libraries starting from plant tissues.

Chromatin

RNA Sequencing Protocols for Short-Read Sequencing.

RNA sequencing (RNA-seq) methodologies allow the discovery of novel variants and transcripts. These comprise three general steps: (1) capture of RNA species of interest, (2) conversion of RNA to complementary DNA (cDNA), and (3) modification of cDNA to fit the sequencing platform. Here we describe four different library preparation protocols for short-read sequencing: cDNA synthesis with poly(A) selection, library preparation with ribosomal depletion, and cDNA synthesis with SMART® (Switching Mechanism at 5' end of RNA Template) technology for low and Pico inputs.

Gene Library

Whole-Genome Bisulfite Sequencing with a Small Amount of DNA.

Whole-genome bisulfite sequencing (WGBS) is the most widely used method to study DNA methylation profiles across the genome. Since the bisulfite reaction causes DNA degradation, a new approach called post-bisulfite adapter tagging (PBAT) was developed to overcome this problem by adding adapters after bisulfite treatment. In mammals, the PBAT method is used for single-cell bisulfite sequencing (scBS-seq), which enables DNA methylation analysis using a very small amount of DNA from only a few cells, including single-cell input. This protocol involves bisulfite conversion, followed by preamplification and tagging with random hexamer primers prior to Illumina library preparation. Since many procedures are completed in one single test tube, the loss of DNA can be minimized, enabling highly sensitive experiments to study DNA methylation profiles from a very small amount of input material.

Sulfites

A trainable language model with potential to modulate translation rates in non-model organisms by generating upstream untranslated region sequence libraries.

Tuning protein expression in non-model organisms is often constrained by the lack of validated genetic parts and predictive design tools. Translational tuning through the modulation of upstream untranslated regions (5'-UTRs) offers a potentially organism-agnostic route, but existing methods typically rely on mechanistic assumptions, prior knowledge that may not be available in non-model contexts, or the screening of sequence libraries. Here, we present a simple generative approach for creating synthetic 5'-UTR libraries based solely on the genomic sequence statistics of any desired organism. The method uses a sliding-window n-gram language model applied to native 5'-UTR sequences to produce novel sequences that preserve organism-specific base distributions and motifs without hard-coding specific motifs or mechanistic rules into inflexible statistical templates. We have applied this approach to the model bacterium Escherichia coli and the non-model probiotic Limosilactobacillus reuteri. Libraries of approximately 1,000 sequences were generated for each organism, from which about 100 unique sequences were experimentally tested for translation of a fluorescent reporter protein. In both organisms, the synthetic libraries yielded a broad range of translation levels from this relatively small number of tested variants. Sequences derived from an organism's own genomic statistics provided a more uniformly distributed range of translation rates in that organism than sequences derived from the other species. Correlations of individual sequence performance across the two species were weak, and thermodynamic predictions of ribosome binding strength showed very little predictive power, especially in the non-model L. reuteri. The results demonstrate that simple statistical language model approaches applied to genomic data can generate functional translational regulatory sequence libraries without detailed mechanistic knowledge or explicit reference to consensus motifs. The approach requires minimal computational resources, avoids reproducing native sequences, and can be readily applied to any organism with a sequenced genome. This strategy may lower technical barriers to expression tuning in non-model organisms.

5' Untranslated Regions

Systematic performance evaluation and application validation of an end-to-end NGS workstation.

Next-generation sequencing (NGS) library preparation is a core component of precision genomics, but it is commonly constrained by inefficiency, variability, and low throughput of manual protocols. To address these limitations, we developed and systematically evaluated a fully automated NGS workstations and further validated its performance across representative application scenarios. The automated system reduced total processing time from 8 to 10 to 4–6 h. At the same time, it maintained similar performance in pre-library metric, including DNA yield and fragment size, as well as post-capture sequencing metrics (Q30 > 90%, mapping rates > 95%, on-target rates 85–90%). The duplication rate was reduced to 5–8%, compared with 10–15% for manual methods, indicating increased library complexity. Bioinformatic evaluation of inter-species read mapping showed minimal cross-contamination, with a maximum contamination ratio of 0.0003%, indicating effective sample isolation in the automated workflow. High concordance in variant detection was observed between automated and manual workflows. Overall, this automated workstation provides a standardized and reproducible workflow that supports scalable precision genomics applications.

High-Throughput Nucleotide Sequencing

Deep learning guided programmable design of Escherichia coli core promoters from sequence architecture to strength control.

Core promoters are essential regulatory elements that control transcription initiation, but accurately predicting and designing their strength remains challenging due to complex sequence-function relationships and the limited generalizability of existing AI-based approaches. To address this, we developed a modular platform integrating rational library design, predictive modelling, and generative optimization into a closed-loop workflow for end-to-end core promoter engineering. Conserved and spacer region of core promoters exert distinct effects on transcriptional strength, with the former driving large-scale variation and the latter enabling finer gradation. Based on this insight, Mutation-Barcoding-Reverse Sequencing approach was used and constructed a synthetic promoter library comprising 112 955 variants with minimal redundancy and a 16 226-fold expression range. A Transformer-based model trained on this dataset achieved a Pearson correlation of 0.87 with experimentally measured promoter strengths. When combined with a conditional diffusion model, the system enabled de novo generation of promoter sequences with defined strengths, achieving a design-to-measurement correlation of 0.95 and maintaining high accuracy (R = 0.93) across varied sequence contexts. The designed promoters consistently preserved their intended strength gradients, demonstrating robust plug-and-play functionality. This work establishes a scalable and extensible platform (www.yudenglab.com) for deep learning-guided programmable design of Escherichia coli core promoters, enabling precise transcriptional control.

Promoter Regions, Genetic

IMPACT OF FLUORESCENT DYES ON MUTATIONS IN NEXT GENERATION SEQUENCING LIBRARY GENERATION.

DNA labelling fluorescent dyes such as ethidium bromide have long been considered to be highly mutagenic during DNA replication. While recent studies have pushed back on this narrative, the intercalative nature of these dyes continues to raise the possibility that these dyes can induce mutations. The iconPCR instrument by n6tec uses fluorescent dyes to measure amplification in real time and to adjust cycling conditions. However, since this use of qPCR is preparative and not analytical, mutations introduced by fluorescent dyes would be propagated into the sequencing reaction. To address the impact of these dyes on downstream analyses, we have performed routine mutation calling as well as mutational signature analysis on samples amplified using the iconPCR in the presence of either SYBR or EvaGreen. Sequence analysis revealed very minimal impacts of dyes on the reactions, largely within the noise regimen with only subtle changes in mutation rates seen. Mutational signature analysis was unable to identify any key signatures assignable to the dyes in either substitutions or indel domains. The mutational impact of intercalating dyes during fluorescence-guided amplification is therefore minimal and can be disregarded in all but the most sensitive NGS applications.

Fluorescent Dyes

A directed evolution approach to select for novel Adeno-associated virus capsids on an HIV-1 producer T cell line.

A directed evolution approach was used to select for Adeno-associated virus (AAV) capsids that would exhibit more tropism toward an HIV-1 producer T cell line with the long-term goal of developing improved gene transfer vectors. A library of AAV variants was used to infect H9 T cells previously infected or uninfected by HIV-1 followed by AAV amplification with wild-type adenovirus. Six rounds of biological selection were performed, including negative selection and diversification after round three. The H9 T cells were successfully infected with all three wild-type viruses (AAV, adenovirus, and HIV-1). Four AAV cap mutants best representing the small number of variants emerging after six rounds of selection were chosen for further study. These mutant capsids were used to package an AAV vector and subsequently used to infect H9 cells that were previously infected or uninfected by HIV-1. A quantitative polymerase chain reaction assay was performed to measure cell-associated AAV genomes. Two of the four cap mutants showed a significant increase in the amount of cell-associated genomes as compared to wild-type AAV2. This study shows that directed evolution can be performed successfully to select for mutants with improved tropism for a T cell line in the presence of HIV-1.

Capsid

Construction of cDNA library of Dalbergia odorifera induced by low temperature stress and screening of low temperature tolerant genes.

To systematically analyze the gene function of Dalbergia odorifera, the seedlings of D. odorifera were treated with low-temperature stress for 6 h. Total RNA was extracted from a mixture of seedling roots, stems, and leaves, and a low-temperature-induced D. odorifera yeast cDNA expression library was constructed. The library volume was 1.032 × 108 CFU, and the PCR (Polymerase Chain Reaction) identification of the library bacterial fluid showed that the amplification was around 1000 bp, with a single randomly distributed band, indicating that the library had been recombinantly inserted into the pYES2 vector. The GO (Gene Ontology) analysis showed that the library genes were mainly involved in metabolic and stress signaling pathways. The KEGG (Kyoto Encyclopedia of Genes and Genomes) pathway enrichment analysis showed that the genes were primarily related to energy and metabolic pathways. Twenty-one genes were screened or obtained at -20°C for low-temperature tolerance. In addition, the organ expression profiles of the candidate genes were analyzed based on RNA-seq data, and the expression profiles of the candidate genes under low-temperature stress were also examined. The construction of the yeast library provides genetic resources for the analysis of the mechanism of low-temperature tolerance of D. odorifera, which is important for comprehending and utilizing the genetic resources of D. odorifera.

Gene Library