Search PubMedSearch

SEARCH · Search PubMed

Results for “Genomic Library”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

RabbitSketch: a high-performance sketching library for genome analysis.

SUMMARY: We present RabbitSketch, a highly optimized library of sketching algorithms such as MinHash, OrderMinHash, and HyperLogLog that can exploit the power of modern multi-core CPUs. It provides significant speedups compared to existing implementations, ranging from 2.30× to 49.55×, as well as flexible and easy-to-use interfaces for both Python and C++. As a result, the similarity analysis of 455GB genomic data can be completed in only 5 minutes using RabbitSketch with merely 20 lines of Python code. As a case study, we enhanced RabbitTClust by integrating RabbitSketch's Kssd algorithm, resulting in a 1.54× speedup with no loss in accuracy. AVAILABILITY AND IMPLEMENTATION: RabbitSketch is available at https://github.com/RabbitBio/RabbitSketch with an archived version at Zenodo: https://doi.org/10.5281/zenodo.14903962. Detailed API documentation is available at https://rabbitsketch.readthedocs.io/en/latest.

Software

DeepGeSeq: deep learning library for genomic sequence modeling and analysis.

MOTIVATION: Deep learning methods have demonstrated significant potential in genomics, enabling broad applications such as sequence activity prediction, regulatory rule identification, and variant effect quantification. However, their widespread adoption is often hindered by the steep computational learning curve required for model construction, training, and downstream biological interpretation. Here, we introduce DeepGeSeq, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis. RESULTS: By integrating state-of-the-art architectural modules, DeepGeSeq streamlines the entire deep learning workflow, requiring minimal user input via a simple configuration file and an intuitive agentic skill. We comprehensively validate the efficacy of DeepGeSeq through diverse case studies, encompassing pipeline verification using synthetic datasets, the reproduction and application of established models, and model fine-tuning coupled with biological interpretation on user-defined data. Furthermore, we demonstrate DeepGeSeq's versatility in domain-specific applications, including single-cell ATAC-seq modeling for cell-type clustering, and MPRA data modeling coupled with in silico saturation mutagenesis to dissect cis-regulatory elements. Ultimately, DeepGeSeq bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development and broad application of deep learning methods in genomics research. AVAILABILITY AND IMPLEMENTATION: https://github.com/JiaqiLi1024/DeepGeSeq.

Deep Learning

Expansion of the functional genomics GRACE library reveals genes relevant for temperature-dependent fitness in Candida albicans.

A small percentage of species in the fungal kingdom can cause devastating infections in humans, with Candida albicans reigning as a leading cause of systemic disease. One of the key virulence phenotypes for pathogenic fungi is the ability to survive at host body temperature; however, a comprehensive understanding of the mechanisms that orchestrate thermal adaptation in fungi remains incomplete. In this study, we expand the largest functional genomics resource in C. albicans, reaching 71.3% coverage of the entire genome, and perform screens under six different temperatures to identify genes important for temperature-dependent fitness. We describe the function of genes involved in translation (GAR1), splicing (C1_11680C or YSF3), and cell cycle progression (C6_00110C or RHT1) in enabling fungal survival at both low and high temperatures. Through experimental evolution, we also show that C. albicans can rapidly overcome deleterious mutations and adapt to extreme temperature environments. Overall, our study highlights the transformative potential of genome-wide functional genomics to uncover critical vulnerabilities in pathogenic fungi.

Genomics

The cost and cost trajectory of genome sequencing and bioinformatics analysis for Indigenous children with suspected rare diseases.

PURPOSE: Indigenous peoples are underrepresented in reference genome libraries. Consequently, rare disease diagnosis may require bespoke bioinformatics analyses of genome sequences. Establishing diagnostic cost is crucial to support policy development for equitable diagnosis of rare diseases. We estimated the cost and cost trajectory of diagnostic genome sequencing and bioinformatics for Indigenous participants with suspected rare diseases. METHODS: We conducted a microcosting study of Indigenous children and their families receiving genome sequencing through Canada's Silent Genomes Project. Invoice data informed the costs of genome sequencing. We conducted a time-and-motion study for bioinformatics analyses, including labor, computing, and data storage costs. RESULTS: With standard bioinformatics, costs ranged from C$3645 (SD: 455) for singletons to C$7402 (SD: 566) for trios. With advanced, bespoke bioinformatics, costs ranged from C$5344 (SD: 634) for singletons to C$9760 (SD: 822) for trios. Genome sequencing was a primary cost driver; however, sequencing costs decreased by 61% over 4 years. Bioinformatics costs ranged from 21.3% to 58.3% of the total costs. The time required for bioinformatics ranged from 71 hours to 215 hours for standard and advanced analyses, respectively. CONCLUSION: Genome sequencing costs decreased over time. Bioinformatics is a significant cost driver, particularly for bespoke analyses arising from nonrepresentative reference libraries.

Humans

High-throughput recovery of integron cassettes for gene discovery screens.

Integrons capture functional genes in mobile genetic elements called integron cassettes, which represent an untapped source of genes of biotechnological interest. Here we present two tools, cassette gatherer and cassette hunter, that enable high-throughput establishment of gene libraries either from genetically tractable strains or directly from DNA. We re-engineered a class 1 integron into counterselection markers on a plasmid or on the chromosome of a naturally competent Vibrio cholerae, which enabled capture of single cassettes in a sequence- and function-independent manner. When applied to Vibrio strains and genomic libraries, our tools recovered hundreds of single cassettes per assay with more than 99% specificity. We further subjected the library of cassettes generated by the hunter and gatherer tools to screens against phages ICP2 and T4, and identified nine phage-defence systems, including five previously undescribed. These tools enable rapid and large-scale recovery of integron cassettes that could be leveraged for functional gene discovery.

Journal Article

Controlling AAV Tropism in the Nervous System with Natural and Engineered Capsids.

More than one hundred naturally occurring variants of adeno-associated virus (AAV) have been identified, and this library has been further expanded by an array of techniques for modification of the viral capsid. AAV capsid variants possess unique antigenic profiles and demonstrate distinct cellular tropisms driven by differences in receptor binding. AAV capsids can be chemically modified to alter tropism, can be produced as hybrid vectors that combine the properties of multiple serotypes, and can carry peptide insertions that introduce novel receptor-binding activity. Furthermore, directed evolution of shuffled genome libraries can identify engineered variants with unique properties, and rational modification of the viral capsid can alter tropism, reduce blockage by neutralizing antibodies, or enhance transduction efficiency. This large number of AAV variants and engineered capsids provides a varied toolkit for gene delivery to the CNS and retina, with specialized vectors available for many applications, but selecting a capsid variant from the array of available vectors can be difficult. This chapter describes the unique properties of a range of AAV variants and engineered capsids, and provides a guide for selecting the appropriate vector for specific applications in the CNS and retina.

Animals

Genomic Tracking of Market-Derived Bull Shark Fins Back to Source Population of Origin.

International trade of shark fins remains difficult to monitor because products are rarely labelled to species and are often highly processed, resulting in severely degraded DNA. For several shark species listed under Appendix II of the Convention on International Trade in Endangered Species of Wild Fauna and Flora (CITES), this limits external verification of source populations supplying global trade hubs. Here, we assess whether nuclear genomic approaches can be applied to market-derived bull shark (Carcharhinus leucas) fins to determine their population of origin. We analysed dried fin trimmings collected from retail vendors in Hong Kong SAR, one of the world's largest dried shark fin trade hubs, using a targeted DArTcap single nucleotide polymorphism (SNP) panel, originally developed for population genomic studies of this species. Despite substantial DNA degradation, genomic libraries were successfully obtained for most samples, yielding sufficient SNP data to perform robust provenance and sex assignment. Using a Bayesian mixed-stock analysis, most fin samples were assigned to the Indo-West Pacific (71.4%), with smaller contributions from the western Atlantic (22.6%) and eastern Pacific (3.0%). Genetic sex assignment revealed twice as many males as females, although results indicated a conservative bias towards male assignment due to the limited number of X-linked markers available in degraded samples. Our results demonstrate that genome-wide targeted approaches can be effectively applied to highly processed shark fin products to infer population sources and sex composition. This study provides proof-of-concept for integrating genomics into shark trade monitoring, highlighting its potential to improve traceability, support CITES implementation and inform conservation and fisheries management, particularly for species with well-resolved population structure.

Animals

Gene overexpression reduces inhibitory metabolites to enhance CHO cell growth and IgG1 production.

Controlling the generation of toxic by-products in mammalian bioprocess to maximize therapeutic protein production and glycosylation patterns is a challenge. Intracellular metabolism is often not well-regulated and known to secrete toxic intermediate by-products which hampers cellular performance and negatively impacts critical quality attributes (CQA) of cells. Previous studies have identified trigonelline (TRI), n-acetyl putrescine (NAP), aconitic acid (AA), and cytidine monophosphate (CMP) generated through CHO cell metabolism and verified their negative impacts on growth and antibody production. In this approach, a genetic engineering strategy was developed to control downstream accumulation of inhibitory metabolites. The study successfully identified four different metabolic genes in CHO cells, including Cat (nicotinate and nicotinamide metabolism) to control the generation of TRI, Got1 and Hoga1 (proline metabolism) to control the generation of NAP, Got1 (TCA cycle) to control the generation of AA, and Slc35a1 (n-glycan biosynthesis) to control the generation of CMP. Each target gene-of-interest (GOI) was cloned from CHO genomic library, inserted into linearized vector plasmid, and subsequently transfected into cells. CQA of the bioprocess realized 22-30% increase in peak cell density, 16-22% increase overall IVCD, with an improving growth rate during cellular expansion phase when comparing engineered cells against control cells. The study also conducted a follow-up quadruple transfection study where all four GOIs were co-transfected into cells at ¼ of the total DNA concentration per GOI. An increase in cellular performance was also realized, as increases in peak VCD (17% increase), cumulative IVCD (17% increase), and growth rate were achieved. Both studies also found higher IgG1 antibody synthesis when cell metabolism was better regulated, as the studies measured 4% to 40% titer increase across all engineered cells when compared against control cells. The study also measured higher levels of G1F and G2F glycans with decreased level of G0F across all transfected cells, further indicating improvement in bioprocess, as cells were able to produce a higher fraction of semi-complex and complex versus simple glycoforms. Further investigation revealed that Cat and Slc35a1 exhibited comparable expression levels in the MG condition to their single-gene conditions (within 1% and 10% difference, respectively), corresponding to modest titer improvements closest to the control. These findings suggest that when all four genes are co-expressed, Cat and Got1 may act as rate-limiting factors influencing both cellular phenotypes and titer production. In both studies, the concentrations of downstream metabolic inhibitors were measured to be significantly decreased when comparing engineered cells against control cells, further demonstrating that overexpression of genes to re-allocate metabolic fluxes away from synthesizing toxic by-products can significantly improve cellular growth and protein synthesis.

Animals

Identification and characterization of Prp45p and Prp46p, essential pre-mRNA splicing factors.

Through exhaustive two-hybrid screens using a budding yeast genomic library, and starting with the splicing factor and DEAH-box RNA helicase Prp22p as bait, we identified yeast Prp45p and Prp46p. We show that as well as interacting in two-hybrid screens, Prp45p and Prp46p interact with each other in vitro. We demonstrate that Prp45p and Prp46p are spliceosome associated throughout the splicing process and both are essential for pre-mRNA splicing. Under nonsplicing conditions they also associate in coprecipitation assays with low levels of the U2, U5, and U6 snRNAs that may indicate their presence in endogenous activated spliceosomes or in a postsplicing snRNP complex.

Base Sequence

Cloned pairs of variable region genes for immunoglobulin heavy chains isolated from a clone library of the entire mouse genome.

To investigate the organization of immunoglobulin genes, we have constructed a clone library containing 10(6) randomly generated fragments of mouse embryo DNA, corresponding to eight equivalents of the genome. The cloning method involved methylation of embryo DNA at EcoRI recognition sites, partial digestion by EcoRI* endonclease activity, and direct ligation of the resulting large fragments to the lambda phage vector Charon 4A. The library was searched for sequences homologous to a cloned complementary DNA copy of a mu heavy chain mRNA. Nine clones bearing variable heavy chain (VH) sequences were isolated, representing at least eight distinct VH genes. Thus, multiple related VH genes are available in the genome to contribute to immunoglobulin diversity. Each of the two clones carries a pair of VH genes, one pair separated by 15 +/- 1 kilobase pairs of mouse DNA and the other by 14 +/- 2 kilobase pairs. This indicates that related VH genes are clustered and may occur in a tandem array having a repeating unit of 14--16 kilobase pairs. The large spacer sequences between VH genes cannot, however, be highly conserved.

Animals

Unraveling cadaverine toxicity effect to guide the engineering of robust strain.

End-product inhibition represents a major challenge in the microbial synthesis of various value-added chemicals. Cadaverine, a key monomer for polyamide synthesis, exhibits severe cytotoxicity, limiting its high-titer biosynthesis. Here, transcriptomic analysis and genome-wide library screening were integrated to systematically elucidate the cytotoxic mechanisms of cadaverine in Escherichia coli (E. coli) and identify beneficial genes for enhanced tolerance and overproduction. Transcriptomic analysis revealed that high concentrations of cadaverine disrupted cell membrane integrity and impaired oxidative phosphorylation, leading to redox imbalance and reactive oxygen species (ROS) accumulation. Subsequent genome-wide screening further confirmed these toxicity mechanisms and uncovered crucial cellular defense strategies. Functional validation highlighted the important role of NikR, UbiE, and YcbX in enhancing membrane integrity, restoring respiratory function and ROS homeostasis, or scavenging 6-N-hydroxylaminopurine (6-HAP) to prevent DNA damage. Among these, YcbX emerged as the most effective target for improving production. Consequently, we constructed a robust E. coli strain by implementing a dynamic regulation system for YcbX expression under cadaverine-responsive promoters, which significantly enhanced cadaverine biosynthesis to 87.2 g/L (a 46.8% enhancement). This work provides an in-depth understanding of cadaverine toxicity and tolerance, offering valuable targets and strategies for the rational design of high-performance microbial cell factories for diamines.

6-HAP clearance

Barcoded mutant library enables high-throughput functional genomics in a filamentous fungus.

Advances in sequencing technology enabling rapid and inexpensive whole-genome sequencing highlight how few genes are functionally characterized. This problem is particularly acute in filamentous fungi, where even in the best studied organisms upward of half of genes are poorly characterized or unannotated. High-throughput tools to identify gene function exist for single-celled organisms, like yeast and bacteria. However, filamentous fungi present challenges to high-throughput gene characterization, including low transformation efficiency and multinucleate cells. Filamentous fungi are critical components of nutrient cycling in ecosystems, form symbioses with plants that improve nutrient uptake, and are devastating human, plant, and animal pathogens causing millions of deaths and substantial crop loss each year. Thus, it is critical to overcome challenges to rapid gene characterization in filamentous fungi. We generated a library of hundreds of millions of uniquely barcoded plasmids containing a broad host-range drug resistance marker for ectopic insertion into filamentous fungal genomes by Agrobacterium tumefaciens. We then optimized A. tumefaciens mediated transformation of the biocontrol agent Trichoderma atroviride and made an insertional mutagenesis library containing 83,311 barcoded insertions, disrupting 5,331 of 11,863 predicted genes. This library enables high-throughput screens to rapidly connect genotype to phenotype. Quantifying relative barcode abundance in the pooled library before and after exposure to experimental conditions identified candidate genes and recovered known pathway components in amino acid biosynthetic, fructose utilization, and xylose utilization pathways. This resource establishes a scalable platform for high-throughput functional genomics in filamentous fungi, enabling investigations of fungal biology to improve medical outcomes, biotechnology, and sustainable agriculture.

Genomics

Extraction, Purification, and Next-Generation Sequencing (NGS) Analysis of DNA and RNA from Formalin-Fixed and Paraffin-Embedded (FFPE) Tissue.

Formalin fixed paraffin embedded (FFPE) tissues have long been used for immunohistological analyses. FFPE tissues can be stored at room temperature for several years enabling analyses to be performed later. Ease of storage and transport makes these tissues an attractive source of biological material. However, formalin fixation results in chemical modifications of proteins and nucleic acids that poses a major challenge to any type of analysis. Recovery of nucleic acids for quantitative assays is rendered difficult due to degradation resulting from fixation and long-term storage, producing low usable yields. Extensive efforts in the last 20 years have led to significant improvements in use of FFPE tissues for DNA and RNA analyses and resulted in development of sensitive assays for a wide range of applications, including next-generation sequencing. In this chapter, we describe the optimization of methods for sequential extraction of DNA and RNA from FFPE tissue and subsequent preparation of DNA-seq and RNA-seq libraries for use with the Illumina platform using commercially available reagents/kits.

Paraffin Embedding

Long-read, high-coverage reference genome of the nymphalid butterfly Catonephele acontius (Nymphalidae: Biblidinae).

Catonephele acontius (Nymphalidae:Biblidinae:Epicalinii) is a butterfly species with a wide distribution across the Neotropics including the Amazon. Here, we present a long-read high-coverage reference genome for this species to serve as a genomic resource for future studies on Biblidinae butterflies, a group that is the subject of ongoing studies of seasonal adaptation under climate change. We used PacBio HiFi and IsoSeq reads to generate a highly contiguous and well-annotated reference genome. Five libraries were constructed, 4 using RNA from different tissues and 1 using high molecular weight (HMW) DNA from a wild-caught female. The DNA was sequenced using PacBio HiFi technology, and the RNA was sequenced using long read PacBio IsoSeq technology. About 20 Gb of raw HiFi data were generated and assembled to an initial size of 520.7 Mb (39 × homozygous coverage) in 90 contigs. The assembly was then polished and decontaminated into 40 contigs with an N50 of 19.927 Mb (BUSCO completeness: 99.0%; duplication: 0.5%; fragmentation: 0.7%; and missing: 0.3%). Final assembly size was 519.2 Mb. Repeats were annotated, showing that the genome consisted of 40.4% transposable elements. IsoSeq transcriptome data from antennae, leg, ovary, and digestive tissue was then used to structurally and functionally annotate gene models for the softmasked genome, uncovering ∼18,500 genes, with 70% of them given functional annotation. This reference assembly joins many published genomes in the Nymphalidae family but represents one of the first high-quality genomes from the Biblidinae subfamily. It provides a valuable resource to study the evolution of plastic and seasonal traits and will help investigate the genetic processes that may influence these species' responses to rapid climate change.

Animals

Proteomic Characterization of 1000 Human and Murine Neutrophils Freshly Isolated From Blood and Sites of Sterile Inflammation.

Neutrophils are indispensable for defense against pathogens. Injured tissue-infiltrated neutrophils can establish a niche of chronic inflammation and promote degeneration. Studies investigated transcriptome of single-infiltrated neutrophils which could misinterpret molecular states of these post mitotic cells. However, neutrophil proteome characterization has been challenging due to low harvests from affected tissues. Here, we present a workflow to obtain proteome of 1000 murine and human tissue-infiltrated neutrophils. We generated spectral libraries containing ∼6200 mouse and ∼5300 human proteins from circulating neutrophils. 4800 mouse and 3400 human proteins were recovered from 1000 cells with 102-108 copies/cell. Neutrophils from stroke-affected mouse brains adapted to the glucose-deprived environment with increased mitochondrial activity and ROS-production, while cells invading inflamed human oral cavities increased phagocytosis and granule release. We provide an extensive protein repository for resting human and mouse neutrophils, identify proteins lost in low input samples, thus enabling the proteomic characterization of limited tissue-infiltrated neutrophils.

Proteomics

Construction of a 2-Mb resolution BAC microarray for CGH analysis of canine tumors.

Recognition of the domestic dog as a model for the comparative study of human genetic traits has led to major advances in canine genomics. The pathophysiological similarities shared between many human and dog diseases extend to a range of cancers. Human tumors frequently display recurrent chromosome aberrations, many of which are hallmarks of particular tumor subtypes. Using a range of molecular cytogenetic techniques we have generated evidence indicating that this is also true of canine tumors. Detailed knowledge of these genomic abnormalities has the potential to aid diagnosis, prognosis, and the selection of appropriate therapy in both species. We recently improved the efficiency and resolution of canine cancer cytogenetics studies by developing a small-scale genomic microarray comprising a panel of canine BAC clones representing subgenomic regions of particular interest. We have now extended these studies to generate a comprehensive canine comparative genomic hybridization (CGH) array that comprises 1158 canine BAC clones ordered throughout the genome with an average interval of 2 Mb. Most of the clones (84.3%) have been assigned to a precise cytogenetic location by fluorescence in situ hybridization (FISH), and 98.5% are also directly anchored within the current canine genome assembly, permitting direct translation from cytogenetic aberration to DNA sequence. We are now using this resource routinely for high-throughput array CGH and single-locus probe analysis of a range of canine cancers. Here we provide examples of the varied applications of this resource to tumor cytogenetics, in combination with other molecular cytogenetic techniques.

Animals

Large-scale screening of genes responsible for silique length and seed size in Brassica Napus via pooled CRISPR library.

BACKGROUND: Enhancing rapeseed (Brassica napus, B. napus) yield is critical for ensuring global vegetable oil security. However, yield is heavily influenced by silique development and seed size, the enhancement of which is limited by scarce genetic resources. The CRISPR/Cas9 system has emerged as a powerful tool for constructing genome-wide mutant libraries, even in polyploid crops with complex genomes. RESULTS: The transcriptome-wide association study (TWAS) data, tissue-specific expression profiles data and reported genes were integrated to identify candidate genes regulating silique development and seed size. We constructed a sgRNA library targeting these genes and generated a CRISPR/Cas9 editing mutant library through genetic transformation. Specifically, 6124 sgRNAs were designed for 1739 candidate genes with ≦ 4 orthologues. 681 T0 plants were obtained through genetic transformation, which harbor 453 sgRNAs. Of 408 T0 plants analyzed, 151 (37.00%) exhibited successful gene editing events, targeting 84 candidate genes. Ten homozygous mutant plants were isolated and preliminary phenotypic analysis was performed in mutants targeting the BnaHRDs. The results suggest that mutations in BnaHRD.A03 and BnaHRD.C03 may modulate plant height (PH), main inflorescence length (MIL), silique length (SL), effective silique number per plant (ENS), seed number per silique (SNPS), and thousand-seed weight (TSW). CONCLUSIONS: This study harnessed the CRISPR/Cas9 technology to establish a preliminary library of gene-edited mutants in B. napus, thereby laying a robust foundation for the future screening of candidate genes pertaining to silique development and seed size. Furthermore, this study provides a methodological framework for rapid functional gene discovery in B. napus through CRISPR-based approaches.

Brassica napus

Global siRNA screen identifies human host factors critical for SARS-CoV-2 replication and late stages of infection.

Defining the subset of cellular factors governing SARS-CoV-2 replication can provide critical insights into viral pathogenesis and identify targets for host-directed antiviral therapies. While a number of genetic screens have previously reported SARS-CoV-2 host dependency factors, most of these approaches relied on utilizing pooled genome-scale CRISPR libraries, which are biased toward the discovery of host proteins impacting early stages of viral replication. To identify host factors involved throughout the SARS-CoV-2 infectious cycle, we conducted an arrayed genome-scale siRNA screen. Resulting data were integrated with published functional screens and proteomics data to reveal (i) common pathways that were identified in all OMICs datasets-including regulation of Wnt signaling and gap junctions, (ii) pathways uniquely identified in this screen-including NADH oxidation, or (iii) pathways supported by this screen and proteomics data but not published functional screens-including arachionate production and MAPK signaling. The identified proviral host factors were mapped into the SARS-CoV-2 infectious cycle, including 32 proteins that were determined to impact viral replication and 27 impacting late stages of infection, respectively. Additionally, a subset of proteins was tested across other coronaviruses revealing a subset of proviral factors that were conserved across pandemic SARS-CoV-2, epidemic SARS-CoV-1 and MERS-CoV, and the seasonal coronavirus OC43-CoV. Further studies illuminated a role for the heparan sulfate proteoglycan perlecan in SARS-CoV-2 viral entry and found that inhibition of the non-canonical NF-kB pathway through targeting of BIRC2 restricts SARS-CoV-2 replication both in vitro and in vivo. These studies provide critical insight into the landscape of virus-host interactions driving SARS-CoV-2 replication as well as valuable targets for host-directed antivirals.

Humans