Search PubMedSearch

SEARCH · Search PubMed

Results for “Genomic Library”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Genes required for Mycobacterium tuberculosis to survive the transition from aerosol to pulmonary alveolar lining fluid and early infection in a model of transmission.

Mycobacterium tuberculosis (Mtb) must withstand physical and chemical stresses during airborne transmission, including during the desiccation of aerosols small enough to reach pulmonary alveoli in a new host. There, Mtb encounters an antimicrobial pulmonary alveolar lining fluid (ALF) before it is engulfed by macrophages. To study the genes involved in Mtb's ability to survive the transition from desiccated droplet to pulmonary alveolus in an in vitro model, we formulated a model alveolar lining fluid (MALF) that mimics the composition of ALF as inferred from human bronchoalveolar lavage fluid (BALF). We compared the transcriptome of log-phase Mtb in MALF to the transcriptome of Mtb in BALF as BALF from the lungs of healthy adults was reconstituted to compensate for the dilution of ALF by lavage (rcBALF). Mtb from log-phase culture in a standard laboratory medium survived quantitatively in MALF and rcBALF for at least 24 hours. In contrast, Mtb that had passed through earlier stages of transmission began to succumb after 3 hours in MALF, past the time when particles have been observed to be phagocytized by alveolar macrophages. Screening of a genome-wide CRISPRi library of Mtb identified 35 genes as uniquely required by Mtb to survive the transition from desiccated microdroplet into rehydration in MALF. Thirty-one of these genes are non-essential under conventional laboratory conditions and seven have unknown functions. Thirteen of the 35 genes were additionally required for Mtb to survive in macrophage-like cells cultured at the air-liquid interface with pulmonary epithelial cells. This study nominates additional members of the transmission survival genome of Mtb, illustrates that different genes may contribute to the survival of Mtb at different stages of transmission, and suggests that modeled transmission can shed light on the functions of Mtb genes whose contributions have been unknown.

Journal Article

Genome-scale overexpression screening identifies product tolerance and efflux transport as key determinants of high-level L-tryptophan production in Escherichia coli.

L-tryptophan is a high-value aromatic amino acid widely used in the food, feed, and pharmaceutical industries. However, large-scale microbial production is constrained by insufficient precursor supply and limited strain tolerance to high product concentrations. In this study, modular metabolic engineering was first employed to enhance the availability of key precursors, including shikimate, serine, and glutamine, yielding strain TRPJ-13 with a 34.6% increase in L-tryptophan titer. To enhance strain tolerance, an indigo-based high-throughput reporter system was constructed and coupled with genome-scale overexpression library screening, leading to the identification of soxS as a tolerance-conferring target. Mechanistic analysis demonstrated that soxS upregulated lpxC to enhance lipopolysaccharide biosynthesis, thereby reinforcing membrane integrity and improving L-tryptophan tolerance. Combinatorial engineering of soxS and lpxC generated strain TRPJ-23, which increased L-tryptophan tolerance by 74.8% and L-tryptophan titer by 10.3%. Furthermore, YicL was identified as a novel transmembrane protein involved in L-tryptophan transport that effectively promoted L-tryptophan efflux, further increasing the titer by 9.0%. After fermentation optimization, strain TRPJ-28 produced 74.3 g/L L-tryptophan in a 5-L bioreactor, with a yield of 0.26 g/g and a productivity of 1.24 g/L/h. In a 1000-L pilot-scale bioreactor, TRPJ-28 reached a titer, yield, and productivity of 70.4 g/L, 0.25 g/g, and 1.17 g/L/h, respectively. This study provides new engineering insights for developing industrially promising L-tryptophan-producing strains.

Genome-scale overexpression screening

CRISPR/Cas9 screening revealed BIRC6-AS1/BIRC6 mediates abiraterone resistance via NHEJ pathway-dependent A20 degradation in prostate cancer.

Abiraterone acetate is a standard-of-care therapy for prostate cancer (PCa). However, resistance frequently emerges, often characterized by the progression to AR-independent phenotypes. Employing a genome-wide CRISPR/Cas9 library screening strategy, we identified 523 long non-coding RNAs (lncRNAs) and 2,183 protein-coding genes as potential candidates associated with abiraterone resistance. Notably, a pair of sense-antisense genes, BIRC6-AS1/BIRC6, was identified as a significant contributor to abiraterone resistance, serving as a critical survival factor in AR-independent contexts. BIRC6-AS1 depletion led to a reduction in both the mRNA and protein levels of BIRC6. Moreover, depletion of either BIRC6-AS1 or BIRC6 enhanced the sensitivity of PCa cells to abiraterone in both in vitro and in vivo settings. Further investigation revealed that BIRC6-AS1 stabilized the mRNA of BIRC6 through interaction with ILF2. Suppression of either BIRC6-AS1 or BIRC6 attenuated non-homologous end joining (NHEJ) repair activity, resulting in the disassembly of 53BP1 foci at DNA damage sites and an increased accumulation of DNA damage, thereby exposing a vulnerability in AR-independent resistant cells. Mechanistically, BIRC6 interacted with A20 and facilitated the K48-linked ubiquitination and subsequent degradation of A20 at the K337 residue. Additionally, A20 knockdown effectively reversed the abiraterone sensitivity induced by BIRC6-AS1 depletion. Collectively, our study provides a landscape of lncRNAs and protein-coding genes associated with abiraterone resistance and suggests that targeting the BIRC6-AS1/BIRC6 axis represents a potential therapeutic strategy to eradicate AR-independent resistant tumors in prostate cancer.

Journal Article

Recurrent reversible mutations at gaf1 driving metastable TORC1 inhibitor resistance in fission yeast.

Metastable phenotypic inheritance is often attributed to epigenetic mechanisms, but reversible genetic alterations can produce similar instability. Here, we investigated the basis of unstable resistance to TORC1 inhibitor (rapamycin plus caffeine) in Schizosaccharomyces pombe. Six independent, metastable resistant mutants were isolated. Genetic mapping positioned the causal lesion to a single Mendelian locus, which sequencing identified as gaf1, encoding a GATA transcription factor and a key negative regulator of growth downstream of TORC1. In each mutant, distinct loss-of-function mutations (insertions, deletions, or point mutations) were found in gaf1 in the resistant state, and these mutations precisely reverted to the wild-type sequence upon loss of resistance. Restoring the wild-type gaf1 allele abolished resistance, indicating that reversible genetic disruption of gaf1 is both necessary and sufficient for the metastable phenotype. Furthermore, strong resistance in several strains from a genome-wide deletion library was due to secondary, inactivating mutations in gaf1, underscoring its role as a recurrent adaptive target under rapamycin plus caffeine treatment. Mechanistically, gaf1 inactivation established a distinct basal transcriptome and pronounced derepression of translation and metabolic programs upon drug treatment. While rapamycin plus caffeine triggered extensive chromatin remodeling and H3K9 methylation contributed partially to resistance, these epigenetic changes were most consistent with a downstream modifying layer. Our study shows that metastable drug resistance in fission yeast is predominantly associated with recurrent, reversible genetic inactivation of the central transcriptional regulator gaf1, demonstrating how rapidly reversible genetic switches can drive adaptive evolution.IMPORTANCEDistinguishing between genetic and epigenetic inheritance is fundamental to understanding how cells adapt to environmental stress. In the fission yeast Schizosaccharomyces pombe, rapid and reversible drug resistance is often assumed to be driven by epigenetic switches that change gene activity without altering DNA. However, our study reveals that this instability can be caused by physical mutations in a single gene, gaf1, which acts as a genetic toggle. These mutations appear under drug pressure and precisely revert to the original sequence when the drug is removed. We also demonstrate that these spontaneous mutations can contaminate standard laboratory yeast collections, leading to potential misinterpretation of experimental data. These findings broaden our understanding of unstable inheritance and show that DNA sequences can be far more dynamic than previously recognized during rapid evolution and the development of drug resistance.

TORC1 signaling

Parsing GTF and FASTA files using the eccLib Library.

SUMMARY: Leveraging the Python/C API, eccLib was developed as a high-performance library designed for parsing genomic files and analysing genomic contexts. To the best of the authors' knowledge, it is the fastest Python-based solution available. With eccLib, users can efficiently parse GTF/GFFv3 and FASTA files and utilize the provided methods for additional analysis. AVAILABILITY AND IMPLEMENTATION: This library is implemented in C and distributed under the GPL-3.0 licence. It is compatible with any system that has the Python interpreter (CPython) installed. The use of C enables numerous optimizations at both the implementation and algorithmic levels, which are either unachievable or impractical in Python.

Software

Improved Genomic Resources for the swordtail cricket, Laupala kohalensis Otte 1994.

Advances in genetic tools such as next and third generation sequencing, paired with a focus on representative clades, provide insight into how processes including adaptation, admixture, and genome structure shape the evolution and maintenance of species. However, our understanding of the genomics of speciation is dominated by systems where ecological adaptations are thought to cause initial barriers to gene exchange. In contrast to other model systems, the 38 species of the genus Laupala constitute a very rapid radiation, where evolution of reproductive barriers and speciation is thought to be driven by sexual selection. Here, with novel PacBio HiFi reads and RNA- and Iso-Seq data, we provide a highly contiguous, chromosome-level genome and markedly improved annotation of the endemic Hawaiian cricket, Laupala kohalensis Otte, 1994. Our new resources advance previous efforts, placing 99% of 47 scaffolds on 7 autosomes and 1 sex chromosome in the 1.67 Gb assembly, with a 98.8% BUSCO score (insecta_db10), N50 of ~268 Mb, and L50 of 3. Using a custom repeat library, we estimate the genome to have 46.09% repeat content, and the new annotation includes an increased estimate of 17,670 genes, which coincides with that known from other Orthopterans. Notably, we find a large nuclear DNA segment of mitochondrial origin on chromosome 7. This new resource provides a powerful tool to identify and compare genomic causes of phenotypic diversification in a system characterized by strong signatures of sexual differentiation, representing an underappreciated but potentially widespread cause of speciation.

Hawaii

A trainable language model with potential to modulate translation rates in non-model organisms by generating upstream untranslated region sequence libraries.

Tuning protein expression in non-model organisms is often constrained by the lack of validated genetic parts and predictive design tools. Translational tuning through the modulation of upstream untranslated regions (5'-UTRs) offers a potentially organism-agnostic route, but existing methods typically rely on mechanistic assumptions, prior knowledge that may not be available in non-model contexts, or the screening of sequence libraries. Here, we present a simple generative approach for creating synthetic 5'-UTR libraries based solely on the genomic sequence statistics of any desired organism. The method uses a sliding-window n-gram language model applied to native 5'-UTR sequences to produce novel sequences that preserve organism-specific base distributions and motifs without hard-coding specific motifs or mechanistic rules into inflexible statistical templates. We have applied this approach to the model bacterium Escherichia coli and the non-model probiotic Limosilactobacillus reuteri. Libraries of approximately 1,000 sequences were generated for each organism, from which about 100 unique sequences were experimentally tested for translation of a fluorescent reporter protein. In both organisms, the synthetic libraries yielded a broad range of translation levels from this relatively small number of tested variants. Sequences derived from an organism's own genomic statistics provided a more uniformly distributed range of translation rates in that organism than sequences derived from the other species. Correlations of individual sequence performance across the two species were weak, and thermodynamic predictions of ribosome binding strength showed very little predictive power, especially in the non-model L. reuteri. The results demonstrate that simple statistical language model approaches applied to genomic data can generate functional translational regulatory sequence libraries without detailed mechanistic knowledge or explicit reference to consensus motifs. The approach requires minimal computational resources, avoids reproducing native sequences, and can be readily applied to any organism with a sequenced genome. This strategy may lower technical barriers to expression tuning in non-model organisms.

5' Untranslated Regions

Uchimata: a toolkit for visualization of 3D genome structures on the web and in computational notebooks.

SUMMARY: Uchimata is a toolkit for visualization of 3D structures of genomes. It consists of two packages: a Javascript library facilitating the rendering of 3D models of genomes, and a Python widget for visualization in Jupyter Notebooks. Main features include an expressive way to specify visual encodings, and filtering of 3D genome structures based on genomic semantics and spatial aspects. Uchimata is designed to be highly integratable with biological tooling available in Python. AVAILABILITY AND IMPLEMENTATION: Uchimata is released under the MIT License. The Javascript library is available on NPM, while the widget is available as a Python package hosted on PyPI. The source code for both is available publicly on Github (https://github.com/hms-dbmi/uchimata and https://github.com/hms-dbmi/uchimata-py) and Zenodo (https://doi.org/10.5281/zenodo.17831959 and https://doi.org/10.5281/zenodo.17832045). The documentation with examples is hosted at https://hms-dbmi.github.io/uchimata/.

Software

Uchimata: a toolkit for visualization of 3D genome structures on the web and in computational notebooks.

SUMMARY: Uchimata is a toolkit for visualization of 3D structures of genomes. It consists of two packages: a Javascript library facilitating the rendering of 3D models of genomes, and a Python widget for visualization in Jupyter Notebooks. Main features include an expressive way to specify visual encodings, and filtering of 3D genome structures based on genomic semantics and spatial aspects. Uchimata is designed to be highly integratable with biological tooling available in Python. AVAILABILITY AND IMPLEMENTATION: Uchimata is released under the MIT License. The Javascript library is available on NPM, while the widget is available as a Python package hosted on PyPI. The source code for both is available publicly on Github (https://github.com/hms-dbmi/uchimata and https://github.com/hms-dbmi/uchimata-py). The documentation with examples is hosted at https://hms-dbmi.github.io/uchimata/. CONTACT: david_kouril@hms.harvard.edu or nils@hms.harvard.edu.

Journal Article

Exploring phage-host interactions in Burkholderia cepacia complex bacterium to reveal host factors and phage resistance genes using CRISPRi functional genomics and transcriptomics.

Complex interactions of bacteriophages with their bacterial hosts determine phage host range and infectivity. While phage defense systems and host factors have been identified in model bacteria, they remain challenging to predict in non-model bacteria. In this paper, we integrate functional genomics and transcriptomics to investigate phage-host interactions, revealing active phage resistance and host factor genes in Burkholderia cenocepacia K56-2. Burkholderia cepacia complex species are commonly found in soil and are opportunistic pathogens in immunocompromised patients. We studied infection of B. cenocepacia K56-2 with Bcep176, a temperate phage isolated from Burkholderia multivorans. A genome-wide dCas9 knockdown library targeting B. cenocepacia K56-2 was constructed, and a pooled infection experiment identified 63 novel genes or operons coding for candidate host factors or phage resistance genes. The activities of a subset of candidate host factor and resistance genes were validated via single-gene knockdowns. Transcriptomics of B. cenocepacia K56-2 during Bcep176 infection revealed that expression of genes coding for host factor and resistance candidates identified in this screen was significantly altered during infection by 4 h post-infection. Identifying which bacterial genes are involved in phage infection is important to understand the ecological niches of B. cenocepacia and its phages, and for designing phage therapies.IMPORTANCEBurkholderia cepacia complex bacteria are opportunistic pathogens inherently resistant to antibiotics, and phage therapy is a promising alternative treatment for chronically infected patients. Burkholderia bacteria are also ubiquitous in soil microbiomes. To develop improved phage therapies for pathogenic Burkholderia bacteria, or engineer phages for applications, such as microbiome editing, it's essential to know the bacterial host factors required by the phage to kill bacteria, as well as how the bacteria prevent phage infection. This work identified 65 genes involved in phage-host interactions in Burkholderia cenocepacia K56-2 and tracked their expression during infection. These findings establish a knowledge base to select and engineer phages infecting or transducing Burkholderia bacteria.

Bacteriophages

Zinc-dependent turnover of ZIP3 transporter mRNA by trypanosome ZNK1.

Like other cells, parasitic and other trypanosomatids sense Zn2+ and regulate Zn2+ transport, but the mechanisms involved remained unknown. Here, we identify a trypanosome RNA-binding protein that specifically eliminates ZIP3 transporter mRNA in Zn2+-replete conditions. We first demonstrate that Trypanosoma brucei ZIP3 mRNA abundance is subject to 3'-untranslated region (3'-UTR) and Zn2+-dependent negative control. A genome-wide RNA interference library screen, using a reporter associated with the ZIP3 3'-UTR, identifies Tb927.11.9510 as a candidate Zn2+-sensor, and we name this protein Zinc Nuclear Knuckles 1 (ZNK1) since it localizes to the nucleus and contains several Zn2+-knuckle motifs. ZNK1 is conserved among trypanosomatids, and a PIN domain suggests a ribonuclease-based mechanism. We use Cas9-editing to knockout ZNK1 and observe specific accumulation of ZIP3 transcripts, and increased intracellular Zn2+, in znk1 null cells. We validate ZNK1 as a ZIP3 3'-UTR-dependent negative regulator and identify a GU-repeat motif in the ZIP3 3'-UTR that is predictive of ZNK1-based negative control. In conclusion, ZNK1 eliminates ZIP3 transporter mRNA in a Zn2+-dependent manner. We suggest that trypanosomatid ZNK1 is a highly selective zinc finger nuclease that binds GU-repeat motifs within ZIP3 3'-UTRs and degrades Zn2+ transporter mRNA only when the tandem sensor knuckle modules are coordinated with Zn2+.

Trypanosoma brucei brucei

Genome-wide CRISPR screen reveals PEX11B as a host restriction factor against ORFV through membrane fluidity regulation.

Host-pathogen interactions are shaped by cellular restriction factors that direct antiviral defenses. We built the first ovine genome-wide CRISPR knockout library in sheep testis (OA3.Ts) cells, targeting all protein-coding genes. Using this platform, we identified PEX11B, a peroxisomal membrane regulatory protein, as a strong restriction factor against orf virus (ORFV) infection. Removing PEX11B increased viral susceptibility and triggered severe cytopathic effects with membrane fusion and syncytia formation. Mechanistic studies showed that PEX11B knockout harmed peroxisomal integrity and disrupted lipid metabolism. This led to greater plasma membrane fluidity, creating a proviral environment that allowed more viral entry and replication. These results reveal a new antiviral function for PEX11B in blocking viral infection and underscore the importance of peroxisomal regulation in host-virus interactions.

Animals

PRMT3 restricts porcine epidemic diarrhea virus replication by disrupting the interaction between VAPA and the viral nucleocapsid protein.

Porcine epidemic diarrhea virus (PEDV) represents a severe threat to the global swine industry. Its infection process involves intricate virus-host interactions and immune evasion mechanisms, but effective therapeutic targets remain elusive. In this study, we identified protein arginine methyltransferase 3 (PRMT3) as a novel regulatory factor that significantly modulates PEDV infection via genome-wide CRISPR/Cas9 knockout library screening. Knockout or inhibition of PRMT3 markedly enhanced PEDV infection in multiple cell lines, including LLC-PK1, IPEC-J2, and primary porcine intestinal epithelial cells. Mechanistic investigations revealed that PRMT3 can restrict PEDV infection by interacting with vesicle-associated membrane protein-associated protein A (VAPA). Further analysis revealed that VAPA facilitates cholesterol transport through binding to oxysterol-binding protein (OSBP) and inhibits the autophagic degradation of the viral nucleocapsid (N) protein, with both processes being critical for promoting PEDV infection in host cells. A detailed analysis revealed that K52 within its major sperm protein (MSP) domain interacts with D404 and D405 in the two phenylalanines in an acidic tract (FFAT)-like motifs of the N protein, and these interactions proved essential for PEDV infection. In summary, this is the first study to identify and validate the PRMT3-VAPA-N protein autophagic degradation axis as a key pathway through which PRMT3 suppresses PEDV infection, with VAPA acting as an essential host factor for PEDV pathogenesis. These findings uncover novel signaling pathways and molecular targets for the development of anti-PEDV therapeutics.

Animals

Epigenetic Repression of TP53 Transcription Underlies Cancer Cell Persistence for Carboplatin Resistance in Non-Small Cell Lung Cancer.

While chemoresistance in non-small cell lung cancer (NSCLC) cells has historically been attributed to permanent genetic mutations, emerging evidence highlights the role of nongenetic transcriptional plasticity and 'drug-tolerant persister' cells. To systematically map these epigenetic vulnerabilities, we utilized a genome-wide CRISPR interference library to screen wild-type TP53 NSCLC (A549) cells under carboplatin selection. Using the DrugZ algorithm and subsequent pathway enrichment analyses, this screen revealed that transcriptional suppression of interstrand crosslink DNA repair networks, including the Fanconi anemia pathway, markedly sensitized cells to carboplatin. Unexpectedly, transcriptional silencing of TP53 and its downstream target CDKN1A emerged as the strongest drivers of resistance, enabling cells to bypass therapy-induced senescence and maintain their proliferative potential later. To validate these findings in a clinically relevant context, we established a chronic carboplatin-resistant cell model (A549CarboR cells). A549CarboR exhibited a reduction in TP53 transcripts, along with decreased H3K27 acetylation and increased DNA hypermethylation on its promoter. Epigenetic remodeling using the DNA methyltransferase inhibitor (DNMTi) was associated with unblocking TP53 transcription, restored p53 signaling, and resensitization of resistant cells to carboplatin. Conversely, histone deacetylase inhibitors induced CDKN1A transcription to bypass TP53, indicating distinct epigenetic circuits. Collectively, the results demonstrate for the first time that TP53 expression is dynamically regulated at the transcriptional level through promoter methylation related to the drug tolerance. These insights emphasize that epigenetic silencing, rather than exclusive genetic loss-of-function, contribute to platinum resistance and underscore the therapeutic potential of pairing platinum regimens with DNMTi to target the transcriptomic plasticity of persistent cancer cell populations.

CRISPR interference screening

Comparative metagenomic assessment of Illumina-compatible library preparation methods, short-read lengths, and PacBio HiFi sequencing reveals differences in microbial and functional diversity recovery from a complex environmental sample.

UNLABELLED: Metagenomics enables comprehensive exploration of microbial communities but is influenced by library preparation and sequencing technologies, affecting recovery of microbial genomes and proteins. Here, we benchmarked six Illumina-compatible short-read library preparation conditions in triplicate at 2 × 150 bp and 2 × 250 bp read lengths alongside PacBio HiFi long-read sequencing using a composite environmental sample of marine mangrove sediment and terrestrial palm tree soil. Longer short reads (2 × 250 bp) combined with optimal library preparation approaches improved assembly quality, protein detection, and metagenome-assembled genome (MAG) recovery, achieving results approaching those of long-read sequencing. TruSeq libraries at 2 × 250 bp recovered more than sevenfold more unique proteins than the same kit at 2 × 150 bp (811,701 vs 110,108) using the same number of sequencing reads, while recovering a comparable number of high-quality MAGs to PacBio HiFi long-read sequencing (11 vs 18) and surpassing it in protein discovery by almost 10-fold (811,701 vs 87,745) at less than half of the sequencing cost. Furthermore, biosynthetic gene cluster analysis identified 46 biosynthetic gene clusters in TruSeq-250PE assemblies compared to 38 in PacBio HiFi, with several showing no close match in the MIBiG database. Although long reads yield more contiguity and complete genomes, longer short reads offer a cost-effective, scalable alternative for uncovering microbial and functional diversity. These findings provide critical guidance for metagenomic experimental design, demonstrating that strategic selection of library preparation chemistry and sequencing parameters can reveal more unknown microbial information in complex biomes without requiring additional sequencing depth. IMPORTANCE: Metagenomic outcomes are strongly influenced by library preparation and sequencing strategies, yet their combined effects in complex environmental samples remain poorly defined. Here, we provide the first direct comparison of Illumina NovaSeq short-read metagenomic sequencing at 2 × 150 bp and 2 × 250 bp across multiple library preparation kits, alongside PacBio HiFi long-read sequencing. We show that sequencing read length and library preparation critically shape assembly quality, protein recovery, and metagenome-assembled genome (MAG) reconstruction. These findings demonstrate that short-read sequencing at 2 × 250 bp, with appropriate library preparation, can match long-read technologies in MAG recovery while substantially surpassing them in protein discovery. With less than half of the sequencing price and a 3.5-fold reduction in cost per gigabase of usable data, this method facilitates more accessible large-scale metagenomic analysis within complex environmental systems.

Metagenomics

Isolation of a genomal clone containing chicken histone genes.

We have used enriched chicken histone cDNA to select genomal clones from a chicken library. Because the cDNA probe also contained other sequences, a further screening of positive plagues with negative probes eliminated most non-histone gene clones. One 'positively-selected' genomal clone, lambda CH-01, hybridised with cloned sea-urchin histone genes and also detected histone genes in EcoRI-digested genomal sea-urchin DNA. Limited DNA sequencing of HaeIII fragments identified two sequences within the coding region of chicken histone H2A. A third fragment predicted an amino acid sequence with strong homology to an H1 histone sequence.

Amino Acid Sequence

The organization of a nuclear DNA sequence from a higher plant: molecular cloning and characterization of soybean ribosomal DNA.

The recombinant DNA vector, lambda Charon 4A, was used to construct a library of DNA sequences from the genomic DNA of soybean (Glycine max). To define the organization of ribosomal DNA (rDNA) in the soybean genome, clones containing sequences complementary to both 17S and 25S rRNA have been isolated from this library and used in conjunction with Southern blot hybridization. The rRNA genes are tandemly reiterated with a relatively small unit repeat length of 7.8 kb. There is no heterogeneity in the length of the rDNA repeat units although they display limited differences in either base sequence or pattern of methylation. The cloned rDNA sequences are shown to comprise the entire repeat unit and have been used to obtain a detailed restriction map as well as an approximate transcription map of soybean rRNA genes. The cloning of rDNA from soybean suggests that recombinant DNA techniques can be successfully applied to the genomic DNA of higher plants despite the high degree of methylation exhibited by plant DNA.

Bacteriophage lambda

polars-bio-fast, scalable, and out-of-core operations on large genomic interval datasets.

MOTIVATION: Genomic studies very often rely on computationally intensive analyses of relationships between features, which are typically represented as intervals along a 1D coordinate system (such as positions on a chromosome). In this context, the Python programming language is extensively used for manipulating and analyzing data stored in a tabular form of rows and columns, called a DataFrame. Pandas is the most widely used Python DataFrame package and has been criticized for inefficiencies and scalability issues, which its modern alternative-Polars-aims to address with a native backend written in the Rust programming language. RESULTS: polars-bio is a Python library that enables fast, parallel and out-of-core operations on large genomic interval datasets. Its main components are implemented in Rust, using the Apache DataFusion query engine and Apache Arrow for efficient data representation. It is compatible with Polars and Pandas DataFrame formats. In a real-world comparison (107 versus 1.2×106 intervals), our library runs overlap queries 6.5×, nearest queries 15.5×, count_overlaps queries 38×, and coverage queries 15× faster than Bioframe. On equally sized synthetic sets (107 versus 107), the corresponding speedups are 1.6×, 5.5×, 6×, and 6×. In streaming mode, on real and synthetic interval pairs, our implementation uses 90× and 15× less memory for overlap, 4.5× and 6.5× less for nearest, 60× and 12× less for count_overlaps, and 34× and 7× less for coverage than Bioframe. Multi-threaded benchmarks show good scalability characteristics. To the best of our knowledge, polars-bio is the most efficient single-node library for genomic interval DataFrames in Python. AVAILABILITY AND IMPLEMENTATION: polars-bio is an open-source Python package distributed under the Apache License available for major platforms, including Linux, macOS, and Windows in the PyPI registry. The online documentation is https://biodatageeks.org/polars-bio/ and the source code is available on GitHub: https://github.com/biodatageeks/polars-bio and Zenodo: https://doi.org/10.5281/zenodo.16374290. are available at Bioinformatics online.

Software