Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Deep sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Deletion of Indian hedgehog gene causes dominant semi-lethal Creeper trait in chicken.

The Creeper trait, a classical monogenic phenotype of chicken, is controlled by a dominant semi-lethal gene. This trait has been widely cited in the genetics and molecular biology textbooks for illustrating autosomal dominant semi-lethal inheritance over decades. However, the genetic basis of the Creeper trait remains unknown. Here we have utilized ultra-deep sequencing and extensive analysis for targeting causative mutation controlling the Creeper trait. Our results indicated that the deletion of Indian hedgehog (IHH) gene was only found in the whole-genome sequencing data of lethal embryos and Creeper chickens. Large scale segregation analysis demonstrated that the deletion of IHH was fully linked with early embryonic death and the Creeper trait. Expression analysis showed a much lower expression of IHH in Creeper than wild-type chickens. We therefore suggest the deletion of IHH to be the causative mutation for the Creeper trait in chicken. Our findings unravel the genetic basis of the longstanding Creeper phenotype mystery in chicken as the same gene also underlies bone dysplasia in human and mouse, and thus highlight the significance of IHH in animal development and human haploinsufficiency disorders.

Animals↗

Pervasive noise in human splice site selection.

RNA splicing has historically been thought to be highly efficient and accurate, with little opportunity for deviation from regulated alternative splicing decisions. This dogma has been challenged by recent observations that suggest that biological noise may contribute substantially to transcriptome diversity. However, quantitative understanding of stochastic variations in splicing is challenging because these transcripts are likely subject to rapid degradation. Here, we use ultra-deep sequencing across RNA compartments to track splicing intermediates in human cells and see abundant cryptic splicing associated with genomic features that promote splicing noise. We observe pervasive usage of low-fidelity splice sites, likely due to stochasticity in recruitment or binding of the spliceosome. These sites are most likely degraded in the nucleus rather than targeted by translation-dependent degradation processes, suggesting widespread surveillance and rapid quality control of non-productive RNA transcripts. Our findings provide unprecedented insights into the propensity for error in RNA processing mechanisms and the regulation of alternative splice sites across a gene.

Journal Article↗

A single-cell transcriptomic atlas of the pigtail macaque placenta in late gestation.

The placenta is a complex organ with multiple immune and non-immune cell types that promote fetal tolerance and facilitate the transfer of nutrients and oxygen. The nonhuman primate (NHP) is a key experimental model for studying human pregnancy complications, in part due to similarities in placental structure, which makes it essential to understand how single-cell populations compare across the human and NHP maternal-fetal interface. We constructed a single-cell RNA-Seq (scRNA-Seq) atlas of the placenta from the pigtail macaque ( Macaca nemestrina ) in the third trimester, comprising three different tissues at the maternal-fetal interface: the chorionic villi (placental disc), chorioamniotic membranes, and the maternal decidua. Each tissue was separately dissociated into single cells and processed through the 10X Genomics and Seurat pipeline, followed by aggregation, unsupervised clustering, and cluster annotation. Next, we determined the maternal-fetal origins of cell populations and analyzed single-cell RNA trajectory, Gene Ontology enrichment, and cell-cell communication. Single-cell populations in the pigtail macaque were strikingly similar in their identity and frequency to those found in the human placenta, including cells from trophoblast, stromal cell, immune, and macrophage lineages. An advantage of our approach was the deep sequencing of three tissues at the maternal-fetal interface, which yielded a rich diversity of common and rare single-cell populations. The third-trimester pigtail macaque single-cell atlas enables the identification of cellular subclusters analogous to those in humans and provides a powerful resource for understanding experimental perturbations on the NHP placenta.

Journal Article↗

Patterns of HIV-1 viral load suppression and drug resistance during the dolutegravir transition: a population-based longitudinal study.

BACKGROUND: Data on the population-scale impact of dolutegravir (DTG)-based HIV regimens in sub-Saharan Africa are extremely limited. We used data from a surveillance cohort in southern Uganda to assess viral suppression and antiretroviral (ART) resistance over 10-years alongside DTG scale-up. METHODS: Consenting participants in the population-based Rakai Community Cohort Study between August 2011 and March 2023 aged 15-59 completed questionnaires and provided samples for HIV testing, viral load quantification, and viral deep-sequencing. We collected data on DTG-utilization at HIV care clinics. We estimated the prevalence of HIV suppression (<1,000 copies/mL) and ART resistance using robust Poisson regression. Bayesian logistic regression quantified associations between resistance and individual-level suppression across surveys. FINDINGS: Among 20,383 people living with HIV (PLHIV), suppression increased from 57.1% (95% confidence interval [CI]: 55.4%-58.8%) to 90.3% (95%CI: 89.2%-91.4%) between 2014 and 2022. By 2020 84.4% (95%CI: 83.7%-85.2%) and 64.6% (95%CI: 63.9%-65.3%) of men and women were on DTG regimens. Among treatment-experienced viremic PLHIV, overall resistance decreased from 51.1% (95%CI: 40.7%-64.1%, 2014) to 27.9% (95%CI: 21.3%-36.5%, 2022). Only two participants harbored intermediate/high-level DTG resistance, attributable to inQ148R, inE138K, and inG140A. Low-level INSTI resistance (inS153Y) was observed in 23/207 (7.5%) of viremic individuals, with putative evidence of transmission. By 2022, suppression was unrelated to prior history of NNRTI/NRTI resistance (risk ratios: 1.14, 95%HPD: 0.96-1.32 and 1.12, 95%HPD: 0.88 - 1.35). INTERPRETATION: Viral suppression increased during the DTG-transition with minimal emerging intermediate/high-level resistance. Falling resistance among treatment-experienced PLHIV underscores the role of ART adherence in reducing viremia. The emergence of inS153Y justifies continued genomic surveillance of ART resistance. FUNDING: National Institutes of Health and the Gates Foundation.

Journal Article↗

High throughput screening of eukaryotic release factor 1 variants to enhance noncanonical amino acid incorporation.

Noncanonical amino acids (ncAAs) enable diversification of protein functions, but the efficiency of genetic code expansion (GCE) in eukaryotes is hindered by competition between suppressor tRNAs and release factors. Prior work has identified eukaryotic release factor 1 (eRF1) mutants that improve ncAA incorporation, suggesting that screens for improved variants may lead to further enhancements. Here, we developed a high-throughput system to screen eRF1 mutants in Saccharomyces cerevisiae where eRF1 mutants are coexpressed on a plasmid alongside genomically encoded, wild-type eRF1. This strategy enabled recovery of live cells expressing eRF1 variants that enhance ncAA incorporation, even with mutants known to severely affect cell viability in the absence of WT eRF1 expression. We prepared and screened a million-member library of randomly mutated eRF1 variants for clones exhibiting improved ncAA integration phenotypes. Deep sequencing revealed a diverse set of enriched mutations across all three major domains of eRF1. Interestingly, several enriched mutations identified here are also found in naturally occurring eRF1 homologs from species that recode canonical stop codons. When eRF1 variants were combined with yeast knockout strains also known to enhance ncAA incorporation, this resulted in further improvements to efficiency, highlighting the complementarity of release factor engineering to other GCE enhancement strategies. This work demonstrates that high-throughput engineering of the eukaryotic translational apparatus is a powerful approach to identify previously unknown solutions for enhancing ncAA incorporation, with implications for elucidating and precisely manipulating the molecular functions of essential translational machinery.

Noncanonical amino acids↗

The fates of cells in the developing cerebral cortex of normal and methylazoxymethanol acetate-lesioned mice.

We are interested in the mechanisms that generate the mature cerebral cortex. We used bromodeoxyuridine (BrdU) to label cortical cells as they were being born. We followed the fates of specific sets of cortical precursors in normal mice and in mice in which other groups of cortical progenitors had been destroyed with the antimitotic agent methylazoxymethanol acetate (MAM Ac). In normal mice, most cells destined for the cerebral cortex were produced from embryonic day 12 (E12) to E16 in the expected inside-to-outside sequence (deep layers first, superficial layers last). Injection of MAM Ac at E13 killed cells that would normally have contributed to the deep cortical layers. As a consequence, the cortex was thinned by approximately 25% at postnatal day 21 (P21). However, all laminae were present and had normal connections with subcortical structures, although all were proportionately thinner. BrdU injected on E16 labelled a normally sized complement of cells that spanned a larger proportion of the depth of the thinned cortex. Thus, the deep cortical layers comprised many cells that were born several days later than normal. At embryonic ages prior to E12, a transient set of cells is produced in the early telencephalon. After injection with MAM Ac at E10, the cortex appeared histologically and histochemically normal at P21. However, many cells that would normally have contributed to superficial cortex (born on E15) were significantly deeper than normal. These results suggest that, during the early stages of cortical development, the nervous system is sufficiently plastic to compensate to some extent for the destruction of specific precursor cells by altering the fates of neurons born later. They indicate that the embryonic date on which a cortical cell is born does not necessarily determine its eventual phenotype.

Animals↗

Principles of bacterial genome organization, a conformational point of view.

Bacterial chromosomes are large molecules that need to be highly compacted to fit inside the cells. Chromosome compaction must facilitate and maintain key biological processes such as gene expression and DNA transactions (replication, recombination, repair, and segregation). Chromosome and chromatin 3D-organization in bacteria has been a puzzle for decades. Chromosome conformation capture coupled to deep sequencing (Hi-C) in combination with other "omics" approaches has allowed dissection of the structural layers that shape bacterial chromosome organization, from DNA topology to global chromosome architecture. Here we review the latest findings using Hi-C and discuss the main features of bacterial genome folding.

Genome, Bacterial↗

Central conducting lymphatic anomaly: from bench to bedside.

Central conducting lymphatic anomaly (CCLA) is a complex lymphatic anomaly characterized by abnormalities of the central lymphatics and may present with nonimmune fetal hydrops, chylothorax, chylous ascites, or lymphedema. CCLA has historically been difficult to diagnose and treat; however, recent advances in imaging, such as dynamic contrast magnetic resonance lymphangiography, and in genomics, such as deep sequencing and utilization of cell-free DNA, have improved diagnosis and refined both genotype and phenotype. Furthermore, in vitro and in vivo models have confirmed genetic causes of CCLA, defined the underlying pathogenesis, and facilitated personalized medicine to improve outcomes. Basic, translational, and clinical science are essential for a bedside-to-bench and back approach for CCLA.

Cell-Free Nucleic Acids↗

A custom library construction method for super-resolution ribosome profiling in Arabidopsis.

BACKGROUND: Ribosome profiling, also known as Ribo-seq, is a powerful technique to study genome-wide mRNA translation. It reveals the precise positions and quantification of ribosomes on mRNAs through deep sequencing of ribosome footprints. We previously optimized the resolution of this technique in plants. However, several key reagents in our original method have been discontinued, and thus, there is an urgent need to establish an alternative protocol. RESULTS: Here we describe a step-by-step protocol that combines our optimized ribosome footprinting in plants with available custom library construction methods established in yeast and bacteria. We tested this protocol in 7-day-old Arabidopsis seedlings and evaluated the quality of the sequencing data regarding ribosome footprint length, mapped genomic features, and the periodic properties corresponding to actively translating ribosomes through open resource bioinformatic tools. We successfully generated high-quality Ribo-seq data comparable with our original method. CONCLUSIONS: We established a custom library construction method for super-resolution Ribo-seq in Arabidopsis. The experimental protocol and bioinformatic pipeline should be readily applicable to other plant tissues and species.

3-nt periodicity↗

Characterization of the brain virome in human immunodeficiency virus infection and substance use disorder.

Viruses can infect the brain in individuals with and without HIV-infection: however, the brain virome is poorly characterized. Metabolic alterations have been identified which predispose people to substance use disorder (SUD), but whether these could be triggered by viral infection of the brain is unknown. We used a target-enrichment, deep sequencing platform and bioinformatic pipeline named "ViroFind", for the unbiased characterization of DNA and RNA viruses in brain samples obtained from the National Neuro-AIDS Tissue Consortium. We analyzed fresh frozen post-mortem prefrontal cortex from 72 individuals without known viral infection of the brain, including 16 HIV+/SUD+, 20 HIV+/SUD-, 16 HIV-/SUD+, and 20 HIV-/SUD-. The average age was 52.3 y and 62.5% were males. We identified sequences from 26 viruses belonging to 11 viral taxa. These included viruses with and without known pathogenic potential or tropism to the nervous system, with sequence coverage ranging from 0.03 to 99.73% of the viral genomes. In SUD+ people, HIV-infection was associated with a higher total number of viruses, and HIV+/SUD+ compared to HIV-/SUD+ individuals had an increased frequency of Adenovirus (68.8 vs 0%; p<0.001) and Epstein-Barr virus (EBV) (43.8 vs 6.3%; p=0.037) as well as an increase in Torque Teno virus (TTV) burden. Conversely, in HIV+ people, SUD was associated with an increase in frequency of Hepatitis C virus, (25 in HIV+/SUD+ vs 0% in HIV+/SUD-; p=0.031). Finally, HIV+/SUD- compared to HIV-/SUD- individuals had an increased frequency of EBV (50 vs 0%; p<0.001) and an increase in TTV viral burden, but a decreased Adenovirus viral burden. These data demonstrate an unexpectedly high variety in the human brain virome, identifying targets for future research into the impact of these taxa on the central nervous system. ViroFind could become a valuable tool for monitoring viral dynamics in various compartments, monitoring outbreaks, and informing vaccine development.

Male↗

Longitudinal analysis of high-risk HPV infections reveals within-host viral genome changes over time.

Persistent infection with high-risk (HR)-HPV causes cervical cancer, however, it is unclear why most infections resolve while a minority progress. We deep sequenced the HPV genomes of 1,228 HR-HPV-positive serial samples from 351 women with persistent infections (2-10 serial samples per woman over 1-8 years), including 279 controls and 72 precancer/cancer cases, to assess HR-HPV genome changes during infection and relation to infection outcomes. Seventy-seven percent of persistent infections (45-97% by HPV type) were infections with the same exact viral genome isolate; for HPV16, only 52% were persistent with the same isolate. This may suggest some infections include a type-specific isolate switch or new isolate infection during persistence. We additionally observed within-host change to the HPV genome estimated as gradual changes to intrahost single nucleotide variant (iSNV) frequency, and changes varied by HPV type, with HPV33 infections showing the most iSNV changes. Cases exhibited fewer viral genome changes during infection compared to controls (OR = 0.31, 95% CI&#x2009;=&#x2009;0.1 - 0.86, p&#x2009;=&#x2009;0.019), suggesting a more stable and clonal viral genome in cases. By viral gene, E7 had fewer nonsynonymous mutations in the cases compared to controls that cleared within 2 years of infection (p&#x2009;=&#x2009;0.012), which confirms the importance of E7 conservation and suggests mutations to E7 reduce persistence associated with progression. There was a similar pattern in E4 (p&#x2009;=&#x2009;0.013), while E5 had more changes in the cases (p&#x2009;=&#x2009;0.008). A subset of 28 infections had an intervening HPV-negative sample between HPV-positive visits; 93% of these infections had the same exact viral genome isolate in the samples before and after the negative, consistent with subclinical persistence and subsequent re-detection. Our data suggests that HR-HPV type-persistence can include a collection of viral isolates, and viral mutations during infection, particularly in E7, reduce HR-HPV persistence and thus carcinogenic potential.

Humans↗

The use of next-generation sequencing in personalized medicine.

The revolutionary progress in development of next-generation sequencing (NGS) technologies has made it possible to deliver accurate genomic information in a timely manner. Over the past several years, NGS has transformed biomedical and clinical research and found its application in the field of personalized medicine. Here we discuss the rise of personalized medicine and the history of NGS. We discuss current applications and uses of NGS in medicine, including infectious diseases, oncology, genomic medicine, and dermatology. We provide a brief discussion of selected studies where NGS was used to respond to wide variety of questions in biomedical research and clinical medicine. Finally, we discuss the challenges of implementing NGS into routine clinical use.

High-throughput sequencing↗

Genome mining reveals an architecturally expanded pyoluteorin-associated biosynthetic gene cluster and a divergent flavin-dependent halogenase-like sequence in deep-sea Pseudomonas Aeruginosa from the Gulf of Guinea.

BACKGROUND: Marine deep-sea environments harbour microorganisms with extraordinary biosynthetic potential, yet their secondary metabolite repertoires remain largely uncharacterised. RESULTS: This study reports the isolation, phenotypic characterisation, and whole-genome analysis of Pseudomonas aeruginosa strain E1, recovered from deep Atlantic seawater (Gulf of Guinea, ~2500&#xa0;m depth), which exhibits antifungal activity against multidrug-resistant Candida parapsilosis. Three presumptive P. aeruginosa isolates (E1, E17, and E44) showed >&#x2009;99% 16S rRNA gene sequence identity to P. aeruginosa reference sequences, while whole-genome dDDH analysis of strain E1 yielded 95.2% (95% CI: 93.6-96.4%; formula d4) relative to the P. aeruginosa type strain DSM 50071&#x1d40; (=&#x2009;ATCC 10145&#x1d40;), supporting its species-level assignment. Antifungal screening and PCR-based detection of flavin-dependent halogenase genes identified strain E1 as the primary candidate for genomic investigation. Illumina whole-genome sequencing produced a 6.33&#xa0;Mb draft genome assembly (113 contigs, 5862 protein-coding genes, 66.4% GC content). Genome mining with antiSMASH 8.0 identified 27 biosynthetic gene clusters (BGCs) spanning nonribosomal peptide synthetase (NRPS), polyketide synthase (PKS), phenazine, terpene, and metallophore pathways. Region 7.1 of strain E1 harbours a predicted 50.8&#xa0;kb pyoluteorin-associated BGC, comprising 34 genes, substantially larger than its terrestrial counterpart (~&#x2009;22&#xa0;kb, ~&#x2009;17 genes), and featuring nine transport genes and three regulatory elements. Phylogenetic analysis resolved three halogenase genes: ctg7_146 showed 98.7% amino acid identity to PltA, and ctg7_149 showed 99.2% amino acid identity to PltM, supporting their annotation as PltA-like and PltM-like components of the predicted pyoluteorin biosynthetic pathway. Among the characterised reference enzymes included in this analysis, ctg7_143 showed the highest amino acid identity to PltM from P. fluorescens Pf-5. However, the identity remained low at approximately 30.4%, supporting its placement as a divergent FDH-like sequence rather than a close PltM orthologue. CONCLUSION: This study provides the first comprehensive genomic characterisation of a pyoluteorin-BGC-harbouring marine P. aeruginosa strain, demonstrating conservation of the core biosynthetic machinery alongside an expanded transport architecture and a divergent FDH-like sequence that may represent a candidate for future biochemical investigation. These findings expand current knowledge of FDH-like sequence diversity in deep-sea bacteria and support further investigation of Gulf of Guinea microorganisms as a potential source of biosynthetic and enzymatic diversity.

Multigene Family↗

DeepGeSeq: deep learning library for genomic sequence modeling and analysis.

MOTIVATION: Deep learning methods have demonstrated significant potential in genomics, enabling broad applications such as sequence activity prediction, regulatory rule identification, and variant effect quantification. However, their widespread adoption is often hindered by the steep computational learning curve required for model construction, training, and downstream biological interpretation. Here, we introduce DeepGeSeq, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis. RESULTS: By integrating state-of-the-art architectural modules, DeepGeSeq streamlines the entire deep learning workflow, requiring minimal user input via a simple configuration file and an intuitive agentic skill. We comprehensively validate the efficacy of DeepGeSeq through diverse case studies, encompassing pipeline verification using synthetic datasets, the reproduction and application of established models, and model fine-tuning coupled with biological interpretation on user-defined data. Furthermore, we demonstrate DeepGeSeq's versatility in domain-specific applications, including single-cell ATAC-seq modeling for cell-type clustering, and MPRA data modeling coupled with in silico saturation mutagenesis to dissect cis-regulatory elements. Ultimately, DeepGeSeq bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development and broad application of deep learning methods in genomics research. AVAILABILITY AND IMPLEMENTATION: https://github.com/JiaqiLi1024/DeepGeSeq.

Deep Learning↗

Deep generative models in biological sequence and structure analysis and design.

Deep generative models have transformed biological sequence modeling from predictive analysis toward increasingly controllable design. Early biological applications of Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) established latent representation learning and sequence synthesis, while recent advances in transformer-based language models, discrete diffusion, flow-matching, and multimodal generative frameworks have substantially expanded the scope of biological design. This review examines generative models for DNA, RNA, and protein sequence design, emphasizing how different model classes represent biological constraints, operate over discrete and continuous spaces, and integrate sequence, structure, and function. We compare VAEs, GANs, autoregressive and masked language models, diffusion models, and flow-based approaches across genomics, transcriptomics, and proteomics, with particular attention to controllability, long-range dependency modeling, structural grounding, generalization, and experimental utility. We further examine evaluation strategies, out-of-distribution generalization, and closed-loop design-build-test-learn workflows that connect in silico generation with empirical validation. We distinguish fundamental modality-dependent constraints including sequence discreteness, context length, structural coupling, and physical or thermodynamic requirements from architecture-dependent advantages that reflect the current state of the field. Current studies suggest that long-context models are particularly useful for genome-scale representation and sequence modeling, whereas structure-aware diffusion, flow-based, and inverse-folding approaches provide better frameworks for geometry-constrained RNA and protein design. This perspective provides a critical framework for understanding the present capabilities, limitations, and convergence of generative approaches toward reliable and experimentally grounded biological design.

Biological sequence analysis↗

Recovery and phylogenetic analysis of novel archaeal rRNA sequences from a deep-sea deposit feeder.

In 1992, two independent reports based on small-subunit rRNA gene (SSU rDNA) cloning revealed the presence of novel Archaea among marine bacterioplankton. Here, we report the presence of further novel Archaea SSU rDNA sequences recovered from the midgut contents of a deep-sea marine holothurian. Phylogenetic analyses show that these abyssal Archaea are a paraphyletic component of a highly divergent clade that also includes some planktonic sequences. Our data confirm that this clade is a deep-branching lineage in the tree of life.

Animals↗

A multi-modal transformer for cell type-agnostic regulatory predictions.

Sequence-based deep learning models have emerged as powerful tools for deciphering the cis-regulatory grammar of the human genome but cannot generalize to unobserved cellular contexts. Here, we present EpiBERT, a multi-modal transformer that learns generalizable representations of genomic sequence and cell type-specific chromatin accessibility through a masked accessibility-based pre-training objective. Following pre-training, EpiBERT can be fine-tuned for gene expression prediction, achieving accuracy comparable to the sequence-only Enformer model, while also being able to generalize to unobserved cell states. The learned representations are interpretable and useful for predicting chromatin accessibility quantitative trait loci (caQTLs), regulatory motifs, and enhancer-gene links. Our work represents a step toward improving the generalization of sequence-based deep neural networks in regulatory genomics.

Humans↗