Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Deep sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

A deep intronic IFT172 variant causing pseudoexon inclusion identified by whole-genome sequencing in nephronophthisis.

Nephronophthisis is an autosomal recessive ciliopathy and a major genetic cause of end-stage kidney disease in children and young adults. Although next-generation sequencing panels have improved diagnostic yield, some patients remain genetically unresolved, partly due to deep intronic variants that disrupt pre-mRNA splicing and are not captured by exon-focused approaches. We report a 13-year-old boy who presented with advanced kidney dysfunction, small renal cysts, and kidney histopathology consistent with nephronophthisis. Targeted gene panel sequencing failed to identify causative pathogenic variants beyond a missense variant of uncertain significance. Whole-genome sequencing subsequently revealed compound heterozygous variants in IFT172 (NM_015662.3): a missense variant (c.4696C > T, p.Arg1566Cys) and a deep intronic variant (c.4915-94A > G). In silico analysis predicted activation of cryptic splice sites leading to inclusion of an 86-bp pseudoexon, which was confirmed by a minigene splicing assay. These findings established a molecular diagnosis of IFT172-related nephronophthisis. To our knowledge, this is the first report demonstrating pseudoexon inclusion in IFT172, thereby expanding its mutational spectrum. Our case underscores the importance of evaluating deep intronic regions using whole-genome sequencing and functional validation in genetically unresolved nephronophthisis.

Humans↗

SIMS: A deep-learning label transfer tool for single-cell RNA sequencing analysis.

Cell atlases serve as vital references for automating cell labeling in new samples, yet existing classification algorithms struggle with accuracy. Here we introduce SIMS (scalable, interpretable machine learning for single cell), a low-code data-efficient pipeline for single-cell RNA classification. We benchmark SIMS against datasets from different tissues and species. We demonstrate SIMS's efficacy in classifying cells in the brain, achieving high accuracy even with small training sets (<3,500 cells) and across different samples. SIMS accurately predicts neuronal subtypes in the developing brain, shedding light on genetic changes during neuronal differentiation and postmitotic fate refinement. Finally, we apply SIMS to single-cell RNA datasets of cortical organoids to predict cell identities and uncover genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Single-Cell Analysis↗

Refining sequence-to-activity models by increasing model resolution.

Decoding the cis-regulatory syntax that controls gene expression is essential for improving our understanding of cell differentiation and disease. To identify regulatory motifs and their regulatory syntax, deep learning based sequence-to-activity (S2A) models learn transcription factor binding motifs and their combinations from DNA sequence by modeling measured chromatin accessibility. Previously, we developed AI-TAC, a S2A model that predicts chromatin accessibility across various immune cell types in multi-task fashion, effectively decoding the regulatory syntax underlying immune cell differentiation. While ATAC-seq is commonly used to measure regional accessibility, it also provides high-resolution profiles, the distribution of Tn5 insertion sites, that offer additional insights into the precise location and strength of TF binding sites. Here we demonstrate that modeling ATAC-seq profiles alongside accessibility consistently improves predictions of differential chromatin accessibility across cell types. Moreover, we also find that multi-task learning across related immune cell types consistently outperforms single-task models. To understand what additional information bpAITAC learns from ATAC-seq profiles, we systematically compare sequence attributions from models trained with and without ATAC-seq profiles. We identify novel motifs with strong effect sizes that emerge only when profile data is included. Our findings suggest that modeling ATAC-seq at base-pair resolution enables the model to learn a more nuanced and sensitive representation of the cis-regulatory syntax driving immune cell-specific chromatin landscapes.

ATAC-seq↗

Panning for genes--A visual strategy for identifying novel gene orthologs and paralogs.

We have developed a rapid visual method for identifying novel members of gene families. Starting with an evolutionary tree, 20-50 protein query sequences for a gene family are selected from different branches of the tree. These query sequences are used to search the GenBank and expressed sequence tag (EST) DNA databases and their nightly updates using the tfastx3 or tfasty3 programs. The results of all 20-50 searches are collated and resorted to highlight EST or genomic sequences that share significant similarity with the query sequences. The statistical significance of each DNA/protein alignment is plotted, highlighting the portion of the query sequence that is present in the database sequence and the percent identity in the aligned region. The collated results for database sequences are linked using the WWW to the underlying scores and alignments; these links can also be used to perform additional searches to characterize the novel sequence further. With traditional "deep" scoring matrices (BLOSUM50) one can search for previously unrecognized families of large protein superfamilies. Alternatively, by using query sequences and EST libraries from the same species (e. g., human or mouse) together with "shallow" scoring matrices and filters that remove high-identity sequences, one can highlight new paralogs of previously described subfamilies. Using query sequences from the glutathione transferase superfamily, we identified two novel mammalian glutathione transferase families that were recognized previously only in plants. Using query sequences from known mammalian glutathione transferase subfamilies, we identified new candidate paralogs from the mouse class-mu, class-pi, and class-theta families.

Amino Acid Sequence↗

Phylogeny of oral asaccharolytic Eubacterium species determined by 16S ribosomal DNA sequence comparison and proposal of Eubacterium infirmum sp. nov. and Eubacterium tardum sp. nov.

16S rRNA gene sequences of Eubacterium brachy, Eubacterium nodatum, Eubacterium saphenum, Eubacterium timidum, and two previously unnamed taxa were determined. The results of a phylogenetic analysis indicated that all of the strains sequenced belonged to a deep branch of the low-G+C-content gram-positive group. The levels of 16S ribosomal DNA sequence similarity between species were low, suggesting that a number of genera may be represented in this group. The representatives of the two unnamed taxa, which were isolated from patients with periodontitis, were clearly distinct from the previously described species, and, therefore, the following two new species are proposed: Eubacterium infirmum (type strain, NCTC 12940) and Eubacterium tardum (type strain, NCTC 12941).

Base Sequence↗

Distribution of the pressure-regulated operons in deep-sea bacteria.

DNA regions corresponding to portions of two different pressure-regulated operons previously identified in two deep-sea barophilic bacteria were separately PCR amplified from a variety of deep-sea microorganisms and sequenced. With the two sets of primers employed, amplification was particularly successful from the more barophilic bacteria examined. 16S rRNA sequence analysis revealed that these bacteria are all phylogenetically related and belong in a sub-branch of the genus Shewanella containing only the deep-sea Shewanella barophilic bacteria. We define this sub-branch as the 'Shewanella barophile branch' containing at least two different species. Our results suggest that the DNA sequences of the pressure-regulated operons can be regarded as marker sequences to identify the Shewanella barophilic strains.

Amino Acid Sequence↗

Cloning, sequencing and overexpression of the gene encoding malate dehydrogenase from the deep-sea bacterium Photobacterium species strain SS9.

The gene encoding malate dehydrogenase (mdhA) was obtained from the psychrophilic, barophilic, deep-sea isolate Photobacterium species strain SS9. The SS9 mdhA gene directed high levels of malate dehydrogenase (MDH) production in Escherichia coli. A comparison of SS9 MDH to three mesophile MDHs, a MDH sequence obtained from another deep-sea bacterium, and to other psychrophile proteins is presented.

Amino Acid Sequence↗

Amino acid sequence of the dimeric hemoglobin (Hb I) from the deep-sea cold-seep clam Calyptogena soyoae and the phylogenetic relationship with other molluscan globins.

The deep-sea cold-seep clam Calyptogena soyoae has two homodimeric hemoglobins (Hbs I and II) in erythrocytes. The complete amino acid sequence of Hb I has been determined. It is composed of 144 amino acid residues, has a high content of hydrophobic residues, and a calculated molecular weight of 16,350 including a heme group. The sequence of Calyptogena Hb I showed high homology (42% identity) with that of Calyptogena Hb II (Suzuki, T., Takagi T. and Ohta, S. (1989) Biochem. J. 260, 177-182), although it has a long insertion of seven residues in the C-terminal region compared with Hb II. On the other hand, it showed low homology (12-20% identity) with other molluscan globins. As well as Hb II, Calyptogena Hb I lacked the N-terminal extension of 7-9 residues characteristic of molluscan intracellular hemoglobins, and the distal (E7) histidine was replaced by glutamine. A phylogenetic tree was constructed from 13 molluscan globins belonging to the five families Aplysiidae, Galeodidae, Potamididae, Arcidae and Vesicomyidae. The globin sequences of Calyptogena (Vesicomyidae) were found to be rather distant from other globin sequences, suggesting that they might conserve a primitive form of molluscan globins.

Amino Acid Sequence↗

Analysis of human immunodeficiency virus type 1 gp160 sequences from a patient with HIV dementia: evidence for monocyte trafficking into brain.

Towards understanding the pathogenesis of HIV dementia, we molecularly cloned and sequenced human immundeficiency virus type 1 (HIV-1) gp160 genes from uncultured post-mortem tissues collected from a patient with HIV dementia. Sequences from bone marrow, lymph node, lung, and four regions of brain - the deep white matter, head of caudate, choroid plexus and meninges - were compared. Also included were gp160 sequences recovered from blood monocytes collected 5 months prior to death. Phylogenetic analyses showed that the sequences from deep white matter were more closely related to those from bone marrow, than to those from the other tissues, and moreover, were most closely related to sequences from the blood monocytes. These findings suggest trafficking of bone marrow-derived monocytes into the deep white matter during this late stage of infection. Another cluster included sequences from choroid plexus, meninges and lymph node, and interestingly, identical patterns of four or nine stop codons were shared among these tissues. These mutations appear to be the consequence of G-->A hypermutation, and could reflect independent events, or the movement of virions or infected cells, from the choroid plexus into the cerebrospinal fluid and ultimately, into the lymph node. We propose that a critical step towards the development of HIV dementia is an increase in monocyte trafficking into the brain, and that this process is either initiated and/or accelerated during late-stage infection, which could explain why dementia occurs primarily during this time.

AIDS Dementia Complex↗

The ribonucleotide sequence of 5s rRNA from two strains of deep-sea barophilic bacteria.

Deep-sea bacteria were isolated from the digestive tract of animals inhabiting depths of 5900 m in the Puerto Rico Trench and 4300 m near the Walvis Ridge. Growth of two bacterial strains was measured in marine broth and in solid media under a range of pressures and temperatures. Both strains were barophilic at 2 degrees C (+/- 1 degrees C) with an optimal growth rate of 0.22 h-1 at a pressure 30% lower than that encountered in situ. At 1 atm they grew at temperatures ranging from 1.2 to 18.2 degrees C (+/- 0.3 degrees C), while in situ pressures increased the upper temperature limit to 23.3 degrees C. Both strains were identified as members of the genus Vibrio, based on standard taxonomic tests and mol% G + C values (47.0 and 47.1). Ribonucleotide sequences determined for 5S ribosomal RNA from each strain confirmed relationship to the Vibrio-Photobacterium group, as represented by V. harveyi and P. phosphoreum, but the barophiles were clearly distinct from these species. Secondary structure conformed to the established model for eubacterial 5S rRNA.

Animals↗

Structural genomics sheds light on protein functions and remote homologs across the insect tree of life.

Protein structure bridges the sequence-function relationship, enabling deep exploration of biological processes across diverse organisms. Insects, the most diverse animal lineage, accounting for over 50% of all described animal species, provide an exceptional system for exploring sequence-structure-function relationships. Here, we reconstructed a comprehensive and well-resolved phylogeny of 4854 insects, spanning all orders. Leveraging this framework, we created an atlas of 13.29 million predicted protein structures from 824 representative species, including 11.63 million newly predicted structures. Structural clustering revealed that proteins with divergent sequences but similar structures could be effectively grouped together. Structural similarity searches against proteins with well-characterized functions yielded annotations for 7.61 million insect proteins, including up to 14% of previously unannotated proteins. We further identified 750 million remote homologs between insect proteins, many of which trace back to ancient branches of the insect phylogeny. Remarkably, despite extensive sequence divergence, cGAS-like receptors (cGLRs) were structurally conserved across all 824 insects. Experimental assays demonstrated that these structurally identified cGLRs play a crucial role in antiviral defense in the yellow fever mosquito. Our findings highlight the significance of structural genomics for understanding protein function and evolution across the tree of life.

Animals↗

Active learning of enhancer and silencer regulatory grammar in photoreceptors.

Cis-regulatory elements (CREs) direct gene expression in health and disease, and models that can accurately predict their activities from DNA sequences are crucial for biomedicine. Deep learning represents one emerging strategy to model the regulatory grammar that relates CRE sequence to function. However, these models require training data on a scale that exceeds the number of CREs in the genome. We address this problem using active machine learning to iteratively train models on multiple rounds of synthetic DNA sequences assayed in live mammalian retinas. During each round of training the model actively selects sequence perturbations to assay, thereby efficiently generating informative training data. We iteratively trained a model that predicts the activities of sequences containing binding motifs for the photoreceptor transcription factor Cone-rod homeobox (CRX) using an order of magnitude less training data than current approaches. The model's internal confidence estimates of its predictions are reliable guides for designing sequences with high activity. The model correctly identified critical sequence differences between active and inactive sequences with nearly identical transcription factor binding sites, and revealed order and spacing preferences for combinations of motifs. Our results establish active learning as an effective method to train accurate deep learning models of cis-regulatory function after exhausting naturally occurring training examples in the genome.

Journal Article↗

Growth-rate-dependent regulation of 6-phosphogluconate dehydrogenase level mediated by an anti-Shine-Dalgarno sequence located within the Escherichia coli gnd structural gene.

Previous work has shown that in Escherichia coli K-12 growth-rate-dependent regulation of expression of 6-phosphogluconate dehydrogenase, encoded by the gnd gene, occurs at the posttranscriptional level and is mediated by a negative control element that lies deep in the coding sequence, somewhere between codons 48 and 118. Deletion analysis of a growth-rate-regulated gnd-lacZ translational fusion showed that the element is the segment of gnd mRNA between codons 67 and 78 that is complementary to an extensive portion of the gnd ribosome-binding site, including its Shine-Dalgarno sequence. The boundaries of the element were further defined by the cloning of a synthetic "internal complementary sequence." The core internal complementary sequence element effected growth-rate-dependent regulation when placed at several sites between codon 40 and codon 69, but it severely reduced gene expression when moved to codon 13. The effect on regulation of single and double mutations introduced into the element by site-directed mutagenesis correlated with the ability of the respective mRNAs to fold into secondary structures that sequester the ribosome-binding site. Thus the gnd gene's internal regulatory element appears to function as a cis-acting antisense RNA.

Base Sequence↗

Genetic diversity among Arthrobacter species collected across a heterogeneous series of terrestrial deep-subsurface sediments as determined on the basis of 16S rRNA and recA gene sequences.

This study was undertaken in an effort to understand how the population structure of bacteria within terrestrial deep-subsurface environments correlates with the physical and chemical structure of their environment. Phylogenetic analysis was performed on strains of Arthrobacter that were collected from various depths, which included a number of different sedimentary units from the Yakima Barricade borehole at the U.S. Department of Energy's Hanford site, Washington, in August 1992. At the same time that bacteria were isolated, detailed information on the physical, chemical, and microbiological characteristics of the sediments was collected. Phylogenetic trees were prepared from the 39 deep-subsurface Arthrobacter isolates (as well as 17 related type strains) based on 16S rRNA and recA gene sequences. Analyses based on each gene independently were in general agreement. These analyses showed that, for all but one of the strata (sedimentary layers characterized by their own unifying lithologic composition), the deep-subsurface isolates from the same stratum are largely monophyletic. Notably, the layers for which this is true were composed of impermeable sediments. This suggests that the populations within each of these strata have remained isolated under constant, uniform conditions, which have selected for a particular dominant genotype in each stratum. Conversely, the few strains isolated from a gravel-rich layer appeared along several lineages. This suggests that the higher-permeability gravel decreases the degree of isolation of this population (through greater groundwater flow), creating fluctuations in environmental conditions or allowing migration, such that a dominant population has not been established. No correlation was seen between the relationship of the strains and any particular chemical or physical characteristics of the sediments. Thus, this work suggests that within sedimentary deep-subsurface environments, permeability of the deposits plays a major role in determining the genetic structure of resident bacterial populations.

Arthrobacter↗

Characterization of attached bacterial populations in deep granitic groundwater from the Stripa research mine by 16S rRNA gene sequencing and scanning electron microscopy.

This paper presents the molecular characterization of attached bacterial populations growing in slowly flowing artesian groundwater from deep crystalline bed-rock of the Stripa mine, south central Sweden. Bacteria grew on glass slides in laminar flow reactors connected to the anoxic groundwater flowing up through tubing from two levels of a borehole, 812-820 m and 970-1240 m. The glass slides were collected, the bacterial DNA was extracted and the 16S rRNA genes were amplified by PCR using primers matching universally conserved positions 519-536 and 1392-1405. The resulting PCR fragments were subsequently cloned and sequenced. The sequences were compared with each other and with 16S rRNA gene sequences in the EMBL database. Three major groups of bacteria were found. Signature bases placed the clones in the appropriate systematic groups. All belonged to the proteobacterial groups beta and gamma. One group was found only at the 812-820 m level, where it constituted 63% of the sequenced clones, whereas the second group existed almost exclusively at the 970-1240 m level, where it constituted 83% of the sequenced clones. The third group was equally distributed between the levels. A few other bacteria were also found. None of the 16S rRNA genes from the dominant bacteria showed more than 88% similarity to any of the others, and none of them resembled anything in the database by more than 96%. Temperature did not seem to have any effect on species composition at the deeper level. SEM images showed rods appearing in microcolonies.(ABSTRACT TRUNCATED AT 250 WORDS)

Anaerobiosis↗

Emergence of master sequences in families of retroposons derived from 7sl RNA.

The past few years have brought new insight into the evolution of families of retroposons. These are composed of a very small number of master sequences able to duplicate, and a large majority of copies that are inactive for retroposition. During the course of time, successive replacements of master sequences have produced waves of amplification that are recognizable as subfamilies. In the Alu and the B1 families, one can distinguish two evolutionary periods. The first involves only monomeric elements that are now extinguished (fossil elements) and is characterized by deep remodeling of the sequences. This period ends, in primates, with the fusion of a free left and a free right Alu monomer, producing the first modern Alu dimeric element; in rodents it ends with a tandem duplication of 29 bp to create the first modern B1 element. The second period is characterized by a great stability of the master sequences. The observed turn-over of master sequences is still an enigma. However, analysis of the contemporary master sequences and of the oldest master sequences provide some clues. Here, we review the very first stages of the appearance of the Alu and the B1 families in mammalian genomes.

Animals↗

Chronic atrophic gastritis: early diagnosis in a population where Helicobacter pylori infection is frequent.

Chronic atrophic gastritis (CAG) is a premalignant condition characterized by loss of gastric antral deep glands. The histologic changes in antral gastric biopsy specimens from 54 Peruvian patients with dyspepsia were studied to detail the development and characteristics of CAG. Ninety-six percent of the biopsies revealed severe superficial mucosal inflammation and 89% showed deep inflammation. Moderate or severe CAG was present in 36 (67%) of the 54 patients. In the early stages of CAG, a glandular lymphoid adherence lesion was noted in 17 (31%) of the 54 biopsy specimens. This lesion consisted of lymphocytes adherent to the antral deep gland cells and was associated with glandular epithelium alterations. The late stage was characterized by small glands, remnants of glands, and gland replacement with a fibrocellular infiltrate or intestinal metaplasia. We propose that the development of CAG probably proceeds via a stereotyped sequence, with an early deep inflammatory component that may trigger local gland destruction and eventual permanent loss.

Adolescent↗