Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Deep sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Short mononucleotide repeat sequence variability in mismatch repair-deficient cancers.

Mismatch repair-deficient cancers are characterized by widespread insertions and deletions in microsatellite sequences, including those comprised of mononucleotide repeats. Such alterations have been observed in relatively short mononucleotide tracts in several genes and often are interpreted to indicate that the affected genes normally act as tumor suppressors. To aid in the interpretation of such changes, we have systematically assessed their frequency within transcribed regions of the genome that are unlikely to play a tumorigenic role. The advent of the complete human genomic sequences of chromosome 22 allowed us to select 29 genes for this analysis, spaced at approximately 1-Mb intervals. Each of the selected genes had an (A)(8) or a (G)(8) tract deep within intronic sequences that was not included in the processed transcript. Surprisingly, we found that there was substantial variation in the prevalence of mutations among these tracts. Some tracts were altered in < 5% of the mismatch repair-deficient cancers studied, whereas other tracts were altered in nearly half of the cancers. In particular, (G)(8) tracts were considerably more prone to mutation than (A)(8) tracts, and the sequences or chromatin structures surrounding the mononucleotide tracts seemed to affect their mutability significantly.

Animals↗

Molecular phylogenetic analyses of reverse-transcribed bacterial rRNA obtained from deep-sea cold seep sediments.

A depth profile of naturally occurring bacterial community structures associated with the deep-sea cold seep push-core sediment in the Japan Trench at a depth of 5343 m were evaluated using molecular phylogenetic analyses of RNA reverse transcription-PCR (RT-PCR) amplified 16S crDNA fragments. A total of 137 clones of bacterial crDNA (complimentary rDNA) phylotypes (phylogenetic types) obtained at three different depths (2-4, 8-10 and 14-16 cm) were identified in partial crDNA sequencings. crDNA phylotypes from the cold seep sediment were dominantly composed of delta- and epsilon-Proteobacteria (36% and 42% respectively). Phylotype analysis of crDNA clone libraries and terminal restriction fragment length polymorphism (T-RFLP) analysis revealed that the majority of bacterial components shifted from delta- Proteobacteria to epsilon-Proteobacteria with increasing depth. Among the delta-proteobacterial crDNA clones, the sequences related to the genus Desulfosarcina were dominant. Almost all sequences of crDNA belonging to epsilon-Proteobacteria were affiliated with the same cluster (epsilon-CSG: epsilon-proteobacterial cold seep group), and were closely related with rDNA sequences from deep-sea hydrothermal vent environments.

Base Sequence↗

Host specificity in the Richelia-diatom symbiosis revealed by hetR gene sequence analysis.

The filamentous heterocyst-forming cyanobacterium Richelia intracellularis forms associations with diatoms and is very abundant in tropical and subtropical seas. The genus Richelia contains only one species, R. intracellularis Schmidt, although it forms associations with several diatom genera and has considerable variation in size and morphology. The genetic diversity, and possible host specificity, within the genus Richelia is unknown. Using primers against hetR, a gene unique for filamentous cyanobacteria, specific polymerase chain reaction (PCR) products were obtained from natural populations of R. intracellularis filaments associated with three diatom genera. Phylogenetic analyses of these sequences showed that they were all in the same clade. This clade contained only the R. intracellularis sequences. The genetic affiliation of hetR sequences of R. intracellularis to those of other heterocystous cyanobacteria strongly suggests that it was not closely related to endosymbiotic Nostoc spp. hetR sequences. Sequences from R. intracellularis-Hemiaulus membranaceus sampled in the Atlantic and Pacific Oceans were almost identical, demonstrating that the genetic relatedness was not dependent on geographical location. All sequences displayed a deep divergence between symbionts from different genera and a high degree of host specificity.

Bacterial Proteins↗

Is the decline of desert bighorn sheep from infectious disease the result of low MHC variation?

Bighorn sheep populations have greatly declined in numbers and distribution since European settlement, primarily because of high susceptibility to infectious diseases transmitted to them from domestic livestock. It has been suggested that low variation at major histocompatibility complex (MHC) genes, the most important genetic aspect of the vertebrate immune system, may result in high susceptibility to infectious disease. Therefore, we examined genetic polymorphism at a MHC gene (Ovca-DRB) in a large sample, both numerically and geographically, of bighorn sheep. Strikingly, there were 21 different alleles that showed extensive nucleotide and amino acid sequence divergence. In other words, low MHC variation does not appear to be the basis of the high disease susceptibility and decline in bighorn sheep. On the other hand, analysis of the pattern of the MHC polymorphism suggested that nonsynonymous substitutions predominated, especially at amino acids in the antigen-binding site. The average overall heterozygosity for the 16 amino acid positions that are part of the antigen binding site is 0.389 whereas that for the 67 amino acid positions not involved with antigen binding is 0.076. These findings imply that the diversity present in this gene is functionally significant and is, or has been, maintained by balancing selection. To examine the evolution of DRB alleles in related species, a phylogenetic analysis including other published ruminant (Bovidae and Cervidae) species, was carried out. An intermixture of sequences from bighorn sheep, domestic sheep, goats, cattle, bison, and musk ox was observed supporting trans-species polymorphism for these species. To reconcile the species and gene trees for the 104 sequences examined, 95 'deep coalescent' events were necessary, illustrating the importance of balancing selection maintaining variation over speciation events.

Amino Acid Sequence↗

Sequence of the ompH gene from the deep-sea bacterium Photobacterium SS9.

In contrast to studies of many other extremophiles, the molecular characterization of the barophilic or high-pressure-adapted bacteria of the deep ocean is virtually nonexistent. One exception is the discovery that the moderate barophile Photobacterium SS9 preferentially synthesizes a 37-kDa outer membrane protein, designated OmpH, in response to elevated hydrostatic pressure. We report here on the molecular characterization of the ompH gene. The deduced amino acid sequence of mature OmpH is similar to a number of porin proteins, including significant similarity to porin protein P2 from Haemophilus influenzae. It appears likely that OmpH is a unique porin whose synthesis is responsive to changes in the pressure regime of the deep-sea bacterium.

Adaptation, Physiological↗

Microdiversity of uncultured marine prokaryotes: the SAR11 cluster and the marine Archaea of Group I.

The SAR11 cluster and the Group I of marine Archaea represent probably the best two examples of uncultured marine prokaryotes of widespread occurrence. To study their microdiversity and distribution, a total of 81 and 48 clones, respectively, were sequenced from Mediterranean and Antarctic waters at different locations and depths. The DNA regions chosen for the analysis were the last third, approximately, of the 16S rRNA gene and the 16S-23S intergenic spacer (also known as internal transcribed spacer [ITS]). There was a high concordance in both, even with the extremely variable ITS, where potential probes have been proposed for the identification and isolation of these micro-organisms. In terms of community structure, our results show that although depth-related factors seem to be predominant in the final associations of the clones, geography also plays a significant role. A major group of surface-associated sequences was found in both SAR11 and marine Archaea. In both cases this group was relatively homogeneous containing little diversity in terms of sequence, while sequences retrieved from deep samples and some surface clones contained much more heterogeneity. As a whole, both groups of prokaryotes seem to fall within the limits of well-defined taxonomic units.

Antarctic Regions↗

Isolation and cDNA-derived amino acid sequences of hemoglobin and myoglobin from the deep-sea clam Calyptogena kaikoi.

The heterodont clam Calyptogena kaikoi, living in the cold-seep area at a depth of 3761 m of the Nankai Trough, Japan, has abundant hemoglobins and myoglobins in erythrocytes and adductor muscle, respectively. Two types of hemoglobins (Hb I and Hb II) were isolated, and the complete amino acid sequences of Hb I (145 residues) and Hb II (137 residues) were obtained with combination of cDNA and protein sequencing. The amino acid sequences of C. kaikoi Hbs I and II differed from homologous chains of the congeneric clam Calyptogena soyoae in eight and five positions, respectively. The distal (E7) His, one of the functionally important residues in hemoglobin and myoglobin, was replaced by Gln in hemoglobins of C. kaikoi. A phylogenetic analysis of clam hemoglobins indicates that the evolutionary rate of Calyptogena hemoglobins is rather faster than those of other clams, suggesting that the mutation rate might be accelerated in the deep-sea animals around the areas of cold seeps or hydrothermal vents. On the other hand, it was found unexpectedly that two myoglobins Mbs I and II, isolated from the red adductor muscle, are identical in amino acid sequence Hbs I and II, respectively. Thus it was assumed that genes for Hbs I and II are also expressed in the muscle of C. kaikoi in substitution for myoglobin gene. This suggests that the major physiological role of globins in C. kaikoi is storage of oxygen under the low oxygen conditions, rather than circulating of oxygen.

Amino Acid Sequence↗

DeepES: deep learning-based enzyme screening to identify orphan enzyme genes.

MOTIVATION: Progress in sequencing technology has led to determination of large numbers of protein sequences, and large enzyme databases are now available. Although many computational tools for enzyme annotation were developed, sequence information is unavailable for many enzymes, known as orphan enzymes. These orphan enzymes hinder sequence similarity-based functional annotation, leading gaps in understanding the association between sequences and enzymatic reactions. RESULTS: Therefore, we developed DeepES, a deep learning-based tool for enzyme screening to identify orphan enzyme genes, focusing on biosynthetic gene clusters and reaction class. DeepES uses protein sequences as inputs and evaluates whether the input genes contain biosynthetic gene clusters of interest by integrating the outputs of the binary classifier for each reaction class. The validation results suggested that DeepES can capture functional similarity between protein sequences, and it can be implemented to explore orphan enzyme genes. By applying DeepES to 4744 metagenome-assembled genomes, we identified candidate genes for 236 orphan enzymes, including those involved in short-chain fatty acid production as a characteristic pathway in human gut bacteria. AVAILABILITY AND IMPLEMENTATION: DeepES is available at https://github.com/yamada-lab/DeepES. Model weights and the candidate genes are available at Zenodo (https://doi.org/10.5281/zenodo.11123900).

Deep Learning↗

The total number, time or origin and kinetics of proliferation of neurons comprising the deep cerebellar nuclei in the rhesus monkey.

The genesis of the neurons that form the cerebellar nuclei was studied by autoradiographic methods in 30 postnatal rhesus monkeys which were exposed to 3H-thymidine at various embryonic (E) and postnatal (P) ages. As a basis for this quantitative analysis, five 2-3 month old monkeys were used for cell counting and estimation of the total number of neurons in each of the cerebellar nuclei. The results show that the cerebellar nuclei on each side contain 131,000 neurons. There are 68,000 neurons in the dentate nucleus, 25,000 neurons in the posterior interposed nucleus, and 19,000 neurons in both the anterior interposed the fastigial nuclei. All of the neurons comprising the deep nuclei are generated during the first half of the 165 days gestation period in this species. Although neurogenesis lasts from E30 through E70, approximately 81% of the neuron population is generated during a one week period between E36 and E40, with the peak of proliferation occurring at E36. Before E45 both large (maximum diameter greater than 35 micrometers) and small (maximum diameter 35 micrometers or less) neurons are produced simultaneously; after this period only small neurons are generated. Although no clearcut spatio-temporal gradients of neurogenesis could be discerned along any of the cardinal axes, each cerebellar nucleus has a somewhat distinctive developmental history in terms of the onset and cessation of neurogenesis and the tempo of cell proliferation. Thus, genesis of neurons destined for the dentate nucleus begins earlier and ends later than proliferation of the neurons that ultimately comprise the fastigial nucleus. Generation of the neurons destined for the anterior and posterior interposed nuclei follows an intermediate time course. The present data on neurogenetic sequences in the deep nuclei could not be correlated with the zonal pattern of reciprocal axonal connections that link the deep nuclei and overlying cerebellar cortex.

Aging↗

N-terminal amino acid sequences of 440 kDa hemoglobins of the deep-sea tube worms, Lamellibrachia sp.1, Lamellibrachia sp.2 and slender vestimentifera gen. sp.1 evolutionary relationship with annelid hemoglobins.

The deep-sea tube worm Lamellibrachia, belonging to the phylum Vestimentifera, contains two types of extracellular hemoglobins, a 3,000 kDa hemoglobin and a 440 kDa hemoglobin. The latter hemoglobin is composed of four heme-containing chains with molecular masses of 16-18 kDa. We have collected Lamellibrachia sp.1, Lamellibrachia sp.2 and Slender vestimentifera gen. sp.1 from the deep-sea cold-seep or hydrothermal areas at a depth of 1100-1400 m. The four constituent chains of the 440 kDa hemoglobin were isolated from each of the three tube worms by reverse-phase chromatography, and the N-terminal amino acid sequences of 16-44 residues were determined by automated protein sequencer. The amino acid sequences of the homologous chains showed high homology (76-85%), suggesting that they are closely related. The sequences also showed 45-49% homology with annelid hemoglobins. A phylogenetic tree constructed from hemoglobin sequences showed that the tube worm Lamellibrachia, the polychaete Tylorrhynchus and the oligochaete Lumbricus diverged from a common ancestor at almost the same time, about 450 million years before present.

Amino Acid Sequence↗

Sequential development of connections between striate and extrastriate visual cortical areas in the rat.

In these experiments we have asked whether the projection from the rat's primary visual cortex, area 17, to the extrastriate visual cortical area 18a is formed in a sequence and whether that sequence resembles the pattern of inside-out cortical neurogenesis. For this purpose fluorescent retrograde tracers were injected into area 18a at different postnatal ages (P1, P5, adult). Animals survived until 3-4 weeks of age, after migration is complete and neurons have arrived at their final laminar location. In the ipsilateral cortex, P1 injections retrogradely labeled cells in layers 5 and 6 of area 17. Labeling after P5 injections extended into more superficial layers and included the bottom of layer 2/3 and layers 4-6. After P5, more labeled cells were found at the top of layer 2/3, producing the adult laminar pattern, where the projection originates predominantly from layer 2/3. A similar sequence of laminar labeling was observed in the transcallosal connection of area 18a. This sequence of labeling, deep layers before superficial, resembles the pattern in which cortical neurons are born and indicates that axons arrive at their cortical targets in the order the cells were generated.

Animals↗

Mechanisms underlying memory impairment in schizophrenia.

BACKGROUND: The purpose of this experiment was to investigate mechanisms underlying commonly-observed verbal memory impairments in schizophrenia, and especially the hypothesized encoding deficit. METHODS: A verbal memory task was administered to 38 patients with schizophrenia and 38 normal controls. Three functions involved in long-term memory-encoding, early phase of storage, retrieval-were investigated. First, non-organizable lists were compared to semantically-organizable lists in a free recall task, in order to vary encoding conditions. Superficial encoding (measured by a 'sequence' index) and deep encoding (measured by a categorization index) were assessed. Secondly, early storage was investigated by varying the delay between learning and recall. Lastly, cues were provided for organizable lists (semantic cues) and non-organizable lists (recognition sheet), in order to vary retrieval conditions. RESULTS: An analysis of variance revealed an interaction between type of list (organizable, non-organizable) and group, showing that patients used organization less than controls. A further analysis showed that deep encoding was impaired. Also, although the propensity to use superficial encoding was unimpaired, its efficiency was less. The analysis of variance revealed no interaction with delay or with either type of cue. A correlation was found between deep processing and memory performance in both groups. CONCLUSIONS: A major deficit in encoding appeared in the patient group, with a lesser use of deep encoding and a lesser efficiency of superficial encoding. On the other hand, the early phase of storage and the retrieval function seemed unaffected. Overall memory performance appeared to be related to the depth of encoding.

Adult↗

Phylogeography of the tidewater goby, Eucyclogobius newberryi (Teleostei, Gobiidae), in coastal California.

The tidewater goby, Eucyclogobius newberryi, inhabits discrete, seasonally closed estuaries and lagoons along approximately 1500 km of California coastline. This species is euryhaline but has no explicit marine stage, yet population extirpation and recolonization data suggest tidewater gobies disperse intermittently via the sea. Analyses of mitochondrial control region and cytochrome b sequences demonstrate a deep evolutionary bifurcation in the vicinity of Los Angeles that separates southern California populations from all more northerly populations. Shallower phylogeographic breaks, in the vicinities of Seacliff, Point Buchon, Big Sur, and Point Arena segregate the northerly populations into five groups in three geographic clusters: the Point Conception and Ventura groups between Los Angeles and Point Buchon, a lone Estero Bay group from central California, and San Francisco and Cape Mendocino groups from northern California. The phylogenetic relationships between and patterns of molecular diversity within the six groups are consistent with repeated, and sometimes rapid, northward and southward range expansions out of central California caused by Quaternary climate change. Plio-Pleistocene tectonism, Quaternary coastal geography and hydrography, and historical human activities probably also influenced the modern geographic and genetic structure of E. newberryi. The phylogeography of E. newberryi is concordant with phylogeographic patterns in several other coastal California taxa, suggesting common extrinsic factors have had similar effects on different species. However, there is no evidence of a phylogeographic break coincident with a biogeographic boundary at Point Conception.

Animals↗

Neurogenesis and identification of developing layers in the visual cortex of the wallaby (Macropus eugenii).

Birthdates of the neurons that comprise the layers of the mature visual cortex in the wallaby (Macropus eugenii) have been determined with the aid of tritiated thymidine autoradiography. The laminar positions of cells, identified by their birthdates, have then been followed at early stages during development and compared with previously published data on the distribution of thalamocortical afferents and corticothalamic projecting cells (Sheng et al. [1991] J. Comp. Neurol. 307:17-38). Neurons are born in a deep to superficial sequence typical of other mammals. The loosely packed zone of cells, which develops at the base of the thin compact zone of cells at the superficial margin of the cortical plate early in development, was identified as being part of the cortical plate. Afferents did not wait below this zone but grew into the developing cortical layers immediately after the cells that form these layers began accumulating in the loosely packed zone, starting with layer 6 on postnatal day 22 (P22). The genesis of layer 4 did not begin until P32, and these cells reached the superficial cortical plate at P54 and entered the loosely packed zone by P65. Cells of layers 5 and 6 formed the initial projection to the thalamus. Despite the protracted development of the wallaby and the large discrepancy between the time of thalamic ingrowth and genesis of layer 4, there was no extended waiting period for afferents in the subplate.

Afferent Pathways↗

Kv3.3b: a novel Shaw type potassium channel expressed in terminally differentiated cerebellar Purkinje cells and deep cerebellar nuclei.

A two-step hybridization/subtraction procedure was employed to isolate markers for the later stages of Purkinje cell differentiation. From this screen, a novel Shaw potassium channel cDNA (Kv3.3b) was identified that is developmentally regulated. Expression of this channel is highly enriched in the brain, particularly in the cerebellum, where its expression is confined to Purkinje cells and deep cerebellar nuclei. Sequence analysis revealed that it is an alternatively spliced form of the mouse Kv3.3 gene, and that the previously reported Kv3.3 mRNA (Ghanshani et al., 1992) is not expressed in cerebellum. Expression of the Kv3.3b mRNA begins in cerebellar Purkinje cells between postnatal day 8 (P8) and P10 and continues through adulthood, coinciding with elaboration of the mature Purkinje cell dendritic arbor. The timing of expression of Kv3.3b mRNA is maintained in mixed, dissociated primary cerebellar cell culture. These results suggest that the Kv3.3b K+ channel function is restricted to terminally differentiated Purkinje cells, and that analysis of the mechanisms governing its expression in vivo and in vitro can reveal molecular mechanisms governing Purkinje cell differentiation.

Amino Acid Sequence↗

A deep intronic IFT172 variant causing pseudoexon inclusion identified by whole-genome sequencing in nephronophthisis.

Nephronophthisis is an autosomal recessive ciliopathy and a major genetic cause of end-stage kidney disease in children and young adults. Although next-generation sequencing panels have improved diagnostic yield, some patients remain genetically unresolved, partly due to deep intronic variants that disrupt pre-mRNA splicing and are not captured by exon-focused approaches. We report a 13-year-old boy who presented with advanced kidney dysfunction, small renal cysts, and kidney histopathology consistent with nephronophthisis. Targeted gene panel sequencing failed to identify causative pathogenic variants beyond a missense variant of uncertain significance. Whole-genome sequencing subsequently revealed compound heterozygous variants in IFT172 (NM_015662.3): a missense variant (c.4696C > T, p.Arg1566Cys) and a deep intronic variant (c.4915-94A > G). In silico analysis predicted activation of cryptic splice sites leading to inclusion of an 86-bp pseudoexon, which was confirmed by a minigene splicing assay. These findings established a molecular diagnosis of IFT172-related nephronophthisis. To our knowledge, this is the first report demonstrating pseudoexon inclusion in IFT172, thereby expanding its mutational spectrum. Our case underscores the importance of evaluating deep intronic regions using whole-genome sequencing and functional validation in genetically unresolved nephronophthisis.

Humans↗

SIMS: A deep-learning label transfer tool for single-cell RNA sequencing analysis.

Cell atlases serve as vital references for automating cell labeling in new samples, yet existing classification algorithms struggle with accuracy. Here we introduce SIMS (scalable, interpretable machine learning for single cell), a low-code data-efficient pipeline for single-cell RNA classification. We benchmark SIMS against datasets from different tissues and species. We demonstrate SIMS's efficacy in classifying cells in the brain, achieving high accuracy even with small training sets (<3,500 cells) and across different samples. SIMS accurately predicts neuronal subtypes in the developing brain, shedding light on genetic changes during neuronal differentiation and postmitotic fate refinement. Finally, we apply SIMS to single-cell RNA datasets of cortical organoids to predict cell identities and uncover genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Single-Cell Analysis↗

Refining sequence-to-activity models by increasing model resolution.

Decoding the cis-regulatory syntax that controls gene expression is essential for improving our understanding of cell differentiation and disease. To identify regulatory motifs and their regulatory syntax, deep learning based sequence-to-activity (S2A) models learn transcription factor binding motifs and their combinations from DNA sequence by modeling measured chromatin accessibility. Previously, we developed AI-TAC, a S2A model that predicts chromatin accessibility across various immune cell types in multi-task fashion, effectively decoding the regulatory syntax underlying immune cell differentiation. While ATAC-seq is commonly used to measure regional accessibility, it also provides high-resolution profiles, the distribution of Tn5 insertion sites, that offer additional insights into the precise location and strength of TF binding sites. Here we demonstrate that modeling ATAC-seq profiles alongside accessibility consistently improves predictions of differential chromatin accessibility across cell types. Moreover, we also find that multi-task learning across related immune cell types consistently outperforms single-task models. To understand what additional information bpAITAC learns from ATAC-seq profiles, we systematically compare sequence attributions from models trained with and without ATAC-seq profiles. We identify novel motifs with strong effect sizes that emerge only when profile data is included. Our findings suggest that modeling ATAC-seq at base-pair resolution enables the model to learn a more nuanced and sensitive representation of the cis-regulatory syntax driving immune cell-specific chromatin landscapes.

ATAC-seq↗