Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Deep sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Root of the universal tree of life based on ancient aminoacyl-tRNA synthetase gene duplications.

Universal trees based on sequences of single gene homologs cannot be rooted. Iwabe et al. [Iwabe, N., Kuma, K.-I., Hasegawa, M., Osawa, S. & Miyata, T. (1989) Proc. Natl. Acad. Sci. USA 86, 9355-9359] circumvented this problem by using ancient gene duplications that predated the last common ancestor of all living things. Their separate, reciprocally rooted gene trees for elongation factors and ATPase subunits showed Bacteria (eubacteria) as branching first from the universal tree with Archaea (archaebacteria) and Eucarya (eukaryotes) as sister groups. Given its topical importance to evolutionary biology and concerns about the appropriateness of the ATPase data set, an evaluation of the universal tree root using other ancient gene duplications is essential. In this study, we derive a rooting for the universal tree using aminoacyl-tRNA synthetase genes, an extensive multigene family whose divergence likely preceded that of prokaryotes and eukaryotes. An approximately 1600-bp conserved region was sequenced from the isoleucyl-tRNA synthetases of several species representing deep evolutionary branches of eukaryotes (Nosema locustae), Bacteria (Aquifex pyrophilus and Thermotoga maritima) and Archaea (Pyrococcus furiosus and Sulfolobus acidocaldarius). In addition, a new valyl-tRNA synthetase was characterized from the protist Trichomonas vaginalis. Different phylogenetic methods were used to generate trees of isoleucyl-tRNA synthetases rooted by valyl- and leucyl-tRNA synthetases. All isoleucyl-tRNA synthetase trees showed Archaea and Eucarya as sister groups, providing strong confirmation for the universal tree rooting reported by Iwabe et al. As well, there was strong support for the monophyly (sensu Hennig) of Archaea. The valyl-tRNA synthetase gene from Tr. vaginalis clustered with other eukaryotic ValRS genes, which may have been transferred from the mitochondrial genome to the nuclear genome, suggesting that this amitochondrial trichomonad once harbored an endosymbiotic bacterium.

Amino Acid Sequence↗

Improving spliced alignment by modeling splice sites with deep learning.

MOTIVATION: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alignment demands sophisticated modeling of splice sites, but current aligners use simple models, which may affect their accuracy given dissimilar sequences. RESULTS: We implemented minisplice to learn splice signals with a one-dimensional convolutional neural network (1D-CNN) and trained a model with 7,026 parameters for vertebrate and insect genomes. It captures conserved splice signals across phyla and reveals GC-rich introns specific to mammals and birds. We used this model to estimate the empirical splicing probability for every GT and AG in genomes, and modified minimap2 and miniprot to leverage pre-computed splicing probability during alignment. Evaluation on human long-read RNA-seq data and cross-species protein datasets showed our method greatly improves the junction accuracy especially for noisy long RNA-seq reads and proteins of distant homology. AVAILABILITY AND IMPLEMENTATION: https://github.com/lh3/minisplice.

Journal Article↗

Exploring the use of machine and deep learning in genome-wide association studies: a comprehensive review.

The advent of high-throughput sequencing technologies has generated increasingly large and complex genomic datasets, necessitating analytical approaches capable of capturing high-dimensional and potentially nonlinear genetic interactions. This situation has significantly impacted the entire field of Genome-Wide Association Study (GWAS), whose primary goal is the identification of genomic traits and variants that are statistically associated with the risk of a disease. However, traditional GWAS methods may show reduced performance when applied to highly polygenic and nonlinear genetic architectures. Computational strategies from Artificial Intelligence (AI) and, in particular, from machine- and deep-learning may provide a powerful tool to overcome such limitations, especially by capturing nonlinear interactions and complex hidden regularities in large-scale data, which traditional GWAS approaches might overlook. To date, only a few approaches have been introduced and systematically assessed. In this review, we describe the main characteristics and limitations of standard statistical approaches for GWAS, the main uses of AI methods in computational genomics, and recent attempts to leverage AI strategies in GWAS. Particular attention will be devoted to key issues, such as the interpretability of methods and results, and the curse of dimensionality. More specifically, the review presents 30 methods designed to leverage AI in GWAS, as well as presenting a comprehensive set of evaluation metrics for their performance, also providing references to the most frequently used databases, and biobanks. Overall, this work may serve as a starting point for both dry- and wet-lab researchers, aiming to extract deeper insights from genomic data by moving beyond traditional linear additive assumptions, and leveraging large-scale datasets through AI-driven approaches.

Artificial intelligence↗

Diffusion-weighted magnetic resonance imaging in early stage of 5-fluorouracil-induced leukoencephalopathy.

We report a case of 5-fluorouracil (5-FU)-induced leukoencephalopathy in which magnetic resonance imaging (MRI) of the brain, including diffusion-weighted imaging (DWI), was performed serially. The initial T2-weighted and FLAIR images showed diffuse mild hyperintensity in bilateral deep cerebral white matter and corpus callosum, which on T1WI appeared as non-enhanced faint hypointensity. Isotropic DWI disclosed the abnormality as well-conspicuous diffuse hyperintensity with decreased ADC. Serial studies revealed that majority of the abnormal signal intensity on these sequences resolved, and the decreased ADC values approached normal. Some hyperintensity remained in the deep cerebral white matter and the splenium, but no further significant ADC change after normalization was noted. Measurement of ADC along the three orthogonal directions showed the presence of directional dependence of diffusion throughout the length of study. These findings suggest that early stage of 5-FU-induced leukoencephalopathy is associated with reversible restricted diffusion and preservation of anisotropy. Diffusion-weighted imaging may be useful for the diagnosis.

Adult↗

NextVir: Enabling classification of tumor-causing viruses with genomic foundation models.

MOTIVATION: Oncoviruses, pathogens known to cause or increase the risk of cancer, include both common viruses such as human papillomaviruses and rarer pathogens such as human T-lymphotropic viruses. Computational methods for detecting viral DNA from data acquired by modern DNA sequencing technologies have enabled studies of the association between oncoviruses and cancers. Those studies are rendered particularly challenging when multiple species of oncovirus are present in a tumor sample. In such scenarios, merely detecting the presence of a sequencing read of viral origin is insufficiently informative-instead, a more precise characterization of the viral content in the sample is required. RESULTS: We address this need with NextVir, to our knowledge the first multi-class viral classification framework that adapts genomic foundation models to detecting and classifying sequencing reads of oncoviral origin. Specifically, NextVir explores several foundation models-DNABERT-S, Nucelotide Transformer, and HyenaDNA-and efficiently fine-tunes them to enable accurate identification of the sequencing reads' origin. The results demonstrate superior performance of the proposed framework over existing deep learning methods and suggest downstream potential for foundational models in genomics.

Humans↗

Crystal structures of the lytic transglycosylase MltA from N.gonorrhoeae and E.coli: insights into interdomain movements and substrate binding.

MltA is a lytic transglycosylase of Gram-negative bacteria that cleaves the beta-1,4 glycosidic linkages between N-acetylmuramic acid (MurNAc) and N-acetylglucosamine (GlcNAc) in peptidoglycan. We have determined the crystal structures of MltA from Neisseria gonorrhoeae and Escherichia coli (NgMltA and EcMltA), which have only 21.5% sequence identity. Both proteins have two main domains separated by a deep groove. Domain 1 shows structural similarity with the so-called double-psi barrel family of proteins. Comparison of the two structures reveals substantial differences in the relative positions of domains 1 and 2 such that the active site groove in NgMltA is much wider and appears more able to accommodate peptidoglycan substrate than EcMltA, suggesting that domain closure occurs after substrate binding. Docking of a peptidoglycan molecule into the structure of NgMltA reveals a number of conserved residues that are likely involved in substrate binding, including a potential binding pocket for the peptidyl moieties. This structure supports the assignment of Asp405 as the acid catalyst responsible for cleavage of the glycosidic bond. In EcMltA, the equivalent residue is Asp328, which has been identified previously. The structures also suggest a catalytic role for Asp393 (Asp317 in EcMltA) in activating the C6 hydroxyl group during formation of the 1,6-anhydro linkage. Finally, in comparison to EcMltA, NgMltA contains a unique third domain that is an insertion within domain 2. The domain is beta in structure and may mediate protein-protein interactions that are specific to peptidoglycan metabolism in N.gonorrhoeae.

Amino Acid Sequence↗

Phylogeographic structure in the bogus yucca moth Prodoxus quinquepunctellus (Prodoxidae): comparisons with coexisting pollinator yucca moths.

The pollination mutualism between yucca moths and yuccas highlights the potential importance of host plant specificity in insect diversification. Historically, one pollinator moth species, Tegeticula yuccasella, was believed to pollinate most yuccas. Recent phylogenetic studies have revealed that it is a complex of at least 13 distinct species, eight of which are specific to one yucca species. Moths in the closely related genus Prodoxus also specialize on yuccas, but they do not pollinate and their larvae feed on different plant parts. Previous research demonstrated that the geographically widespread Prodoxus quinquepunctellus can rapidly specialize to its host plants and may harbor hidden species diversity. We examined the phylogeographic structure of P. quinquepunctellus across its range to compare patterns of diversification with six coexisting pollinator yucca moth species. Morphometric and mtDNA cytochrome oxidase I sequence data indicated that P. quinquepunctellus as currently described contains two species. There was a deep division between moth populations in the eastern and the western United States, with limited sympatry in central Texas; these clades are considered separate species and are redescribed as P. decipiens and P. quinquepunctellus (sensu stricto), respectively. Sequence data also showed a lesser division within P. quinquepunctellus s.s. between the western populations on the Colorado Plateau and those elsewhere. The divergence among the three emerging lineages corresponded with major biogeographic provinces, whereas AMOVA indicated that host plant specialization has been relatively unimportant in diversification. In comparison, the six pollinator species comprise three lineages, one eastern and two western. A pollinator species endemic to the Colorado Plateau has evolved in both of the western lineages. The east-west division and the separate evolution of two Colorado Plateau pollinator species suggest that similar biogeographic factors have influenced diversification in both Tegeticula and Prodoxus. For the pollinators, however, each lineage has produced a monophagous species, a pattern not seen in P. quinquepunctellus.

Analysis of Variance↗

Phylogeny of trichomonads inferred from small-subunit rRNA sequences.

Small subunit (16S-like) ribosomal RNA sequences were obtained from representatives of all four families constituting the order Trichomonadida. Comparative sequence analysis revealed that the Trichomonadida are a monophyletic lineage and a deep branch of the eukaryotic tree. Relative to the early divergent eukaryotic assemblages the branching pattern within the Trichomonadida is very shallow. This pattern suggests the Trichomonadida radiated recently, perhaps in conjunction with their animal hosts. From a morphological perspective the Devescovinidae and Calonymphidae are considered more derived than the Monocercomonadidae and Trichomonadidae. Molecular trees inferred by distance, parsimony and likelihood techniques consistently show the Devescovinidae and Calonymphidae are the earliest diverging lineages within the Trichomonadida, however bootstrap values do not strongly support a particular branching order. In an analysis of all known 16S-like ribosomal RNA sequences, the Trichomonadida share most recent common ancestry with unidentified protists from the hindgut of the termite Reticulitermes flavipes. The position of two putative free-living trichomonads in the tree is indicative of derivation from symbionts rather than direct descent from some free-living ancestral trichomonad.

Animals↗

Genetic heterogeneity of Borrelia burgdorferi sensu lato in the southern United States based on restriction fragment length polymorphism and sequence analysis.

Fifty-six strains of Borrelia burgdorferi sensu lato, isolated from ticks and vertebrate animals in Missouri, South Carolina, Georgia, Florida, and Texas, were identified and characterized by PCR-restriction fragment length polymorphism (RFLP) analysis of rrf (5S)-rrl (23S) intergenic spacer amplicons. A total of 241 to 258 bp of intergenic spacers between tandemly duplicated rrf (5S) and rrl (23S) was amplified by PCR. MseI and DraI restriction fragment polymorphisms were used to analyze these strains. PCR-RFLP analysis results indicated that the strains represented at least three genospecies and 10 different restriction patterns. Most of the strains isolated from the tick Ixodes dentatus in Missouri and Georgia belonged to the genospecies Borrelia andersonii. Excluding the I. dentatus strains, most southern strains, isolated from the ticks Ixodes scapularis and Ixodes affinis, the cotton rat (Sigmodon hispidus), and cotton mouse (Peromyscus gossypinus) in Georgia and Florida, belonged to Borrelia burgdorferi sensu stricto. Seven strains, isolated from Ixodes minor, the wood rat (Neotoma floridana), the cotton rat, and the cotton mouse in South Carolina and Florida, belonged to Borrelia bissettii. Two strains, MI-8 from Florida and TXW-1 from Texas, exhibited MseI and DraI restriction patterns different from those of previously reported genospecies. Eight Missouri tick strains (MOK-3a group) had MseI patterns similar to that of B. andersonii reference strain 21038 but had a DraI restriction site in the spacer. Strain SCGT-8a had DraI restriction patterns identical to that of strain 25015 (B. bissettii) but differed from strain 25015 in its MseI restriction pattern. Strain AI-1 had the same DraI pattern as other southern strains in the B. bissettii genospecies but had a distinct MseI profile. The taxonomic status of these atypical strains needs to be further evaluated. To clarify the taxonomic positions of these atypical Borrelia strains, the complete sequences of rrf-rrl intergenic spacers from 20 southeastern and Missouri strains were determined. The evolutionary and phylogenetic relationships of these strains were compared with those of the described genospecies in the B. burgdorferi sensu lato species complex. The 20 strains clustered into five separate lineages on the basis of sequence analysis. MI-8 and TXW-1 appeared to belong to two different undescribed genospecies, although TXW-1 was closely related to Borrelia garinii. The MOK-3a group separated into a distinct deep branch in the B. andersonii lineage. PCR-RFLP analysis results and the results of sequence analyses of the rrf-rrl intergenic spacer confirm that greater genetic heterogeneity exists among B. burgdorferi sensu lato strains isolated from the southern United States than among strains isolated from the northern United States. The B. andersonii genospecies and its MOK-3a subgroup are associated with the I. dentatus-cottontail rabbit enzootic cycle, but I. scapularis was also found to harbor a strain of this genospecies. Strains that appear to be B. bissettii in our study were isolated from I. minor and the cotton mouse, cotton rat, and wood rat. The B. burgdorferi sensu stricto strains from the south are genetically and phenotypically similar to the B31 reference strain.

Animals↗

In-field spatial variability in the degradation of the phenyl-urea herbicide isoproturon is the result of interactions between degradative Sphingomonas spp. and soil pH.

Substantial spatial variability in the degradation rate of the phenyl-urea herbicide isoproturon (IPU) [3-(4-isopropylphenyl)-1,1-dimethylurea] has been shown to occur within agricultural fields, with implications for the longevity of the compound in the soil, and its movement to ground- and surface water. The microbial mechanisms underlying such spatial variability in degradation rate were investigated at Deep Slade field in Warwickshire, United Kingdom. Most-probable-number analysis showed that rapid degradation of IPU was associated with proliferation of IPU-degrading organisms. Slow degradation of IPU was linked to either a delay in the proliferation of IPU-degrading organisms or apparent cometabolic degradation. Using enrichment techniques, an IPU-degrading bacterial culture (designated strain F35) was isolated from fast-degrading soil, and partial 16S rRNA sequencing placed it within the Sphingomonas group. Denaturing gradient gel electrophoresis (DGGE) of PCR-amplified bacterial community 16S rRNA revealed two bands that increased in intensity in soil during growth-linked metabolism of IPU, and sequencing of the excised bands showed high sequence homology to the Sphingomonas group. However, while F35 was not closely related to either DGGE band, one of the DGGE bands showed 100% partial 16S rRNA sequence homology to an IPU-degrading Sphingomonas sp. (strain SRS2) isolated from Deep Slade field in an earlier study. Experiments with strains SRS2 and F35 in soil and liquid culture showed that the isolates had a narrow pH optimum (7 to 7.5) for metabolism of IPU. The pH requirements of IPU-degrading strains of Sphingomonas spp. could largely account for the spatial variation of IPU degradation rates across the field.

Biodegradation, Environmental↗

Aromatic-degrading Sphingomonas isolates from the deep subsurface.

An obligately aerobic chemoheterotrophic bacterium (strain F199) previously isolated from Southeast Coastal Plain subsurface sediments and shown to degrade toluene, naphthalene, and other aromatic compounds (J. K. Fredrickson, F. J. Brockman, D. J. Workman, S. W. Li, and T. O. Stevens, Appl. Environ. Microbiol. 57:796-803, 1991) was characterized by analysis of its 16S rRNA nucleotide base sequence and cellular lipid composition. Strain F199 contained 2-OH14:0 and 18:1 omega 7c as the predominant cellular fatty acids and sphingolipids that are characteristic of the genus Sphingomonas. Phylogenetic analysis of its 16S rRNA sequence indicated that F199 was most closely related to Sphingomonas capsulata among the bacteria currently in the Ribosomal Database. Five additional isolates from deep Southeast Coastal Plain sediments were determined by 16S rRNA sequence analysis to be closely related to F199. These strains also contained characteristic sphingolipids. Four of these five strains could also grow on a broad range of aromatic compounds and could mineralize [14C]toluene and [14C]naphthalene. S. capsulata (ATCC 14666), Sphingomonas paucimobilis (ATCC 29837), and one of the subsurface isolates were unable to grow on any of the aromatic compounds or mineralize toluene or naphthalene. These results indicate that bacteria within the genus Sphingomonas are present in Southeast Coastal Plain subsurface sediments and that the capacity for degrading a broad range of substituted aromatic compounds appears to be common among Sphingomonas species from this environment.

Bacteria, Aerobic↗

Parallelism of amino acid changes at the RH1 affecting spectral sensitivity among deep-water cichlids from Lakes Tanganyika and Malawi.

Many examples of the appearance of similar traits in different lineages are known during the evolution of organisms. However, the underlying genetic mechanisms have been elucidated in very few cases. Here, we provide a clear example of evolutionary parallelism, involving changes in the same genetic pathway, providing functional adaptation of RH1 pigments to deep-water habitats during the adaptive radiation of East African cichlid fishes. We determined the RH1 sequences from 233 individual cichlids. The reconstruction of cichlid RH1 pigments with 11-cis-retinal from 28 sequences showed that the absorption spectra of the pigments of nine species were shifted toward blue, tuned by two particular amino acid replacements. These blue-shifted RH1 pigments might have evolved as adaptations to the deep-water photic environment. Phylogenetic evidence indicates that one of the replacements, A292S, has evolved several times independently, inducing similar functional change. The parallel evolution of the same mutation at the same amino acid position suggests that the number of genetic changes underlying the appearance of similar traits in cichlid diversification may be fewer than previously expected.

Adaptation, Physiological↗

Thermovibrio ammonificans sp. nov., a thermophilic, chemolithotrophic, nitrate-ammonifying bacterium from deep-sea hydrothermal vents.

A thermophilic, anaerobic, chemolithoautotrophic bacterium was isolated from the walls of an active deep-sea hydrothermal vent chimney on the East Pacific Rise at 9 degrees 50' N. Cells of the organism were Gram-negative, motile rods that were about 1.0 microm in length and 0.6 microm in width. Growth occurred between 60 and 80 degrees C (optimum at 75 degrees C), 0.5 and 4.5% (w/v) NaCl (optimum at 2%) and pH 5 and 7 (optimum at 5.5). Generation time under optimal conditions was 1.57 h. Growth occurred under chemolithoautotrophic conditions in the presence of H2 and CO2, with nitrate or sulfur as the electron acceptor and with concomitant formation of ammonium or hydrogen sulfide, respectively. Thiosulfate, sulfite and oxygen were not used as electron acceptors. Acetate, formate, lactate and yeast extract inhibited growth. No chemoorganoheterotrophic growth was observed on peptone, tryptone or Casamino acids. The genomic DNA G+C content was 54.6 mol%. Phylogenetic analyses of the 16S rRNA gene sequence indicated that the organism was a member of the domain Bacteria and formed a deep branch within the phylum Aquificae, with Thermovibrio ruber as its closest relative (94.4% sequence similarity). On the basis of phylogenetic, physiological and genetic considerations, it is proposed that the organism represents a novel species within the newly described genus Thermovibrio. The type strain is Thermovibrio ammonificans HB-1T (=DSM 15698T=JCM 12110T).

Ammonia↗

Mesozoic origin for West Indian insectivores.

The highly endangered solenodons, endemic to Cuba (Solenodon cubanus) and Hispaniola (S. paradoxus), comprise the only two surviving species of West Indian insectivores. Combined gene sequences (13.9 kilobases) from S. paradoxus established that solenodons diverged from other eulipotyphlan insectivores 76 million years ago in the Cretaceous period, which is consistent with vicariance, though also compatible with dispersal. A sequence of 1.6 kilobases of mitochondrial DNA from S. cubanus indicated a deep divergence of 25 million years versus the congeneric S. paradoxus, which is consistent with vicariant origins as tectonic forces separated Cuba and Hispaniola. Efforts to prevent extinction of the two surviving solenodon species would conserve an entire lineage as old or older than many mammalian orders.

Animals↗

Integrated approaches to terminal Proterozoic stratigraphy: an example from the Olenek Uplift, northeastern Siberia.

In the Olenek Uplift of northeastern Siberia, the Khorbusuonka Group and overlying Kessyusa and Erkeket formations preserve a significant record of terminal Proterozoic and basal Cambrian Earth history. A composite section more than 350 m thick is reconstructed from numerous exposures along the Khorbusuonka River. The Khorbusuonka Group comprises three principal sedimentary sequences: peritidal dolomites of the Mastakh Formation, which are bounded above and below by red beds; the Khatyspyt and most of the overlying Turkut formations, which shallow upward from relatively deep-water carbonaceous micrites to cross-bedded dolomitic grainstones and stromatolites; and a thin upper Turkut sequence bounded by karst surfaces. The overlying Kessyusa Formation is bounded above and below by erosional surfaces and contains additional parasequence boundaries internally. Ediacaran metazoans, simple trace fossils, and vendotaenids occur in the Khatyspyt Formation; small shelly fossils, more complex trace fossils, and acritarchs all appear near the base of the Kessyusa Formation and diversify upward. The carbon-isotopic composition of carbonates varies stratigraphically in a pattern comparable to that determined for other terminal Proterozoic and basal Cambrian successions. In concert, litho-, bio-, and chemostratigraphic data indicate the importance of the Khorbusuonka Group in the global correlation of terminal Proterozoic sedimentary rocks. Stratigraphic data and a recently determined radiometric date on basal Kessyusa volcanic breccias further underscore the significance of the Olenek region in investigations of the Proterozoic-cambrian boundary.

Animals↗

Microbial community structure in three deep-sea carbonate crusts.

Carbonate crusts in marine environments can act as sinks for carbon dioxide. Therefore, understanding carbonate crust formation could be important for understanding global warming. In the present study, the microbial communities of three carbonate crust samples from deep-sea mud volcanoes in the eastern Mediterranean were characterized by sequencing 16S ribosomal RNA (rRNA) genes amplified from DNA directly retrieved from the samples. In combination with the mineralogical composition of the crusts and lipid analyses, sequence data were used to assess the possible role of prokaryotes in crust formation. Collectively, the obtained data showed the presence of highly diverse communities, which were distinct in each of the carbonate crusts studied. Bacterial 16S rRNA gene sequences were found in all crusts and the majority was classified as alpha-, gamma-, and delta- Proteobacteria. Interestingly, sequences of Proteobacteria related to Halomonas and Halovibrio sp., which can play an active role in carbonate mineral formation, were present in all crusts. Archaeal 16S rRNA gene sequences were retrieved from two of the crusts studied. Several of those were closely related to archaeal sequences of organisms that have previously been linked to the anaerobic oxidation of methane (AOM). However, the majority of archaeal sequences were not related to sequences of organisms known to be involved in AOM. In combination with the strongly negative delta 13C values of archaeal lipids, these results open the possibility that organisms with a role in AOM may be more diverse within the Archaea than previously suggested. Different communities found in the crusts could carry out similar processes that might play a role in carbonate crust formation.

Anaerobiosis↗