Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Deep sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Sphingomonas alaskensis sp. nov., a dominant bacterium from a marine oligotrophic environment.

Seven Gram-negative strains, isolated in 1990 from a 10(6)-fold dilution series of seawater from Resurrection Bay, a deep fjord of the Gulf of Alaska, were identified in a polyphasic taxonomic study. Analysis of 16S rDNA sequences and DNA-homology studies confirmed the phylogenetic position of all strains in the genus Sphingomonas and further indicated that all of the strains constitute a single homogeneous genomic species, distinct from all validly described Sphingomonas species. The ability to differentiate the species, both phenotypically and chemotaxonomically, from its nearest neighbours justifies the proposal of a new species name, Sphingomonas alaskensis sp. nov., for this taxon. Strain LMG 18877T (= RB2256T = DSM 13593T) was selected as the type strain.

Bacterial Typing Techniques↗

OpenSpliceAI: An efficient, modular implementation of SpliceAI enabling easy retraining on non-human species.

The SpliceAI deep learning system is currently one of the most accurate methods for identifying splicing signals directly from DNA sequences. However, its utility is limited by its reliance on older software frameworks and human-centric training data. Here we introduce OpenSpliceAI, a trainable, open-source version of SpliceAI implemented in PyTorch to address these challenges. OpenSpliceAI supports both training from scratch and transfer learning, enabling seamless retraining on species-specific datasets and mitigating human-centric biases. Our experiments show that it achieves faster processing speeds and lower memory usage than the original SpliceAI code, allowing large-scale analyses of extensive genomic regions on a single GPU. Additionally, OpenSpliceAI's flexible architecture makes for easier integration with established machine learning ecosystems, simplifying the development of custom splicing models for different species and applications. We demonstrate that OpenSpliceAI's output is highly concordant with SpliceAI. In silico mutagenesis (ISM) analyses confirm that both models rely on similar sequence features, and calibration experiments demonstrate similar score probability estimates.

Journal Article↗

Goblet cell-specific expression mediated by the MUC2 mucin gene promoter in the intestine of transgenic mice.

The regulation of MUC2, a major goblet cell mucin gene, was examined by constructing transgenic mice containing bases -2864 to +17 of the human MUC2 5'-flanking region fused into the 5'-untranslated region of a human growth hormone (hGH) reporter gene. Four of eight transgenic lines expressed reporter. hGH message expression was highest in the distal small intestine, with only one line expressing comparable levels in the colon. This contrasts with endogenous MUC2 expression, which is expressed at its highest levels in the colon. Immunohistochemical analysis indicated that goblet cell-specific expression of reporter begins deep in the crypts, as does endogenous MUC2 gene expression. These results indicate that the MUC2 5'-flanking sequence contains elements sufficient for the appropriate expression of MUC2 in small intestinal goblet cells. Conversely, elements located outside this region appear necessary for efficient colonic expression, implying that the two tissues utilize different regulatory elements. Thus many, but not all, of the elements necessary for MUC2 gene regulation reside between bases -2864 and +17 of the 5'-flanking region.

Animals↗

Bioinformatics for whole-genome shotgun sequencing of microbial communities.

The application of whole-genome shotgun sequencing to microbial communities represents a major development in metagenomics, the study of uncultured microbes via the tools of modern genomic analysis. In the past year, whole-genome shotgun sequencing projects of prokaryotic communities from an acid mine biofilm, the Sargasso Sea, Minnesota farm soil, three deep-sea whale falls, and deep-sea sediments have been reported, adding to previously published work on viral communities from marine and fecal samples. The interpretation of this new kind of data poses a wide variety of exciting and difficult bioinformatics problems. The aim of this review is to introduce the bioinformatics community to this emerging field by surveying existing techniques and promising new approaches for several of the most interesting of these computational problems.

Journal Article↗

Interlaboratory comparison of radioimmunological calcitonin determination.

An interlaboratory study for the radioimmunological determination of calcitonin was performed within the European PTH Study Group (ESPG), to improve comparability using external quality control. Twelve laboratories determined calcitonin in 21 deep frozen samples using their respective calcitonin radioimmunoassay systems. The samples included a standard curves of human calcitonin (sequence 1-32) in serum as well as in assay buffer, several dilutions of a serum from a patient with medullary thyroid carcinoma, and additional sera containing calcitonin levels of clinical importance. For evaluation, the known concentrations were related to the measured values of each laboratory, and of all laboratories together. Except for one, the laboratories recognized the different dilutions used, although absolute values were scattered over a wide range. Satisfactory agreement was reached only when the data were calculated as the percentage of a given value. In future, a serum with a well defined and constant calcitonin concentration should be used as an international standard in all determination of calcitonin by radioimmunoassay.

Calcitonin↗

Integrating genomic distance analyses in the description of a new family, genus, and species of sponge-associated antipatharians (black corals).

Antipatharians (black corals) are among the least studied coral groups, with much of their diversity still undescribed. Here, we present an integrative morphological, phylogenomic and genomic distance study of deep-sea antipatharians sampled in high seas areas of the North Pacific Ocean and from New Zealand's Exclusive Economic Zone. These corals grow on hexactinellid sponges - a unique characteristic in the order Antipatharia. Using a dataset of ultra-conserved elements and exons, combined with morphological analyses, we reconstruct phylogenomic relationships and formally describe a new family (Eidikopathidae fam. nov.), a new genus (Eidikopathesgen. nov.), and two new species (E. korallispongiasp. nov., E. zealandkoralliasp. nov.). Morphologically, the new family is distinguished by a corallum consisting of a network of loose branches that fuse with the sponge skeletal framework. Phylogenomic analyses recovered consistent topologies with strong nodal support, corroborating the distinct evolutionary placement of this sponge-associated lineage. Pairwise genomic distances estimated using the Tamura-Nei model were concordant with patristic genomic distances, identifying Pteridopathidae as the genetically closest family to Eidikopathidae fam. nov., followed by Myriopathidae and Stylopathidae, which were recovered as sister families in the phylogeny. This pattern shows that genomic distance complements, rather than simply mirrors, tree topology by quantifying accumulated sequence divergence among lineages. Together, these results provide the first genomic distance framework for Antipatharia, offering a baseline for future systematic, evolutionary, and biodiversity studies on this fundamental shallow, mesophotic and deep-sea coral group.

Animals↗

Is it possible to construct phylogenetic trees using polypeptide hormone sequences?

Because of the high degree of primary sequence conservation in neuropeptides and low molecular weight polypeptide hormones, these polypeptides are not useful for constructing phylogenetic trees based on maximum parsimony analysis. This review focuses on the organization of neuropeptide and polypeptide hormone precursors and discusses strategies for aligning polypeptide precursors for phylogenetic analysis. Examples are provided to support the hypothesis that some neuropeptide and polypeptide precursors and some high molecular weight polypeptide hormones can be used as data sets for resolving deep divergences among vertebrate taxa.

Amino Acid Sequence↗

Genetic analysis of housekeeping genes reveals a deep-sea ecotype of Alteromonas macleodii in the Mediterranean Sea.

The genetic diversity of 19 strains belonging to Alteromonas macleodii isolated from different geographic areas (Pacific and Indian Ocean, and different parts of the Mediterranean Sea) and at different depths (from the surface down to 3500 m) has been studied. Fragments of the 16S rRNA gene, the internal transcribed spacer (ITS) between 16S and 23S rDNA genes, the gyrB and the rpoB genes, have been sequenced for each strain. Amplified fragment length polymorphisms were used to characterize similarity at the level of the whole genome. Most of the diversity reflected the existence of a cluster of strains isolated from deep Mediterranean waters and two isolates from the Black Sea. Particularly the isolates from the deep sites were consistently different from all the others indicating the existence of a specific ecotype adapted to these conditions. Amplification of gyrB gene and ITS directly from DNA retrieved from deep Mediterreanean waters and one Atlantic sample showed that presence of this deep-sea ecotype is widespread and is not a product of culture bias. On the other hand, strains isolated from surface tropical waters showed a remarkable level of resemblance to the first isolate of this species obtained from Hawaii in 1972. The results indicate the existence of both lineages of global distribution and ecotypes adapted to specific conditions such as deep or more diluted (the Black Sea) waters.

Alteromonas↗

A phylogenetic analysis of Aquifex pyrophilus.

The 16S rRNA of the bacterion Aquifex pyrophilus, a microaerophilic, oxygen-reducing hyperthermophile, has been sequenced directly from the the PCR amplified gene. Phylogenetic analyses show the Aq. pyrophilus lineage to be probably the deepest (earliest) in the (eu)bacterial tree. The addition of this deep branching to the bacterial tree further supports the argument that the Bacteria are of thermophilic ancestry.

Archaea↗

Differential binding of tropane-based photoaffinity ligands on the dopamine transporter.

Benztropine and its analogs are tropane ring-containing dopamine uptake inhibitors that produce behavioral effects markedly different from cocaine and other dopamine transporter blockers. We investigated the benztropine binding site on dopamine transporters by covalently attaching a benztropine-based photoaffinity ligand, [125I]N-[n-butyl-4-(4"'-azido-3"'-iodophenyl)]-4', 4"-difluoro-3alpha-(diphenylmethoxy)tropane ([125I]GA II 34), to the protein, followed by proteolytic and immunological peptide mapping. The maps were compared with those obtained for dopamine transporters photoaffinity labeled with a GBR 12935 analog, [125I]1-[2-(diphenylmethoxy)ethyl]-4-[2-(4-azido-3-iodophenyl)ethy l]p iperazine ([125I]DEEP), and a cocaine analog, [125I]3beta-(p-chlorophenyl)tropane-2beta-carboxylic acid, 4'-azido-3'-iodophenylethyl ester ([125I]RTI 82), which have been shown previously to interact with different regions of the primary sequence of the protein. [125I]GA II 34 became incorporated in a membrane-bound, 14 kDa fragment predicted to contain transmembrane domains 1 and 2. This is the same region of the protein that binds [125I]DEEP, whereas the binding site for [125I]RTI 82 occurs closer to the C terminal in a domain containing transmembrane helices 4-7. Thus, although benztropine and cocaine both contain tropane rings, their binding sites are distinct, suggesting that dopamine transport inhibition may occur by different mechanisms. These results support previously derived structure-activity relationships suggesting that benztropine and cocaine analogs bind to different domains on the dopamine transporter. These differing molecular interactions may lead to the distinctive behavioral profiles of these compounds in animal models of drug abuse and indicate promise for the development of benztropine-based molecules for cocaine substitution therapies.

Azides↗

Leveraging protein language models for cross-variant CRISPR/Cas9 sgRNA activity prediction.

MOTIVATION: Accurate prediction of single-guide RNA (sgRNA) activity is crucial for optimizing the CRISPR/Cas9 gene-editing system, as it directly influences the efficiency and accuracy of genome modifications. However, existing prediction methods mainly rely on large-scale experimental data of a single Cas9 variant to construct Cas9 protein (variants)-specific sgRNA activity prediction models, which limits their generalization ability and prediction performance across different Cas9 protein (variants), as well as their scalability to the continuously discovered new variants. RESULTS: In this study, we proposed PLM-CRISPR, a novel deep learning-based model that leverages protein language models to capture Cas9 protein (variants) representations for cross-variant sgRNA activity prediction. PLM-CRISPR uses tailored feature extraction modules for both sgRNA and protein sequences, incorporating a cross-variant training strategy and a dynamic feature fusion mechanism to effectively model their interactions. Extensive experiments demonstrate that PLM-CRISPR outperforms existing methods across datasets spanning seven Cas9 protein (variants) in three real-world scenarios, demonstrating its superior performance in handling data-scarce situations, including cases with few or no samples for novel variants. Comparative analyses with traditional machine learning and deep learning models further confirm the effectiveness of PLM-CRISPR. Additionally, motif analysis reveals that PLM-CRISPR accurately identifies high-activity sgRNA sequence patterns across diverse Cas9 protein (variants). Overall, PLM-CRISPR provides a robust, scalable, and generalizable solution for sgRNA activity prediction across diverse Cas9 protein (variants). AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/CSUBioGroup/PLM-CRISPR.

CRISPR-Cas Systems↗

Trends in protein evolution inferred from sequence and structure analysis.

Complementary developments in comparative genomics, protein structure determination and in-depth comparison of protein sequences and structures have provided a better understanding of the prevailing trends in the emergence and diversification of protein domains. The investigation of deep relationships among different classes of proteins involved in key cellular functions, such as nucleic acid polymerases and other nucleotide-dependent enzymes, indicates that a substantial set of diverse protein domains evolved within the primordial, ribozyme-dominated RNA world.

Evolution, Molecular↗

Susceptibility patterns and molecular identification of Trichosporon species.

The physiological patterns, the sequence polymorphisms of the internal transcriber spacer (ITS), and intergenic spacer regions (IGS) of the rRNA genes and the antifungal susceptibility profile were evaluated for their ability to identify Trichosporon spp. and their specificity for the identification of 49 clinical isolates of Trichosporon spp. Morphological and biochemical methodologies were unable to differentiate among the Trichosporon species. ITS sequencing was also unable to differentiate several species. However, IGS1 sequencing unambiguously identified all Trichosporon isolates. Following the results of DNA-based identification, Trichosporon asahii was the species most frequently isolated from deep sites (15 of 25 strains; 60%). In the main, other Trichosporon species were recovered from cutaneous samples. The majority of T. asahii, T. faecale, and T. coremiiforme clinical isolates exhibited resistance in vitro to amphotericin B, with geometric mean (GM) MICs >4 mug/ml. The other species of Trichosporon did not show high MICs of amphotericin B, and GM MICs were <1 mug/ml. Azole agents were active in vitro against the majority of clinical strains. The most potent compound in vitro was voriconazole, with a GM MIC </=0.14 mug/ml. The sequencing of IGS correctly identified Trichosporon isolates; however, this technique is not available in many clinical laboratories, and strains should be dispatched to reference centers where these complex methods are available. Therefore, it seems to be more practical to perform antifungal susceptibility testing of all isolates belonging to Trichosporon spp., since correct identification could take several weeks, delaying the indication of an antifungal agent which exhibits activity against the infectious strain.

Amphotericin B↗

Characterization and expression of genes from the RubisCO gene cluster of the chemoautotrophic symbiont of Solemya velum: cbbLSQO.

Chemoautotrophic endosymbionts residing in Solemya velum gills provide this shallow water clam with most of its nutritional requirements. The cbb gene cluster of the S. velum symbiont, including cbbL and cbbS, which encode the large and small subunits of the carbon-fixing enzyme ribulose 1,5-bisphosphate carboxylase/oxygenase (RubisCO), was cloned and expressed in Escherichia coli. The recombinant RubisCO had a high specific activity, approximately 3 micromol min(-1) mg protein (-1), and a KCO2 of 40.3 microM. Based on sequence identity and phylogenetic analyses, these genes encode a form IA RubisCO, both subunits of which are closely related to those of the symbiont of the deep-sea hydrothermal vent gastropod Alviniconcha hessleri and the photosynthetic bacterium Allochromatium vinosum. In the cbb gene cluster of the S. velum symbiont, the cbbLS genes were followed by cbbQ and cbbO, which are found in some but not all cbb gene clusters and whose products are implicated in enhancing RubisCO activity post-translationally. cbbQ shares sequence similarity with nirQ and norQ, found in denitrification clusters of Pseudomonas stutzeri and Paracoccus denitrificans. The 3' region of cbbO from the S. velum symbiont, like that of the three other known cbbO genes, shares similarity to the 3' region of norD in the denitrification cluster. This is the first study to explore the cbb gene structure for a chemoautotrophic endosymbiont, which is critical both as an initial step in evaluating cbb operon structure in chemoautotrophic endosymbionts and in understanding the patterns and forces governing RubisCO evolution and physiology.

Animals↗

Linker chain L1 of earthworm hemoglobin. Structure of gene and protein: homology with low density lipoprotein receptor.

The extracellular hemoglobins (Hbs) of annelids and tube worms are giant multisubunit proteins of up to approximately 200 polypeptides and molecular masses to at least 3,900 kDa. They differ from all other Hbs in having both O2-binding chains and "linker" chains. The latter are required for assembly and structural integrity of the protein and are deficient in or lack heme. We have determined the nucleotide sequences of the cDNA and gene for linker chain L1 of the hemoglobin of Lumbricus terrestris. The cDNA-derived amino acid sequence has 225 residues and a calculated molecular mass of 25,847 Da. The chain is 21-28% identical to linker chains of the related annelid Tylorrhynchus heterochaetus and the deep-sea tube worm Lamellibrachia sp. A remarkable feature of the linker chains is a conserved 38-39-residue segment that contains a repeating pattern of cysteinyl residues: (Cys-X6)3-Cys-X5-Cys-X10-Cys. This pattern, not present in any globin sequence, corresponds exactly to the cysteine-rich repeats of the ligand binding domains of the low density lipoprotein (LDL) receptors of man and Xenopus laevis. Furthermore, the cysteine-rich segment of linker chain L1 has the sequence Asp-Gly-Ser-Asp-Glu which is characteristic of LDL receptor repeats. Similar cysteine-rich sequences also occur in two other mammalian proteins, complement C9 and renal glycoprotein GP330. The results support the conclusion that the cysteine-rich motif of the LDL receptor and annelid Hbs is a multipurpose protein-binding unit of ancient origin which has been incorporated into diverse unrelated proteins, presumably by the process of exon shuffling.

Amino Acid Sequence↗

UFold-X: an enhanced Dual & Dynamic U-Mamba model for long-range RNA secondary structure prediction.

RNA secondary structure is essential for understanding the functions of non-coding RNAs, ribosomal RNAs, and viral genomes. However, accurate prediction of long RNA structures remains challenging due to complex long-range interactions and the limited availability of long-RNA training data. We present UFold-X, a dual-branch deep learning framework that combines a convolutional encoder for local structure modeling with a Mamba-based Visual State Space Module for capturing long-range dependencies. A dynamic gating mechanism adaptively integrates the two branches according to sequence length. UFold-X was evaluated on multiple benchmark datasets containing RNAs up to 5000 nucleotides. To rigorously assess generalization, we introduced a cross-clan benchmark for long RNAs. Under this stringent setting, UFold-X achieved performance comparable to state-of-the-art classical approaches while achieving the best performance among deep learning-based methods. Additional cross-family and within-family evaluations further demonstrated robust transferability and competitive predictive performance. UFold-X also maintained excellent computational efficiency, requiring only 0.08 s per sequence on average. To assess biological consistency, we developed a SHAPE-based reactivity prediction variant (UFold-X-R) and an integrated metric, the Hybrid Reactivity-Pairing Score (HRPS). UFold-X-R showed strong agreement with experimental icSHAPE data and achieved the highest HRPS among all evaluated methods. A user-friendly web server is available at https://ufold-x.ai4bread.com.

Nucleic Acid Conformation↗

Sequencing and modeling of anti-DNA immunoglobulin Fv domains. Comparison with crystal structures.

Models for the three-dimensional structures of the combining regions of six DNA-binding antibodies have been derived from the sequence data for their Fv domains presented here. Using the amino acid sequences and the canonical structure classes described by Chothia and Lesk (Chothia, C., and Lesk, A.M. (1987) J. Mol. Biol. 196, 901), model loops were selected from immunoglobulin domains of known structure for five of the six antibody hypervariable regions. Models for the third complementarity-determining region of the heavy chain were constructed from known immunoglobulin loops of similar length and sequence. Comparison of three of the models with the respective crystal structure indicates that this procedure can generate a working model of the antibody combining region that provides useful information on the nature of the interactions between antibodies and nucleic acids. As part of our continuing investigation into the structural basis of antibody-DNA recognition, the observed and predicted models for the combining regions of nucleic acid-binding antibodies have been examined. In general, single strand-specific antibodies have deep clefts where the antigen might bind, whereas duplex-specific antibodies present a relatively flat surface. In addition, on the basis of both sequence and structure, there is little to distinguish autoimmune antibodies from those produced by immunization. Testable hypotheses for how these antibodies might interact with single- and double-stranded nucleic acids are presented.

Amino Acid Sequence↗

Composition of archaeal, bacterial, and eukaryal RuBisCO genotypes in three Western Pacific arc hydrothermal vent systems.

We studied the diversity of all forms of the RuBisCO large subunit-encoding gene cbbL in three RuBisCO uncharacterized hydrothermal vent communities. This diversity included the archaeal cbbL and the forms IC and ID, which have not previously been studied in the deep-sea environment, in addition to the forms IA, IB and II. Vent plume sites were Fryer and Pika in the Mariana arc and the Suiyo Seamount, Izu-Bonin, Japan. The cbbL forms were PCR amplified from plume bulk microbial DNA and then cloned and sequenced. Archaeal cbbL was detected in the Mariana samples only. Both forms IA and II were amplified from all samples, while the form IC was amplified only from the Pika and Suiyo samples. Only the Suiyo sample showed amplification of the form ID. The form IB was not recorded in any sample. Based on rarefaction analysis, nucleotide diversity and average pairwise difference, the archaeal cbbL was the most diverse form in Mariana samples, while the bacterial form IA was the most diverse form in the Suiyo sample. Also, the Pika sample harbored the highest diversity of cbbL phylogenetic lineages. Based on pairwise reciprocal library comparisons, the Fryer and Pika archaeal cbbL libraries showed the most significant difference, while Pika and Suiyo showed the highest similarity for forms IA and II libraries. This suggested that the Fryer supported the most divergent sequences. All archaeal cbbL sequences formed unique phylogenetic lineages within the branches of anaerobic thermophilic archaea of the genera Pyrococcus, Archaeoglobus, and Methanococcus. The other cbbL forms formed novel phylogenetic clusters distinct from any recorded previously in other deep-sea habitats. This is the first evidence for the diversity of archaeal cbbL in environmental samples.

Archaea↗