Search PubMed⌕ Search

Biomedical subjects

Bao Tran

Publications and source records attributed to Bao Tran.

7 recordsLinked to original sources

Accurate somatic small variant discovery for multiple sequencing technologies with DeepSomatic.

Somatic variant detection is an integral part of cancer genomics analysis. While most methods have focused on short-read sequencing, long-read technologies offer potential advantages in repeat mapping and variant phasing. We present DeepSomatic, a deep-learning method for detecting somatic small nucleotide variations and insertions and deletions from both short-read and long-read data. The method has modes for whole-genome and whole-exome sequencing and can run on tumor-normal, tumor-only and formalin-fixed paraffin-embedded samples. To train DeepSomatic and help address the dearth of publicly available training and benchmarking data for somatic variant detection, we generated and make openly available the Cancer Standards Long-read Evaluation (CASTLE) dataset of six matched tumor-normal cell line pairs whole-genome sequenced with Illumina, PacBio HiFi and Oxford Nanopore Technologies, along with benchmark variant sets. Across samples, both cell line and patient-derived, and across short-read and long-read sequencing technologies, DeepSomatic consistently outperforms existing callers.

Humans↗

Long-read sequencing of single cell-derived melanoma subclones reveals divergent and parallel genomic and epigenomic evolutionary trajectories.

Tumor evolution is driven by various mutational processes, ranging from single-nucleotide variants (SNVs) to large structural variants (SVs) to dynamic shifts in DNA methylation. Current short-read sequencing methods struggle to accurately capture the full spectrum of these genomic and epigenomic alterations due to inherent technical limitations. To overcome that, here we introduce an approach for long-read sequencing of single-cell derived subclones, and use it to profile 23 subclones of a mouse melanoma cell line, characterized with distinct growth phenotypes and treatment responses. We develop a computational framework for harmonization and joint analysis of different variant types in the evolutionary context. Uniquely, our framework enables detection of recurrent amplifications of putative driver genes, generated by independent SVs across different lineages, suggesting parallel evolution. In addition, our approach revealed gradual and lineage-specific methylation changes associated with aggressive clonal phenotypes. We also show our set of phylogeny-constrained variant calls along with openly released sequencing data can be a valuable resource for the development of new computational methods.

Journal Article↗

Severus detects somatic structural variation and complex rearrangements in cancer genomes using long-read sequencing.

For the detection of somatic structural variation (SV) in cancer genomes, long-read sequencing is advantageous over short-read sequencing with respect to mappability and variant phasing. However, most current long-read SV detection methods are not developed for the analysis of tumor genomes characterized by complex rearrangements and heterogeneity. Here, we present Severus, a breakpoint graph-based algorithm for somatic SV calling from long-read cancer sequencing. Severus works with matching normal samples, supports unbalanced cancer karyotypes, can characterize complex multibreak SV patterns and produces haplotype-specific calls. On a comprehensive multitechnology cell line panel, Severus consistently outperforms other long-read and short-read methods in terms of SV detection F1 score (harmonic mean of the precision and recall). We also illustrate that compared to long-read methods, short-read sequencing systematically misses certain classes of somatic SVs, such as insertions or clustered rearrangements. We apply Severus to several clinical cases of pediatric leukemia/lymphoma, revealing clinically relevant cryptic rearrangements missed by standard genomic panels.

Humans↗

DeepSomatic: Accurate somatic small variant discovery for multiple sequencing technologies.

Somatic variant detection is an integral part of cancer genomics analysis. While most methods have focused on short-read sequencing, long-read technologies now offer potential advantages in terms of repeat mapping and variant phasing. We present DeepSomatic, a deep learning method for detecting somatic SNVs and insertions and deletions (indels) from both short-read and long-read data, with modes for whole-genome and exome sequencing, and able to run on tumor-normal, tumor-only, and with FFPE-prepared samples. To help address the dearth of publicly available training and benchmarking data for somatic variant detection, we generated and make openly available a dataset of five matched tumor-normal cell line pairs sequenced with Illumina, PacBio HiFi, and Oxford Nanopore Technologies, along with benchmark variant sets. Across samples and technologies (short-read and long-read), DeepSomatic consistently outperforms existing callers, particularly for indels.

Journal Article↗

Whole genome comparisons of serotype 4b and 1/2a strains of the food-borne pathogen Listeria monocytogenes reveal new insights into the core genome components of this species.

The genomes of three strains of Listeria monocytogenes that have been associated with food-borne illness in the USA were subjected to whole genome comparative analysis. A total of 51, 97 and 69 strain-specific genes were identified in L.monocytogenes strains F2365 (serotype 4b, cheese isolate), F6854 (serotype 1/2a, frankfurter isolate) and H7858 (serotype 4b, meat isolate), respectively. Eighty-three genes were restricted to serotype 1/2a and 51 to serotype 4b strains. These strain- and serotype-specific genes probably contribute to observed differences in pathogenicity, and the ability of the organisms to survive and grow in their respective environmental niches. The serotype 1/2a-specific genes include an operon that encodes the rhamnose biosynthetic pathway that is associated with teichoic acid biosynthesis, as well as operons for five glycosyl transferases and an adenine-specific DNA methyltransferase. A total of 8603 and 105 050 high quality single nucleotide polymorphisms (SNPs) were found on the draft genome sequences of strain H7858 and strain F6854, respectively, when compared with strain F2365. Whole genome comparative analyses revealed that the L.monocytogenes genomes are essentially syntenic, with the majority of genomic differences consisting of phage insertions, transposable elements and SNPs.

Base Composition↗

Increased expression of vasopressin v1a receptors after traumatic brain injury.

Experimental evidence obtained in various animal models of brain injury indicates that vasopressin promotes the formation of cerebral edema. However, the molecular and cellular mechanisms underlying this vasopressin action are not fully understood. In the present study, we analyzed the temporal changes in expression of vasopressin V1a receptors after traumatic brain injury (TBI) in rats. In the intact brain, the V1a receptor was expressed in neurons located in all layers of the frontoparietal cortex. The V1a receptor-immunoreactive product was predominantly localized to neuronal nuclei and had both a diffused and punctate staining pattern. The V1a receptors were also expressed in astrocytes, especially in layer 1 of the frontoparietal cortex. In these cells, two distinctive patterns of immunopositive staining for V1a receptors were observed: a diffused cytosolic staining of cell bodies and processes and a clearly punctate staining pattern that was predominantly localized to the astrocytic cell bodies. The real-time reverse-transcriptase polymerase chain reaction analysis of changes in mRNA for the V1a receptor demonstrated that after TBI, there is an early (4 h post-TBI) increase in the number of transcripts in the ipsilateral frontoparietal cortex, when compared to the contralateral hemisphere or the sham-injured rats. This increase in the message was followed by the up-regulation of expression of the V1a receptors at the protein level. This was most evident in cortical astrocytes in the areas surrounding the lesion. The number of the V1a receptor-immunopositive astrocytes in the traumatized parenchyma gradually increased, starting at 8 h and peaking at 4-6 days after TBI. Furthermore, a redistribution of V1a receptors from the astrocytic cell bodies to the astrocytic processes was observed. In addition to astrocytes, an increased expression of V1a receptors was found in the endothelium of both blood microvessels and the large-diameter blood vessels in the frontoparietal cortex ipsilateral to injury. This increase in the V1a receptor expression was apparent between 2 and 4 days after TBI. As early as 1-2 h following the impact, there was also a striking increase in the number of the V1a receptor-immunopositive beaded axonal processes, with greatly enlarged varicosities, that were localized to various areas of the injured parenchyma. It is suggested that the increased expression of V1a receptors plays an important role in the vasopressin-mediated formation of edema in the injured brain.

Animals↗

The complete genome sequence of the Arabidopsis and tomato pathogen Pseudomonas syringae pv. tomato DC3000.

We report the complete genome sequence of the model bacterial pathogen Pseudomonas syringae pathovar tomato DC3000 (DC3000), which is pathogenic on tomato and Arabidopsis thaliana. The DC3000 genome (6.5 megabases) contains a circular chromosome and two plasmids, which collectively encode 5,763 ORFs. We identified 298 established and putative virulence genes, including several clusters of genes encoding 31 confirmed and 19 predicted type III secretion system effector proteins. Many of the virulence genes were members of paralogous families and also were proximal to mobile elements, which collectively comprise 7% of the DC3000 genome. The bacterium possesses a large repertoire of transporters for the acquisition of nutrients, particularly sugars, as well as genes implicated in attachment to plant surfaces. Over 12% of the genes are dedicated to regulation, which may reflect the need for rapid adaptation to the diverse environments encountered during epiphytic growth and pathogenesis. Comparative analyses confirmed a high degree of similarity with two sequenced pseudomonads, Pseudomonas putida and Pseudomonas aeruginosa, yet revealed 1,159 genes unique to DC3000, of which 811 lack a known function.

Arabidopsis↗