Search PubMed⌕ Search

Biomedical subjects

Xiaotu Ma

Publications and source records attributed to Xiaotu Ma.

6 recordsLinked to original sources

The landscape of structural variation in pediatric cancer.

Structural variants (SVs) account for over 60% of the driver variants in pediatric cancer, and in many cases act as the cancer initiating event. To study SVs from a pan-cancer perspective, we analyzed 1,616 pediatric cancer genomes in 16 major cancer types of hematological malignancies (n = 908), brain tumors (n = 183), and solid tumors (n = 525) and compared their profiles to those of 2,203 adult cancers. The SV burden varied ~100-fold across pediatric cancer types and demonstrated an 8- to 16-fold reduction compared to adult brain and solid tumors but was comparable in pediatric versus adult hematological malignancies. Recurrent SV hotspots occurred uniquely in pediatric acute lymphoblastic leukemias (ALLs) in proximity to RAG-mediated recombination signal sequences (RSS) and disrupted multiple immune-related loci as well as 69 genes, which often involved cryptic RSS sites. By contrast, such hotspots affected only immune-related loci but not driver genes in adult lymphoid cancers. Eight SV signatures extracted from the cohort had varying distributions across cancer types, with clustered translocations reflecting templated insertions in osteosarcoma, and medium-sized deletions (10 kb to 1 Mb) enriched in cancers with RAG-mediated deletions. Intra-patient evolutionary analysis in 13 patients with multiple spatiotemporally distinct samples revealed that RAG-mediated recombination in leukemia and complex rearrangements in solid tumors occurred both early in disease initiation and continuously during later diversification, contributing to clonal heterogeneity. Finally, we found that both driver genes and fragile sites were the two genomic regions most frequently disrupted by SVs. The unique and diverse SV landscapes that emerged from this comprehensive analysis expand the scope of RSS-mediated mutagenesis in pediatric ALL and will be a valuable resource for guiding future functional studies and the design of clinical genomic testing in pediatric cancer.

Journal Article↗

SJPedPanel: A Pan-Cancer Gene Panel for Childhood Malignancies to Enhance Cancer Monitoring and Early Detection.

PURPOSE: The purpose of the study was to design a pan-cancer gene panel for childhood malignancies and validate it using clinically characterized patient samples. EXPERIMENTAL DESIGN: In addition to 5,275 coding exons, SJPedPanel also covers 297 introns for fusions/structural variations and 7,590 polymorphic sites for copy-number alterations. Capture uniformity and limit of detection are determined by targeted sequencing of cell lines using dilution experiment. We validate its coverage by in silico analysis of an established real-time clinical genomics (RTCG) cohort of 253 patients. We further validate its performance by targeted resequencing of 113 patient samples from the RTCG cohort. We demonstrate its power in analyzing low tumor burden specimens using morphologic remission and monitoring samples. RESULTS: Among the 485 pathogenic variants reported in RTCG cohort, SJPedPanel covered 86% of variants, including 82% of 90 rearrangements responsible for fusion oncoproteins. In our targeted resequencing cohort, 91% of 389 pathogenic variants are detected. The gene panel enabled us to detect ∼95% of variants at allele fraction (AF) 0.5%, whereas the detection rate is ∼80% at AF 0.2%. The panel detected low-frequency driver alterations from morphologic leukemia remission samples and relapse-enriched alterations from monitoring samples, demonstrating its power for cancer monitoring and early detection. CONCLUSIONS: SJPedPanel enables the cost-effective detection of clinically relevant genetic alterations including rearrangements responsible for subtype-defining fusions by targeted sequencing of ∼0.15% of human genome for childhood malignancies. It will enhance the analysis of specimens with low tumor burdens for cancer monitoring and early detection.

Humans↗

CGI: a new approach for prioritizing genes by combining gene expression and protein-protein interaction data.

MOTIVATION: Identifying candidate genes associated with a given phenotype or trait is an important problem in biological and biomedical studies. Prioritizing genes based on the accumulated information from several data sources is of fundamental importance. Several integrative methods have been developed when a set of candidate genes for the phenotype is available. However, how to prioritize genes for phenotypes when no candidates are available is still a challenging problem. RESULTS: We develop a new method for prioritizing genes associated with a phenotype by Combining Gene expression and protein Interaction data (CGI). The method is applied to yeast gene expression data sets in combination with protein interaction data sets of varying reliability. We found that our method outperforms the intuitive prioritizing method of using either gene expression data or protein interaction data only and a recent gene ranking algorithm GeneRank. We then apply our method to prioritize genes for Alzheimer's disease. AVAILABILITY: The code in this paper is available upon request.

Algorithms↗

MARD: a new method to detect differential gene expression in treatment-control time courses.

MOTIVATION: Characterizing the dynamic regulation of gene expression by time course experiments is becoming more and more important. A common problem is to identify differentially expressed genes between the treatment and control time course. It is often difficult to compare expression patterns of a gene between two time courses for the following reasons: (1) the number of sampling time points may be different or hard to be aligned between the treatment and the control time courses; (2) estimation of the function that describes the expression of a gene in a time course is difficult and error-prone due to the limited number of time points. We propose a novel method to identify the differentially expressed genes between two time courses, which avoids direct comparison of gene expression patterns between the two time courses. RESULTS: Instead of attempting to 'align' and compare the two time courses directly, we first convert the treatment and control time courses into neighborhood systems that reflect the underlying relationships between genes. We then identify the differentially expressed genes by comparing the two gene relationship networks. To verify our method, we apply it to two treatment-control time course datasets. The results are consistent with the previous results and also give some new biologically meaningful findings. AVAILABILITY: The algorithm in this paper is coded in C++ and is available from http://leili-lab.cmb.usc.edu/yeastaging/projects/MARD/

Algorithms↗

Last intron of the chemokine-like factor gene contains a putative promoter for the downstream CKLF super family member 1 gene.

The genes for chemokine-like factor (CKLF) and four chemokine-like factor super family members (CKLFSF1-4) are tightly linked on chromosome 16, with only 325 bp separating CKLF and CKLFSF1. We used Northern blotting and RT-PCR to show that these two genes are expressed independently of one another. We then used a novel computational promoter prediction method based on the interaction among transcription factor binding sites (TFBSs) to identify a putative promoter region for the CKLFSF1 gene. Our method predicted a promoter region in the last intron of the upstream gene, CKLF. We PCR amplified the predicted promoter region and used a luciferase assay to show that the region was able to drive the luciferase gene. DNA decoy experiments indicated that 214 bp fragment neighboring the TATA box markedly inhibited CKLFSF1 gene expression. Sequence analysis of the region revealed a typical TATA box (TATATAA) and multiple potential transcription factor binding sites, providing further evidence for this being a functional promoter for CKLFSF1. This work provides the first evidence of a promoter from one gene located in an intron of another.

Base Sequence↗