Search PubMed⌕ Search

Biomedical subjects

Yuan Qi

Publications and source records attributed to Yuan Qi.

10 recordsLinked to original sources

An Annotated Biobank of Triple-Negative Breast Cancer Patient-Derived Xenografts Features Treatment-Naïve and Longitudinal Samples during Neoadjuvant Chemotherapy.

UNLABELLED: Triple-negative breast cancer (TNBC) that fails to respond to neoadjuvant chemotherapy (NACT) can be lethal. Developing effective strategies to eradicate chemoresistant disease requires experimental models that recapitulate the heterogeneity characteristic of TNBC. To that end, we established a biobank of 92 orthotopic patient-derived xenograft (PDX) models of TNBC from the tumors of 75 patients enrolled in A Robust TNBC Evaluation fraMework to Improve Survival clinical trial (ARTEMIS, NCT02276443), including 12 longitudinal sets generated from serial patient biopsies collected throughout NACT treatment and from metastatic disease. Models were established from both chemosensitive and chemoresistant tumors, and nearly 30% of the PDX models were capable of metastasizing to the lungs. Comprehensive molecular profiling demonstrated conservation of genomes and transcriptomes between patient and corresponding PDX tumors, with representation of all major transcriptional subtypes. Transcriptional changes observed in the longitudinal PDX models highlighted dysregulation in pathways associated with DNA integrity, extracellular matrix interactions, the ubiquitin-proteasome system, epigenetics, and inflammatory signaling. These alterations revealed a complex network of adaptations associated with chemoresistance. Overall, this PDX biobank provides a valuable tool for tackling the most pressing issues facing the clinical management of TNBC. SIGNIFICANCE: The development of a patient-derived xenograft biobank that comprehensively captures the genomic and transcriptional diversity of triple-negative breast cancer promises to be a robust resource to investigate and overcome chemoresistance and metastasis.

Animals↗

Semi-supervised analysis of gene expression profiles for lineage-specific development in the Caenorhabditis elegans embryo.

MOTIVATION: Gene expression profiling is a powerful approach to identify genes that may be involved in a specific biological process on a global scale. For example, gene expression profiling of mutant animals that lack or contain an excess of certain cell types is a common way to identify genes that are important for the development and maintenance of given cell types. However, it is difficult for traditional computational methods, including unsupervised and supervised learning methods, to detect relevant genes from a large collection of expression profiles with high sensitivity and specificity. Unsupervised methods group similar gene expressions together while ignoring important prior biological knowledge. Supervised methods utilize training data from prior biological knowledge to classify gene expression. However, for many biological problems, little prior knowledge is available, which limits the prediction performance of most supervised methods. RESULTS: We present a Bayesian semi-supervised learning method, called BGEN, that improves upon supervised and unsupervised methods by both capturing relevant expression profiles and using prior biological knowledge from literature and experimental validation. Unlike currently available semi-supervised learning methods, this new method trains a kernel classifier based on labeled and unlabeled gene expression examples. The semi-supervised trained classifier can then be used to efficiently classify the remaining genes in the dataset. Moreover, we model the confidence of microarray probes and probabilistically combine multiple probe predictions into gene predictions. We apply BGEN to identify genes involved in the development of a specific cell lineage in the C. elegans embryo, and to further identify the tissues in which these genes are enriched. Compared to K-means clustering and SVM classification, BGEN achieves higher sensitivity and specificity. We confirm certain predictions by biological experiments. AVAILABILITY: The results are available at http://www.csail.mit.edu/~alanqi/projects/BGEN.html.

Algorithms↗

High-resolution computational models of genome binding events.

Direct physical information that describes where transcription factors, nucleosomes, modified histones, RNA polymerase II and other key proteins interact with the genome provides an invaluable mechanistic foundation for understanding complex programs of gene regulation. We present a method, joint binding deconvolution (JBD), which uses additional easily obtainable experimental data about chromatin immunoprecipitation (ChIP) to improve the spatial resolution of the transcription factor binding locations inferred from ChIP followed by DNA microarray hybridization (ChIP-Chip) data. Based on this probabilistic model of binding data, we further pursue improved spatial resolution by using sequence information. We produce positional priors that link ChIP-Chip data to sequence data by guiding motif discovery to inferred protein-DNA binding sites. We present results on the yeast transcription factors Gcn4 and Mig2 to demonstrate JBD's spatial resolution capabilities and show that positional priors allow computational discovery of the Mig2 motif when a standard approach fails.

Base Sequence↗

Structural classification of thioredoxin-like fold proteins.

Protein structure classification is necessary to comprehend the rapidly growing structural data for better understanding of protein evolution and sequence-structure-function relationships. Thioredoxins are important proteins that ubiquitously regulate cellular redox status and various other crucial functions. We define the thioredoxin-like fold using the structure consensus of thioredoxin homologs and consider all circular permutations of the fold. The search for thioredoxin-like fold proteins in the PDB database identified 723 protein domains. These domains are grouped into eleven evolutionary families based on combined sequence, structural, and functional evidence. Analysis of the protein-ligand structure complexes reveals two major active site locations for the thioredoxin-like proteins. Comparison to existing structure classifications reveals that our thioredoxin-like fold group is broader and more inclusive, unifying proteins from five SCOP folds, five CATH topologies and seven DALI domain dictionary globular folding topologies. Considering these structurally similar domains together sheds new light on the relationships between sequence, structure, function and evolution of thioredoxins.

Amino Acid Motifs↗

4SCOPmap: automated assignment of protein structures to evolutionary superfamilies.

BACKGROUND: Inference of remote homology between proteins is very challenging and remains a prerogative of an expert. Thus a significant drawback to the use of evolutionary-based protein structure classifications is the difficulty in assigning new proteins to unique positions in the classification scheme with automatic methods. To address this issue, we have developed an algorithm to map protein domains to an existing structural classification scheme and have applied it to the SCOP database. RESULTS: The general strategy employed by this algorithm is to combine the results of several existing sequence and structure comparison tools applied to a query protein of known structure in order to find the homologs already classified in SCOP database and thus determine classification assignments. The algorithm is able to map domains within newly solved structures to the appropriate SCOP superfamily level with approximately 95% accuracy. Examples of correctly mapped remote homologs are discussed. The algorithm is also capable of identifying potential evolutionary relationships not specified in the SCOP database, thus helping to make it better. The strategy of the mapping algorithm is not limited to SCOP and can be applied to any other evolutionary-based classification scheme as well. SCOPmap is available for download. CONCLUSION: The SCOPmap program is useful for assigning domains in newly solved structures to appropriate superfamilies and for identifying evolutionary links between different superfamilies.

Algorithms↗

PCOAT: positional correlation analysis using multiple methods.

UNLABELLED: PCOAT (Positional COrrelation Analysis Tool) is a program to perform positional correlation analysis for protein multiple sequence alignment in order to identify structurally or functionally important interactions between positions in a protein family. We implement different statistical methods to detect highly correlated position pairs, amino acid pairs, individual positions and networks of correlated positions, and utilize multiple sequence weighting and sampling methods to eliminate background correlations caused by phylogeny and stochastic events. Our program runs relatively fast and is suitable for analyzing alignments containing large number of sequences. AVAILABILITY: ftp://iole.swmed.edu/pub/PCOAT/. SUPPLEMENTARY INFORMATION: The PCOAT ftp site contains a detailed description of the program, and the results of PCOAT analysis on C2H2 alignment and ACT domain alignment.

Algorithms↗

CASP5 target classification.

This report summarizes the Critical Assessment of Protein Structure Prediction (CASP5) target proteins, which included 67 experimental models submitted from various structural genomics efforts and independent research groups. Throughout this special issue, CASP5 targets are referred to with the identification numbers T0129-T0195. Several of these targets were excluded from the assessment for various reasons: T0164 and T0166 were cancelled by the organizers; T0131, T0144, T0158, T0163, T0171, T0175, and T0180 were not available in time; T0145 was "natively unfolded"; the T0139 structure became available before the target expired; and T0194 was solved for a different sequence than the one submitted. Table I outlines the sequence and structural information available for CASP5 proteins in the context of existing folds and evolutionary relationships. This information provided the basis for a domain-based classification of the target structures into three assessment categories: comparative modeling (CM), fold recognition (FR), and new fold (NF). The FR category was further subdivided into homologues [FR(H)] and analogs [FR(A)] based on evolutionary considerations, and the overlap between assessment categories was classified as CM/FR(H) and FR(A)/NF. CASP5 domains are illustrated in Figure 1. Examples of nontrivial links between CASP5 target domains and existing structures that support our classifications are provided.

Binding Sites↗

CASP5 assessment of fold recognition target predictions.

We present an overview of the fifth round of Critical Assessment of Protein Structure Prediction (CASP5) fold recognition category. Prediction models were evaluated by using six different structural measures and four different alignment measures, and these scores were compared to those assigned manually over a diverse subset of target domains. Scores were combined to compare overall performance of participating groups and to estimate rank significance. The methods used by a few groups outperformed all other methods in terms of the evaluated criteria and could be considered state-of-the-art in structure prediction. We discuss a few examples of difficult fold recognition targets to highlight the progress of ab initio-type methods on difficult structure analogs and the difficulties of predicting multidomain targets and selecting prediction models. We also compared the results of manual groups to those of automatic servers evaluated in parallel by CAFASP, showing that the top performing automated server structure predictions approached those of the best manual predictors.

Algorithms↗

C-terminal domain of gyrase A is predicted to have a beta-propeller structure.

Two different type II topoisomerases are known in bacteria. DNA gyrase (Gyr) introduces negative supercoils into DNA. Topoisomerase IV (Par) relaxes DNA supercoils. GyrA and ParC subunits of bacterial type II topoisomerases are involved in breakage and reunion of DNA. The spatial structure of the C-terminal fragment in GyrA/ParC is not available. We infer homology between the C-terminal domain of GyrA/ParC and a regulator of chromosome condensation (RCC1), a eukaryotic protein that functions as a guanine-nucleotide-exchange factor for the nuclear G protein Ran. This homology, complemented by detection of 6 sequence repeats with 4 predicted beta-strands each in GyrA/ParC sequences, allows us to predict that the GyrA/ParC C-terminal domain folds into a 6-bladed beta-propeller. The prediction rationalizes available experimental data and sheds light on the spatial properties of the largest topoisomerase domain that lacks structural information.

Amino Acid Sequence↗

Rapid genotyping of Bacillus anthracis strains by real-time polymerase chain reaction.

Rapid and accurate identification of Bacillus anthracis is critical for patient care as well as outbreak control. We have developed 3 separate PCR based assays using fluorescence resonance energy transfer (FRET) to detect the presence of pXO1, pXO2 plasmids and a chromosomal marker. A set of amplification primers and probes were used in each assay. The probes were ad jacently placed inside the primer sites and were 1-bp apart. The upstream probe was labeled with fluorescein at the 3' end, and the downstream probe had Cy5 attached at the 5' end. The probes are included in the PCR reactions and hybridize with the PCR products as they are formed. Binding of probes to PCR products results in transfer of energy from fluorescein to Cy5, resulting in emission from Cy5. Increase in fluorescence, indicating amplification, was monitored in real time on a LightCycler((TM)) LC24. Initial denaturation of target sequences was accomplished at 95 degrees C for 1 min, followed by 28 cycles of denaturation at 95 degrees C for 0 sec, annealing at 58 degrees C for 15 sec, and elongation at 72 degrees C for 5 sec. These assays are specific and can be performed on as little as 25 ng of total DNA or crude cell lysate a from fresh colony. It is thus possible to deter mine the genotype of B. anthracis strains in less than 1 hour.

Animals↗