Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Comprehensive genomic profiling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Multiregion profiling of genomic and transcriptional heterogeneity in head and neck squamous-cell carcinoma.

BACKGROUND: Intratumoral heterogeneity (ITH) is thought to contribute to tumour evolution and treatment resistance but its biological and clinical significance in localised head and neck squamous-cell carcinoma (HNSCC) remains incompletely understood. PATIENTS AND METHODS: In the prospective SCANDARE study, we analysed 87 patients with resectable HNSCC treated with upfront surgery. Two to five spatially distinct tumour regions per patient underwent pathological evaluation, targeted DNA sequencing, and bulk RNA sequencing. Genomic ITH (gITH) was quantified using clonal deconvolution and Shannon diversity indices, whereas transcriptional heterogeneity (tITH) was assessed using the intratumour expression distance metric. Associations between ITH, molecular features, tumour microenvironment composition, and clinical outcomes were explored using multivariable statistical models. RESULTS: Pathology-based spatial heterogeneity showed limited prognostic value. gITH was common, with 37% of tumours displaying regionally heterogeneous pathogenic variants, including spatially actionable alterations in 10% of patients. In an initial multivariable Cox model, higher gITH was associated with shorter disease-free survival. However, after Ridge-penalised modelling and bootstrap internal validation, the effect size was attenuated [corrected hazard ratio 1.42, 95% confidence interval (CI) 0.91-2.75]. The overall model retained moderate discriminative performance (optimism-corrected C-index 0.69, 95% CI 0.59-0.79). gITH was associated with tumour cellularity, reduced estimated endothelial cell infiltration, and alterations in KMT2C and PIK3CA. tITH differed according to human papillomavirus (HPV) status, with lower tITH in HPV-positive tumours, and was associated with distinct biological pathways and genomic alterations. Genomic and tITH were not correlated. CONCLUSIONS: This prospective multiregion study provides a comprehensive characterisation of genomic and tITH in localised HNSCC. Our findings highlight substantial spatial molecular diversity within primary tumours and suggest potential associations between heterogeneity, tumour biology, and clinical outcome that warrant validation in independent cohorts.

head and neck squamous-cell carcinoma (HNSCC)↗

arrayCGHbase: an analysis platform for comparative genomic hybridization microarrays.

BACKGROUND: The availability of the human genome sequence as well as the large number of physically accessible oligonucleotides, cDNA, and BAC clones across the entire genome has triggered and accelerated the use of several platforms for analysis of DNA copy number changes, amongst others microarray comparative genomic hybridization (arrayCGH). One of the challenges inherent to this new technology is the management and analysis of large numbers of data points generated in each individual experiment. RESULTS: We have developed arrayCGHbase, a comprehensive analysis platform for arrayCGH experiments consisting of a MIAME (Minimal Information About a Microarray Experiment) supportive database using MySQL underlying a data mining web tool, to store, analyze, interpret, compare, and visualize arrayCGH results in a uniform and user-friendly format. Following its flexible design, arrayCGHbase is compatible with all existing and forthcoming arrayCGH platforms. Data can be exported in a multitude of formats, including BED files to map copy number information on the genome using the Ensembl or UCSC genome browser. CONCLUSION: ArrayCGHbase is a web based and platform independent arrayCGH data analysis tool, that allows users to access the analysis suite through the internet or a local intranet after installation on a private server. ArrayCGHbase is available at http://medgen.ugent.be/arrayCGHbase/.

Base Sequence↗

Identification of a WRKY protein as a transcriptional regulator of benzylisoquinoline alkaloid biosynthesis in Coptis japonica.

Selected cultured Coptis japonica cells produce a large amount of the benzylisoquinoline alkaloid berberine. Previous studies have suggested that berberine productivity is controlled at the transcript level of biosynthetic genes. We have identified a regulator of transcription in berberine biosynthesis using functional genomics with a transient RNA interference (RNAi) and overexpression of the candidate gene. The 24 primary candidate clones were selected from 1,014 expressed sequence tags (ESTs) that were obtained from a C. japonica cell line producing high levels of berberine. Further characterization of the expression profiles of these ESTs suggested that five ESTs would be good candidates as regulators of berberine production. A newly developed transient RNAi system with C. japonica protoplasts indicated that double-stranded RNA of an EST clone significantly reduced the level of transcripts of 3'-hydroxy N-methylcoclaurine 4'-O-methyltransferase. Sequence analysis showed that this EST encoded a group-II WRKY, and we named it CjWRKY1. When the effects of double-stranded RNA of the CjWRKY1 gene were examined in detail, a marked reduction in the transcripts of all genes involved in berberine biosynthesis was detected, whereas little effect was found in the transcript levels of glyceraldehyde-3-phosphate dehydrogenase (GAPDH) and chorismate mutase (CM) that are associated with primary metabolism. Ectopic expression of CjWRKY1 cDNA in C. japonica protoplasts clearly increased the level of transcripts of all berberine biosynthetic genes examined compared with control treatment, whereas the levels of GAPDH and CM were not affected. The functional role of CjWRKY1 as a specific and comprehensive regulator of berberine biosynthesis is discussed.

Alkaloids↗

A comparative analysis of HGSC and Celera human genome assemblies and gene sets.

MOTIVATION: Since the simultaneous publication of the human genome assembly by the International Human Genome Sequencing Consortium (HGSC) and Celera Genomics, several comparisons have been made of various aspects of these two assemblies. In this work, we set out to provide a more comprehensive comparative analysis of the two assemblies and their associated gene sets. RESULTS: The local sequence content for both draft genome assemblies has been similar since the early releases, however it took a year for the quality of the Celera assembly to approach that of HGSC, suggesting an advantage of HGSC's hierarchical shotgun (HS) sequencing strategy over Celera's whole genome shotgun (WGS) approach. While similar numbers of ab initio predicted genes can be derived from both assemblies, Celera's Otto approach consistently generated larger, more varied gene sets than the Ensembl gene build system. The presence of a non-overlapping gene set has persisted with successive data releases from both groups. Since most of the unique genes from either genome assembly could be mapped back to the other assembly, we conclude that the gene set discrepancies do not reflect differences in local sequence content but rather in the assemblies and especially the different gene-prediction methodologies.

Databases, Protein↗

Hidden Markov model variants and their application.

Markov statistical methods may make it possible to develop an unsupervised learning process that can automatically identify genomic structure in prokaryotes in a comprehensive way. This approach is based on mutual information, probabilistic measures, hidden Markov models, and other purely statistical inputs. This approach also provides a uniquely common ground for comparative prokaryotic genomics. The approach is an on-going effort by its nature, as a multi-pass learning process, where each round is more informed than the last, and thereby allows a shift to the more powerful methods available for supervised learning at each iteration. It is envisaged that this "bootstrap" learning process will also be useful as a knowledge discovery tool. For such an ab initio prokaryotic gene-finder to work, however, it needs a mechanism to identify critical motif structure, such as those around the start of coding or start of transcription (and then, hopefully more).For eukaryotes, even with better start-of-coding identification, parsing of eukaryotic coding regions by the HMM is still limited by the HMM's single gene assumption, as evidenced by the poor performance in alternatively spliced regions. To address these complications an approach is described to expand the states in a eukaryotic gene-predictor HMM, to operate with two layers of DNA parsing. This extension from the single layer gene prediction parse is indicated after preliminary analysis of the C. elegans alt-splice statistics. State profiles have made use of a novel hash-interpolating MM (hIMM) method. A new implementation for an HMM-with-Duration is also described, with far-reaching application to gene-structure identification and analysis of channel current blockade data.

Alternative Splicing↗

A global survey of gene regulation during cold acclimation in Arabidopsis thaliana.

Many temperate plant species such as Arabidopsis thaliana are able to increase their freezing tolerance when exposed to low, nonfreezing temperatures in a process called cold acclimation. This process is accompanied by complex changes in gene expression. Previous studies have investigated these changes but have mainly focused on individual or small groups of genes. We present a comprehensive statistical analysis of the genome-wide changes of gene expression in response to 14 d of cold acclimation in Arabidopsis, and provide a large-scale validation of these data by comparing datasets obtained for the Affymetrix ATH1 Genechip and MWG 50-mer oligonucleotide whole-genome microarrays. We combine these datasets with existing published and publicly available data investigating Arabidopsis gene expression in response to low temperature. All data are integrated into a database detailing the cold responsiveness of 22,043 genes as a function of time of exposure at low temperature. We concentrate our functional analysis on global changes marking relevant pathways or functional groups of genes. These analyses provide a statistical basis for many previously reported changes, identify so far unreported changes, and show which processes predominate during different times of cold acclimation. This approach offers the fullest characterization of global changes in gene expression in response to low temperature available to date.

Acclimatization↗

Interactome: gateway into systems biology.

Protein-protein interactions are fundamental to all biological processes, and a comprehensive determination of all protein-protein interactions that can take place in an organism provides a framework for understanding biology as an integrated system. The availability of genome-scale sets of cloned open reading frames has facilitated systematic efforts at creating proteome-scale data sets of protein-protein interactions, which are represented as complex networks or 'interactome' maps. Protein-protein interaction mapping projects that follow stringent criteria, coupled with experimental validation in orthogonal systems, provide high-confidence data sets immanently useful for interrogating developmental and disease mechanisms at a system level as well as elucidating individual protein function and interactome network topology. Although far from complete, currently available maps provide insight into how biochemical properties of proteins and protein complexes are integrated into biological systems. Such maps are also a useful resource to predict the function(s) of thousands of genes.

Animals↗

MGOS: A resource for studying Magnaporthe grisea and Oryza sativa interactions.

The MGOS (Magnaporthe grisea Oryza sativa) web-based database contains data from Oryza sativa and Magnaporthe grisea interaction experiments in which M. grisea is the fungal pathogen that causes the rice blast disease. In order to study the interactions, a consortium of fungal and rice geneticists was formed to construct a comprehensive set of experiments that would elucidate information about the gene expression of both rice and M. grisea during the infection cycle. These experiments included constructing and sequencing cDNA and robust long-serial analysis gene expression libraries from both host and pathogen during different stages of infection in both resistant and susceptible interactions, generating >50,000 M. grisea mutants and applying them to susceptible rice strains to test for pathogenicity, and constructing a dual O. sativa-M. grisea microarray. MGOS was developed as a central web-based repository for all the experimental data along with the rice and M. grisea genomic sequence. Community-based annotation is available for the M. grisea genes to aid in the study of the interactions.

Computational Biology↗

Comprehensive DNA microarray analysis of Bacillus subtilis two-component regulatory systems.

It has recently been shown through DNA microarray analysis of Bacillus subtilis two-component regulatory systems (DegS-DegU, ComP-ComA, and PhoR-PhoP) that overproduction of a response regulator of the two-component systems in the background of a deficiency of its cognate sensor kinase affects the regulation of genes, including its target ones. The genome-wide effect on gene expression caused by the overproduction was revealed by DNA microarray analysis. In the present work, we newly analyzed 24 two-component systems by means of this strategy, leaving out 8 systems to which it was unlikely to be applicable. This analysis revealed various target gene candidates for these two-component systems. It is especially notable that interesting interactions appeared to take place between several two-component systems. Moreover, the probable functions of some unknown two-component systems were deduced from the list of their target gene candidates. This work is heuristic but provides valuable information for further study toward a comprehensive understanding of the B. subtilis two-component regulatory systems. The DNA microarray data obtained in this work are available at the KEGG Expression Database website (http://www.genome.ad.jp/kegg/expression).

Bacillus subtilis↗

A classification-based framework for predicting and analyzing gene regulatory response.

BACKGROUND: We have recently introduced a predictive framework for studying gene transcriptional regulation in simpler organisms using a novel supervised learning algorithm called GeneClass. GeneClass is motivated by the hypothesis that in model organisms such as Saccharomyces cerevisiae, we can learn a decision rule for predicting whether a gene is up- or down-regulated in a particular microarray experiment based on the presence of binding site subsequences ("motifs") in the gene's regulatory region and the expression levels of regulators such as transcription factors in the experiment ("parents"). GeneClass formulates the learning task as a classification problem--predicting +1 and -1 labels corresponding to up- and down-regulation beyond the levels of biological and measurement noise in microarray measurements. Using the Adaboost algorithm, GeneClass learns a prediction function in the form of an alternating decision tree, a margin-based generalization of a decision tree. METHODS: In the current work, we introduce a new, robust version of the GeneClass algorithm that increases stability and computational efficiency, yielding a more scalable and reliable predictive model. The improved stability of the prediction tree enables us to introduce a detailed post-processing framework for biological interpretation, including individual and group target gene analysis to reveal condition-specific regulation programs and to suggest signaling pathways. Robust GeneClass uses a novel stabilized variant of boosting that allows a set of correlated features, rather than single features, to be included at nodes of the tree; in this way, biologically important features that are correlated with the single best feature are retained rather than decorrelated and lost in the next round of boosting. Other computational developments include fast matrix computation of the loss function for all features, allowing scalability to large datasets, and the use of abstaining weak rules, which results in a more shallow and interpretable tree. We also show how to incorporate genome-wide protein-DNA binding data from ChIP chip experiments into the GeneClass algorithm, and we use an improved noise model for gene expression data. RESULTS: Using the improved scalability of Robust GeneClass, we present larger scale experiments on a yeast environmental stress dataset, training and testing on all genes and using a comprehensive set of potential regulators. We demonstrate the improved stability of the features in the learned prediction tree, and we show the utility of the post-processing framework by analyzing two groups of genes in yeast--the protein chaperones and a set of putative targets of the Nrg1 and Nrg2 transcription factors--and suggesting novel hypotheses about their transcriptional and post-transcriptional regulation. Detailed results and Robust GeneClass source code is available for download from http://www.cs.columbia.edu/compbio/robust-geneclass.

Algorithms↗

CRISPR/Cas9-driven double modification of grapevine MLO6-7 imparts powdery mildew resistance, while editing of NPR3 augments powdery and downy mildew tolerance.

The implementation of genome editing strategies in grapevine is the easiest way to improve sustainability and resilience while preserving the original genotype. Among others, the Mildew Locus-O (MLO) genes have already been reported as good candidates to develop powdery mildew-immune plants. A never-explored grapevine target is NPR3, a negative regulator of the systemic acquired resistance. We report the exploitation of a cisgenic approach with the Cre-lox recombinase technology to generate grapevine-edited plants with the potential to be transgene-free while preserving their original genetic background. The characterization of three edited lines for each target demonstrated immunity development against Erysiphe necator in MLO6-7-edited plants. Concomitantly, a significant improvement of resilience, associated with increased leaf thickness and specific biochemical responses, was observed in defective NPR3 lines against E. necator and Plasmopara viticola. Transcriptomic analysis revealed that both MLO6-7 and NPR3 defective lines modulated their gene expression profiles, pointing to distinct though partially overlapping responses. Furthermore, targeted metabolite analysis highlighted an overaccumulation of stilbenes coupled with an improved oxidative scavenging potential in both editing targets, likely protecting the MLO6-7 mutants from detrimental pleiotropic effects. Finally, the Cre-loxP approach allowed the recovery of one MLO6-7 edited plant with the complete removal of transgene. Taken together, our achievements provide a comprehensive understanding of the molecular and biochemical adjustments occurring in double MLO-defective grape plants. In parallel, the potential of NPR3 mutants for multiple purposes has been demonstrated, raising new questions on its wide role in orchestrating biotic stress responses.

Vitis↗

Mass distributed clustering: a new algorithm for repeated measurements in gene expression data.

The availability of whole-genome sequence data and high-throughput techniques such as DNA microarray enable researchers to monitor the alteration of gene expression by a certain organ or tissue in a comprehensive manner. The quantity of gene expression data can be greater than 30,000 genes per one measurement, making data clustering methods for analysis essential. Biologists usually design experimental protocols so that statistical significance can be evaluated; often, they conduct experiments in triplicate to generate a mean and standard deviation. Existing clustering methods usually use these mean or median values, rather than the original data, and take significance into account by omitting data showing large standard deviations, which eliminates potentially useful information. We propose a clustering method that uses each of the triplicate data sets as a probability distribution function instead of pooling data points into a median or mean. This method permits truly unsupervised clustering of the data from DNA microarrays.

Algorithms↗

An annotation update via cDNA sequence analysis and comprehensive profiling of developmental, hormonal or environmental responsiveness of the Arabidopsis AP2/EREBP transcription factor gene family.

AP2/EREBP transcription factors (TFs) play functionally important roles in plant growth and development, especially in hormonal regulation and in response to environmental stress. Here we reported verification and correction of annotation through an exhaustive cDNA cloning and sequence analysis performed on 145 of 147 gene family members. A RACE analysis performed on genes with potential in-frame up-stream ATG codon resulted in identification of At2g28520 as an authentic AP2/EREBP member and corrected ORF annotations for three other members. A further phylogenetic analysis of this updated and likely complete family divided it into three major subfamilies. The expression patterns of the AP2/EREBP family members among the 11 organ or tissue types were examined using an oligo microarray and their hormonal and environmental responsiveness were further characterized using cDNA custom macroarrays. These detailed expression profile results provide strong support for a role for AP2/EREBP family members in development and in response to environmental stimuli, and a foundation for future functional analysis of this gene family.

Algorithms↗

5'SAGE: 5'-end Serial Analysis of Gene Expression database.

To comprehensively identify transcription start sites and the frequencies of individual mRNAs in human cell libraries, a method of 5' end Serial Analysis of Gene Expression (SAGE) was developed recently, which makes it possible to collect a large amount of start site information, and subsequently, we have established a related database server called 5'SAGE. This database displays the observed frequencies of individual 5' end SAGE tags and previously unknown transcription start sites in the promoter regions, introns and intergenic regions of known genes. 5'SAGE will be useful for analyzing promoter regions and start site variation in different tissues, and is freely available at http://5sage.gi.k.u-tokyo.ac.jp/.

5' Flanking Region↗

Systematic discovery of regulatory motifs in human promoters and 3' UTRs by comparison of several mammals.

Comprehensive identification of all functional elements encoded in the human genome is a fundamental need in biomedical research. Here, we present a comparative analysis of the human, mouse, rat and dog genomes to create a systematic catalogue of common regulatory motifs in promoters and 3' untranslated regions (3' UTRs). The promoter analysis yields 174 candidate motifs, including most previously known transcription-factor binding sites and 105 new motifs. The 3'-UTR analysis yields 106 motifs likely to be involved in post-transcriptional regulation. Nearly one-half are associated with microRNAs (miRNAs), leading to the discovery of many new miRNA genes and their likely target genes. Our results suggest that previous estimates of the number of human miRNA genes were low, and that miRNAs regulate at least 20% of human genes. The overall results provide a systematic view of gene regulation in the human, which will be refined as additional mammalian genomes become available.

3' Untranslated Regions↗

A comprehensive rice transcript map containing 6591 expressed sequence tag sites.

To determine the chromosomal positions of expressed rice genes, we have performed an expressed sequence tag (EST) mapping project by polymerase chain reaction-based yeast artificial chromosome (YAC) screening. Specific primers designed from 6713 unique EST sequences derived from 19 cDNA libraries were screened on 4387 YAC clones and used for map construction in combination with genetic analysis. Here, we describe the establishment of a comprehensive YAC-based rice transcript map that contains 6591 EST sites and covers 80.8% of the rice genome. Chromosomes 1, 2, and 3 have relatively high EST densities, approximately twice those of chromosomes 11 and 12, and contain 41% of the total EST sites on the map. Most of the EST-dense regions are distributed on the distal regions of each chromosome arm. Genomic regions flanking the centromeres for most of the chromosomes have lower EST density. Recombination frequency in these regions is suppressed significantly. Our EST mapping also shows that 40% of the assigned ESTs occupy only approximately 21% of the entire genome. The rice transcript map has been a valuable resource for genetic study, gene isolation, and genome sequencing at the Rice Genome Research Program and should become an important tool for comparative analysis of chromosome structure and evolution among the cereals.

Chromosome Mapping↗

Gliogenesis in Drosophila: genome-wide analysis of downstream genes of glial cells missing in the embryonic nervous system.

In Drosophila, the glial cells missing (gcm) gene encodes a transcription factor that controls the determination of glial versus neuronal fate. In gcm mutants, presumptive glial cells are transformed into neurons and, conversely, when gcm is ectopically misexpressed, presumptive neurons become glia. Although gcm is thought to initiate glial cell development through its action on downstream genes that execute the glial differentiation program, little is known about the identity of these genes. To identify gcm downstream genes in a comprehensive manner, we used genome-wide oligonucleotide arrays to analyze differential gene expression in wild-type embryos versus embryos in which gcm is misexpressed throughout the neuroectoderm. Transcripts were analyzed at two defined temporal windows during embryogenesis. During the first period of initial gcm action on determination of glial cell precursors, over 400 genes were differentially regulated. Among these are numerous genes that encode other transcription factors, which underscores the master regulatory role of gcm in gliogenesis. During a second later period, when glial cells had already differentiated, over 1200 genes were differentially regulated. Most of these genes, including many genes for chromatin remodeling factors and cell cycle regulators, were not differentially expressed at the early stage, indicating that the genetic control of glial fate determination is largely different from that involved in maintenance of differentiated cells. At both stages, glial-specific genes were upregulated and neuron-specific genes were downregulated, supporting a model whereby gcm promotes glial development by activating glial genes, while simultaneously repressing neuronal genes. In addition, at both stages, numerous genes that were not previously known to be involved in glial development were differentially regulated and, thus, identified as potential new downstream targets of gcm. For a subset of the differentially regulated genes, tissue-specific in vivo expression data were obtained that confirmed the transcript profiling results. This first genome-wide analysis of gene expression events downstream of a key developmental transcription factor presents a novel level of insight into the repertoire of genes that initiate and maintain cell fate choices in CNS development.

Animals↗

Altered expression of genes involved in hepatic morphogenesis and fibrogenesis are identified by cDNA microarray analysis in biliary atresia.

Biliary atresia (BA) is characterized by a progressive, sclerosing, inflammatory process that leads to cirrhosis in infancy. Although it is the most common indication for liver transplantation in early childhood, little is known about its etiopathogenesis. To elucidate factors involved in this process, we performed comprehensive genome-wide gene expression analysis using complementary DNA (cDNA) microarrays. We compared messenger RNA expression levels of approximately 18,000 human genes from normal, diseased control, and end-stage BA livers. Reverse-transcription polymerase chain reaction (RT-PCR) and Northern blot analysis were performed to confirm changes in gene expression. Cluster and principal component analysis showed that all BA samples clustered together, forming a distinct group well separated from normal and diseased controls. We further identified 35 genes and ESTs whose expression differentiated BA from normal and diseased controls. Most of these genes are known to be associated with cell signaling, transcription regulation, hepatic development, morphogenesis, and fibrogenesis. In conclusion, this study serves to delineate processes that are involved in the pathogenesis of BA.

Adolescent↗