Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Comprehensive genomic profiling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Comprehensive analysis of pathway or functionally related gene expression in the National Cancer Institute's anticancer screen.

We have analyzed the level of gene coregulation, using gene expression patterns measured across the National Cancer Institute's 60 tumor cell panels (NCI(60)), in the context of predefined pathways or functional categories annotated by KEGG (Kyoto Encyclopedia of Genes and Genomes), BioCarta, and GO (Gene Ontology). Statistical methods were used to evaluate the level of gene expression coherence (coordinated expression) by comparing intra- and interpathway gene-gene correlations. Our results show that gene expression in pathways, or groups of functionally related genes, has a significantly higher level of coherence than that of a randomly selected set of genes. Transcriptional-level gene regulation appears to be on a "need to be" basis, such that pathways comprising genes encoding closely interacting proteins and pathways responsible for vital cellular processes or processes that are related to growth or proliferation, specifically in cancer cells, such as those engaged in genetic information processing, cell cycle, energy metabolism, and nucleotide metabolism, tend to be more modular (lower degree of gene sharing) and to have genes significantly more coherently expressed than most signaling and regular metabolic pathways. Hierarchical clustering of pathways based on their differential gene expression in the NCI(60) further revealed interesting interpathway communications or interactions indicative of a higher level of pathway regulation. The knowledge of the nature of gene expression regulation and biological pathways can be applied to understanding the mechanism by which small drug molecules interfere with biological systems.

Algorithms↗

Age-stratified mutation patterns in early-onset colorectal cancer reveal distinct molecular features and therapeutic implications.

BACKGROUND: Colorectal cancer (CRC) is increasingly diagnosed in younger adults, with evidence that early-onset cases (age <50 years) differ in the spectrum of prevalent gene mutations compared with older individuals. To evaluate how these age-related differences may inform testing guidelines and therapeutic development, we examined mutation rates of the most prevalent gene mutations across four age-stratified cohorts. PATIENTS AND METHODS: Clinicogenomic data were obtained from Memorial Sloan Kettering Center for Harmonized Onco-genomic Research Dataset and China Pan-Cancer cohorts available in cBioPortal. A total of 6762 samples were analyzed. Mutation frequencies for a comprehensive panel of the 100 most prevalent CRC genes were compared across four age groups: 18-29 (n = 79), 30-39 (n = 402), 40-49 (n = 1064), and &#x2265;50 (n = 5217) using chi-square analysis. False discovery rate (FDR) correction for multiple comparisons was carried out using Benjamini-Hochberg procedure. RESULTS: Statistically significant variation in mutation frequency across age groups was seen in 22 key genes. APC mutations increased with age and were seen in 49.4% of patients in the 18-29 group, 69.7% in 30-39, 73.3% in 40-49, and 75.25% of patients &#x2265;50 (P < 0.001, FDR < 0.001). The oldest cohort was more than three times more likely to have an APC mutation than the youngest [odds ratio (OR) = 3.74, 95% confidence interval (CI) 2.44-5.74, P < 0.001]. In contrast, SMAD4 mutations were twice as common in the youngest age group at 31.6% compared with those over 40, with a prevalence of 17.29% in patients 40-49, and 18.84% in patients over 50 (OR = 2.03, 95% CI 1.26-3.27, P < 0.001, FDR < 0.001). POLE mutations peaked in the 30-39 age group with a prevalence of 10.7% compared with 6.3% in patients aged 18-29, 4.9% in patients aged 40-49, and 5.9% in patients aged &#x2265;50 (P < 0.001, FDR < 0.001). Individuals in the 30-39 group were nearly twice as likely to carry a POLE mutation compared with those over 40 (OR = 1.96, 95% CI 1.41-2.74, P < 0.001). CONCLUSIONS: Differences in mutations of key genes including a lower prevalence of APC mutations and increased SMAD4 mutations in younger individuals provides further supporting evidence that early-onset CRC may represent a distinct biological subtype of CRC. Enrichment of POLE mutations in younger patients highlights the importance of expanded molecular profiling in early-onset CRC, which could help identify patients most likely to benefit from immunotherapy and advance personalized treatment strategies in CRC. Together, these findings reinforce the need to approach early-onset CRC as a distinct biological entity and ensure that appropriate molecular assays are incorporated to guide care.

APC↗

Comprehensive search for chicken W chromosome-linked genes expressed in early female embryos from the female-minus-male subtracted cDNA macroarray.

In order to seek chicken W chromosome-linked genes expressed significantly earlier than the time of gonadal differentiation, female-minus-male-subtracted cDNA macroarrays were prepared from day 2 (Hamburger-Hamilton stages 12-13), day 3 (stages 19-20) and day 4 (stages 24-25) embryos. From a total of 15-744 macroarrayed cDNA clones, 610 clones exhibiting significantly female-specific expression were selected. When each one of the 610 cDNA clones was used as a probe in Southern blot hybridization with male or female chicken genomic DNA, 62 clones, grouped into eight (A-H) types according to their patterns of hybridization, were considered to be derived from W chromosome-linked genes. When representative cDNA clones in each type were sequenced, clones derived from two known W-linked genes; SPIN-W and ATP5A1W , and from two hitherto unknown W-linked genes, represented by 2d-2D9 and 2d-2F9 clones, were identified and their localizations on the W chromosome were confirmed by fluorescence in-situ hybridization. The 2d-2D9 sequence has no significant homology with other genes in databases but 2d-2F9 has a region which shows partial homology to the consensus sequence of the AAA ATPase superfamily. Both 2d-2D9 and 2d-2F9 sequences are found in contigs of undetermined chromosome-linkage in the Draft Chicken Genome Sequence.

Animals↗

Comprehensive analysis of 19q12 amplicon in human gastric cancers.

Amplification at 19q12 has been observed in multiple tumor types, while cyclin E1 (CCNE1) has been considered to be the key oncogene within this amplicon. We have previously applied cDNA microarray analysis to systematically characterize gene expression patterns of gastric tumor and nontumor samples. We identified a cluster of five tightly coregulated genes all located at chromosome 19q12, including CCNE1. We found that the 19q12 gene cluster is highly expressed in gastric tumors compared to nontumor gastric samples. Array based comparative genomic hybridization and real-time PCR was used to define the boundary of the 19q12 amplicon to a region of approximately 200 kb. Interestingly, we found that in some cases amplification at 19q12 was not associated with DNA copy number gain at CCNE1, suggesting that some other genes within the 19q12 amplicon may also have important function during gastric tumorigenesis. We found high expression of the 19q12 gene cluster to be statistically correlated with the cell proliferation gene signature. Using the SAM software, we identified a set of 577 genes whose expression levels positively correlated with the 19q12 gene cluster. GO term analysis revealed that this genelist is enriched with genes involved in cell cycle regulation and cell proliferation. In conclusion, expression array analysis combined with array comparative genomic hybridization and real-time PCR provides a new and powerful tool to identify clusters of genes which may be regulated by genomic DNA aberrations. In addition, our study indicates that amplification at 19q12 is associated with cell proliferation in vivo.

Adenocarcinoma↗

De novo transcriptome assembly and gene expression analysis of Cnidium officinale under high-temperature conditions.

BACKGROUND: The medicinal plant Cnidium officinale (CO) is widespread in Northeast Asia and vulnerable to heat stress. The naturally occurring composition of pharmacological ingredients of CO results in overall physiological consequences; therefore, it is crucial to have a comprehensive understanding of metabolic response to ambient heat in terms of acclimation to estimate how much CO is exposed to threatening environmental conditions. RESULTS: Transcriptome analysis is critical for understanding the consequences of long-term physiological adaptation of CO to abiotic stress. However, transcriptome analysis on this species, particularly under prolonged stress conditions, has remained limited. We employed a temperature gradient tunnel (TGT) to subject CO to high-temperature exposure for four months, enabling us to observe the cumulative effects of heat and assess its acclimation mechanisms. In the absence of genome sequencing data, we performed de novo transcriptome assembly and compared DEGs from temperature treatment plots of a TGT and a growth chamber (GC). Since interpreting transcriptomic data can be complex, we employed a sequential analytical approach, including DEG clustering, GO enrichment, KEGG pathway mapping, miRNA-target gene analysis, and multiple rounds of RNA sequencing validation. DEGs were classified into two categories: genes exhibiting significant fold changes and genes showing significant count changes rather than fold changes. Then, we analyzed the functional roles&#xa0;of DEGs to determine which pathways respond to ambient and stressful high temperatures and validated the findings through cross-comparison with GC. Additionally, we conducted miRNA analysis to investigate post-transcriptional regulation under high temperatures. CO grown under higher ambient temperatures exhibited slight upregulation of pathways related to protein stability and turnover, ABA biosynthesis, and energy production, such as photosynthesis and oxidative phosphorylation. However, under extreme heat stress, most metabolic pathways were downregulated except for those involved in transcription, translation, oxidative phosphorylation and the biosynthesis of cutin, suberin, and wax. CONCLUSION: This study demonstrated that proper clustering of genes based on expression levels and fold changes in two different experimental conditions, along with pathway mapping, may provide a comprehensive understanding of CO's response to heat stress. These insights could contribute to future research on heat tolerance and crop improvement.

Gene Expression Profiling↗

Genome-wide identification of CHY zinc finger and RING finger (CHYR) genes in pepper and functional characterization of CaCHYR5 in response to Phytophthora capsici infection.

CHY zinc finger and RING finger (CHYR) proteins play crucial roles in the growth and development, as well as stress response. To date, no systematic or comprehensive analysis of the CHYR gene family has been performed in pepper (Capsicum annuum L.). In this study, we identified 8 CaCHYR genes (CaCHYR1-CaCHYR8), which were classified into 3 groups based on phylogenetic relationships. CaCHYR members within the same group exhibited similar distributions of conserved motifs and exon-intron structures. Chromosomal localization analysis showed that 8 CaCHYR genes were unevenly distributed on 6 chromosomes. Segmental duplication, rather than tandem duplication, was found to be the major contributor to the expansion of this gene family. CaCHYR genes feature a variety of cis-elements involved in developmental processes, phytohormone responses, and stress adaptation. Expression analysis based on RNA-seq data revealed that CaCHYR genes exhibited distinct spatial expression patterns across different tissues and in response to Phytophthora capsici infection (PCI), and quantitative real-time PCR (qRT-PCR) further confirmed that three of them (CaCHYR2, CaCHYR3, and CaCHYR5) exhibited altered expression under PCI. Furthermore, transient overexpression of CaCHYR5 in pepper leaves increased susceptibility to PCI, suggesting its potential negative regulatory role in pepper defense against P. capsici. Collectively, these findings reveal the expression patterns and regulatory functions of pepper CHYR genes in growth and development, laying a groundwork for breeding pepper cultivars tolerant to PCI.

Phytophthora capsici infection (PCI)↗

Identification of key genes related to bone metastasis of breast cancer using bioinformatics methods and construction of a prognostic model.

Breast cancer (BC) ranks among the most prevalent cancers in females, with bone metastasis significantly compromising patients' quality of life and survival rates. Enhancing our comprehension of BC bone metastasis mechanisms at the molecular level holds promise for improving BC treatment and prognosis. Leveraging bioinformatics tools, we integrated multiple datasets, conducted comprehensive analyses across various databases, identified biomarkers associated with BC bone metastasis, and constructed a prognostic model. Firstly, 3 BC bone metastasis-related datasets were downloaded from gene expression omnibus, the data were merged, and batch effects were removed, followed by identification of differentially expressed genes (DEGs). Gene ontology and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed on the DEGs. A protein-protein interaction network was constructed using the STRING database to screen hub genes. Then, survival analysis of hub genes was performed using the Cancer Genome Atlas (TCGA) database. A prognostic model was constructed using key genes with survival differences, and the model was evaluated. Two hundred ninety-two DEGs were identified. Gene ontology and KEGG pathway enrichment analysis yielded 769 biological processes (BPs), 78 cellular components, 43 molecular functions, and 50 KEGG pathways. Fifteen hub genes were selected from the protein-protein interaction network. Survival analysis revealed 6 genes related to BC survival. The prognostic model identified 4 genes with important predictive value for BC prognosis. Our study utilized bioinformatics analysis to identify a series of DEGs related to BC bone metastasis. Based on further selection of hub genes, we constructed a relatively ideal prognostic model for BC, and identified 4 genes (DLGAP5, TPX2, PLK1, and CENPN) with valuable predictive value for BC prognosis.

Humans↗

Environmental conditions and transcriptional regulation in Escherichia coli: a physiological integrative approach.

Bacteria develop a number of devices for sensing, responding, and adapting to different environmental conditions. Understanding within a genomic perspective how the transcriptional machinery of bacteria is modulated, as a response for changing conditions, is a major challenge for biologists. Knowledge of which genes are turned on or turned off under specific conditions is essential for our understanding of cell behavior. In this study we describe how the information pertaining to gene expression and associated growth conditions (even with very little knowledge of the associated regulatory mechanisms) is gathered from the literature and incorporated into RegulonDB, a database on transcriptional regulation and operon organization in E. coli. The link between growth conditions, signal transduction, and transcriptional regulation is modeled in the database in a simple format that highlights biological relevant information. As far as we know, there is no other database that explicitly clarifies the effect of environmental conditions on gene transcription. We discuss how this knowledge constitutes a benchmark that will impact future research aimed at integration of regulatory responses in the cell; for instance, analysis of microarrays, predicting culture behavior in biotechnological processes, and comprehension of dynamics of regulatory networks. This integrated knowledge will contribute to the future goal of modeling the behavior of E. coli as an entire cell. The RegulonDB database can be accessed on the web at the URL: http://www.cifn.unam.mx/Computational_Biology/regulondb/.

Adaptation, Physiological↗

The Isolation of differentially expressed mRNA sequences by selective amplification via biotin and restriction-mediated enrichment.

Molecular analysis of development frequently implies the isolation and characterization of genes with specific spatial and temporal expression patterns. Several methods have been developed to identify such DNA sequences. The most comprehensive technique involves the genomewide probing of DNA sequence microarrays with mRNA sequences. However, at present this technology is limited to the few organisms for which the entire genome has been sequenced. Here, we describe a subtractive hybridization technique, called selective amplification via biotin and restriction-mediated enrichment (SABRE), which allows the selective amplification of cDNA fragments representing differentially expressed mRNA species. The method involves the competitive hybridization of an excess of driver cDNA fragments (D) to a trace of tester cDNA fragments (T), and the subsequent purification of tester homohybrids (in which both strands are contributed by the tester cDNA). After competitive hybridization, cDNA fragments that are more abundant in the tester than in the driver are enriched in the tester homohybrids. However, as the fraction of tester homohybrids is very small [T(2)/(D + T)(2)], their purification requires highly efficient procedures. In SABRE, the isolation of tester homohybrids is afforded by a combination of three successive steps: removal of biotinylated terminal sequences from most of the heterohybrids by S1 nuclease digestion, capture of biotinylated hybrids with streptavidin-coated paramagnetic beads, and specific release of homohybrids from the beads by restriction nuclease digestion. If several rounds of SABRE selection are performed in series, even relatively rare differentially expressed mRNA sequences may result in the production of predominant cDNA fragments in the final tester homohybrid population.

Animals↗

Where are we in genomics?

Genomic studies provide scientists with methods to quickly analyse genes and their products en masse. The first high-throughput techniques to be developed were sequencing methods. A great number of genomes from different organisms have thus been sequenced. Genomics is now shifting to the study of gene expression and function. In the past 5-10 years genomics, proteomics and high-throughput microarray technologies have fundamentally changed our ability to study the molecular basis of cells and tissues in health and diseases, giving a new comprehensive view. For example, in cancer research we have seen new diagnostic opportunities for tumour classification, and prognostication. A new exciting development is metabolomics and lab-on-a-chip techniques (which combine miniaturization and automation) for metabolic studies. However, to interpret the large amount of data, extensive computational development is required. In the coming years, we will see the study of biological networks dominating the scene in Physiology. The great accumulation of genomics information will be used in computer programs to simulate biologic processes. Originally developed for genome analysis, bioinformatics now encompasses a wide range of fields in biology from gene studies to integrated biology (i.e. combination of different data sets from genes to metabolites). This is systems biology which aims to study biological organisms as a whole. In medicine, scientific results and applied biotechnologies arising from genomics will be used for effective prediction of diseases and risk associated with drugs. Preventive medicine and medical therapy will be personalized. Widespread applications of genomics for personalized medicine will require associations of gene expression pattern with diagnoses, treatment and clinical data. This will help in the discovery and development of drugs. In agriculture and animal science, the outcomes of genomics will include improvement in food safety, in crop yield, in traceability and in quality of animal products (dairy products and meat) through increased efficiency in breeding and better knowledge of animal physiology. Genomics and integrated biology are huge tasks and no single lab can pursue this alone. We are probably at the end of the beginning rather than at the beginning of the end because Genomics will probably change Biology to a greater extent than previously forecasted. In addition, there is a great need for more information and better understanding of genomics before complete public acceptance.

Animals↗

[Genetic factors in myocardial infarction--Results from a candidate gene and a genome-wide approach between beta blockers].

BACKGROUND: Coronary artery disease and myocardial infarction are the most frequent causes of death in the Western societies. Even nowadays, every second myocardial infarction is lethal and hits the patients unexpectedly without previous signs or symptoms. In order to install preventive measures most efficiently, it is necessary to have a detailed knowledge on the pathophysiology of the disease. The identification of patients who are at high risk for suffering from myocardial infarction can be done with epidemiological methods, such as the determination of "traditional" risk factors, like arterial hypertension, hypercholesterolemia, diabetes mellitus or smoking), or eventually in the future using molecular genetic testing. This is of great importance especially for asymptomatic siblings and children from myocardial infarction patients. POLYMORPHISMS: Although traditional risk factors occur frequently in families, they explain only in part the familial accumulation of coronary artery disease. Furthermore, stron genetic effects on the development of coronary artery disease and myocardial infarction have been demonstrated in several studies. These genetic effects can be examined by 1. a candidate gene approach, or 2. a systematic screening of the whole genome. In the first step, several polymorphisms (sequence variations) wee examined in several candidate genes in which a significant influence on a cardiovascular risk factor or intermediate phenotype (such as atherogeneic lipid profile or arterial hypertension) has been shown in the literature. We thus examined in a large population of patients with myocardial infarction and a sample of the general population the effects of the HindIII polymorphism in the lipoproteinlipase gene, of the -344T/C promoter polymorphism in the aldosterone synthase gene and of the 825C/T polymorphism in the gene of the beta3 subunit of the G protein gene (GNB3). In the general population, we could show an association with unfavorable lipid levels in men and in postmenopausal (but not premenopausal) women for the H2H2 genotype of the HindIII lipoproteinlipase polymorphism. However, the theoretical increase in risk for this genotype is not large enough to demonstrate a significant association with myocardial infarction in the population examined. With the promoter polymorphism in the aldosterone synthase gene, anthropometrical and echocardiographical data did not suggest that the polymorphism is a risk factor for myocardial infarction nor for left ventricular remodeling after myocardial infarction, which was observed in earlier studies. Furthermore, we could show an association with arterial hypertension in our general population sample with the polymorphism in the GNB3 gene. However, no association could be demonstrated for this polymorphism with myocardial infarction. AFFECTED SIB-PAIR APPROACH: In a systematic screening of the genome for genes that are relevant in the pathogenesis of coronary artery disease or myocardial infarction, an affected sib-pair approach was followed. 1,261 families were identified in which at least two brothers or sisters were affected with myocardial infarction or severe coronary artery disease, such as percutaneous coronary intervention or coronary after bypass grafting. In a subpopulation of 513 families and 1,407 individuals, we performed a total genome screening. The analyses using the variance component method and the SOLAR program revealed a susceptibility locus for myocardial infarction of chromosome 14q32 with a lod score of 3.89 (genome-wide p < 0.05). This locus comprises a region of about seven centi-Morgan and contains approximately 150 genes. Furthermore, a comprehensive analysis including the cardiovascular risk factors showed that 1. this myocardial infarction locus is unique and does not overlap with chromosomal loci for well-established risk factors, 2. cardiovascular risk factors, such as Lp(a), diabetes mellitus, serum lipids, or arterial hypertension have strong genetic components. CONCLUSION: These findings do not exclude a role of cardiovascular s do not exclude a role of cardiovascular risk factors or candidate genes in the pathogenesis of myocardial infarction, but rather demonstrate that risk factors may act as surrogates of specific underlying disease mechanisms. It is thus necessary to perform a comprehensive analysis of complex polygenic diseases, such as myocardial infarction, including both, established cardiovascular risk factors and genomic data.

Adrenergic beta-Antagonists↗

Extensive association of functionally and cytotopically related mRNAs with Puf family RNA-binding proteins in yeast.

Genes encoding RNA-binding proteins are diverse and abundant in eukaryotic genomes. Although some have been shown to have roles in post-transcriptional regulation of the expression of specific genes, few of these proteins have been studied systematically. We have used an affinity tag to isolate each of the five members of the Puf family of RNA-binding proteins in Saccharomyces cerevisiae and DNA microarrays to comprehensively identify the associated mRNAs. Distinct groups of 40-220 different mRNAs with striking common themes in the functions and subcellular localization of the proteins they encode are associated with each of the five Puf proteins: Puf3p binds nearly exclusively to cytoplasmic mRNAs that encode mitochondrial proteins; Puf1p and Puf2p interact preferentially with mRNAs encoding membrane-associated proteins; Puf4p preferentially binds mRNAs encoding nucleolar ribosomal RNA-processing factors; and Puf5p is associated with mRNAs encoding chromatin modifiers and components of the spindle pole body. We identified distinct sequence motifs in the 3'-untranslated regions of the mRNAs bound by Puf3p, Puf4p, and Puf5p. Three-hybrid assays confirmed the role of these motifs in specific RNA-protein interactions in vivo. The results suggest that combinatorial tagging of transcripts by specific RNA-binding proteins may be a general mechanism for coordinated control of the localization, translation, and decay of mRNAs and thus an integral part of the global gene expression program.

3' Untranslated Regions↗

Clinical Utility of Next-Generation Sequencing in Tumors Diagnosed as Lung Squamous Cell Carcinoma: Real-World Data of Diagnostic and Therapeutic Implications.

Lung squamous cell carcinoma (LUSC) is the second most common subtype of non-small cell lung carcinoma (NSCLC), typically associated with a poor prognosis. Unlike lung adenocarcinoma, the application of next-generation sequencing (NGS) in LUSC has lagged because of the long-standing perception of low therapeutic yield, primarily based on highly selected, resected cohorts. We sought to determine the real-world clinical utility of NGS in LUSC. We analyzed an institutional cohort of 576 tumors initially diagnosed as LUSC that underwent NGS profiling. We defined "clinical yield" as either diagnostic reclassification or the identification of a targetable mitogenic alteration. Twenty cases (3.5%) were reclassified, including rediagnosis to cutaneous squamous cell carcinoma, transformed adenocarcinoma (post targeted therapy), and rare entities such as nuclear protein of the testis-rearranged carcinoma and lymphoepithelial carcinoma. Primary mitogenic drivers were identified in 83 cases (14.4% of the total cohort), of which 43 (7.5% of the total cohort) harbored alterations with currently Food and Drug Administration-approved therapies for NSCLC (including KRAS, EGFR, MET, ALK, and ROS1). Overall clinical yield-defined as the sum of diagnostic reclassifications and identification of NSCLC-specific targetable alterations-was 11.0% (63/576). Univariate and multivariate analysis demonstrated that never or light smoking history was the strongest independent predictor of clinical yield, with 57.3% of tumors in this subset being reclassified or harboring a strong driver. Our findings demonstrate that NGS provides significant diagnostic and therapeutic value in a real-world LUSC cohort, challenging the historical premise of low yield. Although clinicodemographic features can help prioritize testing in resource-limited settings, the identification of targetable drivers across all smoking groups supports the universal application of comprehensive NGS for all patients diagnosed with LUSC.

Humans↗

From glycomics to functional glycomics of sugar chains: Identification of target proteins with functional changes using gene targeting mice and knock down cells of FUT8 as examples.

Comprehensive analyses of proteins from cells and tissues are the most effective means of elucidating the expression patterns of individual disease-related proteins. On the other hand, the simultaneous separation and characterization of proteins by 1-DE or 2-DE followed by MS analysis are one of the fundamental approaches to proteomic analysis. However, these analyses do not permit the complete structural identification of glycans in glycoproteins or their structural characterization. Over half of all known proteins are glycosylated and glycan analyses of glycoproteins are requisite for fundamental proteomics studies. The analysis of glycan structural alterations in glycoproteins is becoming increasingly important in terms of biomarkers, quality control of glycoprotein drugs, and the development of new drugs. However, usual approach such as proteoglycomics, glycoproteomics and glycomics which characterizes and/or identifies sugar chains, provides some structural information, but it does not provide any information of functionality of sugar chains. Therefore, in order to elucidate the function of glycans, functional glycomics which identifies the target glycoproteins and characterizes functional roles of sugar chains represents a promising approach. In this review, we show examples of functional glycomics technique using alpha 1,6 fucosyltransferase gene (Fut8) in order to identify the target glycoprotein(s). This approach is based on glycan profiling by CE/MS and LC/MS followed by proteomic approaches, including 2-DE/1-DE and lectin blot techniques and identification of functional changes of sugar chains.

Animals↗

A comprehensive catalog of human KRAB-associated zinc finger genes: insights into the evolutionary history of a large family of transcriptional repressors.

Krüppel-type zinc finger (ZNF) motifs are prevalent components of transcription factor proteins in all eukaryotes. KRAB-ZNF proteins, in which a potent repressor domain is attached to a tandem array of DNA-binding zinc-finger motifs, are specific to tetrapod vertebrates and represent the largest class of ZNF proteins in mammals. To define the full repertoire of human KRAB-ZNF proteins, we searched the genome sequence for key motifs and then constructed and manually curated gene models incorporating those sequences. The resulting gene catalog contains 423 KRAB-ZNF protein-coding loci, yielding alternative transcripts that altogether predict at least 742 structurally distinct proteins. Active rounds of segmental duplication, involving single genes or larger regions and including both tandem and distributed duplication events, have driven the expansion of this mammalian gene family. Comparisons between the human genes and ZNF loci mined from the draft mouse, dog, and chimpanzee genomes not only identified 103 KRAB-ZNF genes that are conserved in mammals but also highlighted a substantial level of lineage-specific change; at least 136 KRAB-ZNF coding genes are primate specific, including many recent duplicates. KRAB-ZNF genes are widely expressed and clustered genes are typically not coregulated, indicating that paralogs have evolved to fill roles in many different biological processes. To facilitate further study, we have developed a Web-based public resource with access to gene models, sequences, and other data, including visualization tools to provide genomic context and interaction with other public data sets.

Computational Biology↗

The transcriptome and its translation during recovery from cell cycle arrest in Saccharomyces cerevisiae.

Complete genome sequences together with high throughput technologies have made comprehensive characterizations of gene expression patterns possible. While genome-wide measurement of mRNA levels was one of the first applications of these advances, other important aspects of gene expression are also amenable to a genomic approach, for example, the translation of message into protein. Earlier we reported a high throughput technology for simultaneously studying mRNA level and translation, which we termed translation state array analysis, or TSAA. The current studies test the proposition that TSAA can identify novel instances of translation regulation at the genome-wide level. As a biological model, cultures of Saccharomyces cerevisiae were cell cycle-arrested using either alpha-factor or the temperature-sensitive cdc15-2 allele. Forty-eight mRNAs were found to change significantly in translation state following release from alpha-factor arrest, including genes involved in pheromone response and cell cycle arrest such as BAR1, SST2, and FAR1. After the shift of the cdc15-2 strain from 37 degrees C to 25 degrees C, 54 mRNAs were altered in translation state, including the products of the stress genes HSP82, HSC82, and SSA2. Thus, regulation at the translational level seems to play a significant role in the response of yeast cells to external physical or biological cues. In contrast, surprisingly few genes were found to be translationally controlled as cells progressed through the cell cycle. Additional refinements of TSAA should allow characterization of both transcriptional and translational regulatory networks on a genomic scale, providing an additional layer of information that can be integrated into models of system biology and function.

Cell Cycle↗

An expression-driven approach to the prediction of carbohydrate transport and utilization regulons in the hyperthermophilic bacterium Thermotoga maritima.

Comprehensive analysis of genome-wide expression patterns during growth of the hyperthermophilic bacterium Thermotoga maritima on 14 monosaccharide and polysaccharide substrates was undertaken with the goal of proposing carbohydrate specificities for transport systems and putative transcriptional regulators. Saccharide-induced regulons were predicted through the complementary use of comparative genomics, mixed-model analysis of genome-wide microarray expression data, and examination of upstream sequence patterns. The results indicate that T. maritima relies extensively on ABC transporters for carbohydrate uptake, many of which are likely controlled by local regulators responsive to either the transport substrate or a key metabolic degradation product. Roles in uptake of specific carbohydrates were suggested for members of the expanded Opp/Dpp family of ABC transporters. In this family, phylogenetic relationships among transport systems revealed patterns of possible duplication and divergence as a strategy for the evolution of new uptake capabilities. The presence of GC-rich hairpin sequences between substrate-binding proteins and other components of Opp/Dpp family transporters offers a possible explanation for differential regulation of transporter subunit genes. Numerous improvements to T. maritima genome annotations were proposed, including the identification of ABC transport systems originally annotated as oligopeptide transporters as candidate transporters for rhamnose, xylose, beta-xylan, and beta-glucans and identification of genes likely to encode proteins missing from current annotations of the pentose phosphate pathway. Beyond the information obtained for T. maritima, the present study illustrates how expression-based strategies can be used for improving genome annotation in other microorganisms, especially those for which genetic systems are unavailable.

5' Flanking Region↗

An expressed sequence tag (EST) data mining strategy succeeding in the discovery of new G-protein coupled receptors.

We have developed a comprehensive expressed sequence tag database search method and used it for the identification of new members of the G-protein coupled receptor superfamily. Our approach proved to be especially useful for the detection of expressed sequence tag sequences that do not encode conserved parts of a protein, making it an ideal tool for the identification of members of divergent protein families or of protein parts without conserved domain structures in the expressed sequence tag database. At least 14 of the expressed sequence tags found with this strategy are promising candidates for new putative G-protein coupled receptors. Here, we describe the sequence and expression analysis of five new members of this receptor superfamily, namely GPR84, GPR86, GPR87, GPR90 and GPR91. We also studied the genomic structure and chromosomal localization of the respective genes applying in silico methods. A cluster of six closely related G-protein coupled receptors was found on the human chromosome 3q24-3q25. It consists of four orphan receptors (GPR86, GPR87, GPR91, and H963), the purinergic receptor P2Y1, and the uridine 5'-diphosphoglucose receptor KIAA0001. It seems likely that these receptors evolved from a common ancestor and therefore might have related ligands. In conclusion, we describe a data mining procedure that proved to be useful for the identification and first characterization of new genes and is well applicable for other gene families.

Amino Acid Motifs↗