Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Cell annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Statistical resynchronization and Bayesian detection of periodically expressed genes.

We propose a periodic-normal mixture (PNM) model to fit transcription profiles of periodically expressed (PE) genes in cell cycle microarray experiments. The model leads to a principled statistical estimation procedure that produces more accurate estimates of the mean cell cycle length and the gene expression periodicity than existing heuristic approaches. A central component of the proposed procedure is the resynchronization of the observed transcription profile of each PE gene according to the PNM with estimated periodicity parameters. By using a two-component mixture-Beta model to approximate the PNM fitting residuals, we employ an empirical Bayes method to detect PE genes. We estimate that about one-third of the genes in the genome of Saccharomyces cerevisiae are likely to be transcribed periodically, and identify 822 genes whose posterior probabilities of being PE are greater than 0.95. Among these 822 genes, 540 are also in the list of 800 genes detected by Spellman. Gene ontology annotation analysis shows that many of the 822 genes were involved in important cell cycle-related processes, functions and components. When matching the 822 resynchronized expression profiles of three independent experiments, little phase shifts were observed, indicating that the three synchronization methods might have brought cells to the same phase at the time of release.

Bayes Theorem↗

Combining evidence, biomedical literature and statistical dependence: new insights for functional annotation of gene sets.

BACKGROUND: Large-scale genomic studies based on transcriptome technologies provide clusters of genes that need to be functionally annotated. The Gene Ontology (GO) implements a controlled vocabulary organised into three hierarchies: cellular components, molecular functions and biological processes. This terminology allows a coherent and consistent description of the knowledge about gene functions. The GO terms related to genes come primarily from semi-automatic annotations made by trained biologists (annotation based on evidence) or text-mining of the published scientific literature (literature profiling). RESULTS: We report an original functional annotation method based on a combination of evidence and literature that overcomes the weaknesses and the limitations of each approach. It relies on the Gene Ontology Annotation database (GOA Human) and the PubGene biomedical literature index. We support these annotations with statistically associated GO terms and retrieve associative relations across the three GO hierarchies to emphasise the major pathways involved by a gene cluster. Both annotation methods and associative relations were quantitatively evaluated with a reference set of 7397 genes and a multi-cluster study of 14 clusters. We also validated the biological appropriateness of our hybrid method with the annotation of a single gene (cdc2) and that of a down-regulated cluster of 37 genes identified by a transcriptome study of an in vitro enterocyte differentiation model (CaCo-2 cells). CONCLUSION: The combination of both approaches is more informative than either separate approach: literature mining can enrich an annotation based only on evidence. Text-mining of the literature can also find valuable associated MEDLINE references that confirm the relevance of the annotation. Eventually, GO terms networks can be built with associative relations in order to highlight cooperative and competitive pathways and their connected molecular functions.

Algorithms↗

Integrated multi-omics identification of m6A-SNP-related diagnostic biomarkers in amyotrophic lateral sclerosis.

BACKGROUND: Amyotrophic lateral sclerosis (ALS) lacks reliable and minimally invasive biomarkers for early diagnosis. m6A-associated single-nucleotide polymorphisms (m6A-SNPs) may influence RNA methylation and gene expression, offering opportunities to identify clinically relevant diagnostic markers. METHODS: We integrated eQTLGen cis-eQTL data, RMVar m6A-SNP annotations, and ALS transcriptomic datasets to identify m6A-SNP-related genes. Random Forest and LASSO regression were combined to screen robust diagnostic markers. A nomogram was constructed and validated using independent cohorts. Immune infiltration, predicted m6A modification sites, and potential RBP-SNP interactions were assessed. Peripheral blood samples from ALS patients were used for exploratory validation of gene expression and global m6A levels. RESULTS: We identified 109 ALS-associated m6A-SNP-related genes with cis-eQTL signals and narrowed these to seven candidate diagnostic markers (TMED5, OXR1, BRI3, FEM1C, SUZ12, EIF2AK4, and TJAP1). The seven-gene model outperformed the individual markers in the training cohort and retained moderate discrimination in the independent validation cohort. ALS samples showed differences in inferred immune-cell composition, including monocytes, neutrophils, and T-cell subsets. The selected SNP loci were located near predicted m6A sites and annotated RBP-binding regions. Exploratory clinical validation showed significant upregulation of FEM1C and SUZ12 at both mRNA and protein levels, accompanied by reduced global m6A modification. CONCLUSIONS: Through multi-omics integration and exploratory clinical validation, this study identifies m6A-SNP-related candidate markers associated with ALS. The findings support further evaluation of m6A-related signatures for ALS discrimination and molecular characterization, while larger independent cohorts and additional calibration are required before clinical application.

Humans↗

Functional cloning, sorting, and expression profiling of nucleic acid-binding proteins.

A major challenge in the post-sequencing era is to elucidate the activity and biological function of genes that reside in the human genome. An important subset includes genes that encode proteins that regulate gene expression or maintain the structural integrity of the genome. Using a novel oligonucleotide-binding substrate as bait, we show the feasibility of a modified functional expression-cloning strategy to identify human cDNAs that encode a spectrum of nucleic acid-binding proteins (NBPs). Approximately 170 cDNAs were identified from screening phage libraries derived from a human colorectal adenocarcinoma cell line and from noncancerous fetal lung tissue. Sequence analysis confirmed that virtually every clone contained a known DNA- or RNA-binding motif. We also report on a complementary sorting strategy that, in the absence of subcloning and protein purification, can distinguish different classes of NBPs according to their particular binding properties. To extend our functional annotation of NBPs, we have used GeneChip expression profiling of 14 different breast-derived cell lines to examine the relative transcriptional activity of genes identified in our screen and cluster analysis to discover other genes that have similar expression patterns. Finally, we present strategies to analyze the upstream regulatory region of each gene within a cluster group and select unique combinations of transcription factor binding sites that may be responsible for dictating the observed synexpression.

Adenocarcinoma↗

The current and future perspective of ChickenGTEx project and its applications in precision breeding.

The Chicken Genotype-Tissue Expression (ChickenGTEx) project was established to systematically characterize the regulatory landscape of the chicken genome and to accelerate the translation of functional genomics into precision breeding. By integrating whole-genome sequencing with multi-tissue transcriptomic profiling, ChickenGTEx provides a comprehensive atlas of gene expression regulation across diverse tissues and physiological systems. Current findings demonstrate that complex production traits are governed by coordinated regulatory networks rather than isolated loci, with substantial contributions from tissue-specific gene expression, structural variation, and genotype-by-sex interactions. Sex-dependent regulatory effects further refine the genetic architecture of metabolic, immune, and reproductive traits, highlighting the importance of incorporating sex as a biological variable in genomic analyses. Application of integrative omics frameworks within elite layer populations has revealed multilayer regulatory mechanisms underlying extended laying performance, feed efficiency, metabolic health, and eggshell quality. By partitioning phenotypic variance into genetic, regulatory, and host-microbiome components, these approaches move beyond association-based mapping toward causal inference and biological interpretation. Importantly, validated regulatory loci identified through ChickenGTEx and related analyses provide actionable markers for genomic selection and rational targets for precision genome modification. Looking forward, continued expansion of regulatory atlases, incorporation of single-cell and longitudinal data in diverse environmental conditions, and integration of functional annotation into breeding pipelines will further enhance prediction accuracy and sustainable genetic improvement. The ChickenGTEx project thus represents a foundational platform bridging functional genomics and practical poultry breeding.

Animals↗

Genome-wide protein interaction maps using two-hybrid systems.

Automated sequence technology has rendered functional biology amenable to genomic scale analysis. Among genome-wide exploratory approaches, the two-hybrid system in yeast (Y2H) has outranked other techniques because it is the system of choice to detect protein-protein interactions. Deciphering the cascade of binding events in a whole cell helps define signal transduction and metabolic pathways or enzymatic complexes. The function of proteins is eventually attributed through whole cell protein interaction maps where totally unknown proteins are partnered with fully annotated proteins belonging to the same functional category. Since its first description in the late 1980's, several versions of the Y2H have been developed in order to overcome the major limitations of the system, namely false positives and false negatives. Optimized versions have been recently applied at multi-molecular and genomic scale. These genome-wide surveys can be methodologically divided into two types of approaches: one either tests combinations of predefined polypeptides (the so-called matrix approach) using various short-cuts to speed up the process, or one screens with a given polypeptide (bait) for potential partners (preys) present in complex libraries of genomic or complementary DNA (library screening). In the former strategy, one tests what one knows, for example pair-wise interactions between full-length open reading frames from recently sequenced and annotated genomes. Although based on a one-by-one scheme, this method is reported to be amenable to large-scale genomics thanks to multicloning strategies and to the use of small robotics workstations. In the latter, highly complex cDNA or genomic libraries of protein domains can be screened to saturation with high-throughput screening systems allowing the discovery of yet unidentified proteins. Both approaches have strengths and drawbacks that will be discussed here. None yields a full proteome-wide screening since certain proteins (e.g. some transcription factors) are not usable in Y2H. Novel two-hybrid assays have been recently described in bacteria. Applications of these time- and cost-effective assays to genomic screening will be discussed and compared to the Y2H technology.

Animals↗

Serial analysis of gene expression (SAGE) in Plasmodium falciparum: application of the technique to A-T rich genomes.

The advent of high-throughput methods for the analysis of global gene expression, together with the Malaria Genome Project open up new opportunities for furthering our understanding of the fundamental biology and virulence of the malaria parasite. Serial analysis of gene expression (SAGE) is particularly well suited for malarial systems, as the genomes of Plasmodium species remain to be fully annotated. By simultaneously and quantitatively analyzing mRNA transcript profiles from a given cell population, SAGE allows for the discovery of new genes. In this study, one reports the successful application of SAGE in Plasmodium falciparum, 3D7 strain parasites, from which a preliminary library of 6880 tags corresponding to 4146 different genes was generated. It was demonstrated that P. falciparum is amenable to this technique, despite the remarkably high A-T content of its genome. SAGE tags as short as 10 nucleotides were sufficient to uniquely identify parasite transcripts from both nuclear and mitochondrial genomes. Moreover, the skewed A-T content of parasite sequence did not preclude the use of enzymes that are crucial for generating representative SAGE libraries. Finally, a few modifications to DNA extraction and cloning steps of the SAGE protocol proved useful for circumventing specific problems presented by A-T rich genomes.

Animals↗

Protein sequencing by mass analysis of polypeptide ladders after controlled protein hydrolysis.

The characterization of protein modifications is essential for the study of protein function using functional genomic and proteomic approaches. However, current techniques are not efficient in determining protein modifications. We report an approach for sequencing proteins and determining modifications with high speed, sensitivity and specificity. We discovered that a protein could be readily acid-hydrolyzed within 1 min by exposure to microwave irradiation to form, predominantly, two series of polypeptide ladders containing either the N- or C-terminal amino acid of the protein, respectively. Mass spectrometric analysis of the hydrolysate produced a simple mass spectrum consisting of peaks exclusively from these polypeptide ladders, allowing direct reading of amino acid sequence and modifications of the protein. As examples, we applied this technique to determine protein phosphorylation sites as well as the sequences and several previously unknown modifications of 28 small proteins isolated from Escherichia coli K12 cells. This technique can potentially be automated for large-scale protein annotation.

Algorithms↗

Molecular characterization of mouse gastric epithelial progenitor cells.

The adult mouse gastric epithelium undergoes continuous renewal in discrete anatomic units. Lineage tracing studies have previously disclosed the morphologic features of gastric epithelial lineage progenitors (GEPs), including those of the presumptive multipotent stem cell. However, their molecular features have not been defined. Here, we present the results of an analysis of genes and pathways expressed in these cells. One hundred forty-seven transcripts enriched in GEPs were identified using an approach that did not require physical disruption of the stem cell niche. Real-time quantitative RT-PCR studies of laser capture microdissected cells retrieved from this niche confirmed enriched expression of a selected set of genes from the GEP list. An algorithm that allows quantitative comparisons of the functional relatedness of automatically annotated expression profiles showed that the GEP profile is similar to a dataset of genes that defines mouse hematopoietic stem cells, and distinct from the profiles of two differentiated GEP descendant lineages (parietal and zymogenic cell). Overall, our analysis revealed that growth factor response pathways are prominent in GEPs, with insulin-like growth factor appearing to play a key role. A substantial fraction of GEP transcripts encode products required for mRNA processing and cytoplasmic localization, including numerous homologs of Drosophila genes (e.g., Y14, staufen, mago nashi) needed for axis formation during oogenesis. mRNA targeting proteins may help these epithelial progenitors establish differential communications with neighboring cells in their niche.

Adenosine Triphosphate↗

A systematic approach to reconstructing transcription networks in Saccharomycescerevisiae.

Decomposing regulatory networks into functional modules is a first step toward deciphering the logical structure of complex networks. We propose a systematic approach to reconstructing transcription modules (defined by a transcription factor and its target genes) and identifying conditionsperturbations under which a particular transcription module is activateddeactivated. Our approach integrates information from regulatory sequences, genome-wide mRNA expression data, and functional annotation. We systematically analyzed gene expression profiling experiments in which the yeast cell was subjected to various environmental or genetic perturbations. We were able to construct transcription modules with high specificity and sensitivity for many transcription factors, and predict the activation of these modules under anticipated as well as unexpected conditions. These findings generate testable hypotheses when combined with existing knowledge on signaling pathways and protein-protein interactions. Correlating the activation of a module to a specific perturbation predicts links in the cell's regulatory networks, and examining coactivated modules suggests specific instances of crosstalk between regulatory pathways.

Gene Expression Profiling↗

Assigning functions to genes: identification of S-phase expressed genes in Leishmania major based on post-transcriptional control elements.

Assigning functions to genes is one of the major challenges of the post-genomic era. Usually, functions are assigned based on similarity of the coding sequences to sequences of known genes, or by identification of transcriptional cis-regulatory elements that are known to be associated with specific pathways or conditions. In trypanosomatids, where regulation of gene expression takes place mainly at the post-transcriptional level, new approaches for function assignment are needed. Here we demonstrate the identification of novel S-phase expressed genes in Leishmania major, based on a post-transcriptional control element that was recognized in Crithidia fasciculata as involved in the cell cycle-dependent expression of several nuclear and mitochondrial S-phase expressed genes. Hypothesizing that a similar regulatory mechanism is manifested in L.major, we have applied a computational search for similar control elements in the genome of L.major. Our computational scan yielded 132 genes, of which 33% are homologues of known DNA metabolism genes and 63% lack any annotation. Experimental testing of seven of these genes revealed that their mRNAs cycle throughout the cell cycle, reaching a maximum level during S-phase or just prior to it. It is suggested that screening for post-transcriptional control elements associated with a specific function provides an efficient method for assigning functions to trypanosomatid genes.

Animals↗

Extensive expansion of the claudin gene family in the teleost fish, Fugu rubripes.

In humans, the claudin superfamily consists of 19 homologous proteins that commonly localize to tight junctions of epithelial and endothelial cells. Besides being structural tight-junction components, claudins participate in cell-cell adhesion and the paracellular transport of solutes. Here, we identify and annotate the claudin genes in the whole-genome of the teleost fish, Fugu rubripes (Fugu), and determine their phylogenetic relationships to those in mammals. Our analysis reveals extensive gene duplications in the teleost lineage, leading to 56 claudin genes in Fugu. A total of 35 Fugu claudin genes can be assigned orthology to 17 mammalian claudin genes, with the remaining 21 genes being specific to the fish lineage. Thus, a significant number of the additional Fugu genes are not the result of the proposed whole-genome duplication in the fish lineage. Expression profiling shows that most of the 56 Fugu claudin genes are expressed in a more-or-less tissue-specific fashion, or at particular developmental stages. We postulate that the expansion of the claudin gene family in teleosts allowed the acquisition of novel functions during evolution, and that fish-specific novel members of gene families such as claudins contribute to a large extent to the distinct physiology of fishes and mammals.

Animals↗

Comparative profiling of the sense and antisense transcriptome of maize lines.

BACKGROUND: There are thousands of maize lines with distinctive normal as well as mutant phenotypes. To determine the validity of comparisons among mutants in different lines, we first address the question of how similar the transcriptomes are in three standard lines at four developmental stages. RESULTS: Four tissues (leaves, 1 mm anthers, 1.5 mm anthers, pollen) from one hybrid and one inbred maize line were hybridized with the W23 inbred on Agilent oligonucleotide microarrays with 21,000 elements. Tissue-specific gene expression patterns were documented, with leaves having the most tissue-specific transcripts. Haploid pollen expresses about half as many genes as the other samples. High overlap of gene expression was found between leaves and anthers. Anther and pollen transcript expression showed high conservation among the three lines while leaves had more divergence. Antisense transcripts represented about 6 to 14 percent of total transcriptome by tissue type but were similar across lines. Gene Ontology (GO) annotations were assigned and tabulated. Enrichment in GO terms related to cell-cycle functions was found for the identified antisense transcripts. Microarray results were validated via quantitative real-time PCR and by hybridization to a second oligonucleotide microarray platform. CONCLUSION: Despite high polymorphisms and structural differences among maize inbred lines, the transcriptomes of the three lines displayed remarkable similarities, especially in both reproductive samples (anther and pollen). We also identified potential stage markers for maize anther development. A large number of antisense transcripts were detected and implicated in important biological functions given the enrichment of particular GO classes.

Expressed Sequence Tags↗

Prospective molecular profiling of melanoma metastases suggests classifiers of immune responsiveness.

We amplified RNAs from 63 fine needle aspiration (FNA) samples from 37 s.c. melanoma metastases from 25 patients undergoing immunotherapy for hybridization to a 6108-gene human cDNA chip. By prospectively following the history of the lesions, we could correlate transcript patterns with clinical outcome. Cluster analysis revealed a tight relationship among autologous synchronously sampled tumors compared with unrelated lesions (average Pearson's r = 0.83 and 0.7, respectively, P < 0.0003). As reported previously, two subgroups of metastatic melanoma lesions were identified that, however, had no predictive correlation with clinical outcome. Ranking of gene expression data from pretreatment samples identified approximately 30 genes predictive of clinical response (P < 0.001). Analysis of their annotations denoted that approximately half of them were related to T-cell regulation, suggesting that immune responsiveness might be predetermined by a tumor microenvironment conducive to immune recognition.

Adult↗

Gene expression profiling of leukemic cell lines reveals conserved molecular signatures among subtypes with specific genetic aberrations.

Hematologic malignancies are characterized by fusion genes of biological/clinical importance. Immortalized cell lines with such aberrations are today widely used to model different aspects of leukemogenesis. Using cDNA microarrays, we determined the gene expression profiles of 40 cell lines as well as of primary leukemias harboring 11q23/MLL rearrangements, t(1;19)[TCF3/PBX1], t(12;21)[ETV6/RUNX1], t(8;21)[RUNX1/CBFA2T1], t(8;14)[IGH@/MYC], t(8;14)[TRA@/MYC], t(9;22)[BCR/ABL1], t(10;11)[PICALM/MLLT10], t(15;17)[PML/RARA], or inv(16)[CBFB/MYH11]. Unsupervised classification revealed that hematopoietic cell lines of diverse origin, but with the same primary genetic changes, segregated together, suggesting that pathogenetically important regulatory networks remain conserved despite numerous passages. Moreover, primary leukemias cosegregated with cell lines carrying identical genetic rearrangements, further supporting that critical regulatory pathways remain intact in hematopoietic cell lines. Transcriptional signatures correlating with clinical subtypes/primary genetic changes were identified and annotated based on their biological/molecular properties and chromosomal localization. Furthermore, the expression profile of tyrosine kinase-encoding genes was investigated, identifying several differentially expressed members, segregating with primary genetic changes, which may be targeted with tyrosine kinase inhibitors. The identified conserved signatures are likely to reflect regulatory networks of importance for the transforming abilities of the primary genetic changes and offer important pathogenetic insights as well as a number of targets for future rational drug design.

Acute Disease↗

Genome-scale functional profiling of the mammalian AP-1 signaling pathway.

Large-scale functional genomics approaches are fundamental to the characterization of mammalian transcriptomes annotated by genome sequencing projects. Although current high-throughput strategies systematically survey either transcriptional or biochemical networks, analogous genome-scale investigations that analyze gene function in mammalian cells have yet to be fully realized. Through transient overexpression analysis, we describe the parallel interrogation of approximately 20,000 sequence annotated genes in cancer-related signaling pathways. For experimental validation of these genome data, we apply an integrative strategy to characterize previously unreported effectors of activator protein-1 (AP-1) mediated growth and mitogenic response pathways. These studies identify the ADP-ribosylation factor GTPase-activating protein Centaurin alpha1 and a Tudor domain-containing hypothetical protein as putative AP-1 regulatory oncogenes. These results provide insight into the composition of the AP-1 signaling machinery and validate this approach as a tractable platform for genome-wide functional analysis.

Animals↗

IMGT, the International ImMunoGeneTics database.

IMGT, the international ImMunoGeneTics database, is an integrated database specialising in Immunoglobulins (Ig), T cell Receptors (TcR) and Major Histocompatibility Complex (MHC) of all vertebrate species, created by Marie-Paule Lefranc, CNRS, Montpellier II University, Montpellier, France (lefranc@ligm.crbm.cnrs-mop.fr). IMGT includes three databases: LIGM-DB (for Ig and TcR), MHC/HLA-DB and PRIMER-DB (the last two in development). IMGT comprises expertly annotated sequences and alignment tables. LIGM-DB contains more than 23 000 Immunoglobulin and T cell Receptor sequences from 78 species. MHC/HLA-DB contains Class I and Class II Human Leucocyte Antigen alignment tables. An IMGT tool, DNAPLOT, developed for Ig, TcR and MHC sequence alignments, is also available. IMGT works in close collaboration with the EMBL database. IMGT goals are to establish a common data access to all immunogenetics data, including nucleotide and protein sequences, oligonucleotide primers, gene maps and other genetic data of Ig, TcR and MHC molecules, and to provide a graphical user friendly data access. IMGT has important implications in medical research (repertoire in autoimmune diseases, AIDS, leukemias, lymphomas), therapeutical approaches (antibody engineering), genome diversity and genome evolution studies. IMGT is freely available at http://imgt.cnusc.fr:8104

Amino Acid Sequence↗

A network-based analysis of allergen-challenged CD4+ T cells from patients with allergic rhinitis.

We performed a network-based analysis of DNA microarray data from allergen-challenged CD4(+) T cells from patients with seasonal allergic rhinitis. Differentially expressed genes were organized into a functionally annotated network using the Ingenuity Knowledge Database, which is based on manual review of more than 200,000 publications. The main function of this network is the regulation of lymphocyte apoptosis, a role associated with several genes of the tuber necrosis factor superfamily. The expression of TNFRSF4, one of the genes in this family, was found to be 48 times higher in allergen-challenged cells than in diluent-challenged cells. TNFRSF4 is known to inhibit apoptosis and to enhance Th2 proliferation. Examination of a different material of allergen-stimulated peripheral blood mononuclear cells showed a higher number of interleukin-4(+) type 2 CD4(+) T (Th2) cells in patients than in controls (P<0.01), as well as a higher number of non-apoptotic Th2 cells in patients (P<0.01). The number of Th2 cells expressing TNFRSF4, TNFSF7 and TNFRSF1B was also significantly higher in patients. Treatment with anti-TNFSF4 resulted in a significantly decreased number of Th2 cells (P<0.05). A logical inference from all this is that the proliferation of allergen-challenged Th2 cells is associated with a decreased apoptosis of Th2 cells and an increase in TNFRSF4 signalling.

Adolescent↗