Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Cell annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Long homopurine*homopyrimidine sequences are characteristic of genes expressed in brain and the pseudoautosomal region.

Homo(purine*pyrimidine) sequences (R*Y tracts) with mirror repeat symmetries form stable triplexes that block replication and transcription and promote genetic rearrangements. A systematic search was conducted to map the location of the longest R*Y tracts in the human genome in order to assess their potential function(s). The 814 R*Y tracts with > or =250 uninterrupted base pairs were preferentially clustered in the pseudoautosomal region of the sex chromosomes and located in the introns of 228 annotated genes whose protein products were associated with functions at the cell membrane. These genes were highly expressed in the brain and particularly in genes associated with susceptibility to mental disorders, such as schizophrenia. The set of 1957 genes harboring the 2886 R*Y tracts with > or =100 uninterrupted base pairs was additionally enriched in proteins associated with phosphorylation, signal transduction, development and morphogenesis. Comparisons of the > or =250 bp R*Y tracts in the mouse and chimpanzee genomes indicated that these sequences have mutated faster than the surrounding regions and are longer in humans than in chimpanzees. These results support a role for long R*Y tracts in promoting recombination and genome diversity during evolution through destabilization of chromosomal DNA, thereby inducing repair and mutation.

Animals↗

Immunoinformatics--the new kid in town.

The astounding diversity of immune system components (e.g. immunoglobulins, lymphocyte receptors, or cytokines) together with the complexity of the regulatory pathways and network-type interactions makes im munology a combinatorial science. Currently available data represent only a tiny fraction of possible situations and data continues to accrue at an exponential rate. Computational analysis has therefore become an essential element of immunology research with a main role of immunoinformatics being the management and analysis of immunological data. More advanced analyses of the immune system using computational models typically involve conversion of an immunological question to a computational problem, followed by solving of the computational problem and translation of these results into biologically meaningful answers. Major immunoinformatics developments include immunological databases, sequence analysis, structure modelling, mathematical modelling of the immune system, simulation of laboratory experiments, statistical support for immunological experimentation and immunogenomics. In this paper we describe the status and challenges within these sub-fields. We foresee the emergence of immunomics not only as a collective endeavour by researchers to decipher the sequences of T cell receptors, immunoglobulins, and other immune receptors, but also to functionally annotate the capacity of the immune system to interact with the whole array of selfand non-self entities, including genome-to-genome interactions.

Allergy and Immunology↗

SOX4 expression in bladder carcinoma: clinical aspects and in vitro functional characterization.

The human transcription factor SOX4 was 5-fold up-regulated in bladder tumors compared with normal tissue based on whole-genome expression profiling of 166 clinical bladder tumor samples and 27 normal urothelium samples. Using a SOX4-specific antibody, we found that the cancer cells expressed the SOX4 protein and, thus, did an evaluation of SOX4 protein expression in 2,360 bladder tumors using a tissue microarray with clinical annotation. We found a correlation (P < 0.05) between strong SOX4 expression and increased patient survival. When overexpressed in the bladder cell line HU609, SOX4 strongly impaired cell viability and promoted apoptosis. To characterize downstream target genes and SOX4-induced pathways, we used a time-course global expression study of the overexpressed SOX4. Analysis of the microarray data showed 130 novel SOX4-related genes, some involved in signal transduction (MAP2K5), angiogenesis (NRP2), and cell cycle arrest (PIK3R3) and others with unknown functions (CGI-62). Among the genes regulated by SOX4, 25 contained at least one SOX4-binding motif in the promoter sequence, suggesting a direct binding of SOX4. The gene set identified in vitro was analyzed in the clinical bladder material and a small subset of the genes showed a high correlation to SOX4 expression. The present data suggest a role of SOX4 in the bladder cancer disease.

Apoptosis↗

Gene expression profiling and analysis of signaling pathways involved in priming and differentiation of human neural stem cells.

Human neural stem cells have the ability to differentiate into all three major cell types in the CNS including neurons, astrocytes and oligodendrocytes. The multipotency of human neural stem cells shed a light on the possibility of using stem cells as a therapeutic tool for various neurological disorders including neurodegenerative diseases and neurotrauma that involve a loss of functional neurons. We have discovered previously a priming procedure to direct primarily cultured human neural stem cells to differentiate into almost pure neurons when grafted into adult CNS. However, the molecular mechanism underlying this phenomenon is still unknown. To unravel transcriptional changes of human neural stem cells upon priming, cDNA microarray was used to study temporal changes in human neural stem cell gene expression profile during priming and differentiation. As a result, transcriptional levels of 520 annotated genes were detected changed in at least at two time points during the priming process. In addition, transcription levels of more than 3000 hypothetical protein encoding genes and EST genes were modulated during the priming and differentiation processes of human neural stem cells. We further analyzed the named genes and grouped them into 14 functional categories. Of particular interest, key cell signal transduction pathways, including the G-protein-mediated signaling pathways (heterotrimeric and small monomeric GTPase pathways), the Wnt signaling pathway and the TGF-beta pathway, are modulated by the neural stem cell priming, suggesting important roles of these key signaling pathways in priming and differentiation of human neural stem cells.

Bone Morphogenetic Proteins↗

Crystallization and preliminary X-ray diffraction data of Mycobacterium tuberculosis FbpC1 (Rv3803c).

The heterotrimeric antigen 85 complex (Ag85) is a major component of the cell wall of Mycobacterium tuberculosis and consists of three abundantly secreted proteins (FbpA, FbpB and FbpC2). These play key roles in the pathogenesis of tuberculosis and in maintaining cell-wall integrity. A homologue of the Ag85 subunits ( approximately 40% identity) was recently annotated in the M. tuberculosis genome as FbpC1. Unlike the Ag85-complex components, FbpC1 lacks mycolyltransferase activity and its function remains to be established. In order to aid functional characterization, FbpC1 has been crystallized. At room temperature, tetragonal crystals of FbpC1 were obtained belonging to space group P4(1)2(1)2 (unit-cell parameters a = b = 109.9, c = 61.8 A), yet when frozen the crystals underwent a phase transition to orthorhombic symmetry, space group P2(1)2(1)2(1) (a = 59.9, b = 108.9, c = 109.9 A). Diffraction data complete to 1.7 A resolution were recorded at 100 K at the synchrotron.

Antigens, Bacterial↗

Cholesterol loading augments oxidative stress in macrophages.

To investigate the molecular consequence of loading free cholesterol into macrophages, we conducted a large-scale gene expression study to analyze acetylated-LDL-laden foam cells (AFC) and oxidized-LDL-laden foam cells (OFC) induced from human THP-1 cell lines. Cluster analysis was performed using 9600-gene microarray datasets from time course experiment. AFC and OFC shared common expression profiles; however, there were sufficient differences between these two treatments that AFC and OFC appealed as two separate entities. We identified 80 commonly upregulated genes and 48 commonly downregulated genes in AFC and OFC. Functional annotation of the differentially expressed genes indicated that apoptosis, extracellular matrix, oxidative stress, and cell proliferation was deregulated. We also identified 87 differentially expressed genes unique for AFC and 31 genes for OFC. The uniquely expressed genes of AFC are associated with kinase activity, ATP binding activity, and transporter activity, while unique genes for OFC are associated with cell signaling and adhesion. To validate the hypothesis that oxidative stress is a common feature for AFC and OFC, we performed a cluster analysis employing the genes related to oxidative stress, but we were unable to distinguish AFC from OFC in this manner. We performed real-time RT-PCR and ELISA on foam cells to examine the transcripts and secreted protein of interleukin 1 beta (IL1beta). IL1beta was rapidly induced in foam cells, but for AFC both RNA level and protein level dropped immediately and was attenuated. To detect levels of reactive oxygen species in foam cells we conducted hydroethidine staining and observed high levels of superoxide anion. We conclude that loading free cholesterol induces high levels of superoxide anion, increases oxidative stress, and triggers a transient inflammatory response in macrophages.

Cell Line↗

A large-scale, gene-driven mutagenesis approach for the functional analysis of the mouse genome.

A major challenge of the postgenomic era is the functional characterization of every single gene within the mammalian genome. In an effort to address this challenge, we assembled a collection of mutations in mouse embryonic stem (ES) cells, which is the largest publicly accessible collection of such mutations to date. Using four different gene-trap vectors, we generated 5,142 sequences adjacent to the gene-trap integration sites (gene-trap sequence tags; http://genetrap.de) from >11,000 ES cell clones. Although most of the gene-trap vector insertions occurred randomly throughout the genome, we found both vector-independent and vector-specific integration "hot spots." Because >50% of the hot spots were vector-specific, we conclude that the most effective way to saturate the mouse genome with gene-trap insertions is by using a combination of gene-trap vectors. When a random sample of gene-trap integrations was passaged to the germ line, 59% (17 of 29) produced an observable phenotype in transgenic mice, a frequency similar to that achieved by conventional gene targeting. Thus, gene trapping allows a large-scale and cost-effective production of ES cell clones with mutations distributed throughout the genome, a resource likely to accelerate genome annotation and the in vivo modeling of human disease.

Animals↗

Molecular characterization of prostatic small-cell neuroendocrine carcinoma.

OBJECTIVES: A subset of prostate carcinomas is composed predominantly, even exclusively, of neuroendocrine (NE) cells. In this report, we sought to characterize the gene expression profile of a prostate small cell NE carcinoma by assessing the diversity and abundance of transcripts in the LuCaP 49 prostate small cell carcinoma xenograft. METHODS: We constructed a cDNA library (PRCA3) from the LuCap 49 prostate small cell xenograft. Single pass DNA sequencing of randomly selected cDNA clones followed by sequence assembly and annotation produced a library of Expressed Sequence Tags (ESTs) representing the LuCaP 49 transcriptome. Comparative sequence analysis with ESTs derived from prostate adenocarcinoma libraries was performed using statistical algorithms designed to identify differentially expressed sequences. Putative NE cell-specific genes were further examined by Northern analysis. RESULTS: Sequence assembly and analysis identified 1,447 distinct genes expressed in the LuCaP 49 cDNA library. These include cDNAs encoding the NE markers secretogranin (SCG2), CD24, and ENO2. Northern analysis revealed that three additional genes, ASCL1, INA, and SV2B are expressed in LuCaP 49 but not in various prostate cancer cell lines or xenografts. Fifteen genes were identified with a statistical probability (P > 0.9) of being up-regulated in LuCaP 49 small cell carcinoma relative to prostate adenocarcinoma (two primary prostate adenocarcinomas and the LNCaP prostate adenocarcinoma cell line). CONCLUSIONS: Prostate small cell carcinoma expresses a diverse repertoire of genes that reflect characteristics of their NE cell of origin. ASCL1, INA, and SV2B are potential molecular markers for small cell NE tumors and NE cells of the prostate. This small cell NE carcinoma gene expression profile may yield insights into the development, progression, and treatment of subtypes of prostate cancer.

Aged↗

A catalogue of the effector secretome of plant pathogenic oomycetes.

The oomycetes form a phylogenetically distinct group of eukaryotic microorganisms that includes some of the most notorious pathogens of plants. Oomycetes accomplish parasitic colonization of plants by modulating host cell defenses through an array of disease effector proteins. The biology of effectors is poorly understood but tremendous progress has been made in recent years. This review classifies and catalogues the effector secretome of oomycetes. Two classes of effectors target distinct sites in the host plant: Apoplastic effectors are secreted into the plant extracellular space, and cytoplasmic effectors are translocated inside the plant cell, where they target different subcellular compartments. Considering that five species are undergoing genome sequencing and annotation, we are rapidly moving toward genome-wide catalogues of oomycete effectors. Already, it is evident that the effector secretome of pathogenic oomycetes is more complex than expected, with perhaps several hundred proteins dedicated to manipulating host cell structure and function.

Fungal Proteins↗

A checkpoint protein that scans the chromosome for damage at the start of sporulation in Bacillus subtilis.

In response to DNA damage, cells activate checkpoint signaling cascades to control cell-cycle progression and elicit DNA repair in order to maintain genomic integrity. The sensing and repair of lesions is critical for Bacillus subtilis cells entering the developmental process of sporulation as damaged DNA may prevent the cells from completing spore morphogenesis. We report the identification of the protein DisA (DNA integrity scanning protein, annotated YacK), which is required to delay the initiation of sporulation in response to chromosomal damage. DisA is a nonspecific DNA binding protein that forms a single focus, which moves rapidly within the bacterial cell, pausing at sites of DNA damage. We propose that the DisA focus scans along the chromosomes searching for lesions. Upon encountering a lesion, DisA delays entry into sporulation until the damage is repaired.

Bacillus subtilis↗

A comprehensive two-dimensional gel protein database of noncultured unfractionated normal human epidermal keratinocytes: towards an integrated approach to the study of cell proliferation, differentiation and skin diseases.

A two-dimensional (2-D) gel database of cellular proteins from noncultured, unfractionated normal human epidermal keratinocytes has been established. A total of 2651 [35S]methionine-labeled cellular proteins (1868 isoelectric focusing, 783 nonequilibrium pH gradient electrophoresis) were resolved and recorded using computer-aided 2-D gel electrophoresis. The protein numbers in this database differ from those reported in an earlier version due to changes in the scanning hardware (Celis et al., Electrophoresis 1990, 11, 242-254). Annotation categories reported include: "protein name" (listing 207 known proteins in alphabetical order), "basal cell markers", "differentiation markers", "proteins highly up-regulated in psoriatic skin", "microsequenced proteins" and "human autoantigens". For reference, we have also included 2-D gel (isoelectric focusing) patterns of cultured normal and psoriatic keratinocytes, melanocytes, fibroblasts, dermal microvascular endothelial cells, peripheral blood mononuclear cells and sweat duct cells. The keratinocyte 2-D gel protein database will be updated yearly in the November issue of Electrophoresis.

Biopsy↗

QCatch: a framework for quality control assessment and analysis of single-cell sequencing data.

MOTIVATION: Single-cell sequencing data analysis requires robust quality control (QC) to mitigate technical artifacts and ensure reliable downstream results. While tools like alevin-fry and simpleaf (and augmented execution context for the alevin-fry), offer flexibility and computational efficiency to process single-cell data, this ecosystem will further benefit from a standardized QC reporting tailored for its outputs. RESULTS: We introduce QCatch, a Python-based command-line tool that generates comprehensive and interactive HTML QC reports designed specifically for single-cell quantification results. Taking the output directory of alevin-fry or simpleaf as the input, QCatch is able to perform essential processing steps, like cell calling, and generate detailed QC reports that contain informative visualizations and statistics, including unique molecular identifier (UMI) count distributions, sequencing saturation estimates, and splicing status information, for QC assurance. Built for seamless integration into downstream analysis workflows, QCatch exports the processed results in a richly-annotated H5AD format file, a widely used data format common among many downstream single-cell data analysis tools. AVAILABILITY AND IMPLEMENTATION: The source code and documentation of QCatch are available on GitHub at https://github.com/COMBINE-lab/QCatch. QCatch can be installed via both Bioconda and PyPI.

Single-Cell Analysis↗

A question of size: the eukaryotic proteome and the problems in defining it.

We discuss the problems in defining the extent of the proteomes for completely sequenced eukaryotic organisms (i.e. the total number of protein-coding sequences), focusing on yeast, worm, fly and human. (i) Six years after completion of its genome sequence, the true size of the yeast proteome is still not defined. New small genes are still being discovered, and a large number of existing annotations are being called into question, with these questionable ORFs (qORFs) comprising up to one-fifth of the 'current' proteome. We discuss these in the context of an ideal genome-annotation strategy that considers the proteome as a rigorously defined subset of all possible coding sequences ('the orfome'). (ii) Despite the greater apparent complexity of the fly (more cells, more complex physiology, longer lifespan), the nematode worm appears to have more genes. To explain this, we compare the annotated proteomes of worm and fly, relating to both genome-annotation and genome evolution issues. (iii) The unexpectedly small size of the gene complement estimated for the complete human genome provoked much public debate about the nature of biological complexity. However, in the first instance, for the human genome, the relationship between gene number and proteome size is far from simple. We survey the current estimates for the numbers of human genes and, from this, we estimate a range for the size of the human proteome. The determination of this is substantially hampered by the unknown extent of the cohort of pseudogenes ('dead' genes), in combination with the prevalence of alternative splicing. (Further information relating to yeast is available at http://genecensus.org/yeast/orfome)

Animals↗

In silico analysis of SH3BP2 genomic alterations and expression profiles in CRC.

AIM: Colorectal cancer (CRC) is a widespread health issue that attains high mortality. The adaptor protein SH3BP2 amplification results in metabolic changes, oxidative stress, NK cell activity, and inflammation. The NK cells are capable of destroying tumor cells without prior activation, help prevent metastasis, and have prognostic value. Targeting SH3BP2 to regulate NK cell activity in the TME could enhance CRC-based immunotherapy. MATERIALS AND METHODS: The cancer hallmark tool helps in understanding SH3BP2&#xa0;hallmark annotation. Utilizing the STRING tool and the KEGG pathway, protein functional enrichment and PPI networking were analyzed. TIMER 2.0 was used for immune cell infiltration correlation analysis, and UALCAN was used for CPTAC-based protein expression profiling. RESULTS AND CONCLUSIONS: The GEO (GSE9348) dataset showed SH3BP2 is upregulated in CRC (log2 fold change&#x2009;=&#x2009;1.18). GEO, TCGA, and cBioPortal revealed SH3BP2 alterations in CRC cases, potentially aiding immune evasion. Mutations in SH3BP2 influence cancer growth, suppressing tumors or promoting them by activating NF-&#x3ba;B and affecting immune responses through WNT/&#x3b2;-catenin, PI3K, MAPK, and JAK-STAT pathways. Overall, SH3BP2 plays a key role in cancer growth and immune regulation, making it a promising target for CRC therapy. Further experimental validation is needed to demonstrate its diagnostic and therapeutic potency.

Humans↗

In search of differentially expressed genes and proteins.

A great challenge for modern cell biology is the successful examination of the co-expression of thousands of genes under physiological or pathological conditions and how the expression patterns define the different states of a single cell, tissue or a microorganism. Gene expression can be analyzed today on a large scale by advanced technical approaches for differential screening of proteins and mRNAs. The identification of differentially expressed mRNAs has been successfully applied to understand gene function and the underlying molecular mechanism(-s) of differentiation, development and disease state. Analysis of gene expression by the systematic mapping of thousands of proteins present in a cell or tissue can be achieved by the use of two-dimensional (2D) gel electrophoresis, quantitative computer image analysis, and protein identification techniques. In this article, we comment on some of these techniques and try to stress their advantages and drawbacks. We show how data from RNA/DNA mapping, sequence information from genome projects and protein pattern profiling can be linked with each other and annotated. These comprehensive approaches permit the study of differential gene and protein expressions in cells or tissues.

Animals↗

IMGT, the international ImMunoGeneTics database: a new design for immunogenetics data access.

IMGT, the international ImMunoGeneTics database is an integrated database specializing in Immunoglobulins (Ig), T-cell receptors (TcR) and MHC molecules of all vertebrate species, created by Marie-Paule Lefranc, University of Montpellier, CNRS, Montpellier, France (Nucleic Acids Research, Database issue, Vol 26, January 1998). IMGT includes three databases: LIGM-DB (for Ig and TcR), MHC/HLA-DB and IMGT/PRIMER-DB (an Ig, TcR and MHC-related primer database), the last two in development. IMGT comprises expertly annotated sequences and alignment tables. LIGM-DB contains more than 24.000 Immunoglobulin and T cell Receptor sequences from 81 different species. MHC/HLA-DB contains class I and class II Human Leucocyte Antigen alignment tables. An IMGT tool, DNAPLOT, developed for Ig, TcR and MHC sequence analysis, is also available. IMGT goals are to establish a common data access to all immunogenetics data, including nucleotide and protein sequences, oligonucleotide primers, gene maps and other genetic data of Ig, TcR and MHC molecules, from all species, and to provide a graphical user friendly data access. IMGT has important implications in medical research (repertoire in autoimmune diseases, AIDS, leukemias, lymphomas), therapeutical approaches (antibody engineering), genome diversity and genome evolution studies. In this paper, we describe our approach for the data modelisation, the automation of the annotation procedure and control of data quality in LIGM-DB database. IMGT is freely available on the CNUSC WWW server at Montpellier: http://imgt.cnusc.fr: 8104 (contact: Denys.Chaume@cnusc.fr) and on the EBI servers: http://www.ebi.ac.uk/imgt (contact: malik@ebi.ac.uk) and ftp.ebi.ac.uk/pub/databases/imgt. LIGM-DB users are encouraged to report errors or suggestions to giudi@ligm.crbm.cnrs-mop.fr. IMGT initiator and coordinator: Marie-Paule Lefranc, lefranc@ligm.crbm.cnrs-mop.fr. (fax: +33(0)467040231).

Amino Acid Sequence↗

MILANO--custom annotation of microarray results using automatic literature searches.

BACKGROUND: High-throughput genomic research tools are becoming standard in the biologist's toolbox. After processing the genomic data with one of the many available statistical algorithms to identify statistically significant genes, these genes need to be further analyzed for biological significance in light of all the existing knowledge. Literature mining--the process of representing literature data in a fashion that is easy to relate to genomic data--is one solution to this problem. RESULTS: We present a web-based tool, MILANO (Microarray Literature-based Annotation), that allows annotation of lists of genes derived from microarray results by user defined terms. Our annotation strategy is based on counting the number of literature co-occurrences of each gene on the list with a user defined term. This strategy allows the customization of the annotation procedure and thus overcomes one of the major limitations of the functional annotations usually provided with microarray results. MILANO expands the gene names to include all their informative synonyms while filtering out gene symbols that are likely to be less informative as literature searching terms. MILANO supports searching two literature databases: GeneRIF and Medline (through PubMed), allowing retrieval of both quick and comprehensive results. We demonstrate MILANO's ability to improve microarray analysis by analyzing a list of 150 genes that were affected by p53 overproduction. This analysis reveals that MILANO enables immediate identification of known p53 target genes on this list and assists in sorting the list into genes known to be involved in p53 related pathways, apoptosis and cell cycle arrest. CONCLUSIONS: MILANO provides a useful tool for the automatic custom annotation of microarray results which is based on all the available literature. MILANO has two major advances over similar tools: the ability to expand gene names to include all their informative synonyms while removing synonyms that are not informative and access to the GeneRIF database which provides short summaries of curated articles relevant to known genes. MILANO is available at http://milano.md.huji.ac.il.

Algorithms↗

Comprehensive analysis of pathway or functionally related gene expression in the National Cancer Institute's anticancer screen.

We have analyzed the level of gene coregulation, using gene expression patterns measured across the National Cancer Institute's 60 tumor cell panels (NCI(60)), in the context of predefined pathways or functional categories annotated by KEGG (Kyoto Encyclopedia of Genes and Genomes), BioCarta, and GO (Gene Ontology). Statistical methods were used to evaluate the level of gene expression coherence (coordinated expression) by comparing intra- and interpathway gene-gene correlations. Our results show that gene expression in pathways, or groups of functionally related genes, has a significantly higher level of coherence than that of a randomly selected set of genes. Transcriptional-level gene regulation appears to be on a "need to be" basis, such that pathways comprising genes encoding closely interacting proteins and pathways responsible for vital cellular processes or processes that are related to growth or proliferation, specifically in cancer cells, such as those engaged in genetic information processing, cell cycle, energy metabolism, and nucleotide metabolism, tend to be more modular (lower degree of gene sharing) and to have genes significantly more coherently expressed than most signaling and regular metabolic pathways. Hierarchical clustering of pathways based on their differential gene expression in the NCI(60) further revealed interesting interpathway communications or interactions indicative of a higher level of pathway regulation. The knowledge of the nature of gene expression regulation and biological pathways can be applied to understanding the mechanism by which small drug molecules interfere with biological systems.

Algorithms↗