Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Cell annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

[Cancer genome or the development of molecular portraits of tumors].

The rapid development of cancer genomics is due to important progresses in oncogenesis, human genome sequencing and emergence of new technologies in genome and transcriptome analysis. In this context, the aim of the French program 'Cartes d'Identites des Tumeurs--Molecular Portraits of Tumors' is to build a public data base containing a pan genome assessment of genome and transcriptome alterations in the major types of tumors as well as in relevant normal cells and experimental models. Data mining is done in the context of genome annotations and clinical and biological informations attached to the enrolled samples. The goal of the program is to define new tests useful for diagnostic procedures in clinical laboratories and new targets for biological treatments of tumors.

France↗

LY-6K gene: a novel molecular marker for human breast cancer.

A full-length cDNA was identified using one STS sequence containing an SNP (single nucleotide polymorphism) derived from genomic DNAs of breast cancer patients using a variety of bioinformatics tools. The cDNA encodes LY-6K, a novel member protein of the Ly-6/uPAR superfamily. It has been annotated as a target antigen for the HNSCC (head-and neck squamous cell carcinoma). We isolated the LY-6K gene from genomic DNAs obtained from breast cancer patients through large scale, case-control-screening. We performed northern blot hybridization and semi-quantitative RT-PCR on a human multiple-tissue mRNA blot from several breast cancer patients. We investigated the expression level of the LY-6K gene in human breast cancer, and compared this to expression in human normal breast tissue. We found that LY-6K was more highly expressed in the mRNA of breast tumors compared to its expression in normal breast tissue. These results suggest that LY-6K is not only a target antigen for HNSCC but also a significant new molecular marker for diagnosis and gene therapy in patients with breast cancer.

Antigens, Ly↗

Annotation and analysis of 10,000 expressed sequence tags from developing mouse eye and adult retina.

BACKGROUND: As a biomarker of cellular activities, the transcriptome of a specific tissue or cell type during development and disease is of great biomedical interest. We have generated and analyzed 10,000 expressed sequence tags (ESTs) from three mouse eye tissue cDNA libraries: embryonic day 15.5 (M15E) eye, postnatal day 2 (M2PN) eye and adult retina (MRA). RESULTS: Annotation of 8,633 non-mitochondrial and non-ribosomal high-quality ESTs revealed that 57% of the sequences represent known genes and 43% are unknown or novel ESTs, with M15E having the highest percentage of novel ESTs. Of these, 2,361 ESTs correspond to 747 unique genes and the remaining 6,272 are represented only once. Phototransduction genes are preferentially identified in MRA, whereas transcripts for cell structure and regulatory proteins are highly expressed in the developing eye. Map locations of human orthologs of known genes uncovered a high density of ocular genes on chromosome 17, and identified 277 genes in the critical regions of 37 retinal disease loci. In silico expression profiling identified 210 genes and/or ESTs over-expressed in the eye; of these, more than 26 are known to have vital retinal function. Comparisons between libraries provided a list of temporally regulated genes and/or ESTs. A few of these were validated by qRT-PCR analysis. CONCLUSIONS: Our studies present a large number of potentially interesting genes for biological investigation, and the annotated EST set provides a useful resource for microarray and functional genomic studies.

Aging↗

The mighty microproteins: from versatile cellular regulators to precision medicine therapeutics.

Microproteins, are tiny proteins encoded by small open reading frame (sORF), translation of these non-canonical open reading frames (ncORFs) has been implicated in diverse biological processes and diseases. This review summarizes recent developments in the discovery, biogenesis, and functional characterization of microproteins, and their involvement in various disease, with special focus on their roles in cancer, cardiovascular, metabolic, neurodegenerative and immune-related disorders. We emphasize the regulation of key cellular pathways by microproteins, including mitochondrial homeostasis, apoptosis, metabolic reprogramming, and immune signaling, all of which affect disease initiation and progression. Emerging evidence also supports their potential as disease biomarkers and therapeutic candidates for precision medicine. Finally, the review critically discusses the current challenges including discrepancies in microprotein annotation, the limitations of ribosome profiling and proteogenomic approaches, the gap between computationally predicted and experimentally validated microproteins, and the need for rigorous orthogonal validation by means of CRISPR-based genome editing, ribosome release assays, mutational analysis, high-resolution mass spectrometry, and functional studies. Finally, we review recent development of AI-assisted ORF prediction, single-cell translatomics, spatial proteomics, and integrated multi-omics as emerging technologies reshaping. Microprotein discovery and functional annotation. Finally, we discuss the translational potential of microproteins and highlight the remaining challenges to clinical application, including peptide stability, pharmacokinetics, tissue-specific delivery, immunogenicity, and the need for rigorous preclinical and clinical validation. Together, this review provides an updated and critical overview of the rapidly evolving microprotein field and highlights future research priorities for translating these molecules into clinically useful biomarkers and precision therapeutics.

Microproteins↗

Comparing the protein expression profiles of human mesenchymal stem cells and human osteoblasts using gene ontologies.

One of the hallmark events regulating the process of osteogenesis is the transition of undifferentiated human mesenchymal stem cells (hMSCs) found in the bone marrow into mineralized-matrix producing osteoblasts (hOSTs) through mechanisms that are not entirely understood. With recent developments in mass spectrometry and its potential application to the systematic definition of the stem cell proteome, proteins that govern cell fate decisions can be identified and tracked during this differentiation process. We hypothesize that protein profiling of hMSCs and hOSTs will identify potential osteogenic marker proteins associated with hMSC commitment and hOST differentiation. To identify markers for each cell population, we analyzed the expression of hMSC proteins and compared them to that of hOST by two-dimensional gel electrophoresis and two-dimensional liquid chromatography tandem mass spectrometry (2D LC-MS/MS). The 2D LC-MS/MS data sets were analyzed using the Database for Annotation, Visualization and Integrated Discovery (DAVID). Only 34% of the spots in 2D gels were found in both cell populations; of those that differed between populations, 65% were unique to hOST cells. Of the 755 different proteins identified by 2D LCMS/ MS in both cell populations, two sets of 247 and 158 proteins were found only in hMSCs and hOST cells, respectively. Differential expression of some of the identified proteins was further confirmed by Western blot analyses. Substantial differences in clusters of proteins responsible for calcium- based signaling and cell adhesion were found between the two cell types. Osteogenic differentiation is accompanied by a substantial change in the overall protein expression profile of hMSCs. This study, using gene ontology analysis, reveals that these changes occur in clusters of functionally related proteins. These proteins may serve as markers for identifying stem cell differentiation into osteogenic fates because they promote differentiation by mechanisms that remain to be defined.

Blotting, Western↗

Prokaryotic RNA preparation methods useful for high density array analysis: comparison of two approaches.

High density oligonucleotide arrays have been used extensively for expression studies of eukaryotic organisms. We have designed a prokaryotic high density oligonucleotide array using the complete Escherichia coli genome sequence to monitor expression levels of all genes and intergenic regions in the genome. Because previously described methods for preparing labeled target nucleic acids are not useful for prokaryotic cell analysis using such arrays, a mRNA enrichment and direct labeling protocol was developed together with a cDNA synthesis protocol. The reproducibility of each labeling method was determined using high density oligonucleotide probe arrays as a read-out methodology and the expression results from direct labeling were compared to the expression results from the cDNA synthesis. About 50% of all annotated E.coli open reading frames are observed to be transcribed, as measured by both protocols, when the cells were grown in rich LB medium. Each labeling method individually showed a high degree of concordance in replica experiments (95 and 99%, respectively), but when each sample preparation method was compared to the other, approximately 32% of the genes observed to be expressed were discordant. However, both labeling methods can detect the same relative gene expression changes when RNA from IPTG-induced cells was labeled and compared to RNA from uninduced E.coli cells.

Adenosine Triphosphate↗

Analysis of sample set enrichment scores: assaying the enrichment of sets of genes for individual samples in genome-wide expression profiles.

MOTIVATION: Gene expression profiling experiments in cell lines and animal models characterized by specific genetic or molecular perturbations have yielded sets of genes annotated by the perturbation. These gene sets can serve as a reference base for interrogating other expression datasets. For example, a new dataset in which a specific pathway gene set appears to be enriched, in terms of multiple genes in that set evidencing expression changes, can then be annotated by that reference pathway. We introduce in this paper a formal statistical method to measure the enrichment of each sample in an expression dataset. This allows us to assay the natural variation of pathway activity in observed gene expression data sets from clinical cancer and other studies. RESULTS: Validation of the method and illustrations of biological insights gleaned are demonstrated on cell line data, mouse models, and cancer-related datasets. Using oncogenic pathway signatures, we show that gene sets built from a model system are indeed enriched in the model system. We employ ASSESS for the use of molecular classification by pathways. This provides an accurate classifier that can be interpreted at the level of pathways instead of individual genes. Finally, ASSESS can be used for cross-platform expression models where data on the same type of cancer are integrated over different platforms into a space of enrichment scores. AVAILABILITY: Versions are available in Octave and Java (with a graphical user interface). Software can be downloaded at http://people.genome.duke.edu/assess.

Algorithms↗

Identification of catabolite repression as a physiological regulator of biofilm formation by Bacillus subtilis by use of DNA microarrays.

Biofilms are structured communities of cells that are encased in a self-produced polymeric matrix and are adherent to a surface. Many biofilms have a significant impact in medical and industrial settings. The model gram-positive bacterium Bacillus subtilis has recently been shown to form biofilms. To gain insight into the genes involved in biofilm formation by this bacterium, we used DNA microarrays representing >99% of the annotated B. subtilis open reading frames to follow the temporal changes in gene expression that occurred as cells transitioned from a planktonic to a biofilm state. We identified 519 genes that were differentially expressed at one or more time points as cells transitioned to a biofilm. Approximately 6% of the genes of B. subtilis were differentially expressed at a time when 98% of the cells in the population were in a biofilm. These genes were involved in motility, phage-related functions, and metabolism. By comparing the genes differentially expressed during biofilm formation with those identified in other genomewide transcriptional-profiling studies, we were able to identify several transcription factors whose activities appeared to be altered during the transition from a planktonic state to a biofilm. Two of these transcription factors were Spo0A and sigma-H, which had previously been shown to affect biofilm formation by B. subtilis. A third signal that appeared to be affecting gene expression during biofilm formation was glucose depletion. Through quantitative biofilm assays and confocal scanning laser microscopy, we observed that glucose inhibited biofilm formation through the catabolite control protein CcpA.

Bacillus subtilis↗

Analysis of genes induced in peripheral nerve after axotomy using cDNA microarrays.

One of the most striking features of neurons in the mature peripheral nervous system is their ability to survive and to regenerate their axons following axonal injury. To perform a comprehensive survey of the molecular mechanisms that underlie peripheral nerve regeneration, we analyzed a cDNA library derived from the distal stumps of post-injured sciatic nerve which was enriched in non-myelinating Schwann cells using cDNA microarrays. The number of up- and down-regulated genes in the transected sciatic nerve was 370 and 157, respectively, of the 9596 spotted genes. In the up-regulated group, the number of known genes was 216 and the number of expressed sequence tag (EST) sequences was 154. In the down-regulated group, the number of known genes was 103 and that of EST sequences was 54. We obtained several genes that were previously reported to be involved in regeneration of the injured neurons, such as cathepsin D, ninjurin 1, tenascin C, and co-receptor for glial cell line-derived neurotrophic factor family of trophic factors. In addition to unknown genes, there seemed to be a lot of annotated genes whose role in nerve regeneration remains unknown.

Animals↗

Malaria and the red blood cell membrane.

Malaria is the most serious and widespread parasitic disease of humans and is arguably the commonest disease of red blood cells (RBCs). Malaria has exerted a powerful effect on human evolution and selection for resistance has led to the appearance and persistence of a number of inherited diseases. After parasite invasion, RBCs are progressively and dramatically modified. New structures appear inside the RBC and novel parasite proteins are exported to the erythrocyte cytoplasm and membrane skeleton. Radical biochemical, morphological, and rheological alterations manifest as increased membrane rigidity, reduced cell deformability, and greater adhesiveness for the vascular endothelium and other blood cells. Numerous protein-protein interactions between the malaria-parasite and the host RBC are important for many aspects of parasite biology and the pathogenesis of malaria. In addition, there are many other parasite proteins located within the infected red cell and at the membrane skeleton, for which no precise functional roles have yet been elucidated. Sequencing and annotation of the complete genome of Plasmodium falciparum, the production of proteomic and transcriptomic profiles of parasites, and the development of a transfection system for the asexual stage of the parasite are all recent achievements that should advance understanding of the molecular mechanisms that underlie the parasite-induced functional alterations in red cells.

Animals↗

Genome SEGE: a database for 'intronless' genes in eukaryotic genomes.

BACKGROUND: A number of completely sequenced eukaryotic genome data are available in the public domain. Eukaryotic genes are either 'intron containing' or 'intronless'. Eukaryotic 'intronless' genes are interesting datasets for comparative genomics and evolutionary studies. The SEGE database containing a collection of eukaryotic single exon genes is available. However, SEGE is derived using GenBank. The redundant, incomplete and heterogeneous qualities of GenBank data are a bottleneck for biological investigation in comparative genomics and evolutionary studies. Such studies often require representative gene sets from each genome and this is possible only by deriving specific datasets from completely sequenced genome data. Thus Genome SEGE, a database for 'intronless' genes in completely sequenced eukaryotic genomes, has been constructed. AVAILABILITY: http://sege.ntu.edu.sg/wester/intronless DESCRIPTION: Eukaryotic 'intronless' genes are extracted from nine completely sequenced genomes (four of which are unicellular and five of which are multi-cellular). The complete dataset is available for download. Data subsets are also available for 'intronless' pseudo-genes. The database provides information on the distribution of 'intronless' genes in different genomes together with their length distributions in each genome. Additionally, the search tool provides pre-computed PROSITE motifs for each sequence in the database with appropriate hyperlinks to InterPro. A search facility is also available through the web server. CONCLUSIONS: The unique features that distinguish Genome SEGE from SEGE is the service providing representative 'intronless' datasets for completely sequenced genomes. 'Intronless' gene sets available in this database will be of use for subsequent bio-computational analysis in comparative genomics and evolutionary studies. Such analysis may help to revisit the original genome data for re-examination and re-annotation.

Databases, Genetic↗

Transcriptional analysis of a novel cluster of LY-6 family members in the human and mouse major histocompatibility complex: five genes with many splice forms.

Lymphocyte antigen-6 (LY-6) superfamily members are cysteine-rich, generally GPI-anchored cell surface proteins, which have definite or putative immune related roles. A cluster of five potential LY-6 superfamily members is located in the human and mouse major histocompatibility complex class III region. Comparative analysis of their genomic and cDNA sequences allowed us to carry out detailed annotations of these genes. We analyzed their mRNA expression patterns by RT-PCR performed on human and mouse cell line and tissue RNA. Sequence analysis of the transcripts revealed splice variants of all these genes in humans, and all but one in mouse. These splice forms retained introns or intron fragments, mainly generating premature stop codons, such that the only potentially functional mRNA was the predicted form. In some cases, the mis-spliced form was the most abundant form, suggesting a control mechanism for gene expression. Each gene showed mRNA expression differences between human and mouse.

Alternative Splicing↗

Genome-scale models of microbial cells: evaluating the consequences of constraints.

Microbial cells operate under governing constraints that limit their range of possible functions. With the availability of annotated genome sequences, it has become possible to reconstruct genome-scale biochemical reaction networks for microorganisms. The imposition of governing constraints on a reconstructed biochemical network leads to the definition of achievable cellular functions. In recent years, a substantial and growing toolbox of computational analysis methods has been developed to study the characteristics and capabilities of microorganisms using a constraint-based reconstruction and analysis (COBRA) approach. This approach provides a biochemically and genetically consistent framework for the generation of hypotheses and the testing of functions of microbial cells.

Bacterial Physiological Phenomena↗

Discovery of 342 putative new genes from the analysis of 5'-end-sequenced full-length-enriched cDNA human transcripts.

In this work we describe the process that, starting with the production of human full-length-enriched cDNA libraries using the CAP-Trapper method, led us to the discovery of 342 putative new human genes. Twenty-three thousand full-length-enriched clones, obtained from various cell lines and tissues in different developmental stages, were 5'-end sequenced, allowing the identification of a pool of 5300 unique cDNAs. By comparing these sequences to various human and vertebrate nucleotide databases we found that about 40% of our clones extended previously annotated 5' ends, 662 clones were likely to represent splice variants of known genes, and finally 342 clones remained unknown, with no or poor functional annotation. cDNA-microarray gene expression analysis showed that 260 of 342 unknown clones are expressed in at least one cell line and/or tissue. Further analysis of their sequences and the corresponding genomic locations allowed us to conclude that most of them represent potential novel genes, with only a small fraction having protein-coding potential.

5' Flanking Region↗

CoSAG-nf: A Scalable Nextflow Pipeline for Co-assembly, Optimization, and Interactive Visualization of High-Throughput Single-Cell Genomes.

MOTIVATION: Single-cell amplified genomes (SAGs) are crucial for resolving intra-population microbial heterogeneity and accurately understanding the metabolic potential of microbial dark matter populations. However, SAGs generated through multiple displacement amplification (MDA) of genomic DNA from single cells with single-copy chromosomes are highly fragmented and prone to contamination, severely hindering high-quality genome reconstruction and functional analysis, which greatly limits their scientific utility. Co-assembly of related SAGs can substantially improve genome quality, but to our knowledge no automated pipeline exists for high-throughput processing, forcing manual implementation of complex workflows that scale poorly to modern dataset sizes. RESULTS: We present CoSAG-nf, an automated high-throughput co-assembly and optimization pipeline for SAGs, implemented following the nf-core framework standards. The pipeline performs alignment-free clustering using sourmash MinHash signatures, then employs iterative tetranucleotide frequency profiling to identify and exclude outlier SAGs from co-assembly groups. CheckM2 quality assessment guides dynamic selection of optimal SAG combinations to optimize genome completeness and minimize contamination. Fully containerized, CoSAG-nf ensures reproducibility and scalability for the high-throughput processing of large-scale SAG datasets across diverse computing environments, including HPC and cloud platforms. The pipeline generates comprehensive HTML reports with quality metrics and taxonomic annotations, providing an end-to-end solution for automated high-throughput single-cell genome reconstruction. AVAILABILITY: CoSAG-nf is freely available under the MIT License at: https://github.com/linfengxu/CoSAG-nf. Archival code repository snapshots are published at zenodo with doi: https://doi.org/10.5281/zenodo.21525244. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Journal Article↗

Characterization of soybean genomic features by analysis of its expressed sequence tags.

We analyzed 314,254 soybean expressed sequence tags (ESTs), including 29,540 from our laboratory and 284,714 from GenBank. These ESTs were assembled into 56,147 unigenes. About 76.92% of the unigenes were homologous to genes from Arabidopsis thaliana ( Arabidopsis). The putative products of these unigenes were annotated according to their homology with the categorized proteins of Arabidopsis. Genes corresponding to cell growth and/or maintenance, enzymes and cell communication belonged to the slow-evolving class, whereas genes related to transcription regulation, cell, binding and death appeared to be fast-evolving. Soybean unigenes with no match to genes within the Arabidopsis genome were identified as soybean-specific genes. These genes were mainly involved in nodule development and the synthesis of seed storage proteins. In addition, we also identified 61 genes regulated by salicylic acid, 1,322 transcription factor genes and 326 disease resistance-like genes from soybean unigenes. SSR analysis showed that the soybean genome was more complex than the Arabidopsis and the Medicago truncatula genomes. GC content in soybean unigene sequences is similar to that in Arabidopsis and M. truncatula. Furthermore, the combined analysis of the EST database and the BAC-contig sequences revealed that the total gene number in the soybean genome is about 63,501.

Chromosomes, Artificial, Bacterial↗

Red and far-red light alter the transcript profile in the cyanobacterium Synechocystis sp. PCC 6803: impact of cyanobacterial phytochromes.

Cyanobacteria possess genes encoding phytochrome-related proteins. We used a DNA microarray approach to evaluate the impact of the phytochromes Cph1 and Cph2 on red light (R)- and far-red light (FR)-dependent gene expression in the unicellular cyanobacterium Synechocystis sp. PCC 6803. In cells of wild-type and phytochrome mutants, one-fourth of all 3165 annotated putative protein encoding genes was light-responsive. R predominantly enhanced the expression of genes involved in transcription, translation, and photosynthesis, whereas FR upregulated the transcript level of genes known to be inducible by stress. The absence of Cph1 and/or Cph2 altered the light-dependent expression of about 20 genes. Hence, receptor(s) different from the two phytochromes are supposed to trigger the global R/FR alterations of the expression profile.

Bacterial Proteins↗

CDP-2,3-Di-O-geranylgeranyl-sn-glycerol:L-serine O-archaetidyltransferase (archaetidylserine synthase) in the methanogenic archaeon Methanothermobacter thermautotrophicus.

CDP-2,3-di-O-geranylgeranyl-sn-glycerol:L-serine O-archaetidyltransferase (archaetidylserine synthase) activity in cell extracts of Methanothermobacter thermautotrophicus cells was characterized. The enzyme catalyzed the formation of unsaturated archaetidylserine from CDP-unsaturated archaeol and L-serine. The identity of the reaction products was confirmed by thin-layer chromatography, fast atom bombardment-mass spectrum analysis, and chemical degradation. The enzyme showed maximal activity in the presence of 10 mM Mn2+ and 1% Triton X-100. Among various synthetic substrate analogs, both enantiomers of CDP-unsaturated archaeols with ether-linked geranylgeranyl chains and CDP-saturated archaeol with ether-linked phytanyl chains were similarly active toward the archaetidylserine synthase. The activity on the ester analog of the substrate was two to three times higher than that on the corresponding ether-type substrate. The activity of D-serine with the enzyme was 30% of that observed for L-serine. A trace amount of an acid-labile, unsaturated archaetidylserine intermediate was detected in the cells by a pulse-labeling experiment. A gene (MT1027) in M. thermautotrophicus genome annotated as the gene encoding phosphatidylserine synthase was found to be homologous to Bacillus subtilis pssA but not to Escherichia coli pssA. The substrate specificity of phosphatidylserine synthase from B. subtilis was quite similar to that observed for the M. thermautotrophicus archaetidylserine synthase, while the E. coli enzyme had a strong preference for CDP-1,2-diacyl-sn-glycerol. It was concluded that M. thermautotrophicus archaetidylserine synthase belongs to subclass II phosphatidylserine synthase (B. subtilis type) on the basis of not only homology but also substrate specificity and some enzymatic properties. The possibility that a gene encoding the subclass II phosphatidylserine synthase might be transferred from a bacterium to an ancestor of methanogens is discussed.

Amino Acid Sequence↗