Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Cell annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Whole metagenome sequencing: not deep enough for complete microbial function recovery.

BACKGROUND: Whole metagenome shotgun sequencing (WMS) is widely used to profile microbial function. However, technical variability in sequencing and analysis often obscures true biological patterns. Large-scale studies are particularly susceptible to batch effects, such as differences in sequencing depth and platform and annotation strategies, as well as sample-to-flow-cell assignments. However, the relative effects of these factors on functional inference in such studies have yet to be systematically evaluated. We analyzed oral-rinse WMS data from 671 Nigerian youths aged 9-18, sequenced on two Illumina platforms. Microbial molecular functionality encoded in these data was annotated using the mi-faser/Fusion pipeline, to capture the broad functional repertoire, and HUMAnN 3/EC numbers pipeline to characterize curated enzymatic activities. We then quantified how technical factors and batch effects shaped the recovery of microbial functionality. RESULTS: Three findings of our work were most salient. First, we observed that the choice of annotation strategy traded off between breadth and specificity of functional coverage. Second, we found that low-prevalence functions were disproportionately lost at shallow sequencing depths, indicating that in, e.g., case-control studies with few representatives of the minor class, sequencing depth could critically impact study resolution. Finally, using our newly developed model relating sequencing depth to functional recovery, we demonstrated that increasing sequencing depth does not directly or proportionally improve functional recall. That is, at as little as 10% of this study's sequencing depth, 30% of the estimated complete microbiome functional repertoire was detectable. However, even at the full depth used in this study, we were only able to recover an estimated 60% of that complete functional repertoire. We further showed that despite biomes differences in functional diversity and host contamination levels (e.g., soil, fecal), incomplete functional recovery at commonly used sequencing depths was consistently observed. CONCLUSIONS: Together, these findings and our depth-to-function mapping framework provide practical guidelines for the design and interpretation of WMS studies. Coordinating sequencing depth planning with annotation strategy, experimental design, and rigorous batch control is thus essential for robust detection of microbial functions and for ensuring reproducible microbiome insights. Video Abstract.

Humans↗

Dissecting progressive stages of 5-fluorouracil resistance in vitro using RNA expression profiling.

Resistance to anticancer drugs such as the widely used antimetabolite 5-fluorouracil (FU) is one of the most important obstacles to cancer chemotherapy. Using GeneChip arrays, we compared the expression profile of different stages of FU resistance in colon cancer cells after in vitro selection of low-, intermediate- and high-resistance phenotypes. Drug resistance was associated with significant changes in expression of 330 genes, mainly during early or intermediate stage. Functional annotation revealed a majority of genes involved in signal transduction, cell adhesion and cytoskeleton with subsequent alterations in apoptotic response, cell cycle control, drug transport, fluoropyrimidine metabolism and DNA repair. A set of 33 genes distinguished all resistant subclones from sensitive progenitor cells. In the early stage, downregulation of collagens and keratins, together with upregulation of profilin 2 and ICAM-2, suggested cytoskeletal changes and cell adhesion remodeling. Interestingly, 6 members of the S100 calcium-binding protein family were suppressed. Acquisition of the intermediate-resistance phenotype included upregulation of the well-known drug resistance gene ABCC6 (ATP-binding cassette subfamily C member 6). The very small number of genes affected during transition to high resistance included the primary FU target thymidylate synthase. Although limited to an in vitro model, our data suggest that resistance to FU cannot be explained by known mechanisms alone and substantially involves a wide molecular repertoire. This study emphasizes the understanding of resistance as a time-depending process: the cell is particularly challenged at the beginning of this process, while acquisition of the high-resistance phenotype seems to be less demanding.

Antimetabolites, Antineoplastic↗

The fasciclin-like arabinogalactan proteins of Arabidopsis. A multigene family of putative cell adhesion molecules.

Fasciclin-like arabinogalactan proteins (FLAs) are a subclass of arabinogalactan proteins (AGPs) that have, in addition to predicted AGP-like glycosylated regions, putative cell adhesion domains known as fasciclin domains. In other eukaryotes (e.g. fruitfly [Drosophila melanogaster] and humans [Homo sapiens]), fasciclin domain-containing proteins are involved in cell adhesion. There are at least 21 FLAs in the annotated Arabidopsis genome. Despite the deduced proteins having low overall similarity, sequence analysis of the fasciclin domains in Arabidopsis FLAs identified two highly conserved regions that define this motif, suggesting that the cell adhesion function is conserved. We show that FLAs precipitate with beta-glucosyl Yariv reagent, indicating that they share structural characteristics with AGPs. Fourteen of the FLA family members are predicted to be C-terminally substituted with a glycosylphosphatidylinositol anchor, a cleavable form of membrane anchor for proteins, indicating different FLAs may have different developmental roles. Publicly available microarray and expressed sequence tag data were used to select FLAs for further expression analysis. RNA gel blots for a number of FLAs indicate that they are likely to be important during plant development and in response to abiotic stress. FLAs 1,2, and 8 show a rapid decrease in mRNA abundance in response to the phytohormone abscisic acid. Also, the accumulation of FLA1 and FLA2 transcripts differs during callus and shoot development, indicating that the proteins may be significant in the process of competence acquisition and induction of shoot development.

Abscisic Acid↗

A bioinformatics based approach to discover small RNA genes in the Escherichia coli genome.

The recent explosion in available bacterial genome sequences has initiated the need to improve an ability to annotate important sequence and structural elements in a fast, efficient and accurate manner. In particular, small non-coding RNAs (sRNAs) have been difficult to predict. The sRNAs play an important number of structural, catalytic and regulatory roles in the cell. Although a few groups have recently published prediction methods for annotating sRNAs in bacterial genome, much remains to be done in this field. Toward the goal of developing an efficient method for predicting unknown sRNA genes in the completed Escherichia coli genome, we adopted a bioinformatics approach to search for DNA regions that contain a sigma70 promoter within a short distance of a rho-independent terminator. Among a total of 227 candidate sRNA genes initially identified, 32 were previously described sRNAs, orphan tRNAs, and partial tRNA and rRNA operons. Fifty-one are mRNAs genes encoding annotated extremely small open reading frames (ORFs) following an acceptable ribosome binding site. One hundred forty-four are potentially novel non-translatable sRNA genes. Using total RNA isolated from E. coli MG1655 cells grown under four different conditions, we verified transcripts of some of the genes by Northern hybridization. Here we summarize our data and discuss the rules and advantages/disadvantages of using this approach in annotating sRNA genes on bacterial genomes.

Base Sequence↗

PhenoGO: assigning phenotypic context to gene ontology annotations with natural language processing.

Natural language processing (NLP) is a high throughput technology because it can process vast quantities of text within a reasonable time period. It has the potential to substantially facilitate biomedical research by extracting, linking, and organizing massive amounts of information that occur in biomedical journal articles as well as in textual fields of biological databases. Until recently, much of the work in biological NLP and text mining has revolved around recognizing the occurrence of biomolecular entities in articles, and in extracting particular relationships among the entities. Now, researchers have recognized a need to link the extracted information to ontologies or knowledge bases, which is a more difficult task. One such knowledge base is Gene Ontology annotations (GOA), which significantly increases semantic computations over the function, cellular components and processes of genes. For multicellular organisms, these annotations can be refined with phenotypic context, such as the cell type, tissue, and organ because establishing phenotypic contexts in which a gene is expressed is a crucial step for understanding the development and the molecular underpinning of the pathophysiology of diseases. In this paper, we propose a system, PhenoGO, which automatically augments annotations in GOA with additional context. PhenoGO utilizes an existing NLP system, called BioMedLEE, an existing knowledge-based phenotype organizer system (PhenOS) in conjunction with MeSH indexing and established biomedical ontologies. More specifically, PhenoGO adds phenotypic contextual information to existing associations between gene products and GO terms as specified in GOA. The system also maps the context to identifiers that are associated with different biomedical ontologies, including the UMLS, Cell Ontology, Mouse Anatomy, NCBI taxonomy, GO, and Mammalian Phenotype Ontology. In addition, PhenoGO was evaluated for coding of anatomical and cellular information and assigning the coded phenotypes to the correct GOA; results obtained show that PhenoGO has a precision of 91% and recall of 92%, demonstrating that the PhenoGO NLP system can accurately encode a large number of anatomical and cellular ontologies to GO annotations. The PhenoGO Database may be accessed at the following URL: http://www.phenoGO.org

Computational Biology↗

Gene expression profiles of human proximal tubular epithelial cells in proteinuric nephropathies.

In kidney disease renal proximal tubular epithelial cells (RPTEC) actively contribute to the progression of tubulointerstitial fibrosis by mediating both an inflammatory response and via epithelial-to-mesenchymal transition. Using laser capture microdissection we specifically isolated RPTEC from cryosections of the healthy parts of kidneys removed owing to renal cell carcinoma and from kidney biopsies from patients with proteinuric nephropathies. RNA was extracted and hybridized to complementary DNA microarrays after linear RNA amplification. Statistical analysis identified 168 unique genes with known gene ontology association, which separated patients from controls. Besides distinct alterations in signal-transduction pathways (e.g. Wnt signalling), functional annotation revealed a significant upregulation of genes involved in cell proliferation and cell cycle control (like insulin-like growth factor 1 or cell division cycle 34), cell differentiation (e.g. bone morphogenetic protein 7), immune response, intracellular transport and metabolism in RPTEC from patients. On the contrary we found differential expression of a number of genes responsible for cell adhesion (like BH-protocadherin) with a marked downregulation of most of these transcripts. In summary, our results obtained from RPTEC revealed a differential regulation of genes, which are likely to be involved in either pro-fibrotic or tubulo-protective mechanisms in proteinuric patients at an early stage of kidney disease.

Aged↗

Applications of quantitative digital image analysis to breast cancer research.

Our studies of radiogenic carcinogenesis in mouse and human models of breast cancer are based on the view that cell phenotype, microenvironment composition, communication between cells and within the microenvironment are important factors in the development of breast cancer. This is complicated in the mammary gland by its postnatal development, cyclic evolution via pregnancy and involution, and dynamic remodeling of epithelial-stromal interactions, all of which contribute to breast cancer susceptibility. Microscopy is the tool of choice to examine cells in context. Specific features can be defined using probes, antibodies, immunofluorescence, and image analysis to measure protein distribution, cell composition, and genomic instability in human and mouse models of breast cancer. We discuss the integration of image acquisition, analysis, and annotation to efficiently analyze large amounts of image data. In the future, cell and tissue image-based studies will be facilitated by a bioinformatics strategy that generates multidimensional databases of quantitative information derived from molecular, immunological, and morphological probes at multiple resolutions. This approach will facilitate the construction of an in vivo phenotype database necessary for understanding when, where, and how normal cells become cancer.

Animals↗

Progress in rickettsial genome analysis from pioneering of Rickettsia prowazekii to the recent Rickettsia typhi.

Three rickettsial genomes have been sequenced and annotated. Rickettsia prowazekii and R. typhi have similar gene order and content. The few differences between R. prowazekii and R. typhi include a 12-kb insertion in R. prowazekii, a large inversion close to the origin of replication in R. typhi, and loss of the complete cytochrome c oxidase system by R. typhi. R. prowazekii, R. typhi, and R. conorii have 13, 24, and 560 unique genes, respectively, and share 775 genes, most likely their essential genes. The small genomes contain many pseudogenes and much noncoding DNA, reflecting the process of genome decay. R. typhi contains the largest number of pseudogenes (41), and R. conorii the fewest, in accordance with its larger number of genes and smaller proportion of noncoding DNA. Conversely, typhus rickettsiae contain fewer repetitive sequences. These genomes portray the key themes of rickettsial intracellular survival: lack of enzymes for sugar metabolism, lipid biosynthesis, nucleotide synthesis, and amino acid metabolism, suggesting that rickettsiae depend on the host for nutrition and building blocks; enzymes for the complete TCA cycle and several copies of ATP/ADP translocase genes, suggesting independent synthesis of ATP and acquisition of host ATP; and type IV secretion system. All rickettsiae share two outer membrane proteins (OmpB and Sca 4) and LPS biosynthesis machinery. RickA, unique to spotted fever rickettsiae, plays a role in induction of actin polymerization in R. conorii, but not in R. prowazekii or R. typhi. The genome of R. typhi contains four potentially membranolytic genes (tlyA, tlyC, pldA, and pat-1) and five autotransporter genes, sca 1, sca 2, sca 3, ompA, and ompB. The presence of six 50-amino acid repeat units in Sca 2 suggests function as an adhesin. The high laboratory passage of the sequenced strains raises the issue of the occurrence of laboratory mutations in genes not required for growth in cell culture or eggs. Resequencing revealed that eight annotated pseudogenes of E strain are actually intact genes. Comparative genomics of virulent and avirulent strains of rickettsial species may reveal their virulence factors.

Genome, Bacterial↗

Predicting protein subcellular localization: past, present, and future.

Functional characterization of every single protein is a major challenge of the post-genomic era. The large-scale analysis of a cell's proteins, proteomics, seeks to provide these proteins with reliable annotations regarding their interaction partners and functions in the cellular machinery. An important step on this way is to determine the subcellular localization of each protein. Eukaryotic cells are divided into subcellular compartments, or organelles. Transport across the membrane into the organelles is a highly regulated and complex cellular process. Predicting the subcellular localization by computational means has been an area of vivid activity during recent years. The publicly available prediction methods differ mainly in four aspects: the underlying biological motivation, the computational method used, localization coverage, and reliability, which are of importance to the user. This review provides a short description of the main events in the protein sorting process and an overview of the most commonly used methods in this field.

Computational Biology↗

CMAtlas: a comprehensive DNA methylation atlas for exploring epigenetic alterations in 34 human cancer types.

MOTIVATION: Aberrant DNA methylation is a fundamental epigenetic hallmark of cancer. However, existing resources often lack technological diversity and comprehensive cancer coverage. Furthermore, most platforms fail to achieve deep multi-omics integration and tend to ignore cancer-type-specific methylation features, limiting their utility in precision oncology and drug discovery. RESULTS: We developed Cancer Methylation Atlas (CMAtlas), a comprehensive platform integrating 13 753 samples across 34 cancer types. By applying technology-tailored pipelines to data from various profiling technologies, we identified 830 725 tumor-specific differentially methylated elements (DMEs) and 1 480 098 differentially methylated regions (DMRs), alongside 1 154 256 cancer-type-specific DMEs and 329 154 DMRs. The platform demonstrates high cross-platform consistency and strong concordance between tumor tissues and cell lines, ensuring the robustness of our findings. All DMEs and DMRs are annotated with multi-omics data (RNA expression, somatic mutations, and chromatin accessibility) and clinical relevance (survival associations and cell-free DNA profiling). We further demonstrate the utility of CMAtlas by identifying prognostic aberrant methylation in colorectal cancer driver genes. AVAILABILITY AND IMPLEMENTATION: CMAtlas is freely accessible at {{https://cmatlas.renlab.cn/}}. The platform offers an intuitive web interface supporting gene-centric and cancer-centric queries, alongside customizable analysis modules designed to facilitate user-specific research needs.

Humans↗

IMGT gene identification and Colliers de Perles of human immunoglobulins with known 3D structures.

A new database, IMGT/3Dstructure-DB, was developed and implemented in the IMGT (international ImMunoGeneTics database) information system (http://imgt.cines.fr) to provide a unique expertised resource on immunoglobulin and T-cell receptor structural data. Corresponding protein sequences were annotated with IMGT tools, which allow the precise identification of the genes expressed in these proteins, and the description of framework and complementarity determining regions according to the IMGT standardized nomenclature and IMGT unique numbering. Two-dimensional graphical representations of the V-DOMAINs, designated as Colliers de Perles, are automatically produced. A query Web interface allows interactive search of the IMGT/3D structure-DB data. In this article, IMGT gene identification and Colliers de Perles of human immunoglobulins with known 3D structures in the Protein Data Bank are presented.

Alleles↗

Proteomic analysis on metastasis-associated proteins of human hepatocellular carcinoma tissues.

PURPOSE: A comparative proteomic approach was used to identify and analyze proteins related to metastasis of hepatocellular carcinoma (HCC). METHODS: Proteins extracted from 12 HCC tissue specimens (six with metastases and six without) were separated by two-dimensional gel electrophoresis (2-DE). The protein spots exhibiting statistical alternations between the two groups through computerized image analysis were then identified by mass spectrometry. In addition immunohistochemistry (IHC), Western blotting and RT-PCR were performed to verify the expression of certain candidate proteins. RESULTS: 16 proteins including HSP27, S100A11, CK18 were annotated by mass spectrometry, relevant to chaperone function, cell mobility, cytoskeletal architecture, respectively. Most were previously unconnected with metastasis of HCC. Of these HSP27 was found overexpressed consistently in 2-DE patterns of all metastatic HCC tissues compared with nonmetastatic ones. IHC and Western blotting of HCC tissues confirmed this difference while RT-PCR did not. CONCLUSION: There are various proteins joined together in HCC metastasis. The overexpression of HSP27 may serve as a biomarker for early detection and therapeutic targets unique to the metastatic phenotype of HCC.

Actins↗

Identification of novel steroid target genes through the combination of bioinformatics and functional analysis of hormone response elements.

Steroid hormone receptors including androgen receptor (AR), glucocorticoid receptor (GR), progesterone receptor (PR), and mineralocorticoid receptor (MR) recognize and bind to identical consensus hormone response elements (HREs), which consist of two hexameric half-sites (5'-AGAACA-3') arranged as inverted repeats with a 3-bp spacer. Although only a few near-consensus HRE sequences have been identified in the transcriptional regulatory regions of known steroid target genes, it has been unclear whether the exact consensus sequences function as bona fide HREs in vivo. A genome-wide in silico screening of palindromic HREs identified 565 exact consensus sequences in human genome (NCBI 35 assembly). In this study, of 565 exact consensus elements, functional in vivo receptor binding was evaluated regarding 26 sequences located within 10 kb upstream to the 5' end of annotated genes through chromatin immunoprecipitation (ChIP) assay using cells endogenously expressing steroid hormone receptors. Hormone responsiveness of proximal gene expression was examined through quantitative RT-PCR. As far as performing ChIP assay for AR, GR, and PR, 14 of 26 elements significantly recruited at least one of the receptors by hormone treatment (>2-fold enrichment versus vehicle). In terms of gene expression in the vicinity of the above 14 functional perfect HREs, four genes were upregulated by >2-fold with hormone treatment. The present data suggest that the combination of bioinformatics analysis and quantitative experimental evaluation is useful to identify novel functional HREs that may contribute to the transcriptional regulation of steroid target genes.

Cell Line, Tumor↗

Application of eVOC: controlled vocabularies for unifying gene expression data.

To provide standardised description of gene expression and cross platform querying of databases, we have developed eVOC (http://www.sanbi.ac.za/evoc/), consisting of four orthogonal ontologies which describe Anatomical System, Cell Type, Pathology and Developmental Stage. We have annotated 47 microarray expression data sets and all publicly available human cDNA and SAGE tag libraries. eVOC has been integrated with the public resource EnsMart, which provides linking of transcripts and libraries with expression terms and the human genome sequence (http://www.ensembl.org/Homo_sapiens/martview).

Databases, Genetic↗

The nematode Caenorhabditis elegans as a model to study the roles of proteoglycans.

The nematode Caenorhabditis elegans is a powerful animal model for exploring the genetic basis of metazoan development. Recent genetic and biochemical studies have revealed that the molecular machinery of glycosaminoglycan (GAG) biosynthesis and modification is highly conserved between C. elegans and mammals. In addition, genetic studies have implicated GAGs in vulval morphogenesis and zygotic cytokinesis. The extensive knowledge of C. elegans biology, including its elucidated cell lineage, together with the completed and well annotated DNA sequence and availability of reverse genetic tools, provide a platform for studying the functions of proteoglycans and their GAG modification.

Animals↗

Ontology for immunogenetics: the IMGT-ONTOLOGY.

MOTIVATION: IMGT, the international ImMunoGeneTics database (http:@imgt.cines.fr:8104), created by M.-P. Lefranc, is an integrated database specializing in antigen receptors (immunoglobulins and T-cell receptors) and major histocompatibility complex (MHC) of all vertebrate species. IMGT accurate immunogenetics data are based on the standardization of the biological knowledge provided by the 'ImMunoGeneTics' IMGT-ONTOLOGY. The IMGT-ONTOLOGY describes the classification and specification of terms needed for immunogenetics and bioinformatics. IMGT-ONTOLOGY covers four main concepts: 'IDENTIFICATION', 'DESCRIPTION', 'CLASSIFICATION' and 'OBTENTION'. These concepts allow an extensive and standardized description and characterization of immunoglobulin and T-cell receptor data. The controlled vocabulary and the annotation rules are indispensable to ensure accuracy, consistency and coherence in IMGT. IMGT-ONTOLOGY allows scientists and clinicians to use, for the first time, identical terms with the same meaning in immunogenetics. It provides a semantic repository that will improve interoperability between specialist and generalist databases.

Animals↗

Age- and sex-adjusted genomic differences between Korean and Beat AML cohorts.

Genomic profiling plays a central role in risk stratification and therapeutic decision-making in acute myeloid leukemia (AML), yet the clinical implications of population-specific genomic architectures remain incompletely defined. We conducted a prospective, multicenter study of 603 adults with newly diagnosed AML in Korea, integrating targeted sequencing of 83 recurrently mutated genes with comprehensive clinical annotation across treatment intensities, including allogeneic hematopoietic stem cell transplantation (allo-HSCT). For contextual comparison, genomic profiles were evaluated against the Beat AML cohort. The overall genomic landscape was broadly conserved, supporting shared core disease biology across populations. However, RUNX1::RUNX1T1, CEBPA, GATA2, KIT, and DDX41 mutations were more frequent in the Korean cohort, whereas FLT3 and NPM1 mutations were less common. These differences translated into a distinct distribution of European LeukemiaNet (ELN) 2022 risk categories, with implications for therapeutic stratification. Notably, most DDX41 alterations were germline (3.2%), highlighting the need for systematic germline evaluation with implications for genetic counseling and donor selection. Although unadjusted overall survival appeared longer in the Korean cohort, this difference was not significant after adjustment for key clinical variables. These findings indicate that population-specific genomic distributions reshape the clinical application of risk stratification and support population-aware precision medicine strategies in AML.

Journal Article↗

Phylogenetic study of Rickettsia species using sequences of the autotransporter protein-encoding gene sca2.

The analyses of genome sequences from Rickettsia conorii and R. prowazekii have allowed the identification of five genes encoding autotransporter proteins, including ompA and four genes annotated in the R. prowazekii genome as "surface cell antigen" (sca) genes. Of these, ompA and sca5 (ompB) are known to encode membrane-exposed antigenic proteins playing a major role in the host's immune response, and sca4 encodes a truncated autotransporter protein. In order to study further the phylogeny of the genus Rickettsia, we attempted amplification and sequencing of the sca2 genes from the 20 currently validated Rickettsia species. Sixteen species exhibited a complete sca2 gene, ranging from 2,727 bp in R. bellii to 5,580 bp in R. rhipicephali. R. helvetica and R. canadensis had a split gene, and R. prowazekii and R. typhi had remnant fragments of sca2. We also identified in R. akari, R. prowazekii, and R. typhi a duplication of the sca2 gene. The phylogenetic trees inferred from both the nucleotide and protein sequences of sca2 showed four clusters of rickettsiae, that is, the R. rickettsii, R. massiliae, R. akari, and typhus groups, which were supported by significant bootstrap values.

Antigens, Surface↗