Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Cell annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Insertion sequences in prokaryotic genomes.

Insertion sequences (ISs) are small DNA segments that are often capable of moving neighbouring genes. Over 1500 different ISs have been identified to date. They can have large and spectacular effects in shaping and reshuffling the bacterial genome. Recent studies have provided dramatic examples of such IS activity, including massive IS expansion during the emergence of some pathogenic bacterial species and the intimate involvement of ISs in assembling genes into complex plasmid structures. However, a global understanding of their impact on bacterial genomes requires detailed knowledge of their distribution across the eubacterial and archaeal kingdoms, understanding their partition between chromosomes and extra-chromosomal elements (e.g. plasmids and viruses) and the factors which influence this, and appreciation of the different transposition mechanisms in action, the target preferences and the host factors that influence transposition. In addition, defective (non- autonomous) elements, which can be complemented by related active elements in the same cell, are often overlooked in genome annotations but also contribute to the evolution of genome organisation.

Bacteria↗

An integrated data analysis approach to characterize genes highly expressed in hepatocellular carcinoma.

Hepatocellular carcinoma (HCC) is one of the major causes of cancer deaths worldwide. New diagnostic and therapeutic options are needed for more effective and early detection and treatment of this malignancy. We identified 703 genes that are highly expressed in HCC using DNA microarrays, and further characterized them in order to uncover novel tumor markers, oncogenes, and therapeutic targets for HCC. Using Gene Ontology annotations, genes with functions related to cell proliferation and cell cycle, chromatin, repair, and transcription were found to be significantly enriched in this list of highly expressed genes. We also identified a set of genes that encode secreted (e.g. GPC3, LCN2, and DKK1) or membrane-bound proteins (e.g. GPC3, IGSF1, and PSK-1), which may be attractive candidates for the diagnosis of HCC. A significant enrichment of genes highly expressed in HCC was found on chromosomes 1q, 6p, 8q, and 20q, and we also identified chromosomal clusters of genes highly expressed in HCC. The microarray analyses were validated by RT-PCR and PCR. This approach of integrating other biological information with gene expression in the analysis helps select aberrantly expressed genes in HCC that may be further studied for their diagnostic or therapeutic utility.

Biomarkers, Tumor↗

Charting gene regulatory networks: strategies, challenges and perspectives.

One of the foremost challenges in the post-genomic era will be to chart the gene regulatory networks of cells, including aspects such as genome annotation, identification of cis-regulatory elements and transcription factors, information on protein-DNA and protein-protein interactions, and data mining and integration. Some of these broad sets of data have already been assembled for building networks of gene regulation. Even though these datasets are still far from comprehensive, and the approach faces many important and difficult challenges, some strategies have begun to make connections between disparate regulatory events and to foster new hypotheses. In this article we review several different genomics and proteomics technologies, and present bioinformatics methods for exploring these data in order to make novel discoveries.

Animals↗

Genomic Regions Associated with Resistance to Soybean Cyst Nematode (Heterodera glycines Ichinohe) Population HG Type 1.2.5.7 in Dry Beans (Phaseolus vulgaris L.).

North Dakota, the largest dry bean (Phaseolus vulgaris L.) producing state in the U.S., faces an emerging production threat caused by the soybean cyst nematode (SCN; Heterodera glycines Ichinohe, 1952). Host resistance is an effective management strategy, yet resistance to the virulent SCN population HG type 1.2.5.7 has not been genetically characterized in dry beans. In this study, 170 dry bean genotypes (113 breeding lines/cultivars and 57 germplasm accessions) were evaluated for response to HG type 1.2.5.7 under controlled conditions using female index (FI) as the resistance phenotype. FI values ranged from 4.1% to 78.1%, with one genotype (PI 313733) classified as resistant, 35 moderately resistant, 104 moderately susceptible, and 30 susceptible. Genome-wide association analysis using 2,044 single-nucleotide polymorphism (SNP) markers from the 3.8K Bean Panel chip and the BLINK model identified four significant marker-trait associations on chromosomes Pv02, Pv05, Pv07, and Pv11. Linkage disequilibrium-defined candidate intervals spanned 108 kb (Pv02), 1.50 Mb (Pv05), 798 kb (Pv07), and 1.45 Mb (Pv11), collectively containing 126 annotated genes: 20 on Pv02, 39 on Pv05, 35 on Pv07, and 32 on Pv11. The intervals contained putative genes annotated for signaling and transcriptional regulation, cell wall and carbohydrate metabolism, transport, and secondary metabolism. Together, these findings indicate that the response to HG type 1.2.5.7 in dry bean is quantitative and associated with multiple genomic regions. The identified intervals provide candidate targets for independent validation, fine mapping, functional analysis, and future marker development to support breeding for SCN resistance.

Disease Resistance↗

Prognostic meta-signature of breast cancer developed by two-stage mixture modeling of microarray data.

BACKGROUND: An increasing number of studies have profiled tumor specimens using distinct microarray platforms and analysis techniques. With the accumulating amount of microarray data, one of the most intriguing yet challenging tasks is to develop robust statistical models to integrate the findings. RESULTS: By applying a two-stage Bayesian mixture modeling strategy, we were able to assimilate and analyze four independent microarray studies to derive an inter-study validated "meta-signature" associated with breast cancer prognosis. Combining multiple studies (n = 305 samples) on a common probability scale, we developed a 90-gene meta-signature, which strongly associated with survival in breast cancer patients. Given the set of independent studies using different microarray platforms which included spotted cDNAs, Affymetrix GeneChip, and inkjet oligonucleotides, the individually identified classifiers yielded gene sets predictive of survival in each study cohort. The study-specific gene signatures, however, had minimal overlap with each other, and performed poorly in pairwise cross-validation. The meta-signature, on the other hand, accommodated such heterogeneity and achieved comparable or better prognostic performance when compared with the individual signatures. Further by comparing to a global standardization method, the mixture model based data transformation demonstrated superior properties for data integration and provided solid basis for building classifiers at the second stage. Functional annotation revealed that genes involved in cell cycle and signal transduction activities were over-represented in the meta-signature. CONCLUSION: The mixture modeling approach unifies disparate gene expression data on a common probability scale allowing for robust, inter-study validated prognostic signatures to be obtained. With the emerging utility of microarrays for cancer prognosis, it will be important to establish paradigms to meta-analyze disparate gene expression data for prognostic signatures of potential clinical use.

Bayes Theorem↗

Biomedical literature mining: challenges and solutions in the 'omics' era.

It is now obvious that the rate-limiting step in high throughput experimentation is neither data acquisition nor analysis, but rather our ability to interpret data on a genome-wide scale. Indeed, the explosion of data sampling capacity combined with increasing publication rates greatly impairs our ability to find meaning in vast collections of data. In order to support data interpretation, bioinformatic tools are needed to identify critical information contained in large bodies of literature. However, extracting knowledge embedded in free text is an arduous task, compounded in the biomedical field by an inconsistent gene nomenclature, domain-specific language and restricted access to full text articles. This paper presents a selection of currently available biomedical literature mining software. These tools rely on statistic and, more recently, semantic analyses (Natural Language Processing) to automatically extract information from the literature. In addition, a literature mining strategy has been developed to explore patterns of term occurrences in abstracts. This method automatically identifies relevant keywords in collections of abstracts, and uses a pattern discovery algorithm to generate a visual interface for exploring functional associations among genes. Term occurrence heatmaps can also be combined with gene expression profiles to provide valuable functional annotations. Furthermore, as demonstrated with tumor cell line literature profiling results, this approach can be applied to a variety of themes beyond genomic data analysis. Altogether, these examples illustrate how literature analysis can be employed to support knowledge discovery in biomedical research.

Algorithms↗

A functional approach to questions about life, death, and phosphorylation.

The success of the family of kinases as targets for small-molecule cancer therapeutics is probably best illustrated by the efficacy of the drug Gleevec. In spite of this, the function of many of the kinases in the mammalian genome remains unknown. In a recent paper, MacKeigan and colleagues report a functional genetic screen using RNA interference to identify kinases and phosphatases involved in programmed cell death (MacKeigan et al., 2005). Functional annotation is a prerequisite for selection of new drug targets. Such studies may therefore lay the foundation for the next generation of cancer drugs.

Antineoplastic Agents↗

Modeling a whole organ using proteomics: the avian bursa of Fabricius.

While advances in proteomics have improved proteome coverage and enhanced biological modeling, modeling function in multicellular organisms requires understanding how cells interact. Here we used the chicken bursa of Fabricius, a common experimental system for B cell function, to model organ function from proteomics data. The bursa has two major functional cell types: B cells and the supporting stromal cells. We used differential detergent fractionation-multidimensional protein identification technology (DDF-MudPIT) to identify 5198 proteins from all cellular compartments. Of these, 1753 were B cell specific, 1972 were stroma specific and 1473 were shared between the two. By modeling programmed cell death (PCD), cell differentiation and proliferation, and transcriptional activation, we have improved functional annotation of chicken proteins and placed chicken-specific death receptors into the PCD process using phylogenetics. We have identified 114 transcription factors (TFs); 42 of the bursal B cell TFs have not been reported before in any B cells. We have also improved the structural annotation of a newly sequenced genome by confirming the in vivo expression of 4006 "predicted", and 6623 ab initio, ORFs. Finally, we have developed a novel method for facilitating structural annotation, "expressed peptide sequence tags" (ePSTs) and demonstrate its utility by identifying 521 potential novel proteins from the chicken "unassigned chromosome".

Amino Acid Sequence↗

Catalog of gene expression in adult neural stem cells and their in vivo microenvironment.

Stem cells generally reside in a stem cell microenvironment, where cues for self-renewal and differentiation are present. However, the genetic program underlying stem cell proliferation and multipotency is poorly understood. Transcriptome analysis of stem cells and their in vivo microenvironment is one way of uncovering the unique stemness properties and provides a framework for the elucidation of stem cell function. Here, we characterize the gene expression profile of the in vivo neural stem cell microenvironment in the lateral ventricle wall of adult mouse brain and of in vitro proliferating neural stem cells. We have also analyzed an Lhx2-expressing hematopoietic-stem-cell-like cell line in order to define the transcriptome of a well-characterized and pure cell population with stem cell characteristics. We report the generation, assembly and annotation of 50,792 high-quality 5'-end expressed sequence tag sequences. We further describe a shared expression of 1065 transcripts by all three stem cell libraries and a large overlap with previously published gene expression signatures for neural stem/progenitor cells and other multipotent stem cells. The sequences and cDNA clones obtained within this framework provide a comprehensive resource for the analysis of genes in adult stem cells that can accelerate future stem cell research.

Animals↗

Small interfering RNA screens reveal enhanced cisplatin cytotoxicity in tumor cells having both BRCA network and TP53 disruptions.

RNA interference technology allows the systematic genetic analysis of the molecular alterations in cancer cells and how these alterations affect response to therapies. Here we used small interfering RNA (siRNA) screens to identify genes that enhance the cytotoxicity (enhancers) of established anticancer chemotherapeutics. Hits identified in drug enhancer screens of cisplatin, gemcitabine, and paclitaxel were largely unique to the drug being tested and could be linked to the drug's mechanism of action. Hits identified by screening of a genome-scale siRNA library for cisplatin enhancers in TP53-deficient HeLa cells were significantly enriched for genes with annotated functions in DNA damage repair as well as poorly characterized genes likely having novel functions in this process. We followed up on a subset of the hits from the cisplatin enhancer screen and validated a number of enhancers whose products interact with BRCA1 and/or BRCA2. TP53(+/-) matched-pair cell lines were used to determine if knockdown of BRCA1, BRCA2, or validated hits that associate with BRCA1 and BRCA2 selectively enhances cisplatin cytotoxicity in TP53-deficient cells. Silencing of BRCA1, BRCA2, or BRCA1/2-associated genes enhanced cisplatin cytotoxicity approximately 4- to 7-fold more in TP53-deficient cells than in matched TP53 wild-type cells. Thus, tumor cells having disruptions in BRCA1/2 network genes and TP53 together are more sensitive to cisplatin than cells with either disruption alone.

Antineoplastic Agents↗

cDNA microarray analysis of host-pathogen interactions in a porcine in vitro model for Toxoplasma gondii infection.

Toxoplasma gondii induces the expression of proinflammatory cytokines, reorganizes organelles, scavenges nutrients, and inhibits apoptosis in infected host cells. We used a cDNA microarray of 420 annotated porcine expressed sequence tags to analyze the molecular basis of these changes at eight time points over a 72-hour period in porcine kidney epithelial (PK13) cells infected with T. gondii. A total of 401 genes with Cy3 and Cy5 spot intensities of >/=500 were selected for analysis, of which 263 (65.6%) were induced >/=2-fold (expression ratio, >/=2.0; P </= 0.05 [t test]) over at least one time point and 48 (12%) were significantly down-regulated. At least 12 functional categories of genes were modulated (up- or down-regulated) by T. gondii. The majority of induced genes were clustered as transcription, signal transduction, host immune response, nutrient metabolism, and apoptosis related. The expression of selected genes altered by T. gondii was validated by quantitative real-time reverse transcription-PCR. These results suggest that significant changes in gene expression occur in response to T. gondii infection in PK13 cells, facilitating further analysis of host-pathogen interactions in toxoplasmosis in a secondary host.

Animals↗

A single-cell transcriptomic atlas of the pigtail macaque placenta in late gestation.

The placenta is a complex organ with multiple immune and non-immune cell types that promote fetal tolerance and facilitate the transfer of nutrients and oxygen. The nonhuman primate (NHP) is a key experimental model for studying human pregnancy complications, in part due to similarities in placental structure, which makes it essential to understand how single-cell populations compare across the human and NHP maternal-fetal interface. We constructed a single-cell RNA-Seq (scRNA-Seq) atlas of the placenta from the pigtail macaque ( Macaca nemestrina ) in the third trimester, comprising three different tissues at the maternal-fetal interface: the chorionic villi (placental disc), chorioamniotic membranes, and the maternal decidua. Each tissue was separately dissociated into single cells and processed through the 10X Genomics and Seurat pipeline, followed by aggregation, unsupervised clustering, and cluster annotation. Next, we determined the maternal-fetal origins of cell populations and analyzed single-cell RNA trajectory, Gene Ontology enrichment, and cell-cell communication. Single-cell populations in the pigtail macaque were strikingly similar in their identity and frequency to those found in the human placenta, including cells from trophoblast, stromal cell, immune, and macrophage lineages. An advantage of our approach was the deep sequencing of three tissues at the maternal-fetal interface, which yielded a rich diversity of common and rare single-cell populations. The third-trimester pigtail macaque single-cell atlas enables the identification of cellular subclusters analogous to those in humans and provides a powerful resource for understanding experimental perturbations on the NHP placenta.

Journal Article↗

Novel transcription factors in human CD34 antigen-positive hematopoietic cells.

Transcription factors (TFs) and the regulatory proteins that control them play key roles in hematopoiesis, controlling basic processes of cell growth and differentiation; disruption of these processes may lead to leukemogenesis. Here we attempt to identify functionally novel and partially characterized TFs/regulatory proteins that are expressed in undifferentiated hematopoietic tissue. We surveyed our database of 15 970 genes/expressed sequence tags (ESTs) representing the normal human CD34(+) cells transcriptosome (http://westsun.hema.uic.edu/cd34.html), using the UniGene annotation text descriptor, to identify genes with motifs consistent with transcriptional regulators; 285 genes were identified. We also extracted the human homologues of the TFs reported in the murine stem cell database (SCdb; http://stemcell.princeton.edu/), selecting an additional 45 genes/ESTs. An exhaustive literature search of each of these 330 unique genes was performed to determine if any had been previously reported and to obtain additional characterizing information. Of the resulting gene list, 106 were considered to be potential TFs. Overall, the transcriptional regulator dataset consists of 165 novel or poorly characterized genes, including 25 that appeared to be TFs. Among these novel and poorly characterized genes are a cell growth regulatory with ring finger domain protein (CGR19, Hs.59106), an RB-associated CRAB repressor (RBAK, Hs.7222), a death-associated transcription factor 1 (DATF1, Hs.155313), and a p38-interacting protein (P38IP, Hs. 171185). The identification of these novel and partially characterized potential transcriptional regulators adds a wealth of information to understanding the molecular aspects of hematopoiesis and hematopoietic disorders.

Amino Acid Motifs↗

A High-Resolution Stereo-Seq Spatial Transcriptomic Resource for Adult Holstein Cattle Liver.

The bovine liver is a highly compartmentalized organ that plays essential roles in continuous gluconeogenesis and nitrogen recycling; however, its spatial molecular architecture has remained largely uncharacterized due to the limitations of traditional bulk and single-cell approaches. To address this gap, Spatial Enhanced Resolution Omics-sequencing (Stereo-seq) was utilized to generate a subcellular-resolution (500 nm) transcriptomic map of an adult Holstein cattle liver, and a refined reference-guided workflow was implemented to overcome standard annotation limitations in livestock. Raw sequencing data were processed using the Stereo-seq Analysis Workflow and analyzed with Stereopy, Seurat, SingleR, and reference-guided workflows. Spatial aggregation was evaluated at Bin20, Bin50, Bin100, Bin150, and Bin200. Increasing bin size increased molecular identifier counts and detected-gene complexity while progressively reducing spatial granularity. Bin50, corresponding to 50 &#xd7; 50 DNA nanoballs and an approximate nominal footprint of 25 &#xd7; 25 &#xb5;m, was therefore selected as a practical intermediate aggregation level for the primary analyses. Quality-control assessment, Leiden clustering, UMAP visualization, reference-based cell-type annotation, cluster-marker analysis, and spatial mapping of canonical hepatic genes demonstrated preservation of biologically interpretable liver transcriptional organization. Raw sequencing data processed spatial matrices, annotated objects, and analysis code are publicly available to support reanalysis and computational benchmarking. In summary, we present a Stereo-seq spatial transcriptomic resource generated from liver tissue of an adult Holstein cow. This initial resource provides a valuable foundation for future studies of bovine liver biology, comparative genomics, and the spatial basis of livestock health and production traits.

Animals↗

GeneExt: a gene model extension tool for enhanced single-cell RNA-seq analysis.

MOTIVATION: Incomplete gene models negatively impact single-cell gene expression quantification. This is particularly true in non-model species where often gene 3' ends are inaccurately annotated, while most scRNA-seq methods only capture the 3' transcript region. This results in many genes being incorrectly quantified or not detected. RESULTS: GeneExt leverages scRNA-seq data to refine gene annotations. We exemplify GeneExt usage and its impact on the gene expression quantification of eight non-model organism single-cell atlases. By extending and homogenizing gene annotations, our tool will help improve biological interpretation and cross-species comparisons of cell type expression atlases. AVAILABILITY: GeneExt is available at https://github.com/sebepedroslab/GeneExt (DOI: https://doi.org/10.5281/zenodo.18712940) under a GNU General Public license, together with test data and usage instructions.

Software↗

Saccharomyces Genome Database (SGD) provides secondary gene annotation using the Gene Ontology (GO).

The Saccharomyces Genome Database (SGD) resources, ranging from genetic and physical maps to genome-wide analysis tools, reflect the scientific progress in identifying genes and their functions over the last decade. As emphasis shifts from identification of the genes to identification of the role of their gene products in the cell, SGD seeks to provide its users with annotations that will allow relationships to be made between gene products, both within Saccharomyces cerevisiae and across species. To this end, SGD is annotating genes to the Gene Ontology (GO), a structured representation of biological knowledge that can be shared across species. The GO consists of three separate ontologies describing molecular function, biological process and cellular component. The goal is to use published information to associate each characterized S.cerevisiae gene product with one or more GO terms from each of the three ontologies. To be useful, this must be done in a manner that allows accurate associations based on experimental evidence, modifications to GO when necessary, and careful documentation of the annotations through evidence codes for given citations. Reaching this goal is an ongoing process at SGD. For information on the current progress of GO annotations at SGD and other participating databases, as well as a description of each of the three ontologies, please visit the GO Consortium page at http://www.geneontology.org. SGD gene associations to GO can be found by visiting our site at http://genome-www.stanford.edu/Saccharomyces/.

Animals↗

Novel RNAs identified from an in-depth analysis of the transcriptome of human chromosomes 21 and 22.

In this report, we have achieved a richer view of the transcriptome for Chromosomes 21 and 22 by using high-density oligonucleotide arrays on cytosolic poly(A)(+) RNA. Conservatively, only 31.4% of the observed transcribed nucleotides correspond to well-annotated genes, whereas an additional 4.8% and 14.7% correspond to mRNAs and ESTs, respectively. Approximately 85% of the known exons were detected, and up to 21% of known genes have only a single isoform based on exon-skipping alternative expression. Overall, the expression of the well-characterized exons falls predominately into two categories, uniquely or ubiquitously expressed with an identifiable proportion of antisense transcripts. The remaining observed transcription (49.0%) was outside of any known annotation. These novel transcripts appear to be more cell-line-specific and have lower and less variation in expression than the well-characterized genes. Novel transcripts were further characterized based on their distance to annotations, transcript size, coding capacity, and identification as antisense to intronic sequences. By RT-PCR, 126 novel transcripts were independently verified, resulting in a 65% verification rate. These observations strongly support the argument for a re-evaluation of the total number of human genes and an alternative term for "gene" to encompass these growing, novel classes of RNA transcripts in the human genome.

Cell Line↗

Cell-specific DNA methylation in human alpha and beta cells regulates gene expression in type 2 diabetes.

Epigenome-wide studies of pancreatic islets provide valuable insights into type 2 diabetes (T2D) but lack methylomes from individual cell types. Here we show changes to alpha and beta cell-specific methylomes and transcriptomes from people with or without T2D, using whole-genome bisulfite sequencing and RNA sequencing. We discover 22,544 differentially methylated regions annotated to 7,975 genes in alpha versus beta cells, such as INS, GCG, PDX1 and PCSK1, with ~50% showing differential expression. CRISPR-dCas9-DNMT3A-based epigenetic editing increases INS and TH DNA methylation, while CRISPR-dCas9-TET1-based editing decreases GCG methylation, each altering INS, TH or GCG expression and content in beta cells. Pre-T2D/T2D-associated differentially methylated regions in alpha and beta cells overlap 12-18% of T2D-associated genome-wide association study candidates. Additionally, ONECUT2 is epigenetically upregulated in beta cells from people with pre-T2D/T2D and elevated in male Goto-Kakizaki rat islets. ONECUT2 overexpression in beta cells/islets downregulates gene sets impacting insulin secretion and glucose homeostasis, and reduces mitochondrial activity, ATP/ADP ratio and insulin secretion. We also provide 'alpha-beta-methylome' ( https://alpha-beta-methylome.serve.scilifelab.se/app/alpha-beta-methylome/ ), a resource exploring T2D, age and sex associations on methylation, highlighting cell-specific epigenetic regulation and dysfunctions contributing to T2D.

Humans↗