Search PubMed⌕ Search

Biomedical subjects

Christian J Stoeckert

Publications and source records attributed to Christian J Stoeckert.

16 recordsLinked to original sources

Fidelity and enhanced sensitivity of differential transcription profiles following linear amplification of nanogram amounts of endothelial mRNA.

Although mRNA amplification is necessary for microarray analyses from limited amounts of cells and tissues, the accuracy of transcription profiles following amplification has not been well characterized. We tested the fidelity of differential gene expression following linear amplification by T7-mediated transcription in a well-established in vitro model of cytokine [tumor necrosis factor alpha (TNFalpha)]-stimulated human endothelial cells using filter arrays of 13,824 human cDNAs. Transcriptional profiles generated from amplified antisense RNA (aRNA) (from 100 ng total RNA, approximately 1 ng mRNA) were compared with profiles generated from unamplified RNA originating from the same homogeneous pool. Amplification accurately identified TNFalpha-induced differential expression in 94% of the genes detected using unamplified samples. Furthermore, an additional 1,150 genes were identified as putatively differentially expressed using amplified RNA which remained undetected using unamplified RNA. Of genes sampled from this set, 67% were validated by quantitative real-time PCR as truly differentially expressed. Thus, in addition to demonstrating fidelity in gene expression relative to unamplified samples, linear amplification results in improved sensitivity of detection and enhances the discovery potential of high-throughput screening by microarrays.

Bias↗

Integrating computationally assembled mouse transcript sequences with the Mouse Genome Informatics (MGI) database.

Databases of experimentally generated and computationally derived transcript sequences are valuable resources for genome analysis and annotation. The utility of such databases is enhanced when the sequences they contain are integrated with such biological information as genomic location, gene function, gene expression and phenotypic variation. We present the analysis and results of a semi-automated process of connecting transcript assemblies with highly curated biological information for mouse genes that is available through the Mouse Genome Informatics (MGI) database.

Animals↗

PlasmoDB: the Plasmodium genome resource. A database integrating experimental and computational data.

PlasmoDB (http://PlasmoDB.org) is the official database of the Plasmodium falciparum genome sequencing consortium. This resource incorporates the recently completed P. falciparum genome sequence and annotation, as well as draft sequence and annotation emerging from other Plasmodium sequencing projects. PlasmoDB currently houses information from five parasite species and provides tools for intra- and inter-species comparisons. Sequence information is integrated with other genomic-scale data emerging from the Plasmodium research community, including gene expression analysis from EST, SAGE and microarray projects and proteomics studies. The relational schema used to build PlasmoDB, GUS (Genomics Unified Schema) employs a highly structured format to accommodate the diverse data types generated by sequence and expression projects. A variety of tools allow researchers to formulate complex, biologically-based, queries of the database. A stand-alone version of the database is also available on CD-ROM (P. falciparum GenePlot), facilitating access to the data in situations where internet access is difficult (e.g. by malaria researchers working in the field). The goal of PlasmoDB is to facilitate utilization of the vast quantities of genomic-scale data produced by the global malaria research community. The software used to develop PlasmoDB has been used to create a second Apicomplexan parasite genome database, ToxoDB (http://ToxoDB.org).

Animals↗

Gene discovery in the apicomplexa as revealed by EST sequencing and assembly of a comparative gene database.

Large-scale EST sequencing projects for several important parasites within the phylum Apicomplexa were undertaken for the purpose of gene discovery. Included were several parasites of medical importance (Plasmodium falciparum, Toxoplasma gondii) and others of veterinary importance (Eimeria tenella, Sarcocystis neurona, and Neospora caninum). A total of 55192 ESTs, deposited into dbEST/GenBank, were included in the analyses. The resulting sequences have been clustered into nonredundant gene assemblies and deposited into a relational database that supports a variety of sequence and text searches. This database has been used to compare the gene assemblies using BLAST similarity comparisons to the public protein databases to identify putative genes. Of these new entries, approximately 15%-20% represent putative homologs with a conservative cutoff of p < 10(-9), thus identifying many conserved genes that are likely to share common functions with other well-studied organisms. Gene assemblies were also used to identify strain polymorphisms, examine stage-specific expression, and identify gene families. An interesting class of genes that are confined to members of this phylum and not shared by plants, animals, or fungi, was identified. These genes likely mediate the novel biological features of members of the Apicomplexa and hence offer great potential for biological investigation and as possible therapeutic targets.

Animals↗

Transcriptional program of the endocrine pancreas in mice and humans.

The Endocrine Pancreas Consortium was formed in late 1999 to derive and sequence cDNA libraries enriched for rare transcripts expressed in the mammalian endocrine pancreas. Over the past 3 years, the Consortium has generated 20 cDNA libraries from mouse and human pancreatic tissues and deposited >150,000 sequences into the public expressed sequence tag databases. A special effort was made to enrich for cDNAs from the endocrine pancreas by constructing libraries from isolated islets. In addition, we constructed a library in which fetal pancreas from Neurogenin 3 null mice, which consists of only exocrine and duct cells, was subtracted from fetal wild-type pancreas to enrich for the transcripts from the endocrine compartment. Sequence analysis showed that these clones cluster into 9,464 assembly groups (approximating unique transcripts) for the mouse and 13,910 for the human sequences. Of these, >4,300 were unique to Consortium libraries. We have assembled a core clone set containing one cDNA for each assembly group for the mouse and have constructed the corresponding microarray, termed "PancChip 4.0," which contains >9,000 nonredundant elements. We show that this PancChip is highly enriched for genes expressed in the endocrine pancreas. The mouse and human clone sets and corresponding arrays will be important resources for diabetes research.

Animals↗

A molecular profile of a hematopoietic stem cell niche.

The hematopoietic microenvironment provides a complex molecular milieu that regulates the self-renewal and differentiation activities of stem cells. We have characterized a stem cell supportive stromal cell line, AFT024, that was derived from murine fetal liver. Highly purified in vivo transplantable mouse stem cells are maintained in AFT024 cultures at input levels, whereas other primitive progenitors are expanded. In addition, human stem cells are very effectively supported by AFT024. We suggest that the AFT024 cell line represents a component of an in vivo stem cell niche. To determine the molecular signals elaborated in this niche, we undertook a functional genomics approach that combines extensive sequence mining of a subtracted cDNA library, high-density array hybridization and in-depth bioinformatic analyses. The data have been assembled into a biological process oriented database, and represent a molecular profile of a candidate stem cell niche.

Amino Acid Sequence↗

Comparison of different labeling methods for two-channel high-density microarray experiments.

In this report we evaluate three methods for labeling nucleic acids to be hybridized to a cDNA microarray: direct labeling, indirect amino-allyl labeling, and the dendrimer labeling method (Genisphere). The dendrimer method requires the smallest quantity of sample, 2.5 microg of total RNA compared with 20 microg with the direct or indirect methods. Therefore, we wanted to know whether the performance of the dendrimer method is comparable to the other methods, or whether significant information is lost. Performance can be considered in terms of sensitivity, dynamic range, and reproducibility of the quantitative signals for gene intensity. We compared the three labeling methods by generating three sets of eight self-to-self hybridizations using the same total RNA sample in all cases ("replicate study"). In our analysis, we controlled for the effects of print-tip and background subtraction biases. We also performed a smaller study, namely, a dilution series study with five dilution points per labeling method, to evaluate one aspect of predictive ability. From the replicate study, the dendrimer method appeared to perform as well, and often better, with respect to reproducibility and ability to detect expression. However, in the dilution series study, this method was outperformed by the other two in terms of predictive ability and did not perform very well. These findings are helping to guide our decisions on what labeling method to use for subsequent studies, based on the purpose of a specific study and its limitations in terms of available material.

Fluorescent Dyes↗

Design and implementation of microarray gene expression markup language (MAGE-ML).

BACKGROUND: Meaningful exchange of microarray data is currently difficult because it is rare that published data provide sufficient information depth or are even in the same format from one publication to another. Only when data can be easily exchanged will the entire biological community be able to derive the full benefit from such microarray studies. RESULTS: To this end we have developed three key ingredients towards standardizing the storage and exchange of microarray data. First, we have created a minimal information for the annotation of a microarray experiment (MIAME)-compliant conceptualization of microarray experiments modeled using the unified modeling language (UML) named MAGE-OM (microarray gene expression object model). Second, we have translated MAGE-OM into an XML-based data format, MAGE-ML, to facilitate the exchange of data. Third, some of us are now using MAGE (or its progenitors) in data production settings. Finally, we have developed a freely available software tool kit (MAGE-STK) that eases the integration of MAGE-ML into end users' systems. CONCLUSIONS: MAGE will help microarray data producers and users to exchange information by providing a common platform for data exchange, and MAGE-STK will make the adoption of MAGE easier.

Computer Simulation↗

PlasmoDB: the Plasmodium genome resource. An integrated database providing tools for accessing, analyzing and mapping expression and sequence data (both finished and unfinished).

PlasmoDB (http://PlasmoDB.org) is the official database of the Plasmodium falciparum genome sequencing consortium. This resource incorporates finished and draft genome sequence data and annotation emerging from Plasmodium sequencing projects. PlasmoDB currently houses information from five parasite species and provides tools for cross-species comparisons. Sequence information is also integrated with other genomic-scale data emerging from the Plasmodium research community, including gene expression analysis from EST, SAGE and microarray projects. The relational schemas used to build PlasmoDB [Genomics Unified Schema (GUS) and RNA Abundance Database (RAD)] employ a highly structured format to accommodate the diverse data types generated by sequence and expression projects. A variety of tools allow researchers to formulate complex, biologically based queries of the database. A version of the database is also available on CD-ROM (Plasmodium GenePlot), facilitating access to the data in situations where Internet access is difficult (e.g. by malaria researchers working in the field). The goal of PlasmoDB is to enhance utilization of the vast quantities of data emerging from genome-scale projects by the global malaria research community.

Animals↗

Microarray databases: standards and ontologies.

A single microarray can provide information on the expression of tens of thousands of genes. The amount of information generated by a microarray-based experiment is sufficiently large that no single study can be expected to mine each nugget of scientific information. As a consequence, the scale and complexity of microarray experiments require that computer software programs do much of the data processing, storage, visualization, analysis and transfer. The adoption of common standards and ontologies for the management and sharing of microarray data is essential and will provide immediate benefit to the research community.

Database Management Systems↗

Predicting gene ontology functions from ProDom and CDD protein domains.

A heuristic algorithm for associating Gene Ontology (GO) defined molecular functions to protein domains as listed in the ProDom and CDD databases is described. The algorithm generates rules for function-domain associations based on the intersection of functions assigned to gene products by the GO consortium that contain ProDom and/or CDD domains at varying levels of sequence similarity. The hierarchical nature of GO molecular functions is incorporated into rule generation. Manual review of a subset of the rules generated indicates an accuracy rate of 87% for ProDom rules and 84% for CDD rules. The utility of these associations is that novel sequences can be assigned a putative function if sufficient similarity exists to a ProDom or CDD domain for which one or more GO functions has been associated. Although functional assignments are increasingly being made for gene products from model organisms, it is likely that the needs of investigators will continue to outpace the efforts of curators, particularly for nonmodel organisms. A comparison with other methods in terms of coverage and agreement was performed, indicating the utility of the approach. The domain-function associations and function assignments are available from our website http://www.cbil.upenn.edu/GO.

Algorithms↗

Functional genomics of the endocrine pancreas: the pancreas clone set and PancChip, new resources for diabetes research.

Over the past 5 years, microarrays have greatly facilitated large-scale analysis of gene expression levels. Although these arrays were not specifically geared to represent tissues and pathways known to be affected by diabetes, they have been used in both type 1 and type 2 diabetes research. To prepare a tool that is particularly useful in the study of type 1 diabetes, we have assembled a nonredundant set of 3,400 clones representing genes expressed in the mouse pancreas or pathways known to be affected by diabetes. We have demonstrated the usefulness of this clone set by preparing a cDNA glass microarray, the PancChip, and using it to analyze pancreatic gene expression from embryonic day 14.5 through adulthood in mice. The clone set and corresponding array are useful resources for diabetes research.

Adult↗