Search PubMed⌕ Search

Biomedical subjects

J Quackenbush

Publications and source records attributed to J Quackenbush.

At least 19 recordsLinked to original sources

Identification of tumor markers in models of human colorectal cancer using a 19,200-element complementary DNA microarray.

Metastasis represents a crucial transition in disease development and progression and has a profound impact on survival for a wide variety of cancers. Cell line models of metastasis have played an important role in developing our understanding of the metastatic process. We used a 19,200-element human cDNA microarray to profile transcription in three paired cell-line models of colorectal tumor metastasis. By correlating expression patterns across these cell lines, we have identified 176 genes that appear to be differentially expressed (greater than 2-fold) in all highly metastatic cell lines relative to their reference. An analysis of these genes reiterates much of our understanding of the metastatic process and suggests additional genes, many of previously uncharacterized function, that may be causatively involved in, or at least prognostic of, metastasis. Northern analysis of a limited number of these genes validates the observed pattern of expression and suggests that further investigation and functional characterization of the identified genes is warranted.

Biomarkers, Tumor↗

Exploring the transcriptome of the malaria sporozoite stage.

Most studies of gene expression in Plasmodium have been concerned with asexual and/or sexual erythrocytic stages. Identification and cloning of genes expressed in the preerythrocytic stages lag far behind. We have constructed a high quality cDNA library of the Plasmodium sporozoite stage by using the rodent malaria parasite P. yoelii, an important model for malaria vaccine development. The technical obstacles associated with limited amounts of RNA material were overcome by PCR-amplifying the transcriptome before cloning. Contamination with mosquito RNA was negligible. Generation of 1,972 expressed sequence tags (EST) resulted in a total of 1,547 unique sequences, allowing insight into sporozoite gene expression. The circumsporozoite protein (CS) and the sporozoite surface protein 2 (SSP2) are well represented in the data set. A BLASTX search with all tags of the nonredundant protein database gave only 161 unique significant matches (P(N) < or = 10(-4)), whereas 1,386 of the unique sequences represented novel sporozoite-expressed genes. We identified ESTs for three proteins that may be involved in host cell invasion and documented their expression in sporozoites. These data should facilitate our understanding of the preerythrocytic Plasmodium life cycle stages and the development of preerythrocytic vaccines.

Amino Acid Motifs↗

Functional annotation of a full-length mouse cDNA collection.

The RIKEN Mouse Gene Encyclopaedia Project, a systematic approach to determining the full coding potential of the mouse genome, involves collection and sequencing of full-length complementary DNAs and physical mapping of the corresponding genes to the mouse genome. We organized an international functional annotation meeting (FANTOM) to annotate the first 21,076 cDNAs to be analysed in this project. Here we describe the first RIKEN clone collection, which is one of the largest described for any organism. Analysis of these cDNAs extends known gene families and identifies new ones.

Animals↗

The TIGR Gene Indices: analysis of gene transcript sequences in highly sampled eukaryotic species.

While genome sequencing projects are advancing rapidly, EST sequencing and analysis remains a primary research tool for the identification and categorization of gene sequences in a wide variety of species and an important resource for annotation of genomic sequence. The TIGR Gene Indices (http://www.tigr.org/tdb/tgi. shtml) are a collection of species-specific databases that use a highly refined protocol to analyze EST sequences in an attempt to identify the genes represented by that data and to provide additional information regarding those genes. Gene Indices are constructed by first clustering, then assembling EST and annotated gene sequences from GenBank for the targeted species. This process produces a set of unique, high-fidelity virtual transcripts, or Tentative Consensus (TC) sequences. The TC sequences can be used to provide putative genes with functional annotation, to link the transcripts to mapping and genomic sequence data, to provide links between orthologous and paralogous genes and as a resource for comparative sequence analysis.

Animals↗

Effects of ischemia on gene expression.

Microarray gene expression technology has recently made it feasible to characterize the RNA expression of thousands of genes across numerous tissue samples. We hypothesized that the warm ischemia commonly associated with the surgical extirpation of human tissue would have significant effects on gene expression profiles. To quantitate the effects of warm ischemia on human tissue, we rapidly dissected normal mucosa from a human colon cancer specimen. The specimen was divided and maintained at room temperature until snap-frozen in liquid nitrogen. Aliquots of tissue were frozen at times 5, 10, 15, 20, 40, and 60 min after extirpation. Spotted microarrays composed of 2400 distinct elements were used to assay mRNA derived from each time point in triplicate. Eisen's hierarchical clustering methodology and Bayesean statistical methods were then used to assay the effects of warm ischemia on gene expression. Application of time-course statistical models suggest that three patterns were induced by ischemia, accounting for 68.2, 17.8, and 13.4% of the evaluable genes, respectively. Pattern I corresponds to an average change of 27% over 60 min from 5 min baseline level of expression and 63.8% of the genes with at least 80% probability of membership in this pattern show average increases in expression over 60 min. The remainder decrease on average. Pattern II genes show the least ischemia-related effects, demonstrating an average change of only 12% over 60 min. In contrast to pattern I, we find that 67.5% of the genes with at least 80% probability of membership in this pattern are decreasing in expression on average over time. The remaining 32.5% in this pattern increase an average of 12% over 60 min. Finally, pattern III genes (13.4% of the sample) show the greatest sensitivity to ischemia, changing an average of 50% over 60 min, with about the same number increasing as are decreasing. Fold changes in RNA over- or under-expression were observed up to greater than 20-fold. Warm ischemia associated with the surgical extirpation of human tissues has significant effects on gene expression. These data support the careful monitoring of ischemic time for tissues harvested for the purpose of gene profiling.

Colonic Neoplasms↗

Characterization of a Vibrio vulnificus LysR homologue, HupR, which regulates expression of the haem uptake outer membrane protein, HupA.

In Vibrio vulnificus, the ability to acquire iron from the host has been shown to correlate with virulence. Here, we show that the DNA upstream of hupA (haem uptake receptor) in V. vulnificus encodes a protein in the inverse orientation to hupA (named hupR). HupR shares homology with the LysR family of positive transcriptional activators. A hupA-lacZ fusion contained on a plasmid was transformed into Fur(-), Fur(+)and HupR(-)strains of V. vulnificus. The beta-galactosidase assays and Northern blot analysis showed that transcription of hupA is negatively regulated by iron and the Fur repressor in V. vulnificus. Under low-iron conditions with added haemin, the expression of hupA in the hupR mutant was significantly lower than in the wild-type. This diminished response to haem was detected by both Northern blot and hupA-lacZ fusion analysis. The haem response of hupA in the hupR mutant was restored to wild-type levels when complemented with hupR in trans. These studies suggest that HupR may act as a positive regulator of hupA transcription under low-iron conditions in the presence of haemin.

Amino Acid Sequence↗

Computational analysis of microarray data.

Microarray experiments are providing unprecedented quantities of genome-wide data on gene-expression patterns. Although this technique has been enthusiastically developed and applied in many biological contexts, the management and analysis of the millions of data points that result from these experiments has received less attention. Sophisticated computational tools are available, but the methods that are used to analyse the data can have a profound influence on the interpretation of the results. A basic understanding of these computational tools is therefore required for optimal experimental design and meaningful data analysis.

Algorithms↗

Minimum information about a microarray experiment (MIAME)-toward standards for microarray data.

Microarray analysis has become a widely used tool for the generation of gene expression data on a genomic scale. Although many significant results have been derived from microarray studies, one limitation has been the lack of standards for presenting and exchanging such data. Here we present a proposal, the Minimum Information About a Microarray Experiment (MIAME), that describes the minimum information required to ensure that microarray data can be easily interpreted and that results derived from its analysis can be independently verified. The ultimate goal of this work is to establish a standard for recording and reporting microarray-based gene expression data, which will in turn facilitate the establishment of databases and public repositories and enable the development of data analysis tools. With respect to MIAME, we concentrate on defining the content and structure of the necessary information rather than the technical format for capturing it.

Computational Biology↗

Sequence evaluation of four pooled-tissue normalized bovine cDNA libraries and construction of a gene index for cattle.

An essential component of functional genomics studies is the sequence of DNA expressed in tissues of interest. To provide a resource of bovine-specific expressed sequence data and facilitate this powerful approach in cattle research, four normalized cDNA libraries were produced and arrayed for high-throughput sequencing. The libraries were made with RNA pooled from multiple tissues to increase efficiency of normalization and maximize the number of independent genes for which sequence data were obtained. Target tissues included those with highest likelihood to have impact on production parameters of animal health, growth, reproductive efficiency, and carcass merit. Success of normalization and inter- and intralibrary redundancy were assessed by collecting 6000-23,000 sequences from each of the libraries (68,520 total sequences deposited in GenBank). Sequence comparison and assembly of these sequences was performed in combination with 56,500 other bovine EST sequences present in the GenBank dbEST database to construct a cattle Gene Index (available from The Institute for Genomic Research at http://www.tigr.org/tdb/tgi.shtml). The 124,381 bovine ESTs present in GenBank at the time of the analysis form 16,740 assemblies that are listed and annotated on the Web site. Analysis of individual library sequence data indicates that the pooled-tissue approach was highly effective in preparing libraries for efficient deep sequencing.

Animals↗

Rice bioinformatics. analysis of rice sequence data and leveraging the data to other plant species.

Rice (Oryza sativa) is a model species for monocotyledonous plants, especially for members in the grass family. Several attributes such as small genome size, diploid nature, transformability, and establishment of genetic and molecular resources make it a tractable organism for plant biologists. With an estimated genome size of 430 Mb (Arumuganathan and Earle, 1991), it is feasible to obtain the complete genome sequence of rice using current technologies. An international effort has been established and is in the process of sequencing O. sativa spp. japonica var "Nipponbare" using a bacterial artificial chromosome/P1 artificial chromosome shotgun sequencing strategy. Annotation of the rice genome is performed using prediction-based and homology-based searches to identify genes. Annotation tools such as optimized gene prediction programs are being developed for rice to improve the quality of annotation. Resources are also being developed to leverage the rice genome sequence to partial genome projects such as expressed sequence tag projects, thereby maximizing the output from the rice genome project. To provide a low level of annotation for rice genomic sequences, we have aligned all rice bacterial artificial chromosome/P1 artificial chromosome sequences with The Institute of Genomic Research Gene Indices that are a set of nonredundant transcripts that are generated from nine public plant expressed sequence tag projects (rice, wheat, sorghum, maize, barley, Arabidopsis, tomato, potato, and barrel medic). In addition, we have used data from The Institute of Genomic Research Gene Indices and the Arabidopsis and Rice Genome Projects to identify putative orthologues and paralogues among these nine genomes.

Base Sequence↗

Microarray analysis of the in vivo effects of hypophysectomy and growth hormone treatment on gene expression in the rat.

Complementary DNA microarrays containing 3000 different rat genes were used to study the consequences of severe hormonal deficiency (hypophysectomy) on the gene expression patterns in heart, liver, and kidney. Hybridization signals were seen from a majority of the arrayed complementary DNAs; nonetheless, tissue-specific expression patterns could be delineated. Hypophysectomy affected the expression of genes involved in a variety of cellular functions. Between 16-29% of the detected transcripts from each tissue changed expression level as a reaction to this condition. Chronic treatment of hypophysectomized animals with human GH also caused significant changes in gene expression patterns. The study confirms previous knowledge concerning certain gene expression changes in the above-mentioned situations and provides new information regarding hypophysectomy and chronic human GH effects in the rat. Furthermore, we have identified several new genes that respond to GH treatment. Our results represent a first step toward a more global understanding of gene expression changes in states of hormonal deficiency.

Animals↗

Anchoring of rice BAC clones to the rice genetic map in silico.

A wealth of molecular resources have been developed for rice genomics, including dense genetic maps, expressed sequence tags (ESTs), yeast artificial chromosome maps, bacterial artificial chromosome (BAC) libraries and BAC end sequence databases. Integration of genetic and physical maps involves labor-intensive empirical experiments. To accelerate the integration of the bacterial clone resources with the genetic map for the International Rice Genome Sequencing Project, we cleaned and filtered the available EST and BAC end sequences for repetitive sequences and then searched all available rice genetic markers with our filtered databases. We identified 418 genetic markers that aligned with at least one BAC end sequence with >95% sequence identity, providing a set of large insert clones with an average separation of 1 Mb that can serve as nucleation points for the sequencing phase of the International Rice Genome Sequencing Project.

Chromosome Mapping↗

An optimized protocol for analysis of EST sequences.

The vast body of Expressed Sequence Tag (EST) data in the public databases provide an important resource for comparative and functional genomics studies and an invaluable tool for the annotation of genomic sequences. We have developed a rigorous protocol for reconstructing the sequences of transcribed genes from EST and gene sequence fragments. A key element in developing this protocol has been the evaluation of a number of sequence assembly programs to determine which most faithfully reproduce transcript sequences from EST data. The TIGR Gene Indices constructed using this protocol for human, mouse, rat and a variety of other plant and animal models have demonstrated their utility in a variety of applications and are freely available to the scientific research community.

Algorithms↗

The African trypanosome genome.

The haploid nuclear genome of the African trypanosome, Trypanosoma brucei, is about 35 Mb and varies in size among different trypanosome isolates by as much as 25%. The nuclear DNA of this diploid organism is distributed among three size classes of chromosomes: the megabase chromosomes of which there are at least 11 pairs ranging from 1 Mb to more than 6 Mb (numbered I-XI from smallest to largest); several intermediate chromosomes of 200-900 kb and uncertain ploidy; and about 100 linear minichromosomes of 50-150 kb. Size differences of as much as four-fold can occur, both between the two homologues of a megabase chromosome pair in a specific trypanosome isolate and among chromosome pairs in different isolates. The genomic DNA sequences determined to date indicated that about 50% of the genome is coding sequence. The chromosomal telomeres possess TTAGGG repeats and many, if not all, of the telomeres of the megabase and intermediate chromosomes are linked to expression sites for genes encoding variant surface glycoproteins (VSGs). The minichromosomes serve as repositories for VSG genes since some but not all of their telomeres are linked to unexpressed VSG genes. A gene discovery program, based on sequencing the ends of cloned genomic DNA fragments, has generated more than 20 Mb of discontinuous single-pass genomic sequence data during the past year, and the complete sequences of chromosomes I and II (about 1 Mb each) in T. brucei GUTat 10.1 are currently being determined. It is anticipated that the entire genomic sequence of this organism will be known in a few years. Analysis of a test microarray of 400 cDNAs and small random genomic DNA fragments probed with RNAs from two developmental stages of T. brucei demonstrates that the microarray technology can be used to identify batteries of genes differentially expressed during the various life cycle stages of this parasite.

Animals↗

The TIGR gene indices: reconstruction and representation of expressed gene sequences.

Expressed sequence tags (ESTs) have provided a first glimpse of the collection of transcribed sequences in a variety of organisms. However, a careful analysis of this sequence data can provide significant additional functional, structural and evolutionary information. Our analysis of the public EST sequences, available through the TIGR Gene Indices (TGI; http://www.tigr.org/tdb/tdb.html ), is an attempt to identify the genes represented by that data and to provide additional information regarding those genes. Gene Indices are constructed for selected organisms by first clustering, then assembling EST and annotated gene sequences from GenBank. This process produces a set of unique, high-fidelity virtual transcripts, or tentative consensus (TC) sequences. The TC sequences can be used to provide putative genes with functional annotation, to link the transcripts to mapping and genomic sequence data, and to provide links between orthologous and paralogous genes.

Base Sequence↗

Gene index analysis of the human genome estimates approximately 120,000 genes.

Although sequencing of the human genome will soon be completed, gene identification and annotation remains a challenge. Early estimates suggested that there might be 60,000-100,000 (ref. 1) human genes, but recent analyses of the available data from EST sequencing projects have estimated as few as 45,000 (ref. 2) or as many as 140, 000 (ref. 3) distinct genes. The Chromosome 22 Sequencing Consortium estimated a minimum of 45,000 genes based on their annotation of the complete chromosome, although their data suggests there may be additional genes. The nearly 2,000,000 human ESTs in dbEST provide an important resource for gene identification and genome annotation, but these single-pass sequences must be carefully analysed to remove contaminating sequences, including those from genomic DNA, spurious transcription, and vector and bacterial sequences. We have developed a highly refined and rigorously tested protocol for cleaning, clustering and assembling EST sequences to produce high-fidelity consensus sequences for the represented genes (F.L. et al., manuscript submitted) and used this to create the TIGR Gene Indices-databases of expressed genes for human, mouse, rat and other species (http://www.tigr.org/tdb/tgi.html). Using highly refined and tested algorithms for EST analysis, we have arrived at two independent estimates indicating the human genome contains approximately 120,000 genes.

Algorithms↗