Search PubMed⌕ Search

Biomedical subjects

David J States

Publications and source records attributed to David J States.

4 recordsLinked to original sources

Computationally identifying novel NF-kappa B-regulated immune genes in the human genome.

Identifying novel NF-kappa B-regulated immune genes in the human genome is important to our understanding of immune mechanisms and immune diseases. We fit logistic regression models to the promoters of 62 known NF-kappa B-regulated immune genes, to find patterns of transcription factor binding in the promoters of genes with known immune function. Using these patterns, we scanned the promoters of additional genes to find matches to the patterns, selected those with NF-kappa B binding sites conserved in the mouse or fly, and then confirmed them as NF-kappa B-regulated immune genes based on expression data. Among 6440 previously identified promoters in the human genome, we found 28 predicted immune gene promoters, 19 of which regulate genes with known function, allowing us to calculate specificity of 93%-100% for the method. We calculated sensitivity of 42% when searching the 62 known immune gene promoters. We found nine novel NF-kappa B-regulated immune genes which are consistent with available SAGE data. Our method of predicting gene function, based on characteristic patterns of transcription factor binding, evolutionary conservation, and expression studies, would be applicable to finding genes with other functions.

Animals↗

Comparison of whole genome assemblies of the human genome.

A fundamental problem in the human genome project is uncovering the correct assembly of the human genome. Many studies, including transcriptional analysis, SNP detection and characterization, gene finding and EST clustering, use genome assemblies as templates so it is important to determine the consistency among the various whole genome assemblies. A comparison of the order and orientation of the GenBank entries used to construct the NCBI and UCSC Goldenpath assemblies was made. In addition, a sequence level comparison was performed using MULTI, an efficient database search tool developed to make whole genome comparisons possible. The resulting comparisons show significant discrepancies in the sequence as well as in the order and orientation of GenBank entries used in constructing the NCBI and UCSC assemblies.

Chromosomes, Human↗

Consensus promoter identification in the human genome utilizing expressed gene markers and gene modeling.

Deciphering the human genome includes locating the promoters that initiate transcription and identifying the exons of genes. Many promoter prediction programs have been proposed, but when they are applied to extended regions of the genome, most of their predictions are false-positives. The extensive collection of gene transcript sequences is an important new source of information, which has not been used previously in promoter predictions. Our approach is to enhance the specificity of predictions by restricting the genomic regions that are searched using gene transcript alignments as anchors in the genome for gene modeling. We developed a consensus promoter prediction method combining previously developed algorithms with the GENSCAN gene modeling program. Our method, CONPRO (CONsensus PROmoter), identifies promoters with very high confidence, and the predicted promoters are guaranteed to be associated with genes. On our test data set, the method correctly detects promoters for approximately half of all human genes (37%-71%), and most predictions are true promoters (85%-90%). Applying our method to the human genome and human genes from the Unigene data set, we find the promoters for 13,744 genes. Of these, 6440 are genes with a functionally cloned mRNA, and 7304 are novel genes for which only expressed sequence tags (ESTs) are available. Candidate promoters for many novel genes will be a useful resource in elucidating complex biological response mechanisms.

5' Untranslated Regions↗