Search PubMed⌕ Search

Biomedical subjects

Kenneth H Buetow

Publications and source records attributed to Kenneth H Buetow.

At least 19 recordsLinked to original sources

Reproducible autosomal gene expression changes with loss of typical X and Y complement across tumor types.

Although there are known sex differences in cancer incidence, severity, and treatment, the sex chromosomes are typically excluded from genomic analyses because of the unique technical challenges associated with assessing their copy number, sequence variation, and expression. Here we assess sex chromosome complement in three widely-used human genomics datasets from normal (non-cancerous) tissues, primary tumors, and cancer cell lines and study the effects on genome-wide gene expression. Expected sex chromosome complements based on reported patient sex were observed in non-cancerous tissues, but about half of tumors and cancer cell lines showed loss of typical sex chromosome gene expression across tissue types with three categories: loss of chromosome Y (LOY), loss of chromosome X (LOX) and reactivation of the inactive X chromosome (XaXa). Genes consistently differentially expressed in tumors with loss of chromosome X, loss of chromosome Y, or loss of X chromosome inactivation are associated with the hallmarks of cancer and include both sex-linked and autosomal genes from nearly all chromosomes, druggable genes, and genes with molecular functions relevant to cancer signaling, such as kinase activity. Strikingly, tumors that are X0, including tumors from female patients that have lost an X chromosome and tumors from male patients that have lost a Y chromosome, cluster together by gene expression profile. Patients with tumors that have LOX or LOY had poorer survival outcomes compared to those with tumors that had maintained their sex chromosome complement. Further, LOX and LOY eliminates nearly all of the differential gene expression between tumors from different patient sexes, affecting sex chromosomal and autosomal gene expression. Going forward, considering patient sex as well as the entire genome, including assessment of the sex chromosome complement, will provide additional insights into personalized tumor etiology, progression, treatment, and patient outcome.

Journal Article↗

Genome-wide loss of heterozygosity and copy number alteration in esophageal squamous cell carcinoma using the Affymetrix GeneChip Mapping 10 K array.

BACKGROUND: Esophageal squamous cell carcinoma (ESCC) is a common malignancy worldwide. Comprehensive genomic characterization of ESCC will further our understanding of the carcinogenesis process in this disease. RESULTS: Genome-wide detection of chromosomal changes was performed using the Affymetrix GeneChip 10 K single nucleotide polymorphism (SNP) array, including loss of heterozygosity (LOH) and copy number alterations (CNA), for 26 pairs of matched germ-line and micro-dissected tumor DNA samples. LOH regions were identified by two methods--using Affymetrix's genotype call software and using Affymetrix's copy number alteration tool (CNAT) software--and both approaches yielded similar results. Non-random LOH regions were found on 10 chromosomal arms (in decreasing order of frequency: 17p, 9p, 9q, 13q, 17q, 4q, 4p, 3p, 15q, and 5q), including 20 novel LOH regions (10 kb to 4.26 Mb). Fifteen CNA-loss regions (200 kb to 4.3 Mb) and 36 CNA-gain regions (200 kb to 9.3 Mb) were also identified. CONCLUSION: These studies demonstrate that the Affymetrix 10 K SNP chip is a valid platform to integrate analyses of LOH and CNA. The comprehensive knowledge gained from this analysis will enable improved strategies to prevent, diagnose, and treat ESCC.

Aged↗

SNPdetector: a software tool for sensitive and accurate SNP detection.

Identification of single nucleotide polymorphisms (SNPs) and mutations is important for the discovery of genetic predisposition to complex diseases. PCR resequencing is the method of choice for de novo SNP discovery. However, manual curation of putative SNPs has been a major bottleneck in the application of this method to high-throughput screening. Therefore it is critical to develop a more sensitive and accurate computational method for automated SNP detection. We developed a software tool, SNPdetector, for automated identification of SNPs and mutations in fluorescence-based resequencing reads. SNPdetector was designed to model the process of human visual inspection and has a very low false positive and false negative rate. We demonstrate the superior performance of SNPdetector in SNP and mutation analysis by comparing its results with those derived by human inspection, PolyPhred (a popular SNP detection tool), and independent genotype assays in three large-scale investigations. The first study identified and validated inter- and intra-subspecies variations in 4,650 traces of 25 inbred mouse strains that belong to either the Mus musculus species or the M. spretus species. Unexpected heterozygosity in CAST/Ei strain was observed in two out of 1,167 mouse SNPs. The second study identified 11,241 candidate SNPs in five ENCODE regions of the human genome covering 2.5 Mb of genomic sequence. Approximately 50% of the candidate SNPs were selected for experimental genotyping; the validation rate exceeded 95%. The third study detected ENU-induced mutations (at 0.04% allele frequency) in 64,896 traces of 1,236 zebra fish. Our analysis of three large and diverse test datasets demonstrated that SNPdetector is an effective tool for genome-scale research and for large-sample clinical studies. SNPdetector runs on Unix/Linux platform and is available publicly (http://lpg.nci.nih.gov).

Algorithms↗

ATG deserts define a novel core promoter subclass.

The MHC class I gene, PD1, has neither functional TATAA nor Initiator (Inr) elements in its core promoter and initiates transcription at multiple, dispersed sites over an extended region in vitro. Here, we define a novel core promoter feature that supports regulated transcription through selective transcription start site (TSS) usage. We demonstrate that TSS selection is actively regulated and context dependent. Basal and activated transcriptions initiate from largely nonoverlapping TSS regions. Transcripts derived from multiple TSS encode a single protein, due to the absence of any ATG triplets within approximately 430 bp upstream of the major transcription start site. Thus, the PD1 core promoter is embedded within an "ATG desert". Remarkably, extending this analysis genome-wide, we find that ATG deserts define a novel promoter subclass. They occur nonrandomly, are significantly associated with non-TATAA promoters that use multiple TSS, independent of the presence of CpG islands (CGI). We speculate that ATG deserts may provide a core promoter platform upon which complex upstream regulatory signals can be integrated, targeting multiple TSS whose products encode a single protein.

Animals↗

Cyberinfrastructure: empowering a "third way" in biomedical research.

Biomedicine has experienced explosive growth, fueled in parts by the substantial increase of government support, continued development of the biotechnology industry, and the increasing adoption of molecular-based medicine. At its core, it is composed of fiercely independent, innovative, entrepreneurial individuals, organizations, and institutions. The field has developed unprecedented capacity to characterize biologic systems at their most fundamental levels with the use of tools and technologies almost unimaginable a generation ago. Biomedicine is at the precipice of unlocking the very essence of biologic life and enabling a new generation of medicine. Development and deployment of cyberinfrastructure may prove to be on the critical path to obtaining these goals.

Biomedical Research↗

Genome-wide association study in esophageal cancer using GeneChip mapping 10K array.

Whole genome association studies of complex human diseases represent a new paradigm in the postgenomic era. In this study, we report application of the Affymetrix, Inc. (Santa Clara, CA) high-density single nucleotide polymorphism (SNP) array containing 11,555 SNPs in a pilot case-control study of esophageal squamous cell carcinoma (ESCC) that included the analysis of germ line samples from 50 ESCC patients and 50 matched controls. The average genotyping call rate for the 100 samples analyzed was 96%. Using the generalized linear model (GLM) with adjustment for potential confounders and multiple comparisons, we identified 37 SNPs associated with disease, assuming a recessive mode of transmission; similarly, 48 SNPs were identified assuming a dominant mode and 53 SNPs in a continuous mode. When the 37 SNPs identified from the GLM recessive mode were used in a principal components analysis, the first principal component correctly predicted 46 of 50 cases and 47 of 50 controls. Among all the SNPs selected from GLMs for the three modes of transmission, 39 could be mapped to 1 of 33 genes. Many of these genes are involved in various cancers, including GASC1, shown previously to be amplified in ESCCs, and EPHB1 and PIK3C3. In conclusion, we have shown the feasibility of the Affymetrix 10K SNP array in genome-wide association studies of common cancers and identified new candidate loci to study in ESCC.

Carcinoma, Squamous Cell↗

Interlaboratory comparability study of cancer gene expression analysis using oligonucleotide microarrays.

A key step in bringing gene expression data into clinical practice is the conduct of large studies to confirm preliminary models. The performance of such confirmatory studies and the transition to clinical practice requires that microarray data from different laboratories are comparable and reproducible. We designed a study to assess the comparability of data from four laboratories that will conduct a larger microarray profiling confirmation project in lung adenocarcinomas. To test the feasibility of combining data across laboratories, frozen tumor tissues, cell line pellets, and purified RNA samples were analyzed at each of the four laboratories. Samples of each type and several subsamples from each tumor and each cell line were blinded before being distributed. The laboratories followed a common protocol for all steps of tissue processing, RNA extraction, and microarray analysis using Affymetrix Human Genome U133A arrays. High within-laboratory and between-laboratory correlations were observed on the purified RNA samples, the cell lines, and the frozen tumor tissues. Intraclass correlation within laboratories was only slightly stronger than between laboratories, and the intraclass correlation tended to be weakest for genes expressed at low levels and showing small variation. Finally, hierarchical cluster analysis revealed that the repeated samples clustered together regardless of the laboratory in which the experiments were done. The findings indicate that under properly controlled conditions it is feasible to perform complete tumor microarray analysis, from tissue processing to hybridization and scanning, at multiple independent laboratories for a single study.

Adenocarcinoma↗

Detecting false expression signals in high-density oligonucleotide arrays by an in silico approach.

High-density oligonucleotide arrays have become a popular assay for concurrent measurement of mRNA expression at the genome scale. Much effort has been devoted to the development of statistical analysis tools aimed at reducing experimental noise and normalizing experimental variation in gene expression analysis. However, these investigations do not detect or catalog systematic problems associated with specific oligonucleotide probes. Here, we present an investigation of problematic probes that yield consistent but inaccurate signals across multiple experiments. By evaluating data integrity among gene, probe sequence, and genomic structure we identified a total of 20,696 (10.5%) nonspecific probes that could cross-hybridize to multiple genes and a total of 18,363 (9.3%) probes that miss the target transcript sequences on the Affymetrix GeneChip U95A/Av2 array. The numbers of nonspecific and mistargeted probes on the U133A array are 29,405 (12.1%) and 19,717 (8.0%), respectively. The poor performance of the mistargeted probes was confirmed in two GeneChip experiments, in which these probes showed a 20-30% decrease in detecting present signals compared with normal probes. Comparison of qualitative expression signals obtained from SAGE and EST data with those from GeneChip arrays showed that the consistency of the two platforms is 30% lower in problematic probes than in normal probes. A Web application was developed to apply our results for improving the accuracy of expression analysis.

Expressed Sequence Tags↗

A high-resolution multistrain haplotype analysis of laboratory mouse genome reveals three distinctive genetic variation patterns.

Understanding of the structure and the origin of genetic variation patterns in the laboratory inbred mouse provides insight into the utility of the mouse model for studying human complex diseases and strategies for disease gene mapping. In order to address this issue, we have constructed a multistrain, high-resolution haplotype map for the 99-Mb mouse Chromosome 16 using approximately 70,000 single nucleotide polymorphism (SNP) markers derived from whole-genome shotgun sequencing of five laboratory inbred strains. We discovered that large polymorphic blocks (i.e., regions where only two haplotypes, thus one SNP conformation, are found in the five strains), large monomorphic blocks (i.e., regions where the five strains share the same haplotype), and fragmented blocks (i.e., regions of greater complexity not resembling at all the first two categories) span 50%, 18%, and 32% of the chromosome, respectively. The haplotype map has 98% accuracy in predicting mouse genotypes in two other studies. Its predictions are also confirmed by experimental results obtained from resequencing of 40-kb genomic sequences at 21 distinct genomic loci in 13 laboratory inbred strains and 12 wild-derived strains. We demonstrate that historic recombination, intra-subspecies variations and inter-subspecies variations have all contributed to the formation of the three distinctive genetic signatures. The results suggest that the controlled complexity of the laboratory inbred strains may provide a means for uncovering the biological factors that have shaped genetic variation patterns.

Alleles↗

Hemochromatosis gene mutations and distal adenomatous colorectal polyps.

Iron has been suggested to be a risk factor for colorectal neoplasia. Some individuals who are heterozygous for mutations in the hemochromatosis gene (HFE) have higher than average serologic measures of iron. We therefore investigated whether heterozygosity for HFE mutations was related to risk of advanced distal adenoma and whether the relationship was affected by dietary iron intake. In the Prostate, Lung, Colorectal and Ovarian Cancer Screening Trial, 679 persons with advanced distal adenoma and 697 control persons were genotyped for the two major HFE mutations (C282Y and H63D), one HFE polymorphism (IVS2+4), and one polymorphism (G142S) in the transferrin receptor gene (TFRC). HFE haplotypes were also created to examine the effect of haplotype on risk. Food frequency questionnaire data were used to estimate daily iron intake. There was no relationship between any HFE genotype or haplotype and advanced adenoma. Stratification of HFE genotype by TFRC genotype did not change the results. In addition, there was no relationship between dietary iron intake and risk of adenoma or between HFE genotype and risk of adenoma, stratified by iron intake. These results do not support a relationship between HFE heterozygosity and risk of advanced distal adenoma.

Adenomatous Polyps↗

Large-scale analysis of non-synonymous coding region single nucleotide polymorphisms.

MOTIVATION: Single nucleotide polymorphisms (SNPs) are the most common form of genetic variant in humans. SNPs causing amino acid substitutions are of particular interest as candidates for loci affecting susceptibility to complex diseases, such as diabetes and hypertension. To efficiently screen SNPs for disease association, it is important to distinguish neutral variants from deleterious ones. RESULTS: We describe the use of Pfam protein motif models and the HMMER program to predict whether amino acid changes in conserved domains are likely to affect protein function. We find that the magnitude of the change in the HMMER E-value caused by an amino acid substitution is a good predictor of whether it is deleterious. We provide internet-accessible display tools for a genomewide collection of SNPs, including 7391 distinct non-synonymous coding region SNPs in 2683 genes. AVAILABILITY: http://lpgws.nci.nih.gov/cgi-bin/GeneViewer.cgi

Amino Acid Motifs↗

Direct ampholyte-free liquid-phase isoelectric peptide focusing: application to the human serum proteome.

In this study, we utilized a multidimensional peptide separation strategy combined with tandem mass spectrometry (MS/MS) for the identification of proteins in human serum. After enzymatically digesting serum with trypsin, the peptides were fractionated using liquid-phase isoelectric focusing (IEF) in a novel ampholyte-free format. Twenty IEF fractions were collected and analyzed by reversed-phase microcapillary liquid chromatography (microLC)-MS/MS. Bioinformatic analysis of the raw MS/MS spectra resulted in the identification of 844 unique peptides, corresponding to 437 proteins. This study demonstrates the efficacy of ampholyte-free peptide autofocusing, which alleviates peptide losses in ampholyte removal strategies. The results show that the separation strategy is effective for high-throughput characterization of proteins from complex proteomic mixtures.

Blood Proteins↗

Comparative sequence analysis of imprinted genes between human and mouse to reveal imprinting signatures.

We performed a comparative genomic sequence analysis between human and mouse for 24 imprinted genes on human chromosomes 1, 6, 7, 11, 13, 14, 15, 18, 19, and 20. The MEME program was used to search for motifs within conserved sequences among the imprinted genes and we then used the MAST program to analyze for the presence or absence of motifs in the imprinted genes and 128 nonimprinted genes. Our analysis identified 15 motifs that were significantly enriched in the imprinted genes. We generated a logistic regression model by combining multiple motifs as input variables and the 24 imprinted genes and the 128 nonimprinted genes as a training set. The accuracy, sensitivity, and specificity of our model were 98, 92, and 99%, respectively. The model was further validated by an open test on 12 additional imprinted genes. The motifs identified in this study are novel imprinting signatures, which should improve our understanding of genomic imprinting and the role of genomic imprinting in human diseases.

Amino Acid Motifs↗

A computational approach to measuring coherence of gene expression in pathways.

This study uses a computational approach to analyze coherence of expression of genes in pathways. Microarray data were analyzed with respect to coherent gene expression in a group of genes defined as a pathway in the Kyoto Encyclopedia of Genes and Genomes (KEGG) database. Our hypothesis is that genes in the same pathway are more likely to be coordinately regulated than a randomly selected gene set. A correlation coefficient for each pair of genes in a pathway was estimated based on gene expression in normal or tumor samples, and statistically significant correlation coefficients were identified. The coherence indicator was defined as the ratio of the number of gene pairs in the pathway whose correlation coefficients are significant, divided by the total number of gene pairs in the pathway. We defined all genes that appeared in the KEGG pathways as a reference gene set. Our analysis indicated that the mean coherence indicator of pathways is significantly larger than the mean coherence indicator of random gene sets drawn from the reference gene set. Thus, the result supports our hypothesis. The significance of each individual pathway of n genes was evaluated by comparing its coherence indicator with coherence indicators of 1000 random permutation sets of n genes chosen from the reference gene set. We analyzed three data sets: two Affymetrix microarrays and one cDNA microarray. For each of the three data sets, statistically significant pathways were identified among all KEGG pathways. Seven of 96 pathways had a significant coherence indicator in normal tissue and 14 of 96 pathways had a significant coherence indicator in tumor tissue in all three data sets. The increase in the number of pathways with significant coherence indicators may reflect the fact that tumor cells have a higher rate of metabolism than normal cells. Five pathways involved in oxidative phosphorylation, ATP synthesis, protein synthesis, or RNA synthesis were coherent in both normal and tumor tissue, demonstrating that these are essential genes, a high level of expression of which is required regardless of cell type.

Databases, Genetic↗

Bioinformatics tools for single nucleotide polymorphism discovery and analysis.

Single nucleotide polymorphisms (SNPs) are a valuable resource for investigating the genetic basis of disease. These variants can serve as markers for fine-scale genetic mapping experiments and genome-wide association studies. Certain of these nucleotide polymorphisms may predispose individuals to illnesses such as diabetes, hypertension, or cancer, or affect disease progression. Bioinformatics techniques can play an important role in SNP discovery and analysis. We use computational methods to identify SNPs and to predict whether they are likely to be neutral or deleterious. We also use informatics to annotate genes that contain SNPs. To make this information available to the research community, we provide a variety of Internet-accessible tools for data access and display. These tools allow researchers to retrieve data about SNPs based on gene of interest, genetic or physical map location, or expression pattern.

Chromosome Mapping↗

caCORE: a common infrastructure for cancer informatics.

MOTIVATION: Sites with substantive bioinformatics operations are challenged to build data processing and delivery infrastructure that provides reliable access and enables data integration. Locally generated data must be processed and stored such that relationships to external data sources can be presented. Consistency and comparability across data sets requires annotation with controlled vocabularies and, further, metadata standards for data representation. Programmatic access to the processed data should be supported to ensure the maximum possible value is extracted. Confronted with these challenges at the National Cancer Institute Center for Bioinformatics, we decided to develop a robust infrastructure for data management and integration that supports advanced biomedical applications. RESULTS: We have developed an interconnected set of software and services called caCORE. Enterprise Vocabulary Services (EVS) provide controlled vocabulary, dictionary and thesaurus services. The Cancer Data Standards Repository (caDSR) provides a metadata registry for common data elements. Cancer Bioinformatics Infrastructure Objects (caBIO) implements an object-oriented model of the biomedical domain and provides Java, Simple Object Access Protocol and HTTP-XML application programming interfaces. caCORE has been used to develop scientific applications that bring together data from distinct genomic and clinical science sources. AVAILABILITY: caCORE downloads and web interfaces can be accessed from links on the caCORE web site (http://ncicb.nci.nih.gov/core). caBIO software is distributed under an open source license that permits unrestricted academic and commercial use. Vocabulary and metadata content in the EVS and caDSR, respectively, is similarly unrestricted, and is available through web applications and FTP downloads. SUPPLEMENTARY INFORMATION: http://ncicb.nci.nih.gov/core/publications contains links to the caBIO 1.0 class diagram and the caCORE 1.0 Technical Guide, which provide detailed information on the present caCORE architecture, data sources and APIs. Updated information appears on a regular basis on the caCORE web site (http://ncicb.nci.nih.gov/core).

Animals↗

Genomewide distribution of high-frequency, completely mismatching SNP haplotype pairs observed to be common across human populations.

Knowledge of human haplotype structure has important implications for strategies of disease-gene mapping and for understanding human evolutionary history. Many attributes of SNPs and haplotypes appear to exhibit highly nonrandom behavior, suggesting past operation of selection or other nonneutral forces. We report the exceptional abundance of a particular haplotype pattern in which two high-frequency haplotypes have different alleles at every SNP site (hence the name "yin yang haplotypes"). Analysis of common haplotypes in 62 random genomic loci and 85 gene coding regions in humans shows that the proportion of the genome spanned by yin yang haplotypes is 75%-85%. Population data of 28 genomic loci in Drosophila melanogaster reveal a similar pattern. The high recurrence (>/=85%) of these haplotype patterns in four distinct human populations suggests that the yin yang haplotypes are likely to predate the African diaspora. The pattern initially appeared to suggest deep population splitting or maintenance of ancient lineages by selection; however, coalescent simulation reveals that the yin yang phenomenon can be explained by strictly neutral evolution in a well-mixed population.

Algorithms↗