Search PubMed⌕ Search

Biomedical subjects

David J Cutler

Publications and source records attributed to David J Cutler.

At least 19 recordsLinked to original sources

A powerful framework for differential co-expression analysis of general risk factors.

MOTIVATION: Differential co-expression analysis (DCA) aims to identify genes in a pathway whose shared expression depends on a risk factor. While DCA provides insights into the biological activity of diseases, existing methods are limited to categorical risk factors and/or suffer from bias due to batch and variance-specific effects. We propose a new framework, Kernel-based DCA (KDCA), that harnesses correlation patterns between genes in a pathway to detect differential co-expression arising from general (i.e. continuous, discrete, or categorical) risk factors. RESULTS: Using various simulated pathway architectures, we find that KDCA accounts for common sources of bias to control the type I error rate while substantially increasing the power compared to the standard eigengene approach. We then applied KDCA to The Cancer Genome Atlas thyroid data set and found several differentially co-expressed pathways by age of diagnosis and BRAF mutation status that were undetected by the eigengene method. Collectively, our results demonstrate that KDCA is a powerful testing framework that expands DCA applications in expression studies. AVAILABILITY AND IMPLEMENTATION: KDCA is publicly available in the R package kdca. The package can be downloaded at https://github.com/ajbass/kdca.

Humans↗

Relative contribution of genetic and nongenetic modifiers to intestinal obstruction in cystic fibrosis.

BACKGROUND & AIMS: Neonatal intestinal obstruction (meconium ileus [MI]) occurs in 15% of patients with cystic fibrosis (CF). Our aim was to determine the relative contribution of genetic and nongenetic modifiers to the development of this major complication of CF. METHODS: A total of 65 monozygous twin pairs, 23 dizygous twin/triplet sets, and 349 sets of siblings with CF were analyzed for MI status, significant covariates, and genome-wide linkage. RESULTS: Specific mutations in the CF transmembrane conductance regulator (CFTR), the gene responsible for CF, correlated with MI, indicating a role for CFTR genotype. Monozygous twins showed substantially greater concordance for MI than dizygous twins and siblings (P = 1 x 10(-5)), showing that modifier genes independent of CFTR contribute substantially to this trait. Regression analysis revealed that MI was correlated with distal intestinal obstruction syndrome (P = 8 x 10(-4)). Unlike MI, concordance analysis indicated that the risk for development of distal intestinal obstruction syndrome in CF patients is caused primarily by nongenetic factors. Regions of suggestive linkage (logarithm of the odds of linkage >2.0) for modifier genes that cause MI (chromosomes 4q35.1, 8p23.1, and 11q25) or protect from MI (chromosomes 20p11.22 and 21q22.3) were identified by genome-wide analyses. These analyses did not support the existence of a major modifier gene on chromosome 19 in a region previously linked to MI. CONCLUSIONS: The CFTR gene along with 2 or more modifier genes are the major determinants of intestinal obstruction in newborn CF patients, whereas intestinal obstruction in older CF patients is caused primarily by nongenetic factors.

Chromosomes, Human, Pair 11↗

An oligonucleotide microarray for high-throughput sequencing of the mitochondrial genome.

Previously we developed an oligonucleotide sequencing microarray (MitoChip) as an array-based sequencing platform for rapid and high-throughput analysis of mitochondrial DNA. The first generation MitoChip, however, was not tiled with probes for the noncoding D-loop region, a site frequently mutated in human cancers. Here we report the development of a second-generation MitoChip (v2.0) with oligonucleotide probes to sequence the entire mitochondrial genome. In addition, the MitoChip v2.0 contains redundant tiling of sequences for 500 of the most common haplotypes including single-nucleotide changes, insertions, and deletions. Sequencing results from 14 primary head and neck tumor tissues demonstrated that the v2.0 MitoChips detected a larger number of variants than the original version. Multiple coding region variants detected only in the second generation MitoChips, but not the earlier chip version, were further confirmed with conventional sequencing. Moreover, 31 variations in noncoding region were identified using MitoChips v2.0. Replicate experiments demonstrated >99.99% reproducibility in the second generation MitoChip. In seven head and neck cancer samples with matched lymphocyte DNA, the MitoChip v2.0 detected at least one cancer-associated mitochondrial mutation in four (57%) samples. These results indicate that the second generation MitoChip is a high-throughput platform for identification of mitochondrial DNA mutations in primary tumors.

Base Sequence↗

Genomic alterations in cultured human embryonic stem cells.

Cultured human embryonic stem cell (hESC) lines are an invaluable resource because they provide a uniform and stable genetic system for functional analyses and therapeutic applications. Nevertheless, these dividing cells, like other cells, probably undergo spontaneous mutation at a rate of 10(-9) per nucleotide. Because each mutant has only a few progeny, the overall biological properties of the cell culture are not altered unless a mutation provides a survival or growth advantage. Clonal evolution that leads to emergence of a dominant mutant genotype may potentially affect cellular phenotype as well. We assessed the genomic fidelity of paired early- and late-passage hESC lines in the course of tissue culture. Relative to early-passage lines, eight of nine late-passage hESC lines had one or more genomic alterations commonly observed in human cancers, including aberrations in copy number (45%), mitochondrial DNA sequence (22%) and gene promoter methylation (90%), although the latter was essentially restricted to 2 of 14 promoters examined. The observation that hESC lines maintained in vitro develop genetic and epigenetic alterations implies that periodic monitoring of these lines will be required before they are used in in vivo applications and that some late-passage hESC lines may be unusable for therapeutic purposes.

Cell Culture Techniques↗

On the probability that a novel variant is a disease-causing mutation.

When a novel variant is found in a patient and not in a group of controls, it becomes a candidate for the disease-causing mutation in that patient. At present, no sampling theory exists for assessing the probability that the novel SNP might actually be a neutral variant. We have developed a population genetics-based method for calculating a P-value for a mutation-detection effort. Our method can be applied to a heterozygous patient, a homozygous patient, with or without inbreeding, or to a patient who is a compound heterozygote. Additionally, the method can be used to calculate the probability of finding a neutral variant at frequencies that differ between a group of patients and a group of controls, given some length of sequence examined. This method accounts for the multiple testing that is inherent in identification of variants through sequencing, to be used in subsequent case-control analyses. We show, for example, that for complete resequencing of 10 kb, the probability of finding a neutral variant in a patient and not in 50 controls is about 15%. Thus, discovery of a variant in a patient and not in a group of controls is, on its own, very weak evidence of involvement with disease.

Animals↗

A common sex-dependent mutation in a RET enhancer underlies Hirschsprung disease risk.

The identification of common variants that contribute to the genesis of human inherited disorders remains a significant challenge. Hirschsprung disease (HSCR) is a multifactorial, non-mendelian disorder in which rare high-penetrance coding sequence mutations in the receptor tyrosine kinase RET contribute to risk in combination with mutations at other genes. We have used family-based association studies to identify a disease interval, and integrated this with comparative and functional genomic analysis to prioritize conserved and functional elements within which mutations can be sought. We now show that a common non-coding RET variant within a conserved enhancer-like sequence in intron 1 is significantly associated with HSCR susceptibility and makes a 20-fold greater contribution to risk than rare alleles do. This mutation reduces in vitro enhancer activity markedly, has low penetrance, has different genetic effects in males and females, and explains several features of the complex inheritance pattern of HSCR. Thus, common low-penetrance variants, identified by association studies, can underlie both common and rare diseases.

Animals↗

A note on exact tests of Hardy-Weinberg equilibrium.

Deviations from Hardy-Weinberg equilibrium (HWE) can indicate inbreeding, population stratification, and even problems in genotyping. In samples of affected individuals, these deviations can also provide evidence for association. Tests of HWE are commonly performed using a simple chi2 goodness-of-fit test. We show that this chi2 test can have inflated type I error rates, even in relatively large samples (e.g., samples of 1,000 individuals that include approximately 100 copies of the minor allele). On the basis of previous work, we describe exact tests of HWE together with efficient computational methods for their implementation. Our methods adequately control type I error in large and small samples and are computationally efficient. They have been implemented in freely available code that will be useful for quality assessment of genotype data and for the detection of genetic association or population stratification in very large data sets.

Genetics, Population↗

Microarray-based resequencing of multiple Bacillus anthracis isolates.

We used custom-designed resequencing arrays to generate 3.1 Mb of genomic sequence from a panel of 56 Bacillus anthracis strains. Sequence quality was shown to be very high by replication (discrepancy rate of 7.4 x 10(-7)) and by comparison to independently generated shotgun sequence (discrepancy rate < 2.5 x 10(-6)). Population genomics studies of microbial pathogens using rapid resequencing technologies such as resequencing arrays are critical for recognizing newly emerging or genetically engineered strains.

Bacillus anthracis↗

Exhaustive allelic transmission disequilibrium tests as a new approach to genome-wide association studies.

Genome-wide disease-association mapping has been heralded as the study design of the next generation, but the lack of analytical methods to use genotype data fully is a large stumbling block. Here we describe an algorithm and statistical method that efficiently and exhaustively exploits haplotype information by subjecting alleles (a marker or contiguous sets of markers) from sliding windows of all sizes to transmission disequilibrium tests. By applying our method to simulated data and to Hirschsprung disease, we show that it can detect both common and rare disease variants of small effect. These results show that the theoretical benefits of genome-wide association studies are at last realizable.

Algorithms↗

Haplotype and missing data inference in nuclear families.

Determining linkage phase from population samples with statistical methods is accurate only within regions of high linkage disequilibrium (LD). Yet, affected individuals in a genetic mapping study, including those involving cases and controls, may share sequences identical-by-descent stretching on the order of 10s to 100s of kilobases, quite possibly over regions of low LD in the population. At the same time, inferring phase from nuclear families may be hampered by missing family members, missing genotypes, and the noninformativity of certain genotype patterns. In this study, we reformulate our previous haplotype reconstruction algorithm, and its associated computer program, to phase parents with information derived from population samples as well as from their offspring. In applications of our algorithm to 100-kb stretches, simulated in accordance to a Wright-Fisher model with typical levels of LD in humans, we find that phase reconstruction for 160 trios with 10% missing data is highly accurate (>90%) over the entire length. Furthermore, our algorithm can estimate allelic status for missing data at high accuracy (>95%). Finally, the input capacity of the program is vast, easily handling thousands of segregating sites in > or = 1000 chromosomes.

Algorithms↗

Aberrant gating of photic input to the suprachiasmatic circadian pacemaker of mice lacking the VPAC2 receptor.

VIP acting via the VPAC(2) receptor is implicated as a key signaling pathway in the maintenance and resetting of the hypothalamic suprachiasmatic nuclei (SCN) circadian pacemaker; circadian rhythms in SCN clock gene expression and wheel-running behavior are abolished in mice lacking the VPAC(2) receptor (Vipr2(-/-)). Here, using immunohistochemical detection of pERK (phosphorylated extracellular signal-regulated kinases 1/2) and c-FOS, we tested whether the gating of photic input to the SCN is maintained in these apparently arrhythmic Vipr2(-/-) mice. Under light/dark and constant darkness, spontaneous expression of pERK and c-FOS in the wild-type mouse SCN was significantly elevated during subjective day compared with subjective night; no diurnal or circadian variation in pERK or c-FOS was detected in the SCN of Vipr2(-/-) mice. In constant darkness, light pulses given during the subjective night but not the subjective day significantly increased expression of pERK and c-FOS in the wild-type SCN. In contrast, light pulses given during both subjective day and subjective night robustly increased expression of pERK and c-FOS in the Vipr2(-/-) mouse SCN. Although photic stimuli activate intracellular pathways within the SCN of Vipr2(-/-) mice, they do not engage the core clock mechanisms. The absence of photic gating, together with the general lack of overt rhythms in circadian output, strongly suggests that the SCN circadian pacemaker is completely dysfunctional in the Vipr2(-/-) mouse.

Animals↗

Pharmacokinetic parameter prediction from drug structure using artificial neural networks.

Simple methods for determining the human pharmacokinetics of known and unknown drug-like compounds is a much sought-after goal in the pharmaceutical industry. The current study made use of artificial neural networks (ANNs) for the prediction of clearances, fraction bound to plasma proteins, and volume of distribution of a series of structurally diverse compounds. A number of theoretical descriptors were generated from the drug structures and both automated and manual pruning were used to derive optimal subsets of descriptors for quantitative structure-pharmacokinetic relationship models. Models were trained on one set of compounds and validated with another. Absolute predicted ability was evaluated using a further independent test set of compounds. Correlations for test compounds ranged from 0.855 to 0.992. Predicted values agreed closely with experimental values for total clearance, renal clearance, and volume of distribution, while predictions for protein binding were encouraging. The combination of descriptor generation, ANNs, and the speed and success of this technique compared with conventional methods shows strong potential for use in pharmaceutical product development.

Drug Design↗

Discrepancies in dbSNP confirmation rates and allele frequency distributions from varying genotyping error rates and patterns.

SUMMARY: Three recent publications have examined the quality and completeness of public database single nucleotide polymorphism (dbSNP) and have come to dramatically different conclusions regarding dbSNPs false positive rate and the proportion of dbSNPs that are expected to be common. These studies employed different genotyping technologies and different protocols in determining minimum acceptable genotyping quality thresholds. Because heterozygous sites typically have lower quality scores than homozygous sites, a higher minimum quality threshold reduces the number of false positive SNPs, but yields fewer heterozygotes and leads to fewer confirmed SNPs. To account for the different confirmation rates and distributions of minor allele frequencies, we propose that the three confirmation studies have different false positive and false negative rates. We developed a mathematical model to predict SNP confirmation rates and the apparent distribution of minor allele frequencies under user-specified false positive and false negative rates. We applied this model to the three published studies and to our own resequencing effort. We conclude that the dbSNP false positive rate is approximately 15-17% and that the reported confirmation studies have vastly different genotyping error rates and patterns.

Algorithms↗

Tracking the evolution of the SARS coronavirus using high-throughput, high-density resequencing arrays.

Mutations in the SARS-Coronavirus (SARS-CoV) can alter its clinical presentation, and the study of its mutation patterns in human populations can facilitate contact tracing. Here, we describe the development and validation of an oligonucleotide resequencing array for interrogating the entire 30-kb SARS-CoV genome in a rapid, cost-effective fashion. Using this platform, we sequenced SARS-CoV genomes from Vero cell culture isolates of 12 patients and directly from four patient tissues. The sequence obtained from the array is highly reproducible, accurate (>99.99% accuracy) and capable of identifying known and novel variants of SARS-CoV. Notably, we applied this technology to a field specimen of probable SARS and rapidly deduced its infectious source. We demonstrate that array-based resequencing-by-hybridization is a fast, reliable, and economical alternative to capillary sequencing for obtaining SARS-CoV genomic sequence on a population scale, making this an ideal platform for the global monitoring of SARS-CoV and other small-genome pathogens.

Animals↗

Heterozygous mutations in BBS1, BBS2 and BBS6 have a potential epistatic effect on Bardet-Biedl patients with two mutations at a second BBS locus.

Bardet-Biedl syndrome (BBS) is a pleiotropic genetic disorder with substantial inter- and intrafamilial variability, that also exhibits remarkable genetic heterogeneity, with seven mapped BBS loci in the human genome. Recent data have demonstrated that BBS may be inherited either as a simple Mendelian recessive or as an oligogenic trait, since mutations at two loci are sometimes required for pathogenesis. This observation suggests that genetic interactions between the different BBS loci may modulate the phenotype, thus contributing to the clinical variability of BBS. We present three families with two mutations in either BBS1 or BBS2, in which some but not all patients carry a third mutation in BBS1, BBS2 or the putative chaperonin BBS6. In each example, the presence of three mutant alleles correlates with a more severe phenotype. For one of the missense alleles, we also demonstrate that the introduction of the mutation in mammalian cells causes a dramatic mislocalization of the protein compared with the wild-type. These data suggest that triallelic mutations are not always necessary for disease manifestation, but might potentiate a phenotype that is caused by two recessive mutations at an independent locus, thus introducing an additional layer of complexity on the genetic modeling of oligogenicity.

Bardet-Biedl Syndrome↗

Undetected genotyping errors cause apparent overtransmission of common alleles in the transmission/disequilibrium test.

The transmission/disequilibrium test (TDT), a family-based test of linkage and association, is a popular and intuitive statistical test for studies of complex inheritance, as it is nonparametric and robust to population stratification. We carried out a literature search and located 79 significant TDT-derived associations between a microsatellite marker allele and a disease. Among these, there were 31 (39%) in which the most common allele was found to exhibit distorted transmission to affected offspring, implying that the allele may be associated with either susceptibility to or protection from a disease. In 27 of these 31 studies (87%), the most common allele appeared to be overtransmitted to affected offspring (a risk factor), and, in the remaining 4 studies, the most common allele appeared to be undertransmitted (a protective factor). In a second literature search, we identified 92 case-control studies in which a microsatellite marker allele was found to have significantly different frequencies in case and control groups. Of these, there were 37 instances (40%) in which the most common allele was involved. In 12 of these 37 studies (32%), the most common allele was enriched in cases relative to controls (a risk factor), and, in the remaining 25 studies, the most common allele was enriched in controls (a protective factor). Thus, the most common allele appears to be a risk factor when identified through the TDT, and it appears to be protective when identified through case-control analysis. To understand this phenomenon, we incorporated an error model into the calculation of the TDT statistic. We show that undetected genotyping error can cause apparent transmission distortion at markers with alleles of unequal frequency. We demonstrate that this distortion is in the direction of overtransmission for common alleles. Therefore, we conclude that undetected genotyping errors may be contributing to an inflated false-positive rate among reported TDT-derived associations and that genotyping fidelity must be increased.

Alleles↗

Mometasone furoate degradation and metabolism in human biological fluids and tissues.

The in vitro metabolic and non-metabolic degradation kinetics of mometasone furoate (MF) was investigated in selected human biological fluids and subcellular fractions of tissues. Qualitative and quantitative differences in transformation profiles of MF were observed among human biological media. Degradation was the major event in plasma and urine with four new degradation products identified; A: 21-chloro-17alpha-hydroxy-16alpha-methyl-9beta,11beta-oxidopregna-1,4-diene-3,20-dione 17-(2-furoate), B: 9alpha,21beta-dichloro-11beta,21alpha-dihydroxy-16alpha-methylpregna-1,4,17,20-tetraen-3-one 21-(2-furoate), C: 21beta-chloro-21alpha-hydroxy-16alpha-methyl-9beta,11beta-oxidopregna-1,4,17,20-tetraen-3-one 21-(2-furoate), and D: 21-chloro-17alpha-hydroxy-16alpha-methyl-9beta,11beta-oxidopregna-1,4-diene-3,20-dione. A, B and C were predominant and D was minor in plasma while A and C were predominant in urine. Hydrolysis of the 17-ester bond of MF was not a major event in plasma. The turnover of MF in plasma was faster than that in phosphate buffers of pH 7.4. Metabolism of MF occurred primarily and rapidly in liver, appreciably in intestine, but negligibly in in vitro lung tissue. While 6beta-hydroxylation was a major metabolic pathway for MF in microsomes of both human liver and intestine, other parallel and subsequent metabolism pathways could also be involved. If these degradation and metabolic products are also formed and active in humans in vivo, both MF and its 'active' products need to be taken into account when determining the systemic bioavailability of MF and in establishing concentration-effect relationships with this drug.

Adult↗

Selective descriptor pruning for QSAR/QSPR studies using artificial neural networks.

Selection of optimal descriptors in quantitative structure-activity-property relationship (QSAR/QSPR) studies has been a perennial problem. Artificial Neural Networks (ANNs) have been used widely in QSAR/QSPR studies but less widely in descriptor selection. The current study used ANNs to select an optimal set of descriptors using large numbers of input variables. The effects of clean, noisy, and random input descriptors with linear, nonlinear, and periodic data on synthetic and real data QSAR/QSPR sets were examined. The optimal set of descriptors could be determined using a signal-to-noise ratio method. The optimal values for the rho parameter, which relates sample size to network architecture, were found to vary with the type of data. ANNs were able to detect meaningful descriptors in the presence of large numbers of random false descriptors.

Journal Article↗