Search PubMed⌕ Search

Biomedical subjects

Peter K Rogan

Publications and source records attributed to Peter K Rogan.

17 recordsLinked to original sources

BIPAD: a web server for modeling bipartite sequence elements.

BACKGROUND: Many dimeric protein complexes bind cooperatively to families of bipartite nucleic acid sequence elements, which consist of pairs of conserved half-site sequences separated by intervening distances that vary among individual sites. RESULTS: We introduce the Bipad Server, a web interface to predict sequence elements embedded within unaligned sequences. Either a bipartite model, consisting of a pair of one-block position weight matrices (PWM's) with a gap distribution, or a single PWM matrix for contiguous single block motifs may be produced. The Bipad program performs multiple local alignment by entropy minimization and cyclic refinement using a stochastic greedy search strategy. The best models are refined by maximizing incremental information contents among a set of potential models with varying half site and gap lengths. CONCLUSION: The web service generates information positional weight matrices, identifies binding site motifs, graphically represents the set of discovered elements as a sequence logo, and depicts the gap distribution as a histogram. Server performance was evaluated by generating a collection of bipartite models for distinct DNA binding proteins.

Algorithms↗

Dendrimer FISH detection of single-copy intervals in acute promyelocytic leukemia.

Acute promyelocytic leukemia (AML-M3) is characterized by a translocation between chromosomes 15 and 17 [t(15;17)]. The detection of t(15;17) at the single cell level, is commonly done by fluorescence in situ hybridization (FISH) using recombinant locus specific genomic probes greater than 14 kilobases kb in length. To allow a more thorough study of t(15;17), we designed small (0.9-3.6 kb), target-specific, single-copy probes from the human genome sequence. A novel detection approach was evaluated using moieties possessing more fluorophores, DNA dendrimers (up to 375 fluorophores per dendrimer). Two detection approaches were evaluated using the dendrimers: (1) dendrimers modified with anti-biotin antibodies for detection of biotinylated bound probes, and (2) dendrimers modified with 45-base long oligonucleotides designed from the single-copy probes, for direct detection of the target region. The selectivity of the probes was confirmed via indirect labeling with biotin/digoxigenin by nick translation, with detection efficiencies between 50 and 90%. Furthermore, the scFISH probes were successfully detected on metaphase cells with anti-biotin dendrimer conjugates and on interphase cells with 45-base modified dendrimers. Our results bring up the possibility to detect target regions of less than 1 kb, which will be a great contribution to high-resolution analysis of genomic sequences.

Cell Line, Tumor↗

Splice-site contribution in alternative splicing of PLP1 and DM20: molecular studies in oligodendrocytes.

Mutations in the proteolipid protein 1 (PLP1) gene cause the X-linked dysmyelinating diseases Pelizaeus-Merzbacher disease (PMD) and spastic paraplegia 2 (SPG2). We examined the severity of the following mutations that were suspected of affecting levels of PLP1 and DM20 RNA, the alternatively spliced products of PLP1: c.453G>A, c.453G>T, c.453G>C, c.453+2T>C, c.453+4A>G, c.347C>A, and c.453+28_+46del (the old nomenclature did not include the methionine codon: G450A, G450T, G450C, IVS3+2T>C, IVS3+4A>G, C344A, and IVS3+28-+46del). These mutations were evaluated by information theory-based analysis and compared with mRNA expression of the alternatively spliced products. The results are discussed relative to the clinical severity of disease. We conclude that the observed PLP1 and DM20 splicing patterns correlated well with predictions of information theory-based analysis, and that the relative strength of the PLP1 and DM20 donor splice sites plays an important role in PLP1 alternative splicing.

Alternative Splicing↗

Determination of genomic copy number with quantitative microsphere hybridization.

We developed a novel quantitative microsphere suspension hybridization (QMH) assay for determination of genomic copy number by flow cytometry. Single copy (sc) products ranging in length from 62 to 2,304 nucleotides [Rogan et al., 2001; Knoll and Rogan, 2004] from ABL1 (chromosome 9q34), TEKT3 (17p12), PMP22 (17p12), and HOXB1 (17q21) were conjugated to spectrally distinct polystyrene microspheres. These conjugated probes were used in multiplex hybridization to detect homologous target sequences in biotinylated genomic DNA extracted from fixed cell pellets obtained for cytogenetic studies. Hybridized targets were bound to phycoerythrin-labeled streptavidin; then the spectral emissions of both target and conjugated microsphere were codetected by flow cytometry. Prior amplification of locus-specific target DNA was not required because sc probes provide adequate specificity and sensitivity for accurate copy number determination. Copy number differences were distinguishable by comparing the mean fluorescence intensities (MFI) of test probes with a biallelic reference probe in genomic DNA of patient samples and abnormal cell lines. Concerted 5' ABL1 deletions in patient samples with a chromosome 9;22 translocation and chronic myelogenous leukemia were confirmed by comparison of the mean fluorescence intensities of ABL1 test probes with a HOXB1 reference probe. The relative intensities of the ABL1 probes were reduced to 0.59+/-0.02 fold in three different deletion patients and increased 1.42+/-0.01 fold in three trisomic 9 cell lines. TEKT3 and PMP22 probes detected proportionate copy number increases in five patients with Charcot-Marie-Tooth Type 1a disease and chromosome 17p12 duplications. Thus, the assay is capable of distinguishing one allele and three alleles from a biallelic reference sequence, regardless of chromosomal context.

Cell Line↗

Distortion of quantitative genomic and expression hybridization by Cot-1 DNA: mitigation of this effect.

Cross-hybridization of repetitive sequences in genomic and expression arrays is reported to be suppressed with repeat-blocking nucleic acids (C(o)t-1 DNA). Contrary to expectation, we demonstrated that C(o)t-1 also enhanced non-specific hybridization between probes and genomic targets. When added to target DNA, C(o)t-1 enhanced hybridization (2.2- to 3-fold) to genomic probes containing conserved repetitive elements. In addition to repetitive sequences, C(o)t-1 was found to be enriched for linked single copy (sc) sequences. Adventitious association between these sequences and probes distort quantitative measurements of the probes hybridized to desired genomic targets. Quantitative microarray hybridization studies using C(o)t-1 DNA are also susceptible to these effects, especially for probes that map to genomic regions containing conserved repetitive sequences. Hybridization measurements with such probes are less reproducible in the presence of C(o)t-1 than for probes derived from sc regions or regions containing divergent repeat elements, a finding with significant ramifications for genomic and expression microarray studies. We mitigated the requirement for C(o)t-1 either by hybridizing with computationally defined sc probes lacking repeats or by substituting synthetic repetitive elements complementary to sequences in genomic probes.

DNA↗

A variant form of Oguchi disease mapped to 13q34 associated with partial deletion of GRK1 gene.

PURPOSE: The purpose of this paper is to map the locus for a variant form of Oguchi disease in a Pakistani family and to identify the causative mutation. METHODS: Family 61029 was ascertained in the Punjab province of Pakistan. It includes three 13- to 19-year-old patients with night blindness and 12 unaffected family members. A complete ophthalmological examination including fundus photography and electroretinography (ERG) was performed on each family member. A genome-wide scan was performed using microsatellite markers at about 10 cM intervals, and two-point lod scores were calculated. Polymerase chain reaction (PCR) cycle dideoxynucleotide sequencing was used to screen candidate genes inside the linked region for mutations and to delineate the deletion. Multiplex PCR and long template PCR were used to detect deletions and to define the size of deletions. Evaluation of fundus changes and ERG, lod score estimation, and identification of a mutation in the GRK1 gene were carried out. RESULTS: All patients had night blindness since early childhood. Irregular coarse pigmentation was observed in the peripheral retina of each patient. The fundus appearance before and after 4 h of dark adaptation was similar except that the peripheral retinal pigmentary changes were slightly less evident after extended dark adaptation. Minimal or no rod function with normal cone function on ERG recordings were detected in all three affected members. The rod showed slow recovery to nearly normal amplitude after 4 h in the dark ERG in one individual but not in two other patients. A genome-wide scan showed linkage only to D13S285. Fine mapping defined a region from D13S1315 to 13qter, with a lod score of 2.89 at theta=0 shown by D13S285 and 2.90 at theta=0 by the D13S261-D13S285-D13S1295-D13S293 haplotype. Analysis of the GRK1 gene, which is included in this interval, identified a c.827+623_883del mutation. This intragenic deletion cosegregates with the disease in the family and is only homozygous in affected individuals. This mutation was not detected in 96 controls. CONCLUSIONS: The retinal disease in the family reported here has several features differing from typical Oguchi disease, including an atypical Mizuo-Nakamura phenomenon and a non-recordable rod ERG even after 4 h of dark adaptation. Normal visual acuity, normal caliber of retinal blood vessels, and normal cone response on ERG recording suggest retinal dysfunction rather than degeneration (i.e., a variant form of Oguchi disease but unlikely to be retinitis pigmentosa). The disease in the Pakistani family localizes to 13q34 and is caused by a novel deletion including Exon 3 of the GRK1 gene.

Adolescent↗

Tandem machine learning for the identification of genes regulated by transcription factors.

BACKGROUND: The identification of promoter regions that are regulated by a given transcription factor has traditionally relied upon the identification and distributions of binding sites recognized by the factor. In this study, we have developed a tandem machine learning approach for the identification of regulatory target genes based on these parameters and on the corresponding binding site information contents that measure the affinities of the factor for these cognate elements. RESULTS: This method has been validated using models of DNA binding sites recognized by the xenobiotic-sensitive nuclear receptor, PXR/RXRalpha, for target genes within the human genome. An information theory-based weight matrix was first derived and refined from known PXR/RXRalpha binding sites. The promoter region of candidate genes was scanned with the weight matrix. A novel information density-based clustering algorithm was then used to identify clusters of information rich sites. Finally, transformed data representing metrics of location, strength and clustering of binding sites were used for classification of promoter regions using an ensemble approach involving neural networks, decision trees and Naïve Bayesian classification. The method was evaluated on a set of 24 known target genes and 288 genes known not to be regulated by PXR/RXRalpha. We report an average accuracy (proportion of correctly classified promoter regions) of 71%, sensitivity of 73%, and specificity of 70%, based on multiple cross-validation and the leave-one-out strategy. The performance on a test set of 13 genes showed that 10 were correctly classified. CONCLUSION: We have developed a machine learning approach for the successful detection of gene targets for transcription factors with high accuracy. The method has been validated for the transcription factor PXR/RXRalpha and has the potential to be extended to other transcription factors.

Algorithms↗

Normal and abnormal mechanisms of gene splicing and relevance to inherited skin diseases.

The process of excising introns from pre-mRNA complexes is directed by specific genomic DNA sequences at intron-exon borders known as splice sites. These regions contain well-conserved motifs which allow the splicing process to proceed in a regulated and structured manner. However, as well as conventional splicing, several genes have the inherent capacity to undergo alternative splicing, thus allowing synthesis of multiple gene transcripts, perhaps with different functional properties. Within the human genome, therefore, through alternative splicing, it is possible to generate over 100,000 physiological gene products from the 35,000 or so known genes. Abnormalities in normal or alternative splicing, however, account for about 15% of all inherited single gene disorders, including many with a skin phenotype. These splicing abnormalities may arise through inherited mutations in constitutive splice sites or other critical intronic or exonic regions. This review article examines the process of normal intron-exon splicing, as well as what is known about alternative splicing of human genes. The review then addresses pathological disruption of normal intron-exon splicing that leads to inherited skin diseases, either resulting from mutations in sequences that have a direct influence on splicing or that generate cryptic splice sites. Examples of aberrant splicing, especially for the COL7A1 gene in patients with dystrophic epidermolysis bullosa, are discussed and illustrated. The review also examines a number of recently introduced computational tools that can be used to predict whether genomic DNA sequences changes may affect splice site selection and how robust the influence of such mutations might be on splicing.

Humans↗

Automated splicing mutation analysis by information theory.

Information theory-based software tools have been useful in interpreting noncoding sequence variation within functional sequence elements such as splice sites. Individual information analysis detects activated cryptic splice sites and associated splicing regulatory sites and is capable of distinguishing null from partially functional alleles. We present a server (https://splice.cmh.edu) designed to analyze splicing mutations in binding sites in either human genes, genome-mapped mRNAs, user-defined sequences, or dbSNP entries. Standard HUGO-approved gene symbols and HGVS-approved systematic mutation nomenclature (or dbSNP format) are entered via a web portal. After verifying the accuracy of input variant(s), the surrounding interval is retrieved from the human genome or user-supplied reference sequence. The server then computes the information contents (Ri) of all potential constitutive and/or regulatory splice sites in both the reference and variant sequences. Changes in information content are color-coded, tabulated, and visualized as sequence walkers, which display the binding sites with the reference sequence. The software was validated by analyzing approximately 1,300 mutations from Human Mutation as well as eight mapped SNPs from dbSNP designated as splice site variants. All of the splicing mutations and variants affected splice site strength or activated cryptic splice sites. The server also detected several missense mutations that were unexpectedly predicted to have concomitant effects on splicing or appeared to activate cryptic splicing.

Automation↗

Bipartite pattern discovery by entropy minimization-based multiple local alignment.

Many multimeric transcription factors recognize DNA sequence patterns by cooperatively binding to bipartite elements composed of half sites separated by a flexible spacer. We developed a novel bipartite algorithm, bipartite pattern discovery (Bipad), which produces a mathematical model based on information maximization or Shannon's entropy minimization principle, for discovery of bipartite sequence patterns. Bipad is a C++ program that applies greedy methods to search the bipartite alignment space and examines the upstream or downstream regions of co-regulated genes, looking for cis-regulatory bipartite patterns. An input sequence file with zero or one site per locus is required, and the left and right motif widths and a range of possible gap lengths must be specified. Bipad can run in either single-block or bipartite pattern search modes, and it is capable of comprehensively searching all four orientations of half-site patterns. Simulation studies showed that the accuracy of this motif discovery algorithm depends on sample size and motif conservation level, but results were independent of background composition. Bipad performed equivalent with or better than other pattern search algorithms in correctly identifying Escherichia coli cyclic AMP receptor protein and Bacillus subtilis sigma factor binding site sequences based on experimentally defined benchmarks. Finally, a new bipartite information weight matrix for vitamin D3 receptor/retinoid X receptor alpha (VDR/RXRalpha) binding sites was derived that comprehensively models the natural variability inherent in these sequence elements.

Algorithms↗

Development and refinement of pregnane X receptor (PXR) DNA binding site model using information theory: insights into PXR-mediated gene regulation.

The pregnane X receptor (PXR) acts as a receptor to induce gene expression in response to structurally diverse xenobiotics through binding as a heterodimer with the 9-cis retinoic acid receptor (RXR) to enhancers in target gene promoters. We identified and estimated the affinities of novel PXR/RXR binding sites in regulated genes and additional genomic targets of PXR with an information theory-based model of the PXR/RXR binding site. Our initial PXR/RXR model, the result of the alignment of 15 previously characterized binding sites, was used to scan the promoters of known PXR target genes. Sites from these genes, with information contents of >8 bits bound by PXR/RXR in vitro, were used to revise the information weight matrix; this procedure was repeated by screening for progressively weaker binding sites. After three iterations of refinement, the model was based on 48 validated PXR/RXR binding sites and has an average information content (Rsequence) of 14.43 +/- 3.21 bits. A scan of the human genome predicted novel PXR/RXR binding sites in the promoters of UGT1A3 (19.78 bits at -8040 and 16.37 bits at -6930) and UGT1A6 (12.74 bits at -9216), both of which were identified previously as targets for PXR. These sites were subsequently demonstrated to specifically bind PXR/RXR in competition electrophoretic mobility shift assays. A strong PXR site was also predicted upstream of the CASP10 gene (18.69 bits at -7872) and was validated by binding studies and reporter assays as a PXR responsive element. This suggests that the PXR-mediated response extends beyond genes involved in drug biotransformation and transport.

Binding Sites↗

Hepatic CYP2B6 expression: gender and ethnic differences and relationship to CYP2B6 genotype and CAR (constitutive androstane receptor) expression.

CYP2B6 metabolizes many drugs, and its expression varies greatly. CYP2B6 genotype-phenotype associations were determined using human livers that were biochemically phenotyped for CYP2B6 (mRNA, protein, and CYP2B6 activity), and genotyped for CYP2B6 coding and 5'-flanking regions. CYP2B6 expression differed significantly between sexes. Females had higher amounts of CYP2B6 mRNA (3.9-fold, P < 0.001), protein (1.7-fold, P < 0.009), and activity (1.6-fold, P < 0.05) than did male subjects. Furthermore, 7.1% of females and 20% of males were poor CYP2B6 metabolizers. Striking differences among different ethnic groups were observed: CYP2B6 activity was 3.6- and 5.0-fold higher in Hispanic females than in Caucasian (P < 0.022) or African-American females (P < 0.038). Ten single nucleotide polymorphisms (SNPs) in the CYP2B6 promoter and seven in the coding region were found, including a newly identified 13072A>G substitution that resulted in an Lys139Glu change. Many CYP2B6 splice variants (SV) were observed, and the most common variant lacked exons 4 to 6. A nonsynonymous SNP in exon 4 (15631G>T), which disrupted an exonic splicing enhancer, and a SNP 15582C>T in an intron-3 branch site were correlated with this SV. The extent to which CYP2B6 variation was a predictor of CYP2B6 activity varied according to sex and ethnicity. The 1459C>T SNP, which resulted in the Arg487Cys substitution, was associated with the lowest level of CYP2B6 activity in livers of females. The intron-3 15582C>T SNP (in significant linkage disequilibrium with a SNP in a putative hepatic nuclear factor 4 (HNF4) binding site) was correlated with lower CYP2B6 expression in females. In conclusion, we found several common SNPs that are associated with polymorphic CYP2B6 expression.

Adolescent↗

Genome-wide prediction, display and refinement of binding sites with information theory-based models.

BACKGROUND: We present Delila-genome, a software system for identification, visualization and analysis of protein binding sites in complete genome sequences. Binding sites are predicted by scanning genomic sequences with information theory-based (or user-defined) weight matrices. Matrices are refined by adding experimentally-defined binding sites to published binding sites. Delila-Genome was used to examine the accuracy of individual information contents of binding sites detected with refined matrices as a measure of the strengths of the corresponding protein-nucleic acid interactions. The software can then be used to predict novel sites by rescanning the genome with the refined matrices. RESULTS: Parameters for genome scans are entered using a Java-based GUI interface and backend scripts in Perl. Multi-processor CPU load-sharing minimized the average response time for scans of different chromosomes. Scans of human genome assemblies required 4-6 hours for transcription factor binding sites and 10-19 hours for splice sites, respectively, on 24- and 3-node Mosix and Beowulf clusters. Individual binding sites are displayed either as high-resolution sequence walkers or in low-resolution custom tracks in the UCSC genome browser. For large datasets, we applied a data reduction strategy that limited displays of binding sites exceeding a threshold information content to specific chromosomal regions within or adjacent to genes. An HTML document is produced listing binding sites ranked by binding site strength or chromosomal location hyperlinked to the UCSC custom track, other annotation databases and binding site sequences. Post-genome scan tools parse binding site annotations of selected chromosome intervals and compare the results of genome scans using different weight matrices. Comparisons of multiple genome scans can display binding sites that are unique to each scan and identify sites with significantly altered binding strengths. CONCLUSIONS: Delila-Genome was used to scan the human genome sequence with information weight matrices of transcription factor binding sites, including PXR/RXRalpha, AHR and NF-kappaB p50/p65, and matrices for RNA binding sites including splice donor, acceptor, and SC35 recognition sites. Comparisons of genome scans with the original and refined PXR/RXRalpha information weight matrices indicate that the refined model more accurately predicts the strengths of known binding sites and is more sensitive for detection of novel binding sites.

Binding Sites↗

Sequence-based, in situ detection of chromosomal abnormalities at high resolution.

We developed single copy probes from the draft genome sequence for fluorescence in situ hybridization (scFISH) which precisely delineate chromosome abnormalities at a resolution equivalent to genomic Southern analysis. This study illustrates how scFISH probes detect cryptic and subtle abnormalities and localize the sites of chromosome rearrangements. scFISH probes are substantially shorter than conventional recombinant DNA-derived probes, and C(o)t1 DNA is not required to suppress repetitive sequence hybridization. In this study, 74 single copy sequence probes (>1,500 bp) have been developed from >/=100 kb genomic intervals associated with either constitutional or acquired disorders. Applications of these probes include detection of congenital microdeletion syndromes on chromosomes 1, 4, 7, 15, 17, 22 and submicroscopic deletions involving the imprinting center on chromosome 15q11.2q13. We demonstrate how hybridization with multiple combinations of probes derived from the Smith-Magenis syndrome interval on chromosome 17 identified a patient with an atypical, proximal deletion breakpoint. A similar multi-probe hybridization strategy has also been used to delineate the translocation breakpoint region on chromosome 9 in chronic myelogenous leukemia. Probes have also been designed to hybridize to multiple cis paralogs, both enhancing the chromosomal target size and detecting chromosome rearrangements, for example, by splitting and separating a family of related sequences flanking an inversion breakpoint on chromosome 16 in acute myelogenous leukemia. These novel strategies for rapid and precise characterization of cytogenetic abnormalities are feasible because of the sequence-defined properties and dense euchromatic organization of single copy probes.

Adult↗

Information theory-based analysis of CYP2C19, CYP2D6 and CYP3A5 splicing mutations.

Several mutations are known or suspected to affect mRNA splicing of CYP2C19, CYP2D6 and CYP3A5 genes; however, little experimental evidence exists to support these conclusions. The present study applies mathematical models that measure changes in information content of splice sites in these genes to demonstrate the relationship between the predicted phenotypes of these variants to the corresponding genotypes. Based on information analysis, the CYP2C19*2 variant activates a new cryptic site 40 nucleotides downstream of the natural splice site. CYP2C19*7 abolishes splicing at the exon 5 donor site. The CYP2D6*4 allele similarly inactivates splicing at the acceptor site of exon 4 and activates a new cryptic site one nucleotide downstream of the natural acceptor. CYP2D6*11 inactivates the acceptor site of exon 2. The CYP3A5*3 allele activates a new cryptic site 236 nucleotides upstream of the exon 4 natural acceptor site. CYP3A5*5 inactivates the exon 5 donor site and CYP3A5*6 strengthens a site upstream of the natural donor site, resulting in skipping of exon 7. Other previously described missense and nonsense mutations at terminal codons of exons in these genes affected splicing. CYP2D6*8 and CYP2D6*14 both decrease the strength of the exon 3 donor site, producing transcripts lacking this exon. The results of information analysis are consistent with the poor metabolizer phenotypes observed in patients with these mutations, and illustrate the potential value of these mathematical models to quantitatively evaluate the functional consequences of new mutations suspected of altering mRNA splicing.

Amino Acid Substitution↗

Splice variants but not mutations of DNA polymerase beta are common in bladder cancer.

DNA polymerase beta (POLbeta) is a highly conserved protein that functions in base excision repair. Loss of the POLbeta locus on chromosome 8p is a frequent event in bladder cancer, and loss of POLbeta function could hinder DNA repair leading to a mutator phenotype. Both point mutations and large intragenic deletions of POLbeta have been reported from analysis of various tumor cDNAs but not from genomic DNA. We noticed that the breakpoints of the presumed rearrangements were delineated by exon-exon junctions, which could instead be consistent with alternative splicing of POLbeta mRNA. We tested the hypothesis that the reported intragenic deletion were splice variants by screening genomic DNA of human bladder tumors, bladder cancer cell lines, and normal bladder tissues for mutations or deletions in exons 1-14, exon alpha, and the promoter region of POLbeta. We found no evidence of somatic mutations or deletions in our sample set, although two polymorphisms were identified. Examination of cDNA from a subset of the original sample set revealed that truncated forms of POLbeta were surprisingly common. Forty-eight of 89 (54%) sequenced cDNA clones had large deletions, each beginning and/or ending exactly at exon-exon junctions. Because these deletions occur at exon-exon junctions and are seen in cDNA but not genomic DNA, they are consistent with alternative mRNA splicing. We describe 12 different splicing events occurring in 18 different combinations. Loss of exon 2 was the most frequent, being found in 42 of 49 (86%) of the variant sequenced clones. The splice variants appear to be somewhat more common and variable in bladder cancer cell lines and tumor tissues but occur at a high frequency in normal bladder tissues as well. We examine alternative splicing in terms of the information content of splice donor and acceptor site sequences, and discuss possible explanations for the predominant splicing event, the loss of exon 2.

Alternative Splicing↗