Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Intron annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Comparative sequence analysis of a gene-dense region among closely related species of Drosophila melanogaster.

Comparative sequence analysis among closely related species is essential for investigating the evolution of non-coding sequences, which evolve more rapidly than protein-coding sequences. We sequenced the cytogenetic map 56F10-16, a gene-dense region of D. simulans and D. sechellia, closely related species to D. melanogaster. About 57 kb of the genomic sequences containing 19 genes were annotated from each species according to the corresponding region of the D. melanogaster genome. The order and orientation of genes were perfectly conserved among the three species, and no transposable elements were found. The rate of nucleotide substitutions in the non-coding sequences was lower than that at the fourfold-degenerate sites, implying functional constraints in the non-coding regions. The sequence information from three closely related species, allowed us to estimate the insertions and the deletions that may have occurred in the lineages of D. simulans and D. sechellia using the D. melanogaster sequence as an outgroup. The number of deletions was twice that of insertions for the introns of D. simulans. More remarkably, the deletion outnumbered insertions by 7.5 times for the intergenic sequences of D. sechellia. These results suggest that the non-coding sequences have been shortened by deletion biases. However, the deletion bias was lower than that previously estimated for pseudogenes, suggesting that the non-coding sequences are already rich in functional elements, possibly involved in the regulation of gene expression including transcription and pre-mRNA processing. These features of non-coding sequences may be common to other gene-dense regions contributing to the compactness of the Drosophila genome.

Animals↗

Annotations and functional analyses of the rice WRKY gene superfamily reveal positive and negative regulators of abscisic acid signaling in aleurone cells.

The WRKY proteins are a superfamily of regulators that control diverse developmental and physiological processes. This family was believed to be plant specific until the recent identification of WRKY genes in nonphotosynthetic eukaryotes. We have undertaken a comprehensive computational analysis of the rice (Oryza sativa) genomic sequences and predicted the structures of 81 OsWRKY genes, 48 of which are supported by full-length cDNA sequences. Eleven OsWRKY proteins contain two conserved WRKY domains, while the rest have only one. Phylogenetic analyses of the WRKY domain sequences provide support for the hypothesis that gene duplication of single- and two-domain WRKY genes, and loss of the WRKY domain, occurred in the evolutionary history of this gene family in rice. The phylogeny deduced from the WRKY domain peptide sequences is further supported by the position and phase of the intron in the regions encoding the WRKY domains. Analyses for chromosomal distributions reveal that 26% of the predicted OsWRKY genes are located on chromosome 1. Among the dozen genes tested, OsWRKY24, -51, -71, and -72 are induced by abscisic acid (ABA) in aleurone cells. Using a transient expression system, we have demonstrated that OsWRKY24 and -45 repress ABA induction of the HVA22 promoter-beta-glucuronidase construct, while OsWRKY72 and -77 synergistically interact with ABA to activate this reporter construct. This study provides a solid base for functional genomics studies of this important superfamily of regulatory genes in monocotyledonous plants and reveals a novel function for WRKY genes, i.e. mediating plant responses to ABA.

Abscisic Acid↗

High-throughput expression, purification, and characterization of recombinant Caenorhabditis elegans proteins.

Modern proteomics approaches include techniques to examine the expression, localization, modifications, and complex formation of proteins in cells. In order to address issues of protein function in vitro using classical biochemical and biophysical approaches, high-throughput methods of cloning the appropriate reading frames, and expressing and purifying proteins efficiently are an important goal of modern proteomics approaches. This process becomes more difficult as functional proteomics efforts focus on the proteins from higher organisms, since issues of correctly identifying intron-exon boundaries and efficiently expressing and solubilizing the (often) multi-domain proteins from higher eukaryotes are challenging. Recently, 12,000 open-reading-frame (ORF) sequences from Caenorhabditis elegans have become available for functional proteomics studies [Nat. Gen. 34 (2003) 35]. We have implemented a high-throughput screening procedure to express, purify, and analyze by mass spectrometry hexa-histidine-tagged C. elegans ORFs in Escherichia coli using metal affinity ZipTips. We find that over 65% of the expressed proteins are of the correct mass as analyzed by matrix-assisted laser desorption MS. Many of the remaining proteins indicated to be "incorrect" can be explained by high-throughput cloning or genome database annotation errors. This provides a general understanding of the expected error rates in such high-throughput cloning projects. The ZipTip purified proteins can be further analyzed under both native and denaturing conditions for functional proteomics efforts.

Animals↗

Cis-regulatory variations: a study of SNPs around genes showing cis-linkage in segregating mouse populations.

BACKGROUND: Changes in gene expression are known to be responsible for phenotypic variation and susceptibility to diseases. Identification and annotation of the genomic sequence variants that cause gene expression changes is therefore likely to lead to a better understanding of the cause of disease at the molecular level. In this study we investigate the pattern of single nucleotide polymorphisms (SNPs) in genes for which the mRNA levels show cis-genetic linkage (gene expression quantitative trait loci mapping in cis, or cis-eQTLs) in segregating mouse populations. Such genes are expected to have polymorphisms near their physical location (cis-variations) that affect their mRNA levels by altering one or more of the cis-regulatory elements. This led us to characterize the SNPs in promoter (5 Kb upstream) and non-coding gene regions (introns and 5 Kb downstream) (cis-SNPs) and the effects they may have on putative transcription factor binding sites. RESULTS: We demonstrate that the cis-eQTL genes (CEGs) have a significantly higher frequency of cis-SNPs compared to non-CEGs (when both sets are taken from the non-IBD regions, i.e. regions not identical by descent). Most CEGs having cis-SNPs do not contain these SNPs in the phylogenetically conserved regions. In those CEGs that contain cis-SNPs in the phylogenetically conserved regions, enrichment of cis-SNPs occurs both within and outside of the conserved sequences. A higher fraction of CEGs are also seen to harbor cis-SNP that affect predicted transcription factor binding sites, a likely consequence of the higher cis-SNPs density in these genes. CONCLUSION: This present study provides the first genome-wide investigation of the putative cis-regulatory variations in a large set of genes whose levels of expression give rise to cis-linkage in segregating mammalian populations. Our results provide insights into the challenges that exist in identifying polymorphisms regulating gene expression using bioinformatic sequence analysis approaches. The data provided herein should benefit future investigations in this area.

Adipose Tissue↗

Iterative gene prediction and pseudogene removal improves genome annotation.

Correct gene prediction is impaired by the presence of processed pseudogenes: nonfunctional, intronless copies of real genes found elsewhere in the genome. Gene prediction programs frequently mistake processed pseudogenes for real genes or exons, leading to biologically irrelevant gene predictions. While methods exist to identify processed pseudogenes in genomes, no attempt has been made to integrate pseudogene removal with gene prediction, or even to provide a freestanding tool that identifies such erroneous gene predictions. We have created PPFINDER (for Processed Pseudogene finder), a program that integrates several methods of processed pseudogene finding in mammalian gene annotations. We used PPFINDER to remove pseudogenes from N-SCAN gene predictions, and show that gene prediction improves substantially when gene prediction and pseudogene masking are interleaved. In addition, we used PPFINDER with gene predictions as a parent database, eliminating the need for libraries of known genes. This allows us to run the gene prediction/PPFINDER procedure on newly sequenced genomes for which few genes are known.

Animals↗

SpliceDB: database of canonical and non-canonical mammalian splice sites.

A database (SpliceDB) of known mammalian splice site sequences has been developed. We extracted 43 337 splice pairs from mammalian divisions of the gene-centered Infogene database, including sites from incomplete or alternatively spliced genes. Known EST sequences supported 22 815 of them. After discarding sequences with putative errors and ambiguous location of splice junctions the verified dataset includes 22 489 entries. Of these, 98.71% contain canonical GT-AG junctions (22 199 entries) and 0.56% have non-canonical GC-AG splice site pairs. The remainder (0.73%) occurs in a lot of small groups (with a maximum size of 0.05%). We especially studied non-canonical splice sites, which comprise 3.73% of GenBank annotated splice pairs. EST alignments allowed us to verify only the exonic part of splice sites. To check the conservative dinucleotides we compared sequences of human non-canonical splice sites with sequences from the high throughput genome sequencing project (HTG). Out of 171 human non-canonical and EST-supported splice pairs, 156 (91.23%) had a clear match in the human HTG. They can be classified after sequence analysis as: 79 GC-AG pairs (of which one was an error that corrected to GC-AG), 61 errors corrected to GT-AG canonical pairs, six AT-AC pairs (of which two were errors corrected to AT-AC), one case was produced from a non-existent intron, seven cases were found in HTG that were deposited to GenBank and finally there were only two other cases left of supported non-canonical splice pairs. The information about verified splice site sequences for canonical and non-canonical sites is presented in SpliceDB with the supporting evidence. We also built weight matrices for the major splice groups, which can be incorporated into gene prediction programs. SpliceDB is available at the computational genomic Web server of the Sanger Centre: http://genomic.sanger.ac. uk/spldb/SpliceDB.html and at http://www.softberry. com/spldb/SpliceDB.html.

Animals↗

Two Sample Logo: a graphical representation of the differences between two sets of sequence alignments.

SUMMARY: Two Sample Logo is a web-based tool that detects and displays statistically significant differences in position-specific symbol compositions between two sets of multiple sequence alignments. In a typical scenario, two groups of aligned sequences will share a common motif but will differ in their functional annotation. The inclusion of the background alignment provides an appropriate underlying amino acid or nucleotide distribution and addresses intersite symbol correlations. In addition, the difference detection process is sensitive to the sizes of the aligned groups. Two Sample Logo extends WebLogo, a widely-used sequence logo generator. The source code is distributed under the MIT Open Source license agreement and is available for download free of charge.

Algorithms↗

Fugu and human sequence comparison identifies novel human genes and conserved non-coding sequences.

The compact genome of the pufferfish, Fugu rubripes, has been proposed as a 'reference' genome to aid in annotating and analysing the human genome. We have annotated and compared 85 kb of Fugu sequence containing 17 genes with its homologous loci in the human draft genome and identified three 'novel' human genes that were missed or incompletely predicted by the previous gene prediction methods. Two of the novel genes contain zinc finger domains and are designated ZNF366 and ZNF367. They map to human chromosomes 5q13.2 and 9q22.32, respectively. The third novel gene, designated C9orf21, maps to chromosome 9q22.32. This gene is unique to vertebrates, and the protein encoded by it does not contain any known domains. We could not find human homologs for two Fugu genes, a novel chemokine gene and a kinase gene. These genes are either specific to teleosts or lost in the human lineage. The Fugu-human comparison identified several conserved non-coding sequences in the promoter and intronic regions. These sequences, conserved during 450 million years of vertebrate evolution, are likely to be involved in gene regulation. The 85 kb Fugu locus is dispersed over four human loci, occupying about 1.5 Mb. Contiguity is conserved in the human genome between six out of 16 Fugu gene pairs. These contiguous chromosomal segments should share a common evolutionary history dating back to the common ancestor of mammals and teleosts. We propose contiguity as strong evidence to identify orthologous genes in distant organisms. This study confirms the utility of the Fugu as a supplementary tool to uncover and confirm novel genes and putative gene regulatory regions in the human genome.

Amino Acid Sequence↗

Detailed mapping of the ERG-ETS2 interval of human chromosome 21 and comparison with the region of conserved synteny on mouse chromosome 16.

We have carried out a detailed annotation of 550 kb of genomic DNA on human chromosome 21 containing the ERG and ETS2 genes. Comparative genomic analysis between this region and the interval of conserved synteny on mouse chromosome 16 indicated that the order and orientation of the ERG and ETS2 genes were conserved and revealed several regions containing potential conserved noncoding sequences. Four pseudogenes including those for small protein G, laminin receptor, human transposase protein and meningioma-expressed antigen were identified. A potentially novel gene (C21orf24) with alternative mRNA transcripts, consensus splice donor and acceptor sites, but no coding potential nor murine orthologue, was identified and found to be expressed in a range of human cell lines. We have identified four novel splice variants that arise from a previously undescribed 5' exon of the human ERG gene. Comparison of the cDNA sequences enabled us to determine the complete exon-intron structure of the ERG gene. We have also identified the presence of noncoding RNAs in the first and second introns of the ETS2 gene. Our studies have important implications for Down syndrome as this region contains multiple mRNA transcripts, both coding and potentially noncoding, that may play as yet undescribed roles in the pathogenesis of this disorder.

Alternative Splicing↗

Genetic analysis of four cases of Poirier Bienvenu neurodevelopmental syndrome associated with CSNK2B variant.

BACKGROUND: CSNK2B deficiency underlies the pathogenesis of Poirier-Bienvenu neurodevelopmental syndrome (POBINDS). In this study, we present four cases of pediatric seizures caused by de novo variants in CSNK2B, with the aim to reinforce the clinical and variant data pertaining to early genetic factors associated with epilepsy. METHODS: Trio whole exome sequencing were used to detect variants in the proband and her family members, and bioinformatics annotation was performed for the variant. Sanger sequencing and CSNK2B cDNA sequencing were employed to ascertain the carrier status of additional family members and evaluate the potential impact of variants on splicing. RESULTS: All four cases presented with epilepsy as the initial manifestation, accompanied by global developmental delay, particularly in language and motor developmental delay. Cases 1, 3 and 4 exhibited full-scale tonic-clonic seizures, while case 2 displayed myoclonic and typical absence seizures. Furthermore, case 2 demonstrated delayed growth and development compared to age-matched peers. No abnormality was detected in the head magnetic resonance imaging (MRI). Genetic analysis revealed novel heterozygous variants in the CSNK2B gene in all four cases, including c.175 + 1G > A, c.73-2A > G, c.291 + 1G > A and c.481delA. In case 2, reverse transcription analysis of CSNK2B mRNA revealed the retention of the 3' end sequence of Intron 2 and deletion of the 5' end sequence of Exon 3. In treatment, four case received a combination of one to three types of antiseizure medication and rehabilitation training individually. Case 1 continued to experience seizures to varying degrees, while cases 2-4 demonstrated effective seizure control. Overall motor and intellectual development improved in all four cases, however, there was slow recovery in language function. CONCLUSION: This study elucidates the molecular etiology of epilepsy in four cases with POBINDS and expands the mutational spectrum of pathogenic variants in the CSNK2B, highlighting their impact on splicing. The highly genetic heterogeneous phenotype of POBINDS relies on the detection of pathogenic variants in CSNK2B. Conventional antiseizure medication effectively control seizures, while rehabilitation treatment can significantly improve intelligence and motor function to varying degrees; however, language recovery tends to be relatively slow.

Humans↗

SAGE2Splice: unmapped SAGE tags reveal novel splice junctions.

Serial analysis of gene expression (SAGE) not only is a method for profiling the global expression of genes, but also offers the opportunity for the discovery of novel transcripts. SAGE tags are mapped to known transcripts to determine the gene of origin. Tags that map neither to a known transcript nor to the genome were hypothesized to span a splice junction, for which the exon combination or exon(s) are unknown. To test this hypothesis, we have developed an algorithm, SAGE2Splice, to efficiently map SAGE tags to potential splice junctions in a genome. The algorithm consists of three search levels. A scoring scheme was designed based on position weight matrices to assess the quality of candidates. Using optimized parameters for SAGE2Splice analysis and two sets of SAGE data, candidate junctions were discovered for 5%-6% of unmapped tags. Candidates were classified into three categories, reflecting the previous annotations of the putative splice junctions. Analysis of predicted tags extracted from EST sequences demonstrated that candidate junctions having the splice junction located closer to the center of the tags are more reliable. Nine of these 12 candidates were validated by RT-PCR and sequencing, and among these, four revealed previously uncharacterized exons. Thus, SAGE2Splice provides a new functionality for the identification of novel transcripts and exons. SAGE2Splice is available online at http://www.cisreg.ca.

Algorithms↗

Cloning and initial characterization of the Arabidopsis thaliana endoplasmic reticulum oxidoreductins.

The oxidation and isomerization of disulfide bonds is necessary for the growth of all organisms. In yeast, the oxidative folding of secretory pathway proteins is catalyzed by protein disulfide isomerase (PDI), which requires Ero1p (endoplasmic reticulum oxidoreductin) for its own oxidation. In Homo sapiens, two homologues of Ero1p, Ero1-Lalpha and Ero1-Lbeta, have been cloned. Both Ero1-Lalpha and Ero1-Lbeta interact via disulfide bonds with PDI and support the oxidation of immunoglobulin light chains. However, the function of Ero proteins in plants has not yet been analyzed. In this article, we report the cloning of the two Ero1p homologues present in Arabidopsis thaliana, demonstrating that one of the cDNAs has a shorter terminal exon than predicted and differs from the annotated sequence found in the genome database. Sequence analysis of the Arabidopsis endoplasmic reticulum oxidoreductins (AEROs) reveals that both AERO1 and AERO2 are more closely related to each other than to either of the human Eros. Both in vitro translated AERO proteins are targeted to the endoplasmic reticulum and glycosylated. The ability to use a genetically tractable multicellular organism in combination with biochemical approaches should further our understanding of redox networks and Ero function in both plants and animals.

Arabidopsis↗

A novel gene encoding a putative transmembrane protein with two extracellular CUB domains and a low-density lipoprotein class A module: isolation of alternatively spliced isoforms in retina and brain.

We report herein the cDNA cloning of a novel retina and brain specific gene from mouse and human encoding a putative transmembrane protein with an N-terminal signal sequence and two conserved extracellular CUB domains followed by a single copy of the low-density lipoprotein class A (LDLa) module. The mouse and human genes, termed NETO1 (neuropilin and tolloid like-1), display sequence identities of 87% at the nucleotide and 95% at the protein level. The human NETO1 gene comprises 13 exons on chromosome 18q22-q23 and gives rise to three different mRNA isoforms. Two alternative leader exons 1a and 1b generate transcripts that translate into putative signal peptides with individual sequence composition but otherwise do not affect the primary structure of the mature NETO1 protein. Usage of the internal exon 5 is restricted to the retinal tissue and generates a truncated transcript that codes for a putative soluble protein, termed sNETO1, with only one copy of the CUB domain while lacking the LDLa module. NETO1 exhibits 57% identity to the deduced amino acid sequence of a non-annotated nucleotide sequence in the GenBank database, therefore designated NETO2. Both NETO1 and NETO2 share an identical and unique domain structure thus representing a novel subfamily of CUB- and LDLa-containing proteins. The cytoplasmic domains of NETO1 and NETO2 are not homologous to other known protein sequences but contain a conserved FXNPXY-like motif, which is essential for the internalization of clathrin coated pits during endocytosis or alternatively, may be implicated in intracellular signaling pathways.

Alternative Splicing↗

Computational evidence for hundreds of non-conserved plant microRNAs.

BACKGROUND: MicroRNAs (miRNA) are small (20-25 nt) non-coding RNA molecules that regulate gene expression through interaction with mRNA in plants and metazoans. A few hundred miRNAs are known or predicted, and most of those are evolutionarily conserved. In general plant miRNA are different from their animal counterpart: most plant miRNAs show near perfect complementarity to their targets. Exploiting this complementarity we have developed a method for identification plant miRNAs that does not rely on phylogenetic conservation. RESULTS: Using the presumed targets for the known miRNA as positive controls, we list and filter all segments of the genome of length approximately 20 that are complementary to a target mRNA-transcript. From the positive control we recover 41 (of 92 possible) of the already known miRNA-genes (representing 14 of 16 families) with only four false positives. Applying the procedure to find possible new miRNAs targeting any annotated mRNA, we predict of 592 new miRNA genes, many of which are not conserved in other plant genomes. A subset of our predicted miRNAs is additionally supported by having more than one target that are not homologues. CONCLUSION: These results indicate that it is possible to reliably predict miRNA-genes without using genome comparisons. Furthermore it suggests that the number of plant miRNAs have been underestimated and points to the existence of recently evolved miRNAs in Arabidopsis.

Animals↗

The Anopheles gambiae glutathione transferase supergene family: annotation, phylogeny and expression profiles.

BACKGROUND: Twenty-eight genes putatively encoding cytosolic glutathione transferases have been identified in the Anopheles gambiae genome. We manually annotated these genes and then confirmed the annotation by sequencing of A. gambiae cDNAs. Phylogenetic analysis with the 37 putative GST genes from Drosophila and representative GSTs from other taxa was undertaken to develop a nomenclature for insect GSTs. The epsilon class of insect GSTs has previously been implicated in conferring insecticide resistance in several insect species. We compared the expression level of all members of this GST class in two strains of A. gambiae to determine whether epsilon GST expression is correlated with insecticide resistance status. RESULTS: Two A. gambiae GSTs are alternatively spliced resulting in a maximum number of 32 transcripts encoding cytosolic GSTs. We detected cDNAs for 31 of these in adult mosquitoes. There are at least six different classes of GSTs in insects but 20 of the A. gambiae GSTs belong to the two insect specific classes, delta and epsilon. Members of these two GST classes are clustered on chromosome arms 2L and 3R respectively. Two members of the GST supergene family are intronless. Amongst the remainder, there are 13 unique introns positions but within the epsilon and delta class, there is considerable conservation of intron positions. Five of the eight epsilon GSTs are overexpressed in a DDT resistant strain of A. gambiae. CONCLUSIONS: The GST supergene family in A. gambiae is extensive and regulation of transcription of these genes is complex. Expression profiling of the epsilon class supports earlier predictions that this class is important in conferring insecticide resistance.

Alternative Splicing↗

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting.

A major challenge in clinical genomics is to classify genetic variations correctly, since it directly affects disease diagnosis and personal care. The existing methods tend to be based on the combination of different factors, such as protein structure, population frequencies, phenotypic annotations, and sequence conservation. Nevertheless, these methods often cannot be used to achieve the necessary interpretability, quantify uncertainty, and address rare cases. This paper presents a probabilistic gradient boosting model on variant pathogenicity prediction. The suggested framework applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making. Our machine learning aims to solve the issues of variant interpretation by managing the features and through probability-based pathogenicity prediction. The framework formulation is aimed at generalizing over various datasets and minimizing overfitting. At the same time, it can ensure reasonable performance to facilitate clinical experiments. The model has also been tested on three standard datasets and demonstrated to be more predictive of the pathogenic effect of variants, in comparison with a variety of existing tools. The probabilistic gradient boosting model proposed had ROC AUC values of 0.9293, 0.9610, and 0.9646 on ClinVar variants, GRCh37, and GRCh38 human genome respectively. Furthermore, the dataset was ensured to include both exonic and intronic variants, and Variants of Uncertain Significance were also taken into consideration for Performance Testing. Through this it also aims to provide better clinical significance which will lead to a good interpretable tool for priority of variants for a large variety of disease conditions.

ClinVar↗

The interaction networks of structured RNAs.

All pairwise interactions occurring between bases which could be detected in three-dimensional structures of crystallized RNA molecules are annotated on new planar diagrams. The diagrams attempt to map the underlying complex networks of base-base interactions and, especially, they aim at conveying key relationships between helical domains: co-axial stacking, bending and all Watson-Crick as well as non-Watson-Crick base pairs. Although such wiring diagrams cannot replace full stereographic images for correct spatial understanding and representation, they reveal structural similarities as well as the conserved patterns and distances between motifs which are present within the interaction networks of folded RNAs of similar or unrelated functions. Finally, the diagrams could help devising methods for meaningfully transforming RNA structures into graphs amenable to network analysis.

Base Pairing↗

Analysis of transcriptional regulation of the small leucine rich proteoglycans.

PURPOSE: Small leucine rich proteoglycans (SLRPs) constitute a family of secreted proteoglycans that are important for collagen fibrillogenesis, cellular growth, differentiation, and migration. Ten of the 13 known members of the SLRP gene family are arranged in tandem clusters on human chromosomes 1, 9, and 12. Their syntenic equivalents are on mouse chromosomes 1, 13, and 10, and rat chromosomes 13, 17, and 7. The purpose of this study was to determine whether there is evidence for control elements, which could regulate the expression of these clusters coordinately. METHODS: Promoters were identified using a comparative genomics approach and Genomatix software tools. For each gene a set of human, mouse, and rat orthologous promoters was extracted from genomic sequences. Transcription factor (TF) binding site analysis combined with a literature search was performed using MatInspector and Genomatix' BiblioSphere. Inspection for the presence of interspecies conserved scaffold/matrix attachment regions (S/MARs) was performed using ElDorado annotation lists. DNAseI hypersensitivity assay, chromatin immunoprecipitation (ChIP), and transient transfection experiments were used to validate the results from bioinformatics analysis. RESULTS: Transcription factor binding site analysis combined with a literature search revealed co-citations between several SLRPs and TFs Runx2 and IRF1, indicating that these TFs have potential roles in transcriptional regulation of the SLRP family members. We therefore inspected all of the SLRP promoter sets for matches to IRF factors and Runx factors. Positionally conserved binding sites for the Runt domain TFs were detected in the proximal promoters of chondroadherin (CHAD) and osteomodulin (OMD) genes. Two significant models (two or more transcription factor binding sites arranged in a defined order and orientation within a defined distance range) were derived from these initial promoter sets, the HOX-Runx (homeodomain-Runt domain), and the ETS-FKHD-STAT (erythroblast transformation specific-forkhead-signal transducers and activators of transcription) models. These models were used to scan the genomic sequences of all 13 SLRP genes. The HOX-Runx model was found within the proximal promoter, exon 1, or intron 1 sequences of 11 of the 13 SLRP genes. The ETS-FKHD-STAT model was found in only 5 of these genes. Transient transfections of MG-63 cells and bovine corneal keratocytes with Runx2 isoforms confirmed the relevance of these TFs to expression of several SLRP genes. Distribution of the HOX-Runx and ETS-FKHD-STAT models within 200 kb of genomic sequence on human chromosome 9 and 500 kb sequence on chromosome 12 also were analyzed. Two regions with 3 HOX-Runx matches within a 1,000 bp window were identified on human chromosome 9; one located between OMD and osteoglycin (OGN)/mimecan genes, and the second located upstream of the putative extracellular matrix protein 2 (ECM2) promoter. The intergenic region between OMD and mimecan was shown to coincide with different patterns of DNAse I hypersensitivity sites in MG-63 and U937 cells. ChiP analysis revealed that this region binds Runx2 in U937 cells (mimecan transcript note detectable), but binds Pitx3 in MG-63 cells (expressing high level of mimecan), thereby demonstrating its functional association with mimecan expression. Upon comparing the predictions of S/MARs on the relevant chromosomal context of human chromosomes 9 and 12 and their rodent equivalents, no convincing evidence was found that the tandemly arranged genes build a chromosomal loop. CONCLUSIONS: Twelve of 13 known SLRP genes have at least one HOX-Runx module match in their promoter, exon 1, intron 1, or intergenic region. Although these genes are located in different clusters on different chromosomes, the common HOX-Runx module could be the basis for co-regulated expression.

Animals↗