Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Intron annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Molecular evolution of the AMP-forming Acetyl-CoA synthetase.

Acetyl-CoA-Synthetase (ACS) is involved in the production of acetate, a major metabolite in numerous organisms. There are two forms of this enzyme: ADP-forming ACS and ATP-forming ACS. We focus mainly on the AMP-forming ACS gene, which is relatively well conserved in eubacteria, archeaebacteria, and eukaryotes. BLAST searches in databases showed 30 protein sequences significantly related to the ACS. Most of these sequences were identified as ACS but three of them, belonging to the mammalian species, were annotated as another gene named: the SA gene, which is involved in the essential hypertension. The ACS and SA genes probably derived from a duplication of an ancestral gene but have acquired different functions. Six conserved regions of the ACS protein were defined across the three domains of life. While the precise function of the conserved regions remains unknown, they are probably involved in the enzymatic activity. Among eukaryotes, we found a high variability with respect to the number and the position of introns. However, some positions are conserved between fungi and a nematode. A maximum likelihood tree based upon the conserved regions showed that all sequences except the one from B. subtilis, belong to two basic groups: one the SA-like group including sequences from Archaeoglobus fulgidus and Streptomyces coelicolor, and second, the ACS group. The later can be further divided in two parts: a prokaryotic one including eubacteria and an archaebacterium, and a eukaryotic group within which two proteobacterial sequences branch including ACS from the alpha-proteobacterium Rhodobacter capsulatus. Within the eukaryotic group, bootstrap support is very low, but overall the data are consistent with the view that eukaryotes acquired their ACS gene from the ancestors of mitochondria. The localization of this enzyme in eukaryotic mitochondria is the additional evidence in favor of this interpretation.

Acetate-CoA Ligase↗

Quantitative analysis of nucleic acid three-dimensional structures.

A new computer program to annotate DNA and RNA three-dimensional structures, MC-Annotate, is introduced. The goals of annotation are to efficiently extract and manipulate structural information, to simplify further structural analyses and searches, and to objectively represent structural knowledge. The input of MC-Annotate is a PDB formatted DNA or RNA three-dimensional structure. The output of MC-Annotate is composed of a structural graph that contains the annotations, and a series of HTML documents, one for each nucleotide conformation and base-base interaction present in the input structure. The atomic coordinates of all nucleotides and the homogeneous transformation matrices of all base-base interactions are stored in the structural graph. Symbolic classifications of nucleotide conformations, using sugar puckering modes and nitrogen base orientations around the glycosyl bond, and base-base interactions, using stacking and hydrogen bonding information, are introduced. Peculiarity factors of nucleotide conformations and base-base interactions are defined to indicate their marginalities with all other examples. The peculiarity factors allow us to identify irregular regions and possible stereochemical errors in 3-D structures without interactive visualization. The annotations attached to each nucleotide conformation include its class, its torsion angles, a distribution of the root-mean-square deviations with examples of the same class, the list of examples of the same class, and its peculiarity value. The annotations attached to each base-base interaction include its class, a distribution of distances with examples of the same class, the list of examples of the same class, and its peculiarity value. The distance between two homogeneous transformation matrices is evaluated using a new metric that distinguishes between the rotation and the translation of a transformation matrix in the context of nitrogen bases. MC-Annotate was used to build databases of nucleotide conformations and base-base interactions. It was applied to the ribosomal RNA fragment that binds to protein L11, which annotations revealed peculiar nucleotide conformations and base-base interactions in the regions where the RNA contacts the protein. The question of whether the current database of RNA three-dimensional structures is complete is addressed.

Base Pairing↗

Global identification of human transcribed sequences with genome tiling arrays.

Elucidating the transcribed regions of the genome constitutes a fundamental aspect of human biology, yet this remains an outstanding problem. To comprehensively identify coding sequences, we constructed a series of high-density oligonucleotide tiling arrays representing sense and antisense strands of the entire nonrepetitive sequence of the human genome. Transcribed sequences were located across the genome via hybridization to complementary DNA samples, reverse-transcribed from polyadenylated RNA obtained from human liver tissue. In addition to identifying many known and predicted genes, we found 10,595 transcribed sequences not detected by other methods. A large fraction of these are located in intergenic regions distal from previously annotated genes and exhibit significant homology to other mammalian proteins.

Animals↗

Genome-wide analysis of NBS-LRR-encoding genes in Arabidopsis.

The Arabidopsis genome contains approximately 200 genes that encode proteins with similarity to the nucleotide binding site and other domains characteristic of plant resistance proteins. Through a reiterative process of sequence analysis and reannotation, we identified 149 NBS-LRR-encoding genes in the Arabidopsis (ecotype Columbia) genomic sequence. Fifty-six of these genes were corrected from earlier annotations. At least 12 are predicted to be pseudogenes. As described previously, two distinct groups of sequences were identified: those that encoded an N-terminal domain with Toll/Interleukin-1 Receptor homology (TIR-NBS-LRR, or TNL), and those that encoded an N-terminal coiled-coil motif (CC-NBS-LRR, or CNL). The encoded proteins are distinct from the 58 predicted adapter proteins in the previously described TIR-X, TIR-NBS, and CC-NBS groups. Classification based on protein domains, intron positions, sequence conservation, and genome distribution defined four subgroups of CNL proteins, eight subgroups of TNL proteins, and a pair of divergent NL proteins that lack a defined N-terminal motif. CNL proteins generally were encoded in single exons, although two subclasses were identified that contained introns in unique positions. TNL proteins were encoded in modular exons, with conserved intron positions separating distinct protein domains. Conserved motifs were identified in the LRRs of both CNL and TNL proteins. In contrast to CNL proteins, TNL proteins contained large and variable C-terminal domains. The extant distribution and diversity of the NBS-LRR sequences has been generated by extensive duplication and ectopic rearrangements that involved segmental duplications as well as microscale events. The observed diversity of these NBS-LRR proteins indicates the variety of recognition molecules available in an individual genotype to detect diverse biotic challenges.

Amino Acid Sequence↗

Evolutionary rate variation in eukaryotic lineage specific human intronless proteins.

The present study examines 783 human-mouse orthologous gene pairs for their pattern of sequence evolution, contrasting mammalia, eukaryota, coelomata, and bilateria specific human intronless genes. Such comparisons may be of use in understanding the general evolution of human genome. Evolutionary rate analyses indicate that mammalia specific human intronless genes are evolving faster as compared to other intronless genes specific to eukaryotic lineage, indicating towards their rapid evolution. The observations indicates that the genes conserved in eukaryota, coelomata, and bilateria, that is, proteins that arose earlier in evolution as compared to mammalia specific genes evolve slowly and are subjected to negative selection. The cause underlying rate variations was also explored. Although mutational bias might slightly fasten the nonsynonymous rates in mammalia specific genes, it is unlikely to be major cause of rate difference between the various categories. Furthermore, rate of divergence of mammalia specific intronless genes has been related to functional classification using the protein family annotation. Protein function was found in some cases to have larger impact on the rate of evolution of genes. Also, the codon usage pattern of mammalia specific intronless genes do not seem to differ much from those of other intronless genes conserved solely in eukaryotic lineage.

Animals↗

AluGene: a database of Alu elements incorporated within protein-coding genes.

Alu elements are short interspersed elements (SINEs) approximately 300 nucleotides in length. More than 1 million Alus are found in the human genome. Despite their being genetically functionless, recent findings suggest that Alu elements may have a broad evolutionary impact by affecting gene structures, protein sequences, splicing motifs and expression patterns. Because of these effects, compiling a genomic database of Alu sequences that reside within protein-coding genes seemed a useful enterprise. Presently, such data are limited since the structural and positional information on genes and Alu sequences are scattered throughout incompatible and unconnected databases. AluGene (http://Alugene.tau.ac.il/) provides easy access to a complete Alu map of the human genome, as well as Alu-associated information. The Alu elements are annotated with respect to coding region and exon/intron location. This design facilitates queries on Alu sequences, locations, as well as motifs and compositional properties via a one-stop search page.

Alu Elements↗

Djinn Lite: a tool for customised gene transcript modelling, annotation-data enrichment and exploration.

BACKGROUND: There is an ever increasing rate of data made available on genetic variation, transcriptomes and proteomes. Similarly, a growing variety of bioinformatic programs are becoming available from many diverse sources, designed to identify a myriad of sequence patterns considered to have potential biological importance within inter-genic regions, genes, transcripts, and proteins. However, biologists require easy to use, uncomplicated tools to integrate this information, visualise and print gene annotations. Integrating this information usually requires considerable informatics skills, and comprehensive knowledge of the data format to make full use of this information. Tools are needed to explore gene model variants by allowing users the ability to create alternative transcript models using novel combinations of exons not necessarily represented in current database deposits of mRNA/cDNA sequences. RESULTS: Djinn Lite is designed to be an intuitive program for storing and visually exploring of custom annotations relating to a eukaryotic gene sequence and its modelled gene products. In particular, it is helpful in developing hypothesis regarding alternate splicing of transcripts by allowing the construction of model transcripts and inspection of their resulting translations. It facilitates the ability to view a gene and its gene products in one synchronised graphical view, allowing one to drill down into sequence related data. Colour highlighting of selected sequences and added annotations further supports exploration, visualisation of sequence regions and motifs known or predicted to be biologically significant. CONCLUSION: Gene annotating remains an ongoing and challenging task that will continue as gene structures, gene transcription repertoires, disease loci, protein products and their interactions become more precisely defined. Djinn Lite offers an accessible interface to help accumulate, enrich, and individualize sequence annotations relating to a gene, its transcripts and translations. The mechanism of transcript definition and creation, and subsequent navigation and exploration of features, are very intuitive and demand only a short learning curve. Ultimately, Djinn Lite can form the basis for providing valuable clues to plan new experiments, providing storage of sequences and annotations for dedication to customised projects. The application is appropriate for Windows 98-ME-2000-XP-2003 operating systems.

Alternative Splicing↗

A Drosophila full-length cDNA resource.

BACKGROUND: A collection of sequenced full-length cDNAs is an important resource both for functional genomics studies and for the determination of the intron-exon structure of genes. Providing this resource to the Drosophila melanogaster research community has been a long-term goal of the Berkeley Drosophila Genome Project. We have previously described the Drosophila Gene Collection (DGC), a set of putative full-length cDNAs that was produced by generating and analyzing over 250,000 expressed sequence tags (ESTs) derived from a variety of tissues and developmental stages. RESULTS: We have generated high-quality full-insert sequence for 8,921 clones in the DGC. We compared the sequence of these clones to the annotated Release 3 genomic sequence, and identified more than 5,300 cDNAs that contain a complete and accurate protein-coding sequence. This corresponds to at least one splice form for 40% of the predicted D. melanogaster genes. We also identified potential new cases of RNA editing. CONCLUSIONS: We show that comparison of cDNA sequences to a high-quality annotated genomic sequence is an effective approach to identifying and eliminating defective clones from a cDNA collection and ensure its utility for experimentation. Clones were eliminated either because they carry single nucleotide discrepancies, which most probably result from reverse transcriptase errors, or because they are truncated and contain only part of the protein-coding sequence.

Amino Acid Sequence↗

Y chromosome and other heterochromatic sequences of the Drosophila melanogaster genome: how far can we go?

Whole genome shotgun assemblies have proven remarkably successful in reconstructing the bulk of euchromatic genes, with the only limit appearing to be determined by the sequencing depth. For genes imbedded in heterochromatin, however, the low cloning efficiency of repetitive sequences, combined with the computational challenges, demand that additional clues be used to annotate the sequences. One approach that has proven very successful in identifying protein coding genes in Y-linked heterochromatin of Drosophila melanogaster has been to make a BLASTable database of the small, unmapped contigs and fragments leftover at the end of a shotgun assembly, and to attempt to capture these by blasting with an appropriate query sequence. This approach often yields a staggered alignment of contigs from the unmapped set to the query sequence, as though the disjoint contigs represent small portions of the gene. Further inspection frequently shows that the contigs are broken by very large, heterochromatic introns. Methods of this sort are being expanded to make best use of all available clues to determine which unmapped contigs are associated with genes. These include use of EST libraries, and, in the case of the Y chromosome, testing of male specific genes and reduced shotgun depth of relevant contigs. It appears much more hopeful than anyone would have imagined that whole genome shotgun assemblies can recover the great bulk of even heterochromatic genes.

Animals↗

Lower rate of genomic variation identified in the trans-membrane domain of monoamine sub-class of Human G-Protein Coupled Receptors: the Human GPCR-DB Database.

BACKGROUND: We have surveyed, compiled and annotated nucleotide variations in 338 human 7-transmembrane receptors (G-protein coupled receptors). In a sample of 32 chromosomes from a Nordic population, we attempted to determine the allele frequencies of 80 non-synonymous SNPs, and found 20 novel polymorphic markers. GPCR receptors of physiological and clinical importance were prioritized for statistical analysis. Natural variation and rare mutation information were merged and presented online in the Human GPCR-DB database http://cyrix.cgb.ki.se. RESULTS: The average number of SNPs per 1000 bases of exonic sequence was found to be twice the average number of SNPs per Kilobase of intronic regions (2.2 versus 1.0). Of the 338 genes, 111 were single exon genes, that is, were intronless. The average number of exonic-SNPs per single-exon gene was 3.5 (n = 395) while that for multi-exon genes was 0.8 (n = 1176). The average number of variations within the different protein domain (N-terminus, internal- and external-loops, trans-membrane region, C-terminus) indicates a lower rate of variation in the trans-membrane region of Monoamine GPCRs, as compared to Chemokine- and Peptide-receptor sub-classes of GPCRs. CONCLUSIONS: Single-exon GPCRs on average have approximately three times the number of SNPs as compared to GPCRs with introns. Among various functional classes of GPCRs, Monoamine GPRCs have lower number of natural variations within the trans-membrane domain indicating evolutionary selection against non-synonymous changes within the membrane-localizing domain of this sub-class of GPCRs.

Alleles↗

Enhancer trapping in zebrafish using the Sleeping Beauty transposon.

BACKGROUND: Among functional elements of a metazoan gene, enhancers are particularly difficult to find and annotate. Pioneering experiments in Drosophila have demonstrated the value of enhancer "trapping" using an invertebrate to address this functional genomics problem. RESULTS: We modulated a Sleeping Beauty transposon-based transgenesis cassette to establish an enhancer trapping technique for use in a vertebrate model system, zebrafish Danio rerio. We established 9 lines of zebrafish with distinct tissue- or organ-specific GFP expression patterns from 90 founders that produced GFP-expressing progeny. We have molecularly characterized these lines and show that in each line, a specific GFP expression pattern is due to a single transposition event. Many of the insertions are into introns of zebrafish genes predicted in the current genome assembly. We have identified both previously characterized as well as novel expression patterns from this screen. For example, the ET7 line harbors a transposon insertion near the mkp3 locus and expresses GFP in the midbrain-hindbrain boundary, forebrain and the ventricle, matching a subset of the known FGF8-dependent mkp3 expression domain. The ET2 line, in contrast, expresses GFP specifically in caudal primary motoneurons due to an insertion into the poly(ADP-ribose) glycohydrolase (PARG) locus. This surprising expression pattern was confirmed using in situ hybridization techniques for the endogenous PARG mRNA, indicating the enhancer trap has replicated this unexpected and highly localized PARG expression with good fidelity. Finally, we show that it is possible to excise a Sleeping Beauty transposon from a genomic location in the zebrafish germline. CONCLUSIONS: This genomics tool offers the opportunity for large-scale biological approaches combining both expression and genomic-level sequence analysis using as a template an entire vertebrate genome.

Animals↗

EbEST: an automated tool using expressed sequence tags to delineate gene structure.

Large numbers of expressed sequence tags (ESTs) continue to fill public and private databases with partial cDNA sequences. However, using this huge amount of ESTs to facilitate gene finding in genomic sequence imposes a challenge, especially to wet-lab scientists who often have limited computing resources. In an effort to consolidate the information hidden in the vast number of ESTs into a readable and manageable format, we have developed EbEST-a program that automates the process of using ESTs to help delineate gene structure in long stretches of genomic sequence. The EbEST program consists of three functional modules-the first module separates homologous ESTs into clusters and identifies the most informative ESTs within each cluster; the second module uses the informative ESTs to perform gapped alignment and to predict the exon-intron boundary; and the third module generates text file and graphic outputs that illustrate the orientation, exonic structure, and untranslated regions (UTRs) of putative genes in the genomic sequence being analyzed. Evaluation of EbEST with 176 human genes from the ALLSEQ set indicated that it performed in-line with several existing gene finding programs, but was more tolerant to sequencing errors. Furthermore, when EbEST was challenged with query sequences that harbor more than one gene, it suffered only a slight drop in performance, whereas the performance of the other programs evaluated decreased more. EbEST may be used as a stand-alone tool to annotate human genomic sequences with EST-derived gene elements, or can be used in conjunction with computational gene-recognition programs to increase the accuracy of gene prediction. [EbBEST is available at http://EbEST.ifrc.mcw.edu]

Base Sequence↗

Development and validation of a high-density 'Amahysnp' genotyping array in grain amaranth (Amaranthus hypochondriacus).

BACKGROUND: Grain amaranth has recently gained global attention as a promising crop alternative to traditional cereals due to its nutritional value and adaptability to various growing conditions. Although gene banks conserve extensive collections of amaranth germplasm, the genomic and phenotypic characterization of these resources is limited, which hinders their full utilization in breeding programs. A major challenge is the lack of high-throughput genotyping assays essential for comprehensive genomic characterization and trait mapping. High-density SNP arrays have become standard tools for genome-wide analysis across multiple loci, enabling molecular breeding across a range of crop species. RESULTS: In this study, we developed a 64 K high-throughput SNP genotyping array named "AmahySNP", using Affymetrix® Axiom® technology. The array contains 64,069 high-density SNPs distributed across both genic (55.17%) and non-genic (44.83%) regions of the Amaranthus hypochondriacus genome. The genic region includes 8,879 genes, which consist of 4,830 single-copy genes and 4,049 multi-copy genes distributed across 16 scaffolds. These genes cover various functional regions, including exons (10.5%), introns (40.1%), 5'UTRs (1.6%), and 3'UTRs (2.9%), respectively. The AmahySNP array was effectively utilized for population structure analysis, genetic diversity studies, core development, and genome wide association studies (GWAS) in amaranth germplasm. A representative core set of 112 accessions was identified, which includes two released varieties (Annapurna and Suvarna) and 100 diverse accessions from 12 different regions, representing 12% of the total 917 accessions evaluated. Phylogenetic analysis revealed three major genetic clusters, independent of their geographical origins. GWAS conducted using 22,763 polymorphic SNPs from 540 genotypes identified 13 novel loci associated days to flowering (DTF) trait, seven of which were located within annotated genes. CONCLUSIONS: The AmahySNP 64 K SNP chip a valuable genomic tool for amaranth research and breeding with a strong potential to accelerate its genetic improvement. It enables high-throughput genotyping for a wide range of applications, including GWAS and other genomic studies, and will significantly advance the exploration of natural genetic variations. Ultimately, this resource will empower amaranth breeders to develop improved amaranth cultivars with enhanced crop yield, resilience, and nutritional quality, contributing to global food security and sustainable agriculture.

Amaranthus↗

Gene finding in the chicken genome.

BACKGROUND: Despite the continuous production of genome sequence for a number of organisms, reliable, comprehensive, and cost effective gene prediction remains problematic. This is particularly true for genomes for which there is not a large collection of known gene sequences, such as the recently published chicken genome. We used the chicken sequence to test comparative and homology-based gene-finding methods followed by experimental validation as an effective genome annotation method. RESULTS: We performed experimental evaluation by RT-PCR of three different computational gene finders, Ensembl, SGP2 and TWINSCAN, applied to the chicken genome. A Venn diagram was computed and each component of it was evaluated. The results showed that de novo comparative methods can identify up to about 700 chicken genes with no previous evidence of expression, and can correctly extend about 40% of homology-based predictions at the 5' end. CONCLUSIONS: De novo comparative gene prediction followed by experimental verification is effective at enhancing the annotation of the newly sequenced genomes provided by standard homology-based methods.

Animals↗

Highly specific localization of promoter regions in large genomic sequences by PromoterInspector: a novel context analysis approach.

We present a new algorithm called PromoterInspector to locate eukaryotic polymase II promoter regions in large genomic sequences with a high degree of specificity. PromoterInspector focuses on the genetic context of promoters, rather than their exact location. Application of PromoterInspector can serve as a crucial pre-processing step for other methods to locate exactly, or to analyze promoters. PromoterInspector does not depend on heuristics, because it is purely based on libraries of IUPAC words extracted from training sequences by an unsupervised learning approach. We compared PromoterInspector to in silico promoter prediction tools using the sequences from the review by J.W. Fickett. PromoterInspector compared favourably on Fickett's evaluation scheme. A true positive to false positive ratio of 2.3 was obtained, surpassing the best ratio of 0.6, reported for TSSG. The application of our method to several large genomic sequences of over 1.3 million base-pairs in total resulted in even more specific predictions. The coverage of annotated promoters was comparable to other in silico promoter prediction methods, while the true positive predictions increased by up to 100% of total matches. PromoterInspector scans 100 kb in less than one minute on a workstation, and thus is especially applicable for large genome analysis. The method is available at http://genomatix.gsf. de/cgi-bin/promoterinspector/promoterinspector.pl.

3' Untranslated Regions↗

The ASAP II database: analysis and comparative genomics of alternative splicing in 15 animal species.

We have greatly expanded the Alternative Splicing Annotation Project (ASAP) database: (i) its human alternative splicing data are expanded approximately 3-fold over the previous ASAP database, to nearly 90,000 distinct alternative splicing events; (ii) it now provides genome-wide alternative splicing analyses for 15 vertebrate, insect and other animal species; (iii) it provides comprehensive comparative genomics information for comparing alternative splicing and splice site conservation across 17 aligned genomes, based on UCSC multigenome alignments; (iv) it provides an approximately 2- to 3-fold expansion in detection of tissue-specific alternative splicing events, and of cancer versus normal specific alternative splicing events. We have also constructed a novel database linking orthologous exons and orthologous introns between genomes, based on multigenome alignment of 17 animal species. It can be a valuable resource for studies of gene structure evolution. ASAP II provides a new web interface enabling more detailed exploration of the data, and integrating comparative genomics information with alternative splicing data. We provide a set of tools for advanced data-mining of ASAP II with Pygr (the Python Graph Database Framework for Bioinformatics) including powerful features such as graph query, multigenome alignment query, etc. ASAP II is available at http://www.bioinformatics.ucla.edu/ASAP2.

Alternative Splicing↗

Cloning and functional analysis of cDNAs with open reading frames for 300 previously undefined genes expressed in CD34+ hematopoietic stem/progenitor cells.

Three hundred cDNAs containing putatively entire open reading frames (ORFs) for previously undefined genes were obtained from CD34+ hematopoietic stem/progenitor cells (HSPCs), based on EST cataloging, clone sequencing, in silico cloning, and rapid amplification of cDNA ends (RACE). The cDNA sizes ranged from 360 to 3496 bp and their ORFs coded for peptides of 58-752 amino acids. Public database search indicated that 225 cDNAs exhibited sequence similarities to genes identified across a variety of species. Homology analysis led to the recognition of 50 basic structural motifs/domains among these cDNAs. Genomic exon-intron organization could be established in 243 genes by integration of cDNA data with genome sequence information. Interestingly, a new gene named as HSPC070 on 3p was found to share a sequence of 105bp in 3' UTR with RAF gene in reversed transcription orientation. Chromosomal localizations were obtained using electronic mapping for 192 genes and with radiation hybrid (RH) for 38 genes. Macroarray technique was applied to screen the gene expression patterns in five hematopoietic cell lines (NB4, HL60, U937, K562, and Jurkat) and a number of genes with differential expression were found. The resource work has provided a wide range of information useful not only for expression genomics and annotation of genomic DNA sequence, but also for further research on the function of genes involved in hematopoietic development and differentiation.

Alternative Splicing↗

Comparative sequence analysis of the X-inactivation center region in mouse, human, and bovine.

We have sequenced to high levels of accuracy 714-kb and 233-kb regions of the mouse and bovine X-inactivation centers (Xic), respectively, centered on the Xist gene. This has provided the basis for a fully annotated comparative analysis of the mouse Xic with the 2.3-Mb orthologous region in human and has allowed a three-way species comparison of the core central region, including the Xist gene. These comparisons have revealed conserved genes, both coding and noncoding, conserved CpG islands and, more surprisingly, conserved pseudogenes. The distribution of repeated elements, especially LINE repeats, in the mouse Xic region when compared to the rest of the genome does not support the hypothesis of a role for these repeat elements in the spreading of X inactivation. Interestingly, an asymmetric distribution of LINE elements on the two DNA strands was observed in the three species, not only within introns but also in intergenic regions. This feature is suggestive of important transcriptional activity within these intergenic regions. In silico prediction followed by experimental analysis has allowed four new genes, Cnbp2, Ftx, Jpx, and Ppnx, to be identified and novel, widespread, complex, and apparently noncoding transcriptional activity to be characterized in a region 5' of Xist that was recently shown to attract histone modification early after the onset of X inactivation.

Animals↗