Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Alu elements”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Structure and genetics of the partially duplicated gene RP located immediately upstream of the complement C4A and the C4B genes in the HLA class III region. Molecular cloning, exon-intron structure, composite retroposon, and breakpoint of gene duplication.

The correlation of many HLA-associated autoimmune and genetic diseases with the polymorphic complement C4 genes may be attributed to the presence of disease susceptibility genes in the close proximity of C4. We have cloned and characterized a pair of partially duplicated genes, RP1 and RP2, located 611 base pairs upstream of the human C4A and C4B genes, respectively. The putative RP protein, consisting of 364 amino acid residues, is basic and highly hydrophilic. There is a bipartite nuclear localization signal at residues 114-131 and therefore RP may be a nuclear protein. Northern blot analysis suggested that RP is ubiquitously expressed. The 5' region of the RP1 gene is CpG rich, which is a characteristic of housekeeping genes. The RP1 gene contains nine exons. Located in the fourth intron is a cluster of Alu elements, and a newly defined composite retroposon SVA with a SINE, multiple copies of GC-rich VNTRs and an Alu element altogether enclosed by direct terminal repeats. Members of SVA are also present in the complement C2 gene located about 20 kilobases upstream of RP1 in the HLA and in the cytochrome CYP1A1 gene. Determination of the DNA sequences for RP2 from two different HLA haplotypes revealed identical hybrid sequences which resulted from fusion of RP with the tenascin-like Gene X and truncation of the 5' regions of both genes. Cumulative data suggest that the four tandemly arranged genes RP, complement C4, steroid 21-hydroxylase (CYP21), and Gene X altogether form a modular structure, RCCX. The number of RCCX modules varies from one to three or more in the population. Absence of the truncated genes RP2 and Gene XA have been detected in genomes with single RCCX modules. Duplication of the RCCX modules probably occurred before the speciation of great apes and humans as they contain the same breakpoint region of RP and Gene X gene duplication.

Alleles↗

Fast identification of repetitive elements in biological sequences.

We have developed a fast filtering method for searching repetitive sequences in databases that allows the simultaneous identification of different families of repetitive elements during the same scanning. It discriminates between repetitive elements and non-related sequences by comparing the frequencies of k-words found in both groups of sequences. The distance used to sort out the sequences is based on a weighting of the k-words, which is obtained by performing a correspondence analysis on learning sets of correctly chosen sequences. The identification of Alu elements in human sequences is given as an illustration of the method. The Alu sequences are divided in four distinct groups of elements: the left and right monomers located on the direct and on the complementary strands. The results obtained on the test sets show that a very good discrimination is achieved with a word length of 6 b.p. Indeed, only 0.5% of the non-Alu sequences were incorrectly predicted as Alu elements for a threshold value allowing the identification of all Alu monomers. The misclassification of the different Alu monomers (1.4%) in the four groups of examples occurs only when the left and the right monomers are in the same orientation. Moreover, during the scanning of 63 GenBank sequences longer than 10 Kb, all the Alu elements were correctly identified (616 elements) and only a few non-Alu sequences were wrongly predicted as Alu elements (22 fragments). There is a real need for this kind of method since most of the repetitive elements are not annotated in the database entries. This method can then be used for a systematic screening of new sequences before their insertion in databases. It can also allow the creation of specific databases devoted to repetitive elements, which is a required step for any further analysis of those elements.

Algorithms↗

cDNAs derived from primary and small cytoplasmic Alu (scAlu) transcripts.

We have isolated and sequenced twenty-six cDNAs derived from primary Alu transcripts. Most cDNAs (22/26) sequenced end in multiple T residues, known to be at the termination for RNA polymerase III-directed transcripts. We conclude that these cDNAs were derived from authentic, RNA polymerase III-directed primary Alu transcripts. Sequence alignment of the cDNAs with Alu consensus sequences show that the cDNAs belong to different, previously described Alu subfamilies. The sequence variation observed in the 3' non-Alu regions of each of the cDNAs led us to conclude that they were derived from different genomic loci, thus demonstrating that multiple Alu loci are transcriptionally active. The subfamily distribution of the cDNAs suggests that transcriptional activity is biased towards evolutionarily younger Alu subfamilies, with a strong selection for the consensus sequence in the first 42 bases and the promoter B box. Sequence data from seven cDNAs derived from small cytoplasmic Alu (scAlu) transcripts, a processed form of Alu transcripts, also have a similar bias towards younger Alu subfamilies. About half of these cDNAs are due to processing or degradation, but the other half appear to be due to the formation of a cryptic RNA polymerase III termination signal in multiple loci. Using our sequence data, we have isolated a transcriptionally active genomic Alu element belonging to the Ya5 subfamily. In vitro transcription studies of this element suggest that its flanking sequences contribute to its transcriptional activity. The role of flanking sequences and other factors involved in transcriptional activity of Alu elements are discussed.

Base Sequence↗

A genetic variant of ACE increases cell survival: a new paradigm for biology and disease.

The human angiotensin converting enzyme (ACE) polymorphism is caused by an Alu element insertion resulting in three genotypes (Alu+/+, Alu+/-, Alu-/-, or ACE-II, ACE-ID, and ACE-DD, respectively), with ACE-II displaying lower ACE activity. The polymorphism is associated with athletic performance, aging, and disease. Population studies, however, were confounding because variants of the polymorphism appeared to fortuitously correlate with health and various pathological states. To clarify the functional role of the polymorphism, we studied its direct effect on cell survival. ACE-II (Alu+/+) human endothelial cells (EC) had lower angiotensin-II levels and 20-fold increased viability after slow starvation as compared to ACE-DD cells (Alu-/-). By RT-PCR, only ACE-II cells expressed the pluripotent/stem cell-maintenance factors nanog, numb, and klotho. ACE inhibition by captopril in ACE-DD cells mimicked the ACE-II genotype. These results provide the first evidence of a functional role for a naturally occurring polymorphism, having broad implications for human biology, longevity, and disease.

Alu Elements↗

Splice-mediated insertion of an Alu sequence in the COL4A3 mRNA causing autosomal recessive Alport syndrome.

Alport syndrome is a mainly X-linked hereditary disease of basement membranes characterized by progressive renal failure, deafness, and ocular lesions. The alpha 3(IV) and alpha 4(IV) collagen genes have been recently shown to be involved in the less frequent autosomal recessive form. When screening lymphocyte COL4A3 mRNAs from Alport patients, we found a mutant whose transcripts were disrupted by a 74 bp insertion at the junction of exons IV or V and VI. The insertion derives from an antisense Alu element in COL4A3 intron V, which has been spliced into the alpha 3(IV) mRNA due to a G to T transversion activating a cryptic acceptor splice site in this Alu element. There is complete segregation of this mutation with the disease in the family. Our findings provide the first evidence for the pathogenic role of abnormal splicing of COL4A3. Moreover, we demonstrate the superiority of mutation screening at the mRNA level to detect a hitherto poorly recognized mutation mechanism in humans, splice-mediated insertion of an Alu fragment into a coding sequence.

Antisense Elements (Genetics)↗

Dinucleosome DNA of human K562 cells: experimental and computational characterizations.

Dinucleosome formation is the first step in the organization of the higher order chromatin structure. With the ultimate aim of elucidating the dinucleosome structure, we constructed a library of human dinucleosome DNA. The library consists of PCR-amplifiable DNA fragments obtained by treatment of nuclei of erythroid K562 cells with micrococcal nuclease followed by extraction of DNA and adaptor ligation to the blunt-ended DNA fragments. The library was then cloned using a plasmid vector and the sequences of the clones were determined. The dominating clones containing the Alu elements were removed. A total of 1002 clones, which comprised a dinucleosome database, contained 84 and 918 clones from the clones before and after removing Alu elements, respectively. Approximately 70% of the clones were between 300 and 400 bp in size and they were distributed to various locations of all chromosomes except the Y chromosome. The clones containing A(2)N(8)A(2)N(8)A(2) or T(2)N(8)T(2)N(8)T(2) sequences were classified into three types, Type I (N shape), Type II (V shape) and Type III (M shape) according to DNA curvature plots. The locations of experimentally determined curved DNA segments matched well with the calculated ones though the clones of Types I and III showed additional curved DNA segments as revealed by the curvature plots. The distributions of complementary dinucleotides in the nucleosome DNA, at the ends of the dinucleosome DNA clones, allowed us to predict the positions of the nucleosome dyad axis, and estimate the size of the nucleosome core DNA, 125nt. The distributions of AA and TT dinucleotides, as well as other RR and YY dinucleotides, showed a periodicity with an average period of 10.4 bases, close to the values observed before. Mapping of nucleosome positions in the dinucleosome database based on the observed periodicity revealed that the nucleosomes were separated by a linker of 7.5+ approximately 10 x n nt. This indicates that the nucleosome-nucleosome orientations are, typically, halfway between parallel and antiparallel. Also an important finding is that the distributions of AA/TT and other RR/YY dinucleotides, apparently, reflect both DNA curvature and DNA bendability, cooperatively contributing to the nucleosome formation.

Base Sequence↗

One short well conserved region of Alu-sequences is involved in human gene rearrangements and has homology with prokaryotic chi.

Alu elements have repeatedly been found involved in gene rearrangements in humans. Although these elements have been suggested to stimulate gene rearrangements, sparse information is available for the possible mechanism(s) of these events. Here we present a compilation of Alu elements that have been involved in recombinational events leading to gene rearrangements, indicating the presence of a common 26 bp core sequence at or close to the sites of recombination. Besides the obvious possibility of retrotransposition, gene rearrangements may be induced by sequences that stimulate genetic recombination. We suggest that the core sequence stimulates recombination and may thereby cause the frequent involvement of these elements in gene rearrangements. Curiously, the core sequence contains the pentanucleotide motif CCAGC, which is also part of chi, an 8 bp sequence known to stimulate recBC mediated recombination in Escherichia coli.

Base Sequence↗

Simplified plasmid rescue of host sequences adjacent to integrated proviruses.

We have previously described a Moloney murine leukemia retroviral (MoMLV) vector useful for the generation of anchored long-range maps of complex mammalian genomes. We now report the development of a modified vector carrying the ColE1 origin of replication and the chloramphenicol-resistance (CmR) gene to facilitate the recovery of genomic sequences adjacent to integrated proviruses. We demonstrate the utility of this new vector for the rescue in plasmid form of a 6-kb human fragment containing portions of an Alu element adjacent to the proviral 3'-LTR from an infected human primary fibroblast clone. We generated a sequence-tagged site (STS) from the derived sequence outside the Alu element, and used a somatic cell hybrid mapping panel to assign this STS to human chromosome 10 by a polymerase chain reaction-based assay. We suggest that this new CmR vector will be useful for recovering sequences of genes interrupted by proviral insertion, to generate directional maps with respect to the inserted provirus using the single cleavage sites within the vector, and to generate site-specific STS for defining and mapping the integration site.

Animals↗

Structure, polymorphism, and novel repeated DNA elements revealed by a complete sequence of the human alpha-fetoprotein gene.

The human alpha-fetoprotein gene spans 19,489 base pairs from the putative "Cap" site to the polyadenylation site. It is composed of 15 exons separated by 14 introns, which are symmetrically placed within the three domains of alpha-fetoprotein. In the 5' region, a putative TATAAA box is at position -21, and a variant sequence, CCAAC, of the common CAT box is at -65. Enhancer core sequences GTGGTTTAAAG are found in introns 3 and 4, and several copies of glucocorticoid response sequences AGATACAGTA are found on the template strand of the gene. There are six polymorphic sites within 4690 base pairs of contiguous DNA derived from two allelic alpha-fetoprotein genes. This amounts to a measured polymorphic frequency of 0.13%, or 6.4 X 10(-4)/site, which is about 5-10 times lower than values estimated from studies on polymorphic restriction sites in other regions of the human genome. There are four types of repetitive sequence elements in the introns and flanking regions of the human alpha-fetoprotein gene. At least one of these is apparently a novel structure (designated Xba) and is found as a pair of direct repeats, with one copy in intron 7 and the other in intron 8. It is conceivable that within the last 2 million years the copy in intron 8 gave rise to the repeat in intron 7. Their present location on both sides of exon 8 gives these sequences a potential for disrupting the functional integrity of the gene in the event of an unequal crossover between them. There are three Alu elements, one of which is in intron 4; the others are located in the 3' flanking region. A solitary Kpn repeat is found in intron 3. The Xba and Kpn repeats were only detected by complete sequencing of the introns. Neither X, Xba, nor Kpn elements are present in the related human albumin gene, whereas Alu's are present in different positions. From phylogenetic evidence, it appears that Alu elements were inserted into the alpha-fetoprotein gene at some time postdating the mammalian radiation 85 million years ago.

Base Sequence↗

Large-scale cloning of human chromosome 2-specific yeast artificial chromosomes (YACs) using an interspersed repetitive sequences (IRS)-PCR approach.

We report here an efficient approach to the establishment of extended YAC contigs on human chromosome 2 by using an interspersed repetitive sequences (IRS)-PCR-based screening strategy for YAC DNA pools. Genomic DNA was extracted from 1152 YAC pools comprised of 55,296 YACs mostly derived from the CEPH Mark I library. Alu-element-mediated PCR was performed for each pool, and amplification products were spotted on hybridization membranes (IRS filters). IRS probes for the screening of the IRS filters were obtained by Alu-element-mediated PCR. Of 708 distinct probes obtained from chromosome 2-specific somatic cell hybrids, 85% were successfully used for library screening. Similarly, 80% of 80 YAC walking probes were successfully used for library screening. Each probe detected an average of 6.6 YACs, which is in good agreement with the 7- to 7.5-fold genome coverage provided by the library. In a preliminary analysis, we have identified 188 YAC groups that are the basis for building contigs for chromosome 2. The coverage of the telomeric half of chromosome 2q was considered to be good since 31 of 34 microsatellites and 22 of 23 expressed sequence tags that were chosen from chromosome region 2q13-q37 were contained in a chromosome 2 YAC sublibrary generated by our experiments. We have identified a minimum of 1610 distinct chromosome 2-specific YACs, which will be a valuable asset for the physical mapping of the second largest human chromosome.

Animals↗

Jak3 expression and genomic sequence in pediatric acute lymphoblastic leukemia.

Janus tyrosine kinase 3 (JAK3) is one of several key regulatory enzymes in B-cell precursors which is highly conserved between multiple species. The gene for Jak3 has been mapped to human chromosome 19p12-13.1 and encompasses 23 exons. Constitutively high levels of JAK3 activity may contribute to drug resistance and enhanced clonogenicity of leukemic B-cell precursors from children and infants with acute lymphoblastic leukemia (ALL). As part of a systematic effort to accurately determine the genomic sequence of Jak3 gene in normal and leukemic B-cell precursors, we sequenced a relatively short region of Jak3 spanning two introns, originally termed introns 10 and 11. This genomic sequence appeared in certain RT-PCR products from our analysis of Jak3 gene expression in pediatric, as well as infant, primary ALL cells. Unexpectedly, a gap in the original Jak3 genomic sequence was found in intron 10 across the sequence matching to an Alu element. Furthermore, the sequence obtained from intron 11 did not match at all to that previously reported, and the length of the intron was much larger than expected at 1.1 kb. Homology to Alu elements (three regions, 699 bp total) and a LINE2 element (one region, 189 bp total) were seen across the entire region covering exons 10-12 (2.1 kb total). Two potential single nucleotide polymorphisms (SNPs) were observed in intron 11. No apparent genomic mutation was found across this region in leukemic B-cell precursors from any of the ALL patients examined. This newly described sequence corrects the previous published genomic sequence from this region rather than identifying an insertion or translocation specific to these ALL cases. Our results significantly extend previous efforts to determine the genomic sequence of Jak3 and analyze its expression in childhood pro-B ALL and other forms of ALL.

B-Lymphocytes↗

Isolation and characterization of the human tyrosine aminotransferase gene.

Structure and sequence of the human gene for tyrosine aminotransferase (TAT) was determined by analysis of cDNA and genomic clones. The gene extends over 10.9 kbl and consists of 12 exons giving rise to a 2,754 nucleotide long mRNA (excluding the poly(A)tail). The human TAT gene is predicted to code for a 454 amino acid protein of molecular weight 50,399 dalton. The overall sequence identity within the coding region of the human and the previously characterized rat TAT genes is 87% at the nucleotide and 92% at the protein level. A minor human TAT mRNA results from the use of an alternative polyadenylation signal in the 3' exon which is present but not used at the corresponding position in the rat TAT gene. The non-coding region of the 3' exon contains a complete Alu element which is absent in the rat TAT gene but present in apes and old world monkeys. Two functional glucocorticoid response elements (GREs) reside 2.5 kb upstream of the rat TAT gene. The DNA sequence of the corresponding region of the human TAT gene shows the distal GRE mutated and the proximal GRE replaced by Alu elements.

Amino Acid Sequence↗

A human tRNA gene heterocluster encoding threonine, proline and valine tRNAs.

A cluster of three tRNA genes encoding a tRNA(UGUThr), a tRNA(UGGPro), and a tRNA(AACVal), and two Alu-elements occur in a 6.0-kb human DNA fragment. The tRNA(Thr) gene is 2.7-kb upstream from the tRNA(Pro) gene, which is separated by 367 bp from the tRNA(Val) gene. One Alu-element actually overlaps the tRNA(Val) gene and is of opposite polarity to all three tRNA genes. All three tRNA genes are accurately transcribed in a homologous HeLa cell extract, since the ribonuclease T1 fingerprints of the tRNA transcripts are consistent with the nucleotide sequences of the tRNAs. The upstream region flanking the tRNA(Thr) gene has two tracts of alternating purine/pyrimidine residues potentially capable of adopting the Z-DNA conformation, and presumptive binding sites for two RNA polymerase II transcription factors. The tRNA(Thr) gene apparently has a substantially higher in vitro transcriptional efficiency than the other two tRNA genes in this cluster, and a tRNA(GCCGly) gene from another human DNA segment. Deletion constructs of the tRNA(Thr) gene retaining 272, 168, and 33 bp of original 5'-flanking DNA had about the same in vitro transcriptional efficiency, whereas that of the construct with only 2 bp of 5'-flanking human DNA was drastically reduced. The tRNA(Thr) gene constructs with 272 and 168 bp of original 5'-flanking DNA apparently reduce the transcriptional efficiencies of the proline and glycine tRNA genes, implicating the upstream region from the tRNA(Thr) gene as being crucial for its high transcriptional efficiency.

Base Sequence↗

Mutation analysis in the BRCA2 gene in primary breast cancers.

Breast cancer, one of the most common and deleterious of all diseases affecting women, occurs in hereditary and sporadic forms. Hereditary breast cancers are genetically heterogeneous; susceptibility is variously attributable to germline mutations in the BRCA1 (ref. 1), BRCA2 (ref. 2), TP53 (ref. 3) or ataxia telangiectasia (ATM) genes, each of which is considered to be a tumour suppressor. Recently a number of germline mutations in the BRCA2 gene have been identified in families prone to breast cancer. We screened 100 primary breast cancers from Japanese patients for BRCA2 mutations, using PCR-SSCP. We found two germline mutations and one somatic mutation in our patient group. One of the germline mutations was an insertion of an Alu element into exon 22, which resulted in alternative splicing that skipped exon 22. The presence of a 64-bp polyadenylate tract and evidence for an 8-bp target-site duplication of the inserted DNA implied that the retrotransposal insertion of a transcriptionally active Alu element caused this event. Our results indicate that somatic BRCA2 mutations, like somatic mutations in the BRCA1 gene, are very rare in primary breast cancers.

BRCA2 Protein↗

Sequence correlation between neighboring Alu instances suggests post-retrotransposition sequence exchange due to Alu gene conversion.

Alu elements constitute 10% of the human genome. Alu mobilization is important in the evolution of the human genome. While retrotransposition is the primary pathway of Alu mobilization, Alu gene conversion has been postulated as a secondary pathway for Alu mobilization in human genome. However, the mode and tempo of Alu gene conversions remain a mystery due to lack of sensitive statistical methods. In this paper, we present the first study on sequence correlation between Alu instances, measured by the number of shared mutations away from the Alu consensus, or co-mutations. Our analysis reveals a significantly elevated co-mutation rate between Alu instances that are located in close proximity along a chromosome. This effect is more pronounced outside Alu subfamily diagnostic positions. This effect peaks among immediately adjacent Alu instances, diminishes quickly in increasing distances between Alu instances, and vanishes beyond 5000 bp. Our results suggest that this effect reflects post-retrotransposition sequence exchanges between Alu instances, mainly due to Alu gene conversions.

Alu Elements↗

Long-read sequencing reveals a hidden Alu-mediated splice defect in CPLANE1, causing orofaciodigital syndrome type VI.

Orofaciodigital syndrome type VI (OFD VI) is a recessive ciliopathy characterized by excessive polydactyly, molar tooth sign, cleft lip, and developmental delay, caused by pathogenic variants in CPLANE1. Here, we present a patient with OFD VI that remained genetically unexplained after routine genetic testing, including short-read whole genome sequencing (WGS). Using long-read sequencing, we found two biallelic splice-site variants in CPLANE1, c.8633-4_8633-3del, and an Alu element insertion close to an exon-intron boundary. Transcript analysis showed that each variant independently resulted in exon skipping, and quantitative expression studies revealed reduced total CPLANE1 mRNA levels in patient-derived fibroblasts. Based on these findings, we were able to re-classify the c.8633-4_8633-3del variant from a variant of uncertain significance (VUS) to likely pathogenic. The identification of an Alu element insertion missed by short-read WGS highlights the added diagnostic value of long-read sequencing in uncovering cryptic, transposable element-associated pathogenic variants.

Journal Article↗

Human retroelements may introduce intragenic polyadenylation signals.

In the human genome, the insertion of LINE-1 and Alu elements can affect genes by sequence disruption, and by the introduction of elements that modulate the gene's expression. One of the modulating sequences retroelements may contribute is the canonical polyadenylation signal (pA), AATAAA. L1 elements include these within their own sequence and AATAAA sequences are commonly created in the A-rich tails of both SINEs and LINEs. Computational analysis of 34 genes randomly retrieved from the human genome draft sequence reveals an orientation bias, reflected as a lower number of L1s and Alus containing the pA in the same orientation as the gene. Experimental studies of Alu-based pA sequences when placed in pol II or pol III transcripts suggest that the signal is very weak, or often not used at all. Because the pA signal is highly affected by the surrounding sequence, it is likely that the Alu constructs evaluated did not provide the required recognition signals to the polyadenylation machinery. Although the effect of pA signals contributed by Alus is individually weak, the observed reduction of "sense" oriented pA-containing L1 and Alu elements within genes reflects that even a modest influence causes a change in evolutionary pressure, sufficient to create the biased distribution.

Base Sequence↗

HLA-DRB intron 1 sequences: implications for the evolution of HLA-DRB genes and haplotypes.

Human DRB genes encode beta chains of the major histocompatibility complex (MHC) class II molecules. Although nine DRB loci have been mapped to the short arm of chromosome 6, an individual chromosome contains only one to five loci and is classified into one of five major haplotypes. To elucidate the origin of human DRB loci and haplotypes, intron 1 sequences approximately 5000 bp in length were determined for three DRB1 alleles (DRB1*03, DRB1*04, and DRB1*15) and five DRB genes (DRB2, DRB3, DRB4, DRB5, and DRB7). The sequences were subjected to phylogenetic analyses together with previously determined intron 4 and 5 sequences. The sequences provided two sources of information: Nucleotide substitutions that could be used to construct phylogenetic trees and to estimate divergence times and a set of insertions (mostly Alu elements) that reveal the order of splitting of duplicated genes. The combined data indicate that the ancestor of the human DRB genes was HLA-DRB1*04-like and that the DRB2, DRB7, DRB5, and DRB3 genes arose from this ancestor by four rounds of duplication 58, 56, 53, and 36 million years (MY) ago, respectively. The DRB4 gene may have arisen 46 MY ago by a deletion from the DRB1 and DRB2 genes and the DRB6 gene is probably an allele at the DRB2 locus. During the course of its evolution, the DRB1*04 gene acquired an intron 1 segment (including two Alu elements) from a gene that became the ancestor of DRB1*03. The present-day HLA-DR haplotypes were derived from three principal ancestral haplotypes: DRB1-DRB2, DRB1-DRB5, and DRB1-DRB7.

Base Sequence↗