Search PubMedSearch

SEARCH · Search PubMed

Results for “CAA interruption”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

4 recordsLinked to original sources

Dissecting the relationship between haplotypes around ATXN2 CAG repeats and the number of CAA interruptions by long-read sequencing.

BACKGROUND: CAG repeat expansions in ATXN2 are implicated as risk factors for several neurological diseases, including spinocerebellar ataxia type 2 (SCA2) when >=33 CAG repeats are present, and amyotrophic lateral sclerosis (ALS) when 27-33 CAG repeats are present. However, how haplotypes around the repeats and CAA interruptions within the repeats are associated with disease phenotypes remains poorly understood. Previous studies on haplotypes around ATXN2 were limited to SNPs very close to the repeats (<5kb) or were based on statistical inference only. METHODS: Here, we used long-read sequencing on the Oxford Nanopore Technologies (ONT) platform to simultaneously infer haplotypes around ATXN2, the number of CAG repeats, and the number of CAA interruptions, along with NYGC ALS Consortium NGS dataset. We further sequenced 41 individuals (EUR = 39) with neurological diseases with intermediate repeats by ONT. RESULTS: We found that haplotypes around ATXN2 and the number of interruptions show ethnicity-specific and ALS-specific distribution. Three CAA interruptions are present at low prevalence (~1%) in control populations in multiple ancestry groups, but high prevalence (~55%) in ALS individuals with intermediate repeats. Furthermore, we examined 159 individuals with ALS (~90% European ancestry) with intermediate ATXN2 repeats and found a unique haplotype in ALS individuals with three CAA interruptions, which can be tagged by an SNV, rs148019457. We also validated that the rs148019457-G allele is only present in haplotypes with three CAA interruptions. CONCLUSIONS: In summary, our study shows that 3 CAA interruptions are rarely seen in healthy controls but are common in those with expanded ATXN2 CAG repeats who have neurological disorders, and that rs148019457 tags a specific haplotype with 3 CAA interruptions within expanded ATXN2 CAG repeats in individuals of European ancestry. These results have implications for the development of precision genomic medicine for neurological disorders, and the tag SNP may help identify those with interruptions from existing population genotyping data.

ATXN2

Cloning, structural analysis, and expression of the glycogen phosphorylase-2 gene in Dictyostelium.

The glycogen phosphorylase-2 (GP2) activity that appears during the cell differentiation of Dictyostelium was purified to homogeneity. The molecular weight of the nondenatured enzyme was 200,000 as determined by Sephacryl S-300 gel filtration and was 107,000 on sodium dodecyl sulfate-polyacrylamide gel electrophoresis, suggesting that the native enzyme consists of two similar subunits. The intact protein was digested with trypsin and protease V8, and the resulting peptides were purified by microbore high pressure liquid chromatography. The peptides were sequenced, and oligonucleotides were constructed for polymerase chain reaction amplification of the GP2 gene from Dictyostelium genomic DNA template. The resulting polymerase chain reaction products were sequenced directly and were confirmed to encode portions of the GP2 gene. These fragments were used to probe a partial EcoRI genomic library for the remainder of the GP2 gene. The nucleotide sequence of the GP2-selected clones revealed an open reading frame of 2975 base pairs that was interrupted by two introns of 109 and 105 base pairs, respectively. The open reading frame encoded a protein of 992 amino acids with a calculated molecular mass of 112,500 Da and an isoelectric point of 6.4. An unusual sequence within the second exon of GP2, in which the triplet CAA was repeated 11 times, resulted in 11 in-frame glutamine residues of a possible 15 amino acids coded for by this region. The CAA repeat was transcribed, as shown by the sequence of cDNA. Comparison of the amino acid sequence of Dictyostelium GP2 to the phosphorylases from other organisms revealed that the Dictyostelium protein was 50 and 44% identical to yeast and rabbit muscle phosphorylases, respectively. Northern blot analysis showed that GP2 mRNA was absent in amebas and the early stages of development, reached a maximum level of expression at the slug stage, and then decreased in the terminal stages of development. Comparison of the mRNA expression with the appearance of GP2 enzyme protein and enzyme activity revealed that gp2 mRNA and a 113-kDa GP2 enzyme peptide were expressed concurrently at 10 h of development. However, enzyme activity did not appear until 18 h, coincident with a decrease in the level of the 113-kDa peptide and a corresponding increase in the amount of a 106-kDa GP2 peptide. Addition of cAMP to aggregation-competent cells in liquid culture resulted in the induction of GP2 mRNA, GP2 protein, and GP2 enzyme activity.

Amino Acid Sequence

The structure and organization of a proline-rich protein gene of a mouse multigene family.

One gene of the mouse proline-rich protein multigene family was cloned on a 3.6-kilobase pair EcoRI/BglII DNA fragment from a (partial) Sau3A bacteriophage library of CD-1 mouse chromosomal DNA. Phage harboring the gene were identified by plaque hybridization using 32P-labeled proline-rich protein cDNA inserts from clones pRP33 and pMP1 obtained from rat and mouse, respectively. The transcriptional unit includes three exonic sequences separated by 1434 base pairs (intron I) and 450 base pairs (intron II). The complete primary structure of the gene and the 5' and 3' flanking regions (3595 base pairs) were determined by the Maxam and Gilbert (Maxam, A.M., and Gilbert, W. (1980) Methods Enzymol. 65, 499-560) sequencing method. The DNA on the 5' side of exon I contains several sequences that may be involved in the induction and expression of this mouse gene. These sequences include putative regulatory sites such as those considered to be inducible by cAMP and steroids, Z-DNA and enhancer sequences and the expected TATAA and CAAT boxes. The mature protein coding region, exon II, is not interrupted with intron sequences. Exon III is located in the nontranslated region and contains the poly(A) addition site. The deduced amino acid sequence showed that the protein encoded by this gene contains 13 tandemly repeat regions, each 14 amino acids in length, with the prototype sequence PPPPGGPQPRPPQG. Each amino acid within the repeat has a favored codon. The consensus DNA sequence for each repeat is CCA CCA CCA CCA GGA GGC CCA CAG CCG AGA CCC CCT CAA GGC. The high degree of conservation of both nucleotide and amino acid sequences within the repeat region suggests that proline-rich protein genes likely evolved by gene duplication of a 42-base pair internal repeat.

Amino Acid Sequence

Alpha subunit of mitochondrial F1-ATPase from the fission yeast. Deduced sequence of the wild type and identification of a mutation that alters apparent negative cooperativity.

The nuclear gene atp1 encoding the mitochondrial ATP synthase alpha subunit of the fission yeast Schizosaccharomyces pombe was sequenced. It contains a 1,608-base pair-long open reading frame interrupted by two introns of 175 and 269 base pairs, located near the 5'-end of the gene. The initiation site of transcription AAAC was located 60 nucleotides upstream of the translation initiation codon. The deduced polypeptide sequence contains a 27-amino acid residue presequence, presumably involved in mitochondrial targeting, preceding a mature protein of 509 amino acid residues. The atp1 alleles from mutant A2313 (Bouty, M., and Goffeau, A. (1982) Eur. J. Biochem. 125, 471-477) and its related phenotypic revertant R351 (Falson, P., Di Pietro, A., Darbouret, D., Jault, J. M., Gautheron, D. C., Boutry, M., and Goffeau, A. (1987) Biochem. Biophys. Res. Commun. 148, 1182-1188) were also cloned and sequenced. A single nonsense mutation CAA-TAA (Gln173-stop) in mutant A2313 became a missense mutation TAA-TTA (stop-Leucine) in revertant R351. Glutamine 173 is located in the first putative element of the nucleotide binding site. Its substitution by a leucine residue appears responsible for the lower enzyme affinity toward ADP and for the loss of cooperativity of F1-ATPase activity.

Amino Acid Sequence