Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “pseudogene”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Distinct BRCA1 rearrangements involving the BRCA1 pseudogene suggest the existence of a recombination hot spot.

The 5' end of the breast and ovarian cancer-susceptibility gene BRCA1 has previously been shown to lie within a duplicated region of chromosome band 17q21. The duplicated region contains BRCA1 exons 1A, 1B, and 2 and their surrounding introns; as a result, a BRCA1 pseudogene (PsiBRCA1) lies upstream of BRCA1. However, the sequence of this segment remained essentially unknown. We needed this information to investigate at the nucleotide level the germline deletions comprising BRCA1 exons 1A, 1B, and 2, which we had previously identified in two families with breast and ovarian cancer. We have analyzed the recently deposited nucleotide sequence of the 1.0-Mb region upstream of BRCA1. We found that 14 blocks of homology between the tandemly repeated copies (cumulative length = 11.5 kb) show similarity of 77%-92%. Gaps between blocks result from insertion or deletion, usually of repetitive elements. BRCA1 exon 1A and PsiBRCA1 exon 1A are 44.5 kb apart. In the two families with breast and ovarian cancer mentioned above, distinct homologous recombination events occurred between intron 2 of BRCA1 and intron 2 of PsiBRCA1, leading to 37-kb deletions. Breakpoint junctions were found to be located at close but distinct sites within segments that are 98% identical. The mutant alleles lack the BRCA1 promoter and harbor a chimeric gene consisting of PsiBRCA1 exons 1A, 1B, and 2, which lacks the initiation codon, fused to BRCA1 exons 3-24. Thus, we report a new mutational mechanism for the BRCA1 gene. The presence of a large region homologous to BRCA1 on the same chromosome appears to constitute a hot spot for recombination.

Alleles↗

SAG/ROC2/Rbx2/Hrt2, a component of SCF E3 ubiquitin ligase: genomic structure, a splicing variant, and two family pseudogenes.

We have recently cloned and characterized an evolutionarily conserved gene, Sensitive to Apoptosis Gene (SAG), which encodes a redox-sensitive antioxidant protein that protects cells from apoptosis induced by redox agents. The SAG protein was later found to be the second family member of ROC/Rbx/Hrt, a component of the Skp1-cullin-F box protein (SCF) E3 ubiquitin ligase, being required for yeast growth and capable of promoting cell growth during serum starvation. Here, we report the genomic structure of the SAG gene that consists of four exons and three introns. We also report the characterization of a SAG splicing variant (SAG-v), that contains an additional exon (exon 2; 264 bp) not present in wildtype SAG. The inclusion of exon 2 disrupts the SAG ORF and gives rise to a protein of 108 amino acids that contains the first 59 amino acids identical to SAG and a 49-amino acid novel sequence at the C terminus. The entire RING-finger domain of SAG was not translated because of several inframe stop codons within the exon 2. The SAG-v protein was expressed in multiple human tissues as well as cell lines, but at a much lower level than wildtype SAG. Unlike SAG, SAG-v was not able to rescue yeast cells from lethality in a ySAG knockout, nor did it bind to cullin-1 or have ligase activity, probably because of the lack of the RING-finger domain. Finally, we report the identification of two SAG family pseudogenes, SAGP1 and SAGP2, that share 36% or 47% sequence identity with ROC1/Rbx1/Hrt1 and 30% or 88% with SAG, respectively. Both genes are intronless with two inframe stop codons.

Alternative Splicing↗

Molecular analysis of bovine actin gene and pseudogene sequences: expression of nonmuscle and striated muscle isoforms in adult tissues.

Most studies on the tissue distribution of actin isoform transcripts have been done in small mammals such as rat and mouse. We have begun a characterization of the actin gene family in a large mammal, the bovine. The alpha skeletal gene was isolated, and an isoform-specific probe to the 3' untranslated region of the transcript identified. This probe, in combination with isoform specific probes for alpha cardiac, beta nonmuscle, and gamma nonmuscle actins, was used to examine expression of nonmuscle and striated muscle actin gene transcription in different tissues. In contrast to other species so far examined, striated muscle isoforms were more strictly tissue specific, with virtually no alpha cardiac isoform transcripts detected in skeletal muscle and almost no alpha skeletal transcripts in cardiac tissue. The distribution of the beta and gamma nonmuscle actins was also unique in bovine compared to other species. A partial beta-actin pseudogene, and the chromosomal DNA flanking one end of it, were also cloned and sequenced. This chromosomal site was found to be homologous to a viral integration site previously identified in simian virus 40 (SV40)-transformed rat cells, suggesting that this region of the chromosome may be a preferred target for insertion events.

Actins↗

Expression of a pseudogene for interferon-alpha L.

The human leukocyte interferon-alpha L (IFN-alpha L) appears to be a pseudogene because it contains a termination codon in the DNA coding for the precursor peptide (signal peptide), but the remainder of the gene codes for an interferon protein of normal length. To determine if this gene codes for an active interferon molecule, an expression plasmid was constructed for IFN-alpha L that was produced in Escherichia coli. The IFN-alpha L is active on human and bovine cells, and exhibits a trace of activity on mouse cells.

Cloning, Molecular↗

Identification of a functional allele of a human interferon-alpha gene previously characterized as a pseudogene.

Three recombinant phage lambda L47 clones containing 4 alpha interferon (IFN) genes have been isolated from a newly constructed human genomic library. Each gene is an allele of a previously described IFN gene, three being only minor variants. The fourth gene SMTIII.1A is a functional allele of the psi LeIF-L gene which previously has been described only as a pseudogene. Therefore, it appears likely that other variant alleles may remain to be described and that the IFN system may be able to tolerate some degeneracy as a consequence of the large number of members of the family.

Alleles↗

Characterization of the first nonmammalian T2 cytokine gene cluster: the cluster contains functional single-copy genes for IL-3, IL-4, IL-13, and GM-CSF, a gene for IL-5 that appears to be a pseudogene, and a gene encoding another cytokinelike transcript, KK34.

A genomics approach based on the conservation of synteny was used to develop a bacterial artificial chromosome (BAC) contig across the chicken T2 cytokine gene cluster. Sequencing of representative BACs showed that the chicken genome encodes genes for the homologs of mammalian interleukin-3 (IL-3), IL-4, IL-5, IL-13, and granulocyte-macrophage colony-stimulating factor (GM-CSF). These sequences represent the first T2 cytokines found outside of mammals, and their location demonstrates that the T2 cluster is ancient (at least 300 million years old). Four of these genes (IL-3, IL-4, IL-13, and GM-CSF) are expressed at the mRNA level and can be expressed as recombinant protein. In contrast to the other four genes, the chicken IL-5 (ChIL-5) gene we sequenced lacks a recognizable promoter and regulatory sequences in the predicted 3'-untranslated region (3'-UTR). Further, there is no evidence for its expression at the mRNA level. We, therefore, hypothesize that it is a pseudogene. Genomic analysis revealed that a recently characterized cytokinelike transcript, KK34, not identified in our initial analysis of the BAC sequence, is also encoded in this cluster. This gene may represent a duplication of an ancestral IL-5 gene and may encode the functional homolog of IL-5 in the chicken.

Amino Acid Sequence↗

Ethnic differences in poly(ADP-ribose) polymerase pseudogene genotype distribution and association with lung cancer risk.

Poly(ADP-ribose) polymerase (PADPRP) is a nuclear DNA-binding enzyme that can modulate chromatin structure close to DNA replication, recombination and repair regions. Two-allele polymorphism on the PADPRP chromosome 13 pseudogene has been studied in several ethnic subpopulations, and the association of each allele with different types of cancer has been investigated. To study the frequency of the allele in the context of lung cancer, we performed a PCR assay for the PADPRP polymorphism in 288 lung cancer patients and 292 matched controls and examined the frequency of the alleles in different ethnic groups. Our results showed that the allele distribution was significantly different among members of different ethnic groups. Specifically, the A allele was dominant in Mexican-American and Caucasian groups but not in the African-American group. The frequencies of the B allele in Mexican-American, Caucasian and African-American controls were 0.184, 0.218 and 0.606, respectively, with the Caucasian cases and controls showing an almost identical lower B allele frequency (0.199 in cases versus 0.218 in controls), and the African-American cases and controls showing an almost identical but considerably higher frequency (0.578 in cases versus 0. 606 in controls). In contrast, the Mexican-American cases and controls exhibited a considerable difference in the B allele frequency (0.306 in cases versus 0.184 in controls). When we combined subjects with the AB or BB genotype into a susceptible genotype group and compared them with the AA group using univariate analysis, the susceptible genotype was not shown to be associated with a risk of lung cancer in either the Caucasian or African-American subpopulation but was significantly associated with an increased risk (2.29-fold) of lung cancer in the Mexican-American group. When lung cancer was categorized by histologic type, no elevated risk was noted for squamous cell carcinoma in any ethnicity. However, in Mexican-Americans, susceptible genotypes were associated with significantly increased risks of adenocarcinoma (3. 21-fold) and large cell carcinoma (10.79). Our study and others have demonstrated that the PADPRP polymorphism may modify an individual's susceptibility to certain cancers. Assessment of the interaction between genetic constitution and environmental exposure might expand our understanding of carcinogenesis and enhance our ability to evaluate the populational cancer risk.

Adenocarcinoma↗

Structural organization of the human Elk1 gene and its processed pseudogene Elk2.

In the ets gene family of transcription factors, ELK1 belongs to the subfamily of Ternary Complex Factors (TCFs) which bind to the Serum Response Element (SRE) in conjunction with a dimer of Serum Response Factors (SRFs). The primary structure of the human Elk1 gene was determined by genomic cloning. The gene structure of Elk1 spans 15.2 kb and consists of seven exons and six introns. The coding sequence resides on exons 3, 4, 5, 6 and 7. Sequencing of cDNA clones isolated from human hippocampus library revealed that the second exon was often skipped by an alternative splicing event. All introns commenced with nucleotides GT at the 5' boundary and ended with nucleotides AG at the 3' boundary, in agreement with the proposed consensus sequence for intron spliced donor and acceptance sites. Sequence inspection of the 5'-flanking region revealed the absence of a 'TATA' box and the presence of putative cis-acting regulatory elements such as Sp1, GATA-1, CCAAT, and c-Myb. Moreover, the sequence analysis of Elk2 locus on 14q32.3 confirmed that Elk2 gene corresponds to a processed pseudogene of Elk1 which has been reported between alpha 1 gene (IGHA1) and pseudo gamma gene (IGHGP) of immunoglobulin heavy chain. Furthermore, the results of Southern analysis using DNAs from human-mouse hybrid cell lines carrying a part of 14q32 region revealed that there is another locus hybridizing to Elk1 cDNA on 14q32.2 --> qter region in addition to Elk2 locus between IGHA1 and IGHGP loci.

Alternative Splicing↗

A methylated Neurospora 5S rRNA pseudogene contains a transposable element inactivated by repeat-induced point mutation.

In an analysis of 22 of the roughly 100 dispersed 5S rRNA genes in Neurospora crassa, a methylated 5S rRNA pseudogene, Psi63, was identified. We characterized the Psi63 region to better understand the control and function of DNA methylation. The 120-bp 5S rRNA-like region of Psi63 is interrupted by a 1.9-kb insertion that has characteristics of sequences that have been modified by repeat-induced point mutation (RIP). We found sequences related to this insertion in wild-type strains of N. crassa and other Neurospora species. Most showed evidence of RIP; but one, isolated from the N. crassa host of Psi63, showed no evidence of RIP. A deletion from near the center of this sequence apparently rendered it incapable of participating in RIP with the related full-length copies. The Psi63 insertion and the related sequences have features of transposons and are related to the Fot1 class of fungal transposable elements. Apparently Psi63 was generated by insertion of a previously unrecognized Neurospora transposable element into a 5S rRNA gene, followed by RIP. We name the resulting inactivated Neurospora transposon PuntRIP1 and the related sequence showing no evidence of RIP, but harboring a deletion that presumably rendered it defective for transposition, dPunt.

Amino Acid Sequence↗

Genomic organization of the X-linked gene (PIG-A) that is mutated in paroxysmal nocturnal haemoglobinuria and of a related autosomal pseudogene mapped to 12q21.

The PIG-A gene, whose product is involved in one of the early steps in the synthesis of glycan phosphatidylinositol (GPI) anchors, has been recently found to be defective in all cases of paroxysmal nocturnal haemoglobinuria (PNH). By isolating genomic clones from a human phage library we now show that the PIG-A gene consists of six exons (the first of which is non-coding) spanning 17 kb of DNA, and we have mapped the gene to chromosomal position Xp22.1. The PIG-A promoter has features of a housekeeping gene. We have also isolated additional clones which cross-hybridize to PIG-A cDNA, and we have thus identified an intronless PIG-A pseudogene (psi PIG-A), which we have mapped to chromosomal position 12q21. psi PIG-A cannot be functional because it contains several stop codons and a frameshift. These data make it possible to design primers for amplification of the entire PIG-A coding region, with exclusion of psi PIG-A sequences, which will facilitate characterization of PIG-A mutations in patients with PNH. Database searches revealed that PIG-A contains homologies with a number of glycosyl transferases and is highly homologous (45%) to the protein encoded by the yeast SPT14 gene.

Amino Acid Sequence↗

DNA polymorphism in active gene and pseudogene of the cytosolic phosphoglucose isomerase (PgiC) loci in Arabidopsis halleri ssp. gemmifera.

DNA variations in two PgiC loci were investigated in 15 strains of Arabidopsis halleri ssp. gemmifera. In a 5.5-kb region of the PgiC1 locus, 127 nucleotide substitutions and 33 length variations were observed. In a 6.0-kb region of the PgiC2 locus, 138 nucleotide substitutions and 33 length variations were observed. Frame shift, novel stop codons, and large length variations were observed in the PgiC2 coding region. These findings suggested that PgiC2 may be a pseudogene. The nucleotide diversities (pi) for the entire regions of both PgiC loci were approximately 0.0033. Tajima's test of both PgiC loci yielded significantly negative results. In the coding regions, the high proportions of replacement substitutions caused significant deviations from neutrality in McDonald and Kreitman's test. An excess of singletons and a high proportion of replacement polymorphic sites have been observed in the Adh and ChiA regions of A. halleri ssp. gemmifera. Thus, the A. halleri ssp. gemmifera population may not have reached equilibrium, and thus nonneutral patterns of DNA polymorphism were observed.

Arabidopsis↗

Complex pattern of coalescence and fast evolution of a mitochondrial rRNA pseudogene in a recent radiation of tiger beetles.

Transposed copies of mitochondrial DNA into the nucleus (numts) are widespread, but to date they have not been described from the Coleoptera (beetles). Here we report the discovery of a numt derived from a mitochondrial ribosomal RNA gene in Australian tiger beetles (genus Rivacindela). The loss of function of the numt was confirmed by high proportion of transversions, numerous noncompensatory substitutions in stem regions, and large deletions in functionally important sequences. Phylogenetic analysis of orthologous numt sequences was performed together with the corresponding mtDNA lineage for a study of origination and establishment of the transposed copies in closely related populations and species. All numt sequences were strongly supported to be monophyletic, indicating a single origin of this element. However, populations were polymorphic for the presence of the numt, and phylogenetic trees based on the numt sequences showed inconsistencies with the corresponding mtDNA phylogeny, suggesting slower processes of fixation compared to the mtDNA sequences. In a side-by-side comparison with their mtDNA sister lineage, the nucleotide substitution rate of 1.66 x 10(-8) substitutions/site/year in the numts was approximately equal to the average rate of mtDNA in this group but substantially higher than previous estimates of neutral nuclear rates in vertebrates. The numt clade was affected by several deletions but no insertions, with estimates of nucleotide loss exceeding the rate of nucleotide substitutions by approximately five times. The young age of the Rivacindela numt clade, their absence in species outside of a narrow lineage of related individuals, and the high rate of deletions suggest that insertions do not persist in this group, which is consistent with the view that comparatively small genomes as those of Coleoptera harbor fewer mitochondrial and other nuclear pseudogenes.

Base Sequence↗

Sequence diversity at the proximal 14q32.1 SERPIN subcluster: evidence for natural selection favoring the pseudogenization of SERPINA2.

The superfamily of serine protease inhibitors (SERPINs) plays a key role in controlling the activity of proteinases in diverse biological processes. alpha1-antitrypsin (SERPINA1), the most studied member of this family, is encoded by a gene located within the proximal 14q32.1 SERPIN subcluster, together with the highly homologous alpha1-antitrypsin-like sequence (SERPINA2), which was previously proposed to be a pseudogene. Here, we performed a resequencing study encompassing both SERPINA1 and SERPINA2 as well as the adjacent gene coding for corticosteroid-binding globulin (SERPINA6) in samples from Europe and West Africa. In the African sample, we found that a common haplotype carrying a 2-kb deletion in the SERPINA2 gene is associated with remarkable long-range homozygozity as if it was quickly driven to high frequency by natural selection acting on an advantageous variant. An analysis of the HapMap Phase I data for the Yoruba sample confirmed that variation in this subcluster carries a strong signal of positive selection. We also show that the SERPINA2 gene is expressed and probably encodes a functional SERPIN. Finally, comparisons with orthologous sequences in nonhuman primates showed that SERPINA2 is present in some great apes, but in chimpanzees it was lost by a deletion event independent from that observed in humans. In agreement with the "less is more" hypothesis, we propose that loss of SERPINA2 is an ongoing process associated with a selective advantage during recent primate evolution, possibly because of a role in fertility or in host-pathogen interactions.

Africa↗

Unusual features of CpG-rich (HTF) islands in the human alpha globin complex: association with non-functional pseudogenes and presence within the 3' portion of the zeta gene.

We have characterised a cluster of CpG rich (HTF) islands in the alpha-globin complex and report here two unusual features: The human embryonic zeta 2-globin gene is associated with an HTF island within its 3' portion rather than at the 5' end. Furthermore at least two non-functional pseudogenes within the cluster (psi zeta 1 and psi alpha 2) are associated with CpG rich islands.

Animals↗

Characteristics of a multicopy gene family predominantly consisting of processed pseudogenes.

The monoclonal antibody MOC-32 detected a 40 kDa protein in Western blot analysis. Immunological screening of an expression library of human SCLC cells with MOC-32 led to the isolation of overlapping cDNA clones. One of these, cHD4, was 1.0 kbp long and of about the same size as its corresponding mRNA. Preceded by an in phase stop codon, an open reading frame of 885 bp was present in cHD4 and a translational product of only 33 kDa could be calculated. Biochemical and immunological analysis established the relationship between the 40 kDa antigen and the isolated coding sequences and resolved the apparent discrepancy between the calculated molecular weight and the observed electrophoretic mobility. Nucleotide sequence comparison of cHD4 to the EMBL database revealed that cHD4 was nearly identical to a sequence claimed to encode a laminin binding protein. Southern blot and nucleotide sequence analysis indicated the presence of multiple copies of the gene in the human genome. At least five of these appeared to represent processed pseudogenes.

Amino Acid Sequence↗

The XLR sequence family: dispersion on the X and Y chromosomes of a large set of closely related sequences, most of which are pseudogenes.

The XLR sequence family encodes RNA transcripts specific to late-stage T and B cells and their neoplasms. Only one apparently functional mRNA has been identified thus far and this encodes a novel 25 kDa nuclear protein. In this report, we find that the XLR gene family is composed of 50-75 copies per haploid genome which localize to at least two different portions of the mouse X chromosome. Neither of these locations are near the xid mutation that earlier work had correlated with XLR. In addition, some members of this family are also on the Y chromosome. Another surprising finding is that while the fourteen genomic clones examined to date have the same exon-intron structure and are closely related with respect to sequence conservation (90%), all appear (in most cases by multiple criteria) to be non-functional, raising the possibility that all but one of the members of this large semi-dispersed family are pseudogenes.

Animals↗

Genes, variant genes, and pseudogenes of the human tRNA(Val) gene family are differentially expressed in HeLa cells and in human placenta.

Pre-tRNAs(Val) were identified in unfractionated tRNA preparations from HeLa cells and human placenta and their 5' leader structures were deduced from the nucleotide sequences of the corresponding cDNAs. Several of these precursors can be assigned to nine out of the eleven members of the human tRNA(Val) gene family characterized so far, which demonstrates that these gene loci are actively transcribed in vivo. Among the expressed genes there are (a) genes for the two known tRNA(Val) isoacceptor species from human placenta, (b) gene variants that exhibit sequence alterations as compared to conventional genes, and (c) pseudogenes that produce processing-deficient precursors which are not matured to tRNAs. The transcription products of several yet unknown tRNA(Val) genes have also been detected. Furthermore, different expression patterns are observed in the two cell types studied. These data allow for the first time an insight into the in vivo expression of a human tRNA gene family.

Base Sequence↗

Molecular analysis of a U3 RNA gene locus in tomato: transcription signals, the coding region, expression in transgenic tobacco plants and tandemly repeated pseudogenes.

By screening a tomato genomic library with a tomato U3 RNA probe, we detected a U3 genomic locus whose coding region was determined by primer extension (5' end) and direct RNA sequencing of purified U3 RNA from tomato (3' end). Tomato U3 RNA is 216 nucleotides long, contains all the four evolutionarily highly conserved sequence blocks (Boxes A to D), has at its 5' end a cap not precipitable with anti-m3G antibodies and can be folded into a peculiar secondary structure with two stem-loops at its 5' end. A tagged derivative of the U3 gene was faithfully expressed in transgenic tobacco plants. In the 5' flanking region both plant-specific UsnRNA transcription signals [the TATA-like sequence and the upstream sequence element (USE)] were present, but were positioned closer to each other and also to the cap site in the U3 gene than in the genes for the plant spliceosomal UsnRNAs studied so far. The 3' flanking region of the tomato U3 gene lacked the consensus sequence of the putative termination signal established for the plant spliceosomal UsnRNA genes and contained a pyrimidine-rich tract (R1) followed by four tandemly repeated U3 pseudogenes (U3.1 ps to U3.4 ps) flanked by slightly altered forms (R2 to R5) of R1 and most probably generated by DNA-mediated events. Our results are in line with the conjecture that the enzyme transcribing the tomato U3 gene has different structural requirements for transcriptional activity than the enzyme transcribing plant U1, U2 and U5 genes.

Base Sequence↗