Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “pseudogene”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Homologous nonallelic recombinations between the iduronate-sulfatase gene and pseudogene cause various intragenic deletions and inversions in patients with mucopolysaccharidosis type II.

About 20% of patients with mucopolysaccharidosis type II (MPS II) have gross structural rearrangements involving the iduronate-sulfatase (IDS) gene in Xq27.3-q28. A nearby IDS pseudogene (IDS-2) promotes nonallelic recombination between highly homologous sequences. Here we describe major rearrangements due to gene/pseudogene recombination. In two unrelated patients, partial IDS gene deletions were found joining introns 3 and 7 of the IDS gene together with gene to pseudogene conversion in the area of breakpoints. In a third patient, a junction between intron 3 of IDS-2 and intron 7 of IDS was seen that was due to a deletion and inversion of the 5' part of the gene. Characterisation of breakpoints in six patients with large inversions revealed that all recombinations of this type occurred in the same area of homology between IDS and IDS-2; they were molecularly balanced, and accompanied by gene conversions in most cases. Apart from diagnostic implications, such naturally occurring recombination 'hot spots' may allow some insight into general features of crossover events in mammals.

Alleles↗

Observations on the structure of two human 7SK pseudogenes and on homologous transcripts in vertebrate species.

A comparison of the sequence of two human 7SK RNA pseudogenes, covering approx. 190 and 240 base-pairs of the structural gene, is presented. Both repeated elements are flanked by direct repeats and begin at the 5' end of the gene. Each terminates approx. 90 base-pairs short of the 3' end, the latter representing a continuous sequence and the former carrying an internal deletion of about 40 base-pairs, this region being flanked in the progenitor gene by short repeated sequences. Southern blotting using a human 7SK pseudogene probe illuminated a series of multiple restriction fragments in mammalian genomes, with generally fewer fragments in the genomes of birds and reptiles and a single reactive fragment in DNA from terrapin (Pseudemys scripta elegans) and Xenopus laevis (South African clawed toad). In the latter case this fragment was only detectable on long exposure under the hybridization stringencies employed. 7SK transcripts were readily detectable in all mammalian, avian, reptilian and amphibian species analysed, although the gene appeared to be expressed at rather low levels in the ovaries of Xenopus laevis, possibly accounting for its failure to have become dispersed via 'retroposition' in this species.

Animals↗

Multiple nuclear pseudogenes of mitochondrial cytochrome b in Ctenomys (Caviomorpha, rodentia) with either great similarity to or high divergence from the true mitochondrial sequence.

A fragment of the mitochondrial cytochrome b gene was studied in 13 species of the South American fossorial rodent Ctenomys using PCR with 'universal' primers and DNA sequencing after cloning. Five different groups of sequences were found, one of which corresponds to the functional mitochondrial gene (mt). The other four groups (A, B, C and D) were believed to be nuclear pseudogenes. Sequences A-C were highly divergent from the mt sequences and included substitutions, deletions and insertions such that they could not possibly have coded a functional protein. They all shared a common insertion between positions 15055 and 15056 suggestive of a common origin, although the A, B and C sequences otherwise differed greatly from each other. The D sequences also could not have been functional on the basis of nucleotide sequence, but the differences with the mt sequences were far more subtle and in a more limited study the D sequences could easily have been classified as a true mtDNA sequence. It is suggested that there were two transfers of the cytochrome b gene from the mitochondrion to the nucleus; the first leading to sequences A-C and the second to the D sequence. Subsequent to transfer, a sequence of duplications within the nucleus appears to have generated the full range of pseudogenes that are observed. This study adds to other recent observations suggesting the frequent transfer of mtDNA sequences to the nucleus and reinforces the necessity of great care in interpreting PCR-generated sequences, particularly those produced with universal primers. There are now data from several species of mammals and birds relating to PCR-generated nuclear copies of cytochrome b, which we review.

Animals↗

The opcA and (psi)opcB regions in Neisseria: genes, pseudogenes, deletions, insertion elements and DNA islands.

Previous data have indicated that the opc gene encoding an immunogenic invasin is specific to Neisseria meningitidis (Nm) and is lacking in Neisseria gonorrhoeae (Ng). The data presented here show that Nm and Ng both contain two paralogous opc-like genes, opcA, corresponding to the former opc gene, and (psi)opcB, a pseudogene. The predicted OpcA and OpcB proteins possess transmembrane regions with conserved non-polar faces but differ extensively in four of the five surface-exposed loops. Gonococcal OpcA was expressed weakly under in vitro conditions, and it is unknown whether these bacteria can express this protein at high levels. Analysis of the sequences flanking opcA and (psi)opcB revealed a framework of conserved housekeeping genes interspersed with DNA islands. These regions also contained several pseudogenes, deletions and IS elements, attesting to considerable genome plasticity. Both opcA and (psi)opcB are located on DNA islands that have probably been imported from unrelated bacteria. A third island encodes the dcmD/dcrD R/M genes in Ng versus a small open reading frame in most strains of Nm. Rare strains of Nm were identified in which the R/M island has been imported. DNA islands in Nm and Ng seem to have been acquired by recombination via conserved flanking housekeeping genes rather than by insertion of mobile genetic elements.

Adhesins, Bacterial↗

Evolution of the paralogous hap and iga genes in Haemophilus influenzae: evidence for a conserved hap pseudogene associated with microcolony formation in the recently diverged Haemophilus aegyptius and H. influenzae biogroup aegyptius.

Certain non-capsulate strains belonging to the Haemophilus influenzae/Haemophilus aegyptius complex show unusually high pathogenicity, but the evolutionary origin of these virulent phenotypes, termed H. influenzae biogroup aegyptius, is as yet unknown. The aim of the present study was to elucidate the mechanisms of evolution of two paralogous genes, hap and iga, which encode the adhesion and penetration Hap protein and the IgA1 protease respectively. Partial sequencing of hap and iga genes in a comprehensive collection of strains belonging to the H. influenzae/H. aegyptius complex revealed considerable genetic polymorphism and pronounced mosaic-like patterns in both genes, but no evidence of intrastrain recombination between the two genes. A conserved hap pseudogene was present in all strains of H. aegyptius and H. influenzae biogroup aegyptius, each of which constituted distinct subpopulations as revealed by phylogenetic analysis. There was no evidence for a second, functional copy of the hap gene in these strains. The perturbed expression of the Hap serine protease appears to be associated with the formation of elongated bacterial cells growing in chains and a distinct colonization pattern on conjunctival cells, previously termed microcolony formation. The fact that individual hap pseudogenes differed from the ancestral sequence by zero to two positions within a 1.5 kb stretch suggests that the silencing event happened approximately 2000-11,000 years ago. Divergence of H. aegyptius and H. influenzae biogroup aegyptius occurred subsequent to this genetic event. The loss of Hap protein expression may be one of the genetic events that facilitated exploitation of the conjunctivae as a new niche.

Bacterial Adhesion↗

Exempting homologous pseudogene sequences from polymerase chain reaction amplification allows genomic keratin 14 hotspot mutation analysis.

In patients with the major forms of epidermolysis bullosa simplex, either of the keratin genes KRT5 or KRT14 is mutated. This causes a disturbance of the filament network resulting in skin fragility and blistering. For KRT5, a genomic mutation detection system has been described previously. Mutation detection of KRT14 on a DNA level is, however, hampered by the presence of a highly homologous but nontranscribed KRT14 pseudogene. Consequently, mutation detection in epidermolysis bullosa simplex has mostly been carried out on cDNA synthesized from KRT5 and KRT14 transcripts in mRNA isolated from skin biopsies. Here we present a genomic mutation detection system for exons 1, 4, and 6 of KRT14 that encode the 1A, L1-2, and 2B domains of the keratin 14 protein containing the mutation hotspots. After cutting the KRT14 pseudogene genomic sequences with restriction enzymes while leaving the homologous genomic sequences of the functional gene intact, only the mutation hotspot-containing exons of the functional KRT14 gene are amplified. This is followed by direct sequencing of the polymerase chain reaction products. In this way, three novel mutations could be identified, Y415H, L419Q, and E422K, all located in the helix termination motif of the keratin 14 rod domain 2B, resulting in moderate, severe, and mild epidermolysis bullosa simplex phenotype, respectively. By obviating the need of KRT14 cDNA synthesis from RNA isolated from skin biopsies, this approach substantially facilitates the detection of KRT14 hotspot mutations.

DNA Mutational Analysis↗

Reactivation by exon shuffling of a conserved HLA-DR3-like pseudogene segment in a New World primate species.

The common marmoset (Callithrix jacchus), a New World monkey species with a limited MHC class II repertoire, is highly susceptible to certain bacterial infections. Genomic analysis of exon 2 sequences documented the existence of only one DRB region configuration harboring three loci. Two of these loci display moderate levels of allelic polymorphism, whereas the -DRB*W12 gene appears to be monomorphic. This study shows that only the Caja-DRB*W16 and -DRB*W12 loci produce functional transcripts. The Caja-DRB1*03 locus is occupied by a pseudogene, given that most of the transcripts, if detected at all, show imperfections and are present at low levels. Moreover, two hybrid transcripts were identified that feature the evolutionarily conserved peptide-binding motif characteristic for the Caja-DRB1*03 gene. Thus, the severely reduced MHC class II repertoire in common marmosets has been expanded by reactivation of a pseudogene segment as a result of exon shuffling.

Amino Acid Sequence↗

Isolation and characterization of the human homologue of rig and its pseudogenes: the functional gene has features characteristic of housekeeping genes.

rig (rat insulinoma gene) was first isolated from a cDNA library of rat insulinomas and has been found to be activated in various human tumors such as insulinomas, esophageal cancers, and colon cancers. Here we isolated the human homologue of rig from a genomic DNA library constructed from a human esophageal carcinoma and determined its complete nucleotide sequence. The gene is composed of about 3000 nucleotides and divided into four exons separated by three introns: exon 3 encodes the nuclear location signal and the DNA-binding domain of the RIG protein. The transcription initiation site was located at -46 base pairs upstream from the first ATG codon. The 5'-flanking region of the gene has no apparent TATA-box or CAAT-box sequence. However, two GC boxes are found at -189 and -30 base pairs upstream from the transcription initiation site and five GC boxes are also found in introns 1 and 2. The gene is bounded in the 5' region by CpG islands, regions of DNA with a high GC content and a high frequency of CpG dinucleotides relative to the bulk genome. Furthermore, the human genome contains at least six copies of RIG pseudogenes, and four of them have the characteristics of processed pseudogenes. From these results together with the finding that RIG is expressed in a wide variety of tissues and cells, we speculate that RIG belongs to the class of "housekeeping" genes, whose products are necessary for the growth of all cell types.

Adenoma, Islet Cell↗

Multiple human D5 dopamine receptor genes: a functional receptor and two pseudogenes.

Three genes closely related to the D1 dopamine receptor were identified in the human genome. One of the genes lacks introns and encodes a functional human dopamine receptor, D5, whose deduced amino acid sequence is 49% identical to that of the human D1 receptor. Compared with the human D1 dopamine receptor, the D5 receptor displayed a higher affinity for dopamine and was able to stimulate a biphasic rather than a monophasic intracellular accumulation of cAMP. Neither of the other two genes was able to direct the synthesis of a receptor. Nucleotide sequence analysis revealed that these two genes are 98% identical to each other and 95% identical to the D5 sequence. Relative to the D5 sequence, both contain insertions and deletions that result in several in-frame termination codons. Premature termination of translation is the most likely explanation for the failure of these genes to produce receptors in COS-7 and 293 cells even though their messages are transcribed. We conclude that the two are pseudogenes. Blot hybridization experiments performed on rat genomic DNA suggest that there is one D5 gene in this species and that the pseudogenes may be the result of a relatively recent evolutionary event.

Amino Acid Sequence↗

A novel polymorphic cytochrome P450 formed by splicing of CYP3A7 and the pseudogene CYP3AP1.

The cytochrome P450 3A7 (CYP3A7) is the most abundant CYP in human liver during fetal development and first months of postnatal age, playing an important role in the metabolism of endogenous hormones, drugs, differentiation factors, and potentially toxic and teratogenic substrates. Here we describe and characterize a novel enzyme, CYP3A7.1L, encompassing the CYP3A7.1 protein with the last four carboxyl-terminal amino acids replaced by a unique sequence of 36 amino acids, generated by splicing of CYP3A7 with CYP3AP1 RNA. The corresponding CYP3A7-3AP1 mRNA had a significant expression in liver, kidney, and gastrointestinal tract, and its presence was found to be tissue-specific and dependent on the developmental stage. Heterologous expression in yeast revealed that CYP3A7.1L was a functional enzyme with a specific activity similar to that of CYP3A7.1 and, in some conditions, a different hydroxylation specificity than CYP3A7.1 using dehydroepiandrosterone as a substrate. CYP3A7.1L was found to be polymorphic due to a mutation at position -6 of the first splicing site of CYP3AP1 (CYP3A7_39256T-->A), which abrogates the pseudogene splicing. This polymorphism had pronounced interethnic differences and was in linkage disequilibrium with other functional polymorphisms described in the CYP3A locus: CYP3A7*2 and CYP3A5*1. Therefore, the resulting CYP3A haplotypes express different sets of enzymes within the population. In conclusion, a novel mechanism, consisting of the splicing of the pseudogene CYP3AP1 to CYP3A7, causes the formation of the novel CYP3A7.1L having a different tissue distribution and functional properties than the parent CYP3A7 enzyme, with possible developmental, physiological, and toxicological consequences.

Adult↗

A processed pseudogene codes for a new antigen recognized by a CD8(+) T cell clone on melanoma.

The M88.7 T cell clone recognizes an antigen presented by HLA B*1302 on the melanoma cell line M88. A cDNA encoding this antigen (NA88-A) was isolated using a library transfection approach. Analysis of the genomic gene's sequence identified it is a processed pseudogene, derived from a retrotranscript of mRNA coding for homeoprotein HPX42B. The NA88-A gene exhibits several premature stop codons, deletions, and insertions relative to the HPX42B gene. In NA88-A RNA, a short open reading frame codes for the peptide MTQGQHFLQKV from which antigenic peptides are derived; a stop codon follows the peptide's COOH-terminal Val codon. Part of the HPX42B mRNA's 3' untranslated region codes for a peptide of similar sequence (MTQGQHFSQKV). If produced, this peptide can be recognized by M88.7 T cells. However, in HPX42B mRNA, the peptide's COOH-terminal Val codon is followed by a Trp codon. As a result, expression of HPX42B mRNA does not lead to antigen production. A model is proposed for events that participated in creation of a gene coding for a melanoma antigen from a pseudogene.

Amino Acid Sequence↗

Folate biosynthesis pseudogenes, PsifolP and PsifolK, and an O-sialoglycoprotein endopeptidase gene homolog in the phytoplasma genome.

Phytoplasmas are wall-less phytopathogenic prokaryotes of small genome sizes that are obligate parasites of insect vectors and plant hosts. We have cloned a clover phyllody (CPh) phytoplasma DNA locus containing five potential coding sequences. Two were identified as pseudogenes (PsifolP and PsifolK) homologous to folP and folK genes, which encode dihydropteroate synthase (DHPS) and 6-hydroxymethyl-7,8-dihydropterin pyrophosphokinase (HPPK), respectively, in other bacteria. Evolution of the phytoplasma presumably involved loss of functions through the formation of these and other pseudogenes during adaptation to obligate parasitism. The findings suggest that the phytoplasma lacks capacity for de novo folate biosynthesis and possesses a transport system for absorption of preformed folate from host cells. The PsifolP-PsifolK region was flanked by three open reading frames (ORFs) encoding a DegV family protein, a hypothetical protein with a P60-like lipoprotein domain homologous with the P60-like Mycoplasma hominis protein, and a glycoprotease (Gcp) protein that possibly functions as a host adaptation or virulence factor.

Amino Acid Sequence↗

Human gene encoding the 78,000-dalton glucose-regulated protein and its pseudogene: structure, conservation, and regulation.

The isolation and characterization of a human functional GRP78 gene and a processed pseudogene are described. We present the complete primary structure of the human GRP78 gene, which spans over 5 kb and consists of eight exons. Sequence comparisons reveal that the GRP78 gene shares unusual homology among the human, rat, and hamster in the protein-coding and 3' untranslated regions. In addition, short domains highly conserved with HSP70 isolated from human, Drosophila, Xenopus, yeast, and E. coli DNA are identified within the hydrophobic regions of GRP78. The intronless pseudogene resembles that of a processed gene. It is flanked by a short direct repeat and is embedded within an AT-rich genomic region. The highly active promoter from the functional human GRP78 gene contains a TATA box, five CCAAT sequences, and two potential binding sites for the transcriptional factor Sp1. It consists of a distal domain that enhances basal level expression and a proximal domain essential for responses to calcium ionophore and for a temperature-sensitive mutation which induce the GRP78 gene. Both domains are highly conserved between the rat and the human GRP78 promoters.

Amino Acid Sequence↗

A phospholipase A2-like pseudogene retaining the highly conserved introns of Mojave toxin and other snake venom group II PLA2s, but having different exons.

Mojave toxin is a neurotoxic, heterodimeric phospholipase A2 (PLA2) from the venom of the Mojave rattlesnake (Crotalus scutulatus scutulatus) and is characteristic of all rattlesnake presynaptic neurotoxins. Here, we describe a phospholipase A2 pseudogene (psi-Mtx) located 2,000 nucleotides upstream, and on the opposite DNA strand, from a gene for Mojave toxin acidic subunit (Mtx-a). The pseudogene lacks the first exon and a few segments of noncoding DNA found in functional snake venom PLA2 genes, but does have the coding information for a complete PLA2 protein. psi-Mtx retains the unusual gene sequence similarity pattern found in functional viperid PLA2 genes. When compared to genes from C. s. scutulatus and the Hahn snake (Trimeresurus flavoviridus), psi-Mtx shows strong conservation of nocoding regions and variable protein-coding regions. Although the nocoding regions of psi-Mtx are conserved with respect to other viperid PLA2 genes, the three exons code for a unique PLA2-like protein similar in sequence to ammodytoxin b found in the venom of the western sand viper (Vipera ammodytes ammodytes). The structure of these genes suggests a common ancestor for all viperid PLA2 genes. Phylogenetic analysis of psi-Mtx, Mtx-a, Mtx-b, pgPLA 1a, and pgPLA 1b suggest that psi-Mtx diverged from an ancestral sequence before the presumed gene duplication event leading to Mtx-a and Mtx-b. However, analysis of the basis of coding regions alone gives a conflicting result.

Amino Acid Sequence↗

Primate microRNAs miR-220 and miR-492 lie within processed pseudogenes.

MicroRNAs (miRNAs) are a new and abundant class of small, noncoding RNAs. To date, the evolutionary history of most of these loci appears to be marked by duplication and divergence. The ultimate origin of miRNAs remains an open question. A survey of the genomic context of more than 300 human miRNA loci revealed that two primate-specific miRNAs, miR-220 and miR-492, each lie within a processed pseudogene. In silico and in vitro examinations of these two loci suggest that this is a rare phenomenon requiring the juxtaposition of a specific combination of factors. Thus it appears that, while processed pseudogenes are good candidates for miRNA incubators, it is unlikely that more than a very small percentage of new miRNAs arise this way.

Animals↗

Evolution of the trnF(GAA) gene in Arabidopsis relatives and the brassicaceae family: monophyletic origin and subsequent diversification of a plastidic pseudogene.

Recently, we used the 5'-trnL(UAA)-trnF(GAA) region of the chloroplast DNA for phylogeographic reconstructions and phylogenetic analysis among the genera Arabidopsis, Boechera, Rorippa, Nasturtium, and Cardamine. Despite the fact that extensive gene duplications are rare among the chloroplast genome of higher plants, within these taxa the anticodon domain of the trnF(GAA) gene exhibit extensive gene duplications with one to eight tandemly repeated copies in close 5' proximity of the functional gene. Interestingly, even in Arabidopsis thaliana we found six putative pseudogenic copies of the functional trnF gene within the 5'-intergenic trnL-trnF spacer. A reexamination of trnL(UAA)-trnF(GAA) regions from numerous published phylogenetic studies among halimolobine, cardaminoid, and other cruciferous taxa revealed not only extensive trnF gene duplications but also favor the hypothesis about a single origin of trnF pseudogene formation during evolution of the Brassicaceae family 16-21 MYA. Conserved sequence motifs from this tandemly repeated region are codistributed nonrandomly throughout the plastome, and we found some similarities with a DNA sequence duplication in the rps7 gene and its adjacent spacer. Our results demonstrate the potential evolutionary dynamics of a plastidic region generally regarded as highly conserved and probably cotranscribed and, as shown here for several genera among cruciferous plants, greatly characterized by parallel gains and losses of duplicated trnF copies.

Arabidopsis↗

NKp30 (NCR3) is a pseudogene in 12 inbred and wild mouse strains, but an expressed gene in Mus caroli.

Ancient duplications and rearrangements of protein-coding segments have resulted in complex gene family relationships. As a result, gene products may acquire new specificities, altered recognition properties, modified functions, and even loss of functionality. The natural cytotoxicity receptor (NCR) family are natural killer (NK)-activating receptors whose members are NKp46 (NCR1), NKp44 (NCR2), and NKp30 (NCR3). The NCR proteins are putative immunoglobulin superfamily members whose ligands are unknown. The NKp46 gene is present and expressed in human and mouse, NKp44 is only present and expressed in human, and NKp30 is present and expressed in human but is a nonexpressed pseudogene in mouse. By searching databases we have detected alternatively spliced forms of the three NCR members. In addition, we have shown by reverse transcription-polymerase chain reaction (RT-PCR) analysis that the human NKp30 gene presents differential expression patterns in tissues. However, no expressed sequence tags (ESTs) are detected for mouse NKp30, and the genomic sequence contains two premature stop codons, which would encode a severely truncated nonfunctional protein. We have sequenced genomic DNA from 13 mouse inbred and wild strains and discovered that NKp30 is a pseudogene in every mouse strain sequenced except Mus caroli where two single nucleotide polymorphisms (SNPs) abolished the premature stop codons. We observed that the laboratory-inbred strains are, for the exonic sequences, genetically identical, except Mus m. musculus C3H. The Mus musculus strains only have a few SNPs, but the rest of the Mus strains have accumulated gradually several SNPs, mainly in the functional immunoglobulin and intracellular domains. RT-PCR analysis performed on RNA from M. caroli tissue samples identified two transcripts, one of which would encode a putative soluble NKp30 protein, also detected in rat but not in human. We have observed that the intracellular domains of NKp30 (and NKp46) are not conserved among the different species, with the most striking difference when comparing human against mouse and rat. The NKp44 gene is only found in human and shows three different splice forms varying in their "stalk" and intracellular domains. Searching for NKp44 orthologs, we found similarity to ESTs from a novel rodent TREM family member, which we termed TREM6, and not to any possible NKp44 ortholog.

Alternative Splicing↗

Supernetwork identifies multiple events of plastid trnF(GAA) pseudogene evolution in the Brassicaceae.

The occurrence of nonfunctional trnF pseudogenes has been rarely described in flowering plants. However, we describe the first large-scale supernetwork for the Brassiccaeae built from gene trees for 5 loci (adh, chs, matK, trnL-F, and ITS) and report multiple independent origins for trnF pseudogenes in crucifers. The duplicated regions of the original trnF gene are comprised of its anticodon domain and several other highly structured motifs not related to the original gene. Length variation of the trnL-F intergenic spacer region in different taxa ranges from 219 to 900 bp as a result of differences in pseudocopy number (1-14). It is speculated that functional constraints favor 2-3 or 5-6 copies, as found in Arabidopsis and Boechera. The phylogenetic distribution of microstructural changes for the trnL-F region supports ancient patterns of divergence in crucifer evolution for some but not all gene loci.

Brassicaceae↗