Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “pseudogene”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Identification of 11 pseudogenes in the DNA methyltransferase gene family in rodents and humans and implications for the functional loci.

DNA (cytosine-5-)-methyltransferase genes are important for normal development in mice and humans. We describe here 11 pseudogenes spread among human, mouse, and rat belonging to this gene family, ranging from 1 pseudogene in humans to 7 in rat, all belonging to the Dnmt3 subfamily. All except 1 rat Dnmt3b pseudogene appear to be transcriptionally silent. Dnmt3a2, a transcript variant of Dnmt3a starting at an alternative promoter, had the highest number of processed pseudogenes, while none were found for the canonical Dnmt3a, suggesting the former transcript is more highly expressed in germ cells. Comparison of human, mouse, and rat Dnmt3a2 sequences also suggests that human exon 8 is a recent acquisition. Alignment of the 3'UTR of Dnmt3a2 among the functional genes and the processed pseudogenes suggested that a second polyadenylation site downstream of the RefSeq poly(A) was being used in mice, resulting in a longer 3'UTR, a finding confirmed by RT-PCR in mouse tissues. We also found conserved cytoplasmic polyadenylation elements, usually implicated in regulating translation in oocytes, in Dnmt3b and Dnmt1. Expression of DNMT3B in the mouse oocyte was confirmed by immunocytochemistry. These results clarify the structure of a number of loci in the three species examined and provide some useful insights into the structure and evolution of this gene family.

3' Untranslated Regions↗

Vertebrate pseudogenes.

Pseudogenes are commonly encountered during investigation of the genomes of a wide range of life forms. This review concentrates on vertebrate, and in particular mammalian, pseudogenes and describes their origin and subsequent evolution. Consideration is also given to pseudogenes that are transcribed and to the unusual group of genes that exist at the interface between functional genes and non-functional pseudogenes. As the sequences of different genomes are characterised, the recognition and interpretation of pseudogene sequences will become more important and have a greater impact in the field of molecular genetics.

Animals↗

A human serotonin-7 receptor pseudogene.

Although the serotonin-7 receptor was cloned several years ago, its localization in brain tissues remains confusing because of the existence of a related expressed pseudogene, the sequence of which has not hitherto been reported. During the course of searching for related receptor genes, we also searched for this pseudogene to determine its sequence. Human genomic DNA was screened for dopamine and serotonin receptor-like genes, using the polymerase chain reaction method and degenerate oligonucleotide primers based on the similar sequences in the transmembrane-6 and -7 regions of the serotonin-5A, the serotonin-7, and the dopamine D2, D3 and D4 receptors. This resulted in one of the clones having a 115 bp fragment, of which 89% of the bases were identical to the transmembrane-6 and -7 regions of the serotonin-7 receptor sequence. The fragment was radiolabelled and used to screen a human fetal brain cDNA library. A novel cDNA clone of 1326 bp was isolated. Based on the nucleotide sequence, 88% of the bases in this sequence of the pseudogene are identical to the human serotonin-7 receptor coding sequence. However, compared to the serotonin-7 receptor DNA sequence, the pseudogene sequence has nucleotide deletions and insertions, resulting in frame-shifts and stop codons. It was concluded that this sequence represented a pseudogene related to the serotonin-7 receptor gene.

Amino Acid Sequence↗

Mutational analysis of patients with p47-phox-deficient chronic granulomatous disease: The significance of recombination events between the p47-phox gene (NCF1) and its highly homologous pseudogenes.

OBJECTIVE: The aim of this study was to determine the molecular basis of p47-phox-deficient chronic granulomatous disease (CGD), the most common autosomal recessive form of the disease. CGD is an inherited condition characterized by defective oxygen radical production due to defects in the phagocyte nicotinamide adenine dinucleotide phosphate (NADPH) oxidase. Mutational analysis of p47-phox-deficient CGD patients previously demonstrated that the majority of patients have a GT dinucleotide (Delta GT) deletion at the start of exon 2, a signature sequence also observed in the highly homologous pseudogenes of NCF1. MATERIALS AND METHODS: We performed genetic analysis of NCF1 and its pseudogenes using genomic DNA in 29 p47-phox-deficient CGD patients from 22 separate families. First-strand cDNA analysis was performed in 17 of the 29 patients. RESULTS: We confirmed the significance of the Delta GT mutation; in 27 of 29 patients, only the Delta GT sequence was detectable. All but one of the 27 had at least one additional signature sequence, specific to the pseudogene, in either intron 1 and/or intron 2. We extended our analysis to look at signature sequence differences in exons 6 and 9 and detected both the wild-type and pseudogene sequences in all patients tested. CONCLUSIONS: Although detection of only Delta GT sequence accounts for over 85% of affected patients, the molecular basis is most likely due to partial cross-over events between the wild-type and pseudogene(s) of p47-phox at different recombination sites. Our results suggest that complete gene conversion or deletion of the p47-phox gene (NCF1) occurs rarely, if it all.

Base Sequence↗

Genetic complexity of the human geranylgeranyltransferase I beta-subunit gene: a multigene family of pseudogenes derived from mis-spliced transcripts.

Geranylgeranyltransferase I controls the function of a variety of cellular proteins by attaching a geranylgeranyl group to the carboxy-terminus of proteins. The purified enzyme from rat brain is comprised of two polypeptides, a catalytic alpha-subunit (GGTalpha) and a substrate-binding beta-subunit (GGTbeta). The present paper demonstrates the existence of a GGTbeta multigene family in humans by describing the presence and characterization of at least 13 pseudogenes related to this protein. Sequencing of numerous PCR-derived clones, obtained following amplification of human genomic DNA, revealed multiple, distinct but highly related sequences. All clones had a common deletion of 99-bp that conforms to the GT-AG rule of splicing in eukaryotes, and differed from the human GGTbeta cDNA sequence by multiple nucleotide substitutions. PCR amplification from mRNA, however, yielded only the sequence expected for the expressed GGTbeta protein. This apparent paradox was resolved by cloning and sequencing a complete GGTbeta-specific pseudogene. Multiple features of the cloned gene, in particular the absence of introns, presence of flanking direct repeats, and the lack of sequence similarity with the untranscribed region of the gene, indicate that this clone represents a processed pseudogene possibly resulting from a mis-spliced transcript. Multiple GGTbeta-specific pseudogenes appear to have resulted from more than one retroposition event. These results suggest a potential role for mis-splicing in the evolutionary diversity of pseudogenes.

Alkyl and Aryl Transferases↗

Juxta-centromeric region of human chromosome 21 is enriched for pseudogenes and gene fragments.

A physical map including four pseudogenes and 10 gene fragments and spanning 500 kb in the juxta-centromeric region of the long arm of human chromosome 21 is presented. cDNA fragments isolated from a selected cDNA library were characterized and mapped to the 831B6 YAC and to two BAC contigs that cover 250 kb of the region. An 85 kb genomic sequence located in the proximal region of the map was analyzed for putative exons. Four pseudogenes were found, including psiIGSF3, psiEIF3, psiGCT-rel whose functional copies map to chromosome 1p13, chromosome 2 and chromosome 22q11, respectively. The TTLL1 pseudogene corresponds to a new gene whose functional copy maps to chromosome 22q13. Ten gene fragments represent novel sequences that have related sequences on different human chromosomes and show 97-100% nucleotide identity to chromosome 21. These may correspond to pseudogenes on chromosome 21 and to functional genes in other chromosomes. The 85 kb genomic sequence was analyzed also for GC content, CpG islands, and repetitive sequence distribution. A GC-poor L isochore spanning 40 kb from satellite 1 was observed in the most centromeric region, next to a GC-rich H isochore that is a candidate region for the presence of functional genes. The pericentric duplication of a 7.8 kb region that is derived from the 22q13 chromosome band is described. We showed that the juxta-centromeric region of human chromosome 21 is enriched for retrotransposed pseudogenes and gene fragments transferred by interchromosome duplications, but we do not rule out the possibility that the region harbors functional genes also.

Animals↗

Structure and organization of the human alpha class glutathione S-transferase genes and related pseudogenes.

We have isolated and characterized genomic DNA encoding several human Alpha class glutathione S-transferase genes and pseudogenes. All the genes are composed of seven exons with boundaries identical to those of the Alpha class genes in rats. The GSTA1 gene is approximately 12 kb in length and is closely flanked by other Alpha class gene sequences. The complete sequence of the 1.7-kb intergenic region between exon 7 of an upstream pseudogene and exon 1 of the GSTA1 gene has been determined. An additional gene that encodes an uncharacterized Alpha class glutathione S-transferase has been identified. The protein derived from this gene would have 19 amino acid substitutions compared with the GSTA1 isoenzyme. Several pseudogenes with single-base and/or complete exon deletions have been identified, but no reverse-transcribed pseudogenes have been detected. The occurrence of multiple genes and pseudogenes on a single fragment of cloned genomic DNA and the prior identification of a single chromosomal region (6p12) of hybridization (Board and Webb, 1987, Proc. Natl. Acad. Sci. USA 84:2377-2381) suggest that all the Alpha class genes are members of a closely linked gene family that has evolved by duplication and gene conversion events.

Amino Acid Sequence↗

Human cytosolic sulfotransferase database mining: identification of seven novel genes and pseudogenes.

A total of 10 SULT genes are presently known to be expressed in human tissues. We performed a comprehensive genome-wide search for novel SULT genes using two different but complementary approaches, and developed a novel graphical display to aid in the annotation of the hits. Seven novel human SULT genes were identified, five of which were predicted to be pseudogenes, including two processed pseudogenes and three pseudogenes that contained introns. Those five pseudogenes represent the first unambiguous SULT pseudogenes described in any species. Expression-profiling studies were conducted for one novel gene, SULT6B1, and a series of alternatively spliced transcripts were identified in the human testis. SULT6B1 was also present in chimpanzee and gorilla, differing at only seven encoded amino-acid residues among the three species. The results of these database mining studies will aid in studies of the regulation of these SULT genes, provide insights into the evolution of this gene family in humans, and serve as a starting point for comparative genomic studies of SULT genes.

Amino Acid Sequence↗

Pseudogenes in metazoa: origin and features.

The complete genome sequences with their annotations are a considerable resource in biology, particularly in understanding the global structure of the genetic material at the molecular level. The reason why some eukaryotic genomes contain large quantities of apparently unnecessary DNA, namely pseudogenes, while others seem to invest in more efficient thinning processes or are equipped with protection systems against parasitic elements still remains a mystery. Several genome-wide surveys have been undertaken to identify pseudogenes in the completely sequenced genome, bringing to light some differences both in their amount and distribution. Since pseudogenes are important resources in evolutionary and comparative genomics - as 'molecular fossils' - in this paper, a survey on the origins, features, abundance and localisation of the different pseudogenes is reported. As an example of genes producing processed pseudogenes, some experimental data obtained in the authors' laboratories from the study of a nuclear gene coding for the mitochondrial transcription factor A (mtTFA), a key regulator of mitochondrial biogenesis, are also reported.

Animals↗

Pseudogenes, junk DNA, and the dynamics of Rickettsia genomes.

Studies of neutrally evolving sequences suggest that differences in eukaryotic genome sizes result from different rates of DNA loss. However, very few pseudogenes have been identified in microbial species, and the processes whereby genes and genomes deteriorate in bacteria remain largely unresolved. The typhus-causing agent, Rickettsia prowazekii, is exceptional in that as much as 24% of its 1.1-Mb genome consists of noncoding DNA and pseudogenes. To test the hypothesis that the noncoding DNA in the R. prowazekii genome represents degraded remnants of ancestral genes, we systematically examined all of the identified pseudogenes and their flanking sequences in three additional Rickettsia species. Consistent with the hypothesis, we observe sequence similarities between genes and pseudogenes in one species and intergenic DNA in another species. We show that the frequencies and average sizes of deletions are larger than insertions in neutrally evolving pseudogene sequences. Our results suggest that inactivated genetic material in the Rickettsia genomes deteriorates spontaneously due to a mutation bias for deletions and that the noncoding sequences represent DNA in the final stages of this degenerative process.

Base Sequence↗

Mitochondrial pseudogenes are pervasive and often insidious in the snapping shrimp genus Alpheus.

Here we show that multiple DNA sequences, similar to the mitochondrial cytochrome oxidase I (COI) gene, occur within single individuals in at least 10 species of the snapping shrimp genus Alpheus. Cloning of amplified products revealed the presence of copies that differed in length and (more frequently) in base substitutions. Although multiple copies were amplified in individual shrimp from total genomic DNA (gDNA), only one sequence was amplified from cDNA. These results are best explained by the presence of nonfunctional duplications of a portion of the mtDNA, probably located in the nuclear genome, since transfer into the nuclear gene would render the COI gene nonfunctional due to differences in the nuclear and mitochondrial genetic codes. Analysis of codon variation suggests that there have been 21 independent transfer events in the 10 species examined. Within a single animal, differences between the sequences of these pseudogenes ranged from 0.2% to 20.6%, and those between the real mtDNA and pseudogene sequences ranged from 0.2% to 18.8% (uncorrected). The large number of integration events and the large range of divergences between pseudogenes and mtDNA sequences suggest that genetic material has been repeatedly transferred from the mtDNA to the nuclear genome of snapping shrimp. Unrecognized pseudogenes in phylogenetic or population studies may result in spurious results, although previous estimates of rates of molecular evolution based on Alpheus sister taxa separated by the Isthmus of Panama appear to remain valid. Especially worrisome for researchers are those pseudogenes that are not obviously recognizable as such. An effective solution may be to amplify transcribed copies of protein-coding mitochondrial genes from cDNA rather than using genomic DNA.

Animals↗

Origin and evolution of HLA class I pseudogenes.

The HLA (human major histocompatibility complex, or MHC) region includes three types of class I MHC genes: (1) class Ia loci (HLA-A, -B, and -C), which are highly expressed and polymorphic; (2) class Ib loci, which have much reduced expression and polymorphism; and (3) unexpressed class I pseudogenes. Phylogenetic analysis suggests that both class Ib loci and class I pseudogenes arose by duplication of class Ia loci. 5' regulatory elements conserved in class Ia genes are not conserved in either class Ib or class I pseudogenes, but the former show evidence of purifying selection at nonsynonymous sites in coding regions that is lacking in the latter. In the HLA class I region, there is little correspondence between map distance and phylogenetic distance, suggesting a complex history of tandem gene duplications. Furthermore, separate phylogenetic analysis of different gene regions suggests that several class I genes are evolutionary chaemeras, which presumably have arisen as a result of a process of interlocus recombination. For example, the 5' flanking region of the HLA-92 pseudogene was donated by a gene related to HLA-B and -C; and the HLA-A gene arose when exons 1-3 (and intervening introns) were donated to a gene related to the HLA-70 pseudogene by a gene related to HLA-B and -C.

Animals↗

Inactivation of human keratin genes: the spectrum of mutations in the sequence of an acidic keratin pseudogene.

Keratins are cytoskeletal proteins encoded by a multigene family. We have identified the first human keratin pseudogene and determined its complete nucleotide sequence. Sequence comparisons indicate that the pseudogene arose from a very recent duplication of the 50-kd keratin (K14) gene. The coding and the intron sequences of the two genes are 95% and 93% identical, respectively. Although the sequence of the regulatory region in the pseudogene is virtually identical to that in the 50-kd functional gene, several deleterious mutations have been identified in the pseudogene. There are three frameshifts in the coding regions, one of which is a perfect 8-bp duplication. A single-base-pair deletion in the first exon and a single-base-pair insertion in the penultimate exon also result in frameshifts. The three remaining deleterious mutations interfere with the mRNA processing signals: two alter the intron/exon boundaries, and the third disrupts the polyadenylation signal. These mutations clearly identify the sequence as a human keratin pseudogene.

Base Sequence↗

Processed pseudogenes in Drosophila.

Two species of Drosophila, D. yakuba and D. teissieri, possess pseudogenes of Adh. These pseudogenes lack introns and map to chromosome arm 3R, rather than to chromosome arm 2L, wherein are located the functional Adh genes. Their structure suggests that the pseudogenes arose from reverse transcripts. Because the pseudogenes map to homologous sites in both species, they presumably arose before these species diverged. Remarkably, the pattern of base substitution in the pseudogenes differs between sites that correspond to degenerate and non-degenerate codon positions in their functional paralogs.

Alcohol Dehydrogenase↗

Pattern of organization of human mitochondrial pseudogenes in the nuclear genome.

Mitochondrial pseudogenes in the human nuclear genome have been previously described, mostly as a source of artifacts during the analysis of the mitochondrial genome. With the availability of the complete human genome sequence, we performed a comprehensive analysis of mtDNA insertions into the nucleus. We found 612 independent integrations that are evenly distributed among all chromosomes as well as within each individual chromosome. The identified pseudogenes account for a content of at least 0.016% of the human nuclear DNA. Up to 30% of a chromosome's mtDNA pseudogene content is composed of fragments that encompass two or more adjacent mitochondrial genes, and we found no correlation between the abundance of mitochondrial transcripts and the multiplicity of integrations. These observations indicate that the migrations of mitochondrial DNA sequences to the nucleus were predominantly DNA mediated. Phylogenetic analysis of the mtDNA pseudogenes and mtDNA sequences of primates indicate a continuous transfer into the nucleus. Because of the limited window of opportunity for mtDNA transfer to the germline, sperm mtDNA, which is released from degenerating mitochondria after fertilization, could be an important source of nuclear mtDNA pseudogenes.

Biological Transport↗

Secondary structure of 7SK and 7-2 small RNAs. Possible origin of some 7SK pseudogenes from cDNA formed through self-priming by 7SK RNA.

Pseudogenes having homology to small RNAs, like 7SL, 7SK, 6S, 4.5S, U1, U2, and U3 RNAs, are abundant and dispersed in the genomes of higher eukaryotes [reviewed in Weiner et al. (1986) Annu. Rev. Biochem. 55, 631-661]. To understand better the possible origin of these pseudogenes, we studied the abilities of cytoplasmic 7SL, 7SK, and nucleolar 7-2 RNAs to self-prime and result in the synthesis of cDNAs. When rat 7SK RNA was used as substrate, a 294-nucleotide-long cDNA was synthesized in vitro by reverse transcriptase, indicating that the 3' end of 7SK RNA can act in a self-priming manner to generate 7SK cDNA. When 7-2 RNA was used as a substrate, a cDNA of approximately 235 nucleotides was observed; 7SL RNA did not act as a self-primer. Earlier studies have shown that DNAs homologous to 7SK RNA are represented by a moderately reiterated family in the mammalian genomes and many of these sequences were found to be truncated 7SK pseudogenes [Murphy et al. (1984) J. Mol. Biol. 177, 575-590]. In this study, one 7SK clone from the rat genome was characterized by sequencing. This clone contained 243 base pairs homologous to the 5' end of 7SK RNA, and was flanked by direct repeats. These data suggest that, as previously proposed for some U3 pseudogenes [Bernstein et al. (1983) Cell 32, 461-472], one mechanism for the generation of truncated 7SK pseudogenes may be the integration of self-primed reverse transcripts of 7SK RNA at random genomic sites.

Animals↗

Isolation, characterization and structural organization of the gene and pseudogene for the dihydrolipoamide succinyltransferase component of the human 2-oxoglutarate dehydrogenase complex.

In the present study, the dihydrolipoamide succinyltransferase gene of the 2-oxoglutarate dehydrogenase complex was isolated from a human genomic DNA library and its entire nucleotide sequence was determined. This gene was approximately 23 kbp in size with 15 exons and 14 introns. All of the donor and acceptor splice sites of this gene conformed to the GT/AG rule. A guanine residue 43 bases upstreams of the ATG initiating translation codon was the transcription initiation site of the human dihydrolipoamide succinyltransferase mRNA. Sequence analysis of the promoter-regulatory region showed the presence of a CAAT-box-like sequence but the presence of a TATA-box-like sequence was not evidenced. Also located in this region were sequences resembling glucocorticoid-responsive and cAMP-responsive elements, and an Sp1 binding site. No nucleotide sequence corresponding to the E3-binding and/or E1-binding domain was found in any region of the gene. Therefore, the exon coding for the E3-binding and/or E1-binding domain may have been lost from the gene during evolution. Moreover, a processed pseudogene of dihydrolipoamide succinyltransferase was isolated and sequenced. The nucleotide sequence of the pseudogene is 93% similar to the sequence of the human dihydrolipoamide succinyltransferase cDNA, but the pseudogene is not functional for base changes, deletions and insertions of the pseudogene. Southern-blot analysis showed the presence of a single copy of this gene and a single copy of a pseudogene in the human genome. In addition, a possible relationship between dihydrolipoamide succinyltransferase and familial Alzheimer's disease is discussed.

Acyltransferases↗

Germ line maintenance of the pseudogene donor pool for somatic immunoglobulin gene conversion in chickens.

Somatic immunoglobulin diversity is generated in avian species by sequential gene conversion of variable (V) gene segments of the immunoglobulin heavy- and light-chain loci during B-cell development. The germ line pools of donor sequence information for somatic V-region gene conversion are found in families of V pseudogenes, located 5' of the single functional V gene of each locus. The sequence relationships among the pseudogenes (psi VL) and functional VL1 gene of the chicken light-chain alleles in three inbred strains were compared to determine the extent of diversity within the germ line pseudogene cluster. Numerous differences were observed. For example, compared with the previously reported CB allele and the G4 allele, the S3 allele contains two intact pseudogenes between psi VL16 and psi VL18. These two adjacent psi VL gene segments (psi VL17a and psi VL17b) could have given rise to the psi VL17 segment of the G4 and CB alleles by homologous recombination. The majority of other sequence polymorphisms among the psi VL alleles appear to be the result of meiotic gene conversion. The incidence of untemplated mutations within psi VL segments is significantly lower than the incidence of mutation within the pseudogene flanking regions. Together with the observations that most psi VL segments have open reading frames and lack stop codons, these data support the hypothesis that the psi VL cluster resembles a functional multigene family maintained by evolutionary selection for its functional role in generating somatic antibody diversity. Meiotic gene conversion events within the psi VL cluster serve both to introduce diversity by the exchange of short segments between family members and to prevent the accumulation of random mutations.

Alleles↗