Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “pseudogene”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Identification and characterization of over 100 mitochondrial ribosomal protein pseudogenes in the human genome.

The human (nuclear) genome encodes at least 79 mitochondrial ribosomal proteins (MRPs), which are imported into the mitochondria. Using a comprehensive approach, we find 41 of these give rise to 120 pseudogenes in the genome. The majority of the MRP pseudogenes are of processed origin and can be aligned to match the entire coding region of the functional MRP mRNAs. One processed pseudogene was found to have originated from an alternatively spliced mRNA transcript. We also found two duplicated pseudogenes that are transcribed in the cell as confirmed by screening the human EST database. We observed a significant correlation between the number of processed pseudogenes and the gene CDS length (R = -0.40; p < 0.001), i.e., the relatively shorter genes tend to have more processed pseudogenes. There is also a weaker correlation between the number of processed pseudogenes and the gene CDS GC content. Our study provides a catalogue of human MRP pseudogenes, which will be useful in the study of functional MRP genes. It also provides a molecular record of the evolution of these genes. More details are available at http://pseudogene.org/.

Animals↗

The expression of pseudogene cyclin D2 mRNA in the human ovary may be a novel marker for decreased ovarian function associated with the aging process.

PURPOSE: Our purpose was to investigate the expression pattern of cyclin D2 and pseudogene cyclin D2 mRNA in the human ovary with age. METHODS: After extraction of the total RNA from ovarian tissues of 23 premenopausal patients, cyclin D2 and pseudogene cyclin D2 mRNAs were measured by the reverse transcription-polymerase chain reaction technique using cyclin D2 and pseudogene cyclin D2 specific primers. Analysis of the cyclin D2 and pseudogene cyclin D2 mRNA expression pattern with age and correlation analysis were carried out. RESULTS: A 489-bp cyclin D2 band and a 441-bp pseudogene cycin D2 mRNA band were detected in the human ovarian tissue. While cyclin D2 mRNA expression showed a decreasing tendency with age (P = 0.17), pseudogene cyclin D2 mRNA expression increased with age (P < 0.05). Pseudogene cyclin D2 mRNA expression showed a negative correlation with cyclin D2 mRNA (R = -0.35, P < 0.03). CONCLUSION: The expression of pseudogene cyclin D2 mRNA in the human ovary increases with age, which may be a novel marker for decreased ovarian function associated with the aging process.

Adult↗

Evidence for the presence of HABP1 pseudogene in multiple locations of mammalian genome.

The gene encoding hyaluronan-binding protein 1 (HABP1) is expressed ubiquitously in different rat tissues, and is present in eukaryotic species from yeast to humans. Fluorescence in situ hybridization indicates that this is localized in human chromosome 17p13.3. Here, we report the presence of homologous sequences of HABP1 cDNA, termed processed HABP1 pseudogene in humans. This is concluded from an additional PCR product of ~0.5 kb, along with the expected band at approximately 5 kb as observed by PCR amplification of human genomic DNA with HABP1-specific primers. Partial sequencing of the 5-kb PCR product and comparison of the HABP1 cDNA with the sequence obtained from Genbank accession number AC004148 indicated that the HABP1 gene is comprised of six exons and five introns. The 0.5-kb additional PCR product was confirmed to be homologous to HABP1 cDNA by southern hybridization, sequencing, and by a sequence homology search. Search analysis with HABP1 cDNA sequence further revealed the presence of similar sequence in chromosomes 21 and 11, which could generate ~0.5 kb with the primers used. In this report, we describe the presence of several copies of the pseudogene of HABP1 spread over different chromosomes that vary in length and similarity to the HABP1 cDNA sequence. These are 1013 bp in chromosome 21 with 85.4% similarity, 1071 bp in chromosome 11 with 87.2% similarity, 818 bp in chromosome 15 with 82.3% similarity, and 323 bp in chromosome 4 with 84% similarity to HABP1 cDNA. We have also identified similar HABP1 pseudogenes in the rat and mouse genome. The human pseudogene sequence of HABP1 possesses a 10 base pair direct repeat of "AGAAAAATAA" in chromosome 21, a 12-bp direct repeat of "AG/CAAATTA/CAA/TTA" in chromosome 4, a 8-bp direct repeat of "ACAAAG/TCT" in chromosome 15. In the case of chromosome 11, there is an inverted repeat of "AGCCTGGGCGACAGAGCGAGA" ~50 bp upstream of the HABP1 pseudogene sequence. All of the HABP1 pseudogene sequences lack 5' promoter sequence and possess multiple mutations leading to the insertion of premature stop codons in all three reading frames. Rat and mouse homologs of the HABP1 pseudogene also contain multiple mutations, leading to the insertion of premature stop codons confirming the identity of a processed pseudogene.

Animals↗

Pseudogene evolution and natural selection for a compact genome.

Pseudogenes are nonfunctional copies of protein-coding genes that are presumed to evolve without selective constraints on their coding function. They are of considerable utility in evolutionary genetics because, in the absence of selection, different types of mutations in pseudogenes should have equal probabilities of fixation. This theoretical inference justifies the estimation of patterns of spontaneous mutation from the analysis of patterns of substitutions in pseudogenes. Although it is possible to test whether pseudogene sequences evolve without constraints for their protein-coding function, it is much more difficult to ascertain whether pseudogenes may affect fitness in ways unrelated to their nucleotide sequence. Consider the possibility that a pseudogene affects fitness merely by increasing genome size. If a larger genome is deleterious--for example, because of increased energetic costs associated with genome replication and maintenance--then deletions, which decrease the length of a pseudogene, should be selectively advantageous relative to insertions or nucleotide substitutions. In this article we examine the implications of selection for genome size relative to small (1-400 bp) deletions, in light of empirical evidence pertaining to the size distribution of deletions observed in Drosophila and mammalian pseudogenes. There is a large difference in the deletion spectra between these organisms. We argue that this difference cannot easily be attributed to selection for overall genome size, since the magnitude of selection is unlikely to be strong enough to significantly affect the probability of fixation of small deletions in Drosophila.

Animals↗

Processed pseudogenes are more abundant in human and mouse X chromosomes than in autosomes.

Two different hypotheses have been proposed to explain the observation that some genomes contain more processed pseudogenes than others. One predicts that processed pseudogene abundance is inversely proportional to the substrate specificity of the reverse transcriptase that generates processed pseudogenes. The other predicts that the amount of processed pseudogenes found in genomes is proportional to the length of oogenesis. Here, we test the oogenesis hypothesis by analyzing the data from 6 studies that described the number of pseudogenes on different chromosomes of the human and/or mouse genomes. Our results show a significant overabundance of processed pseudogenes in the X chromosomes and a significant underrepresentation of processed pseudogenes in the Y chromosome of the human genome. These observations support the hypothesis that the number of processed pseudogenes is proportional to the length of oogenesis.

Animals↗

Millions of years of evolution preserved: a comprehensive catalog of the processed pseudogenes in the human genome.

Processed pseudogenes were created by reverse-transcription of mRNAs; they provide snapshots of ancient genes existing millions of years ago in the genome. To find them in the present-day human, we developed a pipeline using features such as intron-absence, frame-disruption, polyadenylation, and truncation. This has enabled us to identify in recent genome drafts approximately 8000 processed pseudogenes (distributed from http://pseudogene.org). Overall, processed pseudogenes are very similar to their closest corresponding human gene, being 94% complete in coding regions, with sequence similarity of 75% for amino acids and 86% for nucleotides. Their chromosomal distribution appears random and dispersed, with the numbers on chromosomes proportional to length, suggesting sustained "bombardment" over evolution. However, it does vary with GC-content: Processed pseudogenes occur mostly in intermediate GC-content regions. This is similar to Alus but contrasts with functional genes and L1-repeats. Pseudogenes, moreover, have age profiles similar to Alus. The number of pseudogenes associated with a given gene follows a power-law relationship, with a few genes giving rise to many pseudogenes and most giving rise to few. The prevalence of processed pseudogenes agrees well with germ-line gene expression. Highly expressed ribosomal proteins account for approximately 20% of the total. Other notables include cyclophilin-A, keratin, GAPDH, and cytochrome c.

Animals↗

Iterative gene prediction and pseudogene removal improves genome annotation.

Correct gene prediction is impaired by the presence of processed pseudogenes: nonfunctional, intronless copies of real genes found elsewhere in the genome. Gene prediction programs frequently mistake processed pseudogenes for real genes or exons, leading to biologically irrelevant gene predictions. While methods exist to identify processed pseudogenes in genomes, no attempt has been made to integrate pseudogene removal with gene prediction, or even to provide a freestanding tool that identifies such erroneous gene predictions. We have created PPFINDER (for Processed Pseudogene finder), a program that integrates several methods of processed pseudogene finding in mammalian gene annotations. We used PPFINDER to remove pseudogenes from N-SCAN gene predictions, and show that gene prediction improves substantially when gene prediction and pseudogene masking are interleaved. In addition, we used PPFINDER with gene predictions as a parent database, eliminating the need for libraries of known genes. This allows us to run the gene prediction/PPFINDER procedure on newly sequenced genomes for which few genes are known.

Animals↗

Pseudogenes: are they "junk" or functional DNA?

Pseudogenes have been defined as nonfunctional sequences of genomic DNA originally derived from functional genes. It is therefore assumed that all pseudogene mutations are selectively neutral and have equal probability to become fixed in the population. Rather, pseudogenes that have been suitably investigated often exhibit functional roles, such as gene expression, gene regulation, generation of genetic (antibody, antigenic, and other) diversity. Pseudogenes are involved in gene conversion or recombination with functional genes. Pseudogenes exhibit evolutionary conservation of gene sequence, reduced nucleotide variability, excess synonymous over nonsynonymous nucleotide polymorphism, and other features that are expected in genes or DNA sequences that have functional roles. We first review the Drosophila literature and then extend the discussion to the various functional features identified in the pseudogenes of other organisms. A pseudogene that has arisen by duplication or retroposition may, at first, not be subject to natural selection if the source gene remains functional. Mutant alleles that incorporate new functions may, nevertheless, be favored by natural selection and will have enhanced probability of becoming fixed in the population. We agree with the proposal that pseudogenes be considered as potogenes, i.e., DNA sequences with a potentiality for becoming new genes.

Animals↗

Pervasive survival of expressed mitochondrial rps14 pseudogenes in grasses and their relatives for 80 million years following three functional transfers to the nucleus.

BACKGROUND: Many mitochondrial genes, especially ribosomal protein genes, have been frequently transferred as functional entities to the nucleus during plant evolution, often by an RNA-mediated process. A notable case of transfer involves the rps14 gene of three grasses (rice, maize, and wheat), which has been relocated to the intron of the nuclear sdh2 gene and which is expressed and targeted to the mitochondrion via alternative splicing and usage of the sdh2 targeting peptide. Although this transfer occurred at least 50 million years ago, i.e., in a common ancestor of these three grasses, it is striking that expressed, nearly intact pseudogenes of rps14 are retained in the mitochondrial genomes of both rice and wheat. To determine how ancient this transfer is, the extent to which mitochondrial rps14 has been retained and is expressed in grasses, and whether other transfers of rps14 have occurred in grasses and their relatives, we investigated the structure, expression, and phylogeny of mitochondrial and nuclear rps14 genes from 32 additional genera of grasses and from 9 other members of the Poales. RESULTS: Filter hybridization experiments showed that rps14 sequences are present in the mitochondrial genomes of all examined Poales except for members of the grass subfamily Panicoideae (to which maize belongs). However, PCR amplification and sequencing revealed that the mitochondrial rps14 genes of all examined grasses (Poaceae), Cyperaceae, and Joinvilleaceae are pseudogenes, with all those from the Poaceae sharing two 4-NT frameshift deletions and all those from the Cyperaceae sharing a 5-NT insertion (only one member of the Joinvilleaceae was examined). cDNA analysis showed that all mitochondrial pseudogenes examined (from all three families) are transcribed, that most are RNA edited, and that surprisingly many of the edits are reverse (U-->C) edits. Putatively nuclear copies of rps14 were isolated from one to several members of each of these three Poales families. Multiple lines of evidence indicate that the nuclear genes are probably the products of three independent transfers. CONCLUSION: The rps14 gene has, most likely, been functionally transferred from the mitochondrion to the nucleus at least three times during the evolution of the Poales. The transfers in Cyperaceae and Poaceae are relatively ancient, occurring in the common ancestor of each family, roughly 80 million years ago, whereas the putative Joinvilleaceae transfer may be the most recent case of functional organelle-to-nucleus transfer yet described in any organism. Remarkably, nearly intact and expressed pseudogenes of rps14 have persisted in the mitochondrial genomes of most lineages of Poaceae and Cyperaceae despite the antiquity of the transfers and of the frameshift and RNA editing mutations that mark the mitochondrial genes as pseudogenes. Such long-term, nearly pervasive survival of expressed, apparent pseudogenes is to our knowledge unparalleled in any genome. Such survival probably reflects a combination of factors, including the short length of rps14, its location immediately downstream of rpl5 in most plants, and low rates of nucleotide substitutions and indels in plant mitochondrial DNAs. Their survival also raises the possibility that these rps14 sequences may not actually be pseudogenes despite their appearance as such. Overall, these findings indicate that intracellular gene transfer may occur even more frequently in angiosperms than already recognized and that pseudogenes in plant mitochondrial genomes can be surprisingly resistant to forces that lead to gene loss and inactivation.

Cell Nucleus↗

Genome-wide survey for biologically functional pseudogenes.

According to current estimates there exist about 20,000 pseudogenes in a mammalian genome. The vast majority of these are disabled and nonfunctional copies of protein-coding genes which, therefore, evolve neutrally. Recent findings that a Makorin1 pseudogene, residing on mouse Chromosome 5, is, indeed, in vivo vital and also evolutionarily preserved, encouraged us to conduct a genome-wide survey for other functional pseudogenes in human, mouse, and chimpanzee. We identify to our knowledge the first examples of conserved pseudogenes common to human and mouse, originating from one duplication predating the human-mouse species split and having evolved as pseudogenes since the species split. Functionality is one possible way to explain the apparently contradictory properties of such pseudogene pairs, i.e., high conservation and ancient origin. The hypothesis of functionality is tested by comparing expression evidence and synteny of the candidates with proper test sets. The tests suggest potential biological function. Our candidate set includes a small set of long-lived pseudogenes whose unknown potential function is retained since before the human-mouse species split, and also a larger group of primate-specific ones found from human-chimpanzee searches. Two processed sequences are notable, their conservation since the human-mouse split being as high as most protein-coding genes; one is derived from the protein Ataxin 7-like 3 (ATX7NL3), and one from the Spinocerebellar ataxia type 1 protein (ATX1). Our approach is comparative and can be applied to any pair of species. It is implemented by a semi-automated pipeline based on cross-species BLAST comparisons and maximum-likelihood phylogeny estimations. To separate pseudogenes from protein-coding genes, we use standard methods, utilizing in-frame disablements, as well as a probabilistic filter based on Ka/Ks ratios.

Animals↗

[son Pseudogenes do not contain five repeating elements of the region of complete tandem repeats present in the homologous sequence of the son gene].

Recently we have published a sequence of the coding region of the son gene, containing at least six areas of the tandem repeats [V.V. Bliskovsky, F.B. Berdichevsky, A.V. Tkachenko, M.E. Belova, I.M. Chumakov--Molecularnaya biologiya, 1992, V. 26. P. 793-806; V.V. Bliskovsky, A.V. Kirillov, V.M. Zacharyev, I.M. Chumakov--Molecularnaya biologiya, 1992, V. 26. P. 807-812]. The presence of several areas of tandem repeats with different nucleotide sequences of the repeated elements within one and the same gene supports the proposition that genomic localization of the sequence influences its duplication. Here we present a nucleotide sequence of the son pseudogene isolated as the result of hybridization screening of a human genomic library using the sequence of the son gene as a probe. The comparison of the gene and the pseudogene nucleotide sequences shows that the sequence of pseudogene does not contain five repeated units of an area of the tandem repeats, which are presented in the homology sequence of the son gene. Because the pseudogene was probably generated by the reverse son-gene transcript insertion to the genome, and so the nucleotide sequences of the coding region of the gene and the sequence of the pseudogene were identical at this moment, the differences between the gene and the pseudogene are the results of their evolution after the generation of the pseudogene. Possible factors, influenced the son gene and the son pseudogene evolution are discussed.

Base Sequence↗

Isolation and characterization of a processed pseudogene for murine cyclin D3.

By using radiolabeled murine cyclin D3 cDNA as a probe, two cyclin D3 genomic clones, MCD3P-117 and MCD3P-327, were isolated from a murine genomic library constructed with murine liver DNA. Physical mapping and DNA sequence analysis revealed that these clones contain approximately 1.5 kb uninterrupted linear sequence similar to murine cyclin D3 cDNA, indicating that the 1.5 kb sequence is a processed pseudogene for cyclin D3. When the nucleotide sequence of the cyclin D3 pseudogene was compared with that of cyclin D3 cDNA at the nucleotide level, the pseudogene contained 229 bp of 5'- and 371 bp of 3'-untranslated regions, and a recognizable complete coding region that is 90% identical to murine cyclin D3. This sequence is bounded by the repeat sequence (GC/AGCTCTCC), which is common to many processed pseudogenes. However, multiple genetic lesions, including substitution, deletion and/or insertion events that result in modification of the reading frame were found in the pseudogene sequence. The pseudogene appeared to accumulate 67 random point mutations in the functional coding region composed of 879 nucleotide positions. It is thus estimated that the cyclin D3 pseudogene arose approximately 11 million years (Myr) ago. These data provide the first characterization of murine cyclin D3 pseudogene and insight into its evolutionary age.

Amino Acid Sequence↗

Transcription factor binding is limited by the 5'-flanking regions of a Drosophila tRNAHis gene and a tRNAHis pseudogene.

We determined the sequence of a Drosophila tRNA gene cluster containing a tRNAHis gene and a tRNAHis pseudogene in close proximity on the same DNA strand. The pseudogene contains eight consecutive base pairs different from the region of the bona fide gene which codes for the 3' portion of the anticodon stem of tRNAHis. The tRNAHis gene is transcribed efficiently in Drosophila Kc cell extract, whereas the pseudogene is not. The pseudogene is also a much poorer competitor than the real gene in a stable transcription complex formation assay, even though the sequence alteration in the pseudogene does not affect the sequence or spacing of the putative internal transcription control regions. Recombinant clones were constructed in which the 5'-flanking regions are exchanged. The transcription efficiencies and competitive abilities of the recombinant clones resemble those of the genes from which the 5' flank was derived; for example, the tRNAHis pseudogene with the 5'-flanking sequence of the tRNAHis gene is now efficiently transcribed. Deletion analysis of the pseudogene 5' flank failed to uncover an inhibitory element. Deletion analysis of the real gene showed very high dependence on the presence of the wild-type 5'-flanking sequence for factor binding to the internal control regions and stable complex formation. The 5'-flanking sequence of a Drosophila tRNAArg gene active in the Drosophila Kc cell extract does not restore transcriptional activity or stable complex formation. The tRNAHis gene and pseudogene behave atypically in HeLa cell extract. Both genes compete for HeLa transcription factors, but neither of them is efficiently transcribed. Removal of the 5'-flanking sequences of each gene and replacement with various sequences, including the tRNAArg gene 5' flank, does not allow increased transcription in HeLa cell extract.

Animals↗

Deletions in processed pseudogenes accumulate faster in rodents than in humans.

The relative rates of point nucleotide substitution and accumulation of gap events (deletions and insertions) were calculated for 22 human and 30 rodent processed pseudogenes. Deletion events not only outnumbered insertions (the ratio being 7:1 and 3:1 for human and rodent pseudogenes, respectively), but also the total length of deletions was greater than that of insertions. Compared with their functional homologs, human processed pseudogenes were found to be shorter by about 1.2%, and rodent pseudogenes by about 2.3%. DNA loss from processed pseudogenes through deletion is estimated to be at least seven times faster in rodents than in humans. In comparison with the rate of point substitutions, the abridgment of pseudogenes during evolutionary times is a slow process that probably does not retard the rate of growth of the genome due to the proliferation of processed pseudogenes.

Animals↗

Age and detection of retroprocessed pseudogenes in murine rodents.

Retroprocessed pseudogenes, calmodulin II (psi1, psi2, and psi3 CALMII), psi alpha-tubulin, pi-glutathione S-transferase (psi pi-GST) from rat, lactic acid dehydrogenase (psi LDH) from mouse, and heat shock protein 60 chaperonin (psi HSP60) from Chinese hamster, were examined for their presence in these species by polymerase chain reaction (PCR). Pseudogenes of these murine rodents were detected by PCR only in those species in which the genes were originally identified, suggesting that the selected pseudogene of one species arose too recently to be detected in the genomes of the other rodent species. The calculated ages of the rodent pseudogenes ranged from 1.7 Myr (psi alpha-tubulin) to 7.5 Myr (psi3 CALMII) when employing a homologous functional gene of the taxon as a reference in the relative rate test with the mouse or rat as the outgroup. Given the high rate of divergence of the genes of rodents relative to other species, selection of an outgroup with similar mutation rates seems warranted. To justify further the conclusion that the selected pseudogenes were indeed retroprocessed after these three taxa diverged, the presence of the pseudogenes in the genome of different rat species was examined. The existence of psi3 CALMII and psi alpha-tubulin pseudogenes of Rattus norvegicus among species belonging to Rattus sensu stricto is evidence for the common ancestry of this group.

Animals↗

Multiple independent pseudogene derivations indicate increased instability of the Mdm2 locus in Mus caroli.

Under conditions of genomic stress, the Mdm locus in human and in mouse is prone to instability manifested as amplification and oncogenesis. The Mdm2 gene is a known oncogene that is amplified in approximately one-third of sarcomas and whose protein product interacts with the tumor suppressor p53. Concimitant with such gene amplification events is the activation and mobilization of endogenous retroelements, typically through the relaxation of epigenetic controlling mechanisms. Processed pseudogenes, which can be formed through endogenous LINE retroelement activity, may indicate increased genomic instability. We have isolated processed pseudogenes for Mdm2 in Mus caroli DNA, likely formed from independent events in different individuals. This is the first identification and characterization of an Mdm2 pseudogene in any organism. Multiple retrotransposition events are suggested by the variable sequence and genomic structure of the identified pseudogenes across all exons and the 3'UTR. The high degree of similarity between the gene and each pseudogene, as well as the lack of evidence for an Mdm2 pseudogene in several other species of Mus, indicate evolutionarily recent retrotransposition events leading to the formation of the Mdm2 pseudogenes in M. caroli. Previous studies on the Mdm2 locus in Mus caroli showed amplification and overexpression of this gene on double minute chromosomes in a Mus musculus x Mus caroli interspecific hybrid. The identification of an Mdm2 retropseudogene within this species further highlights the predisposition to instability for this region of the genome.

3' Untranslated Regions↗

Human lactate dehydrogenase-B processed pseudogene: nucleotide sequence analysis and assignment to the X-chromosome.

Two human genomic clones containing the lactate dehydrogenase-B processed pseudogene were isolated from two patients deficient in lactate dehydrogenase-B isozyme. The sequences of 3,287 nucleotides, including the pseudogenes and its flanking regions, from both clones were found to be identical except for three differences in the pseudogenes. The sequences of 1,286 nucleotides from these two pseudogenes exhibited 93% homology with the cDNA sequence of the lactate dehydrogenase-B functional gene, and the pseudogene contained 75/76 base substitutions, 11/12 single-base deletions, and 5 single-base insertions. This pseudogene was mapped to the x-chromosome by dot-blot analysis using a probe for the pseudogene or its 5' flanking sequence.

Base Sequence↗

The processed pseudogene of mouse thymidine kinase is active after transfection.

Aside of the gene coding for cytoplasmic thymidine kinase, the genome of mouse cells carries two pseudogenes. Both are inactive in situ. One of the pseudogenes is a processed pseudogene in which a two base pair deletion caused a shift of the reading frame and a shortening of the gene product from the 233 amino acids of thymidine kinase to 177 amino acids in the pseudogene product. We report here that introduction of this pseudogene into LTK- cells gave rise to cells with a thymidine kinase positive phenotype. The transformed cells carried multiple copies of the pseudogene the upstream region of which exhibited low but measurable promoter activity. Replacement of the upstream region of the pseudogene by a promoter of Simian virus 40 or of the mammary tumor virus resulted in high transfection efficiencies and in cell lines exhibiting high thymidine kinase activities.

Animals↗