Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “pseudogene”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

HOPPSIGEN: a database of human and mouse processed pseudogenes.

Processed pseudogenes result from reverse transcribed mRNAs. In general, because processed pseudogenes lack promoters, they are no longer functional from the moment they are inserted into the genome. Subsequently, they freely accumulate substitutions, insertions and deletions. Moreover, the ancestral structure of processed pseudogenes could be easily inferred using the sequence of their functional homologous genes. Owing to these characteristics, processed pseudogenes represent good neutral markers for studying genome evolution. Recently, there is an increasing interest for these markers, particularly to help gene prediction in the field of genome annotation, functional genomics and genome evolution analysis (patterns of substitution). For these reasons, we have developed a method to annotate processed pseudogenes in complete genomes. To make them useful to different fields of research, we stored them in a nucleic acid database after having annotated them. In this work, we screened both mouse and human complete genomes from ENSEMBL to find processed pseudogenes generated from functional genes with introns. We used a conservative method to detect processed pseudogenes in order to minimize the rate of false positive sequences. Within processed pseudogenes, some are still having a conserved open reading frame and some have overlapping gene locations. We designated as retroelements all reverse transcribed sequences and more strictly, we designated as processed pseudogenes, all retroelements not falling in the two former categories (having a conserved open reading or overlapping gene locations). We annotated 5823 retroelements (5206 processed pseudogenes) in the human genome and 3934 (3428 processed pseudogenes) in the mouse genome. Compared to previous estimations, the total number of processed pseudogenes was underestimated but the aim of this procedure was to generate a high-quality dataset. To facilitate the use of processed pseudogenes in studying genome structure and evolution, DNA sequences from processed pseudogenes, and their functional reverse transcribed homologs, are now stored in a nucleic acid database, HOPPSIGEN. HOPPSIGEN can be browsed on the PBIL (Pole Bioinformatique Lyonnais) World Wide Web server (http://pbil.univ-lyon1.fr/) or fully downloaded for local installation.

Animals↗

Structural repertoire in VH pseudogenes of immunoglobulins: comparison with human germline genes and human amino acid sequences.

In the pool of human immunoglobulin VH gene segments, pseudogenes amount to roughly 30% of the total number of genes. Some of them are highly conserved among unrelated individuals. These facts suggest a possible functional role for pseudogenes in the human immune response diversity. This paper intends to provide additional information about the structure of VH pseudogene sequences to evaluate the possible role of pseudogenes in the immune response. Mutations capable of altering framework stability in human VH pseudogenes were analyzed. Results indicate that VH pseudogenes are about 14 times as divergent as human VH functional germline genes on the one hand, and four times as divergent in the case of human VH amino acid sequences on the other. The high number of disruptive mutations in pseudogenes is an expected result because of the lack of functionality of these genes. In the second part of the work we analyze whether or not the same takes place in the positions that determine the existence of canonical structures in the hypervariable loops in VH pseudogenes. An extension of such analysis is applied to all species with reported VH pseudogenes. In contrast with results concerning framework positions, 69% of known human VH pseudogenes have canonical structures in the first hypervariable loop, while 48% do so in the second one. Comparison of these results with those found in human VH functional germline genes and human VH amino acid sequences shows that in the former as many as 100% and in the latter 96% have canonical structures. In VH amino acid sequences the result is similar to pseudogenes for H1. For H2, such value lies between the percentage of germline genes (96%) and the percentage of pseudogenes (48%). The possible significance of the existence of canonical structures in the hypervariable loops of VH pseudogenes is discussed.

Amino Acid Sequence↗

Human von Willebrand factor gene and pseudogene: structural analysis and differentiation by polymerase chain reaction.

Structural analysis of the von Willebrand factor gene located on chromosome 12 is complicated by the presence of a partial unprocessed pseudogene on chromosome 22q11-13. The structures of the von Willebrand factor pseudogene and corresponding segment of the gene were determined, and methods were developed for the rapid differentiation of von Willebrand factor gene and pseudogene sequences. The pseudogene is 21-29 kilobases in length and corresponds to 12 exons (exons 23-34) of the von Willebrand factor gene. Approximately 21 kilobases of the gene and pseudogene were sequenced, including the 5' boundary of the pseudogene. The 3' boundary of the pseudogene lies within an 8-kb region corresponding to intron 34 of the gene. The presence of splice site and nonsense mutations suggests that the pseudogene cannot yield functional transcripts. The pseudogene has diverged approximately 3.1% in nucleotide sequence from the gene. This suggests a recent evolutionary origin approximately 19-29 million years ago, near the time of divergence of humans and apes from monkeys. Several repetitive sequences were identified, including 4 Alu, one Line-1, and several short simple sequence repeats. Several of these simple repeats differ in length between the gene and pseudogene and provide useful markers for distinguishing these loci. Sequence differences between the gene and pseudogene were exploited to design oligonucleotide primers for use in the polymerase chain reaction to selectivity amplify sequences corresponding to exons 23-34 from either the von Willebrand factor gene or the pseudogene. This method is useful for the analysis of gene defects in patients with von Willebrand disease, without interference from homologous sequences in the pseudogene.

Amino Acid Sequence↗

A maximum likelihood method for analyzing pseudogene evolution: implications for silent site evolution in humans and rodents.

We present a new likelihood method for detecting constrained evolution at synonymous sites and other forms of nonneutral evolution in putative pseudogenes. The model is applicable whenever the DNA sequence is available from a protein-coding functional gene, a pseudogene derived from the protein-coding gene, and an orthologous functional copy of the gene. Two nested likelihood ratio tests are developed to test the hypotheses that (1) the putative pseudogene has equal rates of silent and replacement substitutions; and (2) the rate of synonymous substitution in the functional gene equals the rate of substitution in the pseudogene. The method is applied to a data set containing 74 human processed-pseudogene loci, 25 mouse processed-pseudogene loci, and 22 rat processed-pseudogene loci. Using the informatics resources of the Human Genome Project, we localized 67 of the human-pseudogene pairs in the genome and estimated the GC content of a large surrounding genomic region for each. We find that, for pseudogenes deposited in GC regions similar to those of their paralogs, the assumption of equal rates of silent and replacement site evolution in the pseudogene is upheld; in these cases, the rate of silent site evolution in the functional genes is approximately 70% the rate of evolution in the pseudogene. On the other hand, for pseudogenes located in genomic regions of much lower GC than their functional gene, we see a sharp increase in the rate of silent site substitutions, leading to a large rate of rejection for the pseudogene equality likelihood ratio test.

Animals↗

Identification and analysis of over 2000 ribosomal protein pseudogenes in the human genome.

Mammals have 79 ribosomal proteins (RP). Using a systematic procedure based on sequence-homology, we have comprehensively identified pseudogenes of these proteins in the human genome. Our assignments are available at http://www.pseudogene.org or http://bioinfo.mbb.yale.edu/genome/pseudogene. In total, we found 2090 processed pseudogenes and 16 duplications of RP genes. In relation to the matching parent protein, each of the processed pseudogenes has an average relative sequence length of 97% and an average sequence identity of 76%. A small number (258) of them do not contain obvious disablements (stop codons or frameshifts) and, therefore, could be mistaken as functional genes, and 178 are disrupted by one or more repetitive elements. On average, processed pseudogenes have a longer truncation at the 5' end than the 3' end, consistent with the target-primed-reverse-transcription (TPRT) mechanism. Interestingly, on chromosome 16, an RPL26 processed pseudogene was found in the intron region of a functional RPS2 gene. The large-scale distribution of RP pseudogenes throughout the genome appears to result, chiefly, from random insertions with the numbers on each chromosome, consequently, proportional to its size. In contrast to RP genes, the RP pseudogenes have the highest density in GC-intermediate regions (41%-46%) of the genome, with the density pattern being between that of LINEs and Alus. This can be explained by a negative selection theory as we observed that GC-rich RP pseudogenes decay faster in GC-poor regions. Also, we observed a correlation between the number of processed pseudogenes and the GC content of the associated functional gene, i.e., relatively GC-poor RPs have more processed pseudogenes. This ranges from 145 pseudogenes for RPL21 down to 3 pseudogenes for RPL14. We were able to date the RP pseudogenes based on their sequence divergence from present-day RP genes, finding an age distribution similar to that for Alus. The distribution is consistent with a decline in retrotransposition activity in the hominid lineage during the last 40 Myr. We discuss the implications for retrotransposon stability and genome dynamics based on these new findings.

Amino Acid Sequence↗

Integrated pseudogene annotation for human chromosome 22: evidence for transcription.

Pseudogenes are inheritable genetic elements formally defined by two properties: their similarity to functioning genes and their presumed lack of activity. However, their precise characterization, particularly with respect to the latter quality, has proven elusive. An opportunity to explore this issue arises from the recent emergence of tiling-microarray data showing that intergenic regions (containing pseudogenes) are transcribed to a great degree. Here we focus on the transcriptional activity of pseudogenes on human chromosome 22. First, we integrated several sets of annotation to define a unified list of 525 pseudogenes on the chromosome. To characterize these further, we developed a comprehensive list of genomic features based on conservation in related organisms, expression evidence, and the presence of upstream regulatory sites. Of the 525 unified pseudogenes we could confidently classify 154 as processed and 49 as duplicated. Using data from tiling microarrays, especially from recent high-resolution oligonucleotide arrays, we found some evidence that up to a fifth of the 525 pseudogenes are potentially transcribed. Expressed sequence tags (EST) comparison further validated a number of these, and overall we found 17 pseudogenes with strong support for transcription. In particular, one of the pseudogenes with both EST and microarray evidence for transcription turned out to be a duplicated pseudogene in the cat eye syndrome critical region. Although we could not identify a meaningful number of transcription factor-binding sites (based on chromatin immunoprecipitation-chip data) near pseudogenes, we did find that approximately 12% of the pseudogenes had upstream CpG islands. Finally, analysis of corresponding syntenic regions in the mouse, rat and chimp genomes indicates, as previously suggested, that pseudogenes are less conserved than genes, but more preserved than the intergenic background (all notation is available from http://www.pseudogene.org).

Animals↗

An unusual adenine phosphoribosyltransferase pseudogene is syntenic with its functional gene and is flanked by highly polymorphic DNAs.

A mouse adenine phosphoribosyltransferase (aprt) pseudogene that had previously been recovered from a BALB/c sperm DNA library possessed several unusual features. Its nucleotide sequence, like that of other processed pseudogenes, was colinear with its corresponding mRNA, but it was truncated at its 3' end and lacked a poly(A) tail. The pseudogene was 82% homologous with corresponding regions of the functional gene and had incurred mutations that included transitions, transversions, deletions, and a point insertion. Even though the pseudogene was truncated within the protein-coding region of the corresponding functional gene, it was flanked at both ends by 13-base-pair direct repeats. Curiously, the direct repeats exhibited homology to APRT mRNA at the site of pseudogene divergence. The pseudogene appeared to be common to BALB/c and A/J mice, but it was contained on a 3-kilobase EcoRI fragment in the former strain and a 4.5-kilobase EcoRI fragment in the latter. The BALB/c and apparently the A/J pseudogene both mapped to chromosome 8, which also contains the functional aprt gene. The DNA sequences immediately surrounding the pseudogene in the two strains appeared to be similar, suggesting that the BALB/c and A/J pseudogenes are allelic. However, DNA sequences more distal to the pseudogene in the two strains appeared to vary. Thus, the EcoRI polymorphism was not due to simple loss of an EcoRI site, but was more complex. The pattern of flanking restriction sites was different for each of several enzymes, consistent with extensive DNA rearrangement. Double digests of BALB/c and A/J genomic DNAs revealed complex polymorphisms on both sides of the pseudogene. The results were consistent with insertion, deletion, or other rearrangement of DNA sequences that flank the pseudogene and suggest that this region of mouse chromosome 8 may be a region active for mutation or recombination.

Adenine Phosphoribosyltransferase↗

Cloning, sequencing, and chromosomal localization of two tandemly arranged human pseudogenes for the proliferating cell nuclear antigen (PCNA).

We have characterized a human genomic clone carrying two pseudogenes for the proliferating cell nuclear antigen (PCNA), which were tandemly arranged on human Chromosome (Chr) 4. One is a processed pseudogene that showed a 73% nucleotide homology to the human PCNA cDNA and possessed none of the introns existing in the functional PCNA gene. This pseudogene presumably arose by reverse transcription of a PCNA mRNA followed by integration of the cDNA into the genome. The other is a 5' and 3' truncated pseudogene that showed a nucleotide homology to a 3' region of the exon 4 and to a 5' region of the exon 5 of the PCNA gene and did not have the intronic sequence between the exons 4 and 5. Both pseudogenes had the same nucleotide deletion as compared with the human functional PCNA gene. A phylogenetic analysis of PCNA gene family, including the functional PCNA gene and another PCNA pseudogene located on a different chromosome, revealed that the truncated pseudogene exhibits the closest evolutionary relationship with the processed pseudogene, suggesting that the truncated pseudogene was generated by duplication of the processed pseudogene after translocation to Chr 4. Furthermore, fluorescence in situ hybridization revealed that these pseudogenes are located on the long arm of Chr 4, 4q24.

Animals↗

Mechanism of spreading of the highly related neurofibromatosis type 1 (NF1) pseudogenes on chromosomes 2, 14 and 22.

Neurofibromatosis type 1 (NF1) is a frequent hereditary disorder that involves tissues derived from the embryonic neural crest. Besides the functional gene on chromosome arm 17q, NF1-related sequences (pseudogenes) are present on a number of chromosomes including 2, 12, 14, 15, 18, 21, and 22. We elucidated the complete nucleotide sequence of the NF1 pseudogene on chromosome 22. Only the middle part of the functional gene but not exons 21-27a, encoding the functionally important GAP-related domain of the NF1 protein, is presented in this pseudogene. In addition to the two known NF1 pseudogenes on chromosome 14 we identified two novel variants. A phylogenetic tree was constructed, from which we concluded that the NF1 pseudogenes on chromosomes 2, 14, and 22 are closely related to each other. Clones containing one of these pseudogenes cross-hybridised with the other pseudogenes in this subset, but did not reveal any in situ hybridisation with the functional NF1 gene or with NF1 pseudogenes on other chromosomes. This suggests that their hybridisation specificity is mainly determined by homologous sequences flanking the pseudogenes. Strong support for this concept was obtained by sequence analysis of the flanking regions, which revealed more than 95% homology. We hypothesise that during evolution this subset of NF1 pseudogenes initially arose by duplication and transposition of the middle part of the functional NF1 gene to chromosome 2. Subsequently, a much larger fragment, including flanking sequences, was duplicated and gave rise to the current NF1 pseudogene copies on chromosomes 14 and 22.

Base Sequence↗

Comprehensive analysis of amino acid and nucleotide composition in eukaryotic genomes, comparing genes and pseudogenes.

Based on searches for disabled homologs to known proteins, we have identified a large population of pseudogenes in four sequenced eukaryotic genomes-the worm, yeast, fly and human (chromosomes 21 and 22 only). Each of our nearly 2500 pseudogenes is characterized by one or more disablements mid-domain, such as premature stops and frameshifts. Here, we perform a comprehensive survey of the amino acid and nucleotide composition of these pseudogenes in comparison to that of functional genes and intergenic DNA. We show that pseudogenes invariably have an amino acid composition intermediate between genes and translated intergenic DNA. Although the degree of intermediacy varies among the four organisms, in all cases, it is most evident for amino acid types that differ most in occurrence between genes and intergenic regions. The same intermediacy also applies to codon frequencies, especially in the worm and human. Moreover, the intermediate composition of pseudogenes applies even though the composition of the genes in the four organisms is markedly different, showing a strong correlation with the overall A/T content of the genomic sequence. Pseudogenes can be divided into 'ancient' and 'modern' subsets, based on the level of sequence identity with their closest matching homolog (within the same genome). Modern pseudogenes usually have a much closer sequence composition to genes than ancient pseudogenes. Collectively, our results indicate that the composition of pseudogenes that are under no selective constraints progressively drifts from that of coding DNA towards non-coding DNA. Therefore, we propose that the degree to which pseudogenes approach a random sequence composition may be useful in dating different sets of pseudogenes, as well as to assess the rate at which intergenic DNA accumulates mutations. Our compositional analyses with the interactive viewer are available over the web at http://genecensus.org/pseudogene.

Amino Acid Sequence↗

Psi-Phi: exploring the outer limits of bacterial pseudogenes.

Because bacterial chromosomes are tightly packed with genes and were traditionally viewed as being optimized for size and replication speed, it was not surprising that the early annotations of sequenced bacterial genomes reported few, if any, pseudogenes. But because pseudogenes are generally recognized by comparisons with their functional counterparts, as more genome sequences accumulated, many bacterial pathogens were found to harbor large numbers of truncated, inactivated, and degraded genes. Because the mutational events that inactivate genes occur continuously in all genomes, we investigated whether the rarity of pseudogenes in some bacteria was attributable to properties inherent to the organism or to the failure to recognize pseudogenes. By developing a program suite (called Psi-Phi, for Psi-gene Finder) that applies a comparative method to identify pseudogenes (attributable both to misannotation and to nonrecognition), we analyzed the pseudogene inventories in the sequenced members of the Escherichia coli/Shigella clade. This approach recovered hundreds of previously unrecognized pseudogenes and showed that pseudogenes are a regular feature of bacterial genomes, even in those whose original annotations registered no truncated or otherwise inactivated genes. In Shigella flexneri 2a, large proportions of pseudogenes are generated by nonsense mutations and IS element insertions, events that seldom produce the pseudogenes present in the other genomes examined. Almost all (>95%) pseudogenes are restricted to only one of the genomes and are of relatively recent origin, suggesting that these bacteria possess active mechanisms to eliminate nonfunctional genes.

Codon, Nonsense↗

The expression of rac1 pseudogene in human tissues and in human brain tumors.

BACKGROUND: Recent studies have demonstrated that Rac is a regulator of cell morphology and growth. Rac1 gene appears to have involvement in tumorigenesis and metastatic potential. In our previous study of rac1 gene in 45 human brain tumors, rac1 pseudogene was found. The rac1 pseudogene is an intronless pseudogene and has a similarity of 86% with rac1 nucleotide sequence. The rac1 pseudogene contains 579 nucleotides and only 46 amino acids can be translated. Little is known about the expression of rac1 pseudogene in human tissues or tumors. MATERIALS AND METHODS: The expression of rac1 gene and rac1 pseudogene in different human tissues and brain tumors was investigated by the use of reverse transcriptase-polymerase chain reaction and Northern blotting. RESULTS: The rac1 gene is apparently expressed in these 8 human tissues. The rac1 pseudogene is also apparently expressed in human tissues except for brain tissue. The overexpression of rac1 gene in brain tumors was 8% (2/25) and the overexpression of rac1 pseudogene was 76.9% (20/26). Only two astrocytomas had overexpression of rac1 gene, compared with normal brain tissues. The overexpression of rac1 pseudogene was 6 of 9 in meningiomas, 7 of 9 in astrocytomas, and 7 of 8 in pituitary adenomas. CONCLUSIONS: High frequency of overexpression of rac1 pseudogene was detected in the human brain tumors when compared with that expressed in the normal brain tissues. Our study suggested that the rac1 pseudogene may play an important role of the tumorigenesis of brain.

Astrocytoma↗

Identification of a transcriptionally active pseudogene in the chorion locus of the silkmoth Bombyx mori. Regional sequence conservation and biological function.

We have determined the primary structure of a 3500 base-pair part of the silkmoth chorion locus mapping in a region containing genes of late developmental specificity. This part of the locus was found to harbour a pseudogene related to one of the families of chorion genes encoding high cysteine proteins, HcB. The pseudogene exhibits an overall sequence identity of 84% to the consensus coding region of HcB chorion genes. A 95% identity was observed over a length of 190 base-pairs of its immediate 5' upstream sequences and the corresponding part of the consensus 5'-intergenic sequences of Hc gene pairs, normally encompassing 270 base-pairs. Thus, the pseudogene has retained part of the promoter region that includes sequence elements whose presence is thought to be necessary for transcriptional competence of HcB genes. The pseudogene is also characterized by the elimination of part of its first exon containing most of the 5' untranslated region, the ATG translation initiation codon and part of the signal peptide sequences. Its intron is longer than that of other HcB genes due to the insertion of a copy of a repetitive element that appears to be transcribed by RNA polymerase III. A previously characterized chorion cDNA clone, m2282, representing a rare mRNA sequence of late developmental specificity, was found to be identical to the pseudogene over its entirety spanning 65% of the pseudogene's second exon. Hybridizations of clones spanning a 260,000 base-pair domain of the chorion locus of Bombyx mori and of total genomic DNA to a subfragment of the cDNA clone containing relatively unique sequences, coupled to primer extension experiments, have demonstrated that m2282 mRNA originated from the pseudogene and that the pseudogene transcripts are initiated at the chorion cap site consensus sequence. We conclude that the 5'-flanking sequences retained by the pseudogene encompass elements necessary and adequate for correct transcriptional activation, but may not include those required for quantitative expression of the promoter. Possible reasons for the observed lack of random drift in the 5'-upstream sequences of the pseudogene and the maintenance of a functional promoter in a non-functional gene are discussed on the basis of the observation that elements resembling scaffold attachment sites are present in these sequences.

Amino Acid Sequence↗

Studying genomes through the aeons: protein families, pseudogenes and proteome evolution.

Protein families can be used to understand many aspects of genomes, both their "live" and their "dead" parts (i.e. genes and pseudogenes). Surveys of genomes have revealed that, in every organism, there are always a few large families and many small ones, with the overall distribution following a power-law. This commonality is equally true for both genes and pseudogenes, and exists despite the fact that the specific families that are enlarged differ greatly between organisms. Furthermore, because of family structure there is great redundancy in proteomes, a fact linked to the large number of dispensable genes for each organism and the small size of the minimal, indispensable sub-proteome. Pseudogenes in prokaryotes represent families that are in the process of being dispensed with. In particular, the genome sequences of certain pathogenic bacteria (Mycobacterium leprae, Yersinia pestis and Rickettsia prowazekii) show how an organism can undergo reductive evolution on a large scale (i.e. the dying out of families) as a result of niche change. There appears to be less pressure to delete pseudogenes in eukaryotes. These can be divided into two varieties, duplicated and processed, where the latter involves reverse transcription from an mRNA intermediate. We discuss these collectively in yeast, worm, fly, and human. The fly has few pseudogenes apparently because of its high rate of genomic DNA deletion. In the other three organisms, the distribution of pseudogenes on the chromosome and amongst different families is highly non-uniform. Pseudogenes tend not to occur in the middle of chromosome arms, and tend to be associated with lineage-specific (as opposed to highly conserved) families that have environmental-response functions. This may be because, rather than being dead, they may form a reservoir of diverse "extra parts" that can be resurrected to help an organism adapt to its surroundings. In yeast, there may be a novel mechanism involving the [PSI+] prion that potentially enables this resurrection. In worm, the pseudogenes tend to arise out of families (e.g. chemoreceptors) that are greatly expanded in it compared to the fly. The human genome stands out in having many processed pseudogenes. These have a character very different from those of the duplicated variety, to a large extent just representing random insertions. Thus, their occurrence tends to be roughly in proportion to the amount of mRNA for a particular protein and to reflect the extent of the intergenic sequences. Further information about pseudogenes is available at http://genecensus.org/pseudogene

Animals↗

Recognizing the pseudogenes in bacterial genomes.

Pseudogenes are now known to be a regular feature of bacterial genomes and are found in particularly high numbers within the genomes of recently emerged bacterial pathogens. As most pseudogenes are recognized by sequence alignments, we use newly available genomic sequences to identify the pseudogenes in 11 genomes from 4 bacterial genera, each of which contains at least 1 human pathogen. The numbers of pseudogenes range from 27 in Staphylococcus aureus MW2 to 337 in Yersinia pestis CO92 (e.g. 1-8% of the annotated genes in the genome). Most pseudogenes are formed by small frameshifting indels, but because stop codons are A + T-rich, the two low-G + C Gram-positive taxa (Streptococcus and Staphylococcus) have relatively high fractions of pseudogenes generated by nonsense mutations when compared with more G + C-rich genomes. Over half of the pseudogenes are produced from genes whose original functions were annotated as 'hypothetical' or 'unknown'; however, several broadly distributed genes involved in nucleotide processing, repair or replication have become pseudogenes in one of the sequenced Vibrio vulnificus genomes. Although many of our comparisons involved closely related strains with broadly overlapping gene inventories, each genome contains a largely unique set of pseudogenes, suggesting that pseudogenes are formed and eliminated relatively rapidly from most bacterial genomes.

Genome, Bacterial↗

Evolution of the NANOG pseudogene family in the human and chimpanzee genomes.

BACKGROUND: The NANOG gene is expressed in mammalian embryonic stem cells where it maintains cellular pluripotency. An unusually large family of pseudogenes arose from it with one unprocessed and ten processed pseudogenes in the human genome. This article compares the NANOG gene and its pseudogenes in the human and chimpanzee genomes and derives an evolutionary history of this pseudogene family. RESULTS: The NANOG gene and all pseudogenes except NANOGP8 are present at their expected orthologous chromosomal positions in the chimpanzee genome when compared to the human genome, indicating that their origins predate the human-chimpanzee divergence. Analysis of flanking DNA sequences demonstrates that NANOGP8 is absent from the chimpanzee genome. CONCLUSION: Based on the most parsimonious ordering of inferred source-gene mutations, the deduced evolutionary origins for the NANOG pseudogene family in the human and chimpanzee genomes, in order of most ancient to most recent, are NANOGP6, NANOGP5, NANOGP3, NANOGP10, NANOGP2, NANOGP9, NANOGP7, NANOGP1, and NANOGP4. All of these pseudogenes were fixed in the genome of the human-chimpanzee common ancestor. NANOGP8 is the most recent pseudogene and it originated exclusively in the human lineage after the human-chimpanzee divergence. NANOGP1 is apparently an unprocessed pseudogene. Comparison of its sequence to the functional NANOG gene's reading frame suggests that this apparent pseudogene remained functional after duplication and, therefore, was subject to selection-driven conservation of its reading frame, and that it may retain some functionality or that its loss of function may be evolutionarily recent.

Animals↗

How many processed pseudogenes are accumulated in a gene family?

A simple kinetic model is developed that describes the accumulation of processed pseudogenes in a functional gene family. Insertion of new pseudogenes occurs at rate v per gene and is countered by spontaneous deletion (at rate delta per DNA segment) of segments containing processed pseudogenes. If there are k functional genes in a gene family, the equilibrium number of processed pseudogenes is k(v/delta), and the percentage of functional genes in the gene family at equilibrium is 1/[1 + (v/delta)]. v/delta values estimated for five gene families ranged from 1.7 to 15. This fairly narrow range suggests that the rates of formation and deletion of processed pseudogenes may be positively correlated for these families. If delta is sufficiently large relative to the per nucleotide mutation rate mu (delta greater than 20 mu), processed pseudogenes will show high homology with each other, even in the absence of gene conversion between pseudogenes. We argue that formation of processed pseudogenes may share common pathways with transposable elements and retroviruses, creating the potential for correlated responses in the evolution of processed pseudogenes due to direct selection for control of transposable elements and/or retroviruses. Finally, we discuss the nature of the selective forces that may act directly or indirectly to influence the evolution of processed pseudogenes.

Animals↗

Calculation and verification of the ages of retroprocessed pseudogenes.

To verify the existence of processed pseudogenes in different primates and their correlation with the estimated age of divergence, selected regions of processed pseudogenes of alpha-enolase, calmodulin II (CALMII), and argininosuccinate synthetase (AS) were amplified by the polymerase chain reaction (PCR) using DNA of blood samples. Published primate divergence times from the accepted paleontological records and the age of the pseudogenes based on molecular clock calculations were compared to data obtained by detection of PCR products exhibiting the expected amplicon size of the pseudogene region. For the alpha-enolase and the CALMII pseudogenes Psi(2), and Psi(3), calculated divergence times were 11, 19, and 36 Myr, respectively. For the AS pseudogenes Psi(1), Psi(3), and Psi(7), the divergence times were calculated to be 21, 25, and 16 Myr, respectively. Primer design and the annealing temperature are critical factors in the detection of pseudogenes in different species and impact greatly on the interpretation of the PCR analysis. The estimated divergence times of the selected pseudogenes utilizing calculations based on the molecular clock theory correlated well with experimental PCR detection of the selected pseudogenes represented in this study.

Animals↗