Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “pseudogene”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Molecular evolution of the Cecropin multigene family in Drosophila. functional genes vs. pseudogenes.

Approximately 4 kb of the Cecropin cluster region have been sequenced in nine lines of Drosophila melanogaster and one line of the sibling species D. simulans, D. mauritiana, and D. sechellia. This region includes three functional genes (CecA1, CecA2, and CecB), which are involved in the insect immune response, and two pseudogenes (CecPsi1 and CecPsi2). The level of silent polymorphism in the three Cec genes is rather high (0.028), and there is no excess of nonsynonymous polymorphism. There is no evidence of gene conversion in the history of these genes. The interspecific comparison has revealed that in the three species of the simulans cluster the CecA2 gene is partially deleted and has therefore lost its function and become a pseudogene; in each of the species, subsequent deletions have accumulated. Divergence estimates indicate that the CecPsi1 and CecPsi2 pseudogenes are highly diverged, both between themselves and relative to the other three Cec genes. However, both CecPsi1 and CecPsi2 have conserved transcriptional signals and splice sites, and they present an open reading frame; also, correctly spliced transcripts have been detected for both CecPsi1 and CecPsi2. The data support that these genes are either active genes with some null alleles or young pseudogenes.

Amino Acid Sequence↗

Comparative analysis of the phosphomannomutase genes PMM1, PMM2 and PMM2psi: the sequence variation in the processed pseudogene is a reflection of the mutations found in the functional gene.

The search for the carbohydrate-deficient glycoprotein syndrome type I (CDG1) gene has revealed the existence of a family of phosphomannomutase (PMM) genes in humans. Two expressed PMM genes, PMM1 and PMM2 , are located on chromosome bands 22q13 and 16p13, respectively, and a processed pseudogene PMM2 psi is located on chromosome 18p. Mutations in PMM2 are the cause of CDG type IA whereas no disorder has been associated with defects in PMM1 as yet. Here, we describe the genomic organization of these paralogous genes. There is a 65% identity of the coding sequence, and all intron/exon boundaries have been conserved. The processed pseudogene is more closely related to PMM2 . Remarkably, several base substitutions in PMM2 that are associated with disease are also present at the corresponding positions in the pseudogene. Thus, mutations that occur at a slow rate in the active gene in the population have also accumulated in the pseudogene.

Amino Acid Sequence↗

Large-scale methylation analysis of human genomic DNA reveals tissue-specific differences between the methylation profiles of genes and pseudogenes.

Cytosine in CpG dinucleotides is frequently found to be methylated in the DNA of higher eukaryotes and differential methylation has been proposed to be a key element in the organization of gene expression in man. To address this question systematically, we used bisulfite genomic sequencing to study the methylation patterns of three X-linked genes and one autosomal pseudogene in two adult individuals and across nine different tissues. Two of the genes, SLC6A8 and MSSK1, are tissue-specifically expressed. CDM is expressed ubiquitously. The pseudogene, psi SLC6A8, is exclusively expressed in the testis. The promoter regions of the SLC6A8, MSSK1 and CDM genes were found to be essentially unmethylated in all tissues, regardless of their relative expression level. In contrast, the pseudogene psi SLC6A8 shows high methylation of the CpG islands in all somatic tissues but complete demethylation in testis. Methylation profiles in different tissues are similar in shape but not identical. The data for the two investigated individuals suggest that methylation profiles of individual genes are tissue specific. Taken together, our findings support a model in which the bodies of the genes are predominantly methylated and thus insulated from the interaction with DNA-binding proteins. Only unmethylated promoter regions are accessible for binding and interaction. Based on this model we propose to use DNA methylation studies in conjunction with large-scale sequencing approaches as a tool for the prediction of cis-acting genomic regions, for the identification of cryptic and potentially active CpG islands and for the preliminary distinction of genes and pseudogenes.

Aged↗

The silencing of pseudogenes.

Pseudogenes are nonfunctional DNA sequences that can accumulate in the genomes of some bacterial species, especially those undergoing processes like niche change, host specialization, or weak selection strength. They may last for long evolutionary periods, opening the question of how the genome prevents expression of these degenerated or disrupted genes that would presumably give rise to malfunctioning proteins. We have investigated ribosomal binding strength at Shine-Dalgarno sequences and the prevalence of sigma70 promoter regions in pseudogenes across bacteria. It is reported that the RNA polymerase-binding sites and more strongly the ribosome-binding regions of pseudogenes are highly degraded, suggesting that transcription and translation are impaired in nonfunctional open reading frames. This would reduce the metabolic investment on faulty proteins because although pseudogenes can persist for long time periods, they would be effectively silenced. It is unclear whether mutation accumulation on regulatory regions is neutral or whether it is accelerated by selection.

Binding Sites↗

Patterns of nucleotide substitution, insertion and deletion in the human genome inferred from pseudogenes.

Nucleotide substitution, insertion and deletion (indel) events are the major driving forces that have shaped genomes. Using the recently identified human ribosomal protein (RP) pseudogene sequences, we have thoroughly studied DNA mutation patterns in the human genome. We analyzed a total of 1726 processed RP pseudogene sequences, comprising more than 700 000 bases. To be sure to differentiate the sequence changes occurring in the functional genes during evolution from those occurring in pseudogenes after they were fixed in the genome, we used only pseudogene sequences originating from parts of RP genes that are identical in human and mouse. Overall, we found that nucleotide transitions are more common than transversions, by roughly a factor of two. Moreover, the substitution rates amongst the 12 possible nucleotide pairs are not homogeneous as they are affected by the type of immediately neighboring nucleotides and the overall local G+C content. Finally, our dataset is large enough that it has many indels, thus allowing for the first time statistically robust analysis of these events. Overall, we found that deletions are about three times more common than insertions (3740 versus 1291). The frequencies of both these events follow characteristic power-law behavior associated with the size of the indel. However, unexpectedly, the frequency of 3 bp deletions (in contrast to 3 bp insertions) violates this trend, being considerably higher than that of 2 bp deletions. The possible biological implications of such a 3 bp bias are discussed.

Algorithms↗

Evidence suggesting that a fifth of annotated Caenorhabditis elegans genes may be pseudogenes.

Only a minority of the genes, identified in the Caenorhabditis elegans genome sequence data by computer analysis, have been characterized experimentally. We attempted to determine the expression patterns for a random sample of the annotated genes using reporter gene fusions. A low success rate was obtained for evolutionarily recently duplicated genes. Analysis of the data suggests that this is not due to conditional or low-level expression. The remaining explanation is that most of the annotated genes in the recently duplicated category are pseudogenes, a proportion corresponding to 20% of all of the annotated C. elegans genes. Further support for this surprisingly high figure was sought by comparing sequences for families of recently duplicated C. elegans genes. Although only a preliminary analysis, clear evidence for a gene having been recently inactivated by genetic drift was found for many genes in the recently duplicated category. At least 4% of the annotated C. elegans genes can be recognized as pseudogenes simply from closer inspection of the sequence data. Lessons learned in identifying pseudogenes in C. elegans could be of value in the annotation of the genomes of other species where, although there may be fewer pseudogenes, they may be harder to detect.

Amino Acid Sequence↗

PCR analysis of the H ferritin multigene family reveals the existence of two classes of processed pseudogenes.

The human gene coding for the apoferritin H subunit belongs to a complex multigene family constituted by the expressed gene and by an undefined number of pseudogenes. We have used a strategy based on PCR to amplify specifically the H pseudogenes from a sample of human genomic DNA. With this approach, three new H pseudogenes have been cloned and characterized by DNA sequence analysis. In addition, we have identified a new type of pseudogene, the size of which (700 bp) is caused by multiple detection events in the putative coding region.

Base Sequence↗

Activation of the silent human cytokeratin 17 pseudogene-promoter region by cryptic enhancer elements of the cytokeratin 17 gene.

We have previously described the three loci CK-CA, CK-CB and CK-CC in the human genome that contain clustered type-I cytokeratin genes and reported the complete nucleic acid sequences of the functional cytokeratin 17 gene located in CK-CA and two closely related pseudogenes present in CK-CB and CK-CC [Troyanovsky, S.M., Leube, R.E. & Franke, W.W. (1992) Eur. J. Cell Biol. 59, 127-137]. By nucleic acid sequence analysis, we now show that extensive similarities between the functional gene and the pseudogenes exist in the 5'-upstream region. However, despite the high degree of nucleic acid identity (94%), only the 5'-upstream region of the functional gene was able to induce significant transcriptional activity in transfected cells of epithelial origin. Using chimeric upstream regions consisting of different fragments from the pseudogene and the functional gene, we made the surprising observation that cis elements in the proximal 5'-upstream region of the pseudogene promoter can cooperate with distal enhancer elements of the functional gene to induce strong chloramphenicol-O-acetyltransferase activity in transfected HeLa cells. A major site in the proximal upstream region was identified by deoxyribonuclease protection experiments to be necessary for this cooperative effect. The structure and properties of this element were further analysed by transfection of different chloramphenicol-O-acetyltransferase gene constructs, and by nucleic acid sequence comparison to corresponding regions of the related cytokeratins 14 and 16. It is concluded that the upstream regions identified in this study contribute to the strong expression of the human cytokeratin 17 gene in a coordinated fashion.

Animals↗

Molecular characterization of nuclear small subunit (18S)-rDNA pseudogenes in a symbiotic dinoflagellate (Symbiodinium, Dinophyta).

For the dinoflagellates, an important group of single-cell protists, some nuclear rDNA phylogenetic studies have reported the discovery of rDNA pseudogenes. However, it is unknown if these aberrant molecules are confined to free-living taxa or occur in other members of the group. We have cultured a strain of symbiotic dinoflagellate, belonging to the genus Symbiodinium, which produces three distinct amplicons following PCR for nuclear small subunit (18S) rDNA genes. These amplicons contribute to a unique restriction fragment length polymorphism pattern diagnostic for this particular strain. Sequence analyses revealed that the largest amplicon was the expected region of 18S-rDNA, while the two smaller amplicons are Symbiodinium nuclear 18S-rDNA genes that contain single long tracts of nucleotide deletions. Reverse transcription (RT)-PCR experiments did not detect RNA transcripts of these latter genes, suggesting that these molecules represent the first report of nuclear 18S-rDNA pseudogenes from the genome of Symbiodinium. As in the free-living dinoflagellates, nuclear rDNA pseudogenes are effective indicators of unique Symbiodinium strains. Furthermore, the evolutionary pattern of dinoflagellate nuclear rDNA pseudogenes appears to be unique among organisms studied to date, and future studies of these unusual molecules will provide insight on the cellular biology and genomic evolution of these protists.

Animals↗

High-level expression of pseudogenes in Mycobacterium leprae.

Recent studies have revealed that some RNAs are transcribed from noncoding DNA regions, including pseudogenes, and are functional as riboregulators. We have attempted to assess the gene expression profile throughout the Mycobacterium leprae genome using an array technique. Twelve highly expressed gene regions were identified that show an alteration in expression levels upon infection. Six of these were pseudogenes. Although M. leprae has an exceptional number and proportion of pseudogenes among species, our results suggest that some of the M. leprae pseudogenes are not just 'decayed' genes, but may have a functional role.

Animals↗

Characterization and chromosome localization of a processed pseudogene related to the bovine laminin receptor gene family.

A bovine BAC clone containing a processed laminin receptor pseudogene (LAMR1P) has been isolated and characterized. A 2,901-bp sequence was produced from the clone, of which 1,187 bp represented seven identifiable exon-like domains, but no intervening sequences. The pseudogene sequence reveals several transversions and transitions, as well as insertions and deletions. A premature stop codon motif is present at nucleotide position 115 located in the exon-2-like domain. Physical mapping of the gene was performed by FISH and RH panel mapping and assigned LAMR1P to BTA4q24-->q26 with the closest linkage to BM6458 (19 cR, LOD score of 11.6). The functional laminin receptor putatively plays an important role in the transmission of bovine spongiform encephalopathy (BSE). In this process, the receptor supposedly acts as the binding site for prion proteins to enter mammalian cells. Considering the existence of several human laminin receptor pseudogenes forming a complex family, any knowledge of even pseudogene sequences might be helpful to isolate the functional bovine laminin receptor gene.

Animals↗

High level transcription of the glucocerebrosidase pseudogene in normal subjects and patients with Gaucher disease.

Gaucher disease is due to mutations involving the glucocerebrosidase gene. A closely homologous pseudogene is located approximately 16 kD downstream from the functional gene. Sequence analysis of clones from cDNA libraries made from skin fibroblast cultures showed several independent clones with the sequence of an aberrantly processed pseudogene message. Examination of cellular RNA from lymphoblasts or fibroblasts obtained from thirteen Gaucher disease patients, one Gaucher disease heterozygote, and four normal subjects showed that the pseudogene was consistently transcribed, and that in some cases the level of transcription seemed to be approximately equal to that of the functional gene. The transcription of the pseudogene must be taken into account when attempting to detect mutations of glucocerebrosidase by the study of cDNA libraries.

Base Sequence↗

Origin and evolution of processed pseudogenes that stabilize functional Makorin1 mRNAs in mice, primates and other mammals.

We investigate the origin and evolution of a mouse processed pseudogene, Makorin1-p1, whose transcripts stabilize functional Makorin1 mRNAs. It is shown that Makorin1-p1 originated almost immediately before the musculus and cervicolor species groups diverged from each other some 4 million years ago and that the Makorin1-p1 orthologs in various Mus species are transcribed. However, Mus caroli in the cervicolor species group expresses not only Makorin1-p1, but also another older Makorin1-derived processed pseudogene, demonstrating the rapid generation and turnover in subgenus Mus. Under this circumstance, transcribed processed pseudogenes (TPPs) of Makorin1 evolved in a strictly neutral fashion even with an enhanced substitution rate at CpG dinucleotide sites. Next, we extend our analyses to rats and other mammals. It is shown that although these species also possess their own Makorin1-derived TPPs, they occur rather infrequently in simian primates. Under this circumstance, it is hypothesized that already existing TPPs must be prevented from accumulating detrimental mutations by negative selection. This hypothesis is substantiated by the presence of two rather old TPPs, MKRNP1 and MKRN4, in humans and New World monkeys. The evolutionary rate and pattern of Makorin1-derived processed pseudogenes depend heavily on how frequently they are disseminated in the genome.

Animals↗

The human genome contains a single processed pseudogene for alpha enolase located on chromosome 1.

We have isolated and characterized human genomic clones containing an alpha enolase pseudogene which lacks introns and has the hallmarks of having been generated by reverse transcription. Two in-frame termination codons renders its coding region incapable of producing a functional protein. An Alu-like sequence is present in the region homologous to the 3' untranslated of the alpha enolase mRNA. Comparison of the two sequences shows that the pseudogene diverged from its functional counterpart about 14 million years ago and interestingly it is the only alpha enolase pseudogene present in the human genome. Chromosomal mapping locates the processed pseudogene to human chromosome 1, the same chromosome where the functional gene has been mapped.

Base Sequence↗

Identification of antisense RNA transcripts from a human DNA topoisomerase I pseudogene.

Eukaryotic topoisomerase I (TOP1), a DNA unwinding enzyme, plays an essential role in several cellular functions; however, regulation of TOP1 activity remains unknown. In an effort to identify potential regulators of TOP1 activity, the transcriptional activity of a TOP1 pseudogene in chromosome 1 was studied. By using primers unique to the TOP1 pseudogene, strand-specific polymerase chain reaction analysis of HeLa RNA amplified products from at least two transcripts oriented in the antisense direction of TOP1 mRNA. In one case, polymerase chain reaction and Northern blot analysis found a 0.7-kilobase antisense transcript. Upon estimation of 5' and 3' boundaries, a 497 base stretch of homology with the TOP1 mRNA was found. While the function of these TOP1 antisense transcripts remains unknown, recent studies of naturally occurring antisense RNA have demonstrated several potential regulatory roles. The production of antisense transcripts from a TOP1 pseudogene was the first example of a naturally occurring antisense RNA transcript produced from a pseudogene.

Base Sequence↗

Gene and pseudogene of the mouse cation-dependent mannose 6-phosphate receptor. Genomic organization, expression, and chromosomal localization.

The cation-dependent mannose 6-phosphate receptor (CD-MPR) is one of the two transmembrane proteins involved in transport of lysosomal enzymes. We have cloned the mouse CD-MPR gene and also a very unusual processed-type CD-MPR pseudogene. They are both present at one copy per haploid genome and map to chromosomes 6 and 3, respectively. Comparison of the complete 10-kilobase (kb) sequence of the functional gene with the cDNA indicates that it contains seven exons. Exon 1 encodes the 5'-untranslated region of the mRNA, the others (exons 2-7) encode the luminal, transmembrane, and cytoplasmic domains of the CD-MPR. Exon 7 also contains a 1.2-kb-long 3'-untranslated region of the mRNA. A unique transcription-initiation site was determined by primer extension of mouse liver mRNA. The promoter elements in the 5' upstream region of this site resemble those contained in genes constitutively transcribed. However, Northern blot analysis demonstrates that the CD-MPR is variably expressed in adult mouse tissues and during mouse development. The pseudogene, which is flanked by direct repeats, is almost colinear with the cDNA indicating that it presumably arose by reverse transcription of an mRNA. However, the pseudogene differs from the cDNA. It contains at its 5' end, an additional 340-nucleotide (nt) sequence homologous to the promoter region of the functional gene. This sequence exhibits some promoter activity in vitro. Furthermore, a 24-nt insertion interrupts the region homologous to the 5'-noncoding region of the cDNA. In the functional gene, this 24-nt sequence occurs between exon 1 and 2, where it is flanked by typical consensus sequences of exon/intron boundaries. Therefore, it may represent an additional exon of the functional gene. These two features of the pseudogene suggest that expression of the CD-MPR gene may be regulated by use of different promoters and/or alternative splicing.

Amino Acid Sequence↗

The HLA class I gene family includes at least six genes and twelve pseudogenes and gene fragments.

We report the characterization of eight HLA class I homologous sequences isolated from cosmid and lambda libraries made from lymphoblastiod cell line 721 DNA. Four of these sequences, each contained within HindIII fragments of 1.7, 2.1, 3.0, and 8.0 kb, have class I homology extending over short intronexon regions. The remaining four are found within 7.5-, 8.0-, 9.0-, and 16.0-kb HindIII fragments, the first having homology to the 5' half of a class I gene whereas the latter three are homologous to the 3' portion of a class I gene. When combined with the characterization of other class I clones, this work brings the total number of HLA class I homologous sequences cloned and characterized to 18. Restriction mapping of cosmid clones showed that some of these sequences are linked to one another and to other class I pseudogenes and genes within 50-kb regions. Reconstruction experiments using the 18 class I genes and pseudogenes were performed that indicated that we had cloned all of the members of the HLA class I gene family detectable using HLA-A2 genomic DNA as probe. An additional 19th member of the class I gene family was identified using an HLA-E cDNA probe. Further Southern analysis with other class I probes indicated the 19 sequences comprise the entire class I gene family in LCL 721. Locus-specific probes were isolated from five of the eight clones and were used in Southern analysis of diverse genomic DNA to examine the polymorphism of the pseudogene sequences, demonstrating that some of them were highly polymorphic and some were missing entirely in certain haplotypes. An additional class I sequence, not contained within the 721 genome, was identified and may be found in association with the HLA-A11-Bw60 haplotype. Sequence comparisons were carried out to examine the evolutionary relationships among the pseudogenes. Hypothetical events in the evolution of the class I region are discussed.

Base Sequence↗

A database of polymorphisms in the von Willebrand factor gene and pseudogene. For the Consortium on von Willebrand Factor Mutations and Polymorphisms and the Subcommittee on von Willebrand Factor of the Scientific and Standardization Committee of the International Society on Thrombosis and Haemostasis.

Nucleotide sequence polymorphisms in the von Willebrand factor (vWF) gene are useful for genetic studies in von Willebrand disease (vWD). This database describes 33 known vWF polymorphisms distributed throughout the vWF gene. DNA sequence information is available for 21 of these sites. The most informative system is a tetranucleotide repeat polymorphism in vWF intron 40. Sixteen of these polymorphisms are within vWF exons, and approximately half of them also alter the encoded amino acid sequence. Many occur close to mutations that cause vWD. The high prevalence of vWF polymorphisms must be considered in the analysis of candidate vWD mutations. In addition to the vWF gene on chromosome 12, there is a partial unprocessed vWF pseudogene on chromosome 22 that corresponds to vWF exons 23 to 34. Three polymorphisms have been assigned to the vWF pseudogene. Because the vWF gene and pseudogene have diverged only approximately 3.1% in DNA sequence, correct assignment of polymorphisms to either locus can be difficult in the region of homology. This problem has been solved in some cases by comparison of the published sequences and predicted restriction maps for the gene and pseudogene.

DNA Mutational Analysis↗