Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “pseudogene”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

The first human genes for tRNA(ArgICG), tRNA(GlyUCC), and tRNA(ThrIGU) and more tRNA(Val) pseudogenes: expression and pre-tRNA maturation in HeLa cell-free extracts.

A functional tRNA(Val) gene, which codes for the major tRNA(ValIAC) isoacceptor species, and three new tRNA(Val) pseudogenes have been isolated from human genomic DNA. Two tRNA(Val) pseudogenes and a tRNA(Val) variant gene were found to be associated with tRNA genes encoding tRNA(ArgICG), tRNA(GlyUCC), and tRNA(ThrIGU), respectively, on distinct DNA fragments. All tRNA genes, including the pseudogenes, are actively transcribed in HeLa nuclear extract. Pre-tRNAs of tRNA(Val), tRNA(Arg), tRNA(Thr), and tRNA(Gly) genes are correctly processed to mature-sized tRNAs, whereas the three tRNA(Val) pseudogenes yield stable pre-tRNAs in vitro. These findings reveal that, together with the three known pseudogenes, half of the members of the human tRNA(Val) gene family are pseudogenes, all of which are active in homologous nuclear extracts in vitro and presumably also in vivo.

Base Sequence↗

Nonneutral evolution of the transcribed pseudogene Makorin1-p1 in mice.

Pseudogenes are nonfunctional relics of formerly functional genes and are thought to evolve neutrally. In some pseudogenes, however, the molecular evolutionary patterns are atypical of neutrally evolving sequences, exhibiting sequence conservation, codon-usage bias, and other features associated with functional genes. Makorin1-p1 is a transcribed pseudogene first identified in the mouse Mus musculus. The transcript of Makorin1-p1 can regulate the stability of the transcript of its paralogous functional gene Makorin1. Specifically, the half-life of Makorin1 mRNA increases significantly in the presence of Makorin1-p1 transcript, and targeted deletion of Makorin1-p1 is lethal in mice. Here, we show that Makorin1-p1 originated after the separation of Mus and Rattus but before the divergence of M. musculus and M. pahari. The transcribed region of Makorin1-p1 exhibits rates of point and indel substitutions that are two to four times lower than those in the untranscribed region, suggesting that the transcribed region is under functional constraint and is not evolving neutrally. Although the transcript of Makorin1-p1 likely functions by its sequence similarity to Makorin1, we find no evidence of gene conversion between them, indicating that functional conservation alone is sufficient to maintain their coordinated evolution. A duplication-degeneration model is proposed to explain how Makorin1-p1 was co-opted into the regulatory system of Makorin1. There are over 10,000 pseudogenes in a typical mammalian genome, and it is plausible that many functional but untranslatable pseudogenes exist. Our results illustrate the potential of using evolutionary analysis to identify such pseudogenes from genome sequences.

Animals↗

Pseudogene.org: a comprehensive database and comparison platform for pseudogene annotation.

The Pseudogene.org knowledgebase serves as a comprehensive repository for pseudogene annotation. The definition of a pseudogene varies within the literature, resulting in significantly different approaches to the problem of identification. Consequently, it is difficult to maintain a consistent collection of pseudogenes in detail necessary for their effective use. Our database is designed to address this issue. It integrates a variety of heterogeneous resources and supports a subset structure that highlights specific groups of pseudogenes that are of interest to the research community. Tools are provided for the comparison of sets and the creation of layered set unions, enabling researchers to derive a current 'consensus' set of pseudogenes. Additional features include versatile search, the capacity for robust interaction with other databases, the ability to reconstruct older versions of the database (accounting for changing genome builds) and an underlying object-oriented interface designed for researchers with a minimal knowledge of programming. At the present time, the database contains more than 100,000 pseudogenes spanning 64 prokaryote and 11 eukaryote genomes, including a collection of human annotations compiled from 16 sources.

Databases, Genetic↗

Comparison of cytochrome P450 (CYP) genes from the mouse and human genomes, including nomenclature recommendations for genes, pseudogenes and alternative-splice variants.

OBJECTIVES: Completion of both the mouse and human genome sequences in the private and public sectors has prompted comparison between the two species at multiple levels. This review summarizes the cytochrome P450 (CYP) gene superfamily. For the first time, we have the ability to compare complete sets of CYP genes from two mammals. Use of the mouse as a model mammal, and as a surrogate for human biology, assumes reasonable similarity between the two. It is therefore of interest to catalog the genetic similarities and differences, and to clarify the limits of extrapolation from mouse to human. METHODS: Data-mining methods have been used to find all the mouse and human CYP sequences; this includes 102 putatively functional genes and 88 pseudogenes in the mouse, and 57 putatively functional genes and 58 pseudogenes in the human. Comparison is made between all these genes, especially the seven main CYP gene clusters. RESULTS AND CONCLUSIONS: The seven CYP clusters are greatly expanded in the mouse with 72 functional genes versus only 27 in the human, while many pseudogenes are present; presumably this phenomenon will be seen in many other gene superfamily clusters. Complete identification of all pseudogene sequences is likely to be clinically important, because some of these highly similar exons can interfere with PCR-based genotyping assays. A naming procedure for each of four categories of CYP pseudogenes is proposed, and we encourage various gene nomenclature committees to consider seriously the adoption and application of this pseudogene nomenclature system.

Alternative Splicing↗

A genome-wide survey of human pseudogenes.

We screened all intergenic regions in the human genome to identify pseudogenes with a combination of homology searches and a functionality test using the ratio of silent to replacement nucleotide substitutions (KA/KS). We identified 19,724 regions of which 95% +/- 3% are estimated to evolve neutrally and thus are likely to encode pseudogenes. Half of these have no detectable truncation in their pseudocoding regions and therefore are not identifiable by methods that require the presence of truncations to prove nonfunctionality. A comparative analysis with the mouse genome showed that 70% of these pseudogenes have a retrotranspositional origin (processed), and the rest arose by segmental duplication (nonprocessed). Although the spread of both types of pseudogenes correlates with chromosome size, nonprocessed pseudogenes appear to be enriched in regions with high gene density. It is likely that the human pseudogenes identified here represent only a small fraction of the total, which probably exceeds the number of genes.

Benchmarking↗

Processed pseudogenes of human endogenous retroviruses generated by LINEs: their integration, stability, and distribution.

We report here the presence of numerous processed pseudogenes derived from the W family of endogenous retroviruses in the human genome. These pseudogenes are structurally colinear with the retroviral mRNA followed by a poly(A) tail. Our analysis of insertion sites of HERV-W processed pseudogenes shows a strong preference for the insertion motif of long interspersed nuclear element (LINE) retrotransposons. The genomic distribution, stability during evolution, and frequent truncations at the 5' end resemble those of the pseudogenes generated by LINEs. We therefore suggest that HERV-W processed pseudogenes arose by multiple and independent LINE-mediated retrotransposition of retroviral mRNA. These data document that the majority of HERV-W copies are actually nontranscribed promoterless pseudogenes. The current search for HERV-Ws associated with several human diseases should concentrate on a small subset of transcriptionally competent elements.

Base Sequence↗

Characterization of the pseudogenic and genic homologous regions of von Willebrand factor.

The homologous pseudogenic and genic regions of von Willebrand factor (vWF) were studied in DNA from a patient with homozygous deletion of vWF genes and compared with a normal control. This analysis indicates informative restriction patterns for the investigation of restriction fragment length polymorphisms (RFLPs) and gene lesions, and for molecular cloning. A useful new genic XbaI RFLP was found and characterized. A large BgIII fragment of the pseudogenic region was cloned and mapped, and single sequences (9 kb) were used as probes. Corresponding genic and pseudogenic fragments, which contain exons 23-28, and specific restriction patterns were identified, including a new polymorphic TaqI site that was mapped in the gene. A cloned fragment contains the 5' boundary of the pseudogene and recognizes an additional and unknown homologous sequence in the genome. The chromosomal localization of the vWF pseudogene and of the breakpoint cluster region (BCR) gene were compared by 'in situ' hybridization: overlapping patterns were detected. The cloning, characterization and mapping of the pseudogenic region improves the analysis of this portion of chromosome 22 affected by several somatic and constitutional alterations, and also of the corresponding genic region on chromosome 12.

Base Sequence↗

The porA pseudogene of Neisseria gonorrhoeae--low level of genetic polymorphism and a few, mainly identical, inactivating mutations.

N. meningitidis is the only Neisseria species known to express two outer membrane porins, PorA and PorB. However, a porA pseudogene has been identified in N. gonorrhoeae. The present study investigated the prevalence and genetic polymorphism of this porA pseudogene in 87 different N. gonorrhoeae strains. The porA pseudogene was identified in all isolates. The pseudogene comprised 12 (5.5%), of which 10 were located in the promoter spacer, and 11 (1.0%) polymorphic nucleotide sites in the upstream segment containing the promoter region, i.e. the putative -10 and -35 sequences and the promoter spacer in-between, and the hypothetical PorA coding sequence, respectively. A phylogenetic analysis of the upstream segment and the hypothetical coding sequence identified 36 sequence variants, of which 30 were not previously described. All strains comprised at least two identical confirmed inactivating deletions, of which one was located in the promoter region and one in the hypothetical PorA coding sequence. In conclusion, the porA pseudogene and its few inactivating mutations are widespread in the N. gonorrhoeae population and the homology with the N. meningitidis porA gene reflects their common evolutionary origin. The highly conserved N. gonorrhoeae porA pseudogene may reflect an evolutionary neutral molecular clock and may be a suitable genetic target for diagnosis of N. gonorrhoeae.

Base Sequence↗

A probabilistic classifier for olfactory receptor pseudogenes.

BACKGROUND: Olfactory receptors (ORs), the largest mammalian gene superfamily (900-1400 genes), has >50% pseudogenes in humans. While most of these inactive genes are identified via coding frame (nonsense) disruptions, seemingly intact genes may also be inactive due to other deleterious (missense) mutations. An ultimate assessment of the actual size of the functional human OR repertoire thus requires an accurate distinction between genes and pseudogenes. RESULTS: To characterize inactive ORs with intact open reading frame, we have developed a probabilistic Classifier for Olfactory Receptor Pseudogenes (CORP). This algorithm is based on deviations from a functionally crucial consensus, constituting sixty highly conserved positions identified by a comparison of two evolutionarily-constrained OR repertoires (mouse and dog) with a small pseudogene fraction. We used a logistic regression analysis to assign appropriate coefficients to the conserved position and thus achieving maximal separation between active and inactive ORs. Consequently, the algorithms identified only 5% of the mouse functional ORs as pseudogenes, setting an upper limit of 0.05 to the false positive detection. Finally we used this algorithm to classify the 384 purportedly intact human OR genes. Of these, 135 were predicted as likely encoding non-functional proteins, and 38 were segregating between active and inactive forms due to missense polymorphisms. CONCLUSION: We demonstrated that the CORP algorithm is capable to distinguish between functional and non-functional OR genes with high precision even when the encoded protein would differ by a single amino acid. Using the CORP algorithm, we predict that approximately 70% of human OR genes are likely non-functional pseudogenes, a much higher number than hitherto suspected. The method we present may be employed for better annotation of inactive members in other gene families as well. CORP algorithm is available at: http://bioportal.weizmann.ac.il/HORDE/CORP/

Algorithms↗

A comparison of variation between a MHC pseudogene and microsatellite loci of the little greenbul (Andropadus virens).

BACKGROUND: We investigated genetic variation of a major histocompatibility complex (MHC) pseudogene (Anvi-DAB1) in the little greenbul (Andropadus virens) from four localities in Cameroon and one in Ivory Coast, West Africa. Previous microsatellite and mitochondrial DNA analyses had revealed little or no genetic differentiation among Cameroon localities but significant differentiation between localities in Cameroon and Ivory Coast. RESULTS: Levels of genetic variation, heterozygosity, and allelic diversity were high for the MHC pseudogene in Cameroon. Nucleotide diversity of the MHC pseudogene in Cameroon and Ivory Coast was comparable to levels observed in other avian species that have been studied for variation in nuclear genes. An excess of rare variants for the MHC pseudogene was found in the Cameroon population, but this excess was not statistically significant. Pairwise measures of population differentiation revealed high divergence between Cameroon and Ivory Coast for microsatellites and the MHC locus, although for the latter distance measures were much higher than the comparable microsatellite distances. CONCLUSION: We provide the first ever comparison of variation in a putative MHC pseudogene to variation in neutral loci in a passerine bird. Our results are consistence with the action of neutral processes on the pseudogene and suggest they can provide an independent perspective on demographic history and population substructure.

Alleles↗

A computational approach for identifying pseudogenes in the ENCODE regions.

BACKGROUND: Pseudogenes are inheritable genetic elements showing sequence similarity to functional genes but with deleterious mutations. We describe a computational pipeline for identifying them, which in contrast to previous work explicitly uses intron-exon structure in parent genes to classify pseudogenes. We require alignments between duplicated pseudogenes and their parents to span intron-exon junctions, and this can be used to distinguish between true duplicated and processed pseudogenes (with insertions). RESULTS: Applying our approach to the ENCODE regions, we identify about 160 pseudogenes, 10% of which have clear 'intron-exon' structure and are thus likely generated from recent duplications. CONCLUSION: Detailed examination of our results and comparison of our annotation with the GENCODE reference annotation demonstrate that our computation pipeline provides a good balance between identifying all pseudogenes and delineating the precise structure of duplicated genes.

Computational Biology↗

The human immunoglobulin kappa locus: pseudogenes, unique and repetitive sequences.

The human kappa locus contains 25 pseudogenes. After seven of them were described earlier the structures of the remaining 18 are reported now, thus completing the description of all human V kappa genes and pseudogenes. Most of the pseudogenes carry several defects each. Alignments of the pseudogene sequences and comparison with the consensus sequences of the potentially functional V kappa genes indicate that, on PCR amplification of genomic DNA aimed at certain genes of the latter class, also some of the pseudogenes would be coamplified. Unique sequences, which qualify as sequence tagged sites (STS), were defined across the locus. The occurrence of 15 repetitive elements of the LINE1 type in the locus is described. The 15 sequenced Alu elements were assigned to the known Alu subfamilies of different evolutionary age. One of the Alu elements was found only in one of the copies of the kappa locus. It must, therefore, have been inserted after the duplication step which may have taken place about one million years ago. This element belongs to an Alu subfamily known to have been mobile until recently. Some aspects of the evolution of the V kappa pseudogenes and orphons (i.e. V kappa genes located outside the kappa locus) are also discussed.

Amino Acid Sequence↗

Chromosomal localization and racial distribution of the polymorphic human dihydrofolate reductase pseudogene (DHFRP1).

The human dihydrofolate reductase (DHFR) gene family comprises one functional gene and at least four intronless processed pseudogenes. The functional DHFR gene is on chromosome 5, and DHFRP4 is on chromosome 3. Using in situ hybridization, we have now localized the functional DHFR gene to the region q11.1-q13.3 on chromosome 5. By genomic DNA analysis of a panel of human X rodent somatic-cell hybrids, we determined the chromosomal assignment of the DHFRP1 pseudogene to chromosome 18 and that of the DHFRP2 pseudogene to chromosome 6. The DHFRP1 pseudogene exhibits a novel form of polymorphism in humans in that it is present in the DNA of some individuals and absent in that of others. We investigated the racial distribution of this pseudogene in five racial groups. The allelic frequency as defined by analysis of 180 chromosomes was found to be 94% in Mediterraneans, 77% in Asian Indians, 67% in Chinese, 57% in Southeast Asians, and 32% in American blacks. These data suggest that the transposition of this "perfect" pseudogene occurred prior to the inception of the human racial groups.

Animals↗

Evolution of cytochrome c genes and pseudogenes.

A statistical analysis of the nucleotide sequences of cytochrome c genes from four species of animals and two of yeast and of cytochrome c pseudogenes from rat, mouse, and human was conducted. It was estimated that animals and yeast diverged 1.2 billion years ago, that the two duplicated genes DC3 and DC4 in Drosophila diverged 520 million years ago, and that the two duplicated genes Iso-1 and Iso-2 in the yeast Saccharomyces cerevisiae diverged 200 million years ago. DC3 is expressed at a low level and has evolved 3 times faster than DC4. This observation supports the neutralist view that relaxation of functional constraints is a more likely cause of accelerated evolution following gene duplication than is advantageous mutation. All the rodent pseudogenes examined appear to be processed pseudogenes derived directly from the functional genes, and most of them apparently arose after the mouse-rat split. No event of gene conversion could be detected between any pair of the rodent pseudogenes. Our analysis suggests that the human cytochrome c gene has evolved at a rate comparable to the average rate for pseudogenes, whereas some human cytochrome c pseudogenes have evolved at an exceptionally low rate.

Animals↗

Amplification of human argininosuccinate synthetase pseudogenes.

The human genome contains multiple pseudogenes for an argininosuccinate synthetase (AS) gene. To elucidate the molecular mechanisms of generation and dispersion, complete nucleotide sequences of four different AS pseudogenes, psi AS-Y, psi AS-A1, psi AS-A2 and psi AS-A3, have been determined. A comparison of these sequences with those of three reported AS pseudogenes, psi AS-1, psi AS-3 and psi AS-7 revealed that two pairs, psi AS-Y/psi AS-7 and psi AS-A3/psi AS-1, are highly homologous but not identical, thereby suggesting that one of the pairs is generated by a duplication of the other member of the pairs. The psi AS-Y, which is probably located on chromosome Y, and the partially sequenced psi AS-7 are both interrupted by an Alu element at exactly the same site in their 3'-end regions. These two Alu elements are located in an opposite orientation relative to the direction of transcription of the pseudogene, and their possible role on pseudogene dispersion was examined. The psi AS-A1 is also accompanied by an Alu element at its 3' end. In this case, the orientation of the Alu element is the same as that of the pseudogene. The psi AS-A1 and the Alu element are flanked with direct repeats, as if they had been inserted into a chromosomal site, as a single unit.

Argininosuccinate Synthase↗

Extraordinarily high evolutionary rate of pseudogenes: evidence for the presence of selective pressure against changes between synonymous codons.

Comparisons of nucleotide sequences of several pseudogenes described to date, including alpha- and beta-globin and immunoglobulin kappa-type variable domain pseudogenes, with those of functional counterparts revealed that pseudogenes accumulate mutations at an extremely high rate uniformly over their entirety. It is remarkable that the evolutionary rate exceeds the rate of changes between synonymous codons, the highest known rate, in functional genes. Because no pseudogenes appear to function, this result strongly supports the neutral theory. In addition this result apparently indicates the presence of selective pressure against changes between synonymous codons in functional genes. Close examinations of codon utilization patterns in pseudogenes and functional genes revealed a significant correlation between the rate of changes at synonymous codon sites and the strength of bias in code word usage. This implies that even synonymous codon changes are not completely free from selective pressure but are constrained in part, although presumably weakly, depending on the degree of bias in code word usage. We also reexamined alignment between mouse beta h3 (pseudogene) and beta maj sequences and found a unique structure of the beta h3 that is homologous in sequence to the beta maj gene overall but contains a long deletion (about 150 base pairs) in the middle of the gene.

Animals↗

pANT: a method for the pairwise assessment of nonfunctionalization times of processed pseudogenes.

We present a method for pairwise Assessment of Nonfunctionalization Times (pANT) in processed pseudogenes. Contrary to existing methods for estimating nonfunctionalization times, pANT utilizes previously calculated probabilities of nucleotide substitution as explicit rate measurements, rather than assume that the substitution rates are the same for all nucleotides. Thus, the method allows a more accurate computation of the time that has elapsed since the nonfunctionalization of a pseudogene. Whereas existing methods require the sequence of an orthologous functional gene, which is not always at hand, pANT only uses the pairwise alignment of the gene/pseudogene pair, thus expanding the range of problems that can be tackled. To estimate evolutionary times in nonfunctional sequences, pANT measures the differences in the pairwise alignment of a gene and its paralogous processed pseudogene, using only the first and second codon positions. It assumes that, because of functional constraints, these positions in the sequence of the functional homolog have not changed since the time of nonfunctionalization of the pseudogene. Hence, the sequence of the gene may be used as the ancestor of the pseudogene. We show that the method's reliance on a detailed substitution matrix, which is derived separately for each species, makes it more accurate than existing methods. We applied pANT to the case of the unitary alpha-1,3-galactosyltransferase human pseudogene and found that our estimate of the nonfunctionalization time was in agreement with that obtained by taxonomic and paleontological considerations pertaining to the divergence between platyrrhines (New World monkeys) and cattarhines (Old World monkeys).

Animals↗

A previously undetected pseudogene in the human alpha globin gene cluster.

The sequence of the DNA between two pseudogenes in the human alpha-like globin gene cluster has been determined. Comparison of this sequence with sequences from other alpha-like globin gene clusters revealed another pseudogene, psi alpha 2, between the previously recognized pseudogenes zeta 1 and psi alpha 1. Therefore, the human alpha-like globin gene family is organized 5'-zeta 2-zeta 1-psi alpha 2-psi alpha 1-alpha 2-alpha 1-3'. The new pseudogene psi alpha 2 is very close to zeta 1, beginning only 65 base pairs 3' to the polyadenylation site of zeta 1. The first exon and the first intron of psi alpha 2 are interrupted by large inserts which are flanked by short (6 to 8 base pairs) direct repeats. The pseudogene psi alpha 2 lacks a promoter for transcription by RNA polymerase II, the first exon is highly divergent, one splice site is mutated, and five different frameshift mutations have occurred in the coding regions. Thus psi alpha 2 cannot encode a globin polypeptide. This pseudogene was not recognized in previous hybridization analyses of the human alpha-like globin gene cluster, and our discovery of it by sequence analysis suggests that divergent copies of a large number of genes may comprise a substantial fraction of the slowly renaturing DNA of mammalian genomes.

Base Sequence↗