Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Duplication”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Gene duplication and gene conversion in the Caenorhabditis elegans genome.

A comprehensive analysis of duplication and gene conversion for 7394 Caenorhabditis elegans genes (about half the expected total for the genome) is presented. Of the genes examined, 40% are involved in duplicated gene pairs. Intrachromosomal or cis gene duplications occur approximately two times more often than expected. In general the closer the members of duplicated gene pairs are, the more likely it is that gene orientation is conserved. Gene conversion events are detectable between only 2% of the duplicated pairs. Even given the excesses of cis duplications, there is an excess of gene conversion events between cis duplicated pairs on every chromosome except the X chromosome. The relative rates of cis and trans gene conversion and the negative correlation between conversion frequency and DNA sequence divergence for unconverted regions of converted pairs are consistent with previous experimental studies in yeast. Three recent, regional duplications, each spanning three genes are described. All three have already undergone substantial deletions spanning hundreds of base pairs. The relative rates of duplication and deletion may contribute to the compactness of the C. elegans genome.

Animals↗

The ghost of selection past: rates of evolution and functional divergence of anciently duplicated genes.

The duplication of genes and even complete genomes may be a prerequisite for major evolutionary transitions and the origin of evolutionary novelties. However, the evolutionary mechanisms of gene evolution and the origin of novel gene functions after gene duplication have been a subject of many debates. Recently, we compiled 26 groups of orthologous genes, which included one gene from human, mouse, and chicken, one or two genes from the tetraploid Xenopus and two genes from zebrafish. Comparative analysis and mapping data showed that these pairs of zebrafish genes were probably produced during a fish-specific genome duplication that occurred between 300 and 450 Mya, before the teleost radiation (Taylor et al. 2001). As discussed here, many of these retained duplicated genes code for DNA binding proteins. Different models have been developed to explain the retention of duplicated genes and in particular the subfunctionalization model of Force et al. (1999) could explain why so many developmental control genes have been retained. Other models are harder to reconcile with this particular set of duplicated genes. Most genes seem to have been subjected to strong purifying selection, keeping properties such as charge and polarity the same in both duplicates, although some evidence was found for positive Darwinian selection, in particular for Hox genes. However, since only the cumulative pattern of nucleotide substitutions can be studied, clear indications of positive Darwinian selection or neutrality may be hard to find for such anciently duplicated genes. Nevertheless, an increase in evolutionary rate in about half of the duplicated genes seems to suggest that either positive Darwinian selection has occurred or that functional constraints have been relaxed at one point in time during functional divergence.

Amino Acid Sequence↗

Evidence that a dodecamer duplication in the gene HOPA in Xq13 is not associated with mental retardation.

A recent study suggested that a dodecamer duplication in exon 42 of the HOPA gene in Xq13 may be a significant factor in the etiology of X-linked mental retardation. In an effort to investigate this possibility, we determined the incidence of the dodecamer duplication in cohorts of non-fragile X males with mental retardation from three countries, cohorts of fragile X males from two countries, 43 probands from families with X-linked mental retardation and control cohorts from three countries. The duplication was found in 3.6-4.0% of male patients from two non-fragile X groups (Italy and South Carolina), in 1.2% from another non-fragile X group (South Africa), but in no male patients from families with X-linked mental retardation (South Carolina). The dodecamer duplication was also found in several white males with fragile X syndrome from France (5%) and South Africa (22.2%). Additionally, the duplication was found in 1.5% of South Carolinian newborn males, 2.5% South Carolinian male college students, 5% Italian male controls and 4.5% of the white South African controls. None of the black South African non-fragile X individuals with mental retardation, the fragile X or the control samples tested carried the duplication, suggesting that the duplication is rare in the black South African population. The incidence of the duplication was not significantly different between any of the groups in the study. Therefore, results of our studies in four different populations do not corroborate the findings of the previous study, and indicate that the HOPA dodecamer duplication does not convey an increased susceptibility to mental retardation.

Adult↗

Protein complexity, gene duplicability and gene dispensability in the yeast genome.

Using functional genomic and protein structural data we studied the effects of protein complexity (here defined as the number of subunit types in a protein) on gene dispensability and gene duplicability. We found that in terms of gene duplicability the major distinction in protein complexity is between hetero-complexes, each of which includes at least two different types of subunits (polypeptides), and homo-complexes, which include monomers and complexes that consist of only subunits of one polypeptide type. However, gene dispensability decreases only gradually as the number of subunit types in a protein complex increases. These observations suggest that the dosage balance hypothesis can explain well gene duplicability of complex proteins, but cannot completely explain the difference in dispensabilities between hetero-complex subunits. It is likely that knocking out a gene coding for a hetero-complex subunit would disrupt the function of the whole complex, so that the deletion effect on fitness would increase with protein complexity. We also found that multi-domain polypeptide genes are less dispensable but more duplicable than single-domain polypeptide genes. Duplicate genes derived from the whole genome duplication event in yeast are more dispensable (except for ribosomal protein genes) than other duplicate genes. Further, we found that subunits of the same protein complex tend to have similar expression levels and similar effects of gene deletion on fitness. Finally, we estimated that in yeast the contribution of duplicate genes to genetic robustness against null mutation is approximately 9%, smaller than previously estimated. In yeast, protein complexity may serve as a better indicator of gene dispensability than do duplicate genes.

Computational Biology↗

Dealing with saturation at the amino acid level: a case study based on anciently duplicated zebrafish genes.

The ray-finned fishes (Actinopterygii) seem to have two copies of many tetrapod (Sarcopterygii) genes. The origin of these duplicate fish genes is the subject of some controversy. One explanation for the existence of these extra fish genes could be an increase in the rate of independent gene duplications in fishes. Alternatively, gene duplicates in fish may have been formed in the ancestor of all or most Actinopterygii during a complete genome duplication event. A third possibility is that tetrapods have lost more genes than fish after gene or genome duplication events in the common ancestor of both lineages. These three hypotheses can be tested by phylogenetic reconstruction. Previously, we found that a large number of anciently duplicated genes of zebrafish are sister sequences in evolutionary trees suggesting that they were produced in Actinopterygii after the divergence of Sarcopterygii [Phil. Trans. R. Soc. Lond. B 356 (2001) 119]. On the other hand, several well-supported trees showed one of the two fish genes as the sister sequence to a monophyletic clade that included the second fish gene and genes from frog, chicken, mouse and human. These so-called outgroup topologies suggest that the origin of many fish duplicates predates the divergence of the Sarcopterygii and Actinopterygii and support the hypothesis that tetrapods have lost duplicates that have been retained in fish. Here we show that many of these 'outgroup' tree topologies are erroneous and can be corrected when mutational saturation is taken into account. To this end, a Java-based application has been developed to visualize the amount of saturation in amino acid sequences. The program graphically displays the number of observed frequent and rare amino acid replacements between pairs of sequences against their overall evolutionary distance. Discrimination between frequent and rare amino acid replacements is based on substitution probability matrices (e.g. PAM and BLOSUM). Evolutionary distances between sequences can be computed from the fraction of unsaturated sites only and evolutionary trees inferred by pairwise distance methods. When trees are computed by omitting the saturated fraction of sites, most fish duplicates are sister sequences.

Amino Acids↗

A genome-wide comparison of recent chimpanzee and human segmental duplications.

We present a global comparison of differences in content of segmental duplication between human and chimpanzee, and determine that 33% of human duplications (> 94% sequence identity) are not duplicated in chimpanzee, including some human disease-causing duplications. Combining experimental and computational approaches, we estimate a genomic duplication rate of 4-5 megabases per million years since divergence. These changes have resulted in gene expression differences between the species. In terms of numbers of base pairs affected, we determine that de novo duplication has contributed most significantly to differences between the species, followed by deletion of ancestral duplications. Post-speciation gene conversion accounts for less than 10% of recent segmental duplication. Chimpanzee-specific hyperexpansion (> 100 copies) of particular segments of DNA have resulted in marked quantitative differences and alterations in the genome landscape between chimpanzee and human. Almost all of the most extreme differences relate to changes in chromosome structure, including the emergence of African great ape subterminal heterochromatin. Nevertheless, base per base, large segmental duplication events have had a greater impact (2.7%) in altering the genomic landscape of these two species than single-base-pair substitution (1.2%).

Animals↗

Gene regulatory network growth by duplication.

We are beginning to elucidate transcriptional regulatory networks on a large scale and to understand some of the structural principles of these networks, but the evolutionary mechanisms that form these networks are still mostly unknown. Here we investigate the role of gene duplication in network evolution. Gene duplication is the driving force for creating new genes in genomes: at least 50% of prokaryotic genes and over 90% of eukaryotic genes are products of gene duplication. The transcriptional interactions in regulatory networks consist of multiple components, and duplication processes that generate new interactions would need to be more complex. We define possible duplication scenarios and show that they formed the regulatory networks of the prokaryote Escherichia coli and the eukaryote Saccharomyces cerevisiae. Gene duplication has had a key role in network evolution: more than one-third of known regulatory interactions were inherited from the ancestral transcription factor or target gene after duplication, and roughly one-half of the interactions were gained during divergence after duplication. In addition, we conclude that evolution has been incremental, rather than making entire regulatory circuits or motifs by duplication with inheritance of interactions.

Bacterial Proteins↗

X inactivation phenotype in carriers of Pelizaeus-Merzbacher disease: skewed in carriers of a duplication and random in carriers of point mutations.

Pelizaeus-Merzbacher disease (PMD) is an X-linked recessive disease caused by coding sequence mutations in the PLP gene, sub-microscopic duplications of variable sizes including the PLP gene or very rarely deletions of the PLP gene. We analysed the X inactivation pattern in blood of PMD female carriers with duplications and with point mutations. In the majority of duplication carriers (7/11), the X chromosome bearing the duplication was preferentially inactivated, whereas a random pattern of X inactivation was detected in point mutation carriers (3/3), a deletion carrier (1/1), affected females (4/4) who did not have a recognised mutation and normal control females. However 2/5 non-carrier female relatives of patients with a duplication, had skewed X inactivation. The skewed pattern of inactivation observed in most duplication carriers and not in mutation carriers suggests a) that there is selection against those cells in which the duplicated X chromosome is active and b) other expressed sequences within the duplicated region rather than mutant PLP may be responsible. Since the skewed X inactivation did not segregate with the disease in two families and the pattern of X inactivation was variable among the duplication carriers, the pattern X inactivation is an unsuitable diagnostic tool for female carriers of PMD.

DNA-Binding Proteins↗

Duplications and copy number variants of 8p23.1 are cytogenetically indistinguishable but distinct at the molecular level.

It has been proposed that duplications of 8p23.1 are either euchromatic variants of the 8p23.1 defensin domain with no phenotypic consequences or true duplications associated with developmental delay and heart defects. Here, we provide evidence for both alternatives in two new families. A duplication of most of band 8p23.1 (circa 5 Mb) was found in a girl of 8 years with pulmonary stenosis and mild language delay. BAC fluorescence in situ hybridisation (FISH) and multiplex amplifiable probe hybridisation (MAPH) showed that the two copies of the duplicated segment were sited, in an alternating fashion, between three copies of a circa 300-450 kb segment from 8p23.1 distal to REPD. Copy number of the variable 8p23.1 defensin domain was consistent with duplication but within the normal range. Duplication of the GATA-binding protein 4 gene (GATA4) in this patient and others with and without heart defects, suggests it is a dosage-sensitive gene with variable penetrance. A cytogenetically similar duplication of 8p23.1 was found at prenatal diagnosis in a fetus, father and grandmother. There was no duplication using BAC FISH but MAPH showed 11 copies of the 360 kb variable defensin domain which is within the expanded range found in previous euchromatic variant carriers. Semiquantitative FISH (SQ-FISH) was consistent with a simultaneous expansion of the adjacent olfactory receptor repeats. These results distinguish duplications of 8p23.1 with clinically significant consequences from benign copy number variants, which have not yet been associated with qualitative or quantitative traits.

Child↗

Quantifying the mechanisms for segmental duplications in mammalian genomes by statistical analysis and modeling.

A large number of the segmental duplications in mammalian genomes have been cataloged by genome-wide sequence analyses. The molecular mechanisms involved in these duplications mostly remain a matter of speculation. To uncover, test, and further quantify the hypotheses on the mechanisms for the recent duplications in the mammalian genomes, we have performed a series of statistical analyses on the sequences flanking the duplicated segments and proposed a dynamic model for the duplication process. The model, when applied to the human duplication data, indicates that approximately 30% of the recent human segmental duplications were caused by a recombination-like mechanism, among which 12% were mediated by the most recently active repeat, Alu. But a significant proportion of the duplications are caused by some mechanism independent of the repeat distribution. A less sure but similar picture is found in the rodent genomes. A further analysis on the physical features of the flanking sequences suggests that one of the uncharacterized duplication mechanisms shared by the mammalian genomes is surprisingly well correlated with the physical instability in the DNA sequences.

Animals↗

The combinatorics of tandem duplication trees.

We developed a recurrence relation that counts the number of tandem duplication trees (either rooted or unrooted) that are consistent with a set of n tandemly repeated sequences generated under the standard unequal recombination (or crossover) model of tandem duplications. The number of rooted duplication trees is exactly twice the number of unrooted trees, which means that on average only two positions for a root on a duplication tree are possible. Using the recurrence, we tabulated these numbers for small values of n. We also developed an asymptotic formula that for large n provides estimates for these numbers. These numbers give a priori probabilities for phylogenies of the repeated sequences to be duplication trees. This work extends earlier studies where exhaustive counts of the numbers for small n were obtained. One application showed the significance of finding that most maximum-parsimony trees constructed from repeat sequences from human immunoglobins and T-cell receptors were tandem duplication trees. Those findings provided strong support to the proposed mechanisms of tandem gene duplication. The recurrence relation also suggests efficient algorithms to recognize duplication trees and to generate random duplication trees for simulation. We present a linear-time recognition algorithm.

Algorithms↗

The probability of preservation of a newly arisen gene duplicate.

Newly emerging data from genome sequencing projects suggest that gene duplication, often accompanied by genetic map changes, is a common and ongoing feature of all genomes. This raises the possibility that differential expansion/contraction of various genomic sequences may be just as important a mechanism of phenotypic evolution as changes at the nucleotide level. However, the population-genetic mechanisms responsible for the success vs. failure of newly arisen gene duplicates are poorly understood. We examine the influence of various aspects of gene structure, mutation rates, degree of linkage, and population size (N) on the joint fate of a newly arisen duplicate gene and its ancestral locus. Unless there is active selection against duplicate genes, the probability of permanent establishment of such genes is usually no less than 1/(4N) (half of the neutral expectation), and it can be orders of magnitude greater if neofunctionalizing mutations are common. The probability of a map change (reassignment of a key function of an ancestral locus to a new chromosomal location) induced by a newly arisen duplicate is also generally >1/(4N) for unlinked duplicates, suggesting that recurrent gene duplication and alternative silencing may be a common mechanism for generating microchromosomal rearrangements responsible for postreproductive isolating barriers among species. Relative to subfunctionalization, neofunctionalization is expected to become a progressively more important mechanism of duplicate-gene preservation in populations with increasing size. However, even in large populations, the probability of neofunctionalization scales only with the square of the selective advantage. Tight linkage also influences the probability of duplicate-gene preservation, increasing the probability of subfunctionalization but decreasing the probability of neofunctionalization.

Alleles↗

Allele-specific expression analysis by RNA-FISH demonstrates preferential maternal expression of UBE3A and imprint maintenance within 15q11- q13 duplications.

15q11- q13 contains many imprinted genes, and undergoes duplicon-mediated rearrangements, including deletions, duplications and triplications, and generation of marker chromosomes. Abnormal phenotypes, including language delays and autism spectrum disorders, are primarily observed with maternal 15q11- q13 duplication. To determine possible epigenetic effects on expression within duplicated 15q11- q13 regions, we utilized RNA-FISH to directly observe gene expression. RNA-FISH, unlike RT-PCR, is polymorphism-independent, and it also detects relative levels of expression at each allele. Unamplified, gene-specific RNA signals were detected using cDNA probes. Subsequent DNA-FISH confirmed RNA signals and assigned parental origin by colocalization of genomic probes. SNRPN and NDN expression was detected primarily from paternal alleles. Control Dystrobrevin transcripts were detected equally from both alleles; however, maternal-UBE3A signals were consistently larger than paternal signals in normal fibroblasts and in neural-precursor cells. Larger UBE3A signals were also observed on one or both maternal alleles in a cell line carrying a maternal interstitial duplication, on both alleles of a maternally derived marker(15) chromosome, and occasionally on a paternal allele in a cell line carrying a paternal interstitial duplication. Expression of NDNL2, just distal to the duplicated region, was not markedly altered but paralleled changes in UBE3A expression. Excess total maternal-UBE3A RNA was confirmed by Northern blot analysis of cell lines carrying 15q11- q13 duplications or triplications. These results demonstrate that: (1) UBE3A is imprinted in fibroblasts, lymphoblasts and neural-precursor cells; (2) allelic imprint status is maintained in the majority of cells upon duplication both in cis and in trans; and (3) alleles on specific types of duplications may exhibit an increase in expression levels/loss of expression constraints.

Autoantigens↗

Phylogenetic dating and characterization of gene duplications in vertebrates: the cartilaginous fish reference.

Vertebrates originated in the lower Cambrian. Their diversification and morphological innovations have been attributed to large-scale gene or genome duplications at the origin of the group. These duplications are predicted to have occurred in two rounds, the "2R" hypothesis, or they may have occurred in one genome duplication plus many segmental duplications, although these hypotheses are disputed. Under such models, most genes that are duplicated in all vertebrates should have originated during the same period. Previous work has shown that indeed duplications started after the speciation between vertebrates and the closest invertebrate, amphioxus, but have not set a clear ending. Consideration of chordate phylogeny immediately shows the key position of cartilaginous vertebrates (Chondrichthyes) to answer this question. Did gene duplications occur as frequently during the 45 Myr between the cartilaginous/bony vertebrate split and the fish/tetrapode split as in the previous approximately 100 Myr? Although the time interval is relatively short, it is crucial to understanding the events at the origin of vertebrates. By a systematic appraisal of gene phylogenies, we show that significantly more duplications occurred before than after the cartilaginous/bony vertebrate split. Our results support rounds of gene or genome duplications during a limited period of early vertebrate evolution and allow a better characterization of these events.

Animals↗

Divergence pattern of duplicate genes in protein-protein interactions follows the power law.

The impact of the biological network structures on the divergence between the two copies of one duplicate gene pair involved in the networks has not been documented on a genome scale. Having analyzed the most recently updated Database of Interacting Proteins (DIP) by incorporating the information for duplicate genes of the same age in yeast, we find that there was a highly significantly positive correlation between the level of connectivity of ancient genes and the number of shared partners of their duplicates in the protein-protein interaction networks. This suggests that duplicate genes with a low ancestral connectivity tend to provide raw materials for functional novelty, whereas those duplicate genes with a high ancestral connectivity tend to create functional redundancy for a genome during the same evolutionary period. Moreover, the difference in the number of partners between two copies of a duplicate pair was found to follow a power-law distribution. This suggests that loss and gain of interacting partners for most duplicate genes with a lower level of ancestral connectivity is largely symmetrical, whereas the "hub duplicate genes" with a higher level of ancient connectivity display an asymmetrical divergence pattern in protein-protein interactions. Thus, it is clear that the protein-protein interaction network structures affect the divergence pattern of duplicate genes. Our findings also provide insights into the origin and development of biological networks.

Databases, Protein↗

Evaluation of whether accelerated protein evolution in chordates has occurred before, after, or simultaneously with gene duplication.

Gene duplication and loss are predicted to be at least of the order of the substitution rate and are key contributors to the development of novel gene function and overall genome evolution. Although it has been established that proteins evolve more rapidly after gene duplication, we were interested in testing to what extent this reflects causation or association. Therefore, we investigated the rate of evolution prior to gene duplication in chordates. Two patterns emerged; firstly, branches, which are both preceded by a duplication and followed by a duplication, display an elevated rate of amino acid replacement. This is reflected in the ratio of nonsynonymous to synonymous substitution (mean nonsynonymous to synonymous nucleotide substitution rate ratio [Ka:Ks]) of 0.44 compared with branches preceded by and followed by a speciation (mean Ka:Ks of 0.23). The observed patterns suggest that there can be simultaneous alteration in the selection pressures on both gene duplication and amino acid replacement, which may be consistent with co-occurring increases in positive selection, or alternatively with concurrent relaxation of purifying selection. The pattern is largely, but perhaps not completely, explained by the existence of certain families that have elevated rates of both gene duplication and amino acid replacement. Secondly, we observed accelerated amino acid replacement prior to duplication (mean Ka:Ks for postspeciation preduplication branches was 0.27). In some cases, this could reflect adaptive changes in protein function precipitating a gene duplication event. In conclusion, the circumstances surrounding the birth of new proteins may frequently involve a simultaneous change in selection pressures on both gene-copy number and amino acid replacement. More precise modeling of the relative importance of preduplication, postduplication, and simultaneous amino acid replacement will require larger and denser genomic data sets from multiple species, allowing simultaneous estimation of lineage-specific fluctuations in mutation rates and adaptive constraints.

Amino Acid Substitution↗

Not born equal: increased rate asymmetry in relocated and retrotransposed rodent gene duplicates.

Duplicated genes frequently evolve at different rates. This asymmetry is evidence of natural selection's ability to discriminate between the 2 copies, subjecting them to different levels of purifying selection or even permitting adaptive evolution of one or both copies. However, if gene duplication creates pairs of protein-coding sequences that are initially identical, this raises the question of how selection tells the 2 copies apart. Here, we investigated asymmetric sequence divergence of recently duplicated genes in rodents and related this to 2 possible sources of such asymmetry: gene relocation as a consequence of duplication and retrotransposition as a mechanism of gene duplication. We found that most young rodent duplicates that have been relocated were created by retrotransposition. The degree of rate asymmetry in gene pairs where one copy has been relocated (either by retrotransposition or DNA-based duplication) is greater than in pairs formed by local DNA-based duplication events. Furthermore, by considering the direction of transposition for distant duplicates, we found a consistent tendency for retrogenes to undergo accelerated protein evolution relative to their static paralogs, whereas DNA-based transpositions showed no such tendency. Finally, we demonstrate that the faster sequence evolution of retrogenes correlates with the profound alteration of their expression pattern that is precipitated by retrotransposition.

Animals↗

Duplication processes in Saccharomyces cerevisiae haploid strains.

Duplication is thought to be one of the main processes providing a substrate on which the effects of evolution are visible. The mechanisms underlying this chromosomal rearrangement were investigated here in the yeast Saccharomyces cerevisiae. Spontaneous revertants containing a duplication event were selected and analyzed. In addition to the single gene duplication described in a previous study, we demonstrated here that direct tandem duplicated regions ranging from 5 to 90 kb in size can also occur spontaneously. To further investigate the mechanisms in the duplication events, we examined whether homologous recombination contributes to these processes. The results obtained show that the mechanisms involved in segmental duplication are RAD52-independent, contrary to those involved in single gene duplication. Moreover, this study shows that the duplication of a given gene can occur in S.cerevisiae haploid strains via at least two ways: single gene or segmental duplication.

Aspartate Carbamoyltransferase↗