Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Duplication”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Role of positive selection in the retention of duplicate genes in mammalian genomes.

The question of how duplicate genes are retained in a population remains controversial. The duplication-degeneration-complementation model, which involves no positive selection, stipulates a higher retention rate of duplicate genes in a small population than in a large one. This model has been accepted by many evolutionists. However, we found considerably more retentions and fewer losses of duplicate genes in the mouse genome than in the human genome, although the population size of rodents is in general larger than that of primates. Indeed, in nearly every interval of synonymous divergence between duplicate genes, the number of gene retentions in mouse is larger than that in human. Our findings suggest a more important role of positive selection in duplicate retention than duplication-degeneration-complementation. In addition, certain functional categories show a higher tendency of lineage-specific expansion than expected, suggesting lineage-specific selection or functional bias in retained duplicates.

Animals↗

HSDSnake: a user-friendly SnakeMake pipeline for analysis of duplicate genes in eukaryotic genomes.

SUMMARY: Gene duplication is a well-known driver of molecular evolution-it acts as a source of genetic novelty, thereby providing the raw substrate for organismal adaption. However, detecting different types of gene duplicates and comparing them in sequence datasets can be difficult. Existing tools can identify and classify gene duplicates that have arisen by various processes, but have limitations; for example, some do not have a user-friendly workflow and can include many intermediate steps requiring manual adjustments of parameters and/or are not maintained for the benefit of research community members. Here, we have developed HSDSnake, a user-friendly SnakeMake pipeline that can detect and classify gene duplications into five categories: dispersed, proximal, tandem, transposed, and whole genome. It also curates and evaluates the highly similar gene duplicates (HSDs) in each gene duplication category with reliance on both sequence similarity and conserved domains. Lastly, the detected gene duplicates can be visualized within a KEGG functional pathway framework and the substitution rates (Ka, Ks, and their Ka/Ks ratio) can be analyzed for all the duplicate gene pairs. We demonstrate HSDSnake's capabilities by analyzing two reference genomes directly downloaded from NCBI and provide detailed instructions for each step. AVAILABILITY AND IMPLEMENTATION: The HSDSnake pipeline uses SnakeMake and Conda to run and install dependencies. The distribution version is available online at GitHub: https://github.com/zx0223winner/HSDSnake and the archived version at Zenodo is https://doi.org/10.5281/zenodo.15521945.

Software↗

Array comparative genomic hybridisation analysis of boys with X linked hypopituitarism identifies a 3.9 Mb duplicated critical region at Xq27 containing SOX3.

INTRODUCTION: Array comparative genomic hybridisation (array CGH) is a powerful method that detects alteration of gene copy number with greater resolution and efficiency than traditional methods. However, its ability to detect disease causing duplications in constitutional genomic DNA has not been shown. We developed an array CGH assay for X linked hypopituitarism, which is associated with duplication of Xq26-q27. METHODS: We generated custom BAC/PAC arrays that spanned the 7.3 Mb critical region at Xq26.1-q27.3, and used them to search for duplications in three previously uncharacterised families with X linked hypopituitarism. RESULTS: Validation experiments clearly identified Xq26-q27 duplications that we had previously mapped by fluorescence in situ hybridisation. Array CGH analysis of novel XH families identified three different Xq26-q27 duplications, which together refine the critical region to a 3.9 Mb interval at Xq27.2-q27.3. Expression analysis of six orthologous mouse genes from this region revealed that the transcription factor Sox3 is expressed at 11.5 and 12.5 days after conception in the infundibulum of the developing pituitary and the presumptive hypothalamus. DISCUSSION: Array CGH is a robust and sensitive method for identifying X chromosome duplications. The existence of different, overlapping Xq duplications in five kindreds indicates that X linked hypopituitarism is caused by increased gene dosage. Interestingly, all X linked hypopituitarism duplications contain SOX3. As mutation of this gene in human beings and mice results in hypopituitarism, we hypothesise that increased dosage of Sox3 causes perturbation of pituitary and hypothalamic development and may be the causative mechanism for X linked hypopituitarism.

Adolescent↗

The pericentromeric region of human chromosome 11: evidence for a chromosome-specific duplication.

We have identified a chromosome duplication in the pericentromeric region of human chromosome 11 located in 11p11 and 11q14. A detailed physical map of each duplicated region was generated to describe the nature of the duplication, the involvement at the centromere and to resolve the correct maps. All clones were evaluated to ensure they were representative of their genetic origin. The order of clones, based on their marker content, as well as the distance covered was determined by SEGMAP. Each duplication encompasses more than 1 Mb of DNA and appears to be chromosome 11 specific. Ten STS markers were mapped within each duplication. Comparative sequence analysis along the duplication identified 35 nucleotide changes in 2,036 bp between the two copies, suggesting the duplication occurred over 14 million years ago. A suggested organization of the pericentromeric region, including the duplications and alpha-related repetitive sequences, is presented.

Centromere↗

Preferential duplication of conserved proteins in eukaryotic genomes.

A central goal in genome biology is to understand the origin and maintenance of genic diversity. Over evolutionary time, each gene's contribution to the genic content of an organism depends not only on its probability of long-term survival, but also on its propensity to generate duplicates that are themselves capable of long-term survival. In this study we investigate which types of genes are likely to generate functional and persistent duplicates. We demonstrate that genes that have generated duplicates in the C. elegans and S. cerevisiae genomes were 25%-50% more constrained prior to duplication than the genes that failed to leave duplicates. We further show that conserved genes have been consistently prolific in generating duplicates for hundreds of millions of years in these two species. These findings reveal one way in which gene duplication shapes the content of eukaryotic genomes. Our finding that the set of duplicate genes is biased has important implications for genome-scale studies.

Animals↗

The duplication of an eight-residue helical stretch in Staphylococcal nuclease is not helical: a model for evolutionary change.

A common method of evolutionary change is gene duplication, followed by other events that lead to new function, decoration of folds, oligomerization, or other changes. As part of a study on the potential for evolutionary change created by duplicated sequences, we have carried out a crystallographic study on a mutant of Staphylococcal nuclease in which residues 55-62 have been duplicated in a wild-type variant termed PHS. In the parental protein (PHS) these residues form the first two turns of a helix running from residue 54 to 68 (hereafter designated as helix I). The crystal structure of the mutant is very similar to that of the parental, with helix I being unaltered. The duplicated residues are accommodated by expanding an existing loop N-terminal to helix I. In addition, circular dichroism (CD) studies have been carried out on a parental peptide containing helix I with six flanking residues at each terminus (residues 48-74) and on the same peptide expanded by the duplication, as a function of 2,2,2-trifluoroethanol (TFE) concentration. Each peptide possesses only modest helical propensity in solution. Our data, which is different from what was observed in T4 lysozyme, show that the conformation of the duplicated sequence is determined by a balance of sequential and longer-range effects. Thus duplicating sequence need not mean duplicating structure. Proteins 2000;40:465-472.

Algorithms↗

Pathogenic role of mtDNA duplications in mitochondrial diseases associated with mtDNA deletions.

We estimated the frequency of multiple mtDNA rearrangements by Southern blot in 32 patients affected by mitochondrial disorders associated with single deletions in order to assess genotype-phenotype correlations and elucidate the pathogenic significance of mtDNA duplications. Muscle in situ hybridization studies were performed in patients showing mtDNA duplications at Southern blot. We found multiple rearrangements in 12/32 (37.5%) patients; in particular, mtDNA duplications were detected in 4/4 Kearns-Sayre syndrome (KSS), in 1 Pearson's syndrome, in 1/3 encephalomyopathies with progressive external ophthalmoplegia (PEO), and in 2/23 PEO. In situ studies documented an exclusive accumulation of deleted mtDNAs in cytochrome c oxidase negative fibers of patients with mtDNA duplications. The presence of mtDNA duplications significantly correlated with onset of symptoms before age 15 and occurrence of clinical multisystem involvement. Analysis of biochemical data documented a predominant reduction of complex III in patients without duplications compared to patients with mtDNA duplications. Our data indicate that multiple mtDNA rearrangements are detectable in a considerable proportion of patients with single deletions and that mtDNA duplications do not cause any oxidative impairment. They more likely play a pathogenic role in the determination of clinical expression of mitochondrial diseases associated with single mtDNA deletions, possibly generating deleted mtDNAs in embryonic tissues by homologous recombination.

Adolescent↗

Genomic instability in multiple myeloma: evidence for jumping segmental duplications of chromosome arm 1q.

Multiple myeloma (MM) is a malignant plasma cell disorder characterized by complex karyotypes and chromosome 1 instability at the cytogenetic level. Chromosome 1 instability generally involves partial duplications, whole-arm translocations, or jumping translocations of 1q, identified by G-banding. To characterize this instability further, we performed spectral karyotyping and fluorescence in situ hybridization with probes for satII/III (1q12), BCL9 (1q21), and IL6R (1q21) on the karyotypes of 44 patients with known 1q aberrations. In eight patients, segmental duplication of 1q12-21 and adjacent bands occurred on nonhomologous chromosomes. In five cases, the 1q first jumped to a nonhomologous chromosome, after which the 1q12-21 segment again duplicated itself 1-3 times. In three other cases, segmental duplications occurred after the 1q first jumped to a nonhomologous chromosome, where the proximal adjacent nonhomologous chromosome segment was duplicated prior to the 1q jumping or inserting itself into a new location. These cases demonstrate that satII/III DNA sequences are not only associated not only with the duplication of adjacent distal chromosome segments after translocation, but are also associated with the duplication and jumping/insertion of proximal nonhomologous chromosome segments. We have designated this type of instability as a jumping segmental duplication.

Adult↗

Rapid evolution in a pair of recent duplicate segments of rice.

Gene duplication has been considered the most important way of generating genetic novelties. The subsequent evolution right after gene duplication is critical for new function to occur. Here we analyzed the evolutionary pattern for a recently duplicated segment between rice chromosomes 11 and 12. This duplication event was estimated to occur about 6 million years ago, during the divergence of the B- and C-genome rice species. The duplicate segment in chromosome 12 has significantly higher frequency of sequence rearrangement rate than non-duplicated regions. The rearrangement rate is approximately 6.5 breakages/Mb per million years, about six times higher than the fastest rate ever reported in eukaryotes. The genes within both segments experienced accelerated nucleotide substitution rates revealed by synonymous (Ks) and non-synonymous divergence (Ka) between Oryza sativa indica and O. sativa japonica. Analysis using EST data also implicates rapid divergence in expression between these segmental duplicate genes. These overall rapid changes from different perspective for the first time provide evidence that relaxation of selection also occurs in large-scale duplications.

Base Sequence↗

A large duplication associated with dominant white color in pigs originated by homologous recombination between LINE elements flanking KIT.

The Dominant White (I/KIT) locus is one of the major coat color loci in the pig. Previous studies showed that the Dominant White (I) and Patch (IP) alleles are both associated with a duplication including the entire KIT coding sequence. We have now constructed a BAC contig spanning the three closely linked tyrosine kinase receptor genes PDGFRA-KIT-KDR. The size of the duplication was estimated at about 450 kb and includes KIT, but not PDGFRA and KDR. Sequence analysis revealed that the duplication arose by unequal homologous recombination between two LINE elements flanking KIT. The same unique duplication breakpoint was identified in animals carrying the I and IP alleles across breeds, implying that Dominant White and Patch alleles are descendants of a single duplication event. An unexpected finding was that Piétrain pigs carry the KIT duplication, since this breed was previously assumed to be wild type at this locus. Comparative sequence analysis indicated that the distinct phenotypic effect of the duplication occurs because the duplicated copy lacks some regulatory elements located more than 150 kb upstream of KIT exon 1 and necessary for normal KIT expression.

Alleles↗

The evolution of engrailed genes after duplication and speciation events.

Members of the engrailed class encode transcription factors involved in major steps of metazoan development. Few developmental regulatory genes have been studied in such a wide range of animals. Furthermore duplications of an ancestral engrailed gene independently generated multiple engrailed paralogues in several organisms. This offers the opportunity to reconstruct the evolution of the engrailed family and to study the processes involved in the functional diversification following speciation or duplication events. The ancestral function of engrailedis very likely involved in neurogenesis. Recent studies in Drosophila and mice have shown its crucial role in neuronal connectivity and neuromuscular targeting. engrailed was probably recruited very early for a role in segmentation through intercalary evolution. Several new functions were acquired later on in specific phyla. Some duplication events have been followed by the loss of one paralogue, whereas others have led to the functional diversification of the paralogues. The Duplication-Degenerescence-Complementation model recently proposed by Force et al. seems to be the main process involved in functional diversification after duplication events. This does not exclude acquisition of new functions for one or both paralogues after duplication. The acquisition of such new functions principally involves the evolution of cis-regulatory sequences, but evolution of the coding sequence has also been revealed. However, in all engrailed duplications studied, even in ancient chromosomal duplications, the paralogues have kept redundant functions. In fact, selection seems to maintain a certain redundancy between engrailed paralogues.

Animals↗

Sharing of transcription factors after gene duplication in the yeast Saccharomyces cerevisiae.

In a set of 190 duplicate gene pairs in yeast Saccharomyces cerevisiae, the sharing of transcription factors tended to decrease with increased divergence in coding sequence, at both synonymous and nonsynonymous sites. Our results showed a significantly higher sharing of transcription factors by duplicated gene pairs falling within duplicated genomic blocks than in other duplicated gene pairs; and genes in duplicated blocks also showed significantly greater conservation at the coding sequence level. In spite of the overall trends, there were certain gene pairs, both in duplicated blocks and in other genomic regions, which were highly divergent in coding sequence and yet had identical patterns of transcription factor binding. These results suggest that functional differentiation of genes after duplication is a multi-dimensional process, with different duplicate pairs differentiating in different ways.

Base Sequence↗

Diagnosing duplications--can it be done?

New genes arise through duplication and modification of DNA sequences on a range of scales: single gene duplication, duplication of large chromosomal fragments and whole-genome duplication. Each duplication mechanism has specific characteristics that influence the fate of the resulting duplicates, such as the size of the duplicated fragment, the potential for dosage imbalance, the preservation or disruption of regulatory control and genomic context. The ability to diagnose or identify the mechanism that produced a pair of paralogs has the potential to increase our ability to reconstruct evolutionary history, to understand the processes that govern genome evolution and to make functional predictions based on paralogy. The recent availability of large amounts of whole-genome sequence, often from several closely related species, has stimulated a wealth of new computational methods to diagnose gene duplications.

Animals↗

Recent duplication, domain accretion and the dynamic mutation of the human genome.

An estimated 5% of the human genome consists of interspersed duplications that have arisen over the past 35 million years of evolution. Two categories of such recently duplicated segments can be distinguished: segmental duplications between nonhomologous chromosomes (transchromosomal duplications) and duplications mainly restricted to a particular chromosome (chromosome-specific duplications). Many of these duplications exhibit an extraordinarily high degree of sequence identity at the nucleotide level (>95%) and span large genomic distances (1-100 kb). Preliminary analyses indicate that these same regions are targets for rapid evolutionary turnover among the genomes of closely related primates. The dynamic nature of these regions because of recurrent chromosomal rearrangement, and their ability to create fusion genes from juxtaposed cassettes suggest that duplicative transposition was an important force in the evolution of our genome.

Biological Evolution↗

Updated map of duplicated regions in the yeast genome.

We have updated the map of duplicated chromosomal segments in the Saccharomyces cerevisiae genome originally published by Wolfe and Shields in 1997 (Nature 387, 708-713). The new analysis is based on the more sensitive Smith Waterman search method instead of BLAST. The parameters used to identify duplicated chromosomal regions were optimized such as to maximize the amount of the genome placed into paired regions, under the assumption that the hypothesis that the entire genome was duplicated in a single event is correct. The core of the new map, with 52 pairs of regions containing three or more duplicated genes, is largely unchanged from our original map. 39 tRNA gene pairs and one snRNA pair have been added. To find additional pairs of genes that may have been formed by whole genome duplication, we searched through the parts of the genome that are not covered by this core map, looking for putative duplicated chromosomal regions containing only two duplicate genes instead of three, or having lower-scoring gene pairs. This approach identified a further 32 candidate paired regions, bringing the total number of protein-coding genes on the duplication map to 905 (16% of the proteome). The updated map suggests that a second copy of the ribosomal DNA array has been deleted from chromosome IV.

Chromosomes, Fungal↗

Gene and genome duplications in vertebrates: the one-to-four (-to-eight in fish) rule and the evolution of novel gene functions.

One important mechanism for functional innovation during evolution is the duplication of genes and entire genomes. Evidence is accumulating that during the evolution of vertebrates from early deuterostome ancestors entire genomes were duplicated through two rounds of duplications (the 'one-to-two-to-four' rule). The first genome duplication in chordate evolution might predate the Cambrian explosion. The second genome duplication possibly dates back to the early Devonian. Recent data suggest that later in the Devonian, the fish genome was duplicated for a third time to produce up to eight copies of the original deuterostome genome. This last duplication took place after the two major radiations of jawed vertebrate life, the ray-finned fish (Actinopterygia) and the sarcopterygian lineage, diverged. Therefore the sarcopterygian fish, which includes the coelacanth, lungfish and all land vertebrates such as amphibians, reptiles, birds and mammals, tend to have only half the number of genes compared with actinopterygian fish. Although many duplicated genes turned into pseudogenes, or even 'junk' DNA, many others evolved new functions particularly during development. The increased genetic complexity of fish might reflect their evolutionary success and diversity.

Animals↗

Tandem duplication mosaicism: characterization of a mosaic dup(5q) and review.

Mosaicism for tandem duplications is rare. Most patients reported had abnormal phenotypes of varying severity, depending on the chromosomal imbalance involved and the level of mosaicism. Post-zygotic unequal sister-chromatid exchange has been proposed as the main mechanism for tandem duplication mosaicism. However, previous molecular analyses have implicated both meiotic and post-zygotic origins for the duplication. We describe a newborn male who was originally diagnosed in utero with arrhythmia and tetralogy of Fallot. He had multiple dysmorphic features including telecanthus, blepharophimosis, high broad nasal bridge with a square-shaped nose, flat philtrum, thin upper lip, down-turned corners of the mouth, high-arched palate, micrognathia, asymmetric ears, and long, thin fingers and toes. Karyotyping of peripheral blood lymphocytes showed mosaicism for a tandem duplication of part of the long arm of one chromosome 5: mos46,XY,dup(5)(q13q33)[6]/46,XY[45]. Fibroblast cultures had the same mosaic karyotype with a higher frequency of the dup(5) clone: mos46,XY,dup(5)(q13q33)[9]/46,XY[21]. Fluorescence in situ hybridization analysis with a wcp5 confirmed the chromosome 5 origin of the additional material. Parental karyotypes were normal indicating a de novo origin of the dup(5) in the proband. Molecular analyses of chromosome 5 sequence-tagged-site (STS) markers in our family were consistent with a post-zygotic origin for the duplication. Therefore, mosaicism for tandem duplications can arise both through meiotic or mitotic errors, as a result of unequal crossing over or unequal sister-chromatid exchange, respectively. Our review indicates that mosaicism for tandem duplications is likely under-ascertained and that parental karyotyping of probands with non-mosaic tandem duplications should be performed.

Adult↗

Age distribution of human gene families shows significant roles of both large- and small-scale duplications in vertebrate evolution.

The classical (two-round) hypothesis of vertebrate genome duplication proposes two successive whole-genome duplication(s) (polyploidizations) predating the origin of fishes, a view now being seriously challenged. As the debate largely concerns the relative merits of the 'big-bang mode' theory (large-scale duplication) and the 'continuous mode' theory (constant creation by small-scale duplications), we tested whether a significant proportion of paralogous genes in the contemporary human genome was indeed generated in the early stage of vertebrate evolution. After an extensive search of major databases, we dated 1,739 gene duplication events from the phylogenetic analysis of 749 vertebrate gene families. We found a pattern characterized by two waves (I, II) and an ancient component. Wave I represents a recent gene family expansion by tandem or segmental duplications, whereas wave II, a rapid paralogous gene increase in the early stage of vertebrate evolution, supports the idea of genome duplication(s) (the big-bang mode). Further analysis indicated that large- and small-scale gene duplications both make a significant contribution during the early stage of vertebrate evolution to build the current hierarchy of the human proteome.

Animals↗