Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Duplication”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

The evolution of gene duplicates.

Gene and genome duplications have given rise to enormous variability among species in the number of genes within their genomes. Gene copies have in turn played important roles in adaptation, having been implicated in the evolution of the immune response, insecticide resistance, efficient protein synthesis, and vertebrate body plans. In this chapter, we discuss the life history of gene duplications, from their first appearance within a population, through the period during which they rise in frequency or disappear, to their long-term fate. At each phase, we discuss the evolutionary processes that have influenced the dynamics of gene duplications and shaped their ultimate roles within a population. We argue that there is no evidence that organisms have evolved strategies to promote gene duplication in order to permit adaptive evolution. In contrast, many mechanisms exist to silence or eliminate duplicated genes, suggesting that selection has acted largely to reduce the rate of gene duplication. We also argue that natural selection has functioned as an effective sieve, increasing the representation of beneficial gene duplicates among those that establish within a population and that play a long-term role in evolution. To refine our understanding of how selection acts on new gene duplications, we provide a model incorporating a single-copy gene, its gene duplicate, and selection either favoring heterozygotes or eliminating deleterious mutations. Although both forms of selection can increase the initial rate of spread of a gene duplicate, the efficacy with which they do so differs dramatically. Heterozygote advantage always increases the rate of spread and can have a large impact. In contrast, masking deleterious mutations never has a large effect on the rate of spread of the duplicate, and this minor effect can be negative as well as positive. In both cases, the degree of linkage between the two gene copies affects the rate of spread of the duplication. Finally, we discuss evolutionary processes that occur over longer periods after a gene duplication has become established within a population. These long-term processes include maintenance, inactivation, and diversification in function. Consideration of each of the short-term and long-term processes affecting duplicated genes illustrates the subtle ways in which selection has acted to shape genomic structure.

Animals↗

Duplicate publication in the field of otolaryngology-head and neck surgery.

OBJECTIVE: This study establishes the approximate prevalence and patterns of duplicate publication in the medical literature in the specialty of otolaryngology-head and neck surgery. STUDY DESIGN AND SETTING: All of the authors and articles published in the American Medical Association Archives of Otolaryngology-Head and Neck Surgery were identified and listed for an 8-year period. During this time, 1965 authors published 1082 articles in the Archives, and this same set of authors published a total of almost 50,000 articles during the 12-year period between January 1977 and December 1988. Of the same set of 1965 authors, we picked 1000 at random and found that they had published a total of 24,353 articles. The titles of these articles were then screened for similar titles, and when similarities were noted, the complete articles were obtained when possible and compared for the degree and pattern of duplicate publication. RESULTS: Of the 1000 authors studied, we found that 228 authors had published 938 articles with similar titles. We were able to obtain the full copy of 886 (94%) of the 938 articles in question, which were written by 226 (99%) of the 228 authors. We found that in the case of 25 authors, there was no duplication despite the similar titles, but in the case of 201 (20% of the 1000) authors, 644 articles were published with some degree of duplication (1.8% duplication rate). CONCLUSIONS: The most common duplicate publication involves sequential publication of very similar data and conclusions. Duplicate publications failed to reference prior articles by the same author 32% of the time or referenced the prior articles only partially (11% of the time). Artificial segmentation of a single study into multiple arbitrary segments composed 20% of the duplicate publication. Duplicate publication across different specialties was noted to account for 4% of the instances. Most of the authors duplicated only once or twice, and most duplicators do reference their prior publications. SIGNIFICANCE: Duplicate publication is an example of inappropriate academic conduct. Because it tarnishes the reputation of the duplicating author and represents an unfair practice in terms of displacing the work of others, efforts should continue to educate authors, particularly young academicians, to avoid the practice of duplicate publication.

Duplicate Publications as Topic↗

Evolution of duplicate control regions in the mitochondrial genomes of metazoa: a case study with Australasian Ixodes ticks.

To investigate the evolution pattern and phylogenetic utility of duplicate control regions (CRs) in mitochondrial (mt) genomes, we sequenced the entire mt genomes of three Ixodes species and part of the mt genomes of another 11 species. All the species from the Australasian lineage have duplicate CRs, whereas the other species have one CR. Sequence analyses indicate that the two CRs of the Australasian Ixodes ticks have evolved in concert in each species. In addition to the Australasian Ixodes ticks, species from seven other lineages of metazoa also have mt genomes with duplicate CRs. Accumulated mtDNA sequence data from these metazoans and two recent experiments on replication of mt genomes in human cell lines with duplicate CRs allowed us to re-examine four intriguing questions about the presence of duplicate CRs in the mt genomes of metazoa: (1) Why do some mt genomes, but not others, have duplicate CRs? (2) How did mt genomes with duplicate CRs evolve? (3) How could the nucleotide sequences of duplicate CRs remain identical or very similar over evolutionary time? (4) Are duplicate CRs phylogenetic markers? It appears that mt genomes with duplicate CRs have a selective advantage in replication over mt genomes with one CR. Tandem duplication followed by deletion of genes is the most plausible mechanism for the generation of mt genomes with duplicate CRs. Once duplicate CRs occur in an mt genome, they tend to evolve in concert, probably by gene conversion. However, there are lineages where gene conversion may not always occur, and, thus, the two CRs may evolve independently in these lineages. Duplicate CRs have much potential as phylogenetic markers at low taxonomic levels, such as within genera, within families, or among families, but not at high taxonomic levels, such as among orders.

Animals↗

Evolution of alternative splicing after gene duplication.

Alternative splicing and gene duplication are two major sources of proteomic function diversity. Here, we study the evolutionary trend of alternative splicing after gene duplication by analyzing the alternative splicing differences between duplicate genes. We observed that duplicate genes have fewer alternative splice (AS) forms than single-copy genes, and that a negative correlation exists between the mean number of AS forms and the gene family size. Interestingly, we found that the loss of alternative splicing in duplicate genes may occur shortly after the gene duplication. These results support the subfunctionization model of alternative splicing in the early stage after gene duplication. Further analysis of the alternative splicing distribution in human duplicate pairs showed the asymmetric evolution of alternative splicing after gene duplications; i.e., the AS forms between duplicates may differ dramatically. We therefore conclude that alternative splicing and gene duplication may not evolve independently. In the early stage after gene duplication, young duplicates may take over a certain amount of protein function diversity that previously was carried out by the alternative splicing mechanism. In the late stage, the gain and loss of alternative splicing seem to be independent between duplicates.

Alternative Splicing↗

Duplicated genes evolve slower than singletons despite the initial rate increase.

BACKGROUND: Gene duplication is an important mechanism that can lead to the emergence of new functions during evolution. The impact of duplication on the mode of gene evolution has been the subject of several theoretical and empirical comparative-genomic studies. It has been shown that, shortly after the duplication, genes seem to experience a considerable relaxation of purifying selection. RESULTS: Here we demonstrate two opposite effects of gene duplication on evolutionary rates. Sequence comparisons between paralogs show that, in accord with previous observations, a substantial acceleration in the evolution of paralogs occurs after duplication, presumably due to relaxation of purifying selection. The effect of gene duplication on evolutionary rate was also assessed by sequence comparison between orthologs that have paralogs (duplicates) and those that do not (singletons). It is shown that, in eukaryotes, duplicates, on average, evolve significantly slower than singletons. Eukaryotic ortholog evolutionary rates for duplicates are also negatively correlated with the number of paralogs per gene and the strength of selection between paralogs. A tally of annotated gene functions shows that duplicates tend to be enriched for proteins with known functions, particularly those involved in signaling and related cellular processes; by contrast, singletons include an over-abundance of poorly characterized proteins. CONCLUSIONS: These results suggest that whether or not a gene duplicate is retained by selection depends critically on the pre-existing functional utility of the protein encoded by the ancestral singleton. Duplicates of genes of a higher biological import, which are subject to strong functional constraints on the sequence, are retained relatively more often. Thus, the evolutionary trajectory of duplicated genes appears to be determined by two opposing trends, namely, the post-duplication rate acceleration and the generally slow evolutionary rate owing to the high level of functional constraints.

Animals↗

Difference in the centrosome duplication regulatory activity among p53 'hot spot' mutants: potential role of Ser 315 phosphorylation-dependent centrosome binding of p53.

The p53 tumor suppressor protein regulates centrosome duplication through multiple pathways, and p21(Waf1/Cip1) (Waf1), a major target of p53's transactivation function, has been shown to be one of the effectors. However, it had been unclear whether the p53's Waf1-independent centrosome duplication regulatory pathways require its transactivation function. In human cancers, specific residues of p53 are mutated at a high frequency. These 'hot spot' mutations abrogate p53's transactivation function. If p53 regulates centrosome duplication in a transactivation-independent manner, different 'hot spot' mutants may regulate centrosome duplication differently. To test this, we examined the effect of two 'hot spot' mutants (R175H and R249S) for their centrosome duplication regulatory activities. We found that R175H lost the ability to regulate centrosome duplication, while R249S partially retained it. Moreover, R249S associates with both unduplicated and duplicated centrosomes similar to wild-type p53, while R175H only associates with duplicated, but not unduplicated centrosomes. Since cyclin-dependent kinase 2 (CDK2) triggers initiation of centrosome duplication, and p53 is phosphorylated on Ser 315 by CDK2, we examined the p53 mutants with a replacement of Ser 315 to Ala (A) and Asp (D), both of which retain the transactivation function. We found that S315D retained a complete centrosome duplication activity, while S315A only partially retained it. Moreover, S315D associates with both unduplicated and duplicated centrosomes, while S315A associates with only duplicated, but not unduplicated centrosomes. Thus, p53 controls the centrosome duplication cycle both in transactivation-dependent and transactivation-independent manners, and the ability to bind to unduplicated centrosomes, which is controlled by phosphorylation on Ser 315, may be important for the overall p53-mediated regulation of centrosome duplication.

Aneuploidy↗

Alimentary tract duplications in children: report of 26 years' experience.

Duplications of the alimentary tract are one of the rare anomalies of the gastrointestinal system. Because of the wide spectrum of the signs and symptoms, preoperative diagnosis frequently cannot be made. A close familiarity with clinical and surgical characteristics provides appropriate management and treatment of duplications. A retrospective clinical study was conducted to evaluate clinical and surgical characteristics and the treatment of duplications of the alimentary tract. During a 26-year period between 1971 and 1997, 38 patients with duplications of alimentary tract underwent operation at the Hacettepe University Department of Pediatric Surgery. Forty-two duplications in 38 patients (20 male, 53%; 18 female, 47%) were encountered. Sixty-nine percent of the patients were symptomatic under the age of one year, with 24 percent presenting with symptoms in the neonatal period. There were one sublingual, nine intrathoracic (including 2 thoracoabdominal) and 32 intraabdominal duplications. Abdominal mass, abdominal distention, constipation, vomiting and respiratory distress were the most frequently encountered signs and symptoms. Plain thoracic and abdominal X-rays, ultrasonography, and computed tomography of the chest and abdomen were the most commonly used diagnostic radiological methods. Thirty-three duplications (79%) were spherical and nine (21%) were tubular. Multiple duplications were encountered in two patients (5.3%). Fourteen duplications (33%) contained heterotopic mucosa, mostly gastric type. More than one type of heterotopic mucosa in the same duplication was encountered in four duplications (10%). Additional malformations were encountered in 26 percent of patients. Six patients (15.8%) died from unrelated causes. The signs and symptoms vary among duplications. Signs and symptoms leading to diagnosis and surgery varied according to the age of patient, location of the duplication, type of mucosal lining, duration of disease and presence of complication. The ideal surgical treatment of duplication is complete excision. However, the other treatment options should be well known.

Digestive System Abnormalities↗

Modeling gene and genome duplications in eukaryotes.

Recent analysis of complete eukaryotic genome sequences has revealed that gene duplication has been rampant. Moreover, next to a continuous mode of gene duplication, in many eukaryotic organisms the complete genome has been duplicated in their evolutionary past. Such large-scale gene duplication events have been associated with important evolutionary transitions or major leaps in development and adaptive radiations of species. Here, we present an evolutionary model that simulates the duplication dynamics of genes, considering genome-wide duplication events and a continuous mode of gene duplication. Modeling the evolution of the different functional categories of genes assesses the importance of different duplication events for gene families involved in specific functions or processes. By applying our model to the Arabidopsis genome, for which there is compelling evidence for three whole-genome duplications, we show that gene loss is strikingly different for large-scale and small-scale duplication events and highly biased toward certain functional classes. We provide evidence that some categories of genes were almost exclusively expanded through large-scale gene duplication events. In particular, we show that the three whole-genome duplications in Arabidopsis have been directly responsible for >90% of the increase in transcription factors, signal transducers, and developmental genes in the last 350 million years. Our evolutionary model is widely applicable and can be used to evaluate different assumptions regarding small- or large-scale gene duplication events in eukaryotic genomes.

Arabidopsis↗

Human-specific duplication and mosaic transcripts: the recent paralogous structure of chromosome 22.

In recent decades, comparative chromosomal banding, chromosome painting, and gene-order studies have shown strong conservation of gross chromosome structure and gene order in mammals. However, findings from the human genome sequence suggest an unprecedented degree of recent (<35 million years ago) segmental duplication. This dynamism of segmental duplications has important implications in disease and evolution. Here we present a chromosome-wide view of the structure and evolution of the most highly homologous duplications (> or = 1 kb and > or = 90%) on chromosome 22. Overall, 10.8% (3.7/33.8 Mb) of chromosome 22 is duplicated, with an average sequence identity of 95.4%. To organize the duplications into tractable units, intron-exon structure and well-defined duplication boundaries were used to define 78 duplicated modules (minimally shared evolutionary segments) with 157 copies on chromosome 22. Analysis of these modules provides evidence for the creation or modification of 11 novel transcripts. Comparative FISH analyses of human, chimpanzee, gorilla, orangutan, and macaque reveal qualitative and quantitative differences in the distribution of these duplications--consistent with their recent origin. Several duplications appear to be human specific, including a approximately 400-kb duplication (99.4%-99.8% sequence identity) that transposed from chromosome 14 to the most proximal pericentromeric region of chromosome 22. Experimental and in silico data further support a pericentromeric gradient of duplications where the most recent duplications transpose adjacent to the centromere. Taken together, these data suggest that segmental duplications have been an ongoing process of primate genome evolution, contributing to recent gene innovation and the dynamic transformation of genome architecture within and among closely related species.

Animals↗

Recent segmental and gene duplications in the mouse genome.

BACKGROUND: The high quality of the mouse genome draft sequence and its associated annotations are an invaluable biological resource. Identifying recent duplications in the mouse genome, especially in regions containing genes, may highlight important events in recent murine evolution. In addition, detecting recent sequence duplications can reveal potentially problematic regions of the genome assembly. We use BLAST-based computational heuristics to identify large (>/= 5 kb) and recent (>/= 90% sequence identity) segmental duplications in the mouse genome sequence. Here we present a database of recently duplicated regions of the mouse genome found in the mouse genome sequencing consortium (MGSC) February 2002 and February 2003 assemblies. RESULTS: We determined that 33.6 Mb of 2,695 Mb (1.2%) of sequence from the February 2003 mouse genome sequence assembly is involved in recent segmental duplications, which is less than that observed in the human genome (around 3.5-5%). From this dataset, 8.9 Mb (26%) of the duplication content consisted of 'unmapped' chromosome sequence. Moreover, we suspect that an additional 18.5 Mb of sequence is involved in duplication artifacts arising from sequence misassignment errors in this genome assembly. By searching for genes that are located within these regions, we identified 675 genes that mapped to duplicated regions of the mouse genome. Sixteen of these genes appear to have been duplicated independently in the human genome. From our dataset we further characterized a 42 kb recent segmental duplication of Mater, a maternal-effect gene essential for embryogenesis in mice. CONCLUSION: Our results provide an initial analysis of the recently duplicated sequence and gene content of the mouse genome. Many of these duplicated loci, as well as regions identified to be involved in potential sequence misassignment errors, will require further mapping and sequencing to achieve accuracy. A Genome Browser database was set up to display the identified duplication content presented in this work. This data will also be relevant to the growing number of investigators who use the draft genome sequence for experimental design and analysis.

Animals↗

Duplications in the DMD gene.

The detection of duplications in Duchenne (DMD)/Becker Muscular Dystrophy (BMD) has long been a neglected issue. However, recent technological advancements have significantly simplified screening for such rearrangements. We report here the detection and analysis of 118 duplications in the DMD gene of DMD/BMD patients. In an unselected patient series the duplication frequency was 7%. In patients already screened for deletions and point mutations, duplications were detected in 87% of cases. There were four complex, noncontiguous rearrangements, with two also involving a partial triplication. In one of the few cases where RNA was analyzed, a seemingly contiguous duplication turned out to be a duplication/deletion case generating a transcript with an unexpected single-exon deletion and an initially undetected duplication. These findings indicate that for clinical diagnosis, duplications should be treated with special care, and without further analysis the reading frame rule should not be applied. As with deletions, duplications occur nonrandomly but with a dramatically different distribution. Duplication frequency is highest near the 5' end of the gene, with a duplication of exon 2 being the single most common duplication identified. Analysis of the extent of 11 exon 2 duplications revealed two intron 2 recombination hotspots. Sequencing four of the breakpoints showed that they did not arise from unequal sister chromatid exchange, but more likely from synthesis-dependent nonhomologous end joining. There appear to be fundamental differences therefore in the origin of deletions and duplications in the DMD gene.

Cohort Studies↗

DNA sequence duplication in Rhodobacter sphaeroides 2.4.1: evidence of an ancient partnership between chromosomes I and II.

The complex genome of Rhodobacter sphaeroides 2.4.1, composed of chromosomes I (CI) and II (CII), has been sequenced and assembled. We present data demonstrating that the R. sphaeroides genome possesses an extensive amount of exact DNA sequence duplication, 111 kb or approximately 2.7% of the total chromosomal DNA. The chromosomal DNA sequence duplications were aligned to each other by using MUMmer. Frequency and size distribution analyses of the exact DNA duplications revealed that the interchromosomal duplications occurred prior to the intrachromosomal duplications. Most of the DNA sequence duplications in the R. sphaeroides genome occurred early in species history, whereas more recent sequence duplications are rarely found. To uncover the history of gene duplications in the R. sphaeroides genome, 44 gene duplications were sampled and then analyzed for DNA sequence similarity against orthologous DNA sequences. Phylogenetic analysis revealed that approximately 80% of the total gene duplications examined displayed type A phylogenetic relationships; i.e., one copy of each member of a duplicate pair was more similar to its orthologue, found in a species closely related to R. sphaeroides, than to its duplicate, counterpart allele. The data reported here demonstrate that a massive level of gene duplications occurred prior to the origin of the R. sphaeroides 2.4.1 lineage. These findings lead to the conclusion that there is an ancient partnership between CI and CII of R. sphaeroides 2.4.1.

Bacterial Proteins↗

Stability of large segmental duplications in the yeast genome.

The high level of gene redundancy that characterizes eukaryotic genomes results in part from segmental duplications. Spontaneous duplications of large chromosomal segments have been experimentally demonstrated in yeast. However, the dynamics of inheritance of such structures and their eventual fixation in populations remain largely unsolved. We analyzed the stability of a vast panel of large segmental duplications in Saccharomyces cerevisiae (from 41 kb for the smallest to 268 kb for the largest). We monitored the stability of three different types of interchromosomal duplications as well as that of three intrachromosomal direct tandem duplications. In the absence of any selective advantage associated with the presence of the duplication, we show that a duplicated segment internally translocated within a natural chromosome is stably inherited both mitotically and meiotically. By contrast, large duplications carried by a supernumerary chromosome are highly unstable. Duplications translocated into subtelomeric regions are lost at variable rates depending on the location of the insertion sites. Direct tandem duplications are lost by unequal crossing over, both mitotically and meiotically, at a frequency proportional to their sizes. These results show that most of the duplicated structures present an intrinsic level of instability. However, translocation within another chromosome significantly stabilizes a duplicated segment, increasing its chance to get fixed in a population even in the absence of any immediate selective advantage conferred by the duplicated genes.

Fungal Proteins↗

Tandem duplication of the FLT3 gene is found in acute lymphoblastic leukaemia as well as acute myeloid leukaemia but not in myelodysplastic syndrome or juvenile chronic myelogenous leukaemia in children.

We examined mRNA expression and internal tandem duplication of the Fms-like tyrosine kinase 3 (FLT3) gene in haematological malignancies by reverse transcriptase-polymerase chain reaction (RT-PCR) and genomic PCR followed by sequencing. By RT-PCR, expression of FLT3 was detected in 45/74 (61%) leukaemia cell lines and the frequency of expression of FLT3 was significantly higher in undifferentiated type (B-precursor acute lymphoblastic leukaemia; ALL) than in differentiated type cell lines (B-ALL) (P = 0.0076). Using the genomic PCR method, 194 fresh samples including 87 acute myeloid leukaemias, 60 ALLs, 32 myelodysplastic syndromes (MDSs) and 15 juvenile chronic myelogenous leukaemias (JCMLs) were examined. Tandem duplication was found in 12 (13.8%) AMLs and two (3.3%) ALLs. Sequence analyses of the 14 samples with the duplication revealed that eight showed a simple tandem duplication and six a tandem duplication with insertion. Most of these tandem duplications occurred within exon 11, and two duplications occurred from exon 11 to intron 11 and exon 12. No tandem duplications of FLT3 gene were detected in MDS or JCML. The frequency of tandem duplication of FLT3 gene in childhood AML was lower than that in adult AML so far reported. All of the 12 AML patients with the duplication died within 47 months after diagnosis, whereas two ALL patients with the duplication have survived 44 and 72 months, respectively. These two ALL patients expressed both lymphoid and myeloid antigens and were considered to have biphenotypic leukaemia. These results suggest that tandem duplication is involved in ALL in addition to AML, but not in childhood MDS or JCML, and that childhood AML patients with the tandem duplication have a poor prognosis.

Acute Disease↗

Extensive gene duplication in the early evolution of animals before the parazoan-eumetazoan split demonstrated by G proteins and protein tyrosine kinases from sponge and hydra.

To know whether genes involved in cell-cell communication typical of multicellular animals dramatically increased in concert with the Cambrian explosion, the rapid evolutionary burst in the major groups of animals, and whether these genes exist in the sponge lacking cell cohesiveness and coordination typical of eumetazoans, we have carried out cloning of the G-protein alpha subunit (Galpha) and the protein tyrosine kinase (PTK) cDNAs from Ephydatia fluviatilis (freshwater sponge) and Hydra magnipapillata strain 105 (hydra). We obtained 13 Galpha and 20 PTK cDNAs. Generally animal gene families diverged first by gene duplication (subtype duplication) that gave rise to diverse subtypes with different primary functions, followed by further gene duplication in the same subtype (isoform duplication) that gave rise to isoform genes with virtually identical function. Phylogenetic trees of Galpha and PTK families including cDNAs from sponge and hydra revealed that most of the present-day subtypes had been established in the very early evolution of animals before the parazoan-eumetazoan split, the earliest branching among the extant animal phyla, by extensive subtype duplication: for PTK and Galpha families, 23 and 9 subtype duplications were observed in the early stage before the parazoan-eumetazoan split, respectively, and after that split, only 2 and 1 subtype duplications were found, respectively. After the separation from arthropods, vertebrates underwent frequent isoform duplications before the fish-tetrapod split. Furthermore, rapid amino acid changes appear to have occurred in concert with the extensive subtype duplication and isoform duplication. Thus the pattern of gene diversification during animal evolution might be characterized by bursts of gene duplication interrupted by considerably long periods of silence, instead of proceeding gradually, and there might be no direct link between the Cambrian explosion and the extensive gene duplication that generated diverse functions (subtypes) of these families.

Amino Acid Substitution↗

The early stages of duplicate gene evolution.

Gene duplications are one of the primary driving forces in the evolution of genomes and genetic systems. Gene duplicates account for 8-20% of the genes in eukaryotic genomes, and the rates of gene duplication are estimated at between 0.2% and 2% per gene per million years. Duplicate genes are believed to be a major mechanism for the establishment of new gene functions and the generation of evolutionary novelty, yet very little is known about the early stages of the evolution of duplicated gene pairs. It is unclear, for example, to what extent selection, rather than neutral genetic drift, drives the fixation and early evolution of duplicate loci. Analysis of recently duplicated genes in the Arabidopsis thaliana genome reveals significantly reduced species-wide levels of nucleotide polymorphisms in the progenitor and/or duplicate gene copies, suggesting that selective sweeps accompany the initial stages of the evolution of these duplicated gene pairs. Our results support recent theoretical work that indicates that fates of duplicate gene pairs may be determined in the initial phases of duplicate gene evolution and that positive selection plays a prominent role in the evolutionary dynamics of the very early histories of duplicate nuclear genes.

Arabidopsis↗

The structure and early evolution of recently arisen gene duplicates in the Caenorhabditis elegans genome.

The significance of gene duplication in provisioning raw materials for the evolution of genomic diversity is widely recognized, but the early evolutionary dynamics of duplicate genes remain obscure. To elucidate the structural characteristics of newly arisen gene duplicates at infancy and their subsequent evolutionary properties, we analyzed gene pairs with < or =10% divergence at synonymous sites within the genome of Caenorhabditis elegans. Structural heterogeneity between duplicate copies is present very early in their evolutionary history and is maintained over longer evolutionary timescales, suggesting that duplications across gene boundaries in conjunction with shuffling events have at least as much potential to contribute to long-term evolution as do fully redundant (complete) duplicates. The median duplication span of 1.4 kb falls short of the average gene length in C. elegans (2.5 kb), suggesting that partial gene duplications are frequent. Most gene duplicates reside close to the parent copy at inception, often as tandem inverted loci, and appear to disperse in the genome as they age, as a result of reduced survivorship of duplicates located in proximity to the ancestral copy. We propose that illegitimate recombination events leading to inverted duplications play a disproportionately large role in gene duplication within this genome in comparison with other mechanisms.

Animals↗

Segmental duplications: organization and impact within the current human genome project assembly.

Segmental duplications play fundamental roles in both genomic disease and gene evolution. To understand their organization within the human genome, we have developed the computational tools and methods necessary to detect identity between long stretches of genomic sequence despite the presence of high copy repeats and large insertion-deletions. Here we present our analysis of the most recent genome assembly (January 2001) in which we focus on the global organization of these segments and the role they play in the whole-genome assembly process. Initially, we considered only large recent duplication events that fell well-below levels of draft sequencing error (alignments 90%-98% similar and > or =1 kb in length). Duplications (90%-98%; > or =1 kb) comprise 3.6% of all human sequence. These duplications show clustering and up to 10-fold enrichment within pericentromeric and subtelomeric regions. In terms of assembly, duplicated sequences were found to be over-represented in unordered and unassigned contigs indicating that duplicated sequences are difficult to assign to their proper position. To assess coverage of these regions within the genome, we selected BACs containing interchromosomal duplications and characterized their duplication pattern by FISH. Only 47% (106/224) of chromosomes positive by FISH had a corresponding chromosomal position by comparison. We present data that indicate that this is attributable to misassembly, misassignment, and/or decreased sequencing coverage within duplicated regions. Surprisingly, if we consider putative duplications >98% identity, we identify 10.6% (286 Mb) of the current assembly as paralogous. The majority of these alignments, we believe, represent unmerged overlaps within unique regions. Taken together the above data indicate that segmental duplications represent a significant impediment to accurate human genome assembly, requiring the development of specialized techniques to finish these exceptional regions of the genome. The identification and characterization of these highly duplicated regions represents an important step in the complete sequencing of a human reference genome.

Base Sequence↗