Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Copy number variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Extensive normal copy number variation of a beta-defensin antimicrobial-gene cluster.

Using a combination of multiplex amplifiable probe hybridization and semiquantitative fluorescence in situ hybridization (SQ-FISH), we analyzed DNA copy number variation across chromosome band 8p23.1, a region that is frequently involved in chromosomal rearrangements. We show that a cluster of at least three antimicrobial beta-defensin genes (DEFB4, DEFB103, and DEFB104) at 8p23.1 are polymorphic in copy number, with a repeat unit >/=240 kb long. Individuals have 2-12 copies of this repeat per diploid genome. By segregation, microsatellite dosage, and SQ-FISH chromosomal signal intensity ratio analyses, we deduce that individual chromosomes can have one to eight copies of this repeat unit. Chromosomes with seven or eight copies of this repeat unit are identifiable by cytogenetic analysis as a previously described 8p23.1 euchromatic variant. Analysis of RNA from different individuals by semiquantitative reverse-transcriptase polymerase chain reaction shows a significant correlation between genomic copy number of DEFB4 and levels of its messenger RNA (mRNA) transcript. The peptides encoded by these genes are potent antimicrobial agents, especially effective against clinically important pathogens, such as Pseudomonas aeruginosa and Staphylococcus aureus, and DEFB4 has been shown to act as a cytokine linking the innate and adaptive immune responses. Therefore, a copy number polymorphism involving these genes, which is reflected in mRNA expression levels, is likely to have important consequences for immune system function.

Alleles↗

Segmental duplications and copy-number variation in the human genome.

The human genome contains numerous blocks of highly homologous duplicated sequence. This higher-order architecture provides a substrate for recombination and recurrent chromosomal rearrangement associated with genomic disease. However, an assessment of the role of segmental duplications in normal variation has not yet been made. On the basis of the duplication architecture of the human genome, we defined a set of 130 potential rearrangement hotspots and constructed a targeted bacterial artificial chromosome (BAC) microarray (with 2,194 BACs) to assess copy-number variation in these regions by array comparative genomic hybridization. Using our segmental duplication BAC microarray, we screened a panel of 47 normal individuals, who represented populations from four continents, and we identified 119 regions of copy-number polymorphism (CNP), 73 of which were previously unreported. We observed an equal frequency of duplications and deletions, as well as a 4-fold enrichment of CNPs within hotspot regions, compared with control BACs (P < .000001), which suggests that segmental duplications are a major catalyst of large-scale variation in the human genome. Importantly, segmental duplications themselves were also significantly enriched >4-fold within regions of CNP. Almost without exception, CNPs were not confined to a single population, suggesting that these either are recurrent events, having occurred independently in multiple founders, or were present in early human populations. Our study demonstrates that segmental duplications define hotspots of chromosomal rearrangement, likely acting as mediators of normal variation as well as genomic disease, and it suggests that the consideration of genomic architecture can significantly improve the ascertainment of large-scale rearrangements. Our specialized segmental duplication BAC microarray and associated database of structural polymorphisms will provide an important resource for the future characterization of human genomic disorders.

Chromosomes, Artificial, Bacterial↗

Mapping tumor-suppressor genes with multipoint statistics from copy-number-variation data.

Array-based comparative genomic hybridization (arrayCGH) is a microarray-based comparative genomic hybridization technique that has been used to compare tumor genomes with normal genomes, thus providing rapid genomic assays of tumor genomes in terms of copy-number variations of those chromosomal segments that have been gained or lost. When properly interpreted, these assays are likely to shed important light on genes and mechanisms involved in the initiation and progression of cancer. Specifically, chromosomal segments, deleted in one or both copies of the diploid genomes of a group of patients with cancer, point to locations of tumor-suppressor genes (TSGs) implicated in the cancer. In this study, we focused on automatic methods for reliable detection of such genes and their locations, and we devised an efficient statistical algorithm to map TSGs, using a novel multipoint statistical score function. The proposed algorithm estimates the location of TSGs by analyzing segmental deletions (hemi- or homozygous) in the genomes of patients with cancer and the spatial relation of the deleted segments to any specific genomic interval. The algorithm assigns, to an interval of consecutive probes, a multipoint score that parsimoniously captures the underlying biology. It also computes a P value for every putative TSG by using concepts from the theory of scan statistics. Furthermore, it can identify smaller sets of predictive probes that can be used as biomarkers for diagnosis and therapeutics. We validated our method using different simulated artificial data sets and one real data set, and we report encouraging results. We discuss how, with suitable modifications to the underlying statistical model, this algorithm can be applied generally to a wider class of problems (e.g., detection of oncogenes).

Gene Deletion↗

A hierarchical clustering method for estimating copy number variation.

Microarray technologies allow for simultaneous measurement of DNA copy number at thousands of positions in a genome. Gains and losses of DNA sequences reveal themselves through characteristic patterns of hybridization intensity. To identify change points along the chromosomes, we develop a marker clustering method which consists of 2 parts. First, a "circular clustering tree test statistic" attaches a statistic to each marker that measures the likelihood that it is a change point. Then construction of the marker statistics is followed by outlier detection approaches. The method provides a new way to build up a binary tree that can accurately capture change-point signals and is easy to perform. A simulation study shows good performance in change-point detection, and cancer cell line data are used to illustrate performance when regions of true copy number changes are known.

Algorithms↗

Ribosomal DNA copy number variation associates with hematological profiles and renal function in the UK Biobank.

The phenotypic impact of genetic variation of repetitive features in the human genome is currently understudied. One such feature is the multi-copy 47S ribosomal DNA (rDNA) that codes for rRNA components of the ribosome. Here, we present an analysis of rDNA copy number (CN) variation in the UK Biobank (UKB). From the first release of UKB whole-genome sequencing (WGS) data, a discovery analysis in White British individuals reveals that rDNA CN associates with altered counts of specific blood cell subtypes, such as neutrophils, and with the estimated glomerular filtration rate, a marker of kidney function. Similar trends are observed in other ancestries. A range of analyses argue against reverse causality or common confounder effects, and all core results replicate in the second UKB WGS release. Our work demonstrates that rDNA CN is a genetic influence on trait variance in humans.

Humans↗

Large scale copy number variation (CNV) at 14q12 is associated with the presence of genomic abnormalities in neoplasia.

BACKGROUND: Advances made in the area of microarray comparative genomic hybridization (aCGH) have enabled the interrogation of the entire genome at a previously unattainable resolution. This has lead to the discovery of a novel class of alternative entities called large-scale copy number variations (CNVs). These CNVs are often found in regions of closely linked sequence homology called duplicons that are thought to facilitate genomic rearrangements in some classes of neoplasia. Recently, it was proposed that duplicons located near the recurrent translocation break points on chromosomes 9 and 22 in chronic myeloid leukemia (CML) may facilitate this tumor-specific translocation. Furthermore, approximately 15-20% of CML patients also carry a microdeletion on the derivative 9 chromosome (der(9)) and these patients have a poor prognosis. It has been hypothesised that der(9) deletion patients have increased levels of chromosomal instability. RESULTS: In this study aCGH was performed and identified a CNV (RP11-125A5, hereafter called CNV14q12) that was present as a genomic gain or loss in 10% of control DNA samples derived from cytogenetically normal individuals. CNV14q12 was the same clone identified by Iafrate et al. as a CNV. Real-time polymerase chain reaction (Q-PCR) was used to determine the relative frequency of this CNV in DNA from a series of 16 CML patients (both with and without a der(9) deletion) together with DNA derived from 36 paediatric solid tumors in comparison to the incidence of CNV in control DNA. CNV14q12 was present in approximately 50% of both tumor and CML DNA, but was found in 72% of CML bearing a der(9) microdeletion. Chi square analysis found a statistically significant difference (p <or= 0.001) between the incidence of this CNV in cancer and normal DNA and a slightly increased incidence in CML with deletions in comparison to those CML without a detectable deletion. CONCLUSION: The increased incidence of CNV14q12 in tumor samples suggests that either acquired or inherited genomic variation of this new class of variation may be associated with onset or progression of neoplasia.

Child↗

Ribosomal DNA and Stellate gene copy number variation on the Y chromosome of Drosophila melanogaster.

Multigene families on the Y chromosome face an unusual array of evolutionary forces. Both ribosomal DNA and Stellate, the two families examined here, have multiple copies of similar sequences on the X and Y chromosomes. Although the rate of sequence divergence on the Y chromosome depends on rates of mutation, gene conversion and exchange with the X chromosome, as well as purifying selection, the regulation of gene copy number may also depend on other pleiotropic functions, such as maintenance of chromosome pairing. Gene copy numbers were estimated for a series of 34 Y chromosome replacement lines using densitometric measurements of slot blots of genomic DNA from adult Drosophila melanogaster. Scans of autoradiographs of the same blots probed with the cloned alcohol dehydrogenase gene, a single copy gene, served as internal standards. Copy numbers span a 6-fold range for ribosomal DNA and a 3-fold range for Stellate DNA. Despite this magnitude of variation, there was no association between copy number and segregation variation of the sex chromosomes.

Alcohol Dehydrogenase↗

Copy number variation in regions flanked (or unflanked) by duplicons among patients with developmental delay and/or congenital malformations; detection of reciprocal and partial Williams-Beuren duplications.

Duplicons, that is, DNA sequences with minimum length 10 kb and a high sequence similarity, are known to cause unequal homologous recombination, leading to deletions and the reciprocal duplications. In this study, we designed a Multiplex Amplifiable Probe Hybridisation (MAPH) assay containing 63 exon-specific single-copy sequences from within a selection of the 169 regions flanked by duplicons that were identified, at a first pass, in 2001. Subsequently, we determined the frequency of chromosomal rearrangements among patients with developmental delay (DD) and/or congenital malformations (CM). In addition, we tried to identify new regions involved in DD/CM using the same assay. In 105 patients, six imbalances (5.8%) were detected and verified. Three of these were located in microdeletion-related regions, two alterations were polymorphic duplications and the effect of the last alteration is currently unknown. The same study population was tested for rearrangements in regions with no known duplicons nearby, using a set of probes derived from 58 function-selected genes. The latter screening revealed two alterations. As expected, the alteration frequency per unit of DNA is much higher in regions flanked by duplicons (fraction of the genome tested: 5.2%) compared to regions without known duplicons nearby (fraction of the genome tested: 24.5-90.2%). We were able to detect three novel rearrangements, including the previously undescribed reciprocal duplication of the Williams Beuren critical region, a subduplicon alteration within this region and a duplication on chromosome band 16p13.11. Our results support the hypothesis that regions flanked by duplicons are enriched for copy number variations.

Child↗

Genome-wide copy number variation association study in anorexia nervosa.

This study represents the first large-scale investigation of rare (<1% population frequency) copy number variants (CNVs) in anorexia nervosa (AN). Large, rare CNVs are reported to be causally associated with anthropometric traits, neurodevelopmental disorders, and schizophrenia, yet their role in the genetic basis of AN is unclear. Using genome-wide association study (GWAS) array data from the Anorexia Nervosa Genetics Initiative (ANGI), which included 7414 AN case and 5044 controls, we investigated the association of 67 well-established syndromic CNVs and 178 pleiotropic disease-risk dosage-sensitive CNVs with AN. To identify novel CNV regions (CNVRs) that increase the risk of AN, we conducted genome-wide association studies with a focus on rare CNV-breakpoints (CNV-GWAS). We found no net enrichment of rare CNVs, either deletions or duplications, in AN, and none of the well-established syndromic or pleiotropic CNVs had a significant association with AN status. However, the CNV-GWAS found 21 nominally associated CNVRs that contribute to AN risk, covering protein-coding genes implicated in synaptic function, metabolic/mitochondrial factors, and lipid characteristics, like the CD36 (7q21.11) gene, which transports long-chain fatty acids into cells. CNVRs intersecting genes previously related to neurodevelopmental traits include deletions of NRXN1 intron 5 (2p16.3), IMMP2L (7q31.1), and PTPRD (9p23). Overall, given that our study is well powered to detect the CNV burden level reported for schizophrenia, we can conclude that rare CNVs have a limited role in the etiology of AN, as reported for bipolar disorder. Our nominal associations for the 21 discovered CNVRs are consistent with AN being a metabo-psychiatric trait, as demonstrated by the common genetic architecture of AN, and we provide association results to allow for replication in future research.

Humans↗

A high-resolution map of segmental DNA copy number variation in the mouse genome.

Submicroscopic (less than 2 Mb) segmental DNA copy number changes are a recently recognized source of genetic variability between individuals. The biological consequences of copy number variants (CNVs) are largely undefined. In some cases, CNVs that cause gene dosage effects have been implicated in phenotypic variation. CNVs have been detected in diverse species, including mice and humans. Published studies in mice have been limited by resolution and strain selection. We chose to study 21 well-characterized inbred mouse strains that are the focus of an international effort to measure, catalog, and disseminate phenotype data. We performed comparative genomic hybridization using long oligomer arrays to characterize CNVs in these strains. This technique increased the resolution of CNV detection by more than an order of magnitude over previous methodologies. The CNVs range in size from 21 to 2,002 kb. Clustering strains by CNV profile recapitulates aspects of the known ancestry of these strains. Most of the CNVs (77.5%) contain annotated genes, and many (47.5%) colocalize with previously mapped segmental duplications in the mouse genome. We demonstrate that this technique can identify copy number differences associated with known polymorphic traits. The phenotype of previously uncharacterized strains can be predicted based on their copy number at these loci. Annotation of CNVs in the mouse genome combined with sequence-based analysis provides an important resource that will help define the genetic basis of complex traits.

Animals↗

Mouse genomic representational oligonucleotide microarray analysis: detection of copy number variations in normal and tumor specimens.

Genomic amplifications and deletions, the consequence of somatic variation, are a hallmark of human cancer. Such variation has also been observed between "normal" individuals, as well as in individuals with congenital disorders. Thus, copy number measurement is likely to be an important tool for the analysis of genetic variation, genetic disease, and cancer. We developed representational oligonucleotide microarray analysis, a high-resolution comparative genomic hybridization methodology, with this aim in mind, and reported its use in the study of humans. Here we report the development of a representational oligonucleotide microarray analysis microarray for the genomic analysis of the mouse, an important model system for many genetic diseases and cancer. This microarray was designed based on the sequence assembly MM3, and contains approximately 84,000 probes randomly distributed throughout the mouse genome. We demonstrate the use of this array to identify copy number changes in mouse cancers, as well to determine copy number variation between inbred strains of mice. Because restriction endonuclease digestion of genomic DNA is an integral component of our method, differences due to polymorphisms at the restriction enzyme cleavage sites are also observed between strains, and these can be useful to follow the inheritance of loci between crosses of different strains.

Animals↗

High resolution analysis of DNA copy number variation using comparative genomic hybridization to microarrays.

Gene dosage variations occur in many diseases. In cancer, deletions and copy number increases contribute to alterations in the expression of tumour-suppressor genes and oncogenes, respectively. Developmental abnormalities, such as Down, Prader Willi, Angelman and Cri du Chat syndromes, result from gain or loss of one copy of a chromosome or chromosomal region. Thus, detection and mapping of copy number abnormalities provide an approach for associating aberrations with disease phenotype and for localizing critical genes. Comparative genomic hybridization (CGH) was developed for genome-wide analysis of DNA sequence copy number in a single experiment. In CGH, differentially labelled total genomic DNA from a 'test' and a 'reference' cell population are cohybridized to normal metaphase chromosomes, using blocking DNA to suppress signals from repetitive sequences. The resulting ratio of the fluorescence intensities at a location on the 'cytogenetic map', provided by the chromosomes, is approximately proportional to the ratio of the copy numbers of the corresponding DNA sequences in the test and reference genomes. CGH has been broadly applied to human and mouse malignancies. The use of metaphase chromosomes, however, limits detection of events involving small regions (of less than 20 Mb) of the genome, resolution of closely spaced aberrations and linking ratio changes to genomic/genetic markers. Therefore, more laborious locus-by-locus techniques have been required for higher resolution studies. Hybridization to an array of mapped sequences instead of metaphase chromosomes could overcome the limitations of conventional CGH (ref. 6) if adequate performance could be achieved. Copy number would be related to the test/reference fluorescence ratio on the array targets, and genomic resolution could be determined by the map distance between the targets, or by the length of the cloned DNA segments. We describe here our implementation of array CGH. We demonstrate its ability to measure copy number with high precision in the human genome, and to analyse clinical specimens by obtaining new information on chromosome 20 aberrations in breast cancer.

Animals↗

ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencing.

MOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome&#x2011;wide coverage of only 0.1-5&#xd7;, sWGS data display a pronounced zero&#x2011;inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several&#x2011;fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy&#x2011;number gains (false positives), and true deletions often become indistinguishable from pervasive zero&#x2011;coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv.

DNA Copy Number Variations↗

Detecting single DNA copy number variations in complex genomes using one nanogram of starting DNA and BAC-array CGH.

Comparative genomic hybridization to bacterial artificial chromosome (BAC)-arrays (array-CGH) is a highly efficient technique, allowing the simultaneous measurement of genomic DNA copy number at hundreds or thousands of loci, and the reliable detection of local one-copy-level variations. We report a genome-wide amplification method allowing the same measurement sensitivity, using 1 ng of starting genomic DNA, instead of the classical 1 microg usually necessary. Using a discrete series of DNA fragments, we defined the parameters adapted to the most faithful ligation-mediated PCR amplification and the limits of the technique. The optimized protocol allows a 3000-fold DNA amplification, retaining the quantitative characteristics of the initial genome. Validation of the amplification procedure, using DNA from 10 tumour cell lines hybridized to BAC-arrays of 1500 spots, showed almost perfectly superimposed ratios for the non-amplified and amplified DNAs. Correlation coefficients of 0.96 and 0.99 were observed for regions of low-copy-level variations and all regions, respectively (including in vivo amplified oncogenes). Finally, labelling DNA using two nucleotides bearing the same fluorophore led to a significant increase in reproducibility and to the correct detection of one-copy gain or loss in >90% of the analysed data, even for pseudotriploid tumour genomes.

Cell Line, Tumor↗

Copy number variation in the genome; the human DMD gene as an example.

Recent developments have yielded new technologies that have greatly simplified the detection of deletions and duplications, i.e., copy number variants (CNVs). These technologies can be used to screen for CNVs in and around specific genomic regions, as well as genome-wide. Several genome-wide studies have demonstrated that CNV in the human genome is widespread and may include millions of nucleotides. One of the questions that emerge is which sequences, structures and/or processes are involved in their generation. Using as an example the human DMD gene, mutations in which cause Duchenne and Becker muscular dystrophy, we review the current data, determine the deletion and duplication profile across the gene and summarize the information that has been collected regarding their origin. In addition we discuss the methods most frequently used for their detection, in particular MAPH and MLPA.

Alleles↗

Spontaneous rDNA copy number variation modulates Sir2 levels and epigenetic gene silencing.

We show that in budding yeast large rDNA deletions arise frequently and cause an increase in telomeric and mating-type gene silencing proportional to repeat loss. Paradoxically, this increase in silencing is correlated with a highly specific down-regulation of SIR2, which encodes a deacetylase enzyme required for silencing. These apparently conflicting observations suggest that a large nucleolar pool of Sir2 is released upon rDNA loss and made available for telomeric and HM silencing, as well as down-regulation of SIR2 itself. Indeed, we present evidence for a reduction in the fraction of Sir2 colocalizing with the nucleolar marker Nop1, and for SIR2 autoregulation. Despite a decrease in the fraction of nucleolar Sir2, and in overall Sir2 protein levels, short rDNA strains display normal rDNA silencing and a lifespan indistinguishable from wild type. These observations reveal an unexpectedly large clonal variation in rDNA cluster size and point to the existence of a novel regulatory circuit, sensitive to rDNA copy number, that balances nucleolar and nonnucleolar pools of Sir2 protein.

DNA, Fungal↗

Representational oligonucleotide microarray analysis: a high-resolution method to detect genome copy number variation.

We have developed a methodology we call ROMA (representational oligonucleotide microarray analysis), for the detection of the genomic aberrations in cancer and normal humans. By arraying oligonucleotide probes designed from the human genome sequence, and hybridizing with "representations" from cancer and normal cells, we detect regions of the genome with altered "copy number." We achieve an average resolution of 30 kb throughout the genome, and resolutions as high as a probe every 15 kb are practical. We illustrate the characteristics of probes on the array and accuracy of measurements obtained using ROMA. Using this methodology, we identify variation between cancer and normal genomes, as well as between normal human genomes. In cancer genomes, we readily detect amplifications and large and small homozygous and hemizygous deletions. Between normal human genomes, we frequently detect large (100 kb to 1 Mb) deletions or duplications. Many of these changes encompass known genes. ROMA will assist in the discovery of genes and markers important in cancer, and the discovery of loci that may be important in inherited predispositions to disease.

Aneuploidy↗

Tandem repeat copy-number variation in protein-coding regions of human genes.

BACKGROUND: Tandem repeat variation in protein-coding regions will alter protein length and may introduce frameshifts. Tandem repeat variants are associated with variation in pathogenicity in bacteria and with human disease. We characterized tandem repeat polymorphism in human proteins, using the UniGene database, and tested whether these were associated with host defense roles. RESULTS: Protein-coding tandem repeat copy-number polymorphisms were detected in 249 tandem repeats found in 218 UniGene clusters; observed length differences ranged from 2 to 144 nucleotides, with unit copy lengths ranging from 2 to 57. This corresponded to 1.59% (218/13,749) of proteins investigated carrying detectable polymorphisms in the copy-number of protein-coding tandem repeats. We found no evidence that tandem repeat copy-number polymorphism was significantly elevated in defense-response proteins (p = 0.882). An association with the Gene Ontology term 'protein-binding' remained significant after covariate adjustment and correction for multiple testing. Combining this analysis with previous experimental evaluations of tandem repeat polymorphism, we estimate the approximate mean frequency of tandem repeat polymorphisms in human proteins to be 6%. Because 13.9% of the polymorphisms were not a multiple of three nucleotides, up to 1% of proteins may contain frameshifting tandem repeat polymorphisms. CONCLUSION: Around 1 in 20 human proteins are likely to contain tandem repeat copy-number polymorphisms within coding regions. Such polymorphisms are not more frequent among defense-response proteins; their prevalence among protein-binding proteins may reflect lower selective constraints on their structural modification. The impact of frameshifting and longer copy-number variants on protein function and disease merits further investigation.

Frameshift Mutation↗