Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Structural genome variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Protocol for haplotype-resolved structural variant detection via long-read sequencing using cuteHap.

Long-read sequencing technologies have revolutionized human genome exploration at an unparalleled resolution, particularly facilitating the analysis of structural variation (SV) at haplotype resolution. Here, we present a protocol for using cuteHap, a robust framework for haplotype-aware SV detection through phased alignment reads generated by diverse long-read sequencing platforms. We describe procedures for single-nucleotide variant (SNV) calling, read phasing, SV calling, and genotyping. We also establish a benchmarking pipeline to evaluate the detected SV callsets. For complete details on the use and execution of this protocol, please refer to Cao et al.1.

Bioinformatics↗

Structural variation in the Waxy gene and differentiation in foxtail millet [Setaria italica (L.) P. Beauv.]: implications for multiple origins of the waxy phenotype.

The origin and evolution of the waxy type of foxtail millet [Setaria italica (L.) P. Beauv] were studied by analyzing structural variation in the Waxy gene. Initially, the Waxy gene was amplified by RT-PCR, RACE and genomic PCR from a non-waxy strain to determine the structure of the wild-type gene. Secondly, we screened by PCR for polymorphisms at the Waxy locus in 79 strains with various waxy phenotypes. We then carried out genomic Southern analysis on 67 strains and identified seven RFLP classes which were designated as types I-VII. RFLP type was correlated with phenotype, such that types I and II corresponded to non-waxy, types III and VI to low-amylose, and types IV, V and VII to waxy phenotypes. The differences between RFLP types could be attributed to insertions in the Waxy gene. Types II and VI were caused by the insertion of a Tourist element into intron 1 and a SINE-like sequence into intron 12, respectively. Types III, IV, V and VII were characterized by the insertion of large sequences into the Waxy gene that may alter the expression of the gene. Thus, multiple, independent insertions in the Waxy gene appear to have caused the loss-of-function waxy phenotypes. Furthermore, the geographical distributions of the three RFLP types associated with the waxy phenotype (types IV, V and VII) were distinct, with type IV being found mainly in Taiwan and Japan, type V in Korea, and type VII in Myanmar. These results indicate a polyphyletic origin for the waxy phenotype in landraces of foxtail millet.

Base Sequence↗

A family of retrotransposons and associated genomic variation in wheat.

A family of related retroelements was characterized in the genomes of some Graminease species. The structure of these retroelements indicates that they are retrotransposons containing reading frames with sequence similarity to the polyproteins of copia and Ty. This family of retroelements (termed WIS-2) occurs in the genomes of barley, wheat, rye, oats, and Aegilops species. Ongoing genomic variation both within individual plants of a wheat variety and within and between varieties of wheat is associated with some members of the WIS-2 family.

Amino Acid Sequence↗

Subpathways of nucleotide excision repair and their regulation.

Nucleotide excision repair provides an important cellular defense against a large variety of structurally unrelated DNA alterations. Most of these alterations, if unrepaired, may contribute to mutagenesis, oncogenesis, and developmental abnormalities, as well as cellular lethality. There are two subpathways of nucleotide excision repair; global genomic repair (GGR) and transcription coupled repair (TCR), that is selective for the transcribed DNA strand in expressed genes. Some of the proteins involved in the recognition of DNA damage (including RNA polymerase) are also responsive to natural variations in the secondary structural features of DNA. Gratuitous repair events in undamaged DNA might then contribute to genomic instability. However, damage recognition enzymes for GGR are normally maintained at very low levels unless the cells are genomically stressed. GGR is controlled through the SOS stress response in E. coli and through the activated p53 tumor suppressor in human cells. These inducible responses in human cells are important, as they have been shown to operate upon chemical carcinogen DNA damage at levels to which humans are environmentally exposed. Interestingly, most rodent tissues are deficient in the p53-dependent GGR pathway. Since rodents are used as surrogates for environmental cancer risk assessment, it is essential that we understand how they differ from humans with respect to DNA repair and oncogenic responses to environmental genotoxins. In the case of terminally differentiated mammalian cells, a new paradigm has appeared in which GGR is attenuated but both strands of expressed genes are repaired efficiently.

Animals↗

Distribution and sequence analysis of a family of type ill-dependent effectors correlate with the phylogeny of Ralstonia solanacearum strains.

In Ralstonia solanacearum, we previously have reported on the characterization of popP1 and popP2 genes. These genes encode type III-dependent pathogenicity effectors related to the large family of AvrRxv/YopJ cysteine proteases that are shared among pathogens of plants and animals. In this study, we identify a third gene, named popP3, that is inactivated in the genome sequence of strain GMI1000 by insertion of a copy of the insertion sequence ISRso13. The three popP genes are localized on two large chromosomal pathogenicity islands, with popP1 and popP2 being present on the same island. Phylogenic analysis demonstrated that the PopP2 and PopP3 proteins are clearly distinct from other effectors of this family previously characterized in plant and animal pathogens. Analysis of the distribution and allelic variations of the three genes in 30 strains representative of the biodiversity of R. solanacearum established that popP genes are distributed widely among strains from two of the three phyla previously defined on the basis of the structure of the core genome. Sequencing of the popP genes from the different strains revealed limited allelic variations at the three loci but did not show evidence of recombination between the popP genes. Limited allelic variation together with occurrence of insertion sequences within or in the close vicinity of popP genes and the presence of gene duplications in these pathogenicity islands suggest that genomic rearrangements might be a major evolutionary driving force controlling evolution of the genes encoded in these regions. The implications of these observations in terms of bacterial evolution, gene acquisition, and horizontal gene transfers are discussed.

Alleles↗

Polymorphic variations in the ori sequences from the mitochondrial genomes of different wild-type yeast strains.

We determined the restriction maps and primary structures of two as yet poorly characterized regions of the mitochondrial genomes of different wild-type strains of Saccharomyces cerevisiae. These regions respectively comprised the ori1 sequence and the newly identified ori8 sequence. Ori1 and ori8, together with their flanking sequences, exhibit a large polymorphism, resulting from specific variations due to insertions or deletions of optional GC clusters at different locations. The mechanisms underlying such sequence rearrangements are discussed.

Base Sequence↗

High density linkage disequilibrium mapping using models of haplotype block variation.

MOTIVATION: The presence of millions of single nucleotide polymorphisms (SNPs) in the human genome has spurred interest in genetic mapping methods based on linkage disequilibrium. The recently discovered haplotype block structure of human variation promises to improve the effectiveness of these methods. A key difficulty for mapping techniques is the cost involved in separately identifying the haplotypes on each of an individual's chromosomes. RESULTS: We present a new approach for performing linkage disequilibrium mapping using high density haplotype or genotype data. Our method is based on a statistical model of haplotype block variation, which takes account of recombination hotspots, bottlenecks, genetic drift and mutation. We test our technique on two empirically determined high density datasets, attempting to recover the location of an SNP which was hidden and converted into phenotype information. We compare the results against a mapping method based on individual SNPs as well as a competing haplotype-based approach. We show that our strategy significantly outperforms these other approaches when used as a guide for resequencing and that it can also deal with both unphased genotype data and low penetrance diseases. AVAILABILITY: HaploBlock executables for Linux, Mac OS X and Sun OS, as well as user documentation, are available online at http://bioinfo.cs.technion.ac.il/haploblock/

Artificial Intelligence↗

Genomic mechanisms and measurement of structural and numerical instability in cancer cells.

The progression to cancer is often associated with instability and the acquisition of genomic heterogeneity, generating both clonal and non-clonal populations. Chromosomal instability (CIN) describes the excessive rate of numerical and structural genomic change in tumors. Mitotic segregation errors strongly influences copy number, while structural aberrations can occur at unstable genomic regions, or through aberrant DNA repair or methylation. Combined molecular cytogenetic analyses can evaluate cell-to-cell variation, and define the complexity of numerical and structural alterations. Because structural change may occur independently of numerical alteration, we propose the term structural chromosomal instability [(S)-CIN] to distinguish numerical from structural CIN.

Alu Elements↗

A high-density genetic recombination map of sequence-tagged sites for sorghum, as a framework for comparative structural and evolutionary genomics of tropical grains and grasses.

We report a genetic recombination map for Sorghum of 2512 loci spaced at average 0.4 cM ( approximately 300 kb) intervals based on 2050 RFLP probes, including 865 heterologous probes that foster comparative genomics of Saccharum (sugarcane), Zea (maize), Oryza (rice), Pennisetum (millet, buffelgrass), the Triticeae (wheat, barley, oat, rye), and Arabidopsis. Mapped loci identify 61.5% of the recombination events in this progeny set and reveal strong positive crossover interference acting across intervals of </=50 cM. Significant variations in DNA marker density are related to possible centromeric regions and to probable chromosome structural rearrangements between Sorghum bicolor and S. propinquum, but not to variation in levels of intraspecific allelic richness. While cDNA and genomic clones are similarly distributed across the genome, SSR-containing clones show different abundance patterns. Rapidly evolving hypomethylated DNA may contribute to intraspecific genomic differentiation. Nonrandom distribution patterns of multiple loci detected by 357 probes suggest ancient chromosomal duplication followed by extensive rearrangement and gene loss. Exemplifying the value of these data for comparative genomics, we support and extend prior findings regarding maize-sorghum synteny-in particular, 45% of comparative loci fall outside the inferred colinear/syntenic regions, suggesting that many small rearrangements have occurred since maize-sorghum divergence. These genetically anchored sequence-tagged sites will foster many structural, functional and evolutionary genomic studies in major food, feed, and biomass crops.

Biological Evolution↗

Comparative genomics reveals population structure and functional differentiation in Limosilactobacillus fermentum.

Limosilactobacillus fermentum is a widely distributed lactic acid bacterium frequently detected in fermented foods and host-associated microbiota, yet its global genomic diversity and functional variability remain insufficiently characterized. Here, we performed a large-scale comparative genomic analysis of 336 high-quality L. fermentum genomes curated from public databases. Species identity was validated using average nucleotide identity (ANI), and population structure was examined using pairwise ANI comparisons together with Mash-based phylogenetic reconstruction. Clustering at &#x2265;&#x2009;99% ANI resolved the dataset into 15 genomic clusters, with four dominant lineages comprising the majority of genomes. Pangenome reconstruction identified 5,853 gene clusters, including 1,325 core genes (22.6%) and a large accessory component dominated by low-frequency genes. Heap's law modeling (&#x3bb;&#x2009;=&#x2009;0.19) indicated a weakly open pangenome, suggesting ongoing gene acquisition as additional genomes are sampled. Functional annotation revealed that core genes were primarily associated with essential cellular processes, whereas accessory genes were enriched in carbohydrate metabolism, membrane-associated functions, and defense-related systems. Variation in carbohydrate-active enzymes (CAZymes), transport systems, and stress-response genes was observed across lineages, indicating strain-level functional diversity. Although genomes from human and food sources were broadly distributed across phylogenetic lineages, multivariate analysis showed that gene-content variation was more strongly associated with genomic lineage than with isolation source. These results provide a population genomic framework for understanding genomic diversity and functional potential in L. fermentum.

Phylogeny↗

Variations of endogenous chicken proviruses: characterization of new loci of endogenous proviruses in the genome of Italian partridge chickens.

The composition and structure of endogenous proviruses present in the genome of Italian Partridge chickens were studied by the method of blot hybridization using RAV-2 [32P]DNA or LTR of RSV as hybridization probes. The genomes of 5 out of 39 chickens analyzed did not contain endogenous proviruses related to RAV-2. Different sets of five so far undescribed endogenous proviruses, differing in the structure and location, were detected in the DNA of other IP chickens. None of them is identical in its structure to the DNA of the endogenous chicken virus RAV-0, all five loci of endogenous proviruses of IP chickens were defective. The origin, the patterns of genetic variation and the function of endogenous proviruses are discussed.

Animals↗

A variation in the structure of the protein-coding region of the human p53 gene.

An extensive analysis of genomic DNA preparations from a number of normal and malignant tissues revealed BglII site polymorphism of the human p53 gene. Approximately 10% of p53 gene alleles were found to contain an additional BglII site localized in a region of intron I. This allelic form of p53 gene was also responsible for p53 protein having altered electrophoretic mobility. Molecular cloning and sequencing of both the alleles of p53 gene revealed a base-pair change in codon 72 causing arginine----proline substitution in the allele with the additional BglII site. Both variants of the p53 gene may occur in homozygous state and are therefore functional.

Amino Acid Sequence↗

Genetic relatedness of hepatitis B viral strains of diverse geographical origin and natural variations in the primary structure of the surface antigen.

A 681 nucleotide fragment of the hepatitis B virus (HBV) genome was sequenced that corresponded to the complete gene for hepatitis B surface antigen (HBsAg) in 80 HBsAg- and hepatitis B e antigen (HBeAg)-positive sera of diverse geographical origins. These and 42 previously published HBV sequences within the S gene were used for the construction of a dendrogram. In this comparison, each of the 122 HBsAg genes was found to be related to one or other of the six previously identified genomic groups of HBV, A to F. The HBV strains within each genomic group showed a characteristic geographical distribution. Group A genomes were represented by 23 strains mainly originating in northern Europe and sub-Saharan Africa. The group B and C genomes, represented by 17 and 28 strains respectively, were confined to populations with origins in eastern Asia and the Far East. The group D genomes, represented by 38 strains, were found worldwide, but were the predominant strains in the Mediterranean area, the Near and Middle East, and in south Asia. Group E genomes, represented by nine strains, were indigenous to western sub-Saharan Africa as far south as Angola. There were indications that the F group, made up of six strains, represented the genomic group of HBV among populations with origins in the New World. Thus, HBV has diverged into genomic groups according to the distribution of mankind in the different continents. As well as giving information on the genetic relationship of HBV strains of different geographical origin, this study also provides information on the primary structure of HBsAg in different regions of the world. Such data might prove valuable in explaining the reported failures to obtain protection with current HBV vaccines.

Amino Acid Sequence↗

PangyPlot: multi-scale interactive visualization of pangenome variation graphs.

SUMMARY: Pangenome variation graphs integrate multiple samples into a unified representation, mitigating the reference bias inherent to linear genomes. However, these graphs can be large and structurally complex. Existing visualization tools are each confined to a fixed scale of resolution, requiring researchers to switch between multiple tools to examine variation at different levels of detail. PangyPlot is an interactive pangenome browser designed for multi-scale exploration of reference variation graphs from full chromosome to nucleotide-level sequence segments. PangyPlot anchors navigation to linear reference coordinates, organizes variation into hierarchical bubble structures, and uses a force-directed layout engine for automatic node arrangement. AVAILABILITY AND IMPLEMENTATION: An instance preloaded with data is available at https://pangyplot.research.sickkids.ca. Source code and documentation are openly available at https://github.com/strug-hub/pangyplot under the MIT License.

Software↗

Giant G+C% mosaic structures of the human genome found by arrangement of GenBank human DNA sequences according to genetic positions.

To determine the overall variation in the G+C% distribution over long ranges of the human genome, DNA sequences of human genes, which were closely linked genetically or physically, were surveyed from the GenBank Data Bank. A total of 72 sequences longer than 2 kb, which were mutually linked within 500 kb, were identified. The sequences belonged to 17 linkage groups and were ordered in each group according to their genetic positions. Analyses of the G+C% distribution along the ordered sequences showed that sequences within each group almost always had similar G+C% levels, but those belonging to different groups often had different levels. Similar analyses of more distantly linked sequences (e.g., greater than 10 Mb) showed mosaic structures of G+C% distribution. These findings are consistent with predictions made from the "isochore" structures found by CsCl equilibrium centrifugation, in that the structures having homogeneous base compositions stretched over at least several hundred kilobases. A possible boundary of the giant G+C% mosaic structures was identified between X-linked G6PD and F8C.

Base Composition↗

[Does genomics determine efficacy of analgesics?].

Recent advances in knowledge about gene structure derived from the human genome project has also revealed data on genomic variation and their possible impact on complex and acute diseases as well as pharmacotherapy. The hypothesis of a genetic predisposition for complex diseases such as pain syndromes, side effects, and adverse outcomes challenging the clinician is ready to be tested by advanced genetic-epidemiologic study designs employing the latest genotyping technology. In pain therapy, the genetic background of the efficacy of analgesics, especially of opioids, is of particular interest. Genetic differences in drug kinetics and dynamics, e.g., differences in metabolism or genetic variations of the drug target (e.g., receptors) will be of importance in the future. Pharmacogenetics can individualize pharmacotherapy and improve care by predicting the optimal dose and avoiding side effects and toxicity in individual patients.

Analgesia↗

Transient expression analysis of allelic variants of a VNTR in the dopamine transporter gene (DAT1).

BACKGROUND: The 10-repeat allele of a variable number tandem repeat (VNTR) polymorphism in the 3'-untranslated region of the dopamine transporter gene (DAT1) has been associated with a range of psychiatric phenotypes, most notably attention-deficit hyperactivity disorder. The mechanism for this association is not yet understood, although several lines of evidence implicate variation in gene expression. In this study we have characterised the genomic structure of the 9- and 10-repeat VNTR alleles, and directly examined the role of the polymorphism in mediating gene expression by measuring comparative in vitro cellular expression using a reporter-gene assay system. RESULTS: Differences in the sequence of the 9- and 10- repeat alleles were confirmed but no polymorphic differences were observed between individuals. There was no difference in expression of reporter gene constructs containing the two alleles. CONCLUSIONS: Our data suggests that this VNTR polymorphism may not have a direct effect on DAT1 expression and that the associations observed with psychiatric phenotypes may be mediated via linkage disequilibrium with other functional polymorphisms.

Alleles↗

Long terminal repeat retrotransposons of Oryza sativa.

BACKGROUND: Long terminal repeat (LTR) retrotransposons constitute a major fraction of the genomes of higher plants. For example, retrotransposons comprise more than 50% of the maize genome and more than 90% of the wheat genome. LTR retrotransposons are believed to have contributed significantly to the evolution of genome structure and function. The genome sequencing of selected experimental and agriculturally important species is providing an unprecedented opportunity to view the patterns of variation existing among the entire complement of retrotransposons in complete genomes. RESULTS: Using a new data-mining program, LTR_STRUC, (LTR retrotransposon structure program), we have mined the GenBank rice (Oryza sativa) database as well as the more extensive (259 Mb) Monsanto rice dataset for LTR retrotransposons. Almost two-thirds (37) of the 59 families identified consist of copia-like elements, but gypsy-like elements outnumber copia-like elements by a ratio of approximately 2:1. At least 17% of the rice genome consists of LTR retrotransposons. In addition to the ubiquitous gypsy- and copia-like classes of LTR retrotransposons, the rice genome contains at least two novel families of unusually small, non-coding (non-autonomous) LTR retrotransposons. CONCLUSIONS: Each of the major clades of rice LTR retrotransposons is more closely related to elements present in other species than to the other clades of rice elements, suggesting that horizontal transfer may have occurred over the evolutionary history of rice LTR retrotransposons. Like LTR retrotransposons in other species with relatively small genomes, many rice LTR retrotransposons are relatively young, indicating a high rate of turnover.

Animals↗