Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75Linked to original sources

Locus-specific genetic differentiation at Rw among warfarin-resistant rat (Rattus norvegicus) populations.

Populations may diverge at fitness-related genes as a result of adaptation to local conditions. The ability to detect this divergence by marker-based genomic scans depends on the relative magnitudes of selection, recombination, and migration. We survey rat (Rattus norvegicus) populations to assess the effect that local selection with anticoagulant rodenticides has had on microsatellite marker variation and differentiation at the warfarin resistance gene (Rw) relative to the effect on the genomic background. Initially, using a small sample of 16 rats, we demonstrate tight linkage of microsatellite D1Rat219 to Rw by association mapping of genotypes expressing an anticoagulant-rodenticide-insensitive vitamin K 2,3-epoxide reductase (VKOR). Then, using allele frequencies at D1Rat219, we show that predicted and observed resistance levels in 27 populations correspond, suggesting intense and recent selection for resistance. A contrast of F(ST) values between D1Rat219 and the genomic background revealed that rodenticide selection has overwhelmed drift-mediated population structure only at Rw. A case-controlled design distinguished these locus-specific effects of selection at Rw from background levels of differentiation more effectively than a population-controlled approach. Our results support the notion that an analysis of locus-specific population genetic structure may assist the discovery and mapping of novel candidate loci that are the object of selection or may provide supporting evidence for previously identified loci.

Animals↗

Unexpected complexity in the haplotypes of commonly used inbred strains of laboratory mice.

Investigation of sequence variation in common inbred mouse strains has revealed a segmented pattern in which regions of high and low variant density are intermixed. Furthermore, it has been suggested that allelic strain distribution patterns also occur in well defined blocks and consequently could be used to map quantitative trait loci (QTL) in comparisons between inbred strains. We report a detailed analysis of polymorphism distribution in multiple inbred mouse strains over a 4.8-megabase region containing a QTL influencing anxiety. Our analysis indicates that it is only partly true that the genomes of inbred strains exist as a patchwork of segments of sequence identity and difference. We show that the definition of haplotype blocks is not robust and that methods for QTL mapping may fail if they assume a simple block-like structure.

Alleles↗

Bioinformatics pipeline for the systematic mining genomic and proteomic variation linked to rare diseases: The example of monogenic diabetes.

Monogenic diabetes is characterized as a group of diseases caused by rare variants in single genes. Like for other rare diseases, multiple genes have been linked to monogenic diabetes with different measures of pathogenicity, but the information on the genes and variants is not unified among different resources, making it challenging to process them informatically. We have developed an automated pipeline for collecting and harmonizing data on genetic variants linked to monogenic diabetes. Furthermore, we have translated variant genetic sequences into protein sequences accounting for all protein isoforms and their variants. This allows researchers to consolidate information on variant genes and proteins linked to monogenic diabetes and facilitates their study using proteomics or structural biology. Our open and flexible implementation using Jupyter notebooks enables tailoring and modifying the pipeline and its application to other rare diseases.

Humans↗

Hierarchical analysis of population genetic variation in mitochondrial and nuclear genes of Daphnia pulex.

The geographic structure of Daphnia pulex populations from the central United States is analyzed with respect to isozyme and mitochondrial DNA variation. The species complex consists of cyclic and obligate parthenogens. A hierarchical analysis of population structure in the cyclic parthenogens by using a fixation-index approach indicates that this is one of the most extremely subdivided species yet studied. This genetic structure, much of which accrues within 100 km, is certainly due in part to the limited dispersal ability of Daphnia. However, previous work has shown that fluctuating selection can account for the spatial heterogeneity in isozyme frequencies in these populations. This may explain why the population subdivision for the mitochondrial genome increases approximately three times as rapidly with distance as does that for nuclear genes, which is slower than the neutral expectation. The obligate parthenogens are shown to be polyphyletic in origin, evolutionarily young, and, in some cases, geographically widespread.

Animals↗

Structure and development of neuronal connections in isogenic organisms: variations and similarities in the optic system of Daphnia magna.

It is readily apparent upon examination and comparison of organisms of the same and related species that to a large extent genes control morphological features. One means to ascertain the role and degree of genetic control is to study in detail the anatomy and development of a particular structure, say an identifiable neuron, in a population with a fixed genome, and at different stages of morphogenesis. The small crustacean, Daphnia, is well suited for this type of study, since it reproduces parthenogenically, its development is easily staged, and its nervous system is reasonably simple. In this paper comparisons of detailed morphology of neurons in the visual system of adult Daphnia magna are considered. Results indicate that gross features of the system are reproduced within the clone, but some of the finer details are not reproduced.

Animals↗

Evolution of naturally occurring 5' non-translated region variants of hepatitis C virus genotype 1b in selectable replicons.

Quasispecies shifts are essential for the development of persistent hepatitis C virus (HCV) infection. Naturally occurring sequence variations in the 5' non-translated region (NTR) of the virus could lead to changes in protein expression levels, reflecting selective forces on the virus. The extreme 5' end of the virus' genome, containing signals essential for replication, is followed by an internal ribosomal entry site (IRES) essential for protein translation as well as replication. The 5' NTR is highly conserved and has a complex RNA secondary structure consisting of several stem-loops. This report analyses the quasispecies distribution of the 5' NTR of an HCV genotype 1b clinical isolate and found a number of sequences differing from the consensus sequence. The consensus sequence, as well as a major variant located in stem-loop IIIa of the IRES, was investigated using self-replicating HCV RNA molecules in human hepatoma cells. The stem-loop IIIa mutation, which is predicted to disrupt the stem structure, showed slightly lower translation efficiency but was severely impaired in the colony formation of selectable HCV replicons. Interestingly, during selection of colonies supporting autonomous replication, mutations emerged that restored the base pairing in the stem-loop. Recloning of these altered IRESs confirmed that these second site revertants were more efficient in colony formation. In conclusion, naturally occurring variants in the HCV 5' NTR can lead to changes in their replication ability. Furthermore, IRES quasispecies evolution was observed in vitro under the selective pressure of the replicon system.

5' Untranslated Regions↗

Human microsatellites applicable for analysis of genetic variation in apes and Old World monkeys.

In studies of the genetics and social structure of primate populations there is a need to develop highly variable genetic markers for characterizing mating success and the nature of population movement or change through time. Because of their highly polymorphic nature, relatively simple amplification and typing, and the possibility of noninvasive sampling, microsatellites have become the molecular tool of choice in such studies. However, until recently it was assumed that many microsatellite loci, which are primarily situated in noncoding regions of the genome, evolve too rapidly to be applicable in evolutionarily divergent species. This has often resulted in the time-consuming process of cloning and sequencing microsatellites in new species. Here we describe the application of 11 human microsatellite primer pairs to a large group of primate species. The loci described are informative in all major groups of apes and Old World monkeys, although levels of allelic variability and heterozygosity differ across species. We confirm that with the use of appropriate universally applicable PCR conditions, a subset of human microsatellites are informative genetic markers in a wide range of divergent primate taxa.

Animals↗

A quantitative analysis of secondary RNA structure using domination based parameters on trees.

BACKGROUND: It has become increasingly apparent that a comprehensive database of RNA motifs is essential in order to achieve new goals in genomic and proteomic research. Secondary RNA structures have frequently been represented by various modeling methods as graph-theoretic trees. Using graph theory as a modeling tool allows the vast resources of graphical invariants to be utilized to numerically identify secondary RNA motifs. The domination number of a graph is a graphical invariant that is sensitive to even a slight change in the structure of a tree. The invariants selected in this study are variations of the domination number of a graph. These graphical invariants are partitioned into two classes, and we define two parameters based on each of these classes. These parameters are calculated for all small order trees and a statistical analysis of the resulting data is conducted to determine if the values of these parameters can be utilized to identify which trees of orders seven and eight are RNA-like in structure. RESULTS: The statistical analysis shows that the domination based parameters correctly distinguish between the trees that represent native structures and those that are not likely candidates to represent RNA. Some of the trees previously identified as candidate structures are found to be "very" RNA like, while others are not, thereby refining the space of structures likely to be found as representing secondary RNA structure. CONCLUSION: Search algorithms are available that mine nucleotide sequence databases. However, the number of motifs identified can be quite large, making a further search for similar motif computationally difficult. Much of the work in the bioinformatics arena is toward the development of better algorithms to address the computational problem. This work, on the other hand, uses mathematical descriptors to more clearly characterize the RNA motifs and thereby reduce the corresponding search space. These preliminary findings demonstrate that graph-theoretic quantifiers utilized in fields such as computer network design hold significant promise as an added tool for genomics and proteomics.

Algorithms↗

How the global structure of protein interaction networks evolves.

Two processes can influence the evolution of protein interaction networks: addition and elimination of interactions between proteins, and gene duplications increasing the number of proteins and interactions. The rates of these processes can be estimated from available Saccharomyces cerevisiae genome data and are sufficiently high to affect network structure on short time-scales. For instance, more than 100 interactions may be added to the yeast network every million years, a fraction of which adds previously unconnected proteins to the network. Highly connected proteins show a greater rate of interaction turnover than proteins with few interactions. From these observations one can explain (without natural selection on global network structure) the evolutionary sustenance of the most prominent network feature, the distribution of the frequency P(d) of proteins with d neighbours, which is broad-tailed and consistent with a power law, that is: P(d) proportional, variant d (-gamma).

Evolution, Molecular↗

Simple sequence repeats and compositional bias in the bipartite Ralstonia solanacearum GMI1000 genome.

BACKGROUND: Ralstonia solanacearum is an important plant pathogen. The genome of R. solananearum GMI1000 is organised into two replicons (a 3.7-Mb chromosome and a 2.1-Mb megaplasmid) and this bipartite genome structure is characteristic for most R. solanacearum strains. To determine whether the megaplasmid was acquired via recent horizontal gene transfer or is part of an ancestral single chromosome, we compared the abundance, distribution and composition of simple sequence repeats (SSRs) between both replicons and also compared the respective compositional biases. RESULTS: Our data show that both replicons are very similar in respect to distribution and composition of SSRs and presence of compositional biases. Minor variations in SSR and compositional biases observed may be attributable to minor differences in gene expression and regulation of gene expression or can be attributed to the small sample numbers observed. CONCLUSIONS: The observed similarities indicate that both replicons have shared a similar evolutionary history and thus suggest that the megaplasmid was not recently acquired from other organisms by lateral gene transfer but is a part of an ancestral R. solanacearum chromosome.

Base Composition↗

Use of variable simple sequence motifs as genetic markers: application to study of myotonic dystrophy.

Among the many classes of repetitive elements present in the human genome, the ubiquitous "simple sequence motifs" (SSMs) composed of [A]n, [TG]n, [AG]n or codon-tandem repeats form a major source of genetic variation. Here we report a detailed molecular-genetic study of a "variable simple sequence motif" (VSSM) in the apolipoprotein C2 (apoC2) gene, which maps to the 19q13.2 region in the vicinity of the myotonic dystrophy (DM) locus. By combining in vitro DNA-amplification using the polymerase chain reaction and high-resolution gel electrophoresis, we could demonstrate a high degree of allelic variation with at least ten alleles, which differ in the number of repeated [TG] or [AG] dinucleotide units. Similar results were found for the somatostatin I gene locus. To evaluate the usefulness of SSM-length polymorphisms as genetic markers, the apoC2-VSSM was employed for linkage analysis in DM families. Our results establish that the orientation of the apolipoprotein gene cluster on 19q is cenapoE-apoC2-ter and indicate that the many thousands of structurally similar VSSMs in the human genome represent a rich source of highly informative genetic and diagnostic markers.

Apolipoprotein C-II↗

Haplotype reconstruction from genotype data using Imperfect Phylogeny.

UNLABELLED: Critical to the understanding of the genetic basis for complex diseases is the modeling of human variation. Most of this variation can be characterized by single nucleotide polymorphisms (SNPs) which are mutations at a single nucleotide position. To characterize the genetic variation between different people, we must determine an individual's haplotype or which nucleotide base occurs at each position of these common SNPs for each chromosome. In this paper, we present results for a highly accurate method for haplotype resolution from genotype data. Our method leverages a new insight into the underlying structure of haplotypes that shows that SNPs are organized in highly correlated 'blocks'. In a few recent studies, considerable parts of the human genome were partitioned into blocks, such that the majority of the sequenced genotypes have one of about four common haplotypes in each block. Our method partitions the SNPs into blocks, and for each block, we predict the common haplotypes and each individual's haplotype. We evaluate our method over biological data. Our method predicts the common haplotypes perfectly and has a very low error rate (<2% over the data) when taking into account the predictions for the uncommon haplotypes. Our method is extremely efficient compared with previous methods such as PHASE and HAPLOTYPER. Its efficiency allows us to find the block partition of the haplotypes, to cope with missing data and to work with large datasets. AVAILABILITY: The algorithm is available via a Web server at http://www.calit2.net/compbio/hap/

Algorithms↗

Analysis of a dispersed repetitive DNA sequence in isogenic lines of Drosophila.

The location of sequences homologous to a cloned D. melanogaster DNA segment, Dm 25, has been examined in polytene chromosomes by hybridization in situ. Dm 25 localizes to multiple sites and shows variation in patterns between different strains and among individuals within wild-type laboratory strains. Analysis of numerous geographically distinct isogenic lines suggests that Dm 25 patterns are determined by germ-line factors and are not the product of strictly somatic events. In general there is wide variation in Dm 25 patterns among different lines, but a significant number of sites are common to two or more distinct lines. Hybridization to restriction digests of genomic DNA suggests that Dm 25 is a moderately repetitive, conserved sequence whose copies are dispersed throughout the genome. Analysis of species other than melanogaster indicates a significant divergence in structure of sequences homologous to Dm 25 as well as a drastic reduction in amount of homology to the melanogaster sequence.

Animals↗

Population dynamics of HIV-1 inferred from gene sequences.

A method for the estimation of population dynamic history from sequence data is described and used to investigate the past population dynamics of HIV-1 subtypes A and B. Using both gag and env gene alignments the effective population size of each subtype is estimated and found to be surprisingly small. This may be a result of the selective sweep of mutations through the population, or may indicate an important role of genetic drift in the fixation of mutations. The implications of these results for the spread of drug-resistant mutations and transmission dynamics, and also the roles of selection and recombination in shaping HIV-1 genetic diversity, are discussed. A larger estimated effective population size for subtype A may be the result of differences in time of origin, transmission dynamics, and/or population structure. To investigate the importance of population structure a model of population subdivision was fitted to each subtype, although the improvement in likelihood was found to be nonsignificant.

Genes, Viral↗

Quantitative trait locus mapping of pyrethroid resistance in Colorado potato beetle, Leptinotarsa decemlineata (Say) (Coleoptera: Chrysomelidae).

Quantitative trait locus (QTL) mapping is a valuable new tool for locating genomic regions that underlie variation in important traits such as insecticide resistance. Because QTL mapping complements a candidate gene strategy for understanding the genetic architecture of important traits, it may also facilitate the identification of genes causing important variation. After mapping the QTL locations, markers closely linked to QTL can be used for genetic analysis of population structure and to measure the spread and increase of resistance-causing QTL alleles. In this study, QTL influencing resistance to the pyrethroid insecticide esfenvalerate were mapped in the Colorado potato beetle Leptinotarsa decemlineata (Say) (CPB). Three QTL contributing to esfenvalerate resistance were identified from a mapping population of 79 individuals analyzed at 90 marker loci. One QTL had a large effect and two QTL had smaller effects. The major QTL occurs on the X chromosome, overlapping the position of a candidate gene (Leptinotarsa decemlineata Voltage sensitive sodium channel [LdVssc1]) previously implicated in pyrethroid resistance. Resistance-increasing alleles at the two minor-effect QTL originated with the susceptible parent, suggesting that alleles of small effect may be segregating in susceptible populations. Comparison of the New York population from which the susceptible parent originated with a more-susceptible population from North Carolina suggests that the minor-effect loci identified here may explain some of the variation in tolerance observed among susceptible populations. DNA sequencing of a portion of LdVssc1 shows that the resistance-conferring allele from the resistant parent does not contain the kdr mutation previously found in CPB and typically observed in other insects that are resistant to pyrethroid insecticides because of changes in this gene.

Animals↗

Influence of volcanic activity on the population genetic structure of Hawaiian Tetragnatha spiders: fragmentation, rapid population growth and the potential for accelerated evolution.

Volcanic activity on the island of Hawaii results in a cyclical pattern of habitat destruction and fragmentation by lava, followed by habitat regeneration on newly formed substrates. While this pattern has been hypothesized to promote the diversification of Hawaiian lineages, there have been few attempts to link geological processes to measurable changes in population structure. We investigated the genetic structure of three species of Hawaiian spiders in forests fragmented by a 150-year-old lava flow on Mauna Loa Volcano, island of Hawaii: Tetragnatha quasimodo (forest and lava flow generalist), T. anuenue and T. brevignatha (forest specialists). To estimate fragmentation effects on population subdivision in each species, we examined variation in mitochondrial and nuclear genomes (DNA sequences and allozymes, respectively). Population subdivision was higher for forest specialists than for the generalist in fragments separated by lava. Patterns of mtDNA sequence evolution also revealed that forest specialists have undergone rapid expansion, while the generalist has experienced more gradual population growth. Results confirm that patterns of neutral genetic variation reflect patterns of volcanic activity in some Tetragnatha species. Our study further suggests that population subdivision and expansion can occur across small spatial and temporal scales, which may facilitate the rapid spread of new character states, leading to speciation as hypothesized by H. L. Carson 30 years ago.

Animals↗

Type I antifreeze proteins expressed in snailfish skin are identical to their plasma counterparts.

Type I antifreeze proteins (AFPs) are usually small, Ala-rich alpha-helical polypeptides found in right-eyed flounders and certain species of sculpin. These proteins are divided into two distinct subclasses, liver type and skin type, which are encoded by separate gene families. Blood plasma from Atlantic (Liparis atlanticus) and dusky (Liparis gibbus) snailfish contain type I AFPs that are significantly larger than all previously described type I AFPs. In this study, full-length cDNA clones that encode snailfish type I AFPs expressed in skin tissues were generated using a combination of library screening and PCR-based methods. The skin clones, which lack both signal and pro-sequences, produce proteins that are identical to circulating plasma AFPs. Although all fish examined consistently express antifreeze mRNA in skin tissue, there is extreme individual variation in liver expression - an unusual phenomenon that has never been reported previously. Furthermore, genomic Southern blot analysis revealed that snailfish AFPs are products of multigene families that consist of up to 10 gene copies per genome. The 113-residue snailfish AFPs do not contain any obvious amino acid repeats or continuous hydrophobic face which typify the structure of most other type I AFPs. These structural differences might have implications for their ice-crystal binding properties. These results are the first to demonstrate a dual liver/skin role of identical type I AFP expression which may represent an evolutionary intermediate prior to divergence into distinct gene families.

Amino Acid Sequence↗

Genome-wide detection and analysis of cell wall-bound proteins with LPxTG-like sorting motifs.

Surface proteins of gram-positive bacteria often play a role in adherence of the bacteria to host tissue and are frequently required for virulence. A specific subgroup of extracellular proteins contains the cell wall-sorting motif LPxTG, which is the target for cleavage and covalent coupling to the peptidoglycan by enzymes called sortases. A comprehensive set of putative sortase substrates was identified by in silico analysis of 199 completely sequenced prokaryote genomes. A combination of detection methods was used, including secondary structure prediction, pattern recognition, sequence homology, and genome context information. With the hframe algorithm, putative substrates were identified that could not be detected by other methods due to errors in open reading frame calling, frameshifts, or sequencing errors. In total, 732 putative sortase substrates encoded in 49 prokaryote genomes were identified. We found striking species-specific variation for the LPxTG motif. A hidden Markov model (HMM) based on putative sortase substrates was created, which was subsequently used for the automatic detection of sortase substrates in recently completed genomes. A database was constructed, LPxTG-DB (http://bamics3.cmbi.kun.nl/sortase_substrates), containing for each genome a list of putative sortase substrates, sequence information of these substrates, the organism-specific HMMs based on the consensus sequence of the sortase recognition motif, and a graphic representation of this consensus.

Algorithms↗