Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Molecular epidemiology and diagnosis of Leishmania: what have we learnt from genome structure, dynamics and function?

This paper reviews our exploration of the dynamics of the Leishmania genome and its contribution to epidemiology and diagnosis. We used as a model Peruvian populations of L. (Viannia) braziliensis and L. (V.) peruviana, 2 species very close phylogenetically, but phenotypically very different in biotope and pathology. We initially focused on karyotype analysis. Our data showed that chromosomes were subject to a fast rate of evolution, and were sensitive indicators of genetic drift. Therefore, molecular karyotyping appeared an adequate tool for monitoring (i) emergence of close species, (ii) ecogeographical differentiation at the intraspecific level, and (iii) strain 'fingerprinting'. Chromosome size variation was mostly due to the number of tandemly repeated genes (rDNA, mini-exon, gp63, and cysteine proteinase genes), and could involve the deletion of unique genes (L. (V.) braziliensis-specific gp63 families). Considering the importance of these genes in parasitism, their rearrangement might have functional implications: adaptation to different environments and pleomorphic pathogenicity. Our knowledge of genome structure and dynamics was used to develop new polymerase chain reaction (PCR) techniques. Amplification of gp63 genes followed by cleavage with restriction enzymes and study of restriction fragment length polymorphism (gp63 PCR-RFLP) allowed the discrimination of all species tested, even directly in biopsies with 95% sensitivity (compared with PCR amplification of kinetoplast deoxyribonucleic acid). At the intra-specific level, RFLP was also observed and corresponded to mutations in major immunogen domains of gp63. These seem to be under strong selection pressure, and the technique should facilitate addressing how the host's immune pressure may modulate parasite population structure. Altogether, gp63 PCR-RFLP represents a significant operational improvement over the other techniques for molecular epidemiology and diagnosis: it combines sensitivity, discriminatory power and prognostic value.

Animals↗

Molecular structure and chromosomal localization of major repetitive DNA families in the chickpea (Cicer arietinum L.) genome.

Three major repetitive DNA sequences were isolated from a genomic library of chickpea (Cicer arietinum L.) and characterized with respect to their genomic organization and chromosomal localization. All repetitive elements are genus-specific and mostly located in the AT-rich pericentric heterochromatin. Two families are organized as satellite DNAs with repeat lengths of 162-168 bp (CaSat1) and 100 bp (CaSat2). CaSat1 is mainly located adjacent to the 18S rDNA clusters on chromosomes A and B, whereas CaSat2 is a major component of the pericentric heterochromatin on all chromosomes. The high abundance of these sequences in closely related species of the genus Cicer as well as their variation in structure and copy number among the annual species provide useful tools for taxonomic studies. The retrotransposon-like sequences of the third family (CaRep) display a more complex organization and are represented by two independent sets of clones (CaRep1 and CaRep2) with homology to different regions of Ty3-gypsy-like retrotransposons. They are distributed over the pericentric heterochromatin block on all chromosomes with extensions into euchromatic regions. Conserved structures within different crossability groups of related Cicer species suggest independent amplification or transposition events during the evolution of the annual species of the genus.

Amino Acid Sequence↗

Large Haplotypes Linked to Climate and Life History Variation in Divergent Lineages of Atlantic Salmon (Salmo salar).

Advances in sequencing are revealing that linked genomic architectures, enabling the evolution of co-adapted alleles at multiple loci, often shape complex phenotypes. Several recent studies have identified such architectures (e.g., chromosomal rearrangements and supergenes) contributing to adaptation or divergence across diverse species, from plants to mammals. Specifically, within Atlantic salmon (Salmo salar ), genomic studies are revealing large haplotypes and structural variants that may underpin local adaptation in the species. Using data from > 4000 individuals from 134 locations spanning the North Atlantic Ocean, we identify a large (~3 Mbp) genomic region on Ssa18 showing patterns of differentiation and linkage disequilibrium (LD) indicative of a large haplotype block containing three divergent haplotypes (herein A, B and C haplotypes). In Europe, haplotypes A and B were common, whereas A and C were more common within North America, suggesting a shared 'ancestral' A haplotype, with different continent-specific alternative haplotypes. Data support independent origins of divergent haplotypes in each continent, as well as signals of trans-oceanic introgression of haplotypes. Haplotype frequency is strongly associated with latitude, climate and life history (smolt age); however, the strength and direction of these relationships vary across continents. Overall, our analyses were consistent with other studies that identify chromosomal rearrangements; however, long-read sequence data did not find evidence of a structural variant, and instead an ancestral fusion may explain the formation and maintenance of the observed haplotypes. Our study contributes to ongoing efforts to understand the evolutionary role of linked genomic architecture in Atlantic salmon and its significance in salmonid diversification.

Climate Change↗

Molecular analysis of the ORFs 3 to 7 of porcine reproductive and respiratory syndrome virus, Québec reference strain.

The cDNA sequence of the 3'-terminal genomic region of the Québec IAF-exp91 strain of porcine reproductive and respiratory syndrome virus (PRRSV) was determined and compared to those of other reference strains from Europe (Lelystad virus) and US (ATCC VR2385, MN-1b). The sequence (2834 nucleotides) which encompassed ORFs 3 to 7 revealed extensive genomic variations between the Québec strain and Lelystad virus (LV), resulting from high number of base substitutions, additions and deletions. The ORFs 5, 3, and 7 seemed to be relatively the most variable; the predicted encoding products of the Québec and LV strains displayed only 52%, 54%, and 59% amino acid identities, respectively. Nevertheless, in vitro translation experiments of the structural genes (ORFs 5, 6, and 7) and radioimmunoprecipitation assays with extracellular virions gave results similar to those previously reported for LV. In contrast, close genomic relationships were demonstrated between Québec and US strains. Taking together, these results indicate that, although structurally similar, North American PRRSV strains belong to a genotype distinct from that of the LV, thus supporting previous findings that allowed to divide PRRSV isolates into two antigenic subgroups (U.S. and European).

Amino Acid Sequence↗

Characterization of the topoisomerase I locus in human colorectal cancer.

DNA topoisomerase I (topo I) is the principle target for Camptothecin and its analogues. The topo I gene is located on chromosome 20q11.2-q13.1 and variation in topo I gene copy number has been shown to have impact on the in vitro sensitivity to topoisomerase I inhibitor chemotherapy. Fluorescence in situ hybridization (FISH) was used to detect and compare the TOPO I gene copy number between metaphase and interphase nuclei in a panel of 7 colorectal cancer cell lines. TOPO I gene copy number varied from 2 to 8 between cell lines, and signal in interphase nuclei demonstrated a linear relationship with that detected in metaphase nuclei. The structure of gene amplification included isochromosome formation, amplicon extension, and marker chromosome generation. Comparative genomic hybridization (CGH) was then used to further define the region of gain on chromosome 20. The region of gain contained the topo I gene and involved nearly all of 20q in most cases. This demonstrates a high degree of intrinsic variation in topo I gene copy number and the involvement of a 20q amplicon in colorectal cancer, which may have important implications for colorectal tumorigenesis and the use of chemotherapy.

Caco-2 Cells↗

Phages of dairy bacteria.

Bacteriophages of lactic acid bacteria are a threat to industrial milk fermentation. Owing to their economical importance, dairy phages became the most thoroughly sequenced phage group in the database. Comparative genomics identified related cos-site and pac-site phages, respectively, in lactococci, lactic streptococci and lactobacilli. Each group was represented with closely related temperate and virulent phages. Over the structural genes their gene maps resembled that of lambdoid coliphages, suggesting distant evolutionary relationships. Despite a lack of sequence similarity, a number of biochemical characteristics of these dairy phages are lambda-like (genetic switch, DNA packaging, head and tail morphogenesis, and integration, but not excision). These dairy phages thus provide interesting variations to the phage lambda paradigm. The structural gene cluster of Lactococcus phage r1t resembled that of phages from mycobacteria. Virulent lactococcal phages with prolate heads (c2-like genus of Siphoviridae), in contrast, have no known counterparts in other bacterial genera.

Animals↗

Selective distribution of histone H1 variants and high mobility group proteins in chromosomes.

Recent results on the differential distribution of sequence variants of histone H1, of proteins of the HMG 1/2 family, and of HMG1 in polytene chromosomes are reviewed. Several organisms are known to contain two different HMG 1/2 proteins. In Chironomus, one of them is restricted to decondensed puffs and may have a specific function. One of the H1 variants of Chironomus is found only in a minority of chromosome bands and differs from the other H1 proteins of the organism by genomic organization and by an inserted structural motif that is also present in single H1 variants of other organisms.

Amino Acid Sequence↗

Measures of diversity for populations and distances between individuals with highly reorganizable genomes.

In this paper we address the problem of defining a measure of diversity for a population of individuals whose genome can be subjected to major reorganizations during the evolutionary process. To this end, we introduce a measure of diversity for populations of strings of variable length defined on a finite alphabet, and from this measure we derive a semi-metric distance between pairs of strings. The definitions are based on counting the number of substrings of the strings, considered first separately and then collectively. This approach is related to the concept of linguistic complexity, whose definition we generalize from single strings to populations. Using the substring count approach we also define a new kind of Tanimoto distance between strings. We show how to extend the approach to representations that are not based on strings and, in particular, to the tree-based representations used in the field of genetic programming. We describe how suffix trees can allow these measures and distances to be implemented with a computational cost that is linear in both space and time relative to the length of the strings and the size of the population. The definitions were devised to assess the diversity of populations having genomes of variable length and variable structure during evolutionary computation runs, but applications in quantitative genomics, proteomics, and pattern recognition can be also envisaged.

Computational Biology↗

Human pigmentation genes: identification, structure and consequences of polymorphic variation.

The synthesis of the visible pigment melanin by the melanocyte cell is the basis of the human pigmentary system, those genes directing the formation, transport and distribution of the specialised melanosome organelle in which melanin accumulates can legitimately be called pigmentation genes. The genes involved in this process have been identified through comparative genomic studies of mouse coat colour mutations and by the molecular characterisation of human hypopigmentary genetic diseases such as OCA1 and OCA2. The melanocyte responds to the peptide hormones alpha-MSH or ACTH through the MC1R G-protein coupled receptor to stimulate melanin production through induced maturation or switching of melanin type. The pheomelanosome, containing the key enzyme of the pathway tyrosinase, produces light red/yellowish melanin, whereas the eumelanosome produces darker melanins via induction of additional TYRP1, TYRP2, SILV enzymes, and the P-protein. Intramelanosomal pH governed by the P-protein may act as a critical determinant of tyrosinase enzyme activity to control the initial step in melanin synthesis or TYRP complex formation to facilitate melanogenesis and melanosomal maturation. The search for genetic variation in these candidate human pigmentation genes in various human populations has revealed high levels of polymorphism in the MC1R locus, with over 30 variant alleles so far identified. Functional correlation of MC1R alleles with skin and hair colour provides evidence that this receptor molecule is a principle component underlying normal human pigment variation.

Genes↗

Genomic structure of murine methylmalonyl-CoA mutase: evidence for genetic and epigenetic mechanisms determining enzyme activity.

Methylmalonyl-CoA mutase (MCM) is a nuclear-encoded mitochondrial matrix enzyme. We have reported characterization of murine MCM and cloning of a murine MCM cDNA and now describe the murine Mut locus, its promoter and evidence for tissue-specific variation in MCM mRNA, enzyme and holo-enzyme levels. The Mut locus spans 30 kb and contains 13 exons constituting a unique transcription unit. A B1 repeat element was found in the 3' untranslated region (exon 13). The transcription initiation site was identified and upstream sequences were shown to direct expression of a reporter gene in cultured cells. The promoter contains sequence motifs characteristic of: (1) TATA-less housekeeping promoters; (2) enhancer elements purportedly involved in co-ordinating expression of nuclear-encoded mitochondrial proteins; and (3) regulatory elements including CCAAT boxes, cyclic AMP-response elements and potential AP-2-binding sites. Northern blots demonstrate a greater than 10-fold variation in steady-state mRNA levels, which correlate with tissue levels of enzyme activity. However, the ratio of holoenzyme to total enzyme varies among different tissues, and there is no correlation between steady-state mRNA levels and holoenzyme activity. These results suggest that, although there may be regulation of MCM activity at the level of mRNA, the significance of genetic regulation is unclear owning to the presence of epigenetic regulation of holoenzyme formation.

Animals↗

Highly repetitive DNA sequence elements from Orseolia oryzae (Wood-Mason) discriminate between the Indian isolates and the Asian rice gall midge and the paspalum midge.

We described multicopy DNA clones isolated from a partial genomic library of Orseolia oryzae, based on reverse genomic hybridization, suitable for studying genetic variation in the Asian rice gall midge and other isomorphic species. Three clones produced monomorphic DNA band patterns between biotypes of O. oryzae but polymorphic patterns were produced between O. oryzae and O. fluvialis, the paspalum midge. These probes detect changes in the repetitive sequence structure between species and constitute the first genetic markers for distinguishing between field isolates of rice gall midge and related species of Orseolia. These will be useful in identifying and perhaps eradicating alternative hosts for this pest, and detecting early-season outbreaks of O. oryzae from light trap collections.

Animals↗

Comparisons of the genomic cis-elements and coding regions in RNA beta components of the hordeiviruses barley stripe mosaic virus, lychnis ringspot virus, and poa semilatent virus.

Nucleotide sequences of the genomic RNA beta components of hordeiviruses poa semilatent virus (PSLV) and lychnis ringspot virus (LRSV) were determined. PSLV and LRSV closely resemble barley stripe mosaic virus (BSMV), type hordeivirus, in the gene arrangement of their RNAs beta, comprising 5'-proximal beta a (coat protein) gene and downstream triple gene block (TGB) coding for the beta b, beta c, and beta d putative transport proteins. The beta a, beta b, beta c, and beta d proteins of the three hordeiviruses showed significant sequence similarity, with the respective proteins of PSLV and BSMV being closer to each other than to their counterparts of LSRV. Comparisons of the TGB-encoded proteins of hordeiviruses, potexviruses, carlaviruses, and furoviruses indicate that the first and second TGB genes belong to the monophyletic groups, whereas the third gene may have multiple ancestry. LRSV, PSLV, and BSMV showed remarkable variation in the 3'-untranslated regions of their genomic RNAs. Among the three hordeiviruses, LRSV has the shortest 3'-noncoding region that lacks tentative pseudoknot-forming elements conserved upstream of the 3'-tRNA-like structure in the BSMV and PSLV genomes. On the other hand, LRSV RNA beta, like that of BSMV, contained the internal poly(A) sequence that is absent from PSLV RNA.

Adaptation, Physiological↗

Bacteriophage MS2: molecular weight and spatial distribution of the protein and RNA components by small-angle neutron scattering and virus counting.

Small-angle neutron scattering (SANS) has been used to extend the structural characterization of the MS2 phage by examining its physical characteristics in solution. Specifically, the contrast variation technique was employed to determine the molecular weight of the individual components of the MS2 virion (protein shell and genomic RNA) and the spatial relationship of the genomic RNA to its protein shell. A consequence of this work was to evaluate a novel particle counting instrument, the integrated virus detection system (IVDS) that, in combination with SANS, has the potential to provide rapid quantitative physical characterization of unidentified viruses and phage.

Electrophoresis, Polyacrylamide Gel↗

Comparisons of the genetic structure of populations of Turnip mosaic virus in West and East Eurasia.

The genetic structure of populations of Turnip mosaic virus in Eurasia was assessed by making host range and gene sequence comparisons of 142 isolates. Most isolates collected in West Eurasia infected Brassica plants whereas those from East Eurasia infected both Brassica and Raphanus plants. Analyses of recombination sites (RSs) in five regions of the genome (one third of the full sequence) showed that the protein 1 (P1 gene) had recombined more frequently than the other gene regions in both subpopulations, but that the RSs were located in different parts of the genomes of the subpopulations. Estimates of nucleotide diversity showed that the West Eurasian subpopulation was more diverse than the East Eurasian subpopulation, but the Asian-BR group of the genes from the latter subpopulation had a greater nonsynonymous/synonymous substitution ratio, especially in the P1, viral genome-linked protein (VPg) and nuclear inclusion a proteinase (NIa-Pro) genes. These subpopulations seem to have evolved independently from the ancestral European population, and their genetic structure probably reflects founder effects.

Amino Acid Substitution↗

Motif prediction in ribosomal RNAs Lessons and prospects for automated motif prediction in homologous RNA molecules.

The traditional way to infer RNA secondary structure involves an iterative process of alignment and evaluation of covariation statistics between all positions possibly involved in basepairing. Watson-Crick basepairs typically show covariations that score well when examples of two or more possible basepairs occur. This is not necessarily the case for non-Watson-Crick basepairing geometries. For example, for sheared (trans Hoogsteen/Sugar edge) pairs, one base is highly conserved (always A or mostly A with some C or U), while the other can vary (G or A and sometimes C and U as well). RNA motifs consist of ordered, stacked arrays of non-Watson-Crick basepairs that in the secondary structure representation form hairpin or internal loops, multi-stem junctions, and even pseudoknots. Although RNA motifs occur recurrently and contribute in a modular fashion to RNA architecture, it is usually not apparent which bases interact and whether it is by edge-to-edge H-bonding or solely by stacking interactions. Using a modular sequence-analysis approach, recurrent motifs related to the sarcin-ricin loop of 23S RNA and to loop E from 5S RNA were predicted in universally conserved regions of the large ribosomal RNAs (16S- and 23S-like) before the publication of high-resolution, atomic-level structures of representative examples of 16S and 23S rRNA molecules in their native contexts. This provides the opportunity to evaluate the predictive power of motif-level sequence analysis, with the goal of automating the process for predicting RNA motifs in genomic sequences. The process of inferring structure from sequence by constructing accurate alignments is a circular one. The crucial link that allows a productive iteration of motif modeling and realignment is the comparison of the sequence variations for each putative pair with the corresponding isostericity matrix to determine which basepairs are consistent both with the sequence and the geometrical data.

Base Pairing↗

Restricted structural gene polymorphism in the Mycobacterium tuberculosis complex indicates evolutionarily recent global dissemination.

One-third of humans are infected with Mycobacterium tuberculosis, the causative agent of tuberculosis. Sequence analysis of two megabases in 26 structural genes or loci in strains recovered globally discovered a striking reduction of silent nucleotide substitutions compared with other human bacterial pathogens. The lack of neutral mutations in structural genes indicates that M. tuberculosis is evolutionarily young and has recently spread globally. Species diversity is largely caused by rapidly evolving insertion sequences, which means that mobile element movement is a fundamental process generating genomic variation in this pathogen. Three genetic groups of M. tuberculosis were identified based on two polymorphisms that occur at high frequency in the genes encoding catalase-peroxidase and the A subunit of gyrase. Group 1 organisms are evolutionarily old and allied with M. bovis, the cause of bovine tuberculosis. A subset of several distinct insertion sequence IS6110 subtypes of this genetic group have IS6110 integrated at the identical chromosomal insertion site, located between dnaA and dnaN in the region containing the origin of replication. Remarkably, study of approximately 6,000 isolates from patients in Houston and the New York City area discovered that 47 of 48 relatively large case clusters were caused by genotypic group 1 and 2 but not group 3 organisms. The observation that the newly emergent group 3 organisms are associated with sporadic rather than clustered cases suggests that the pathogen is evolving toward a state of reduced transmissability or virulence.

Alleles↗

Functions of the gene products of Escherichia coli.

A list of currently identified gene products of Escherichia coli is given, together with a bibliography that provides pointers to the literature on each gene product. A scheme to categorize cellular functions is used to classify the gene products of E. coli so far identified. A count shows that the numbers of genes concerned with small-molecule metabolism are on the same order as the numbers concerned with macromolecule biosynthesis and degradation. One large category is the category of tRNAs and their synthetases. Another is the category of transport elements. The categories of cell structure and cellular processes other than metabolism are smaller. Other subjects discussed are the occurrence in the E. coli genome of redundant pairs and groups of genes of identical or closely similar function, as well as variation in the degree of density of genetic information in different parts of the genome.

Bacterial Proteins↗