Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

Pronounced genetic population structure in a potentially vagile fish species (Pristipomoides multidens, Teleostei; Perciformes; Lutjanidae) from the East Indies triangle.

The East Indies triangle, bordered by the Phillipines, Malay Peninsula and New Guinea, has a high level of tropical marine species biodiversity. Pristipomoides multidens is a large, long-lived, fecund snapper species that is distributed throughout the East Indies and Indo-Pacific. Samples were analysed from central and eastern Indonesia and northern Australia to test for genetic discontinuities in population structure. Fish (n = 377) were collected from the Indonesian islands of Bali, Sumbawa, Flores, West Timor, Tanimbar and Tual along with 131 fish from two northern Australian locations (Arafura and Timor Seas) from a previous study. Genetic variation in the control region of the mitochondrial genome was assayed using restriction fragment length polymorphism and direct sequencing. Haplotype diversity was high (0.67-0.82), as was intraspecific sequence divergence (range 0-5.8%). F(ST) between pairs of populations ranged from 0 to 0.2753. Genetic subdivision was apparent on a small spatial scale; F(ST) was 0.16 over 191 km (Bali/Sumbawa) and 0.17 over 491 km (Bali/Flores). Constraints to dispersal that contribute to, and maintain, the observed degree of genetic subdivision are experienced presumably by all life history stages of this tropical marine finfish. The constraints may include (1) little or no movement of eggs or larvae, (2) little or no home range or migratory movement of adults and (3) loss of larval cohorts due to transport of larvae away from suitable habitat by prevailing currents.

Analysis of Variance↗

DNA methylation and chromosome instability in lymphoblastoid cell lines.

In order to gain more insight into the relationships between DNA methylation and genome stability, chromosomal and molecular evolutions of four Epstein-Barr virus-transformed human lymphoblastoid cell lines were followed in culture for more than 2 yr. The four cell lines underwent early, strong overall demethylation of the genome. The classical satellite-rich, heterochromatic,juxtacentromeric regions of chromosomes 1, 9, and 16 and the distal part of the long arm of the Y chromosome displayed specific behavior with time in culture. In two cell lines, they underwent a strong demethylation, involving successively chromosomes Y, 9, 16, and 1, whereas in the two other cell lines, they remained heavily methylated. For classical satellite 2-rich heterochromatic regions of chromosomes 1 and 16, a direct relationship could be established between their demethylation, their undercondensation at metaphase, and their involvement in non-clonal rearrangements. Unstable sites distributed along the whole chromosomes were found only when the heterochromatic regions of chromosomes 1 and 16 were unstable. The classical satellite 3-rich heterochromatic region of chromosomes 9 and Y, despite their strong demethylation, remained condensed and stable. Genome demethylation and chromosome instability could not be related to variations in mRNA amounts of the DNA methyltransferases DNMT1, DNMT3A, and DNMT3B and DNA demethylase. These data suggest that the influence of DNA demethylation on chromosome stability is modulated by a sequence-specific chromatin structure.

Ataxia Telangiectasia↗

Evolutionary dynamics of the chloroplast genome in Abutilon (Malvoideae, Malvaceae).

The genus Abutilon Mill. (Malvaceae) comprises approximately 178 species distributed across tropical and subtropical regions, many of which hold significant ornamental, economic, and medicinal value; yet its taxonomic classification remains challenging. In this study, six species were sequenced from herbarium specimens, and the chloroplast (cp.) genomes of ten additional species were assembled de novo from publicly available raw data. Three previously reported cp. genomes were also incorporated to characterise cp. genome structure, identify polymorphic loci, and perform phylogenetic analyses. The cp. genomes ranged from 159,458 to 160,454 bp and exhibited the typical quadripartite structure, with each genome containing 112 unique genes (78 protein-coding, 30 tRNA, and 4 rRNA) that showed conserved content and organisation. These genomes exhibited high similarity in GC content, inverted repeat boundaries, relative synonymous codon usage, amino acid frequencies, and substitution patterns. However, notable variation was observed in the total number of simple sequence repeats, ranging from 70 to 97 per genome. Selection analyses indicated predominant purifying selection, with evidence of episodic positive selection detected in rpoC2, rbcL, and ycf1. Two codons in rbcL were clade-specific and provided phylogenetic signal distinguishing Australian and Old World pantropical species. Nucleotide diversity analysis identified six highly polymorphic intergenic spacers (trnH-psbA, rps19-rpl2, psbT-pbf1, psaC-ndhD, trnR-atpA, and ndhJ-ndhK) that may be suitable for taxonomic studies. The phylogeny from maximum likelihood (ML) and Bayesian inference (BI) resolved two major clades: one comprising an exclusively Australian lineage occurring predominantly in arid and semi-arid environments, and the other a pantropical lineage spanning multiple continents. Abutilon grandifolium was recovered as sister to the remaining sampled Abutilon taxa in both ML and BI analyses, although no biogeographic origin inference can be drawn from this placement pending broader taxon sampling and integration of nuclear genomic data. These findings provide insights into the evolutionary dynamics of the cp. genome in Abutilon and offer a foundational genomic framework for refining Abutilon taxonomy.

Genome, Chloroplast↗

Amino acid coupling patterns in thermophilic proteins.

Structural analysis is useful in elucidating structural features responsible for enhanced thermal stability of proteins. However, due to the rapid increase of sequenced genomic data, there are far more protein sequences than the corresponding three-dimensional (3D) structures. The usual sequence-based amino acid composition analysis provides useful but simplified clues about the amino acid types related to thermal stability of proteins. In this work, we developed a statistical approach to identify the significant amino acid coupling sequence patterns in thermophilic proteins. The amino acid coupling sequence pattern is defined as any 2 types of amino acids separated by 1 or more amino acids. Using this approach, we construct the rho profiles for the coupling patterns. The rho value gives a measure of the relative occurrence of a coupling pattern in thermophiles compared with mesophiles. We found that thermophiles and mesophiles exhibit significant bias in their amino acid coupling patterns. We showed that such bias is mainly due to temperature adaptation instead of species or GC content variations. Though no single outstanding coupling pattern can adequately account for protein thermostability, we can use a group of amino acid coupling patterns having strong statistical significance (p values < 10(-7)) to distinguish between thermophilic and mesophilic proteins. We found a good correlation between the optimal growth temperatures of the genomes and the occurrences of the coupling patterns (the correlation coefficient is 0.89). Furthermore, we can separate the thermophilic proteins from their mesophilic orthologs using the amino acid coupling patterns. These results may be useful in the study of the enhanced stability of proteins from thermophiles-especially when structural information is scarce. Proteins 2005. (c) 2005 Wiley-Liss, Inc.

Amino Acid Sequence↗

Genetic effects of contaminant exposure--towards an assessment of impacts on animal populations.

This review aims both to identify the potential risks to animal populations as a consequence of exposure to genotoxins and to identify the techniques most useful in assessing these risks. These evaluations are complicated by the fact that contaminant exposure acts both to restructure naturally occurring genetic diversity and, when contaminants have mutagenic activity, to enhance the rate of introduction of new variation. There is now evidence that contaminant exposure often leads to change in the genetic attributes of natural populations. Short-lived organisms often develop resistance to contaminants, with only modest impacts on diversity in the balance of the genome, although massive mortality occurs during the gene replacement. Resistance is, however, less likely to evolve in species with small population size, such as many wildlife species. Such species will experience population declines or extinction as the impact of contaminants on physiological systems is not counteracted by gene replacements. Even when adaptation to exposure occurs, populations may suffer diminished fitness as a consequence of the mutagenic effects of contaminants. The expression of these effects range from an increase in the incidence of developmental abnormalities to shifts in chromosomal and gene structure. The assessment of this broad range of impacts can only be accomplished with a spectrum of analytical approaches. However, recent advances in molecular and developmental genetics are now making possible the detailed assessment of these mutagenic impacts in natural populations.

Animals↗

Characterization of the 5' internal ribosome entry site of Plautia stali intestine virus.

The RNA genome of Plautia stali intestine virus (PSIV; Cripavirus, Dicistroviridae) contains two open reading frames, the first of which is preceded by a 570 nt untranslated region (5' UTR). The 5' UTR was confirmed to be an internal ribosome entry site (IRES) using an insect cell lysate translation system: translation of a second cistron increased 14-fold in the presence of the 5' UTR and a cap analogue did not inhibit translation of the second cistron. Deletion analysis showed that 349 bases corresponding to nt 225-573 in the PSIV genome were necessary for internal initiation. The PSIV 5' IRES did not function in rabbit reticulocyte lysate or wheatgerm translation systems; however, the intergenic IRES for capsid translation of PSIV was functional in both systems, indicating that the 5' IRES and the intergenic IRES have distinct requirements for their activities. Chemical and enzymic analyses of the 5' IRES of PSIV indicate that its structure is distinct from that of Rhopalosiphum padi virus. Because 5' IRES elements in some dicistroviruses have been reported to be active in plant and mammalian cell-free translation systems, there appears to be variation among dicistroviruses in the mechanism of translation initiation mediated by 5' IRES elements.

5' Untranslated Regions↗

The MMA1 gene family of cancer-testis antigens has multiple alternative splice variants: characterization of their expression profile, the genomic organization, and transcript properties.

Previously, we reported the identification of MMA1A by screening for differential gene expression in two human melanoma cell lines displaying diverse metastatic behavior after subcutaneous inoculation into nude mice. Splice variant MMA1B, which also was identified through database homology searches, showed a high degree of similarity with the MMA1A for exons 1, 2, and 4, but was missing exon 3. Through extensive expression profiling among normal and tumor samples, both MMA1A and -1B were found to belong to the family of cancer-testis antigens. In this study, we identified four additional alternatively spliced MMA1 variants, named MMA1C, MMA1D, MMA1E, and MMA1F. Generally, these novel MMA1 transcripts differ from MMA1A in that exon 2 or exon 3 is enlarged because of the use of alternative splice sites in intron 2 of the MMA1 gene. Moreover, MMA1E also lacks exon 3, as was previously seen in MMA1B. In screening for expression of the novel MMA1 transcripts in normal and tumor tissues, we demonstrated that MMA1C, -1D, and -1E also are members of the cancer-testis antigen family. MMA1F was found in only one melanoma metastasis sample and therefore is believed to have been expressed incidentally. Furthermore, we comprehensively elucidated the genomic structure of the MMA1 gene and the characteristic features of the alternatively spliced MMA1 transcripts.

Alternative Splicing↗

Evolutionary coupling between the deleteriousness of gene mutations and the amount of non-coding sequences.

The phenotypic effects of random mutations depend on both the architecture of the genome and the gene-trait relationships. Both levels thus play a key role in the mutational variability of the phenotype, and hence in the long-term evolutionary success of the lineage. Here, by simulating the evolution of organisms with flexible genomes, we show that the need for an appropriate phenotypic variability induces a relationship between the deleteriousness of gene mutations and the quantity of non-coding sequences maintained in the genome. The more deleterious the gene mutations, the shorter the intergenic sequences. Indeed, in a shorter genome, fewer genes are affected by rearrangements (duplications, deletions, inversions, translocations) at each replication, which compensates for the higher impact of each gene mutation. This spontaneous adjustment of genome structure allows the organisms to retain the same average fitness loss per replication, despite the higher impact of single gene mutations. These results show how evolution can generate unexpected couplings between distinct organization levels.

Animals↗

Evolutionary pathways of the calcitonin (CALC) genes.

Recombinant DNA techniques have made it possible to establish the structure of various genes encoding polypeptide hormones. Comparison of nucleotide sequences of the calcitonin (CALC) genes in man has revealed surprising similarities and variations. These findings and the homologies among the sequences in different species offered an opportunity for speculation about relationships between these genes and about their evolutionary origin. The first gene (CALC-I) directing the synthesis of calcitonin (CT) or CT gene-related peptide (CGRP) comprises six exons and gives rise to two mRNAs by an alternative RNA-processing mechanism. The homology between CGRP and CT reflects their common origin. The human genome contains a second gene (CALC-II) that is structurally related to the CALC-I gene. The CALC-II RNA transcripts do not appear to be differentially processed, as only preproCGRP-II mRNA and not preproCT-II is detected. The first and second CT/CGRP genes probably have evolved from a common ancestor gene early in evolution. Meanwhile, a third genomic locus containing nucleotide sequences highly homologous to exons 2 and 3 of both CALC genes was detected and probably generated by duplication of a part of CALC-II. This locus is not likely to encode a CT- or CGRP-related polypeptide hormone. The CALC genes and this last (pseudo) gene are located on the short arm of chromosome 11. Recently, islet- or insulinoma-amyloid polypeptide (IAPP) was isolated as a major constituent of amyloid present in human insulinoma and in pancreatic islet amyloid in noninsulin-dependent diabetes mellitus. IAPP shows 46% amino acid sequence homology with human CGRP-II.(ABSTRACT TRUNCATED AT 250 WORDS)

Amyloid↗

The non-producer phenotype of the human immunodeficiency virus type 1 provirus F12/HIV-1 is the result of multiple genetic variations.

A cell clone (Hut-78/F12) chronically infected with a non-producer human immunodeficiency virus type 1 (HIV-1) variant showed an abnormal pattern of virus structural proteins and released no detectable virus particles. Exchanges of homologous parts of the F12/HIV provirus and a replication-competent HIV (strain NL4-3) were undertaken to define the genetic determinants of the F12/HIV phenotype. The non-infectious phenotype was reproduced by replacing an NL4-3 genomic fragment encoding the C terminus of gp 120 and the N terminus of gp41 with the corresponding parts of the F12/HIV provirus. Conversely, a much more extended genomic fragment (encompassing the vif, pol and env genes) was necessary to convert the F12/HIV phenotype. These results demonstrate that the F12/HIV non-producer phenotype is the result of mutations scattered along most of the genome, rendering the conversion to an infectious phenotype a very unlikely event. The F12/HIV genome is thus a reliable model for preclinical studies of anti-HIV gene therapy.

Defective Viruses↗

Evolution of function in protein superfamilies, from a structural perspective.

The recent growth in protein databases has revealed the functional diversity of many protein superfamilies. We have assessed the functional variation of homologous enzyme superfamilies containing two or more enzymes, as defined by the CATH protein structure classification, by way of the Enzyme Commission (EC) scheme. Combining sequence and structure information to identify relatives, the majority of superfamilies display variation in enzyme function, with 25 % of superfamilies in the PDB having members of different enzyme types. We determined the extent of functional similarity at different levels of sequence identity for 486,000 homologous pairs (enzyme/enzyme and enzyme/non-enzyme), with structural and sequence relatives included. For single and multi-domain proteins, variation in EC number is rare above 40 % sequence identity, and above 30 %, the first three digits may be predicted with an accuracy of at least 90 %. For more distantly related proteins sharing less than 30 % sequence identity, functional variation is significant, and below this threshold, structural data are essential for understanding the molecular basis of observed functional differences. To explore the mechanisms for generating functional diversity during evolution, we have studied in detail 31 diverse structural enzyme superfamilies for which structural data are available. A large number of variations and peculiarities are observed, at the atomic level through to gross structural rearrangements. Almost all superfamilies exhibit functional diversity generated by local sequence variation and domain shuffling. Commonly, substrate specificity is diverse across a superfamily, whilst the reaction chemistry is maintained. In many superfamilies, the position of catalytic residues may vary despite playing equivalent functional roles in related proteins. The implications of functional diversity within supefamilies for the structural genomics projects are discussed. More detailed information on these superfamilies is available at http://www.biochem.ucl.ac.uk/bsm/FAM-EC/.

Binding Sites↗

The A locus that controls anthocyanin accumulation in pepper encodes a MYB transcription factor homologous to Anthocyanin2 of Petunia.

Pepper plants containing the dominant A gene accumulate anthocyanin pigments in the foliage, flower and immature fruit. We previously mapped A to pepper chromosome 10 in the F(2) progeny of a cross between 5226 (purple-fruited) and PI 159234 (green-fruited) to a region that corresponds, in tomato, to the location of Petunia anthocyanin 2 ( An2), a regulator of anthocyanin biosynthesis. This suggested that A encodes a homologue of Petunia An2. Using the sequences of An2 and a corresponding tomato expressed sequence tag, we isolated a pepper cDNA orthologous to An2 that cosegregated with A. We subsequently determined the expression of A by Northern analysis, using RNA extracted from fruits, flowers and leaves of 5226 and PI 159234. In 5226, expression was detected in all stages of fruit development and in both flower and leaf. In contrast, A was not expressed in the sampled tissues in PI 159234. Genomic sequence comparison of A between green- and purple-fruited genotypes revealed no differences in the coding region, indicating that the lack of expression of A in the green genotypes can be attributed to variation in the promoter region. By analyzing the expression of the structural genes in the anthocyanin biosynthetic pathway in 5226 and PI 159234, it was determined that, similar to Petunia, the early genes in the pathway are regulated independently of A, while expression of the late genes is A-dependent.

Anthocyanins↗

Putative DNA quadruplex formation within the human c-kit oncogene.

The DNA sequence, d(AGGGAGGGCGCTGGGAGGAGGG), occurs within the promoter region of the c-kit oncogene. We show here, using a combination of NMR, circular dichroism, and melting temperature measurements, that this sequence forms a four-stranded quadruplex structure under physiological conditions. Variations in the sequences that intervene between the guanine tracts have been examined, and surprisingly, none of these modified sequences forms a quadruplex arrangement under these conditions. This suggests that the occurrence of quadruplex-forming sequences within the human and other genomes is less than was hitherto expected. The c-kit quadruplex may be a new target for therapeutic intervention in cancers where there is elevated expression of the c-kit gene.

Base Sequence↗

An evolutionarily mobile antigen receptor variable region gene: doubly rearranging NAR-TcR genes in sharks.

Distinctive Ig and T cell receptor (TcR) chains define the two major lineages of vertebrate lymphocyte yet similarly recognize antigen with a single, membrane-distal variable (V) domain. Here we describe the first antigen receptor chain that employs two V domains, which are generated by separate VDJ gene rearrangement events. These molecules have specialized "supportive" TcRdeltaV domains membrane-proximal to domains with most similarity to IgNAR V. The ancestral NAR V gene encoding this domain is hypothesized to have recombined with the TRD locus in a cartilaginous fish ancestor >200 million years ago and encodes the first V domain shown to be used in both Igs and TcRs. Furthermore, these data support the view that gamma/delta TcRs have for long used structural conformations recognizing free antigen.

Amino Acid Sequence↗

Nucleotide sequence of EV1, a British isolate of maedi-visna virus.

We have isolated a maedi-visna-like virus from the peripheral blood mononuclear cells of a British sheep displaying symptoms of arthritis and pneumonia. After brief passage in fibroblasts this virus (designated EV1) was used to infect choroid plexus cells. cDNA clones of the virus were prepared from these cells and sequenced. Gaps between non-overlapping clones were filled using gene amplification by the polymerase chain reaction. The genome structure is similar to that described for visna virus strain 1514, and differs from that described for visna virus strain SA-OMVV in not having a W reading frame. Overall the genome differs by about 20% between each of these strains, but there is fivefold variation in the amount of divergence of derived amino acid sequences of different open reading frames. Two sequenced EV1 clones each contain only one copy of the 43 bp repeat, with paired AP-1 sites, which is a feature of other ruminant lentiviral long terminal repeats (LTRs). However, analysis of viral DNA in infected cells by gene amplification shows that LTRs with two repeats do occur, albeit at a relatively low frequency.

Amino Acid Sequence↗

Haplotype block partitioning and tag SNP selection using genotype data and their applications to association studies.

Recent studies have revealed that linkage disequilibrium (LD) patterns vary across the human genome with some regions of high LD interspersed by regions of low LD. A small fraction of SNPs (tag SNPs) is sufficient to capture most of the haplotype structure of the human genome. In this paper, we develop a method to partition haplotypes into blocks and to identify tag SNPs based on genotype data by combining a dynamic programming algorithm for haplotype block partitioning and tag SNP selection based on haplotype data with a variation of the expectation maximization (EM) algorithm for haplotype inference. We assess the effects of using either haplotype or genotype data in haplotype block identification and tag SNP selection as a function of several factors, including sample size, density or number of SNPs studied, allele frequencies, fraction of missing data, and genotyping error rate, using extensive simulations. We find that a modest number of haplotype or genotype samples will result in consistent block partitions and tag SNP selection. The power of association studies based on tag SNPs using genotype data is similar to that using haplotype data.

Algorithms↗

Analysis of involvement of the RecF pathway in p44 recombination in Anaplasma phagocytophilum and in Escherichia coli by using a plasmid carrying the p44 expression and p44 donor loci.

Anaplasma phagocytophilum, the etiologic agent of human granulocytic anaplasmosis, has a large paralog cluster (approximate 90 members) that encodes the 44-kDa major outer membrane proteins (P44s). Gene conversion at a single p44 expression locus leads to P44 antigenic variation. Homologs of genes for the RecA-dependent RecF pathway, but not the RecBCD or RecE pathways, of recombination were detected in the A. phagocytophilum genome. In the present study, we examined whether the RecF pathway is involved in p44 gene conversion. The recombination intermediate structure between a donor p44 and the p44 expression locus of A. phagocytophilum was detected in an HL-60 cell culture by Southern blot analysis followed by sequencing the band and in blood samples from infected SCID mice by PCR, followed by sequencing. The sequences were consistent with the RecF pathway recombination: a half-crossover structure, consisting of the donor p44 locus connected to the 3' conserved region of the recipient p44 in the p44 expression locus in direct orientation. To determine whether the p44 recombination intermediate structure can be generated in a RecF-active Escherichia coli strain, we constructed a double-origin plasmid carrying the p44 expression locus and a donor p44 locus and introduced the plasmid into various E. coli strains. The recombination intermediate was recovered in an E. coli strain with active RecF recombination pathway but not in strains with deficient RecF pathway. Our results support the view that the p44 gene conversion in A. phagocytophilum occurs through the RecF pathway.

Anaplasma phagocytophilum↗

Molecular-genetic evidence for the relationship of Mycobacterium leprae to slow-growing pathogenic mycobacteria.

A total of 1170 nucleotides of the 16S rRNA from Mycobacterium leprae were compared to the homologous regions of M. tuberculosis, M. bovis Vallée, M. avium, M. scrofulaceum, M. phlei, M. fortuitum and one representative each of the genera Corynebacterium, Nocardia, and Rhodococcus. Homology values were calculated and a phylogenetic tree was constructed from the evolutionary distance values. Despite differences in DNA G + C content and genome size, M. leprae is a true member of the slow-growing pathogenic mycobacteria, branching off intermediate to the other members of this subgroup. Slow- and fast-growing mycobacteria are phylogenetically well separated but constitute an individual branch of the actinomycetes proper. Significant structural variation of certain regions of the 16S rRNA may allow construction of M. leprae-specific probes used for rapid identification.

Animals↗