Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

Characterisation of Haemophilus influenzae proteins by two-dimensional gel electrophoresis.

The proteins of nontypable and type b Haemophilus influenzae isolates were characterised using two-dimensional polyacrylamide gel electrophoresis (2-D PAGE). Coomassie Brilliant. Blue R-250 was used for protein detection. Two hundred and twenty eight proteins were resolved from whole cell lysates prepared from a standard nontypable H. influenzae strain (designated HI-64443) when isoelectric focusing was used for the first-dimensional separation of 2-D PAGE. When nonequilibrium pH gel electrophoresis (NEPHGE) was used to separate basic proteins in the first dimension, 50 proteins were detected for HI-64443; 20 of the basic proteins detected were considered to be unique for this separation protocol. The apparent molecular weights and isoelectric points were determined for 82 of the proteins resolved for HI-64443. The variation of the proteins from the standard bacterial strain (HI-64443) was determined for nontypable H. influenzae isolates. On the basis of their electrophoretic mobilities, 17.5% of the proteins of HI-64443 were shared by four other nontypable H. influenzae strains analysed. These data identified both conserved and variable proteins among the nontypable H. influenzae isolates analysed. The results obtained indicated that 2-D PAGE was able to discriminate nontypable H. influenzae into population clones identified by other procedures. The 2-D protein profiles obtained for type b H. influenzae strains were similar to those obtained for nontypable H. influenzae strains. The extent of the protein variation observed between type b and nontypable H. influenzae strain was similar to that observed among nontypable strains alone. These data are discussed in relation to the application of 2-D PAGE as a tool for studies on bacterial epidemiology and for the analysis of the genome structure and gene expression of Haemophilus influenzae.

Bacterial Proteins↗

Genetic heterogeneity within the coding regions of E2 and NS3 in strains of bovine viral diarrhea virus.

We have amplified and sequenced parts of the genomes of eleven laboratory strains of bovine viral diarrhea (BVD) virus originating from North America, New Zealand and Europe. The cumulative nucleotide (nt) sequence heterogeneity of the amplified fragments located in the analysed region of the gene encoding the nonstructural protein NS3 (P80) was 24% as compared to 47% for E2 (Gp53). The nt substitutions in the E2 region resulted in replacements in 42% of amino acid (aa) positions, while the deduced aa sequence of all BVD virus strains remained identical in NS3 and differed from the corresponding region of classical swine fever viruses. This makes possible the differentiation of bovine and porcine pestiviruses. It is suggested that genetic heterogeneity results from passage in transiently infected animals.

Amino Acid Sequence↗

Detecting false expression signals in high-density oligonucleotide arrays by an in silico approach.

High-density oligonucleotide arrays have become a popular assay for concurrent measurement of mRNA expression at the genome scale. Much effort has been devoted to the development of statistical analysis tools aimed at reducing experimental noise and normalizing experimental variation in gene expression analysis. However, these investigations do not detect or catalog systematic problems associated with specific oligonucleotide probes. Here, we present an investigation of problematic probes that yield consistent but inaccurate signals across multiple experiments. By evaluating data integrity among gene, probe sequence, and genomic structure we identified a total of 20,696 (10.5%) nonspecific probes that could cross-hybridize to multiple genes and a total of 18,363 (9.3%) probes that miss the target transcript sequences on the Affymetrix GeneChip U95A/Av2 array. The numbers of nonspecific and mistargeted probes on the U133A array are 29,405 (12.1%) and 19,717 (8.0%), respectively. The poor performance of the mistargeted probes was confirmed in two GeneChip experiments, in which these probes showed a 20-30% decrease in detecting present signals compared with normal probes. Comparison of qualitative expression signals obtained from SAGE and EST data with those from GeneChip arrays showed that the consistency of the two platforms is 30% lower in problematic probes than in normal probes. A Web application was developed to apply our results for improving the accuracy of expression analysis.

Expressed Sequence Tags↗

Molecular comparison of dengue type 1 Mochizuki strain virus and other selected viruses concerning nucleotide and amino acid sequences of genomic RNA: a consideration of viral epidemiology and variation.

Dengue-1 (D1) Mochizuki strain was examined for its nucleotide and amino acid sequences of genomic RNA and the data obtained were compared with those of other selected virus strains reported previously. Genomic regions corresponding to C, preM and M proteins were the major subjects of study. Parts of E protein were additionally examined. Among the D1 viruses investigated, the Mochizuki virus which was isolated in 1943 in Japan was shown to be close to Philippine 836-1 strain isolated in 1984 and Nauru Island strain isolated in 1974 at the respective places, in contrast with Thai AHF 82-80 strain isolated in 1980 and Caribbean CV1636/77 strain isolated in 1977. At the same time, a difference was noted between the Mochizuki and Philippine/Nauru strains at the cleavage site of preM/M junction: Mochizuki possessed RRGKR/S sequence whereas the Philippine/Nauru had RRDKR/S. The glycosylation site in preM and hydrophobic regions at the carboxyl termini of M and E were well conserved. Significances of the data are discussed in connection with viral epidemiology and variation.

Amino Acid Sequence↗

Logos: a modular bayesian model for de novo motif detection.

The complexity of the global organization and internal structure of motifs in higher eukaryotic organisms raises significant challenges for motif detection techniques. To achieve successful de novo motif detection, it is necessary to model the complex dependencies within and among motifs and to incorporate biological prior knowledge. In this paper, we present LOGOS, an integrated LOcal and GlObal motif Sequence model for biopolymer sequences, which provides a principled framework for developing, modularizing, extending and computing expressive motif models for complex biopolymer sequence analysis. LOGOS consists of two interacting submodels: HMDM, a local alignment model capturing biological prior knowledge and positional dependency within the motif local structure; and HMM, a global motif distribution model modeling frequencies and dependencies of motif occurrences. Model parameters can be fit using training motifs within an empirical Bayesian framework. A variational EM algorithm is developed for de novo motif detection. LOGOS improves over existing models that ignore biological priors and dependencies in motif structures and motif occurrences, and demonstrates superior performance on both semi-realistic test data and cis-regulatory sequences from yeast and Drosophila genomes with regard to sensitivity, specificity, flexibility and extensibility.

Algorithms↗

Mitochondrial DNA variation in bull trout (Salvelinus confluentus) from northwestern North America: implications for zoogeography and conservation.

Bull trout, Salvelinus confluentus (Salmonidae), are distributed in northwestern North America from Nevada to Yukon Territory, largely in interior drainages. The species is of conservation concern owing to declines in abundance, particularly in southern portions of its range. To investigate phylogenetic structure within bull trout that might form the basis for the delineation of major conservation units, we conducted a mitochondrial DNA (mtDNA) survey in bull trout from throughout its range. Restriction fragment length polymorphism (RFLP) analysis of four segments of the mtDNA genome with 11 restriction enzymes resolved 21 composite haplotypes that differed by an average of 0.5% in sequence. One group of haplotypes predominated in 'coastal' areas (west of the coastal mountain ranges) while another predominated in 'interior' regions (east of the coastal mountains). The two putative lineages differed by 0.8% in sequence and were also resolved by sequencing a portion of the ND1 gene in a representative of each RFLP haplotype. Significant variation existed within individual sample sites (12% of total variation) and among sites within major geographical regions (33%), but most variation (55%) was associated with differences between coastal and interior regions. We concluded that: (i) bull trout are subdivided into coastal and interior lineages; (ii) this subdivision reflects recent historical isolation in two refugia south of the Cordilleran ice sheet during the Pleistocene: the Chehalis and Columbia refugia; and (iii) most of the molecular variation resides at the interpopulation and inter-region levels. Conservation efforts, therefore, should focus on maintaining as many populations as possible across as many geographical regions as possible within both coastal and interior lineages.

Animals↗

The complete genomic sequence of 424,015 bp at the centromeric end of the HLA class I region: gene content and polymorphism.

We report here the genomic sequence of the centromeric portion of HLA class I, extending 424,015 bp from tumor necrosis factor alpha to a newly identified gene approximately 20 kb telomeric of Otf-3. As a source of DNA, we used cosmids centromeric of HLA-B that had been mapped previously with conventional restriction digestion and fingerprinting and previously characterized yeast artificial chromosomes subcloned into cosmids and mapped with multiple complete digest methodologies. The data presented provide a description of the gene content of centromeric HLA class I including new data on intron, promoter and flanking sequences of previously described genes, and a description of putative new genes that remain to be characterized beyond the structural information uncovered. A complete accounting of the repeat structure including abundant di-, tri-, and tetranucleotide microsatellite loci yielded access to precisely localized mapping tools for the major histocompatibility complex. Comparative analysis of a highly polymorphic region between HLA-B and -C was carried out by sequencing over 40 kb of overlapping sequence from two haplotypes. The levels of variation observed were much higher than those seen in other regions of the genome and indeed were higher than those observed between allelic HLA class I loci.

Centromere↗

The diversity of LTR retrotransposons.

Eukaryotic genomes are full of long terminal repeat (LTR) retrotransposons. Although most LTR retrotransposons have common structural features and encode similar genes, there is nonetheless considerable diversity in their genomic organization, reflecting the different strategies they use to proliferate within the genomes of their hosts.

Animals↗

Evaluation of alpha hemoglobin stabilizing protein (AHSP) as a genetic modifier in patients with beta thalassemia.

Although beta thalassemia is considered to be a classic monogenic disease, it is clear that there is considerable clinical variability between patients who inherit identical beta globin gene mutations, suggesting that there may be a variety of genetic determinants influencing different clinical phenotypes. It has been suggested that variations in the structure or amounts of a highly expressed red cell protein (alpha hemoglobin stabilizing protein [AHSP]), which can stabilize free alpha globin chains in vitro, could influence disease severity in patients with beta thalassemia. To address this hypothesis, we studied 120 patients with Hb E-beta thalassemia with mild, moderate, or severe clinical phenotypes. Using gene mapping, direct genomic sequencing, and extended haplotype analysis, we found no mutation or specific association between haplotypes of AHSP and disease severity in these patients, suggesting that AHSP is not a disease modifier in Hb E-beta thalassemia. It remains to be seen if any association between AHSP and clinical severity is present in other population groups with a high frequency of beta thalassemia.

Adolescent↗

Intraspecific variation in symbiont genomes: bottlenecks and the aphid-buchnera association.

Buchnera are maternally transmitted bacterial endosymbionts that synthesize amino acids that are limiting in the diet of their aphid hosts. Previous studies demonstrated accelerated sequence evolution in Buchnera compared to free-living bacteria, especially for nonsynonymous substitutions. Two mechanisms may explain this acceleration: relaxed purifying selection and increased fixation of slightly deleterious alleles under drift. Here, we test the divergent predictions of these hypotheses for intraspecific polymorphism using Buchnera associated with natural populations of the ragweed aphid, Uroleucon ambrosiae. Contrary to expectations under relaxed selection, U. ambrosiae from across the United States yielded strikingly low sequence diversity at three Buchnera loci (dnaN, trpBC, trpEG), revealing polymorphism three orders of magnitude lower than in enteric bacteria. An excess of nonsynonymous polymorphism and of rare alleles was also observed. Local sampling of additional dnaN sequences revealed similar patterns of polymorphism and no evidence of food plant-associated genetic structure. Aphid mitochondrial sequences further suggested that host bottlenecks and large-scale dispersal may contribute to genetic homogenization of aphids and symbionts. Together, our results support reduced N(e) as a primary cause of accelerated sequence evolution in Buchnera. However, our study cannot rule out the possibility that mechanisms other than bottlenecks also contribute to reduced N(e) at aphid and endosymbiont loci.

Alleles↗

Allelic variation and light-responsive regulation of FaMYB10-2 underlie tissue-specific anthocyanin accumulation in strawberry.

Anthocyanins critically determine fruit color, nutrition, and stress resilience in cultivated strawberry (Fragaria × ananassa), directly influencing consumer preference. Despite complex genetic and environmental regulation of their biosynthesis, the basis for tissue-specific pigmentation, notably the widespread occurrence of red skin and pale flesh, remains poorly understood. We integrated genomic, transcriptomic, and functional analyses across 200 cultivars to dissect receptacle pigmentation regulation. Approaches included FaMYB10-2 allele mining, promoter structural variant (SV) identification, expression profiling, regulatory interaction assays, and characterization of upstream light-responsive factors. FaMYB10-2 was identified as the key R2R3-MYB regulator of fruit anthocyanin biosynthesis. Alleles FaMYB10-2.2 and FaMYB10-2.3 encode truncated proteins retaining bHLH-binding capacity but lacking activation domains, functioning as dominant-negative repressors. A promoter SV 986 bp upstream of FaMYB10-2 was associated with reduced pale fruit due to cis-regulatory divergence. The SV (Alt) allele is prevalent in Asian cultivars, while the Ref allele is enriched in Western germplasm. Crucially, a light-responsive FaHYH-FaWRKY71 cascade activates FaMYB10-2 and structural genes haplotype-dependently, compensating for weak MYB activity in the skin. Our findings reveal a multilayered regulatory system integrating allelic variation, cis-regulatory divergence, and environmental signals, advancing anthocyanin understanding and providing engineering targets for polyploid crop color improvement.

Fragaria↗

Evolutionary changes in the expression pattern of a developmentally essential gene in three Drosophila species.

The hypothesis that morphological evolution may largely result from changes in gene regulation rather than gene structure has been difficult to test. Morphological differences among insects are often apparent in the cuticle structures produced. The dopa decarboxylase (Ddc) and alpha-methyldopa hypersensitive (amd) genes arose from an ancient gene duplication. In Drosophila, they have evolved nonoverlapping functions, including the production of distinct types of cuticle, and for Ddc, the production of the neurotransmitters, dopamine and serotonin. The amd gene is particularly active in the production of specialized flexible cuticles in the developing embryo. We have compared the pattern of amd expression in three Drosophila species. Several regions of expression conserved in all three species but, surprisingly, a unique domain of expression is found in Drosophila simulans that does occur in the closely related (2-5 million years) Drosophila melanogaster or in the more remote (40-50 million years) Drosophila virilis. The "sudden" appearance of a completely new and robust domain of expression provides a glimpse of evolutionary variation resulting from changes in regulation of structural gene expression.

Animals↗

Genetic locus for the biosynthesis of the variable portion of Neisseria gonorrhoeae lipooligosaccharide.

A locus involved in the biosynthesis of gonococcal lipooligosaccharide (LOS) has been cloned from gonococcal strain F62. The locus contains five open reading frames. The first and second reading frames are homologous, but not identical, to the fourth and fifth reading frames, respectively. Interposed is an additional reading frame which has distant homology to the Escherichia coli rfaI and rfaI genes, both glucosyl transferases involved in lipopolysaccharide core biosynthesis. The second and fifth reading frames show strong homology to the lex-1 or lic2A gene of Haemophilus influenzae, but do not contain the CAAT repeats found in this gene. Deletions of each of these five genes, of combinations of genes, and of the entire locus were constructed and introduced into parental gonococcal strain F62 by transformation. The LOS phenotypes were then analyzed by SDS-PAGE and reactivity with monoclonal antibodies. Analysis of the gonococcal mutants indicates that four of these genes are the glycosyl transferases that add GalNAc beta 1-->3Gal beta 1-->4GlcNAc beta 1-->3 Gal beta 1--4 to the substrate Glc beta 1-->4Hep--R of the inner core region. The gene with homology to E. coli rfaI/rfaI is involved with the addition of the alpha-linked galactose residue in the biosynthesis of the alternative LOS structure Gal alpha 1-->4Gal beta 1-->4Glc beta 1-->4Hep-->R. Since these genes encode LOS glycosyl transferases they have been named lgtA, lgtB, lgtC, lgtD, and lgtE. The DNA sequence analysis revealed that lgtA, lgtC, and lgtD contained poly-G tracts, which, in strain F62 were, respectively, 17, 10, and 11 bp. Thus, three of the LOS biosynthetic enzymes are potentially susceptible to premature termination by reading frame changes. It is likely that these structural features are responsible for the high-frequency genetic variation of gonococcal LOS.

Amino Acid Sequence↗

Intragenic Hill-Robertson interference influences selection intensity on synonymous mutations in Drosophila.

Natural selection influences synonymous mutations and synonymous codon usage in many eukaryotes to improve the efficiency of translation in highly expressed genes. Recent studies of gene composition in eukaryotes have shown that codon usage also varies independently of expression levels, both among genes and at the intragenic level. Here, we investigate rates of evolution (Ks) and intensity of selection (gamma(s)) on synonymous mutations in two groups of genes that differ greatly in the length of their exons, but with equivalent levels of gene expression and rates of crossing-over in Drosophila melanogaster. We estimate gamma(s) using patterns of divergence and polymorphism in 50 Drosophila genes (100 kb of coding sequence) to take into account possible variation in mutation trends across the genome, among genes or among codons. We show that genes with long exons exhibit higher Ks and reduced gamma(s) compared to genes with short exons. We also show that Ks and gamma(s) vary significantly across long exons, with higher Ks and reduced gamma(s) in the central region compared to flanking regions of the same exons, hence indicating that the difference between genes with short and long exons can be mostly attributed to the central region of these long exons. Although amino acid composition can also play a significant role when estimating Ks and gamma(s), our analyses show that the differences in Ks and gamma(s) between genes with short and long exons and across long exons cannot be explained by differences in protein composition. All these results are consistent with the Interference Selection (IS) model that proposes that the Hill-Robertson (HR) effect caused by many weakly selected mutations has detectable evolutionary consequences at the intragenic level in genomes with recombination. Under the IS model, exon size and exon-intron structure influence the effectiveness of selection, with long exons showing reduced effectiveness of selection when compared to small exons and the central region of long exons showing reduced intensity of selection compared to flanking coding regions. Finally, our results further stress the need to consider selection on synonymous mutations and its variation--among and across genes and exons--in studies of protein evolution.

Animals↗

Population genetics of geographically restricted and widespread species of Myrica (Myricaceae).

Allozyme variation of 11 putative loci in five populations of the rare Myrica adenophora Hance, and four populations of its widespread congeneric species, M. rubra (Lour.) Sieb. & Zucc. was studied. Among the 21 alleles studied, no unique allele was detected for M. adenophora, whereas M. rubra had 3 alleles not found in the former species. In terms of genetic diversity, populations of the rare species contained fewer alleles per locus (1.5 versus 1.7), fewer effective number of alleles per locus (1.12 versus 1.20), fewer number of alleles per polymorphic locus (2.14 versus 2.46), lower percentage of polymorphic loci (30.9 versus 40.9), and lower expected heterozygosity (0.106 versus 0.163) than populations of the widespread species. Genetic distances within species average 0.043 for M. adenophora and 0.045 for M. rubra, and between species ranged from 0.052 to 0.177, with a mean of 0.103, which agrees with the very similar gross morphologies of these two species. Intrapopulation differentiation was similar in both species: G(ST) = 0.152 for M. adenophora, and 0.146 for M. rubra, whereas estimated gene flow based on G(ST) values were moderate in these two species (Nm = 1.39 versus 1.46). We inferred that M. rubra and M. adenophora are a progenitor-derivative species pair that emerged before migrating into Taiwan during the last glacial period. We consider the Hengchun population (Chiupeng, Hsuhai, and Chufengpi) and Taitung population (Tienkuan and Lanshan) of M. adenophora which probably arose from two subsets of the genome of M. rubra. Genetic drift was inferred to be one of the forces shaping the observed genetic structure in M. adenophora and M. rubra.

Alleles↗

Genomics and variation of ionotropic glutamate receptors: implications for neuroplasticity.

We used two approaches to identify sequence variants in ionotropic glutamate receptor (IGR) genes: high-throughput screening and resequencing techniques, and "information mining" of public (e.g. dbSNP, ENSEMBL) and private (i.e. Celera Discovery System) sequence databases. Each of the 16 known IGRs is represented in these databases, their positions on a canonical physical map are established. Comparisons of mouse, rat, and human sequences revealed substantial conservation among these genes, which are located on different chromosomes but found within syntenic groups of genes. The IGRs are members of a phylogenetically ancient gene family, sharing similarities with glutamate-like receptors in plants. Parsimony analysis of amino acid sequences groups the IGRs into three distinct clades based on ligand-binding specificity and structural features, such as the channel pore and membrane spanning domains. A collection of 38 variants with amino acid changes was obtained by combining screening, resequencing, and informatics approaches for several of the IGR genes. This represents only a fraction of the sequence variation across these genes, but in fact these may constitute a large fraction of the common polymorphisms at these genes and these polymorphisms are a starting point for understanding the role of these variants in function. Genetically influenced human neurobehavioral phenotypes are likely to be linked to IGR genetic variants. Because ionotropic glutamate receptor activation leads to calcium entry, which is fundamental in brain development and in forms of synaptic plasticity essential for learning and memory and is essential for neuronal survival, it is likely that sequence variants in IGR genes may have profound functional roles in neuronal activation and survival mechanisms.

Amino Acid Substitution↗

Rare pathogenic NR2F2 (COUP-TFII) variants as potential etiological causes in pediatric patients with congenital heart diseases (CHDs).

OBJECTIVES: Congenital heart diseases (CHDs) are complex genetic disorders, and their genetic basis is not yet fully understood. Nuclear receptor subfamily 2 group F member 2 (NR2F2 or COUP-TFII) encodes a transcription factor which is expressed at high levels during mammalian development. Few studies have identified heterozygous and rare variants in the NR2F2 gene in individuals with CHD. This study aimed to evaluate the association between pathogenic genetic alterations in NR2F2 with CHD risk. METHODS: A case-control study was conducted on a group of 135 patients (83 boys and 52 girls) with various types of non-hereditary, isolated CHD who were undergoing open-heart surgery. Additionally, 95 matched healthy children without syndromic or isolated heart abnormalities were selected. RESULTS: Using Sanger sequencing, we identified 5 heterozygous single nucleotide variants in exons 2 and 3 of the NR2F2 gene. These variations were novel and not present in any genomic variation databases. Four of the variations were missense mutations (p.Pro159Arg, p.Ser329Phe, p.Qln338Pro, and p.Tyr348Ser) and one was a synonymous variant (p.G361 = ) in the coding region. Importantly, in silico results indicated that the missense variants had pathogenic effects on protein function. Additionally, the missense variants substantially altered the predicted structure of COUP-TFII. CONCLUSION: The results we obtained not only validate the correlation between NR2F2 mutations and CHDs but also have significant potential for guiding new preventive and therapeutic strategies. This could contribute to the advancement of medical interventions in the fields of cardiology and genetics.

Humans↗

Integrative genomics elucidates the evolutionary, temporal, and developmental origins of a hydrocephalus risk gene.

INTRODUCTION: A prior integrative, multi-omics human genetics and functional genomics study identified maelstrom (MAEL), a gene involved in regulation of DNA transposon activity and genome structure, as a transcriptome-wide predictor of hydrocephalus (HC) in the brain cortex. Here we expand on this discovery and further characterize the evolutionary origin and expression of MAEL across developmental timescales and cell-lineages in the neonatal human brain towards a mechanistic understanding how variation in MAEL expression may cause HC. OBJECTIVE: To characterize the evolutionary, temporal, developmental, and lineages of MAEL expression in HC and the developing human brain. METHODS: Ensembl was used to delineate the evolution and taxonomy of MAEL across species. Analysis of single-cell RNA sequencing (scRNA-seq) of 49 brain regions across pre- and post-natal timescales from the Developing Human Brain Atlas (Allen Institute) identified temporal and spatial MAEL expression patterns. We quantified MAEL expression in primary cortical brain tissue obtained during the surgical treatment of HC. RESULTS: We performed taxonomic gene-mapping to define the evolutionary origin of MAEL to assess suitability for mechanistic characterization in vitro and in vivo across species. We find that MAEL is among the top 0.01% human-specific genes and < 50% sequence homology among commonly used model organisms with highly divergent functions, necessitating mechanistic validation in human tissue. scRNA-seq of the non-disease prenatal human brain identified MAEL expression enriched in cortical excitatory neurons, which was recapitulated in primary HC brain tissue obtained during surgery. Finally, using scRNA-seq of primary HC brain tissue, we functionally validated reduced MAEL expression, consistent with a prior human TWAS analysis. CONCLUSIONS: We identify the evolutionary, temporal, and developmental expression pattern of MAEL in the neonatal human brain. We also provide direct evidence for reduced MAEL expression in human HC brain tissue. These data, at least in part, implicate reduced MAEL expression underlying human HC across etiologies.

Journal Article↗