Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Structural genome variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Impact of natural variation in bacterial F17G adhesins on crystallization behaviour.

Since the introduction of structural genomics, the protein has been recognized as the most important variable in crystallization. Recent strategies to modify a protein to improve crystal quality have included rationally engineered point mutations, truncations, deletions and fusions. Five naturally occurring variants, differing in 1-18 amino acids, of the 177-residue lectin domain of the F17G fimbrial adhesin were expressed and purified in identical ways. For four out of the five variants crystals were obtained, mostly in non-isomorphous space groups, with diffraction limits ranging between 2.4 and 1.1 A resolution. A comparative analysis of the crystal-packing contacts revealed that the variable amino acids are often involved in lattice contacts and a single amino-acid substitution can suffice to radically change crystal packing. A statistical approach proved reliable to estimate the compatibilities of the variant sequences with the observed crystal forms. In conclusion, natural variation, universally present within prokaryotic species, is a valuable genetic resource that can be favourably employed to enhance the crystallization success rate with considerably less effort than other strategies.

Adhesins, Bacterial↗

Analyses of variable antigen gene rearrangements in Trypanosoma brucei.

For the purpose of investigating the genetic basis of antigenic variation in Trypanosoma brucei, we have analyzed the structure of the genome surrounding the gene coding for one T. brucei variable antigen (ILTat 1.2) in several T. brucei clones expressing this and other variable antigens. In each case there are two copies of the gene. We found no evidence of an extra copy associated with the expression of this gene. Differences were found between the two copies in a single clone, and between the copies in different clones. The differences could be explained by insertion and deletion of various lengths of DNA in a region beyond the C-terminal end of the gene. Differences in genomic structure were found even between clones expressing the same antigen, whether ILTat 1.2 or another. Thus, no feature of the rearrangements observed can be correlated with the expression of a particular antigen.

Animals↗

Conservation of a gene conversion mechanism in two distantly related paralogues of Anaplasma marginale.

Anaplasmataceae, the causative agents of anaplasmosis and ehrlichiosis, persist in the bloodstream of their mammalian hosts, allowing acquisition and transmission by tick vectors. Anaplasma marginale establishes persistent infection characterized by sequential cycles of rickettsaemia in which new antigenic variants emerge. The two most immunodominant outer membrane proteins, MSP2 and MSP3, are paralogues, each encoded by a distinct family of related genes. This study demonstrates that, although the two gene families have diverged substantially, each has maintained a similar mechanism to generate structurally and antigenically polymorphic surface antigens. Like MSP2, MSP3 is expressed from a single locus in which variation of the expressed msp3 gene is generated by recombination using msp3 pseudogenes. Each of the msp3 pseudogenes encodes a unique central variable region (CVR) flanked by conserved 5' and 3' regions. Changes in the CVR of the expressed msp3, concomitant with invariance of the pseudogenes, indicate that expression site variation is generated using gene conversion. A. marginale thus maintains two large, separate systems within its small genome to generate antigenic variation of its surface proteins, while analogous structural elements indicate a common mechanism.

Acute Disease↗

The wheat cytochrome oxidase subunit II gene has an intron insert and three radical amino acid changes relative to maize.

We have determined the sequence of the wheat mitochondrial gene for cytochrome oxidase subunit II (COII) and find that its derived protein sequence differs from that of maize at only three amino acid positions. Unexpectedly, all three replacements are non-conservative ones. The wheat COII gene has a highly-conserved intron at the same position as in maize, but the wheat intron is 1.5 times longer because of an insert relative to its maize counterpart. Hybridization analysis of mitochondrial DNA from rye, pea, broad bean and cucumber indicates strong sequence conservation of COII coding sequences among all these higher plants. However, only rye and maize mitochondrial DNA show homology with wheat COII intron sequences and rye alone with intron-insert sequences. We find that a sequence identical to the region of the 5' exon corresponding to the transmembrane domain of the COII protein is present at a second genomic location in wheat mitochondria. These variations in COII gene structure and size, as well as the presence of repeated COII sequences, illustrate at the DNA sequence level, factors which contribute to higher plant mitochondrial DNA diversity and complexity.

Journal Article↗

Spatial patterns of mitochondrial and nuclear gene pools in chamois (Rupicapra r. rupicapra) from the Eastern Alps.

We have assessed the variability of maternally (mtDNA) and biparentally (allozymes) inherited genes of 443 chamois (Rupicapra r. rupicapra) from 19 regional samples in the Eastern Alps, to estimate the degree and patterns of spatial gene pool differentiation, and their possible causes. Based on a total mtDNA-RFLP approach with 16 hexanucleotide-recognizing restriction endonucleases, we found marked substructuring of the maternal gene pool into four phylogeographic groups. A hierarchical AMOVA revealed that 67.09% of the variance was partitioned among these four mtDNA-phylogroups, whereas only 8.04% were because of partitioning among regional samples within the populations, and 24.86% due to partitioning among individuals within regional samples. We interpreted this spatial pattern of mtDNA variability as a result of immigration of chamois from different Pleistocene refugia surrounding the Alps after the withdrawal of glaciers, rather than from topographic barriers to gene flow, such as Alpine valleys, extended glaciers or woodlands. However, this striking geographical structuring of the maternal genome was not paralleled by allelic variation at 33 allozyme loci, which were used as nuclear DNA markers. Wright's hierarchical F-statistics revealed that only < or =0.45% of the explained allozymic diversity was because of partitioning among the four mtDNA-phylogroups. We conclude that this discordance of spatial patterns of nuclear and mtDNA gene pools results from a phylogeographic background and sex-specific dispersal, with higher levels of philopatry in females.

Animals↗

The E protein is a multifunctional membrane protein of SARS-CoV.

The E (envelope) protein is the smallest structural protein in all coronaviruses and is the only viral structural protein in which no variation has been detected. We conducted genome sequencing and phylogenetic analyses of SARS-CoV. Based on genome sequencing, we predicted the E protein is a transmembrane (TM) protein characterized by a TM region with strong hydrophobicity and alpha-helix conformation. We identified a segment (NH2-_L-Cys-A-Y-Cys-Cys-N_-COOH) in the carboxyl-terminal region of the E protein that appears to form three disulfide bonds with another segment of corresponding cysteines in the carboxyl-terminus of the S (spike) protein. These bonds point to a possible structural association between the E and S proteins. Our phylogenetic analyses of the E protein sequences in all published coronaviruses place SARS-CoV in an independent group in Coronaviridae and suggest a non-human animal origin.

Amino Acid Sequence↗

Length variation in mitochondrial DNA of the minnow Cyprinella spiloptera.

Length differences in animal mitochondrial DNA (mtDNA) are common, frequently due to variation in copy number of direct tandem duplications. While such duplications appear to form without great difficulty in some taxonomic groups, they appear to be relatively short-lived, as typical duplication products are geographically restricted within species and infrequently shared among species. To better understand such length variation, we have studied a tandem and direct duplication of approximately 260 bp in the control region of the cyprinid fish, Cyprinella spiloptera. Restriction site analysis of 38 individuals was used to characterize population structure and the distribution of variation in repeat copy number. This revealed two length variants, including individuals with two or three copies of the repeat, and little geographic structure among populations. No standard length (single copy) genomes were found and heteroplasmy, a common feature of length variation in other taxa, was absent. Nucleotide sequence of tandem duplications and flanking regions localized duplication junctions in the phenylalanine tRNA and near the origin of replication. The locations of these junctions and the stability of folded repeat copies support the hypothesized importance of secondary structures in models of duplication formation.

Animals↗

The role of geographic analysis in locating, understanding, and using plant genetic diversity.

The genetic structure of an organism is shaped by various factors, many of which vary significantly over space. In this chapter, we provide insight on how studying geographic patterns may contribute to an improved understanding of variability in genetic structure. We first review the theoretical background on how differences in genetic structure may be generated through processes that are inherently variable over space. We then present novices with some basics on how geographic information systems (GIS) may be adopted to study this variation, including advice on software, data, and the type of research questions that might be addressed. The chapter finishes with a brief review of how spatial analysis has contributed to the conservation and use of plant genetic resources, through an understanding of spatial patterns in species distribution and genetic structure. We conclude that spatial variation is a factor often overlooked in genetic studies and one that merits greater consideration. With the advent of functional genomics and improved quantification of adaptive traits, spatial analysis may be key in understanding variation in genetic structure through careful analysis of genotype-environment interactions.

Algorithms↗

Single nucleotide polymorphisms (SNPs) in human lactoferrin gene.

The lactoferrin protein possesses antimicrobial and antiviral activities. It is also involved in the modulation of the immune response. In a normal healthy individual, lactoferrin plays a role in the front-line host defense against infection and in immune and inflammatory responses. Whether genomic variations, such as single nucleotide polymorphisms (SNPs), have an effect on the structure and function of lactoferrin protein and whether these variations contribute to the different susceptibility of individuals in response to environmental insults are interesting health-related issues. In this study, the lactoferrin gene was resequenced as part of the Environmental Genome Project of the National Institute of Environmental Health Sciences, which operates within the National Institutes of Health. Ninety-one healthy donors of different ethnicities were used to establish common SNPs in the exons of the lactoferrin gene in the general population. The data will serve as a basis from which study the association of lactoferrin polymorphism and disease.

Ethnicity↗

Biochemical aspects of variation in foot-and-mouth disease virus.

The biochemical basis for variation in foot-and-mouth disease virus (FMDV) has been explored by analysis of the virus RNA and the virus-induced and structural proteins of three isolates of the virus. Two of the isolates were from serotype A and the third was from serotype O. Hybridization studies of the RNAs showed greater than 80% homology between the two type A viruses and about 65% homology between the two type A viruses and the virus of type O. The ribonuclease T1 maps of the three viruses gave distinct patterns typical of FMDV, but did not show that any two of the three viruses were more closely related. The virus-induced primary translation products, P88, P52 and P100 isolated from infected cells, were compared by tryptic peptide analysis. Combinations of 3H- and 14C-leucine-labelled polypeptides were hydrolysed with trypsin and resolved on an ion-exchange column. Much greater differences were found in P88 than in P52 or P100, indicating that the major variation occurs in the region of the genome coding for the structural proteins. Similar analysis of combinations of the structural proteins of the three viruses showed that there were differences in VP1, VP2 and VP3 and these results were supported by those obtained by PAGE analysis of the Staphylococcus aureus V8 protease cleavage products.

Aphthovirus↗

Case-control association tests correcting for population stratification.

In case-control association studies unobserved population stratification may act as a confounder, leading to an increased number of false positive results. Methods accounting for population structure by using additional genetic markers broadly follow one of two concepts: Genomic Control (GC) and Structured Association (SA). While extending existing methods of Structured Association we show that it is necessary to incorporate phenotypic information when inferring population structure, otherwise a systematic bias is introduced. Moreover, for moderate population stratification a Wald test statistic should be preferred as a Structured Association test statistic in comparison to a likelihood ratio test. The introduced extensions are compared to existing methods of Structured Association, as well as to Genomic Control, in a simulation study which is based on realistic situations of large case-control studies with moderate population stratification. A disadvantage of Genomic Control turns out to be the large variation in estimating the variance inflation factor, as well as the power loss if population structure increases. We come to the overall conclusion that Structured Association, if applied correctly, is superior to Genomic Control, at least in the case of simple population structure as simulated here.

Case-Control Studies↗

Distinguishable haplotype blocks in the HTR3A and HTR3B region in the Japanese reveal evidence of association of HTR3B with female major depression.

BACKGROUND: Genetic variations in the serotonin receptor 3A (HTR3A) and 3B (HTR3B) genes, positioned in tandem on chromosome 11q23.2, have been shown to be associated with psychiatric disorders in samples of European ancestry. But the polymorphisms highlighted in these reports map to different locations in the two genes, therefore it is unclear which gene exerts a stronger effect on susceptibility. METHODS: To determine the haplotype block structure in the genomic regions of HTR3A and HTR3B, and to examine whether genetic variations in the region show evidence of association with schizophrenia and affective disorder in the Japanese, we performed haplotype-based case-control analysis using 29 polymorphisms. RESULTS: Two haplotype blocks each were revealed for HTR3A and HTR3B in Japanese samples. In HTR3B, haplotype block 2 that included a nonsynonymous single nucleotide polymorphism (SNP), yielded evidence of association with major depression in females (global p = .0023). Analysis employing genome-wide SNPs using the STRUCTURE program did not detect population stratification in the samples. CONCLUSIONS: Our results suggest an important role for HTR3B in major depression in women and also raise the possibility that previously proposed disease-associated SNPs in the HTR3A/B region in Caucasians are in linkage disequilibrium with haplotype block 2 of HTR3B in the Japanese.

Adult↗

A chromosome-scale genome of Capsicum pubescens provides insights into candidate terpene-associated gene clusters and pan variation of terpene synthases.

A chromosome-scale genome of Capsicum pubescens and comparative pan-TPS analysis support structural characterization and gene-level prioritization of a chromosome-9 terpene-associated candidate locus in this accession. Capsicum pubescens is one of the five domesticated Capsicum species, mainly cultivated in mid- to high-elevation regions of the Americas. Despite its distinctive morphology and fruit traits, genomic resources for C. pubescens remain less developed than those for the widely cultivated C. annuum. Here, we assembled a chromosome-scale reference genome for accession HNUCP0001, spanning 3.70&#xa0;Gb with a scaffold N50 of 278.01&#xa0;Mb. Comparative genomics revealed 679 significantly expanded gene families enriched in sesquiterpenoid and triterpenoid biosynthesis. Genome-wide biosynthetic gene-cluster mining identified multiple terpene-associated candidate loci, which were subsequently prioritized using genome-derived structural criteria and Capsicum pubescens-specific expression evidence. Subsequently, we curated the terpene synthase (TPS) repertoire and, across 16 Capsicum genomes, resolved 36 TPS orthogroups with pronounced presence/absence variation, highlighting dynamic lineage-specific diversification. Together, these analyses establish HNUCP0001 as an accession-specific genomic resource and provide a comparative framework for prioritizing terpene-associated TPS genes and candidate BGCs in Capsicum. These candidate loci, together with accession-level transcriptomic and metabolomic evidence, offer testable hypotheses for future functional studies of specialized terpenoid metabolism in C. pubescens.

Alkyl and Aryl Transferases↗

The Indian Genome Variation database (IGVdb): a project overview.

Indian population, comprising of more than a billion people, consists of 4693 communities with several thousands of endogamous groups, 325 functioning languages and 25 scripts. To address the questions related to ethnic diversity, migrations, founder populations, predisposition to complex disorders or pharmacogenomics, one needs to understand the diversity and relatedness at the genetic level in such a diverse population. In this backdrop, six constituent laboratories of the Council of Scientific and Industrial Research (CSIR), with funding from the Government of India, initiated a network program on predictive medicine using repeats and single nucleotide polymorphisms. The Indian Genome Variation (IGV) consortium aims to provide data on validated SNPs and repeats, both novel and reported, along with gene duplications, in over a thousand genes, in 15,000 individuals drawn from Indian subpopulations. These genes have been selected on the basis of their relevance as functional and positional candidates in many common diseases including genes relevant to pharmacogenomics. This is the first large-scale comprehensive study of the structure of the Indian population with wide-reaching implications. A comprehensive platform for Indian Genome Variation (IGV) data management, analysis and creation of IGVdb portal has also been developed. The samples are being collected following ethical guidelines of Indian Council of Medical Research (ICMR) and Department of Biotechnology (DBT), India. This paper reveals the structure of the IGV project highlighting its various aspects like genesis, objectives, strategies for selection of genes, identification of the Indian subpopulations, collection of samples and discovery and validation of genetic markers, data analysis and monitoring as well as the project's data release policy.

Databases, Genetic↗

Integrated multi-omics analyses provide new insights into genomic variation landscape and regulatory network candidate genes associated with walnut endocarp.

Persian walnut (Juglans regia) is an economically important nut oil tree; the fruit has a hard endocarp/shell to protect seeds, thus playing a key role in its evolution, and the shell thickness is an important trait for walnut breeding. However, the genomic landscape and the gene regulatory networks associated with walnut shell development remain to be systematically elucidated. Here, we report a high-quality genome assembly of the walnut cultivar 'Xiangling' and construct a graphic structure pan-genome of eight Juglans species to reveal the genetic variations at the genome level. We re-sequence 285 accessions to characterize the genomic variation landscape. Through genome-wide association studies (GWAS), we identified 19 loci associated with more than 268 loci that underwent selection during walnut domestication and improvement. Multi-omics analyses, including transcriptomics, metabolomics, DNA methylation, and spatial transcriptomics across eleven developmental stages, revealed several candidate genes related to secondary cell biosynthesis and lignin accumulation. This integrated multi-omics approach revealed several candidate genes associated with secondary cell biosynthesis and lignin accumulation, such as UGP, MYB308, MYB83, NAC043, NAC073, CCoAOMT1, CCoAOMT7, CHS2, CESA7, LAC7, COBL4, and IRX12. Overexpression of JrUGP and JrMYB308 in Arabidopsis thaliana confirmed their roles in lignin biosynthesis and cell wall thickening. Consequently, our comprehensive multi-omics findings offer novel insights into walnut genetic variation and network regulation of endocarp development and shell thickness, which enable further genome-informed breeding strategies for walnut cultivar improvement.

Juglans↗

Polygenic and monogenic adaptation drive evolutionary rescue at different magnitudes of environmental change.

Understanding the genetic basis of rapid adaptation is key to predicting species' evolutionary responses to environmental change. However, it is still debatable whether many small-effect mutations or a few large-effect mutations underlie rapid adaptation, and how this knowledge can predict population survival or extinction. To address this question, we performed a series of ecologically grounded forward-in-time genetic simulations to study rapid adaptation and extinction with increasing magnitudes of environmental change. These simulations were seeded with genomic variation of the plant Arabidopsis thaliana to have a realistic genomic structure, with one (monogenic) to 1,000 (polygenic) variants with varying heritabilities contributing to an environmental adaptive trait. Our results revealed two distinct scenarios of rapid adaptation and population rescue. Under small-to-moderate environmental shifts, high polygenic traits increased evolutionary rescue probability. Under extreme environmental shifts, high polygenic traits lead predictably to extinction, yet monogenic traits sometimes produce one-off winning adaptive genotypes. We interpret our rapid evolutionary rescue findings in terms of the fundamental theorem of natural selection, where trait polygenicity shapes the distribution of genetic variance in fitness across replicates and, in turn, the probability of population survival, with polygenic architectures producing more stable and predictable fitness variance and monogenic architectures generating highly skewed and variable outcomes. These results highlight the insights genomics gives us into the (un)predictability of species' evolutionary responses to global change, with management implications for assisted adaptation and conservation.

Arabidopsis↗

Conservation and variation of gene regulation in embryonic stem cells assessed by comparative genomics.

We have examined the gene structure and regulatory regions of octamer-binding transcription factor 3/4 (Oct 3/4), sex determining region Y box 2 (Sox2), signal transducer and activator of transcription 3 (Stat3), embryonal stem cell-specific gene 1 (ESG), Nanog homeobox (Nanog), and several other genes highly expressed in embryonic stem (ES) cells across different species. Our analysis showed that ES cell-expressed Ras (ERAS) was orthologous to a human pseudogene Harvey Ras (HRASP) and that the promoter and other regulatory sequences were highly divergent. No ortholog of (ES) cell-derived homeobox containing gene (Ehox) could be identified in human, and the closest paralogs PEPP gene subfamily 1 (PEPP1), PEPP2, and extraembryonic, spermatogenesis, homeobox 1 (Esx1) were not expressed by ES cells and shared little homology. The Sox2 promoter was the most conserved across species and the Oct3/4 promoter region showed significant homology particularly in the distal enhancer active in ES cells. Analysis suggested common and divergent pathways of regulation. Conserved Oct3/4 and Sox2 co-binding domains were identified in most ES expressed genes, highlighting the importance of this transcriptional pathway. Conserved fibroblast growth factor response element sites were identified in regulatory regions, suggesting a potential parallel pathway for regulation by FGFs. A central role of Stat3 activation in self-renewal and in a regulatory feedback loop was suggested by the identification of the conserved binding sites in most pathways. Although most pathways were evolutionarily conserved, promoters and genomic structure of the leukemia inhibitory factor (LIF) pathway components were divergent, likely explaining the differential requirement of LIF for human and rodent cells. Our analysis further suggested that the Nanog regulatory pathway was relatively independent of the LIF/Oct pathway and may interact with the Nodal/transforming growth factor-beta pathway. These results provide a framework for examining the current reported differences between rodent and human ES cells and define targets for future perturbation studies.

Amino Acid Sequence↗

Heavy chain joining region segments of the channel catfish. Genomic organization and phylogenetic implications.

The JH locus of the channel catfish has been characterized to determine the organization and structural diversity of JH segments. These analyses indicate that there are a total of nine JH segments tightly clustered within a region spanning about 2.2 kb. The JH locus is closely linked to the CH 1 domain of the expressed catfish H chain; the distance between the CH proximal JH segment (JH9) and the CH 1 domain is about 1.8 kb. Each JH segment has an upstream recombination sequence, which includes a T-rich nonamer, a 22- to 24-bp spacer, and a phylogenetically conserved heptamer. Each JH segment also has an open reading frame that encodes the conserved framework region 4 tryptophan (Trp-103) and terminates with a RNA donor splice site. The catfish JH locus contains an internal repetitive sequence region characterized by a short (183-188 bp) repeat that occurs sequentially five times. Strong sequence homology as well as the unified length of the repeated sequences indicate that JH segments JH3-JH7 probably arose as the result of a series of homologous but unequal crossover events. Sequence alignments of the duplicated JH segments indicates that there is diversity within the 5-11 nucleotides located immediately downstream from the heptamer, an observation which indicates that closely related JH segments can serve to enhance CDR3 diversity in the expressed H chain. Comparisons of the genomic JH sequences with different cDNA clones indicate that each JH segment is probably functional and that junctional diversity serves an important role in the generation of CDR3 diversity. In addition, single base differences observed in comparisons of JH-encoded regions indicate that there is probably somatic mutation or allelic variation of genomic JH segments. These studies suggest that the characteristic structure and organizational pattern of JH segments in higher vertebrates may have evolved early in vertebrate phylogeny at the level of the bony fish.

Amino Acid Sequence↗