Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Molecular scanning of the human PPARa gene: association of the L162v mutation with hyperapobetalipoproteinemia.

Peroxisome proliferator-activated receptor alpha (PPARalpha) is a member of the steroid hormone receptor super family involved in the control of cellular lipid utilization. This makes PPARalpha a candidate gene for type 2 diabetes and dyslipidemia. The aim of this study was to investigate whether genetic variation in the human PPARalpha gene can influence the risk of type 2 diabetes and dyslipidemia among French Canadians. We therefore first determined the genomic structure of human PPARalpha, and then designed intronic primers to sequence the coding region and the exon-intron boundaries of the gene in 12 patients with type 2 diabetes and in 2 nondiabetic subjects. Sequence analysis revealed the presence of a L162V missense mutation in exon 5 of one diabetic patient. Leucine 162 is contained within the DNA binding domain of the human PPARalpha gene, and is conserved among humans, mice, rats, and guinea pigs. We subsequently screened a sample of 121 patients newly diagnosed with type 2 diabetes and their age and sex-matched nondiabetic controls, recruited from the Saguenay-Lac-St-Jean region of Northeastern Quebec, for the presence of the L162V mutation by a PCR-RFLP based method. There was no difference in L162 homozygote or V162 carrier frequencies between diabetics and nondiabetics. However, whether diabetic or not, carriers of the V162 allele had higher plasma apolipoprotein B levels compared to noncarriers (P 5 0.05). To further this association, we screened another sample of 193 nondiabetic subjects recruited in the greater Quebec City area. Carriers of the V162 allele compared with homozygotes of the L162 allele had significantly higher concentrations of plasma total and LDL-apolipoprotein B as well as LDL cholesterol (P </= 0.02). These results suggest an association between the PPARalpha V162 allele and the atherogenic/hyperapolipoprotein B dyslipidemia.

Animals↗

Genome-wide SNP-based genomic diversity and population structure analysis in alpaca populations from Europe and Peru.

This study aimed to analyze the genetic diversity and population structure of alpacas in Germany, Switzerland, and Austria (German-speaking regions, GSR) and to compare with that of the country of origin of the species (Peru). A total of 179 animals from GSR and 151 from Peru were genotyped with a species-specific 76k SNP array. The observed and expected heterozygosity was 0.305 and 0.311 for GSR and 0.310 and 0.312 for Peru. The mean FROH values were 0.029 for GSR and 0.023 for Peru. In general, results show that breeders in both analyzed regions efficiently maintain genetic diversity. Principal component analysis identified the GSR and Peru populations as separate from each other, but the relative proximity of both clusters indicates the shared genetic heritage. FST and XPEHH methods identified genomic regions under selection for traits such as coat color and adaptation. Genome-wide association studies comparing black and brown with white or gray alpacas identified associated genome regions containing the ASIP and KIT genes, respectively. The association of a recently identified keratin locus on chromosome 16 with differences in fleece type in alpacas was confirmed, while the putative causality of a TRPV3 variant was rejected.

Animals↗

The effects of alternative splicing on transmembrane proteins in the mouse genome.

Alternative splicing is a major source of variety in mammalian mRNAs, yet many questions remain on its downstream effects on protein function. To this end, we assessed the impact of gene structure and splice variation on signal peptide and transmembrane regions in proteins. Transmembrane proteins perform several key functions in cell signaling and transport, with their function tied closely to their transmembrane architecture. Signal peptides and transmembrane regions both provide key information on protein localization. Thus, any modification to such regions will likely alter protein destination and function. We applied TMHMM and SignalP to a nonredundant set of proteins, and assessed the effects of gene structure and alternative splicing on predicted transmembrane and signal peptide regions. These regions were altered by alternative splicing in roughly half of the cases studied. Transmembrane regions are divided by introns slightly less often than expected given gene structure and transmembrane region size. However, the transmembrane regions in single-pass transmembranes are divided substantially less often than expected. This suggests that intron placement might be subject to some evolutionary pressure to preserve function in these signaling proteins. The data described in this paper is available online at http://www.affymetrix.com/community/publications/affymetrix/tmsplice/.

Alternative Splicing↗

Genomic structure and organization of kringles type 3 to 10 of the apolipoprotein(a) gene in 6q26-27.

Apolipoprotein(a) [apo(a)] is a highly polymorphic glycoprotein covalently linked to the apolipoprotein B-100 of LDL in a particle called lipoprotein(a) [Lp(a)]. High plasma levels of Lp(a) are associated with coronary as well as peripheral atherosclerosis. Plasma levels of Lp(a) show a remarkable variation ranging from 0.1 mg/dl to over 100 mg/dl. The apo(a) gene shows a size polymorphism which resides in the variable number of kringle domains which resemble plasminogen kringle IV. Ten different types of kringle IV repeats have been described, nine of which (kringle IV type 1 and type 3-10) are each supposed to be present in a single copy. The other kringles, namely kringle IV type 2 repeats, vary in number from 3 to 42 between apo(a) alleles and form the basis for the apo(a) size polymorphism. Although an inverse relationship has been observed between the number of kringle type 2 repeats and plasma levels of Lp(a), there are exceptions to this general finding. Indeed, several individuals have been described with similar apo(a) size alleles but very different plasma levels of Lp(a). Genetic studies have linked these differences to the apo(a) locus on 6q26-27, outlining the importance, besides the kringle type 2 repeats, of other regions of the apo(a) gene in contributing to the interindividual differences in the plasma concentration of Lp(a). One of the candidate regions is represented by the non-repeated type-3 to type-10 kringles which are invariably present in each apo(a) allele and whose structural integrity is playing a critical role in the correct assembly of the Lp(a) particle. Biochemical studies with recombinant wild type and mutagenized apo(a) cDNAs with several alterations of the non-repeated kringles have well documented this latter point. As a starting point to search for genetic variations in these kringles associated with different levels of Lp(a), we are presenting the genome organization of type-3 to 10 kringle along with specific PCR primers for easy analysis from genomic DNA. Restriction as well as partial sequencing analyses of the type-3 to 10 kringles region has also provided interesting clues as to the different evolutionary origin of these types of kringle with respect to the polymorphic type-2 kringles.

Apolipoproteins A↗

Structure and functional genomics of lipopolysaccharide expression in Haemophilus influenzae.

The involvement of genes in the lic loci in H. influenzae LPS expression has been known for some time. However, it was not until recently that it was shown that the lic1 locus contains genes required for phase variable expression of phosphocholine substituents, while genes in the lic2 locus and lgtC are required for expression of the globoside trisaccharide, alpha-D-Galp-(1 --> 4)-beta-D-Galp-(1 --> 4)-beta-D-Glcp (i.e., the pK blood group epitope). The availability of the complete sequence of the H. influenzae strain Rd genome has facilitated significant progress in understanding the role of these and other genes in the expression and biosynthesis of LPS. We have employed a comparative structural fingerprinting strategy to establish the structural relationships among LPS from H. influenzae mutant strains in which putative biosynthesis genes were inactivated. Using this functional genomics approach, we have gained considerable insight into the genetic basis for intra-strain and strain-to-strain variation in epitope expression.

Base Sequence↗

DNA structure constraint is probably a fundamental factor inducing CpG deficiency in bacteria.

MOTIVATION: It has been speculated that CpG dinucleotide deficiency in genomes is a consequence of DNA methylation. However, this hypothesis does not adequately explain CpG deficiency in bacteria. The hypothesis based on DNA structure constraint as an alternative explanation was therefore examined. RESULTS: By comparing real bacterial genomes and Markov artificial genomes in the second order, we found that the core structure of a restricted pattern, the TTCGAA pattern, was under represented in low GC content bacterial genomes regardless of CpG dinucleotide level. This is in contrast to the AACGTT pattern, indicating that the counterselection is context-dependent. Further study discovered nine underrepresented patterns that were supposed to be capable of inducing DNA structure constraint. In summary, most of them are in TTCGNA and TTCGAN patterns in both DNA strands. An explanation is also proposed for the strong correlation between GC content and CpG deficiency. The result of random sequence simulation showed that the occurrences of these patterns were correlated with GC content, as well as the percentage of CpG dinucleotides being trapped in these patterns. Finally, we suggest that the degree of counter-selection against these restricted patterns could be influenced by global GC content of a genome.

Base Sequence↗

Evolution of the cetacean mitochondrial D-loop region.

We sequenced the mitochondrial DNA D-loop regions from two cetacean species and compared these with the published D-loop sequences of several other mammalian species, including one other cetacean. Nucleotide substitution rates, DNA sequence simplicity, possible open reading frames (ORFs), and potential RNA secondary structure were investigated. The substitution rate is an order of magnitude lower than would be expected on the basis of reports on human sequence variation in this region but are consistent with interspecific primate and rodent D-loop sequence variation and with estimates of substitution rates from whole mitochondrial genomes. Deletions/insertions are less common in the cetacean D-loop than in other vertebrate species. Areas of high sequence simplicity (clusters of short repetitive motifs) across the region correspond to areas of high sequence divergence. Three regions predicted to form secondary structures are homologous to such putative structures in other species; however, the presumptive structures most conserved in cetaceans are different from those reported for other taxa. While all three species have possible long ORFs, only a short sequence of seven amino acids is shared with other mammalian species, and those changes that had occurred within it are all nonsynonymous. We conclude that DNA slippage, in addition to point mutation, contributes to the evolution of the D-loop and that regions of conserved secondary structure in cetaceans and an ORF are unlikely to contribute significantly to the conservation of the central region.

Amino Acid Sequence↗

Complete sequence of the mitochondrial DNA of the annelid worm Lumbricus terrestris.

We have determined the complete nucleotide (nt) sequence of the mitochondrial genome of an oligochaete annelid, the earthworm Lumbricus terrestris. This genome contains the 37 genes typical of metazoan mitochondrial DNA (mtDNA), including ATPase8, which is missing from some invertebrate mtDNAs. ATPase8 is not immediately upstream of ATPase6, a condition found previously only in the mtDNA of snails. All genes are transcribed from the same DNA strand. The largest noncoding region is 384 nt and is characterized by several homopolymer runs, a tract of alternating TA pairs, and potential secondary structures. All protein-encoding genes either overlap the adjacent downstream gene or end at an abbreviated stop codon. In Lumbricus mitochondria, the variation of the genetic code that is typical of most invertebrate mitochondrial genomes is used. Only the codon ATG is used for translation initiation. Lumbricus mtDNA is A + T rich, which appears to affect the codon usage pattern. The DHU arm appears to be unpaired not only in tRNAser(AGN), as is typical for metazoans, but perhaps also in tRNAser(UCN), a condition found previously only in a chiton and among nematodes. Relating the Lumbricus gene organization to those of other major protostome groups requires numerous rearrangements.

Amino Acid Sequence↗

Constrained genomic and conformational variability of the hypervariable region 1 of hepatitis C virus in chronically infected patients.

We analysed the genomic and conformational variability of the hypervariable region 1 (HVR1) of the hepatitis C virus (HCV) to evaluate the importance of its biological role. A total of 865 genotype 1b HVR1 subclones were collected from serially sampled sera in 11 patients with chronic hepatitis C, four of whom received interferon therapy. Consequently, 169 distinct sequences were examined for amino acid substitutions as well as hydrophilic or hydrophobic profile at each amino acid position within HVR1. Secondary structure of HVR1 was also predicted by the method of Robson in 90 distinct sequences from eight patients, including three interferon-treated patients. Some positions within the HVR1 were invariable or nearly so as to amino acid substitution. Hydrophilic or hydrophobic residues exclusively predominated at several positions. These constrained amino acid replacement and hydrophilic or hydrophobic profiles were conserved irrespective of interferon therapy, though the frequency of amino acid replacement was greater at almost all amino acid positions within the HVR1 in interferon-treated patients. The quasispecies of HCV showed various secondary structures of HVR1, but many sequences seemed to have common characteristics. beta sheet conformations around both the N-terminus and position 20 (numbered from the NH2 terminus of E2 envelope glycoprotein), and/or coil structures around the C-terminus of HVR1 could be identified. These results suggest that HVR1 amino acid replacements are strongly constrained by a well-ordered structure, in spite of being tolerant to amino acid substitutions, and imply an important biological role of the HVR1 protein in HCV replication.

Adult↗

The rat STSL locus: characterization, chromosomal assignment, and genetic variations in sitosterolemic hypertensive rats.

BACKGROUND: Elevated plant sterol accumulation has been reported in the spontaneously hypertensive rat (SHR), the stroke-prone spontaneously hypertensive rat (SHRSP) and the Wistar-Kyoto (WKY) rat. Additionally, a blood pressure quantitative trait locus (QTL) has been mapped to rat chromosome 6 in a New Zealand genetically hypertensive rat strain (GH rat). ABCG5 and ABCG8 (encoding sterolin-1 and sterolin-2 respectively) have been shown to be responsible for causing sitosterolemia in humans. These genes are organized in a head-to-head configuration at the STSL locus on human chromosome 2p21. METHODS: To investigate whether mutations in Abcg5 or Abcg8 exist in SHR, SHRSP, WKY and GH rats, we initiated a systematic search for the genetic variation in coding and non-coding region of Abcg5 and Abcg8 genes in these strains. We isolated the rat cDNAs for these genes and characterized the genomic structure and tissue expression patterns, using standard molecular biology techniques and FISH for chromosomal assignments. RESULTS: Both rat Abcg5 and Abcg8 genes map to chromosome band 6q12. These genes span ~40 kb and contain 13 exons and 12 introns each, in a pattern identical to that of the STSL loci in mouse and man. Both Abcg5 and Abcg8 were expressed only in liver and intestine. Analyses of DNA from SHR, SHRSP, GH, WKY, Wistar, Wistar King A (WKA) and Brown Norway (BN) rat strains revealed a homozygous G to T substitution at nucleotide 1754, resulting in the coding change Gly583Cys in sterolin-1 only in rats that are both sitosterolemic and hypertensive (SHR, SHRSP and WKY). CONCLUSIONS: The rat STSL locus maps to chromosome 6q12. A non-synonymous mutation in Abcg5, Gly583Cys, results in sitosterolemia in rat strains that are also hypertensive (WKY, SHR and SHRSP). Those rat strains that are hypertensive, but not sitosterolemic (e.g. GH rat) do not have mutations in Abcg5 or Abcg8. This mutation allows for expression and apparent apical targeting of Abcg5 protein in the intestine. These rat strains may therefore allow us to study the pathophysiological mechanisms involved in the human disease of sitosterolemia.

Animals↗

Sweeps in Space: Leveraging Geographic Data to Identify Beneficial Alleles in Anopheles gambiae.

As organisms adapt to environmental changes, natural selection modifies the frequency of nonneutral alleles. For beneficial mutations, the outcome of this process may be a selective sweep, in which an allele rapidly increases in frequency and perhaps reaches fixation within a population. Selective sweeps have well-studied effects on patterns of local genetic variation in panmictic populations, but much less is known about the dynamics of sweeps in continuous space. In particular, because limited movement across a landscape leads to unique patterns of population structure, spatial dynamics may influence the trajectory of selected mutations. Here, we use forward-in-time, individual-based simulations in continuous space to study the impact of space on beneficial mutations as they sweep through a population. In particular, we show that selection changes the joint distribution of allele frequency and geographic range occupied by a focal allele and demonstrate that this signal can be used to identify selective sweeps. We then leverage this signal to identify in-progress selective sweeps within the malaria vector Anopheles gambiae, a species under strong selection pressure from vector control measures. By considering space, we identify multiple previously undescribed variants with potential phenotypic consequences, including mutations impacting known IR-associated genes and altering protein structure and properties. Our results demonstrate a novel signal for detecting selection in spatial population genetic data that may have implications for genomic surveillance and understanding geographic patterns of genetic variation.

Animals↗

Sweeps in space: leveraging geographic data to identify beneficial alleles in Anopheles gambiae.

As organisms adapt to environmental changes, natural selection modifies the frequency of non-neutral alleles. For beneficial mutations, the outcome of this process may be a selective sweep, in which an allele rapidly increases in frequency and perhaps reaches fixation within a population. Selective sweeps have well-studied effects on patterns of local genetic variation in panmictic populations, but much less is known about the dynamics of sweeps in continuous space. In particular, because limited movement across a landscape leads to unique patterns of population structure, spatial dynamics may influence the trajectory of selected mutations. Here, we use forward-in-time, individual-based simulations in continuous space to study the impact of space on beneficial mutations as they sweep through a population. In particular, we show that selection changes the joint distribution of allele frequency and geographic range occupied by a focal allele and demonstrate that this signal can be used to identify selective sweeps. We then leverage this signal to identify in-progress selective sweeps within the malaria vector Anopheles gambiae , a species under strong selection pressure from vector control measures. By considering space, we identify multiple previously undescribed variants with potential phenotypic consequences, including mutations impacting known IR-associated genes and altering protein structure and properties. Our results demonstrate a novel signal for detecting selection in spatial population genetic data that may have implications for genomic surveillance and understanding geographic patterns of genetic variation.

Journal Article↗

Differences in the organization and methylation patterns of integrated avian sarcoma proviral DNA sequences in nonpermissive and permissive mammalian cells.

Four types of avian sarcoma virus (ASV)-transformed mammalian cells were analyzed for the presence of ASV-specific sequences in their genome DNA. A great variability in the number of proviral copies and their structure within the DNA of these lines was observed. In all cells tested gag and src sequences were present in a flexible arrangement. The greatest variation was detected within proviral sequences corresponding to pol and env regions. The number of integrated ASV proviral copies do not correlate with the capability of these cells to produce viral particles. Virus-producing (K2S and K12) and virogenic (XC) cells contain in their chromosomal DNA at least one complete proviral genome, whereas proviral sequences in helper-dependent nonvirogenic cells are substantially changed. Provirus expression level does not correlate with the number of integrated virus copies. In the non-virus-producing cells the proviral sequences are hypermethylated.

Animals↗

Complex haplotype structure of the human GNAS gene identifies a recombination hotspot centred on a single nucleotide polymorphism widely used in association studies.

The alpha subunit of the heterotrimeric G protein Gs (Gsalpha) is involved in numerous physiological processes and is a primary determinant of cellular responses to extracellular signals. Genetic variations in the Gsalpha gene may play an important role in complex diseases and drug responses. To characterize the genetic diversity in this locus, we resequenced exons and flanking introns of the gene in 44 genomic samples and analysed the haplotype structure of the gene in an additional 50 African-Americans and 50 Caucasians. Significant differences in allele frequency for nearly all the genotyped single nucleotide polymorphism (SNPs) were detected between the two ethnic groups. Linkage disequilibrium (LD) analysis of this locus revealed two haplotype blocks characterized by strong LD and reduced haplotype diversity, especially in Caucasians. Between the two blocks is a narrow (approximately 3 kb) recombination hotspot centred on exons 4 and 5, and a widely used genetic marker in association studies in this region (rs7121) was in linkage equilibrium with the rest of the gene. The haplotype structure of the GNAS locus warrants reevaluation of previous association studies that used marker rs7121 and affects choice of SNP markers to be used in future studies of this locus.

Black or African American↗

Pangenome of Streptomyces sampsonii and Relatives Highlights Horizontal Gene Transfer and Secondary Metabolism in Environmental Adaptation and Ecological Significance.

Streptomyces sampsonii is a promising biocontrol bacterium, but its genomic basis of adaptation and secondary metabolism remains unclear. Here, we present a chromosome-level genome assembly of S. sampsonii (7.20&#x2009;Mb, 6015 protein-coding genes) and perform comparative analyses with 95 related Streptomyces species. Phylogenomic and synteny analyses revealed its closest relationship with S. albidoflavus, while extensive structural variations distinguished more distant lineages. Pangenome analysis uncovered 84,178 gene clusters, with pan_shell and pan_cloud genes predominantly enriched in xenobiotic biodegradation, metabolism, and antibiotic biosynthesis, highlighting their roles in ecological adaptation and biocontrol potential. Biosynthetic gene cluster (BGC) analysis identified numerous NRPS, PKS, and terpene pathways, many of which belong to pan_shell and pan_cloud regions, suggesting dynamic evolutionary origins. We further detected 66,260 horizontally transferred (HGT) genes, including 438 in BGCs, underscoring HGT as a major driver of metabolic innovation. Together, these findings provide novel insights into the genomic diversity, adaptive capacity, and secondary metabolic potential of S. sampsonii and its close relatives.

BGCs↗

Microgeographic variation in rDNA intergenic spacers of Anopheles gambiae in western Kenya.

The genetic population structure of Anopheles gambiae (Diptera: Culicidae) in western Kenya was investigated by hybridizing a rapidly evolving rDNA intergenic spacer sequence to restriction endonuclease digests of genomic DNA extracted from single mosquitoes from seven localities. Significantly different distributions of restriction fragment arrays were obtained from field sites less than 10 km apart, which suggests restricted gene flow and a subdivided population structure. Eight of twenty-one possible comparisons between pairs of populations yielded significant differences. An eastern Kenya coastal population did not share its restriction fragment arrays with any of the western populations, suggesting that isolation by distance can be complete on a relatively small geographic scale (700 km).

Animals↗

Low sequence variation among isolates of infectious hypodermal and hematopoietic necrosis virus (IHHNV) originating from Hawaii and the Americas.

A 2.9 kb fragment of the infectious hypodermal and hematopoietic necrosis virus (IHHNV) genome, which contains the coding sequence of putative non-structural and capsid proteins, was amplified and sequenced from each of 14 IHHNV isolates collected from cultured penaeid shrimp stocks in Hawaii and various sites in the Americas between 1982 and 1997. The sequence comparison indicates that the IHHNV genome is very stable, with 99.6 to 100% similarity among these 14 isolates. Only nucleotide substitutions were found. The percentage of substitution was higher in the putative capsid proteins region (1.3%) than in the putative non-structural proteins region (0.6%). Out of 25 substitutions found, 14 resulted in amino acid changes. There is no apparent association between clinical outcomes and particular amino acid substitutions. Based on genetic distances, the isolates were clustered into 3 groups that generally correspond with their geographic origins.

Animals↗

Numerous group I introns with variable distributions in the ribosomal DNA of a lichen fungus.

The length of the small subunit ribosomal DNA (SSU rDNA) differs significantly among individuals from natural populations of the ascomycetous lichen complex Cladonia chlorophaea. The sequence of the 3' region of the SSU rDNA from two individuals, chosen to represent the shortest and longest sequences, revealed multiple insertions within a region that otherwise aligned with a 520-nucleotide sequence of the SSU rDNA in Saccharomyces cerevisiae. The high degree of variability in SSU rDNA size can be accounted for by different numbers of insertions; one individual had two group I introns and the second had five introns, two of which were clearly related to introns at identical positions in the other individual. Yet, introns in different positions, whether within an individual or between individuals, were not similar in sequence. The distribution of introns at three of the positions is consistent with either intron loss or acquisition, and clearly indicates the dynamic variability in this region of the nuclear genome. All seven insertions, which ranged in size from 210 to 228 nucleotides, had the conserved sequence and secondary structural elements of group I introns. The variation in distribution and sequence of group I introns within a short highly conserved region of rDNA presents a unique opportunity for examining the molecular evolution and mobility of group I introns within a systematics framework.

Ascomycota↗