Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Comprehensive plastome variation and RNA editing in Mentha: insights into phylogenetic relationships and candidate DNA barcodes.

INTRODUCTION: Mentha is an economically and medicinally important genus in Lamiaceae, but its taxonomy and species delimitation remain challenging because of frequent hybridization, polyploidy, and marked morphological plasticity. METHODS: In this study, we comparatively analyzed 12 plastomes representing major Mentha species, hybrid taxa, and unresolved accessions, including four newly assembled genomes, to characterize plastome structure, repeat composition, sequence divergence, phylogenetic relationships, and plastid RNA editing. The M. arvensis plastome and RNA-seq datasets originated from independent Swiss and Indian accessions, respectively. RESULTS: The plastomes were highly conserved in overall organization, ranging from 151,824 to 152,154 bp and displaying the typical quadripartite structure. Gene content and order were largely stable across taxa, with only minor variation likely associated with annotation differences at IR/SC boundary regions. Codon usage analysis revealed a clear bias toward A/U-ending synonymous codons, and most shared protein-coding genes showed low Ka/Ks ratios, indicating predominant purifying selection. Repeat analyses showed that simple sequence repeats were mainly composed of A/T-rich mononucleotide motifs, whereas long repeats were concentrated in the 30-40 bp size class. Comparative analyses identified six hypervariable regions, namely ccsA-ndhD, ycf1, ndhD, rpl32-trnL-UAG, rbcL-accD, and petA-psbJ, which represent promising candidate plastid markers for species discrimination. Phylogenetic analysis based on complete plastomes provided strong support for relationships among the sampled taxa and recovered a close affinity among M. aquatica, M. arvensis, and M. canadensis. In addition, RNA-seq analysis of M. arvensis identified 17 candidate plastid RNA editing sites, most of which were C-to-U conversions and nonsynonymous events. DISCUSSION: Together, these results expand plastid genomic resources for Mentha and provide a useful framework for phylogenetic inference, species identification, and future germplasm utilization.

RNA editing↗

Deciphering Campylobacter jejuni cell surface interactions from the genome sequence.

The completion of the Campylobacter jejuni genome sequence is a landmark in Campylobacter research. Discoveries directly arising from these data include the identification of a capsular polysaccharide, extensive capacity for phase variable gene expression and lipo-oligosaccharide structural phase variation. The recent identification of a unique system of general protein glycosylation in C. jejuni, a C. jejuni protein that is translocated into eukaryotic cells, and plasmid-encoded components of a putative type IV secretion system are likely to be significant in terms of the host-pathogen interaction.

Antigenic Variation↗

PCR amplification of tandemly repeated DNA: analysis of intra- and interchromosomal sequence variation and homologous unequal crossing-over in human alpha satellite DNA.

Tandemly repeated DNA can comprise several percent of total genomic DNA in complex organisms and, in some instances, may play a role in chromosome structure or function. Alpha satellite DNA is the major family of tandemly repeated DNA found at the centromeres of all human and primate chromosomes. Each centromere is characterized by a large contiguous array of up to several thousand kb which can contain several thousand highly homogeneous repeat units. By using a novel application of the polymerase chain reaction (repPCR), we are able to amplify a representative sampling of multiple repetitive units simultaneously, allowing rapid analysis of chromosomal subsets. Direct sequence analysis of repPCR amplified alpha satellite from chromosomes 17 and X reveals positions of sequence heterogeneity as two bands at a single nucleotide position on a sequencing ladder. The use of TdT in the sequencing reactions greatly reduces the background associated with polymerase pauses and stops, allowing visualization of heterogeneous bases found in as little as 10% of the repeat units. Confirmation of these heterogeneous positions was obtained by comparison to the sequence of multiple individual cloned copies obtained both by PCR and non-PCR based methods. PCR amplification of alpha satellite can also reveal multiple repeat units which differ in size. Analysis of repPCR products from chromosome 17 and X allows rapid determination of the molecular basis of these repeat unit length variants, which appear to be a result of unequal crossing-over. The application of repPCR to the study of tandemly repeated DNA should allow in-depth analysis of intra- and interchromosomal variation and unequal crossing-over, thus providing insight into the biology and genetics of these large families of DNA.

Base Sequence↗

Genome-wide association mapping in Arabidopsis identifies previously known flowering time and pathogen resistance genes.

There is currently tremendous interest in the possibility of using genome-wide association mapping to identify genes responsible for natural variation, particularly for human disease susceptibility. The model plant Arabidopsis thaliana is in many ways an ideal candidate for such studies, because it is a highly selfing hermaphrodite. As a result, the species largely exists as a collection of naturally occurring inbred lines, or accessions, which can be genotyped once and phenotyped repeatedly. Furthermore, linkage disequilibrium in such a species will be much more extensive than in a comparable outcrossing species. We tested the feasibility of genome-wide association mapping in A. thaliana by searching for associations with flowering time and pathogen resistance in a sample of 95 accessions for which genome-wide polymorphism data were available. In spite of an extremely high rate of false positives due to population structure, we were able to identify known major genes for all phenotypes tested, thus demonstrating the potential of genome-wide association mapping in A. thaliana and other species with similar patterns of variation. The rate of false positives differed strongly between traits, with more clinal traits showing the highest rate. However, the false positive rates were always substantial regardless of the trait, highlighting the necessity of an appropriate genomic control in association studies.

Arabidopsis↗

Human protein tyrosine phosphatase-like gene: expression profile, genomic structure, and mutation analysis in families with ARVD.

The mouse protein tyrosine phosphatase-like gene (Ptpla) was recently cloned and data suggested that it plays a role in myogenesis and cardiogenesis. The human homologue (PTPLA) was mapped to chromosome 10p13-14, a region where we have mapped a locus responsible for arrhythmogenic right ventricular dysplasia (ARVD). As a positional candidate gene, we characterized PTPLA by determining its tissue expression, its genomic structure, and we also screened for mutations in the ARVD patients. Northern analysis demonstrated PTPLA is preferentially expressed in both adult and fetal heart. A much lower expression was detected in skeletal and smooth muscle tissues. Virtually no expression was observed in other tissues. The protein-encoding sequences of PTPLA consist of seven exons. A sequence variation (Lys64Gln) was found in all the affecteds in a large ARVD family. However, the same variant was also detected in normal control subjects (three alleles/100 chromosomes). Thus, the variant (Lys64Gln) is not responsible for ARVD in our family and is a benign polymorphism. Nevertheless, its tissue-specific expression in the developing and adult heart suggest PTPLA has a role in regulating cardiac development, differentiation, or other cellular events. The genomic structure and intragenic polymorphism of PTPLA should be useful for further clinical and genetic studies such as gene targeting of PTPLA.

Arrhythmogenic Right Ventricular Dysplasia↗

The transposable elements of the Drosophila melanogaster euchromatin: a genomics perspective.

BACKGROUND: Transposable elements are found in the genomes of nearly all eukaryotes. The recent completion of the Release 3 euchromatic genomic sequence of Drosophila melanogaster by the Berkeley Drosophila Genome Project has provided precise sequence for the repetitive elements in the Drosophila euchromatin. We have used this genomic sequence to describe the euchromatic transposable elements in the sequenced strain of this species. RESULTS: We identified 85 known and eight novel families of transposable element varying in copy number from one to 146. A total of 1,572 full and partial transposable elements were identified, comprising 3.86% of the sequence. More than two-thirds of the transposable elements are partial. The density of transposable elements increases an average of 4.7 times in the centromere-proximal regions of each of the major chromosome arms. We found that transposable elements are preferentially found outside genes; only 436 of 1,572 transposable elements are contained within the 61.4 Mb of sequence that is annotated as being transcribed. A large proportion of transposable elements is found nested within other elements of the same or different classes. Lastly, an analysis of structural variation from different families reveals distinct patterns of deletion for elements belonging to different classes. CONCLUSIONS: This analysis represents an initial characterization of the transposable elements in the Release 3 euchromatic genomic sequence of D. melanogaster for which comparison to the transposable elements of other organisms can begin to be made. These data have been made available on the Berkeley Drosophila Genome Project website for future analyses.

Animals↗

Diploid genome assembly of human fibroblast cell lines enables clone specific variant calling, improved read mapping and accurate phasing.

Human cell lines are fundamental tools in biomedical research and are widely used in disease modeling, drug development, and many other domains. Here, we present chromosome-level, phased diploid genome assemblies of two popular human cell lines: the BJ foreskin fibroblast line and the IMR-90 fetal lung fibroblast line. Our high-quality assemblies, generated using long-read and Hi-C sequencing data, reveal substantial structural variation, including more than 50,000 insertions, deletions, duplications, and inversions compared to the recent T2T-CHM13v2.0 reference. Our assemblies provide detailed maps of genetic variation, enabling more accurate variant calling and the ability to phase reads when using newly generated or historical sequencing data on these cell lines or their derivatives. All assemblies and associated data have been made available as a resource for the research community. We envision that diploid genome assembly will become a cornerstone approach for personalized medicine in the near future.

Journal Article↗

Sequences associated with human iris pigmentation.

To determine whether and how common polymorphisms are associated with natural distributions of iris colors, we surveyed 851 individuals of mainly European descent at 335 SNP loci in 13 pigmentation genes and 419 other SNPs distributed throughout the genome and known or thought to be informative for certain elements of population structure. We identified numerous SNPs, haplotypes, and diplotypes (diploid pairs of haplotypes) within the OCA2, MYO5A, TYRP1, AIM, DCT, and TYR genes and the CYP1A2-15q22-ter, CYP1B1-2p21, CYP2C8-10q23, CYP2C9-10q24, and MAOA-Xp11.4 regions as significantly associated with iris colors. Half of the associated SNPs were located on chromosome 15, which corresponds with results that others have previously obtained from linkage analysis. We identified 5 additional genes (ASIP, MC1R, POMC, and SILV) and one additional region (GSTT2-22q11.23) with haplotype and/or diplotypes, but not individual SNP alleles associated with iris colors. For most of the genes, multilocus gene-wise genotype sequences were more strongly associated with iris colors than were haplotypes or SNP alleles. Diplotypes for these genes explain 15% of iris color variation. Apart from representing the first comprehensive candidate gene study for variable iris pigmentation and constituting a first step toward developing a classification model for the inference of iris color from DNA, our results suggest that cryptic population structure might serve as a leverage tool for complex trait gene mapping if genomes are screened with the appropriate ancestry informative markers.

Chromosomes, Human, Pair 10↗

Helicobacter pylori: recombination, population structure and human migrations.

Helicobacter pylori shows extensive genetic diversity and variability due to frequent intraspecific recombination during mixed infection. In the last years, modern genetic and genomic technology as well as cutting-edge population genetic analysis have been used to investigate the population structure and genetic variability of this pathogen. This review article summarizes recent developments in this rapidly moving field.

Ecosystem↗

Ancient haplotypes of the HLA Class II region.

Allelic variation in codons that specify amino acids that line the peptide-binding pockets of HLA's Class II antigen-presenting proteins is superimposed on strikingly few deeply diverged haplotypes. These haplotypes appear to have been evolving almost independently for tens of millions of years. By complete resequencing of 20 haplotypes across the approximately 100-kbp region that spans the HLA-DQA1, -DQB1, and -DRB1 genes, we provide a detailed view of the way in which the genome structure at this locus has been shaped by the interplay of selection, gene-gene interaction, and recombination.

Alleles↗

Evolutionary genomics of Culex pipiens: global and local adaptations associated with climate, life-history traits and anthropogenic factors.

We present the first genome-wide study of recent evolution in Culex pipiens species complex focusing on the genomic extent, functional targets and likely causes of global and local adaptations. We resequenced pooled samples of six populations of C. pipiens and two populations of the outgroup Culex torrentium. We used principal component analysis to systematically study differential natural selection across populations and developed a phylogenetic scanning method to analyse admixture without haplotype data. We found evidence for the prominent role of geographical distribution in shaping population structure and specifying patterns of genomic selection. Multiple adaptive events, involving genes implicated with autogeny, diapause and insecticide resistance were limited to specific populations. We estimate that about 5-20% of the genes (including several histone genes) and almost half of the annotated pathways were undergoing selective sweeps in each population. The high occurrence of sweeps in non-genic regions and in chromatin remodelling genes indicated the adaptive importance of gene expression changes. We hypothesize that global adaptive processes in the C. pipiens complex are potentially associated with South to North range expansion, requiring adjustments in chromatin conformation. Strong local signature of adaptation and emergence of hybrid bridge vectors necessitate genomic assessment of populations before specifying control agents.

Adaptation, Biological↗

Hidden diversity in Enterococcus faecalis revealed by CRISPR2 screening: eco-evolutionary insights into a novel subspecies.

Enterococcus faecalis is a commensal bacterium that colonizes the gut of humans and animals and is a major opportunistic pathogen, known for causing multidrug-resistant healthcare-associated infections (HAIs). Its ability to thrive in diverse environments and disseminate antimicrobial resistance genes (ARGs) across ecological niches highlights the importance of understanding its ecological, evolutionary, and epidemiological dynamics. The CRISPR2 locus has been used as a valuable marker for assessing clonality and phylogenetic relationships in E. faecalis. In this study, we identified a group of E. faecalis strains lacking CRISPR2, forming a distinct, well-supported clade. We demonstrate that this clade meets the genomic criteria for classification as a novel subspecies, here referred to as "subspecies B." Through a comprehensive pangenome analysis and comparative genomics, we explored the adaptive ecological traits underlying this diversification process, identifying clade-specific features and their predicted functional roles. Our findings suggest that the frequent isolation of subspecies B from meat products and processing facilities may reflect dissemination routes involving environmental contamination (e.g., water, plants, soil) from avian species. The absence of key virulence traits required for pathogenicity in mammals, particularly humans, and the lack of clinically relevant resistance determinants indicate that subspecies B currently poses minimal threat to public health compared with the broadly disseminated "subspecies A." Nevertheless, the unclear potential for genetic exchange between these subspecies and the frequent association of subspecies B with food sources calls for continued genomic surveillance of E. faecalis from a One Health perspective to detect and mitigate the emergence of high-risk variants in advance.IMPORTANCEExploring intraspecific genetic variability in generalist bacteria with pathogenic potential, such as Enterococcus faecalis, is a key to uncovering stable evolutionary trends. By screening the CRISPR2 locus across a representative set of genomes from diverse sources, this study reveals a previously unrecognized lineage within the population structure of E. faecalis, associated with underexplored nonhuman and nonhospital reservoirs. These findings broaden our knowledge of the species' genetic landscape and shed light on its adaptive strategies and patterns of ecological dissemination. By bridging phylogenetic patterns with variation in genetic defense systems and accessory traits, the study generates testable hypotheses about the genomic determinants and corresponding selective pressures that shape the species' behavior and long-term dissemination. This work offers new perspectives on the eco-evolutionary dynamics of E. faecalis and highlights the value of genomic surveillance beyond clinical settings, in alignment with One Health principles.

Enterococcus faecalis↗

The structural genes encoding P450scc and P450arom are closely linked on mouse chromosome 9.

The chromosomal location of the two genes that encode the cytochrome P450 enzymes, P450SCC (cholesterol side-chain cleavage) and P450arom (aromatase), was identified in the mouse. Genomic DNA from several progenitor strains of recombinant inbred (RI) strains of mice was tested with various restriction endonucleases for restriction fragment length variations. Variation in Bam HI fragment length was detected between A/J and C57BL/6J. Genomic DNA from 43 RI strains derived from A/J and C57BL/6J was analyzed in a similar manner. Complete concordance of the strain distribution pattern for P450SCC and that of P450arom was observed for 43 RI strains. The lack of recombination indicates that the structural genes encoding P450SCC and P450arom are closely linked. The strain distribution patterns of the P450SCC and P450arom genes were compared with other markers previously mapped in these RI lines. The results demonstrate that both P450SCC and P450arom are found on mouse chromosome 9. Of the other loci on mouse chromosome 9, P450SCC and P450arom are most closely linked to the gene encoding P1450. Among 31 RI strains for which the three loci were analyzed, only one example of discordance was found. Human P450SCC, P450arom and P1450 have been mapped to human chromosome 15. However, the distance between the human P450SCC gene and other loci has not been determined. The information presented in this report, along with other studies, indicate conservation between homologous human and mouse chromosomal regions and suggest that human P450SCC will be found to be closely linked with human P450arom.

Animals↗

Anther culture-derived regenerants of durum wheat and their cytological characterization.

Anther culture is being increasingly used in cereal crop improvement both as a source of haploids and for inducing new genetic variation. We studied the androgenetic ability and regenerability of 10 cultivars of durum wheat (Triticum turgidum L., 2n = 4x = 28; AABB), using three different growth conditions and four media. From a total of 86,400 anthers cultured, 324 plants were obtained: 248 green and 76 albino. Genotype, growth condition, and media significantly affected anther response and callus production; interactions were also significant. Green plant regeneration was influenced significantly by genotype and growth condition, as well as by genotype and growth condition interactions. Albino plant regeneration was significantly affected only by growth condition. Regenerants showed gametoclonal/somaclonal variation. Differences in morphology, growth habit, adult plant height, spike size, and development of spikes at nodes were observed. Mitotic and meiotic chromosomes were studied by conventional staining and fluorescent genomic in situ hybridization techniques. Chromosome numbers of the regenerants ranged from 14 to 70. All 76 haploid plantlets (2n = 2x = 14; AB) were albino. Some of the 28-chromosome regenerants were also albino. Chromosome number in the green plantlets ranged from 28 to 70. Chromosome number also varied in regenerants originating from the same callus. Both intergenomic and intragenomic multivalents were observed. An interesting feature was the preferential multiplication of B-genome chromosomes, which formed multivalents (trivalents, quadrivalents, and hexavalents). We observed several chromosomal abnormalities, which seemed to increase with the level of polyploidy. Translocations, dicentric chromosomes, chromatid exchanges, and Robertsonian translocations involving the A- and B-genome chromosomes were observed. Chromosome breakages resulting in centric and acentric fragments, and telocentrics were observed. Chromosome multiplication and structural aberrations induced during culture may constitute the bases of gametoclonal and somaclonal variations.

Chromosome Banding↗

The heritage of pathogen pressures and ancient demography in the human innate-immunity CD209/CD209L region.

The innate immunity system constitutes the first line of host defense against pathogens. Two closely related innate immunity genes, CD209 and CD209L, are particularly interesting because they directly recognize a plethora of pathogens, including bacteria, viruses, and parasites. Both genes, which result from an ancient duplication, possess a neck region, made up of seven repeats of 23 amino acids each, known to play a major role in the pathogen-binding properties of these proteins. To explore the extent to which pathogens have exerted selective pressures on these innate immunity genes, we resequenced them in a group of samples from sub-Saharan Africa, Europe, and East Asia. Moreover, variation in the number of repeats of the neck region was defined in the entire Human Genome Diversity Panel for both genes. Our results, which are based on diversity levels, neutrality tests, population genetic distances, and neck-region length variation, provide genetic evidence that CD209 has been under a strong selective constraint that prevents accumulation of any amino acid changes, whereas CD209L variability has most likely been shaped by the action of balancing selection in non-African populations. In addition, our data point to the neck region as the functional target of such selective pressures: CD209 presents a constant size in the neck region populationwide, whereas CD209L presents an excess of length variation, particularly in non-African populations. An additional interesting observation came from the coalescent-based CD209 gene tree, whose binary topology and time depth (approximately 2.8 million years ago) are compatible with an ancestral population structure in Africa. Altogether, our study has revealed that even a short segment of the human genome can uncover an extraordinarily complex evolutionary history, including different pathogen pressures on host genes as well as traces of admixture among archaic hominid populations.

Bacterial Infections↗

Genetic polymorphism of human interleukin-1 alpha.

Interleukin-1 alpha (IL-1 alpha) has been implicated in the pathogenesis of infectious, autoimmune and inflammatory diseases. There is, however, very little information on the cis-acting sequences involved in IL-1 alpha regulation or whether there is any variation in the structure of the gene. It is known that intron 6 of IL-1 alpha shows a 5 x 46 bp tandem repeat in the genomic sequence. We have studied this region of the gene. Amplification by polymerase chain reaction showed different sized products from different individuals, most being of higher molecular weight than the expected size of 620 bp. Sequencing demonstrated that the polymorphism was due to a variable number of repeats of the 46 bp sequence. This was confirmed by restriction fragment length analysis of genomic DNA. Altogether, 72 unrelated individuals were tested and 6 alleles ranging from 5 to 18 repeats were found, the most frequent allele (62%) containing 9 repeats. This polymorphism may be of interest in gene function, since each repeat contains three potential binding sites for transcriptional factors: an SP1 site, a viral enhancer element and a glucocorticoid-responsive element. The latter, at least, demonstrates site-specific protein binding by electromobility shift assay. The functional significance of the polymorphism and its allelic frequency in inflammatory and autoimmune diseases are currently under investigation.

Base Sequence↗

Structure and chromosomal localization of the human glycogenin-2 gene GYG2.

Glycogenin-2 is one of two self-glucosylating proteins involved in the initiation phase of the synthesis of the storage polysaccharide glycogen. Cloning of the human glycogenin-2 gene, GYG2, has revealed the presence of 11 exons and a gene of more than 46 kb in size. The structure of the gene explains much of the observed diversity in glycogenin-2 cDNA sequences as being due to alternate exon usage. In some cases, there is variation in the splice junctions used. Over regions of protein sequence similarity, the GYG2 gene structure is similar to that of the other glycogenin gene, GYG. A genomic GYG2 clone was used to localize the gene to Xp22.3 by fluorescence in-situ hybridization. Localization close to the telomere of the short arm of the X chromosome is consistent with mapping information obtained from glycogenin-2 STS sequences. Glycogenin-2 maps between the microsatellite anchor markers AFM319te9 (DXS7100) and AFM205tf2 (DXS1060), and its 3' end is 34.5 kb from the 3' end of the arylsulphatase gene ARSD. GYG2 is outside the pseudoautosomal region PAR1 but still in a region of X-Y shared genes. As is true for several other genes in this location, an inactive remnant of GYG2, consisting of exons 1-3, may be present on the Y chromosome.

Amino Acid Sequence↗

Codon usage in nucleopolyhedroviruses.

Phylogenetic analyses based on baculovirus polyhedrin nucleotide and amino acid sequences revealed two major nucleopolyhedrovirus (NPV) clades, designated Group I and Group II. Subsequent phylogenetic analyses have revealed three Group II subclades, designated A, B and C. Variations in amino acid frequencies determine the extent of dissimilarity for divergent but structurally and functionally conserved genes and therefore significantly influence the analysis of phylogenetic relationships. Hence, it is important to consider variations in amino acid codon usage. The Genome Hypothesis postulates that genes in any given genome use the same coding pattern with respect to synonymous codons and that genes in phylogenetically related species generally show the same pattern of codon usage. We have examined codon usage in six genes from six NPVs and found that: (1) there is significant variation in codon use by genes within the same virus genome; (2) there is significant variation in the codon usage of homologous genes encoded by different NPVs; (3) there is no correlation between the level of gene expression and codon bias in NPVs; (4) there is no correlation between gene length and codon bias in NPVs; and (5) that while codon use bias appears to be conserved between viruses that are closely related phylogenetically, the patterns of codon usage also appear to be a direct function of the GC-content of the virus-encoded genes.

Animals↗