Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Annexin A11 (ANXA11) gene structure as the progenitor of paralogous annexins and source of orthologous cDNA isoforms.

The genomic organization of the annexin A11 gene was determined in mouse and human to assess its congruity with other family members and to examine the species variation in alternative splicing patterns. Mouse annexin A11 genomic clones were characterized by restriction analysis, Southern blotting, and DNA sequencing, and the homologous human gene (HGMW-approved gene symbol ANXA11) was deciphered from high-throughput genomic sequence with coanalysis of expressed sequence tags. Exons 6-15 of the tetrad core repeat region differ from annexins A7 and A13 but are spliced identically to other phylogenetic descendents, making annexin A11 the putative primary progenitor of up to nine paralogous human annexins. The 5' regions consist of untranslated exon 1, followed by an extensive intron 1 comprising almost half the total gene length of >40 kb, and additional GC-rich exons 2-5 encoding the proline- and glycine-rich amino-terminus. Distinct cDNA isoforms in cow and human were determined to be unique to each species and hence of dubious general significance for this gene's function. Multiple transcription start sites were revealed by primer extension analysis of the mouse gene, and transfection constructs containing the prospective promoter generated transcriptional activity comparable to that of the SV40 promoter. Internal repetitive elements and vicinal gene markers were mapped for the complete human annexin A11 gene sequence to characterize the surrounding genomic environment.

3T3 Cells↗

Genome-wide SNP data support species boundaries in sympatric Polylepis Ruiz & Pav. (Rosaceae) species from Bolivia and Ecuador.

Species delimitation in the South American genus Polylepis is notoriously challenging due to high morphological similarity and phenotypic plasticity, likely driven by hybridization and gene flow. Previous phylogenetic studies suggested that genetic structure aligns more strongly with geography than with taxonomy, questioning existing species concepts and hampering conservation efforts. We used double-digest RAD sequencing (ddRADseq) to generate genome-wide SNP data for 11 Polylepis species sampled across multiple localities in Bolivia and Ecuador. Population genetic analyses, phylogenetic inference, and network approaches were combined to assess whether genetic structure aligns more closely with taxonomy or geography. Morphologically defined species formed largely cohesive genetic lineages across regions, with species identity explaining substantially more genetic variation than locality. While localized admixture and reticulation were detected among closely related taxa, widespread species showed strong genetic cohesion and clear separation from congeners. Our results indicate that the sampled Polylepis species from Bolivia and Ecuador maintain distinct genetic identities despite localized signals consistent with gene flow. This genome-wide support for current taxonomy highlights Polylepis as a valuable model for studying speciation under gene flow and indicates that multiple geographic sampling will be essential in reconstructing a robust phylogeny of the genus, with important implications for conservation planning in Andean montane forests.

Bolivia↗

The diverse and dynamic structure of bacterial genomes.

Bacterial genome sizes, which range from 500 to 10,000 kbp, are within the current scope of operation of large-scale nucleotide sequence determination facilities. To date, 8 complete bacterial genomes have been sequenced, and at least 40 more will be completed in the near future. Such projects give wonderfully detailed information concerning the structure of the organism's genes and the overall organization of the sequenced genomes. It will be very important to put this incredible wealth of detail into a larger biological picture: How does this information apply to the genomes of related genera, related species, or even other individuals from the same species? Recent advances in pulsed-field gel electrophoretic technology have facilitated the construction of complete and accurate physical maps of bacterial chromosomes, and the many maps constructed in the past decade have revealed unexpected and substantial differences in genome size and organization even among closely related bacteria. This review focuses on this recently appreciated plasticity in structure of bacterial genomes, and diversity in genome size, replicon geometry, and chromosome number are discussed at inter- and intraspecies levels.

Bacteria↗

Structure and evolution of transcriptional regulatory networks.

The regulatory interactions between transcription factors and their target genes can be conceptualised as a directed graph. At a global level, these regulatory networks display a scale-free topology, indicating the presence of regulatory hubs. At a local level, substructures such as motifs and modules can be discerned in these networks. Despite the general organisational similarity of networks across the phylogenetic spectrum, there are interesting qualitative differences among the network components, such as the transcription factors. Although the DNA-binding domains of the transcription factors encoded by a given organism are drawn from a small set of ancient conserved superfamilies, their relative abundance often shows dramatic variation among different phylogenetic groups. Large portions of these networks appear to have evolved through extensive duplication of transcription factors and targets, often with inheritance of regulatory interactions from the ancestral gene. Interactions are conserved to varying degrees among genomes. Insights from the structure and evolution of these networks can be translated into predictions and used for engineering of the regulatory networks of different organisms.

Amino Acid Motifs↗

How much expression divergence after yeast gene duplication could be explained by regulatory motif evolution?

We used the yeast genome sequences of gene families, microarray profiles and regulatory motif data to test the current wisdom that there is a strong correlation between regulatory motif structure and gene expression profile. Our results suggest that duplicate genes tend to be co-expressed but the correlation between motif content and expression similarity is generally poor, only approximately 2-3% of expression variation can be explained by the motif divergence. Our observations suggest that, in addition to the cis-regulatory motif structure in the upstream region of the gene, multiple trans-acting factors in the gene network can influence the pattern of gene expression significantly.

Evolution, Molecular↗

Functional genomic analysis of the rates of protein evolution.

The evolutionary rates of proteins vary over several orders of magnitude. Recent work suggests that analysis of large data sets of evolutionary rates in conjunction with the results from high-throughput functional genomic experiments can identify the factors that cause proteins to evolve at such dramatically different rates. To this end, we estimated the evolutionary rates of >3,000 proteins in four species of the yeast genus Saccharomyces and investigated their relationship with levels of expression and protein dispensability. Each protein's dispensability was estimated by the growth rate of mutants deficient for the protein. Our analyses of these improved evolutionary and functional genomic data sets yield three main results. First, dispensability and expression have independent, significant effects on the rate of protein evolution. Second, measurements of expression levels in the laboratory can be used to filter data sets of dispensability estimates, removing variates that are unlikely to reflect real biological effects. Third, structural equation models show that although we may reasonably infer that dispensability and expression have significant effects on protein evolutionary rate, we cannot yet accurately estimate the relative strengths of these effects.

Evolution, Molecular↗

Diversity of mosaic structures and common ancestry of human immunodeficiency virus type 1 BF intersubtype recombinant viruses from Argentina revealed by analysis of near full-length genome sequences.

The findings that BF intersubtype recombinant human immunodeficiency type 1 viruses (HIV-1) with coincident breakpoints in pol are circulating widely in Argentina and that non-recombinant F subtype viruses have failed to be detected in this country were reported recently. To analyse the mosaic structures of these viruses and to determine their phylogenetic relationship, near full-length proviral genomes of eight of these recombinant viruses were amplified by PCR and sequenced. Intersubtype breakpoints were analysed by bootscanning and examining the signature nucleotides. Phylogenetic relationships were determined with neighbour-joining trees. Five viruses, each with predominantly subtype F genomes, exhibited mosaic structures that were highly similar. Two intersubtype breakpoints were shared by all viruses and seven by the majority. Of the consensus breakpoints, all nine were present in two viruses, which exhibited identical recombinant structures, and four to eight breakpoints were present in the remaining viruses. Phylogenetic analysis of partial sequences supported both a common ancestry, at least in part of their genomes, for all recombinant viruses and the phylogenetic relationship of F subtype segments with F subtype viruses from Brazil. A common ancestry of the recombinants was supported also by the presence of shared signature amino acids and nucleotides, either unreported or highly unusual in F and B subtype viruses. These results indicate that HIV-1 BF recombinant viruses with diverse mosaic structures, including a circulating recombinant form (which are widespread in Argentina) derive from a common recombinant ancestor and that F subtype segments of these recombinants are related phylogenetically to the F subtype viruses from Brazil.

Argentina↗

Identification, typing, and insecticidal activity of Xenorhabdus isolates from entomopathogenic nematodes in United Kingdom soil and characterization of the xpt toxin loci.

Xenorhabdus strains from entomopathogenic nematodes isolated from United Kingdom soils by using the insect bait entrapment method were characterized by partial sequencing of the 16S rRNA gene, four housekeeping genes (asd, ompR, recA, and serC) and the flagellin gene (fliC). Most strains (191/197) were found to have genes with greatest similarity to those of Xenorhabdus bovienii, and the remaining six strains had genes most similar to those of Xenorhabdus nematophila. Generally, 16S rRNA sequences and the sequence types based on housekeeping genes were in agreement, with a few notable exceptions. Statistical analysis implied that recombination had occurred at the serC locus and that moderate amounts of interallele recombination had also taken place. Surprisingly, the fliC locus contained a highly variable central region, even though insects lack an adaptive immune response, which is thought to drive flagellar variation in pathogens of higher organisms. All the X. nematophila strains exhibited a consistent pattern of insecticidal activity, and all contained the insecticidal toxin genes xptA1A2B1C1, which were present on a pathogenicity island (PAI). The PAIs were similar among the X. nematophila strains, except for partial deletions of a peptide synthetase gene and the presence of insertion sequences. Comparison of the PAI locus with that of X. bovienii suggested that the PAI integrated into the genome first and then acquired the xpt genes. The independent mobility of xpt genes was further supported by the presence of xpt genes in X. bovienii strain I73 on a type 2 transposon structure and by the variable patterns of insecticidal activity in X. bovienii isolates, even among closely related strains.

Alleles↗

[Mechanisms of antigenic variation in influenza virus].

Human influenza A viruses evolve rapidly by antigenic shift and antigenic drift. Antigenic shift occurs by genetic reassortment between currently circulating human viruses and influenza viruses of other origin, by re-emergence of a previously circulating virus, and by invasion of animal influenza viruses. The segmental structure of the virus genome enables reassortment in multiply infected cells and also promotes multiple infection because it results in yielding noninfectious particles which randomly lack some genome segment and only become infectious by complementation with others. Antigenic drift is due to an accumulation of nonsynonymous substitutions in the genes encoding the HA and NA proteins. The limited nature of infection to the respiratory epithelium which is the border of the immune system enables the virus to reinfect and grow under the partial immune pressure which results in selecting and expanding antigenic mutants.

Animals↗

Searching for RNA genes using base-composition statistics.

The hypothesis that genomic regions rich in non-protein-coding RNAs (ncRNAs) can be identified using local variations in single-base and dinucleotide statistics has been investigated. (G+C)%, (G-C)% difference, (A-T)% difference and dinucleotide-frequency statistics were compared among seven classes of ncRNAs and three genomes. Significant variations were observed in (G+C)% and, in Methanococcus jannaschii, in the frequency of the dinucleotide 'CG'. Screening programs based on these two base-composition statistics were developed. With (G+C)% screening alone, a 1% fraction of the M.jannaschii genome containing all 44 known transfer RNAs, ribosomal RNAs and signal recognition particle RNAs could be identified. When (G+C)% combined with CG dinucleotide-frequency screening was used, 43 of the 44 known M.jannaschii structural ncRNAs were again identified, while the number of presumably false hits overlapping a known or putative protein-coding gene was reduced from 15 to 6. In addition, 19 candidate ncRNAs were identified including one with significant homology to several known archaeal RNaseP RNAs.

Animals↗

A single expression site with a conserved leader sequence regulates variation of expression of the Pneumocystis carinii family of major surface glycoprotein genes.

The major surface glycoprotein (MSG) of Pneumocystis carinii is encoded by a family of related but distinct genes distributed throughout the P. carinii genome. Previous reports of the genomic and mRNA MSG structure suggested that there was a highly conserved 5'-untranslated region and a highly variable translated region. In the current study, we demonstrate that there is a single expression site for MSG expression and that different MSG genes are located downstream of this expression site. Isolation of a genomic clone containing the putative 5'-untranslated region has demonstrated that there was a single base sequencing error in what was considered to be the untranslated region. The corrected sequence reveals an extended open reading frame encoding a constant amino-terminal leader domain, with a typical signal peptide, for the MSG protein family. Since this constant amino-terminal domain is encoded by a single copy genomic sequence, a recombination/gene conversion-mediated antigenic switching event is required to effect the known variability in expressed MSG sequences. Therefore, like some bacterial and protozoan pathogens, the opportunistic fungal pathogen P. carinii contains a constant genomic site dedicated to MSG expression and a switchable downstream region for the variable part of the MSG gene family.

Amino Acid Sequence↗

A graph-based approach for the visualisation and analysis of bacterial pangenomes.

BACKGROUND: The advent of low cost, high throughput DNA sequencing has led to the availability of thousands of complete genome sequences for a wide variety of bacterial species. Examining and interpreting genetic variation on this scale represents a significant challenge to existing methods of data analysis and visualisation. RESULTS: Starting with the output of standard pangenome analysis tools, we describe the generation and analysis of interactive, 3D network graphs to explore the structure of bacterial populations, the distribution of genes across a population, and the syntenic order in which those genes occur, in the new open-source network analysis platform, Graphia. Both the analysis and the visualisation are scalable to datasets of thousands of genome sequences. CONCLUSIONS: We anticipate that the approaches presented here will be of great utility to the microbial research community, allowing faster, more intuitive, and flexible interaction with pangenome datasets, thereby enhancing interpretation of these complex data.

Bacteria↗

Population structure and genetic diversity within California Citrus tristeza virus (CTV) isolates.

The Closterovirus, Citrus tristeza virus (CTV) is an aphid-borne RNA virus that is the causal agent of important worldwide economic losses in citrus. Biological and molecular variation has been observed for many CTV isolates. In this work we detected and analyzed sequence variants (haplotypes) within individual CTV isolates. We studied the population structure of five California CTV isolates by single strand conformation polymorphism (SSCP) analysis of four CTV genomic regions. Also, we estimated the genetic diversity within and between isolates by analysis of haplotype nucleotide sequences. Most CTV isolates were composed of a population of genetically related variants (haplotypes), one being predominant. However in one case, we found a high nucleotide divergence between haplotypes of the same isolate. Comparison of these haplotypes with those from other isolates suggests that some CTV isolates could have arisen as result of a mixed infection of two divergent isolates.

Base Sequence↗

Founder mitochondrial haplotypes in Amerindian populations.

It had been proposed that the colonization of the New World took place by three successive migrations from northeastern Asia. The first one gave rise to Amerindians (Paleo-Indians), the second and third ones to Nadene and Aleut-Eskimo, respectively. Variation in mtDNA has been used to infer the demographic structure of the Amerindian ancestors. The study of RFLP all along the mtDNA and the analysis of nucleotide substitutions in the D-loop region of the mitochondrial genome apparently indicate that most or all full-blooded Amerindians cluster in one of four different mitochondrial haplotypes that are considered to represent the founder maternal lineages of Paleo-Indians. We have studied the mtDNA diversity in 109 Amerindians belonging to 3 different tribes, and we have reanalyzed the published data on 482 individuals from 18 other tribes. Our study confirms the existence of four major Amerindian haplotypes. However, we also found evidence supporting the existence of several other potential founder haplotypes or haplotype subsets in addition to the four ancestral lineages reported. Confirmation of a relatively high number of founder haplotypes would indicate that early migration into America was not accompanied by a severe genetic bottleneck.

Americas↗

Genomic analysis of breed composition and population structure in Montana composite cattle.

The Montana composite was developed in Brazil from crosses between Bos indicus and Bos taurus and structured into four biological types: Zebu (N), adapted taurine (A), British taurine (B), and continental taurine (C). This study aimed to characterize the genetic diversity and population structure of the Montana composite using genomic data through principal component analysis (PCA), admixture analysis, and Wright's FST statistic. The PCA revealed a clear separation between Bos indicus and Bos taurus groups, with Montana animals distributed in an intermediate position. The first two principal components explained 69.48% and 3.45% of the total variation, respectively. Supervised admixture estimates indicated a predominance of taurine contribution, with type A accounting for 34.47%, 52.64%, and 51.71% at K&#x2009;=&#x2009;4, 9, and 11, respectively. Increasing the ancestry resolution refined the contribution of individual founder breeds without changing the overall predominance of taurine ancestry. Comparisons between breed proportions obtained from pedigree and genomic data revealed significant differences, for most biological types and ancestry models (P&#x2009;<&#x2009;0.001), indicating that realized breed composition deviates from theoretical expectations. Estimates of genetic differentiation confirmed greater divergence between Zebu and taurine groups, as well as reduced distances among populations sharing common ancestry. Specific relationships were identified between the composite and some of its founder breeds, particularly Belmont Red, Senepol, and Tuli. Overall, the results demonstrate that the Montana composite has a complex genomic structure, with genomic ancestry varying according to the resolution adopted and differing from pedigree-based expectations.

Animals↗

Chompy: an infestation of MITE-like repetitive elements in the crocodilian genome.

Interspersed repeats are a major component of most eukaryotic genomes and have an impact on genome size and stability, but the repetitive element landscape of crocodilian genomes has not yet been fully investigated. In this report, we provide the first detailed characterization of an interspersed repeat element in any crocodilian genome. Chompy is a putative miniature inverted-repeat transposable element (MITE) family initially recovered from the genome of Alligator mississippiensis (American alligator) but also present in the genomes of Crocodylus moreletii (Morelet's crocodile) and Gavialis gangeticus (Indian gharial). The element has all of the hallmarks of MITEs including terminal inverted repeats, possible target site duplications, and a tendency to form secondary structures. We estimate the copy number in the alligator genome to be approximately 46,000 copies. As a result of their size and unique properties, Chompy elements may provide a useful source of genomic variation for crocodilian comparative genomics.

Alligators and Crocodiles↗

Phenotypic profiles and functional genomics in Alzheimer's disease and in dementia with a vascular component.

Alzheimer's disease (AD) and dementia with vascular component (DVC) are the most prevalent forms of dementia. Both clinical entities share many similarities, but they differ in major phenotypic and genotypic profiles as revealed by structural and functional genomics studies. Comparative phenotypic studies have identified significant differences in 25% of more than 100 parametric variables, including anthropometry, cardiovascular function, aortic atherosclerosis, brain atrophy, blood pressure, blood biochemistry, hematology, thyroid function, folate and vitamin B12 levels, brain hemodynamics and lymphocyte markers. The phenotypic profile of patients with DVC differs from that of AD patients in the following: anthropometric values (weight, height); cardiovascular function (ECG, heart rate); blood pressure; lipid metabolism (HDL-CHO, TGs); uric acid metabolism; peripheral calcium homeostasis; liver function (GOT, GPT, GGT); alkaline phosphatase; lactate dehydrogenase; red and white blood cells; regional brain atrophy (left temporal region, inter-hippocampal distance); and left anterior blood flow velocity. Functional genomics studies incorporating APOE-related changes in biological markers extended the difference between AD and DVC up to 57%. Brain perfusion studies show a severe brain hypoperfusion in dementia associated with enlarged age-dependent arterial perfusion times. Structural genomics studies with AD-related genes, including APP, MAPT, APOE, PS1, PS2, A2M, ACE, AGT, cFOS and PRNP genes, demonstrate different genetic profiles in AD and DVC, with an absolute genetic variation rate ranging from 30% to 80%, depending upon genes and genetic clusters. Single gene analysis identifies relative genetic variations ranging from 0% to 5%. The relative polymorphic variation in genetic clusters integrated by two, three or four genes associated with AD ranges from 1% to 3%. The main phenotypic differences between AD and DVC are genotype-dependent, especially in AD, probably indicating that different genomic factors are determinant for the expression of dementia symptoms which might be accelerated or induced by environmental and/or cerebrovascular factors.

Alzheimer Disease↗