Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

A personalized multi-platform assessment of somatic mosaicism in the human frontal cortex.

Somatic mutations in individual cells create genomic mosaicism, influencing genetic disorders and cancers. While clonal mutations in cancers are well-studied, rarer somatic variants in normal tissues remain poorly characterized. This study systematically evaluates detection methods using a personalized donor-specific assembly (DSA) from a neurotypical individual's dorsolateral prefrontal cortex assessed with Oxford Nanopore, NovaSeq, linked-read sequencing, Cas9-targeted long-read sequencing (TEnCATS), and single-neuron MALBAC amplification. The haplotype-resolved DSA improved cross-platform analysis, dramatically increasing phasing rates. Germline SNVs, structural variations (SVs), and transposable elements (TEs) were recalled with 99.4%-99.7% accuracy in bulk tissue, and phased haplotype analysis reduced false positives by 15.4%-75.1% for putative somatic candidates. Long-read single-neuron sequencing detected nine somatic SV candidates, demonstrating enhanced sensitivity for rare variants, while TEnCATS identified eight low-frequency somatic TE candidates. These findings highlight advanced methodologies for precise somatic variant detection, critical for understanding mosaicism's role in health and disease.

Multi-platform Sequencing↗

Efficient reconstruction of haplotype structure via perfect phylogeny.

Each person's genome contains two copies of each chromosome, one inherited from the father and the other from the mother. A person's genotype specifies the pair of bases at each site, but does not specify which base occurs on which chromosome. The sequence of each chromosome separately is called a haplotype. The determination of the haplotypes within a population is essential for understanding genetic variation and the inheritance of complex diseases. The haplotype mapping project, a successor to the human genome project, seeks to determine the common haplotypes in the human population. Since experimental determination of a person's genotype is less expensive than determining its component haplotypes, algorithms are required for computing haplotypes from genotypes. Two observations aid in this process: first, the human genome contains short blocks within which only a few different haplotypes occur; second, as suggested by Gusfield, it is reasonable to assume that the haplotypes observed within a block have evolved according to a perfect phylogeny, in which at most one mutation event has occurred at any site, and no recombination occurred at the given region. We present a simple and efficient polynomial-time algorithm for inferring haplotypes from the genotypes of a set of individuals assuming a perfect phylogeny. Using a reduction to 2-SAT we extend this algorithm to handle constraints that apply when we have genotypes from both parents and child. We also present a hardness result for the problem of removing the minimum number of individuals from a population to ensure that the genotypes of the remaining individuals are consistent with a perfect phylogeny. Our algorithms have been tested on real data and give biologically meaningful results. Our webserver (http://www.cs.columbia.edu/compbio/hap/) is publicly available for predicting haplotypes from genotype data and partitioning genotype data into blocks.

Adult↗

The population biology and evolutionary significance of Ty elements in Saccharomyces cerevisiae.

The basic structure and properties of Ty elements are considered with special reference to their role as agents of evolutionary change. Ty elements may generate genetic variation for fitness by their action as mutagens, as well as by providing regions of portable homology for recombination. The mutational spectra generated by Ty1 transposition events may, due to their target specificity and gene regulatory capabilities, possess a higher frequency of adaptively favorable mutations than spectra resulting from other types of mutational processes. Laboratory strains contain between 25-35 elements, and in both these and industrial strains the insertions appear quite stable. In contrast, a wide variation in Ty number is seen in wild isolates, with a lower average number/genome. Factors which may determine Ty copy number in populations include transposition rates (dependent on Ty copy number and mating type), and stabilization of Ty elements in the genome as well as selection for and against Ty insertions in the genome. Although the average effect of Ty transpositions are deleterious, populations initiated with a single clone containing a single Ty element steadily accumulated Ty elements over 1,000 generations. Direct evidence that Ty transposition events can be selectively favored is provided by experiments in which populations containing large amounts of variability for Ty1 copy number were maintained for approximately 100 generations in a homogeneous environment. At their termination, the frequency of clones containing 0 Ty elements had decreased to approximately 0.0, and the populations had became dominated by a small number of clones containing > 0 Ty elements. No such reduction in variability was observed in populations maintained in a structured environment, though changes in Ty number were observed. The implications of genetic (mating type and ploidy) changes and environmental fluctuations for the long-term persistence of Ty elements within the S. cerevisiae species group are discussed.

Biological Evolution↗

Whole-genome analysis of Oryza sativa reveals similar architecture of two-component signaling machinery with Arabidopsis.

The two-component system (TCS), which works on the principle of histidine-aspartate phosphorelay signaling, is known to play an important role in diverse physiological processes in lower organisms and has recently emerged as an important signaling system in plants. Employing the tools of bioinformatics, we have characterized TCS signaling candidate genes in the genome of Oryza sativa L. subsp. japonica. We present a complete overview of TCS gene families in O. sativa, including gene structures, conserved motifs, chromosome locations, and phylogeny. Our analysis indicates a total of 51 genes encoding 73 putative TCS proteins. Fourteen genes encode 22 putative histidine kinases with a conserved histidine and other typical histidine kinase signature sequences, five phosphotransfer genes encoding seven phosphotransfer proteins, and 32 response regulator genes encoding 44 proteins. The variations seen between gene and protein numbers are assumed to result from alternative splicing. These putative proteins have high homology with TCS members that have been shown experimentally to participate in several important physiological phenomena in plants, such as ethylene and cytokinin signaling and phytochrome-mediated responses to light. We conclude that the overall architecture of the TCS machinery in O. sativa and Arabidopsis thaliana is similar, and our analysis provides insights into the conservation and divergence of this important signaling machinery in higher plants.

Amino Acid Sequence↗

Molecular analysis of the proviral DNA of equine infectious anemia virus in mules in Greece.

Molecular analysis of the regulatory and structurally important genetic segments of equine infectious anemia virus (EIAV) in mules is presented. We have previously reported clinicopathological and laboratory findings in mules infected with EIAV, both naturally and after experimental inoculation. In this study the fragment coding for integrase, gp90, tat and the fusion domain of gp45 of the proviral genome from these animals was sequenced and compared with one another and with that of EIAV strains already published in the literature. Significant variations were observed mainly in the sequences of the gp90 surface protein. In the two wild type sequences, there were substitutions in the V5 hypervariable domain of this protein. In the sequences of the experimentally inoculated animals and the donor strain, variations were due to insertions/duplications in the V3 principal neutralizing domain (PND) and substitutions in the V5 hypervariable domain. Finally, when compared with the already published strains, the wild type sequences had single amino acid substitutions across the whole protein and multiple substitutions in the V4-V6 variable domains. In general, the two Greek wild type sequences were closer to two of the American strains (WSU5 and Massachusetts), than to the two Japanese (V26 and V70) or the third American strain (Wyoming_wi) used in this study.

Amino Acid Sequence↗

Evolution of myelin proteolipid proteins: gene duplication in teleosts and expression pattern divergence.

The coevolution of neurons and their supporting glia to the highly specialized axon-myelin unit included the recruitment of proteolipids as neuronal glycoproteins (DMbeta, DMgamma) or myelin proteins (DMalpha/PLP/DM20). Consistent with a genome duplication at the root of teleosts, we identified three proteolipid pairs in zebrafish, termed DMalpha1 and DMalpha2, DMbeta1 and DMbeta2, DMgamma1 and DMgamma2. The paralogous amino acid sequences diverged remarkably after gene duplication, indicating functional specialization. Each proteolipid has adopted a distinct spatio-temporal expression pattern in neural progenitors, neurons, and in glia. DMalpha2, the closest homolog to mammalian PLP/DM20, is coexpressed with P0 in oligodendrocytes and upregulated after optic nerve lesion. DMgamma2 is expressed in multipotential stem cells, and the other four proteolipids are confined to subsets of CNS neurons. Comparing protein sequences and gene structures from birds, teleosts, one urochordate species, and four invertebrates, we have reconstructed major steps in the evolution of proteolipids.

Amino Acid Sequence↗

Genomic diversity of Erwinia carotovora subsp. carotovora and its correlation with virulence.

We used genetic and biochemical methods to examine the genomic diversity of the enterobacterial plant pathogen Erwinia carotovora subsp. carotovora. The results obtained with each method showed that E. carotovora subsp. carotovora strains isolated from one ecological niche, potato plants, are surprisingly diverse compared to related pathogens. A comparison of 23 partial mdh sequences revealed a maximum pairwise difference of 10.49% and an average pairwise difference of 2.13%, values which are much greater than the maximum variation (1.81%) and average variation (0.75%) previously reported for Escherichia coli. Pulsed-field gel electrophoresis analysis of I-CeuI-digested genomic DNA revealed seven rrn operons in all E. carotovora subsp. carotovora strains examined except strain WPP17, which had only six copies. We identified 26 I-CeuI restriction fragment length polymorphism patterns and observed significant polymorphism in fragment sizes ranging from 100 to 450 kb for all strains. We detected large plasmids in two strains, including the model strain E. carotovora subsp. carotovora 71. The two least virulent strains had an unusual chromosomal structure, suggesting that a particular pulsotype is correlated with virulence. To compare chromosomal organization of multiple enterobacterial genomes, several genes were mapped onto I-CeuI fragments. We identified portions of the genome that appear to be conserved across enterobacteria and portions that have undergone genome rearrangements. We found that the least virulent strain, WPP17, failed to oxidize cellobiose and was missing several hrp and hrc genes. The unexpected variability among isolates obtained from clonal hosts in one region and in one season suggests that factors other than the host plant, potato, drive the evolution of this common environmental bacterium and key plant pathogen.

Aconitate Hydratase↗

Genomic location and variation of the gene for CRS, a complement binding protein in the M57 strains of Streptococcus pyogenes.

All isolates of serotype M1 of group A streptococci possess a gene for streptococcal inhibitor of complement (SIC) in the mga regulon, which harbors genes for other virulence factors, such as M and M-like proteins, C5a peptidase, and a regulator. In serotype M57 the gene for a protein that is closely related to SIC (crs57) is located outside the mga regulon. We mapped the location of the crs57 gene in six strains of emm57 (gene encoding the M57 protein) sequence types to an intergenic region between the ABC transporter gene (SPy0778) and the gene for a small ribosomal protein (rpsU). The noncoding sequences on both sides of crs57 exhibited high degrees of identity to the corresponding regions of sic from M1 strains. This included one of the inverted repeat sequences of IS1562 but not the insertion element itself. These observations suggest that crs57 was recently acquired by serotype M57 or its progenitor via horizontal acquisition from serotype M1. The six emm57 sequence type isolates analyzed in this study belong to two distinct molecular types (vir types VT8 and VT101). Although the crs57 sequences from VT8 strains had very few substitution mutations, the VT101 crs57 sequence had a large number of such mutations. The CRS57 proteins from these strains are secretory products and have the ability to bind to complement proteins. All these proteins contain several tryptophan-rich repeats designated DWS motifs and internal repeat sequences. In all of these structural and biochemical characteristics CRS57 resembles SIC from M1 strains. Hence, CRS57 has a functional role similar to that of SIC in an M1 strain.

Amino Acid Sequence↗

Characterization of the preprotein and amino acid transporter gene family in Arabidopsis.

Seventeen loci encode proteins of the preprotein and amino acid transporter family in Arabidopsis (Arabidopsis thaliana). Some of these genes have arisen from recent duplications and are not in annotated duplicated regions of the Arabidopsis genome. In comparison to a number of other eukaryotic organisms, this family of proteins has greatly expanded in plants, with 24 loci in rice (Oryza sativa). Most of the Arabidopsis and rice genes are orthologous, indicating expansion of this family before monocot and dicot divergence. In vitro protein uptake assays, in vivo green fluorescent protein tagging, and immunological analyses of selected proteins determined either mitochondrial or plastidic localization for 10 and six proteins, respectively. The protein encoded by At5g24650 is targeted to both mitochondria and chloroplasts and, to our knowledge, is the first membrane protein reported to be targeted to mitochondria and chloroplasts. Three genes encoded translocase of the inner mitochondrial membrane (TIM)17-like proteins, three TIM23-like proteins, and three outer envelope protein16-like proteins in Arabidopsis. The identity of Arabidopsis TIM22-like proteins is most likely a protein encoded by At3g10110/At1g18320, based on phylogenetic analysis, subcellular localization, and complementation of a yeast (Saccharomyces cerevisiae) mutant and coexpression analysis. The lack of a preprotein and amino acid transporter domain in some proteins, localization in mitochondria, plastids, or both, variation in gene structure, and the differences in expression profiles indicate that the function of this family has diverged in plants beyond roles in protein translocation.

Amino Acid Sequence↗

Developmental roles for chromatin and chromosomal structure.

Chromosomal architecture is emerging as a key controlling influence in the developmental regulation of gene expression. Recent genetic experiments using Caenorhabditis elegans, Drosophila melanogaster, and the mouse have provided clear evidence for the functional differentiation of chromosomal structures during development. Chromosomes are visualized as highly specialized entities, within which the activity of particular domains is largely determined by defined structural proteins. At a more local level, the mechanisms regulating gene transcription during early embryogenesis in Xenopus and the mouse have been found to be dependent on the biochemical composition of individual nucleosomes. Thus, variation in the type and modification of chromosomal and chromatin structural proteins provides a dominant means of controlling the transcriptional activity of individual genes, individual chromosomal domains, and of entire chromosomes.

Animals↗

Can composition and structural features of oligonucleotides contribute to their wide-scale applicability as random PCR primers in mapping bacterial genome diversity?

Among current genotypic methodologies, random amplification of polymorphic DNA (RAPD or AP-PCR) represents a widely employed assay for the evaluation of bacterial genomic diversity. A common bottleneck of this technique, however, is represented by the screening of useful informative primers to discriminate among isolates of a particular bacterial species. In an attempt to simplify this process, we evaluated here the utility of degenerate oligonucleotides to act as informative AP-PCR primers. For this purpose, a number of features (G+C contents, degeneracy rate, modifications at the 5' end) of related degenerate primers was tested for their effects in the generation of informative arrays from a set of bacterial genomes. Our results indicate that a combination of a wide base composition and a common palindromic structure at the 5' end of the sequences that compose the degenerate primers tested here beneficially resulted for the generation of informative arrays aimed to evaluate the bacterial genome heterogeneity.

5' Flanking Region↗

Structure and polymorphism of the Chironomus thummi gene encoding special lobe-specific silk protein, ssp160.

cDNA encoding Chironomus thummi ssp160 was used to isolate a genomic clone that hybridized in situ to band A2b on polytene chromosome IV, the site of the ssp160 gene. DNA sequencing, primer extension and gene/cDNA nucleotide sequence alignment revealed the gene contains six exons and five introns; 70% of ssp160 is encoded in exon 3. Variations between cDNA and gene sequences led to the design of a polymerase chain reaction, restriction fragment length polymorphism assay that was subsequently used to demonstrate the existence of polymorphic alleles whose distribution varied between geographically separated populations of larvae. The polymorphism is associated with codon deletions in a six-amino-acid repeat containing an N-linked glycosylation motif. These deletions may have resulted from slipped-strand mispairing during DNA replication.

Amino Acid Sequence↗

Molecular determinants and guided evolution of species-specific RNA editing.

Most RNA editing systems are mechanistically diverse, informationally restorative, and scattershot in eukaryotic lineages. In contrast, genetic recoding by adenosine-to-inosine RNA editing seems common in animals; usually, altering highly conserved or invariant coding positions in proteins. Here I report striking variation between species in the recoding of synaptotagmin I (sytI). Fruitflies, mosquitoes and butterflies possess shared and species-specific sytI editing sites, all within a single exon. Honeybees, beetles and roaches do not edit sytI. The editing machinery is usually directed to modify particular adenosines by information stored in intron-mediated RNA structures. Combining comparative genomics of 34 species with mutational analysis reveals that complex, multi-domain, pre-mRNA structures solely determine species-appropriate RNA editing. One of these is a previously unreported long-range pseudoknot. I show that small changes to intronic sequences, far removed from an editing site, can transfer the species specificity of editing between RNA substrates. Taken together, these data support a phylogeny of sytI gene editing spanning more than 250 million years of hexapod evolution. The results also provide models for the genesis of RNA editing sites through the stepwise addition of structural domains, or by short walks through sequence space from ancestral structures.

Adenosine↗

Nucleotide sequences of pigeon feather keratin genes.

We analyzed two pigeon feather keratin clones from a cosmid pigeon genomic library. Each of the clones contained three feather keratin genes that had the same general structure: a 5' non-coding region separated by an intron, a protein-coding region encoding a protein of 100 amino acids, and a 3' non-coding region. Length and transcriptional organization of the genes were variable. The length variation, about 1.2-3.7 kb, was mainly due to the difference in the length of the 3' non-coding region, and the longer genes had opposite transcriptional organization in contrast to the shorter genes. The nucleotide sequences of the coding region were very similar among the six genes but not the same.

Amino Acid Sequence↗

Studies of the structure and functional organization of foreign DNA integrated into the genome of Nicotiana tabacum.

In transgenic plants obtained either by Agrobacterium tumefaciens-mediated transformation or by direct DNA transfer, the structure of integrated chimeric donor DNA remains stable during vegetative proliferation, during sexual transmission, and under various selection conditions. We correlate the level of expression of the introduced gene in independent transformants and their offspring with the particular arrangement and modification of their integrated DNAs. Genetic analysis of transgenic plants shows that the chimeric gene is transmitted in a Mendelian fashion to the F1 and F2 progeny as a single dominant trait. Deviations from the expected segregation pattern are discussed with respect to different levels of gene activity. We compare the gene activity in heterozygotes versus homozygotes, and show variation in activity between plants regenerated independently from the same transformed callus. Cotransformation studies with two physically unlinked and partly homologous plasmids carrying two different marker genes indicate that they are physically linked after integration into the host genome.

Gene Expression Regulation↗

Identification of genetic variants in the neuronal form of tryptophan hydroxylase (TPH2).

OBJECTIVE: We screened the complete protein coding sequence of the newly identified neuronal form of tryptophan hydroxylase (TPH2) for genetic variants. METHODS: Genomic DNA samples from 24 African-Americans and 24 Caucasian-Americans in the Coriell human variation collection were screened by denaturing high-performance liquid chromatography followed by sequencing. RESULTS: We identified a number of genetic variants in both the coding and exon-flanking intronic sequences. Only one variant was identified that predicts a structural change in the TPH2 protein, and this was seen in only one out of 96 chromosomes. CONCLUSIONS: The gene for TPH2 contains a number of polymorphisms that might serve as useful markers for association analyses of complex behavioral phenotypes or as actual risk factors. Structural polymorphisms are extremely rare in TPH2 and cannot therefore act as substantial risk factors for behavioral disorders in African-American and Caucasian populations.

Black or African American↗

Cyanobacterial peptides - nature's own combinatorial biosynthesis.

Cyanobacterial secondary metabolites have attracted increasing scientific interest due to bioactivity of many compounds in various test systems. Among the known structures, oligopeptides are often found with many congeners sharing conserved substructures, while being highly variable in others. A major part of known oligopeptides are of non-ribosomal origin and can be grouped into classes with conserved structural properties. Thus, the overall structural diversity of cyanobacterial oligopeptides only seemingly suggests an equally high diversity of biosynthetic pathways and respective genes. For each class of peptides, some of which have been found in all major branches of the cyanobacterial evolutionary tree, homologous synthetases and genes can be inferred. This implies that non-ribosomal peptide synthetase genes are a very ancient part of the cyanobacterial genome and presumably have evolved by recombination and duplication events to reach the present structural diversity of cyanobacterial oligopeptides. In addition, peptide synthetases would appear to be an essential part of the cyanobacterial evolution and physiology. The present review presents an overview of the biosynthesis of cyanobacterial peptides and corresponding gene clusters, the structural diversity of structural types and structural variations within peptide classes, and implications for the evolution and plasticity of biosynthetic genes and the potential function of cyanobacterial peptides.

Cyanobacteria↗

Extended repertoire of genes encoding variable surface lipoproteins in Mycoplasma bovis strains.

A genomic cluster of vsp genes was previously shown to mediate high-frequency phenotypic switching of surface lipoprotein antigens in the bovine pathogen Mycoplasma bovis. This study revealed that field strains of M. bovis possess modified versions of the vsp gene complex in which extensive sequence variations occur primarily in the reiterated coding sequences of the vsp structural genes. These findings demonstrate that there is a vastly expanded potential for antigenic variation within populations of this organism.

Amino Acid Sequence↗