Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Underlying regularity in the shapes of nucleoids of Escherichia coli: implications for nucleoid organization and partition.

The genomic DNA of Escherichia coli is localized in one or a few compact nucleoids. Nucleoids in rapidly grown cells appear in complex shapes; the relationship of these shapes to underlying arrangements of the DNA is of structural interest and of potential importance in gene localization and nucleoid partition studies. To help assess this variation in shape, limited three-dimensional information on individual nucleoids was obtained by DNA fluorescence microscopy of cells as they reoriented in solution or by optical sectioning. These techniques were also applied to enlarged nucleoids within swollen cells or spheroplasts. The resulting images indicated that much of the apparent variation was due to imaging from different directions and at different focal planes of more regular underlying nucleoid shapes. Nucleoid images could be transformed into compact doublet shapes by exposure of cells to chloramphenicol or puromycin, consistent with a preexisting bipartite nucleoid structure. Isolated nucleoids and nucleoids in stationary-phase cells also assumed a doublet shape, supporting such a structure. The underlying structure is suggested to be two subunits joined by a linker. Both the subunits and the linker appear to deform to accommodate the space available within cells or spheroplasts ("flexible doublet" model).

Chloramphenicol↗

Yeast-based functional genomics and proteomics technologies: the first 15 years and beyond.

Yeast-based functional genomics and proteomics technologies developed over the past decade have contributed greatly to our understanding of bacterial, yeast, fly, worm, and human gene functions. In this review, we highlight some of these yeast-based functional genomic and proteomic technologies that are advancing the utility of yeast as a model organism in molecular biology and speculate on their future uses. Such technologies include use of the yeast deletion strain collection, large-scale determination of protein localization in vivo, synthetic genetic array analysis, variations of the yeast two-hybrid system, protein microarrays, and tandem affinity purification (TAP)-tagging approaches. The integration of these advances with established technologies is invaluable in the drive toward a comprehensive understanding of protein structure and function in the cellular milieu.

Forecasting↗

The ribosomal DNA loci in Plasmodium falciparum accumulate mutations independently.

Homogeneity of rDNA sequence within a cell is maintained by mechanisms working at the DNA level. The imperative to maintain homogeneity is thought to result from pressure to maintain the sequence of the rRNA transcript. We have investigated the extent of sequence variation within and between members of a species that is unable to utilize some standard mechanisms of rDNA sequence correction. We have compared the sequence of the internal transcribed spacer (ITS1) located between the 18 S rRNA and 5.8 S rRNA genes of five different loci of a single Plasmodium falciparum genotype. The ITS1 sequences are identical at 80 to 91% of the positions among the three asexually expressed genes (A-types) and 75% between the two genes expressed during sporogony (S-types), with only 42 to 57% identity between the types. This is rather startling in that the differences described here for a single genome are greater than those normally seen when comparing rDNA units from distantly related organisms. We observe an apparent conservation of secondary structure within ITS1 sequences from the different transcription units, which would reflect a level of selection at the rRNA but the organism seems to be quite tolerant of primary sequence variation. Investigation of the mature coding region within the 18 S rRNA genes did not reveal sequence variation within A- and S-types from a single genotype. However, comparison of the 18 S rRNA coding region from 17 geographically distinct strains reveals up to 10% sequence variation within a 400 nucleotide region. Hence homogeneity of rRNA units within a species does not seem to be an imperative driven totally by selection at the RNA level. The extraordinary maintenance of homogeneity within rDNA units normally seen within a species appears to have significance beyond those that can be ascribed to the events involved in processing, assembly and function of the ribosome.

Animals↗

Tracking the invasive history of the green alga Codium fragile ssp. tomentosoides.

The spread of nonindigenous species into new habitats is having a drastic effect on natural ecosystems and represents an increasing threat to global biodiversity. In the marine environment, where data on the movement of invasive species is scarce, the spread of alien seaweeds represents a particular problem. We have employed a combination of plastid microsatellite markers and DNA sequence data from three regions of the plastid genome to trace the invasive history of the green alga Codium fragile ssp. tomentosoides. Extremely low levels of genetic variation were detected, with only four haplotypes present in the species' native range in Japan and only two of these found in introduced populations. These invasive populations displayed a high level of geographical structuring of haplotypes, with one haplotype localized in the Mediterranean and the other found in Northwest Atlantic, northern European and South Pacific populations. Consequently, we postulate that there have been at least two separate introductions of C. fragile ssp. tomentosoides from its native range in the North Pacific.

Base Sequence↗

The genetic structure of tetraploid Avena: a comparison of isozyme and RAPD markers.

Isozymes were the first widely used molecular markers in plant population analysis. They yielded valuable information on the amount and the structure of genetic variability. DNA technology has provided new types of markers based on DNA sequence, which make it possible to study polymorphisms in a much greater proportion of the genome. This is the reason why the use of isozymes is less popular nowadays. This effect would be justified if all markers provided the same type of information on polymorphism and genetic relationships among populations; otherwise, it would be necessary to use different markers to obtain the complete picture of the genetic structure of populations and species. In this study, we compared data of isozyme and RAPD markers in the populations of two tetraploid species of wild oats: Avena barbata populations collected in Argentina, and Avena murphyi populations collected in Spain and Morocco. The samples were evaluated for 9 isozymatic systems and 10 primers. The structure of genetic variability was studied using Nei's method, and the relationships between populations were estimated using Hedrick and Jaccard's similarities for isozymes and RAPDs, respectively. As expected, RAPDs were more polymorphic than isozymes, but the information obtained from both markers was weakly correlated. The various reasons for this observation are discussed, but our conclusion is that in order to study the structure of genetic variability, several types of markers should be used.

Argentina↗

The Human Genome Project--an overview.

The human genome sequence will underpin human biology and medicine in the next century, providing a single, essential reference to all genetic information. The international program to determine the complete DNA sequence (3,000 million bases) is well underway. As of January 2000, 50% of the sequence is available in the public domain. A comprehensive working draft is expected this year, and the entire sequence is projected to be finished in 2003. DNA sequencing is carried out on mapped, overlapping bacterial clones of 150-200 kb. The working draft comprises assembled unfinished sequence and is released immediately in the public domain. The draft sequence of each clone is then completed, by closing any remaining gaps and resolving any ambiguities, before the entire sequence is checked, annotated, and submitted to the public databases. The sequence of each clone is finished to an accuracy of >99.99%. The availability of a reference sequence of the genome provides the basis for studying the nature of sequence variation, particularly single nucleotide polymorphisms (SNPs), in human populations. SNP typing is a powerful tool for genetic analysis, and will enable us to uncover the association of loci at specific sites in the genome with many disease traits. SNPs occur at a frequency of approximately 1 SNP/kb throughout the genome when the sequence of any two individuals is compared. Programs to detect and map SNPs in the human genome are underway with the aim of establishing a SNP map of the genome during the next two years. The human genome sequence will provide a complete description of all the genes. Annotation of the sequence with the gene structures is achieved by a combination of computational analysis (predictive and homology-based) and experimental confirmation by cDNA sequencing. Detecting homologies between newly defined gene products and proteins of known function helps to postulate biochemical functions for them, which can then be tested. Establishing the association of specific genes with disease phenotypes by mutation screening, particularly for monogenic disorders, provides further assistance in defining the functions of some gene products, as well as helping to establish the cause of the disease. As our knowledge of gene sequences and sequence variation in populations increases, we will pinpoint more and more of the genes and proteins that are important in common, complex diseases. A more detailed understanding of the function of the human genome will be achieved as we identify sequences that control gene expression. Given the availability of gene sequences, the expression status of genes in particular tissues can be monitored in parallel. By comparing corresponding genomic sequences in different species (for example: man, mouse, chicken, and zebrafish), regions that have been highly conserved during evolution can be identified, many of which reflect conserved functions such as gene regulation. These approaches promise to greatly accelerate our interpretation of the human genome sequence.

Human Genome Project↗

Alu RNA transcripts in human embryonal carcinoma cells. Model of post-transcriptional selection of master sequences.

Alu master sequences colonized the human genome using RNA as amplification intermediate. To understand this phenomenon better we isolated and analyzed Alu RNA from NTera2D1 pluripotential cells. Northern hybridization, primer extension, cDNA cloning and sequencing data are congruent and demonstrate a low level of Alu specific transcription. These bona fide RNA Polymerase III Alu transcripts, although enriched in the cytoplasm, are not dominated by a single master species but rather originate from a variety of loci. However, when compared with the genomic average, or to repeats from RNA Polymerase II co-transcripts, they belong to the youngest group of Alu subfamilies (p less than 0.001) and have a higher content of intact CpG-dinucleotides. This suggests that Alu transcription is influenced both by mutations and the genomic context, and points to a possible role of DNA methylation in silencing the bulk of genomic repeats. Because of the heterogeneity of Alu transcripts a post-transcriptional selection mechanism recruiting Alu master sequences for retroposition is required. We propose that Alu RNA masters could have evolved as selfish satellites to a more complex retroposition system equipped with a reverse transcriptase activity and that their structure was conserved through "phenotypic" selection of the RNA level.

Base Sequence↗

The regulation of albumin and alpha-fetoprotein gene expression in mammals.

Albumin and alpha-fetoprotein (AFP) are two plasma proteins synthesized by the liver and the yolk sac. The production of these major proteins is subject to considerable and characteristic variations during both the course of development and hepatic carcinogenesis. It is therefore a system of choice for the analysis of genetic expression during normal differentiation and the cancerous state of eukaryotic cells. The knowledge of regulatory mechanisms at the cellular and molecular levels of the albumin and AFP genes has recently made great progress: 1) the cells which are responsible for the synthesis of albumin and AFP in the liver and other organs have been defined by conjointly using in vitro and in vivo molecular hybridization techniques; 2) the organization of these genes and their adjoining regions has been established in the rat, the mouse and man; 3) the level at which the synthesis of these two proteins is regulated has been determined; it is the transcriptional level. The transcriptional regulation of the albumin and AFP genes could be the result of genome and/or chromatin conformation level modifications. Different groups have shown that: 1) the global structure of the albumin and AFP genes does not change during the course of development and hepatic carcinogenesis; 2) modifications at the level of the methylation of certain specific cytosines could be associated with the variations in the transcription of these genes; 3) global or local (hypersensitive sites with DNase I) changes of chromatin conformation could be correlated to the potential or the overt activity of the transcription of these genes. Very recently certain 'regulatory' regions having cis 'enhancer' or 'silencer' properties have been detected upstream from the albumin and AFP genes. These regions are hypothesized to be DNA 'target' sequences on which trans-acting regulatory factors are fixed and which control the transcription of these genes. Starting from the framework of this recent work, a model of albumin and AFP gene regulation is proposed.

Albumins↗

HapScope: a software system for automated and visual analysis of functionally annotated haplotypes.

We have developed a software analysis package, HapScope, which includes a comprehensive analysis pipeline and a sophisticated visualization tool for analyzing functionally annotated haplotypes. The HapScope analysis pipeline supports: (i) computational haplotype construction with an expectation-maximization or Bayesian statistical algorithm; (ii) SNP classification by protein coding change, homology to model organisms or putative regulatory regions; and (iii) minimum SNP subset selection by either a Brute Force Algorithm or a Greedy Partition Algorithm. The HapScope viewer displays genomic structure with haplotype information in an integrated environment, providing eight alternative views for assessing genetic and functional correlation. It has a user-friendly interface for: (i) haplotype block visualization; (ii) SNP subset selection; (iii) haplotype consolidation with subset SNP markers; (iv) incorporation of both experimentally determined haplotypes and computational results; and (v) data export for additional analysis. Comparison of haplotypes constructed by the statistical algorithms with those determined experimentally shows variation in haplotype prediction accuracies in genomic regions with different levels of nucleotide diversity. We have applied HapScope in analyzing haplotypes for candidate genes and genomic regions with extensive SNP and genotype data. We envision that the systematic approach of integrating functional genomic analysis with population haplotypes, supported by HapScope, will greatly facilitate current genetic disease research.

Algorithms↗

Fungal transposable elements and genome evolution.

The transposable elements (TEs) identified in fungal genomes reflect the whole spectrum of eukaryotic transposable elements. Most of our knowledge comes from species representing different ecological situations: plant pathogens, industrial, and field strains, most of them lacking the sexual stage. A number of changes in gene structure and function has been shown to be TE-mediated: inactivation of gene expression upon insertion within or adjacent to a gene, DNA sequence variation through excision and probably extensive chromosomal rearrangements due to recombination between members of a particular family. Moreover, TEs may have other roles in evolution related to their ability to be horizontally transferred and to capture and transpose chromosomal host sequences, thus providing a mechanism for dispersing sequences to new sites. However, the activity of transposable elements and consequently their proliferation within a host genome can be affected, in some fungal species which undergo meiosis, by silencing processes. Our understanding of the biological effects of TEs on the fungal genome has increased dramatically in the past few years but elucidation of the extent to which transposons contribute to genetic variation in nature, providing the flexibility for populations to adapt successfully to environmental changes is an important area for future research.

DNA Transposable Elements↗

Integration of GWAS and WGCNA reveals novel candidate genes for cottonseed oil content in Gossypium hirsutum L.

Genetic improvement of cottonseed oil content represents a crucial strategy for enhancing the comprehensive utilization of cotton. Here, genome-wide association study (GWAS) and weighted gene co-expression network analysis (WGCNA) were integrated to elucidate the genetic control underlying oil content. Phenotypic evaluation of 159 cotton accessions revealed extensive genetic variation, with kernel oil content ranging from 17.81% to 39.50%. Population structure analysis based on 20,213 single nucleotide polymorphisms (SNPs) classified the accessions into two major subpopulations. A total of 18 SNPs exhibited significant associations with oil content, two of which were stably detected across multiple environments using the FarmCPU model. Further haplotype analysis within linkage disequilibrium (LD) blocks confirmed a favorable haplotype on chromosome A05 that was strongly correlated with elevated oil content. Integration of publicly available transcriptome data from 11 ovule developmental stages with WGCNA identified modules significantly linked to oil content. Of the 74 candidate genes within LD intervals, 17 were assigned to WGCNA modules. Functional annotation and enrichment analyses highlighted four putative candidate genes (GH_A05G1503, GH_A05G1506, GH_A05G1531, and GH_A10G2150) involved in oil biosynthesis. These findings deepen our understanding of the genetic mechanisms governing cottonseed oil biosynthesis and lay a foundation for breeding high-oil cotton varieties.

Gossypium↗

Phylogeny and natural history of the primate lentiviruses, SIV and HIV.

Studies of primate lentivirus phylogeny over the past decade have established a minimum of five related, but genetically distinct, groups of simian immunodeficiency virus (SIV), each originating from a different African primate species. The hypothesis that HIV-2 (and SIVmac) arose by cross-species transmission from sooty mangabeys (Cercocebus atys has been strengthened by a more detailed characterization of the SIVsm/SIVmac/HIV-2 group of viruses. SIV from all four subspecies of African green monkeys (SIVagm) have been characterized with an apparent chimeric genome structure of SIVagm from West African green monkeys. Although these naturally infected primates remain healthy, cross-species transmission to other primate species may result in immunodeficiency, as caused by SIVsm infection of macaque monkeys (Macaca sp.) and recently, SIVagm infection of pig-tailed macaques (M. nemestrina). Studies of variation within infected individuals have been facilitated by adaptation of the techniques of heteroduplex analysis and single-stranded conformational polymorphism of PCR generated fragments.

Africa↗

Defining the fold space of membrane proteins: the CAMPS database.

Recent progress in structure determination techniques has led to a significant growth in the number of known membrane protein structures, and the first structural genomics projects focusing on membrane proteins have been initiated, warranting an investigation of appropriate bioinformatics strategies for optimal structural target selection for these molecules. What determines a membrane protein fold? How many membrane structures need to be solved to provide sufficient structural coverage of the membrane protein sequence space? We present the CAMPS database (Computational Analysis of the Membrane Protein Space) containing almost 45,000 proteins with three or more predicted transmembrane helices (TMH) from 120 bacterial species. This large set of membrane proteins was subjected to single-linkage clustering using only sequence alignments covering at least 40% of the TMH present in a given family. This process yielded 266 sequence clusters with at least 15 members, roughly corresponding to membrane structural folds, sufficiently structurally homogeneous in terms of the variation of TMH number between individual sequences. These clusters were further subdivided into functionally homogeneous subclusters according to the COG (Clusters of Orthologous Groups) system as well as more stringently defined families sharing at least 30% identity. The CAMPS sequence clusters are thus designed to reflect three main levels of interest for structural genomics: fold, function, and modeling distance. We present a library of Hidden Markov Models (HMM) derived from sequence alignments of TMH at these three levels of sequence similarity. Given that 24 out of 266 clusters corresponding to membrane folds already have associated known structures, we estimate that 242 additional new structures, one for each remaining cluster, would provide structural coverage at the fold level of roughly 70% of prokaryotic membrane proteins belonging to the currently most populated families.

Bacterial Proteins↗

Evolution of transglutaminase genes: identification of a transglutaminase gene cluster on human chromosome 15q15. Structure of the gene encoding transglutaminase X and a novel gene family member, transglutaminase Z.

We isolated and characterized the gene encoding human transglutaminase (TG)(X) (TGM5) and mapped it to the 15q15.2 region of chromosome 15 by fluorescence in situ hybridization. The gene consists of 13 exons separated by 12 introns and spans about 35 kilobases. Further sequence analysis and mapping showed that this locus contained three transglutaminase genes arranged in tandem: EPB42 (band 4.2 protein), TGM5, and a novel gene (TGM7). A full-length cDNA for the novel transglutaminase (TG(Z)) was obtained by anchored polymerase chain reaction. The deduced amino acid sequence encoded a protein with 710 amino acids and a molecular mass of 80 kDa. Northern blotting showed that the three genes are differentially expressed in human tissues. Band 4.2 protein expression was associated with hematopoiesis, whereas TG(X) and TG(Z) showed widespread expression in different tissues. Interestingly, the chromosomal segment containing the human TGM5, TGM7, and EPB42 genes and the segment containing the genes encoding TG(C),TG(E), and another novel gene (TGM6) on chromosome 20q11 are in mouse all found on distal chromosome 2 as determined by radiation hybrid mapping. This finding suggests that in evolution these six genes arose from local duplication of a single gene and subsequent redistribution to two distinct chromosomes in the human genome.

5' Untranslated Regions↗

The evolution of Ca2+-ATPases across plants with profiles in Rhododendron and the function of key members in alleviating high calcium stress.

Ca2+-ATPase (CAP) is a key Ca2+ efflux protein in plants. Our previous research suggests that CAPs may play a crucial role in the adaptation of rhododendrons to high calcium environments. However, the evolution, variation, characteristic expression, and subfunctionalization of this gene family in Rhododendron remain unknown. Through the analysis of pan-genomes and pan-transcriptomes, we elucidated the systematic evolution of CAPs in plants, as well as their characteristic expression patterns in Rhododendron. During the evolutionary process from lower to higher plants, CAPs can be divided into six clades and exhibit structural conservation. CAPs have emerged and differentiated in lower plants such as algae, and they have undergone significant amplification in Eudicots plants like rhododendrons. Three Rhododendron species (Rhododendron bailiense, R. delavayi, and R. irroratum) located in the karst province of Guizhou in Southwest China exhibit the highest copies of CAPs, suggesting a strong association between CAP copy number variation and habitat, particularly in high calcium environments. Through multiple transcriptome analyses, we revealed that CAPs are induced under various environmental/developmental conditions (e.g. karst environments, high altitude, early flower development, hormones, etc.). Co-expression network analysis highlighted key members of calcineurin B-like protein (CBL) and CBL-interacting protein kinases (CIPK) that are associated with the high expression of CAPs. Experimental validation demonstrated that CAPb1 and CAPd1 significantly alleviate high calcium stress, and the CAPb1-CIPK1-CBL1 and CAPd1-CIPK2-CBL1 modules can further enhance the alleviation. These findings provide new insights into the evolution, characteristic expression, and function of CAPs, as well as new perspectives on the high calcium adaptability of rhododendrons.

Journal Article↗

Functional variability of Rev response element in HIV-1 primary isolates.

We have previously studied sequence heterogeneity of HIV-1 Rev response element (RRE), and showed uneven variations in different stem-loops of both primary sequence and secondary structure. Here we studied the functional variation of RRE clones from a set of 10 primary isolates, and demonstrated a variation in the function of these RRE clones on the expression of Gag proteins from a truncated HIV-1 genome. The difference in Gag level was, in part, if not exclusively, resulted from the differential efficiency of RNA transport and enhancing of translation. These data suggested that variation of HIV-1 RRE may play a role in regulation of viral replication rate in HIV-1 primary isolates.

Base Sequence↗

Evidence for multiple hybrid groups in Trypanosoma cruzi.

A role for parasite genetic variability in the spectrum of Chagas disease is emerging but not yet evident, in part due to an incomplete understanding of the population structure of Trypanosoma cruzi. To investigate further the observed genotypic variation at the sequence and chromosomal levels in strains of standard and field-isolated T. cruzi we have undertaken a comparative analysis of 10 regions of the genome from two isolates representing T. cruzi I (Dm28c and Silvio X10) and two from T. cruzi II (CL Brener and Esmeraldo). Amplified regions contained intergenic (non-coding) sequences from tandemly repeated genes. Multiple nucleotide polymorphisms correlated with the T. cruzi I/T. cruzi II classification. Two intergenic regions had useful polymorphisms for the design of classification probes to test on genomic DNA from other known isolates. Two adjacent nucleotide polymorphisms in HSP 60 correlated with the T. cruzi I and T. cruzi II distinction. 1F8 nucleotide polymorphisms revealed multiple subdivisions of T. cruzi II: subgroups IIa and IIc displayed the T. cruzi I pattern; subgroups IId and IIe possessed both the I and II patterns. Furthermore, isolates from subgroups IId and IIe contained the 1F8 polymorphic markers on different chromosome bands supporting a genetic exchange event that resulted in chromosomes V and IX of T. cruzi strain CL Brener. Based on these analyses, T. cruzi I and subgroup IIb appear to be pure lines, while subgroups IIa/IIc and IId/IIe are hybrid lines. These data demonstrate for the first time that IIa/IIc are hybrid, consistent with the hypothesis that genetic recombination has occurred more than once within the T. cruzi lines.

Animals↗

Compositional structure of repetitive elements is quantitatively related to co-expression of gene pairs.

A sequence similarity metric operating on 10 kb upstream regions of gene pairs quantitatively predicts a portion of co-variation of expression of gene pairs in large-scale gene expression studies in human tumors and tumor-derived cell lines. The signal on which the metric depends most strongly originates in the compositional structure of repetitive genomic sequences (particularly Alu elements) present in these upstream regions. This effect is completely separable from effects of isochore composition on gene expression. The results implicate repetitive elements with some functional role in transcriptional regulation of the specific genes in whose promoter regions they reside and lend credence to suggestions that the general phenomenon of repetitive element insertions may be a fundamental evolutionary mechanism for modulating gene transcription.

Base Composition↗