Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

Quantitative mutant analysis of viral quasispecies by chip-based matrix-assisted laser desorption/ ionization time-of-flight mass spectrometry.

RNA viruses exist as quasispecies, heterogeneous and dynamic mixtures of mutants having one or more consensus sequences. An adequate description of the genomic structure of such viral populations must include the consensus sequence(s) plus a quantitative assessment of sequence heterogeneities. For example, in quality control of live attenuated viral vaccines, the presence of even small quantities of mutants or revertants may indicate incomplete or unstable attenuation that may influence vaccine safety. Previously, we demonstrated the monitoring of oral poliovirus vaccine with the use of mutant analysis by PCR and restriction enzyme cleavage (MAPREC). In this report, we investigate genetic variation in live attenuated mumps virus vaccine by using both MAPREC and a platform (DNA MassArray) based on matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) mass spectrometry. Mumps vaccines prepared from the Jeryl Lynn strain typically contain at least two distinct viral substrains, JL1 and JL2, which have been characterized by full length sequencing. We report the development of assays for characterizing sequence variants in these substrains and demonstrate their use in quantitative analysis of substrains and sequence variations in mixed virus cultures and mumps vaccines. The results obtained from both the MAPREC and MALDI-TOF methods showed excellent correlation. This suggests the potential utility of MALDI-TOF for routine quality control of live viral vaccines and for assessment of genetic stability and quantitative monitoring of genetic changes in other RNA viruses of clinical interest.

Animals↗

Genome structure and divergence of nucleotide sequences in echinodermata.

The arrangement of repetitive and single-copy DNA sequences has been studied in DNA of some species of Echinodermata--sea urchin, starfishes and sea-cucumber. Comparison of the reassociation kinetics of short and long DNA fragments indicates that the pattern of DNA sequence organization of all these species is similar to the so-called "Xenopus pattern" characteristic of the genomes of most animals and plants. However, substantional variations have been found in the amount of repetitive nucleotide sequences in DNA of different species and in the length of DNA regions containing adjacent single-copy and repetitive sequences. Measurements of the size of S1-nuclease resistant reassociated repetitive DNA sequences show a variability of ratios between long and short repetitive DNA sequences of different species.--The degree of divergence of short and long repetitive DNA sequences and single-copy DNA was studied by molecular hybridization of the sea urchin Strongylocentrotus intermedius 3H-DNA with the DNA of other species and by determination of the thermostability of the hybridized molecules so obtained. All three fractions of S. intermedius DNA contain sequences homologous to DNA of the other echinoderm species studied. The results obtained suggest that short repetitive DNA sequences are those which have been most highly conserved throughout the evolution of Echinodermata. A new hypothesis is proposed to explain the nature of the evolutionary changes in DNA sequence interspersion patterns.

Animals↗

The genetical history of humans and the great apes.

When and where did modern humans evolve? How did our ancestors spread over the world? Traditionally, answers to questions such as these have been sought in historical, archaeological, and fossil records. However, increasingly genetic data provide information about the evolution of our species. In this review, we focus on the comparison of the variation in the human gene pool to that of our closest evolutionary relatives, the great apes, because this provides a relevant perspective on human genetical evolution. For instance, comparisons to the great apes show that humans are unique in having little genetic variation as well as little genetic structure in their gene pool. Furthermore, genetic data indicate that humans, but not the great apes, have experienced a period of dramatic growth in their early history.

Animals↗

The Role of Small Segmental Duplications in Generating Identical Isoforms Through Alternative Splicing Sites.

Alternative splicing plays a crucial role in expanding proteomic diversity but can also generate identical isoforms under certain conditions. While mutually exclusive splicing of tandem exons has occasionally been reported to produce identical isoforms, the extent to which other splicing events contribute to this phenomenon remains unclear. In this study, we demonstrate that alternative 5' and 3' splice site selection can also lead to the formation of identical isoforms, providing an additional type of splicing event for functional redundancy in transcriptomes. To address this, we analyzed reference genome annotations from 15 plant species, including Arabidopsis thaliana and wheat (Triticum aestivum), obtained from the RefSeq database. Identical isoforms were computationally defined as transcripts with distinct exon-intron structures but identical coding sequences. Our analysis reveals that the majority of alternative 5' and 3' fragments originate from small segmental duplications, suggesting that sequence repetition within gene regions facilitates the emergence of such splicing patterns. We also observed differences in the annotated 5' UTRs of some identical isoforms. However, since the alternative splicing sites themselves were not located within UTRs, these differences may reflect annotation uncertainty rather than genuine AS-derived variation. Given that UTR predictions in reference databases are not always precise, such observations should be interpreted cautiously. Expression analysis using an isoform-specific k-mer approach confirmed that identical isoforms can be differentially regulated. These findings suggest that, beyond expanding protein diversity, alternative splicing can also generate redundant isoforms that are differentially expressed at the RNA level, indicating potential regulatory roles. By elucidating the structural and regulatory factors contributing to the formation and retention of identical isoforms, our study provides new insights into the evolutionary and functional significance of alternative splicing in plants.

Alternative Splicing↗

Immune escape by hepatitis B viruses.

Hepatitis B viruses are DNA viruses characterized by their very small genome size and their unique replication via reverse transcription. The circular genome has been efficiently exploited, thereby limiting genome variation, and leaves no space for genes in addition to those essentially needed during the viral live cycle. Hepatitis B viruses are prototype non-cytopathic viruses causing persistent infection. Human hepatitis B virus (HBV), as well as the closely related animal viruses, most frequently are transmitted vertically from mothers to their offspring. Because infection usually persists for many years, if not lifelong, hepatitis B viruses need efficient mechanisms to hide from the immune response of the host. To escape the immune response, they exploit different strategies. Firstly, they use their structural and non-structural proteins multiplely. One of the purposes is to alter the immune response. Secondly, they replicate by establishing a pool of stable extrachromosomal transcription templates, which allow the virus to react sensitively to changes in its microenvironment by up- or downregulating gene expression. Thirdly, hepatitis B viruses replicate in the liver which is an immunopriviledged site.

Animals↗

The genome of Entamoeba histolytica.

Estimation of genome size of Entamoeba histolytica by different methods has failed to give comparable values due to the inherent complexities of the organism, such as the uncertain level of ploidy, presence of multinucleated cells and a poorly demarcated cell division cycle. The genome of E. histolytica has a low G+C content (22.4%), and is composed of both linear chromosomes and a number of circular plasmid-like molecules. The rRNA genes are located exclusively on some of the circular DNAs. Karyotype analysis by pulsed field gel electrophoresis suggests the presence of 14 conserved linkage groups and an extensive size variation between homologous chromosomes from different isolates. Several repeat families have been identified, some of which have been shown to be present in all the electrophoretically separated chromosomes. The typical nucleosomal structure has not been demonstrated, though most of the histone genes have been identified. Most Entamoeba genes lack introns, have short 3' and 5' untranslated regions, and are tightly packed. Promoter analysis revealed the presence of three conserved motifs and several upstream regulatory elements. Unlike typical eukaryotes, the transcription of protein coding genes is alpha-amanitin resistant. Expressed Sequence Tag analysis has identified a group of highly abundant polyadenylated RNAs which are unlikely to be translated. The Expressed Sequence Tag approach has also helped identify several important genes which encode proteins that may be involved in different biochemical pathways, signal transduction mechanisms and organellar functions.

Animals↗

The distance between bacterial species in sequence space.

Despite the revolution caused by information from macromolecular sequences, the basis of bacterial classification remains the genus and the species. How do these terms relate to the variety of bacteria that exist on earth? In this paper, the inter- and intraspecies differences in amino acid sequence of several bacterial electron transport proteins, cytochromes c, and blue copper proteins are compared. For the soil and water organisms studied, bacterial species can be classed as "tight" when there is little intraspecies variation, or "loose" when this variation is large. For this set of proteins and organisms, interspecies variation is much larger than that within a species. Examples of "tight" species are Pseudomonas aeruginosa and Rhodobacter sphaeroides, while Pseudomonas stutzeri and Rhodopseudomonas palustris are loose species. The results are discussed in the context of the origin and age of bacterial species, and the distribution of genomes in "sequence space." The situation is probably different for commensal or pathogenic bacteria, whose population structure and evolution are linked to the properties of another organism.

Amino Acid Sequence↗

The LQT syndromes--current status of molecular mechanisms.

Our knowledge on the molecular genetics of inherited cardiac arrhythmias is very recent in comparison to the advances of genetics achieved in other inherited cardiac disorders. This is related to the high mortality and early disease onset of these arrhythmias resulting in mostly small nucleus families. Thus, traditional genetic linkage studies that are based on the genetic information obtained from large multi-generation families were made difficult. In 1991, the first chromosomal locus for congenital long-QT (LQT) syndrome was identified on chromosome 11p15.5 (LQT1 locus) by linkage analysis. Meanwhile, the disease-causing gene at the LQT1 locus (KCNQ1), a gene encoding a K+ channel subunit of the IKs channel, and three other, major genes, all encoding cardiac ion channel components, have been identified. Taken together, LQT syndrome turned out to be a heterogeneous channelopathy. Moreover, the power of linkage studies to reveal the genetic causes of the LQT syndrome was also important to identify unknown but fundamental channel components that contribute to the ion currents tuning ventricular repolarization. In-vitro expression of the altered ion channel genes demonstrated in each case that the altered ion channel function produces prolongation of the action potential and thus the increasing propensity to ventricular tachyarrhythmias. Since these ion channels are pharmacological targets of many antiarrhythmic (and other) drugs, individual and potentially deleterious drug responses may be related to genetic variation in ion channel genes. Very recently, also in acquired LQT syndrome, which is a frequent clinical disorder in cardiology a genetic basis has been proposed in part since mutations in LQT genes have been specifically found. The discovery of ion channel defects in LQT syndrome represents the major achievement in our understanding and implies potential therapeutic options. The knowledge of the genomic structure of the LQT genes now offers the possibility to detect the underlying genetic defect in 80-90% of all patients. With this specific information, containing the type of ion channel (Na+ versus K+ channel) and electrophysiological alteration by the mutation (loss-of-function versus change-of-function mutation), gene-directed, elective drug therapies have been initiated in genotyped LQT patients. Based on preliminary data, that were supported by in vitro studies, this approach may be useful in recompensating the characteristic phenotypes in some LQT patients. Mutation detection is a new diagnostic tool which may become of more increasing importance in patients with a normal QTc or just a borderline prolongation of the QTc interval at presentation. These patients represent approximately 40% of all familial cases. Moreover, LQT3 syndrome and idiopathic ventricular fibrillation are allelic disorders and genetically overlap. In both mutations in the LQT3 gene SCN5A encoding the Na+ channel alpha-subunit for INa have been reported. Thus, the clinical nosology of inherited arrhythmias may be reconsidered after elucidation of the underlying molecular bases. Meanwhile, genotype-phenotype correlations in large families are on the way to evaluate intergene, interfamilial, and intrafamilial differences in the clinical phenotype reflecting gene specific, gene-site specific, and individual consequences of a given mutation. LQT syndrome is phenotypically heterogeneous due to the reduced penetrance and variable expressivity associated with the mutations. This paper discusses the current data on molecular genetics and genotype-phenotype correlations and the implications for diagnosis and treatment.

Chromosome Mapping↗

Solution structure of the X4 protein coded by the SARS related coronavirus reveals an immunoglobulin like fold and suggests a binding activity to integrin I domains.

The SARS related Coronavirus genome contains a variety of novel accessory genes. One of these, called ORF7a or ORF8, code for a protein, known as 7a, U122 or X4. We set out to determine the three-dimensional structure of the soluble ectodomain of this type-I transmembrane protein by nuclear magnetic resonance spectroscopy. The fold of the protein is the first member of a further variation of the immunoglobulin like beta-sandwich fold. Because X4 does not reveal significant sequence homologies to proteins in the data bases, we carried out a structure based similarity search for proteins with known function. High structural similarity to Dl domains of ICAM-1 and ICAM-2, and common features in amino acid sequence between X4 and ICAM-1, suggest X4 to possess binding activity for the alpha(L) integrin I domain of LFA-1. Further, based on this structure based prediction, potential functions of X4 in virus replication and pathogenesis are discussed.

Amino Acid Sequence↗

The artificial worlds approach to emergent evolution.

Artificial worlds models of evolutionary systems are computer models that map the essential logical structure of ecological systems, defined as self-sustaining biological organizations. The artificial world comprises an artificial environment, with mass components, energy input, and physical states. It also comprises artificial organisms, including a genome, a phenome, and a (developmental) map that connects the genome to the phenome. Mass components are cycled and space is limited. The evolution process results, as in nature, from genetic variation combined with natural selection imposed by the finiteness of the environment. The selection criteria (fitness values) are not imposed, but rather emerge from the interactions of the organisms with each other and with the environment. The dynamics at the population level also emerges from these basic interactions. In this paper we describe the comparative properties of the EVOLVE family of artificial worlds models.

Artificial Intelligence↗

Genetics and evolution of multilocus isozymes in hexaploid wheat.

Aneuploid genetic studies of isozyme variation in cv Chinese Spring have disclosed that numerous enzymes of hexaploid wheat exist in multiple molecular forms as a direct consequence of polyploidy. Sixty-nine isozyme structural genes have been identified to date. Two of these belong to a duplicate set and at least 54 to triplicate sets of paralogous genes that are located one each in related chromosomes in different genomes. Each of these gene sets encodes either two or three isozymes. The role of regional gene duplication in the production of multilocus isozymes in hexaploid wheat is as yet poorly understood, although a considerable amount of indirect evidence suggests that a large number of isozymes are encoded by genes that were produced by ancient regional gene duplication events in a genome ancestral to the genomes now present in the species. A full assessment of the role of regional gene duplication in the production of hexaploid wheat isozymes must await further studies. The isozyme structural gene locations thus far determined indicate that the gene synteny relationships that existed in the ancestral wheat genome are in large part conserved in each of the three genomes of cv Chinese Spring and that the genetic content of most individual chromosome arms has also been in large part conserved.

Biological Evolution↗

Fine-scale map of encyclopedia of DNA elements regions in the Korean population.

The International HapMap Project aims to generate detailed human genome variation maps by densely genotyping single-nucleotide polymorphisms (SNPs) in CEPH, Chinese, Japanese, and Yoruba samples. This will undoubtedly become an important facility for genetic studies of diseases and complex traits in the four populations. To address how the genetic information contained in such variation maps is transferable to other populations, the Korean government, industries, and academics have launched the Korean HapMap project to genotype high-density Encyclopedia of DNA Elements (ENCODE) regions in 90 Korean individuals. Here we show that the LD pattern, block structure, haplotype diversity, and recombination rate are highly concordant between Korean and the two HapMap Asian samples, particularly Japanese. The availability of information from both Chinese and Japanese samples helps to predict more accurately the possible performance of HapMap markers in Korean disease-gene studies. Tagging SNPs selected from the two HapMap Asian maps, especially the Japanese map, were shown to be very effective for Korean samples. These results demonstrate that the HapMap variation maps are robust in related populations and will serve as an important resource for the studies of the Korean population in particular.

Asian People↗

Nucleotide diversity and haplotype structure of the human angiotensinogen gene in two populations.

Variation in the angiotensinogen gene, AGT, has been associated with variation in plasma angiotensinogen levels. In addition, the T235M polymorphism in the AGT product is associated with an increased risk of essential hypertension in multiple populations, making AGT a good example of a quantitative-trait locus underlying susceptibility to a common disease. To better understand genetic variation in AGT, we sequenced a 14.4-kb genomic region spanning the entire AGT and identified 44 single-nucleotide polymorphisms (SNPs). Forty-two SNPs were observed both in 88 white and in 77 Japanese unselected subjects. Six major haplotypes accounted for most of the variation in this region, indicating less allelic complexity than in many other genomic regions. Although the two populations were found to share all of the major AGT haplotypes, there were substantial differences in haplotype frequencies. Pairwise linkage disequilibrium (LD), measured by the D', r(2), and d(2) statistics, demonstrated a general pattern of decline with increasing distance, but, as expected in a small genomic region, individual LD values were highly variable. LD between T235M and each of the other 39 SNPs was assessed in order to model the usefulness of LD to detect a disease-associated mutation. Among the Japanese subjects, 13 (33%) of the SNPs had r(2) values >0.1, whereas this statistic was substantially higher for the white subjects (occurring in 35/39 [90%]). LD between a hypertension-associated promoter mutation, A-6G, and 39 SNPs was also measured. Similar results were obtained, with 33% of the SNPs showing r(2)>0.1 in the Japanese subjects and 92% of the SNPs showing r(2)>0.1 in the white subjects. This difference, which occurs despite an overall similarity in LD patterns in the two populations, reflects a much higher frequency of the M235-associated haplotype in the white sample. These results have important implications for the usefulness of LD approaches in the mapping of genes underlying susceptibility to complex diseases.

Angiotensinogen↗

The evolutionary history of the common chloroplast genome of Arabidopsis thaliana and A. suecica.

The evolutionary history of the common chloroplast (cp) genome of the allotetraploid Arabidopsis suecica and its maternal parent A. thaliana was investigated by sequencing 50 fragments of cpDNA, resulting in 98 polymorphic sites. The variation in the A. suecica sample was small, in contrast to that of the A. thaliana sample. The time to the most recent common ancestor (T(MRCA)) of the A. suecica cp genome alone was estimated to be about one 37th of the T(MRCA) of both the A. thaliana and A. suecica cp genomes. This corresponds to A. suecica having a MRCA between 10 000 and 50 000 years ago, suggesting that the entire species originated during, or before, this period of time, although the estimates are sensitive to assumptions made about population size and mutation rate. The data was also consistent with the hypothesis of A. suecica being of single origin. Isolation-by-distance and population structure in A. thaliana depended upon the geographical scale analysed; isolation-by-distance was found to be weak on the global scale but locally pronounced. Within the genealogical cp tree of A. thaliana, there were indications that the root of the A. suecica species is located among accessions of A. thaliana that come primarily from central Europe. Selective neutrality of the cp genome could not be rejected, despite the fact that it contains several completely linked protein-coding genes.

Arabidopsis↗

Structure of a human histone cDNA: evidence that basally expressed histone genes have intervening sequences and encode polyadenylylated mRNAs.

We have isolated and sequenced full-length cDNA clones encoding the human basally expressed H3.3 histone from a human fibroblast cDNA library. Several features of this atypical cDNA distinguish it and its gene from the well-characterized cell-cycle regulated histone genes and their RNA transcripts. The H3.3 mRNA is approximately equal to 1200 bases long, contains unusually long 5' and 3' untranslated regions, and has a 3' polyadenylylated terminus. In addition, we have isolated and characterized a cDNA clone that is a precursor to the H3.3 mRNA and contains an intervening sequence interrupting its 5' untranslated region. Hybridization of subsegments of the cDNA to human genomic DNA reveals a complex multigene family. The differences in the structures of basal and cell-cycle histone genes suggest a model to explain the differences in their expression.

Base Sequence↗

Genetic diversity and relationships of Campylobacter species and subspecies.

The existence of tremendous genetic diversity within Campylobacter species has been well documented. To analyse the population structure of Campylobacter and determine whether or not a clonal population structure could be detected, genetic diversity was assessed within the genus Campylobacter by multilocus enzyme electrophoresis of 156 isolates representing 11 species and subspecies from disparate sources. Analyses of electrophoretic mobility of 11 enzymes revealed 109 electrophoretic types (ETs) and 118 ETs when nulls were counted as an allele. Cluster analysis placed most ETs into groups that correlated with species. With nulls counted as alleles, 19 ETs were identified among 33 isolates of Campylobacter lari, 31 ETs among 34 isolates of Campylobacter coli and 43 ETs among 59 isolates of Campylobacter jejuni subsp. jejuni. Nine C. jejuni subsp. jejuni isolates, confirmed as this species by DNA-DNA hybridization, were hippuricase-negative. Reported linkage analyses were done with nulls ignored. Scores for mean genetic diversity (H) were high for the total population (mean H = 0.802). Allelic mismatch-frequency distributions and allelic tracing pointed to possible genetic exchange between subpopulations. C. lari appears to be a panmictic species. Some pairs of species shared multiple alleles of certain loci, possibly indicating genetic exchange between species. Of the species tested, C. jejuni appeared to be the most active in sharing alleles. However, there was evidence of variable involvement in recombination by the different loci. Linkage analysis of loci in C. jejuni and C. coli revealed a clonal framework, with some loci tightly linked to each other. The loci appeared to occur in linkage groups or islands. Campylobacter may have a clonal framework with other portions of the genome involved in frequent recombination. Population genetic structure among Campylobacter is inconclusive and it remains to be seen if pathogenic types can be identified.

Alleles↗

Increased virulence of a mouse-adapted variant of influenza A/FM/1/47 virus is controlled by mutations in genome segments 4, 5, 7, and 8.

To cause disease, influenza virus must possess several genetically determined abilities that mediate stages in pathogenesis. The virulent mouse-adapted variant A/FM/1/47-MA (FM-MA), derived from the avirulent A/FM/1/47 (FM) strain, had acquired mutations in genes that control virulence. The purpose of this study was to identify those genes that had mutated to result in increased virulence and to obtain viruses that differed in virulence because of differences in individual genome segments. The genes that had mutated to increase virulence were initially identified by genetic analysis of reassortants obtained by crossing FM-MA with the avirulent strain A/HK/1/68 (HK). FM-MA genome segments 4, 5, 7, and 8 were significantly associated with virulence, as determined by using the Wilcoxon ranked sum analysis. The role of FM-MA segments 4, 7, and 8 was confirmed by reintroduction of these genes into the parental strain, which also provided virus strains that differed in virulence because of mutations in individual genome segments. Segments 4, 7, and 8 were responsible for a 10(3.6)-fold increase in virulence that was proportioned 10(2.2)-, 10(0.7)-, and 10(0.8)-fold, respectively. The role of segment 5 could not be confirmed on transfer back into the parental strain because of reversion during preparation of such reassortants. The incidence of reversion was shown to be significantly associated with culturing of FM-MA in chicken embryo cells but was not associated with growth in MDCK cells. The genetic analysis of FM-MA suggests that adaptation to increased virulence is an incremental process that involves the acquisition of mutations in multiple genes, each of which plays an individual role in pathogenesis. The structural and functional properties of segments 4, 7, and 8 that control the virulence of FM-MA can now be determined by using viruses that differ in virulence because of mutations in these individual genome segments.

Animals↗

How is it that microsatellites and random oligonucleotides uncover DNA fingerprint patterns?

Minisatellites, microsatellites, and short random oligonucleotides all uncover highly polymorphic DNA fingerprint patterns in Southern analysis of genomic DNA that has been digested with a restriction enzyme having a 4-bp specificity. The polymorphic nature of the fragments is attributed to tandem repeat number variation of embedded minisatellite sequences. This explains why DNA fingerprint fragments are uncovered by minisatellite probes, but does not explain how it is that they are also uncovered by microsatellite and random oligonucleotide probes. To clarify this phenomenon, we sequenced a large bovine genomic BamHI restriction fragment hybridizing to the Jeffreys 33.6 minisatellite probe and consisting of small and large Sau3A-resistant subfragments. The large Sau3A subfragment was found to have a complex architecture, consisting of two different minisatellites, flanked and separated by stretches of unique DNA. The three unique sequences were characterized by sequence simplicity, that is, a higher than chance occurrence of tandem or dispersed repetition of simple sequence motifs. This complex repetitive structure explains the absence of Sau3A restriction sites in the large Sau3A subfragment, yet provides this subfragment with the ability to hybridize to a variety of probe sequences. It is proposed that a large class of interspersed tracts sharing this complex yet simplified sequence structure is found in the genome. Each such tract would have a broad ability to hybridize to a variety of probes, yet would exhibit a dearth of restriction sites. For each restriction enzyme having 4-bp specificity, a subclass of such tracts, completely lacking the corresponding restriction sites, will be present. On digestion with the given restriction enzyme, each such tract would form a large fragment.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals↗