Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequence diversity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Sequence diversity of Treponema pallidum subsp. pallidum tprK in human syphilis lesions and rabbit-propagated isolates.

The tprK gene of Treponema pallidum subsp. pallidum, the causative agent of venereal syphilis, belongs to a 12-member gene family and encodes a protein with a predicted cleavable signal sequence and predicted transmembrane domains. Except for the Nichols type strain, all rabbit-propagated isolates of T. pallidum examined thus far are comprised of mixed populations of organisms with heterogeneous tprK sequences. We show that tprK sequences in treponemes obtained directly from syphilis patients are also heterogeneous. Clustering analysis demonstrates that primary chancre tprK sequences are more likely to cluster within a sample than among samples and that tighter clustering is seen within chancre samples than within rabbit-propagated isolates. Closer analysis of tprK sequences from a rabbit-propagated isolate reveals that individual variable regions have different levels of diversity, suggesting that variable regions may have different intrinsic rates of sequence change or may be under different levels of selection. Most variable regions show increased sequence diversity upon passage. We speculate that the diversification of tprK during infection allows organisms to evade the host immune response, contributing to reinfection and persistent infection.

Amino Acid Sequence↗

Sequence variability in the 5' non-coding region of hepatitis C virus: identification of a new virus type and restrictions on sequence diversity.

We have analysed the pattern of nucleotide sequence variability in the 5' non-coding region (5' NCR) of geographically dispersed variants of hepatitis C virus (HCV). Phylogenetic analysis of sequences in this region indicated the existence of a new virus type, provisionally termed type 4, the identity of which was confirmed by further analysis of the more variable part of the HCV core protein coding region. The geographical distribution of HCV type 4 was distinct from that of other HCV types, it being particularly widespread in Africa and absent or rare in Europe and the Far East. Much of the variability in the 5' NCR appears to be constrained by a requirement for specific secondary structures in the viral RNA. In one of the most variable regions of the 5' NCR (positions -169 to -114), most of the nucleotide changes that are characteristic of different HCV types were covariant, with complementary substitutions at other positions. According to the proposed secondary structure of the 5' NCR, such changes preserved base pairing within a stem-loop structure, whereas the nucleotide insertions found in a proportion of 5' NCR sequences, including those of type 4, localized exclusively to the non-base-paired terminal loop. The specific nucleotide substitutions in the 5' NCR that differentiate each of the four HCV types can be detected by restriction enzyme cleavage, providing a rapid and reliable method for virus typing.

Amino Acid Sequence↗

Aryl hydrocarbon-induced interactions at multiple DNA elements of diverse sequence--a multicomponent mechanism for activation of cytochrome P4501A1 (CYP1A1) gene transcription.

In vivo footprinting experiments, augmented with gel shift and transfection analyses suggest that activation of the CYP1A1 gene by aryl hydrocarbons may be a multicomponent process. During the first 30 minutes of exposure to aryl hydrocarbon carcinogens and environmental contaminants, in vivo footprints appear at nine distinct sites within a 281 bp region centered 950 bp upstream of the CYP1A1 transcription start site. Six of these sites are unrelated in sequence to the three xenobiotic response elements (XREs) within this region, at which the aryl hydrocarbon (AH) receptor is known to bind. These six display a variety of footprint patterns, are diverse in sequence and range in G-C content from 60 to 75%. This diversity suggests that multiple nuclear factors may be responsible for these six in vivo footprints. These observations are consistent with competition gel shift experiments showing that the nuclear factors binding at two of these sites are different from each other, as well as from the AH receptor. Gel shifts also indicate that the sequence-specific factors binding at these sites are expressed constitutively. This is consistent with a model in which in vivo footprints are induced at these six sites, not through direct activation or de novo synthesis of DNA-binding factors, but through a two phase mechanism in which binding of the nuclear AH receptor complex to XREs facilitates the binding of constitutive factors at these sites. This facilitation could be mediated either through specific protein-protein interactions or through alterations in chromatin structure that make these sites accessible to constitutive nuclear factors. A function for the sequences at which aryl hydrocarbons induce in vivo footprints is suggested by transfection experiments showing that one of these sequences cooperates with a weak XRE to confer on a reporter gene responsiveness to aryl hydrocarbons.

Animals↗

Revisiting the rDNA sequence diversity of a natural population of the arbuscular mycorrhizal fungus Acaulospora colossica.

In 1999, the diversity of a field population of the arbuscular mycorrhizal (AM) fungus Acaulospora colossica was characterized using DNA sequence data. Since 1999, AM fungal sequences have accumulated rapidly within public databases. Moreover, novel phylogenetic tools have been developed and can be used to interpret the data. A second analysis of those sequences collected in 1999 demonstrates that while the majority of the sequences are, in fact, sequences of A. colossica; a minority of the sequences still cannot be identified with confidence. Those sequences identified as A. colossica can be used to show that (1) the nuclear rDNA ITS regions are remarkably diverse, and (2) sequences isolated from different spores of the same site may be more closely related to each other than to sequences of other sites, so that the genetic diversity of an AM fungal field population may be spatially structured; however, identical sequences can also be recovered from different sites.

DNA, Fungal↗

Interaction and sequence diversity among T15 VH genes in CBA/J mice.

Nucleotide sequences of the four genes composing the T15 heavy chain variable region (VH) family of the CBA/J mouse have been determined. Comparison of these sequences with their published BALB/c and C57BL/10 homologues reveals that nucleotide differences found between given alleles of two strains, i.e., CBA/J and BALB/c, are observed in other family members of the same strain. We suggest that these patterns of sequence variation are most readily explained by gene interaction (conversion). Additionally, the sequence of a CBA/J hybridoma, 6G6, proposed to have been generated by gene conversion, is directly encoded by the CBA/J V11 gene indicating that the putative conversion has occurred meiotically in the germline. These results are consistent with the premise that gene correction is occurring frequently among members of this family and that such processes may contribute significantly to the evolution of Ig variable region genes even in the relatively short time frame of inbred strain derivation.

Alleles↗

Lack of sequence diversity in the gene encoding merozoite surface protein 5 of Plasmodium falciparum.

The gene encoding merozoite surface protein 5 (MSP5) of Plasmodium falciparum is situated between the genes encoding MSP2 and MSP4 on chromosome 2. Both MSP4 and MSP5 encode proteins that contain hydrophobic signal and glycosylphosphatidylinositol (GPI) attachment signals and a single epidermal growth factor (EGF)-like domain at their carboxyl termini. The similar gene organization, location and similar structural features of the two genes suggest that they have arisen from a gene duplication event. In this study we provide further evidence for the merozoite surface location of MSP5 by demonstrating that MSP5 is present in isolated merozoites, partitions in the detergent-enriched phase following Triton X-114 fractionation and shows a staining pattern consistent with merozoite surface location by indirect immunofluorescence confocal microscopy. Analysis of antigenic diversity of MSP5 shows a lack of sequence variation between various isolates of P. falciparum from different geographical locations, a feature unusual for surface proteins of merozoites and one that may simplify vaccine formulation.

Amino Acid Sequence↗

Mitochondrial cytochrome b gene sequence diversity in the Korean hare, Lepus coreanus Thomas (Mammalia, Lagomorpha).

Partial sequences of the mitochondrial cytochrome b gene of the Korean hare (Lepus coreanus) were analyzed to determine the degree of genetic diversity. Nine haplotypes were observed, and the maximum Tamura-Nei nucleotide distance among them was 2.8%, indicating that genetic diversity of L. coreanus is moderate. In order to clarify the Korean hare's taxonomic status and relationship with the Manchurian hare (L. mandshuricus) and the Chinese hare (L. sinensis), these nine haplotypes of the Korean hare were compared with 13 haplotypes from five other species of eastern Asian Lepus including L. mandshuricus and L. sinensis. The Korean hare was distinct in its cytochrome b gene, and it is confirmed that L. coreanus is a valid species, as noted by Jones and Johnson (1965, Univ. Kansas Publ. (Mus. Nat. Hist.) 16:357). Further analyses of mtDNA cytochrome b gene with additional specimens of L. coreanus from North Korea and other species of Lepus from eastern Asia are needed to clarify the taxonomic status of the divergent mtDNA clades of L. mandshuricus and L. sinensis.

Animals↗

A computer program for the estimation of protein and nucleic acid sequence diversity in random point mutagenesis libraries.

A computer program for the generation and analysis of in silico random point mutagenesis libraries is described. The program operates by mutagenizing an input nucleic acid sequence according to mutation parameters specified by the user for each sequence position and type of point mutation. The program can mimic almost any type of random mutagenesis library, including those produced via error-prone PCR (ep-PCR), mutator Escherichia coli strains, chemical mutagenesis, and doped or random oligonucleotide synthesis. The program analyzes the generated nucleic acid sequences and/or the associated protein library to produce several estimates of library diversity (number of unique sequences, point mutations, and single point mutants) and the rate of saturation of these diversities during experimental screening or selection of clones. This information allows one to select the optimal screen size for a given mutagenesis library, necessary to efficiently obtain a certain coverage of the sequence-space. The program also reports the abundance of each specific protein mutation at each sequence position, which is useful as a measure of the level and type of mutation bias in the library. Alternatively, one can use the program to evaluate the relative merits of preexisting libraries, or to examine various hypothetical mutation schemes to determine the optimal method for creating a library that serves the screen/selection of interest. Simulated libraries of at least 10(9) sequences are accessible by the numerical algorithm with currently available personal computers; an analytical algorithm is also available which can rapidly calculate a subset of the numerical statistics in libraries of arbitrarily large size. A multi-type double-strand stochastic model of ep-PCR is developed in an appendix to demonstrate the applicability of the algorithm to amplifying mutagenesis procedures. Estimators of DNA polymerase mutation-type-specific error rates are derived using the model. Analyses of an alpha-synuclein ep-PCR library and NNS synthetic oligonucleotide libraries are given as examples.

Algorithms↗

Novel alpha-conotoxins identified by gene sequencing from cone snails native to Hainan, and their sequence diversity.

Conotoxins (CTX) from the venom of marine cone snails (genus Conus) represent large families of proteins, which show a similar precursor organization with surprisingly conserved signal sequence of the precursor peptides, but highly diverse pharmacological activities. By using the conserved sequences found within the genes that encode the alpha-conotoxin precursors, a technique based on RT-PCR was used to identify, respectively, two novel peptides (LiC22, LeD2) from the two worm-hunting Conus species Conus lividus, and Conus litteratus, and one novel peptide (TeA21) from the snail-hunting Conus species Conus textile, all native to Hainan in China. The three peptides share an alpha4/7 subfamily alpha-conotoxins common cysteine pattern (CCX(4)CX(7)C, two disulfide bonds), which are competitive antagonists of nicotinic acetylcholine receptor (nAChRs). The cDNA of LiC22N encodes a precursor of 40 residues, including a propeptide of 19 residues and a mature peptide of 21 residues. The cDNA of LeD2N encodes a precursor of 41 residues, including a propeptide of 21 residues and a mature peptide of 16 residues with three additional Gly residues. The cDNA of TeA21N encodes a precursor of 38 residues, including a propeptide of 20 residues and a mature peptide of 17 residues with an additional residue Gly. The additional residue Gly of LeD2N and TeA21N is a prerequisite for the amidation of the preceding C-terminal Cys. All three sequences are processed at the common signal site -X-Arg- immediately before the mature peptide sequences. The properties of the alpha4/7 conotoxins known so far were discussed in detail. Phylogenetic analysis of the new conotoxins in the present study and the published homologue of alpha4/7 conotoxins from the other Conus species were performed systematically. Patterns of sequence divergence for the three regions of signal, proregion, and mature peptides, both nucleotide acids and residue substitutions in DNA and peptide levels, as well as Cys codon usage were analyzed, which suggest how these separate branches originated. Percent identities of the DNA and amino acid sequences of the signal region exhibited high conservation, whereas the sequences of the mature peptides ranged from almost identical to highly divergent between inter- and intra-species. Notably, the diversity of the proregion was also high, with an intermediate percentage of divergence between that observed in the signal and in the toxin regions. The data presented are new and are of importance, and should attract the interest of researchers in this field. The elucidated cDNAs of these toxins will facilitate a better understanding of the relationship of their structure and function, as well as the process of their evolutionary relationships.

Amino Acid Sequence↗

Sequence diversity in the intron of the calmodulin gene from Plasmodium falciparum.

Sequence variation in the single intron of the calmodulin gene of Plasmodium falciparum has been examined following amplification using the polymerase chain reaction (PCR). The intron has 4 repeating motifs varying in length: 3 of these contain a dinunucleotide repeat, dA-dT, the fourth is a pentameric repeating unit, dA-dT-dA-dT-dT. These DNA polymorphisms can be applied to the study of parasite populations in mixed infections and in strain identification. The function of these repeating motifs is unknown. Computer modelling of the possible intron structures demonstrates that each of these repeating motifs forms individual stem loops such that any changes in repeat number are self-compensatory and do not change the overall intron structure. The implications of this sequence variation are discussed.

Animals↗

Demographic history of India and mtDNA-sequence diversity.

The demographic history of India was examined by comparing mtDNA sequences obtained from members of three culturally divergent Indian subpopulations (endogamous caste groups). While an inferred tree revealed some clustering according to caste affiliation, there was no clear separation into three genetically distinct groups along caste lines. Comparison of pairwise nucleotide difference distributions, however, did indicate a difference in growth patterns between two of the castes. The Brahmin population appears to have undergone either a rapid expansion or steady growth. The low-ranking Mukri caste, however, may have either maintained a roughly constant population size or undergone multiple bottlenecks during that period. Comparison of the Indian sequences to those obtained from other populations, using a tree, revealed that the Indian sequences, along with all other non-African samples, form a starlike cluster. This cluster may represent a major expansion, possibly originating in southern Asia, taking place at some point after modern humans initially left Africa.

Base Sequence↗

Sequence diversity analysis of dihydroflavonol 4-reductase intron 1 in common bean.

Variation in common bean (Phaseolus vulgaris L.) was investigated by sequencing intron 1 of the dihydroflavonol 4-reductase (DFR) gene for 92 genotypes that represent both landraces and cultivars. We were also interested in determining if introns provide sufficient variation for genetic diversity studies and if the sequence data could be used to develop allele-specific primers that could differentiate genotypes using a standard PCR assay. Sixty-nine polymorphic sites were observed. Nucleotide variation (pi/bp) was 0.0481, a value higher than that reported for introns from other plant species. Tests for significant deviation from the mutation drift model were positive for the population as a whole, the cultivar and landrace subsets, and the Middle American landrace set. Significant linkage disequilibrium extended about 300 nucleotides. Twenty haplotypes were detected among the cultivated genotypes. Seven recombination events were detected for the whole population, and six events for the landraces. Recombination was not observed among the landraces within either the Middle American or Andean gene pools. Evidence for hybridization between the two gene pools was discovered. Five allele-specific primers were developed that could distinguish 56 additional genotypes. The allele-specific primers were used to map duplicate DFR genes on linkage group B8.

Alcohol Oxidoreductases↗

Identification of genetically diverse sequences (ORF 5) of porcine reproductive and respiratory syndrome virus in a swine herd.

The ability of genetically diverse strains of porcine reproductive and respiratory syndrome virus (PRRSV) to coexist in a 1750-sow farm was assessed through the case study describing a chronically infected farm, and also by an animal experiment involving the use of swine bioassay. The case study employed a program of monitoring sera from suckling, nursery, and finishing pigs for the presence of PRRSV by polymerase chain reaction (PCR) and virus isolation (VI). The swine bioassay tested homogenates, consisting of lymphoid and pulmonary tissues, collected from 60 breeding animals from the same farm. The open reading frame (ORF) 5 portion of selected positive PRRSV detected from sera or tissues were nucleic acid sequenced and their phylogenies compared. The results indicated the presence of 3 genetically diverse groups, designated PRRSV-A, -B, and -C. Sequence heterology ranged from 5.8 to 11% between groups. Sequence homology ranged from 98.7 to 99.8% within groups. Swine bioassay verified the presence of PRRSV-A in 1 of 60 animals, and no evidence of strains B or C were detected. This paper indicates that based on the evaluation of ORF 5, genetically diverse strains of PRRSV appear to coexist, although the frequency and significance of this observation is not understood.

Amino Acid Sequence↗

Mitochondrial DNA sequence diversity in Russians.

The article presents the results of the first regular study of Russian populations by sequencing the control region of mitochondrial DNA (mtDNA). The sequenced region is the most variable on mtDNA molecule and is commonly used for population and evolutionary studies. Russians form one of the largest ethnic groups (more than 129 million). However, their genetic diversity had only been characterized with RFLP and biochemical markers, although there are already established mtDNA sequence databases for many ethnic groups of the world. We have obtained sequence data from 103 individuals living in three Russian regions: Kostroma, Kursk, and Rjazan. The sequenced fragment analyzed is 360 bp in length (positions from 16024 to 16383). Fifty nine nucleotide positions have been found polymorphic in Russians, among those were 57 transitions and two transversions. One individual is found having two insertions of two cytosines between positions 16184 and 16193. Among 64 different mitotypes identified in the study 52 were unique in these samples. The index of genetic diversity (Nei, 1987) for Russians is 0.96. This value is within the established range for European populations (0.93 to 0.98). Genetic distances calculated from our data show that Russians form a cluster with Germans, Bulgarians, Swedes, Estonians, and Volgo-Finns are more distant from Karelians and Finns, and much more differ from Turks and especially Mongolians.

Base Sequence↗

Analysis of sequence diversity in hypervariable regions of the external glycoprotein of human immunodeficiency virus type 1.

Nucleotide sequences in three hypervariable regions of the human immunodeficiency virus type 1 (HIV-1) env gene were obtained by sequencing provirus present in peripheral blood mononuclear cells of HIV-infected individuals. Single molecules of target sequences were isolated by limiting dilution and amplified in two stages by the polymerase chain reaction, using nested primers. The product was directly sequenced to avoid errors introduced by Taq polymerase during the amplification process. There was extensive variation between sequences from the same individual as well as between sequences from different individuals. Interpatient variability was markedly less in individuals infected from a common source. A high proportion of amino acid substitutions in the hypervariable regions altered the number and positions of potential N-linked glycosylation sites. Sequences in two hypervariable regions frequently contained short (3- to 15-bp) duplications or deletions, and by amplifying peripheral blood mononuclear cell DNA containing 10(2) or 10(3) proviral molecules and analyzing the product by high-resolution electrophoresis, the total number and abundance of distinct length variants within an individual could be estimated, providing a more comprehensive analysis of the variants present than would be obtained by sequencing alone. Sequences from many individuals showed frequent amino acid substitutions at certain key positions for neutralizing-antibody and cytotoxic T-cell recognition in the immunodominant loop. The rates of synonymous and nonsynonymous nucleotide substitution in the region of this and flanking regions indicate that strong positive selection for amino acid change is operating in the generation of antigenic diversity.

Amino Acid Sequence↗

Sequence diversity and virulence in Zea mays of Maize streak virus isolates.

Full genomic sequences were determined for 12 Maize streak virus (MSV) isolates obtained from Zea mays and wild grass species. These and 10 other publicly available full-length sequences were used to classify a total of 66 additional MSV isolates that had been characterized by PCR-restriction fragment length polymorphism and/or partial nucleotide sequence analysis. A description is given of the host and geographical distribution of the MSV strain and subtype groupings identified. The relationship between the genotypes of 21 fully sequenced virus isolates and their virulence in differentially MSV-resistant Z. mays genotypes was examined. Within the only MSV strain grouping that produced severe symptoms in maize, highly virulent and widely distributed genotypes were identified that are likely to pose the most serious threat to maize production in Africa. Evidence is presented that certain of the isolates investigated may be the products of either intra- or interspecific recombination.

Geminiviridae↗