Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequence diversity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Sequence diversity within a subgroup of mouse immunoglobulin kappa chains controlled by the IgK-Ef2 locus.

We previously showed that a chromosome 6 locus, IgK-Ef2, controls a pair of prominent bands in normal mouse light-chain isoelectric focusing profiles. Screening of myeloma light chains derived from BALB/c mice (an IgK-EF2 alpha strain) led to the identification of seven light chains cofocusing with the polymorphic bands controlled by IgK-Ef2. Complete sequencing of the variable (V) regions of four of the light chains indicates that they are all members of the same subgroup (Vk-1A) and they differ from one another by 1--3 substitutions. One of the protein differs from the prototype V-region sequence only in the deletion of a single residue at position 95 immediately preceding of J region. The other two differ from the protype V region by 3 (two framework [fr], one complementarity-determined [cdr]) and one (fr) residues, respectively. Complete V-region sequences of two closely related light chains derived from NZB mice (an IgK-Ef2b strain) indicate the NZB proteins are derived from a distinct Vk gene (Vk-1B), differing by four substitutions from the Vk-1A sequence. The results suggest that the IgK-Ef2 polymorphism may be a result of, at least in part, the loss of the gene(s) coding for the Vk-1A subgroups in IgK-Ef2b strains of mice. The nature of the sequence diversity found in the Vk-1A subgroup indicates that either it is coded by a repeated series of virtually identical genes or that somatic mutation of a single Vk-1A gene may give rise to substitutions in framework as well as cdr regions.

Amino Acid Sequence↗

Extensive nuclear DNA sequence diversity among chimpanzees.

Although data on nucleotide sequence variation in the human nuclear genome have begun to accumulate, little is known about genomic diversity in chimpanzees (Pan troglodytes) and bonobos (Pan paniscus). A 10,154-base pair sequence on the chimpanzee X chromosome is reported, representing all major subspecies and bonobos. Comparison to humans shows the diversity of the chimpanzee sequences to be almost four times as high and the age of the most recent common ancestor three times as great as the corresponding values of humans. Phylogenetic analyses show the sequences from the different chimpanzee subspecies to be intermixed and the distance between some chimpanzee sequences to be greater than the distance between them and the bonobo sequences.

Animals↗

Alternative splicing and sequence diversity of transcripts from the oncosphere stage of Taenia solium with homology to the 45W antigen of Taenia ovis.

Genes and transcripts which show homology to the host-protective 45W antigen of Taenia ovis have been cloned from the human parasite Taenia solium. The T. solium genes cloned in this study (TSO45) show conserved genomic structural features which are also features of the T. ovis 45W gene family. The TSO45 genes consist of a four exon and three intron structure. Eight TSO45 transcripts, encoded by at least five genes, were cloned from T. solium oncospheres and comparison of their DNA sequence indicates that some transcripts have arisen by alternative splicing, the first demonstration of exon inclusion/exon skipping in cestodes. Alternative splicing occurred with respect to both exons II and III with three splice variants identified from the TSO45-1 gene and two splice variants from TSO45-5. The proteins encoded by this family of genes contain putative N-linked glycosylation sites, an amino terminal secretory signal, a hydrophobic carboxy terminal sequence characteristic of GPI-anchored proteins and fibronectin type III motifs. These features are common to their T. ovis and Taenia saginata homologues. The similarities of the TSO45 genes cloned in this study with genes encoding host-protective antigens of T. ovis and T. saginata indicates that the encoded T. solium proteins are quite possibly antigenic and have potential use as a vaccine to prevent T. solium infection in the parasite's intermediate host. In this respect, the generation of sequence diversity and hence potential antigenic diversity through alternative splicing of TSO45 genes may have implications for the use of these proteins in vaccines against T. solium cysticercosis.

Alternative Splicing↗

Sequence diversity of TT virus in geographically dispersed human populations.

TT virus (TTV) is a newly discovered DNA virus originally classified as a member of the Parvoviridae. TTV is transmitted by blood transfusion where it has been reported to be associated with mild post-transfusion hepatitis. TTV can cause persistent infection, and is widely distributed geographically; we recently reported extremely high prevalences of viraemia in individuals living in tropical countries (e.g. 74% in Papua New Guinea, 83% in Gambia; Prescott & Simmonds, New England Journal of Medicine 339, 776, 1998). In the current study we have compared nucleotide sequences from the N22 region of TTV (222 bases) detected in eight widely dispersed human populations. Some variants of TTV, previously classified as genotypes 1a, 1b and 2, were widely distributed throughout the world, while others, such as a novel subtype of type 1 in Papua New Guinea, were confined to a single geographical area. Five of the 122 sequences obtained in this study (from Gambia, Nigeria, Papua New Guinea, Brazil and Ecuador) could not be classified as types 1, 2 or 3, with the variant from Brazil displaying only 46-50% nucleotide (32-35% amino acid) sequence similarity to other variants. This study provides an indication of the extreme sequence diversity of TTV, a characteristic which is untypical of parvoviruses.

Africa, Western↗

Unisexuality and molecular drive: bag320 sequence diversity in bacillus taxa (insecta phasmatodea).

Satellite DNA variability follows a pattern of concerted evolution through homogenization of new variants by genomic turnover mechanisms and variant fixation by chromosome redistribution into new combinations with the sexual process. Bacillus taxa share the same Bag320 satellite family and their reproduction ranges from strict bisexuality (B. grandii) to automictic (B. atticus) and apomictic (B. whitei = rossius/ grandii; B. lynceorum = rossius/grandii/atticus) unisexuality. Thelytokous reproduction clearly allows uncoupling of homogenization from fixation. Both trends and absolute values of satellite variability were analyzed in all Bacillus taxa but B. rossius, on 906 sequenced monomers at all level of comparisons: intraspecimen, intrapopulation, interpopulation, intersubspecies, and interspecies. For unisexuals, allozymic and mitochondrial clones were also taken into account. Different reproductive modes (sexual/parthenogenetic) appear to explain observed variability trends, supporting Dover's hypothesis of sexuality acting as a driving force in the fixation of sequence variants, but the present analyses also highlight current spreading of new variants in B. grandii maretimi specimens and point to a biased sequence inheritance at the time of hybrid onset in the apomictic hybrids B. whitei and B. lynceorum. Evidence of biased gene conversion events suggests that, given enough time, sequence homogenization can take place in a unisexual such as B. lynceorum. On the contrary, the absolute values of sequence diversity in each taxon are linked to the species' range, time of divergence, and repeat copy number and, possibly, to transposon features. Satellite dynamics appears therefore to be the outcome of both general molecular processes and specific organismal traits.

Animals↗

Sequence diversity of nuclear and polysomal polyadenylated and non-polyadenylated RNA in normal and regenerating rat liver.

A DNA probe purified from RNA . DNA hybrids of total sham-operated liver RNA and non-repetitive DNA was used to show that nuclear poly(A)-rich RNA from sham-operated liver and from 2.5-h and 48-h regenerating liver contains about 50% of the complexity of total liver RNA. It was further shown that the differences between normal and regenerating liver, reported earlier from this laboratory, occur in the poly(A)-free fraction of nuclear RNA. At the polysome level it was found that polysomal RNA has one-third of the sequence diversity of total RNA and of this, approximately 65% can be accounted for by poly(A)-rich and 55% by poly(A)-free molecules. When DNA probes were prepared from hybrids formed using polysomal RNA from sham-operated liver and regenerating liver at 2.5 h and 48 h post-hepatectomy and then employed in reactions with homologous and heterologous RNAs, no differences were detectable between either normal and regenerating liver or regenerating liver at times during hypertrophy and hyperplasia.

Animals↗

Sequence diversity in the 5'-UTR region of GB virus C/hepatitis G virus assessed using sequencing, heteroduplex mobility analysis and single-strand conformation polymorphism.

GB virus C/hepatitis G virus (GBV-C/HGV) is a positive-sense RNA virus belonging to the Flaviviridae family identified recently. Reverse transcription polymerase chain reaction (RT-PCR) was used to detect GBV-C/HGV RNA using nested primers designed to amplify 245 bp of the 5'-untranslated region (UTR). GBV-C/HGV RNA was detected in 20.7% of 101 HCV-RNA positive and 6.8% of 44 HCV-RNA negative specimens. Sequencing of the PCR products demonstrated they had between 84.3 and 100% nucleotide identity. Most of the diversity corresponded to two variable regions identified within the 5'-UTR. Phylogenetic analysis indicated that GBV-C/HGV subtypes present in Australia belonged to group 2 and were closest in evolutionary terms to isolates from the USA and Europe. All isolates were analysed using single-strand conformation polymorphism (SSCP) and heteroduplex mobility analysis (HMA) on 8% non-denaturing polyacrylamide gels. SSCP of the isolates identified a number of distinct conformation polymorphisms that corresponded with sequence-determined genetic diversity. HMA was developed to assess the amount of genetic diversity between isolates without the need for sequencing. The average difference between the predicted divergence of two isolates calculated from the mobility of the heteroduplex and the actual value (based on nucleotide sequence) was 2.3% in this sample of isolates, where the mean sequence divergence was 8.52%.

5' Untranslated Regions↗

Estimation of DNA sequence diversity in bovine cytokine genes.

DNA sequence variation provides the fundamental material for improving livestock through selection. In cattle, single nucleotide polymorphisms and small insertions/deletions (collectively referred to here as SNPs) have been identified in cytokine genes and scored in a reference population to determine linkage map positions. The aim of the present study was twofold: first, to estimate the SNP frequency in a reference population of beef cattle, and second, to determine cytokine haplotypes in a group of sires from commercial populations. Forty-five SNP markers in DNA segments from nine cytokine gene loci were analyzed in 26 reference parents. Comparison of all 52 haploid genomes at each PCR amplicon locus revealed an average of one SNP per 143 bp of sequence, whereas comparison of any two chromosomes identified heterozygous sites, on average, every 443 bp. The combination of these 45 SNP genotypes was sufficient to uniquely identify each of the 26 animals. The average number of haplotype alleles (4.4) per PCR amplicon (688 bp) and the percentage heterozygosity among founding parents (50%) were similar to those for microsatellite markers in the same population. For 49 sires from seven common breeds of beef cattle, SNP genotypes (1,225 total) were obtained by matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS) at three amplicon loci. All three of the amplicon haplotypes were correctly deduced for each sire without the use of parent or progeny genotypes. The latter allows a wide range of genetic studies in commercial populations of cattle where genotypic information from relatives may not be available.

Alleles↗

The carboxylesterase family exhibits C-terminal sequence diversity reflecting the presence or absence of endoplasmic-reticulum-retention sequences.

Resident proteins of the endoplasmic reticulum lumen are continuously retrieved from an early Golgi compartment by a receptor-mediated mechanism. The sorting or retention sequence on the endoplasmic reticulum proteins is located at the C-terminus and was initially shown to be the tetrapeptide KDEL in mammalian cells and HDEL in Saccharomyces cerevisiae. The carboxylesterases are a large family of enzymes primarily localized to the lumen of the endoplasmic reticulum. Retention sequences in these proteins have been difficult to identify due to atypical and heterogeneous C-terminal sequences. Utilizing the polymerase chain reaction with degenerate primers, we have identified and characterized the C-termini of four members of the carboxylesterase family from rat liver. Three of the carboxylesterases sequences contained C-terminal sequences (HVEL, HNEL or HTEL) resembling the yeast sorting signal which were reported to be non-functional in mammalian cells. A fourth carboxylesterase contained a distinct C-terminal sequence, TEHT. A full-length esterase cDNA clone, terminating in the sequence HVEL, was isolated and was used to assess the retention capabilities of the various esterase C-terminal sequences. This esterase was retained in COS-1 cells, but was secreted when its C-terminal tetrapeptide, HVEL, was deleted. Addition of C-terminal sequences containing HNEL and HTEL resulted in efficient retention. However, the C-terminal sequence containing TEHT was not a functional retention signal. Both HDEL, the authentic yeast retention signal, and KDEL were efficient retention sequences for the esterase. These studies show that some members of the rat liver carboxylesterase family contain novel C-terminal retention sequences that resemble the yeast signal. At least one member of the family does not contain a C-terminal retention signal and probably represents a secretory form.

Amino Acid Sequence↗

Sequence diversity within the reovirus S2 gene: reovirus genes reassort in nature, and their termini are predicted to form a panhandle motif.

To better understand genetic diversity within mammalian reoviruses, we determined S2 nucleotide and deduced sigma 2 amino acid sequences of nine reovirus strains and compared these sequences with those of prototype strains of the three reovirus serotypes. The S2 gene and sigma 2 protein are highly conserved among the four type 1, one type 2, and seven type 3 strains studied. Phylogenetic analyses based on S2 nucleotide sequences of the 12 reovirus strains indicate that diversity within the S2 gene is independent of viral serotype. Additionally, we found marked topological differences between phylogenetic trees generated from S1 and S2 gene nucleotide sequences of the seven type 3 strains. These results demonstrate that reovirus S1 and S2 genes have distinct evolutionary histories, thus providing phylogenetic evidence for lateral transfer of reovirus genes in nature. When variability among the 12 sigma 2-encoding S2 nucleotide sequences was analyzed at synonymous positions, we found that approximately 60 nucleotides at the 5' terminus and 30 nucleotides at the 3' terminus were markedly conserved in comparison with other sigma 2-encoding regions of S2. Predictions of RNA secondary structures indicate that the more conserved S2 sequences participate in the formation of an extended region of duplex RNA interrupted by a pair of stem-loops. Among the 12 deduced sigma 2 amino acid sequences examined, substitutions were observed at only 11% of amino acid positions. This finding suggests that constraints on the structure or function of sigma 2, perhaps in part because of its location in the virion core, have limited sequence diversity within this protein.

Amino Acid Sequence↗

Amino acid sequence diversity between bovine epidermal cytokeratin polypeptides of the basic (type II) subfamily as determined from cDNA clones.

The nucleotide sequences of four cDNA clones, each representing the carboxyterminal portion of a bovine epidermal cytokeratin of the "basic" (type II) subfamily, were determined, i.e., components Ia (Mr 68,000), Ib (Mr 68,000), III (Mr 60,000), and IV (Mr 59,000). The comparison of the sequences with each other and with the human type-II cytokeratin of Mr 56,000 reported by Hanukoglu and Fuchs [24] allows the following conclusions: The four major epidermal keratins of the basic (type II) subfamily, which are co-expressed in keratinocytes of the bovine muzzle, exhibit a high homology (greater than 90%) in the alpha-helical portion, but differ considerably in their nonhelical carboxy-terminal regions. The nonhelical carboxyterminal regions of all four cytokeratins are exceptionally rich in glycine and serine. Within the extrahelical tail, three different domains can be distinguished. The consensus sequence TYR(X)LLEGE which demarcates the end of the alpha-helical rod in all intermediate filaments is followed by a relatively short (22-27 amino acids) intercept rich in hydroxy amino acids and valine (carboxyterminal tail domain C1). This is followed by a long region that is variable in size and sequence, rich in glycine di-, tri-, and tetrapeptides, and contains diverse repeated sequences (domain C2). This is followed by another short (20 residues) hydroxy-amino-acid-rich intercept (domain C3) that ends with a conspicuously basic sequence of approximately four to six carboxyterminal amino acids. The first half of domain C1 is also homologous in all four keratins, suggesting that this region also assumes a common conformation and/or serves a special common function.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Diverse sequences within Tlr elements target programmed DNA elimination in Tetrahymena thermophila.

Tlr elements are a novel family of approximately 30 putative mobile genetic elements that are confined to the germ line micronuclear genome in Tetrahymena thermophila. Thousands of diverse germ line-limited sequences, including the Tlr elements, are specifically eliminated from the differentiating somatic macronucleus. Macronucleus-retained sequences flanking deleted regions are known to contain cis-acting signals that delineate elimination boundaries. It is unclear whether sequences within deleted DNA also play a regulatory role in the elimination process. In the current study, an in vivo DNA rearrangement assay was used to identify internal sequences required in cis for the elimination of Tlr elements. Multiple, nonoverlapping regions from the approximately 23-kb Tlr elements were independently sufficient to stimulate developmentally regulated DNA elimination when placed within the context of flanking sequences from the most thoroughly characterized family member, Tlr1. Replacement of element DNA with macronuclear or foreign DNA abolished elimination activity. Thus, diverse sequences dispersed throughout Tlr DNA contain cis-acting signals that target these elements for programmed elimination. Surprisingly, Tlr DNA was also efficiently deleted when Tlr1 flanking sequences were replaced with DNA from a region of the genome that is not normally associated with rearrangement, suggesting that specific flanking sequences are not required for the elimination of Tlr element DNA.

Animals↗

DNA recombination and natural selection pressure sustain genetic sequence diversity of the feline MHC class I genes.

Sequence comparisons of seven distinct MHC class I cDNA clones revealed that feline class I molecules have a remarkable similarity to human HLA genes in their organization of functional domains as well as in the nonrandom partitioning of genetic variability according to the functional constraints ascribed to different regions of the MHC molecule. The distribution of the pattern of sequence polymorphism in the cat as compared with genetic diversity of human and mouse class I genes provides evidence for four coordinate factors that contribute to the origin and sustenance of abundant allele diversity that characterizes the MHC in the species. These include: (a) a gradual accumulation of spontaneous mutational substitution over evolutionary time; (b) selection against mutational divergence in regions of the class I molecule involved in T cell receptor interaction and also in certain regions that interact with common features of antigens; (c) positive selection pressure in favor of persistence of polymorphism and heterozygosity at 57 nucleotide residues that comprise the antigen recognition site; and (d) periodic intragenic (interallelic) and intergenic recombination within the class I genes. We describe a highly conserved 23-bp nucleotide sequence within the coding region of the first alpha-helix that separates two relatively polymorphic segments located in the alpha 1 domain that may act as a template or "hot spot" for homologous recombination between class I alleles.

Amino Acid Sequence↗

Molecular and population genetic analysis of allelic sequence diversity at the human beta-globin locus.

Allelic sequence polymorphism at the beta-globin locus was investigated in a group of 36 Melanesians. A 3-kilobase fragment containing the gene and its flanking regions was sequenced in 60 normal (beta A) and 12 thalassemic (intron 1, position 5, G-->C) chromosomes. Haplotype relationships between linked polymorphisms were derived by allele-specific PCR amplification and sequencing. Seventeen nucleotide polymorphisms and 2 length variants were identified, and these sites segregated as 17 sequence haplotypes in the normal chromosomes. This haplotype diversity is higher than that expected on the basis of the nucleotide polymorphism observed and is probably due to recombination and gene conversion. Nucleotide diversity at synonymous sites in the sample is 0.14%, suggesting an average age of sequence divergence of approximately 450,000 years, consistent with that expected for a neutrally evolving human nuclear locus.

Alleles↗

Sequence diversity in the kinetoplast DNA minicircles of Trypanosoma cruzi.

Minicircles are the most abundant component of the mitochondrially located kinetoplast DNA in the members of the order Kinetoplastida. Minicircle sequences differ among most trypanosomatid species. To learn about the molecular mechanisms that give rise to this diversity, we sequenced a complete minicircle (pTckAWP-2) and two homologous but polymorphic minicircle fragments isolated from different Trypanosoma cruzi clones. Comparison of these sequences revealed 23 point mutations, 19 of which were transitions. A single base pair insertion was also detected in one of the two minicircle fragments sequenced. Analysis of pTckAWP-2 sequence showed the following features: the presence of four internal 118 base pairs conserved regions with 80% or higher homology; the fact that these four conserved regions also differed mainly by point mutations, although in this case a bias in favor of transversions was observed; the existence in each of these four regions of the highly conserved 13 bp sequence 5'GGGGTTGGTGTAA3', detected in all trypanosomatid minicircles, which is thought to be the origin of replication; and the presence of several direct and inverted repeat sequences of 8 base pairs or longer, scattered throughout the minicircle molecule. Comparison of the T. cruzi conserved minicircle region with that of other trypanosomatids showed a higher homology of T. cruzi with T. lewisi, another stercorarian trypanosome, than with African trypanosomes or Leishmania.

Animals↗

Sequence diversity in the 5' untranslated region of rabbit muscle phosphofructokinase mRNA.

DNA sequences of two full-length rabbit muscle phosphofructokinase (RMPFK) cDNAs (A and B) show identical coding sequence but heterogeneous 5' untranslated regions. cDNA-A is formed by removal of a 1.7 kb upstream intron while cDNA-B retains the 3' region of this intron. A 2.8 kb upstream sequence of RMPFK gene contains several features characteristic of housekeeping genes: high GC content (67%) at its 5' end, with 50 CpG sites; five Sp1 sites; and no functional TATA box. Comparison of the 5' sequences of the two RMPFK cDNAs and three human muscle PFK cDNAs (Nakajima, H., et al., (1990) Biochem. Biophys. Res. Com. 166, 637) suggests that a single splicing event occurs involving different splicing donor sites but the same splicing acceptor site, resulting in diversity in the upstream sequence. These observations suggest that transcription of muscle PFK gene may start at multiple sites, another feature of housekeeping genes.

Animals↗

Sequence diversity of yeast 2 microns RAF gene and its co-evolution with STB and REP1.

Despite the extensive study of yeast 2 microns plasmid, the exact function of plasmid-encoded RAF gene is not clear. Variants of 2 microns plasmids from industrial Saccharomyces cerevisiae yeasts were isolated and characterized. Sequencing of RAF alleles revealed about 8% nucleotide and 10% amino acid diversities between 2 microns variants of closely related strains, RAF sequence variations were correlated with STB-REP1 sequence diversity. We also used restriction fragment length polymorphism linkage to screen a large number of yeast strains from different fermentation industries. The results clearly show a tight linkage of STB-REP1-RAF variations. Thus, our observations suggest that plasmid-borne cis- and trans-acting elements co-evolved to form an optimal molecular parasite and that RAF may play a role in active plasmid partitioning.

Alleles↗

Sequence diversity within the reovirus S3 gene: reoviruses evolve independently of host species, geographic locale, and date of isolation.

To better understand genetic diversity of mammalian reoviruses, we studied sequence variability in the S3 gene segment of 17 field-isolate reovirus strains and prototype strains of the three reovirus serotypes. Strains studied were isolated over a 37-year period from different mammalian hosts and geographic locations. A high degree of variability was observed in the nucleotide sequences of the S3 gene, whereas the deduced amino acid sequences of the S3 gene product, sigma NS, were highly conserved. When variability among the S3 nucleotide sequences was analyzed using pairwise comparisons, we found that 5' and 3' noncoding regions were significantly more conserved than the remainder of the gene. This high degree of sequence conservation was also observed within the first 15 nucleotides of the 5' coding region. Phylogenetic analyses showed that multiple alleles of the S3 gene cocirculate and that genetic diversity in the S3 gene does not correlate with host species, geographic locale, or date of isolation. Phylogenetic trees constructed from variation in the S3 sequences are distinct from those previously generated from sequences that encode attachment protein sigma 1, core protein sigma 2, and outer capsid protein sigma 3, which supports the hypothesis that reovirus gene segments reassort in nature. These findings suggest that reovirus gene segments are well-adapted to mammalian hosts and that reovirus evolution has reached an equilibrium.

Animals↗