Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequence diversity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Identifying functionally important mutations from phenotypically diverse sequence data.

Here we present a simple statistical method to determine the phenotypic contribution of a single mutation from libraries of mutants with diverse phenotypes in which each mutant contains a multitude of mutations. The central premise of this method is that, given M phenotypic classes, mutations that do not affect the phenotype should partition among the M classes according to a multinomial distribution. Deviations from this distribution are indicative of a link between specific mutations and phenotypes. We suggest that this method will aid the engineering of functional nucleic acids, proteins, and other biomolecules by uncovering target sites for rational mutagenesis. As a proof of the principle, we show how the method can be used to deduce the individual effects of mutations in a set of 69 P(L)-lambda promoter variants. Each of these promoters was generated by error-prone PCR and incorporated numerous mutations. The activity of the promoters was assayed using flow cytometry to measure the fluorescence of a green fluorescent protein reporter gene. Our analysis of the sequences of these mutants revealed seven positions having a statistically significant correlation with promoter activity. Using site-directed mutagenesis, we constructed point mutations for several sites, both statistically significant and insignificant, and combinations of these sites. Our results show that the statistical method correctly elucidated the phenotypic manifestations of these mutations. We suggest that this method may be useful for expediting directed evolution experiments by allowing both desired and undesired mutations to be identified and incorporated between rounds of mutagenesis.

Base Sequence↗

Sequence diversity of South Pacific isolates of Taro bacilliform virus and the development of a PCR-based diagnostic test.

We have analysed the sequence variability in the putative reverse transcriptase (RT)/ribonuclease H (RNaseH) and the C-terminal coat protein (CP)-coding regions from Taro bacilliform virus (TaBV) isolates collected throughout the Pacific Islands. When the RT/RNaseH-coding region of 22 TaBV isolates from Fiji, French Polynesia, New Caledonia, Papua New Guinea (PNG), Samoa, Solomon Islands and Vanuatu was examined, maximum variability at the nucleotide and amino acid level was 22.9% and 13.6%, respectively. Within the CP-coding region of 13 TaBV isolates from Fiji, New Caledonia, PNG, Samoa and the Solomon Islands, maximum variability at the nucleotide and amino acid level was 30.7% and 19.5%, respectively. Phylogenetic analysis showed that TaBV isolates from the Solomon Islands showed greatest variability while those from New Caledonia and PNG showed least variability. Based on the sequences of the TaBV RT/RNaseH-coding region, we have developed a PCR-based diagnostic test that specifically detects all known TaBV isolates. Preliminary indexing has revealed that TaBV is widespread throughout Pacific Island countries. A sequence showing approximately 50% nucleotide identity to TaBV in the RT/RNaseH-coding region was also detected in all taro samples tested. The possibility that this may represent either an integrated sequence or the genome of an additional badnavirus infecting taro is discussed.

Badnavirus↗

Amino acid sequence diversity within the family of antibodies bearing the major antiarsonate cross-reactive idiotype of the A strain mouse.

VH region amino acid sequences are described for five A/J anti-p-azophenylarsonate (anti-Ars) hybridoma antibodies for which the VL region sequences have previously been determined, thus completing the V domain sequences of these molecules. These antibodies all belong to the family designated Ars-A which bears the major anti-arsonate cross-reactive idiotype (CRI) of the A strain mouse. However, they differ in the degree to which they express the CRI in standard competition radioimmunoassays. Although the sequences are closely related, all are different from each other. Replacements are distributed throughout the VH region and occur in positions of the chain encoded by all three gene segments, VH, DH, and JH. It is likely that somatic diversification processes play a dominant role in producing the sequence variability in each of these segments. The number of differences from the sequence encoded by the germline is smallest for antibodies that express the CRI most strongly, suggesting that somatic diversification is responsible for loss of the CRI in members of the Ars-A antibody family. There is an unusual degree of clustering of differences in both CDR2 and CDR3 and many of the substitutions are located in "hot spots" of variation. The large number of differences between the chains prohibits the unambiguous identification of positions at which alterations play a major role in reducing the expression of the CRI. However, the data suggest that the loss of the CRI is associated with a definable repertoire of somatic changes at a restricted number of highly variable sites.

Amino Acid Sequence↗

Extensive 16S rRNA gene sequence diversity in Campylobacter hyointestinalis strains: taxonomic and applied implications.

Phylogenetic relationships of Campylobacter hyointestinalis subspecies were examined by means of 16S rRNA gene sequencing. Sequence similarities among C. hyointestinalis subsp. lawsonii strains exceeded 99.0%, but values among C. hyointestinalis subsp. hyointestinalis strains ranged from 96.4 to 100%. Sequence similarities between strains representing the two different subspecies ranged from 95.7 to 99.0%. An intervening sequence was identified in certain of the C. hyointestinalis subsp. lawsonii strains. C. hyointestinalis strains occupied two distinct branches in a phylogenetic analysis of the genus Campylobacter, emphasizing the need for multiple strain analysis when using 16S rRNA gene sequence comparisons for taxonomic investigations.

Animals↗

Sequence diversity of the control region of mitochondrial DNA in Tuscany and its implications for the peopling of Europe.

The control region of mitochondrial DNA has been widely studied in various human populations. This paper reports sequence data for hypervariable segments 1 and 2 of the control region from a population from southern Tuscany (Italy). The results confirm the high variability of the control region, with 43 different haplotypes in 49 individuals sampled. The comparison of this set of data with other European populations allows the reconstruction of the population history of Tuscany. Independent approaches, such as the estimation of haplotype diversity, mean pairwise differences, genetic distances and discriminant analysis, place the Tuscan sample in an intermediate position between sequences from culturally or geographically isolated regions of Europe (Sardinia, the Basque Country, Britain) and those from the Middle East. In spite of the remarkable genetic homogeneity in Europe, a degree of variability is shown by local European populations and homogeneity increases with the relative isolation of the population. The pattern of mitochondrial variation in Tuscany indicates the persistence of an ancient European component subsequently enriched by migrational waves, possibly from the Middle East.

Base Sequence↗

Sequence diversity and chromosomal distribution of "young" Alu repeats.

Members of the recently inserted human-specific (HS)/predicted variant (PV) subfamily of Alu elements were sequenced. A number of these Alu elements share greater than 98% sequence identity with the subfamily consensus sequence, and they are flanked by perfect 5' and 3' direct repeats ranging in size from 6 to 15 nucleotides (nt). Based on the low number of random mutations, the estimated average age of these elements was calculated to be 1.5 million years (Myr). All the young Alu subfamily members were restricted to the human genome, as judged by polymerase chain reaction (PCR) amplification of human and non-human primate DNA samples using the unique flanking sequences specific for each Alu element. The chromosomal locations of several Alu elements belonging to the young subfamilies, designated as HS/PV and Sb2, were determined by PCR amplification of DNA samples from human/rodent somatic cell hybrid panels. A statistical analysis of the chromosomal distribution pattern showed that the recently inserted Alu elements appear to integrate randomly in the human genome.

Animals↗

Inter- and intra-patient sequence diversity among parainfluenza virus-type 1 nucleoprotein genes.

Parainfluenza viruses (PIV) have been categorized into four discrete types (types 1-4), based on antigenic similarities. Here is described an evaluation of nucleoprotein (NP) sequence variability among nine patients infected with the type 1 virus. The examination of short segments of the NP sequence was sufficient to define significant variability both within and between patient samples. These data, in conjunction with previous studies of hemagglutinin-neuraminidase and fusion protein sequences from PIV-infected patient populations suggest a lack of absolute stability among isolates within each virus type. Potentially, antigenic variability exists to the extent that an immune response elicited toward one isolate may not be fully protective against another of the same type. Thus, sequence variability could contribute to natural re-infections with PIV, as well as to previous vaccine failures. Results highlight the importance of analyzing viruses that break through vaccine-induced immunity, in order to measure the influence of virus diversity on PIV vaccine outcome.

Amino Acid Sequence↗

Sequence diversity and large-scale typing of SNPs in the human apolipoprotein E gene.

A common strategy for genotyping large samples begins with the characterization of human single nucleotide polymorphisms (SNPs) by sequencing candidate regions in a small sample for SNP discovery. This is usually followed by typing in a large sample those sites observed to vary in a smaller sample. We present results from a systematic investigation of variation at the human apolipoprotein E locus (APOE), as well as the evaluation of the two-tiered sampling strategy based on these data. We sequenced 5.5 kb spanning the entire APOE genomic region in a core sample of 72 individuals, including 24 each of African-Americans from Jackson, Mississippi; European-Americans from Rochester, Minnesota; and Europeans from North Karelia, Finland. This sequence survey detected 21 SNPs and 1 multiallelic indel, 14 of which had not been previously reported. Alleles varied in relative frequency among the populations, and 10 sites were polymorphic in only a single population sample. Oligonucleotide ligation assays (OLA) were developed for 20 of these sites (omitting the indel and a closely-linked SNP). These were then scored in 2179 individuals sampled from the same three populations (n = 843, 884, and 452, respectively). Relative allele frequencies were generally consistent with estimates from the core sample, although variation was found in some populations in the larger sample at SNPs that were monomorphic in the corresponding smaller core sample. Site variation in the larger samples showed no systematic deviation from Hardy-Weinberg expectation. The large OLA sample clearly showed that variation in many, but not all, of OLA-typed SNPs is significantly correlated with the classical protein-coding variants, implying that there may be important substructure within the classical epsilon 2, epsilon 3, and epsilon 4 alleles. Comparison of the levels and patterns of polymorphism in the core samples with those estimated for the OLA-typed samples shows how nucleotide diversity is underestimated when only a subset of sites are typed and underscores the importance of adequate population sampling at the polymorphism discovery stage. [The sequence data described in this paper have been submitted to the GenBank data library under accession no. AF261279.]

Alleles↗

The human beta-myosin heavy chain gene: sequence diversity and functional characteristics of the protein.

The beta-myosin heavy chain gene (MYH7) encodes the motor protein that drives myocardial contraction. It has been proven to be a disease gene for hypertrophic cardiomyopathy (HCM). We analyzed the DNA sequence variation of MYH7 (about 16 kb) of eight individuals: six patients with HCM and two healthy controls. The overall DNA sequence identity was up to 97.2% compared to Jaenicke and coworkers (Jaenicke et al. [1990] Genomics 8:194-206), while the corresponding amino acid sequences revealed 100% identity. In HCM patients, eleven nucleotide substitutions were identified but no causative disease mutation was found: six were detected in coding, four in intronic, and one in 5' regulatory regions. The average nucleotide diversity across this locus was 0.015% with an average of 0.02% in the coding and 0.012% in the noncoding sequence. Analysis of the kinetic behaviour of beta-MHC in the intact contractile structure of normal individuals and HCM patients revealed apparent rate constants of tension development ranging between 1.58 s(-1) and 1.48 s(-1).

Base Sequence↗

Cloning, sequencing, expression and allelic sequence diversity of ERG3 (C-5 sterol desaturase gene) in Candida albicans.

The C-5 sterol desaturase gene (ERG3), essential for yeast ergosterol biosynthesis, was cloned and sequenced from Candida albicans by homology with the Saccharomyces cerevisiae ERG3. The ERG3 ORF contained 1158bp and encoded 386 deduced amino acids. The clone was used to transform a gal1 mutant derived from the Darlington strain of C. albicans, using galactose selection. The Darlington strain is known to lack Delta(5,6) sterols, i.e. to have an erg3 phenotype (Howell, S.A., et al., 1990. J. Appl. Bacteriol. 69, 692-696). The transformant (CDTR1) contained six tandem integrated ERG3GAL1 repeats, had double the abundance of ERG3 transcript found in the host strain, and synthesized ergosterol, a Delta(5,6) sterol.The Darlington strain was noted to have an abundance of ERG3 transcript. Both ERG3 alleles in Darlington were cloned and sequenced in order to look for changes that might explain the erg3 phenotype. One allele, called Dar-2, contained a stop codon in place of tryptophan-292. The other ERG3 allele, called Dar-1, had changes in three amino acids, two of which were conserved in three fungal and one plant species. EcoRI genomic fragments containing ERG3 from the Dar-1 allele and from B311, the wild-type strain, were inserted into the plasmid pRS316 and used to transform a Saccharomyces cerevisiae erg3,ura3 mutant using uracil selection. The 4.1kb ERG3 fragments from the B311 and Dar-1 both contained 1. 4kb 5' and 1.5kb 3' flanking sequences around the coding region. Transformants with ERG3 from B311 but not from Dar-1 showed restored ergosterol synthesis. One or more of these three deduced amino acids in the Dar-1 allele of ERG3 appeared critical for function.

Amino Acid Sequence↗

Sequence diversity of V1 and V2 domains of gp120 from human immunodeficiency virus type 1: lack of correlation with viral phenotype.

We analyzed by PCR and direct sequencing 57 viral sequences from 47 individuals infected with human immunodeficiency virus type 1, focussing on the V1 and V2 regions of gp120. There was extensive length polymorphism in the V1 region, which rendered sequence alignment difficult. The V2 hypervariable locus also displayed considerable length variations, whereas flanking regions were relatively conserved. Two-thirds of the amino acid residues in these flanking regions were highly conserved (> 80%), presumably reflecting their critical contribution to V2 structure or function. We also characterized the syncytium-inducing properties of the isolates from which we derived sequence information. There was no correlation between V1 or V2 sequences and the viral phenotype, contrary to a previous report (M. Groenink, R. A. M. Fouchier, S. Broersen, C. H. Baker, M. Koot, A. B. van't Wout, H. G. Huisman, F. Miedema, M. Tersmette, and H. Schuitemaker, Science 260:1513-1516, 1993). The sequence heterogeneity described in this study provides information to suggest that it would be most difficult to exploit the V1 and V2 domains for vaccine development.

Amino Acid Sequence↗

Detection and sequence diversity of Kaposi's sarcoma-associated herpesvirus (KSHV)/human herpesvirus 8 (HHV-8) DNA.

Detection of Kaposi's sarcoma (KS)-associated herpesvirus (KSHV)/human herpesvirus (HHV-8) has been reported frequently in patients with KS associated with acquired immunodeficiency syndrome (AIDS). We examined the presence of the KSHV sequence in 8 HIV-positive patients comprising 5 with KS, 2 with syphilis, 1 with prurigo, and 2 HIV-negative patients with angiosarcoma. Using the polymerase chain reaction, we observed amplification of a DNA fragment of the expected size in 4 patients with KS. Sequencing analysis of the amplified fragments revealed several base substitutions upon comparison with the originally reported sequence. Our results support the hypothesis of a pathogenic role of KSHV in the development of skin lesions in HIV-positive patients with KS, and the sequences of KSHV DNA fragments isolated in this study also demonstrated strain diversity similar to that reported previously.

AIDS-Related Opportunistic Infections↗

Genetic recombination between two genotypes of genogroup III bovine noroviruses (BoNVs) and capsid sequence diversity among BoNVs and Nebraska-like bovine enteric caliciviruses.

To determine the genogroups and genotypes of bovine enteric caliciviruses (BECVs) circulating in calves, we determined the complete capsid gene sequences of 21 BECVs. The nucleotide and predicted amino acid sequences were compared phylogenetically with those of known human and animal enteric caliciviruses. Based on these analyses, 15 BECVs belonged to Norovirus genogroup III and genotype 2 (GIII/2) and were genetically distinct from human Norovirus GI and GII. Six BECVs had capsid gene sequences similar to that of the unclassified Nebraska (NB)-like BECV. The 15 bovine noroviruses (BoNVs) were more closely related to Bo/NLV/Newbury-2/76/UK (GIII/2) and other known genotype 2 BoNVs than to genotype 1 Bo/NLV/Jena/80/DE. The BoNV Bo/CV521-OH/02/US showed high nucleotide and amino acid identities (84 and 94%, respectively) with the capsid gene of Bo/NLV/Newbury-2/76/UK, whereas the nucleotide and amino acid sequences of the RNA polymerase gene were more closely related to those of Bo/NLV/Jena/80/DE (77 and 87% identities, respectively) than to those of Bo/NLV/Newbury-2/76/UK (69 and 69% identities, respectively), suggesting that Bo/CV521-OH/02/US is a genotype 1-2 recombinant. Gene conversion analysis by the recombinant identification program and SimPlot also predicted that Bo/CV521-OH/02/US was a recombinant. Six NB-like BECVs shared 88 to 92% nucleotide and 94 to 99.5% amino acid identities with the NB BECV in the capsid gene. The results of this study demonstrate genetic diversity in the capsid genes of BECVs circulating in Ohio veal calves, provide new data for coinfections with distinct BECV genotypes or genogroups, and describe the first natural BoNV genotype 1-2 recombinant, analogous to the previously reported human norovirus recombinants.

Animals↗

Cloning and characterization of cDNAs encoding S-RNases from almond (Prunus dulcis): primary structural features and sequence diversity of the S-RNases in Rosaceae.

cDNAs encoding three S-RNases of almond (Prunus dulcis), which belongs to the family Rosaceae, were cloned and sequenced. The comparison of amino acid sequences between the S-RNases of almond and those of other rosaceous species showed that the amino acid sequences of the rosaceous S-RNases are highly divergent, and intra-subfamilial similarities are higher than inter-subfamilial similarities. Twelve amino acid sequences of the rosaceous S-RNases were aligned to characterize their primary structural features. In spite of their high level of diversification, the rosaceous S-RNases were found to have five conserved regions, C1, C2, C3, C5, and RC4 which is Rosaceae-specific conserved region. Many variable sites fall into one region, named RHV. RHV is located at a similar position to that of the hypervariable region a (HVa) of the solanaceous S-RNases, and is assumed to be involved in recognizing S-specificity of pollen. On the other hand, the region corresponding to another solanaceous hypervariable region (HVb) was not variable in the rosaceous S-RNases. In the phylogenetic tree of the T2/S type RNase, the rosaceous S-RNase fall into two subfamily-specific groups (Amygdaloideae and Maloideae). The results of sequence comparisons and phylogenetic analysis imply that the present S-RNases of Rosaceae have diverged again relatively recently, after the divergence of subfamilies.

Amino Acid Sequence↗

Fluconazole-resistant pathogens Candida inconspicua and C. norvegensis: DNA sequence diversity of the rRNA intergenic spacer region, antifungal drug susceptibility, and extracellular enzyme production.

The opportunistic fungal pathogens Candida inconspicua and C. norvegensis are very rarely isolated from patients and are resistant to fluconazole. We collected 38 strains of the two microorganisms isolated from Europe and Japan, and compared the polymorphism of the rRNA intergenic spacer (IGS) and internal transcribed spacer (ITS) regions, antifungal drug susceptibility, and extracellular enzyme production as a potential virulence factor. While the IGS sequences of C. norvegensis were not very divergent (more than 96.7% sequence similarity among the strains), those of C. inconspicua showed remarkable diversity, and were divided into four genotypes with three subtypes. In the ITS region, no variation was found in either species. Since the sequence similarity of the two species is approximately 70% at the ITS region, they are closely related phylogenetically. Fluconazole resistance was reconfirmed for the two microorganisms but they were susceptible to micafungin and amphotericin B. No strain of either species secreted aspartyl proteinase or phospholipase B. These results provide basal information for accurate identification, which is of benefit to global molecular epidemiological studies and facilitates our understanding of the medical mycological characteristics of C. inconspicua and C. norvegensis.

Amphotericin B↗

Global sequence diversity of BRCA2: analysis of 71 breast cancer families and 95 control individuals of worldwide populations.

The aim of this study was to evaluate the prevalence of simple sequence variation in the BRCA2 gene. To this end, 71 breast and breast-ovarian cancer (HBC/HBOC) families along with 95 control individuals from a wide range of ethnicities were analyzed by means of denaturing high-performance liquid chromatography (DHPLC) and direct sequence analysis. In the coding (10 257 bp) and non-coding (2799 bp) sequences of BRCA2, 82 sequence variants were identified. Three different, apparently disease-associated BRCA2 mutations were found in six HBC/HBOC families (8%): two splice site mutations in introns 5 and 21, and one frameshift mutation in exon 11. In the coding region, 53 simple sequence variants were found: 35 missense mutations, one 2 bp deletion (CT) resulting in a stop at codon 3364, one nonsense mutation with a stop at codon 3326, one deletion of a complete codon (AAA) resulting in the loss of leucine, and 15 silent mutations. In the non-coding region, 26 polymorphisms were detected. Of the 79 sequence variants that were not obviously disease-associated, eight were detected only in HBC/HBOC families. The remaining 71 variants were identified in both HBC/HBOC families and control individuals. Sixty three sequence variants (80%) were specific for a continent. Forty two percent (33 out of 79) of the sequence variants were detected exclusively in Africa, though only 13% of the 332 chromosomes screened were of African origin. Our data indicate that, in BRCA2, simple sequence variation is frequent [in the coding region 1 in 194 bp (straight theta = 2.2 x 10(-4)), and in the non-coding region 1 in 108 bp (straight theta = 4.4 x 10(-4)), respectively].

Africa↗